Want the long version? Five guides go deeper than these definitions:
the AI-Use Score ↗,
your Toolbox ↗,
capabilities ↗,
data sources ↗, and
data access & privacy ↗.
Topics
Every conversation gets exactly one topic, decided by reading what it was actually about rather than by who filed it or what they called it — so nobody can change their numbers by renaming their work. Personal conversations are set aside from every summary by default; the note at the top of the page says how many there were.
Sub-topics
A finer split within each topic. Unlike topics, these are not a fixed list — they're drawn from your own work, then near-duplicates are merged so the same thing under three names doesn't look like three things. That means this list is yours and will look nothing like another company's. Click any sub-topic to read its conversations.
Output types
What the conversation actually produced — the thing you'd point at afterwards. Each conversation gets one, chosen from the fixed list below, so "we made 40 decks this month" is a question with an answer.
Complexity
How much the AI actually did — how many tools it reached for, whether it stopped to think things through, and whether it had files to work from. It deliberately ignores how much typing was involved: getting a hard job done in one well-aimed sentence is complex work, not simple work. That's effort, measured separately below.
how many tool calls (each one counts 1, up to 30)
+ extended thinking (counts 8 if it happened at all)
+ attachments (each one counts 2, up to 5)
- Simple
- Under 5 · a question and an answer, nothing else involved.
- Medium
- 5–14 · everyday assisted work with a few steps to it.
- Complex
- 15–29 · multi-step work with several tools in play.
- Deep
- 30 and up · long, tool-heavy work, usually with the AI thinking things through.
Effort
How much the person put in — how many times they came back, how much they wrote and handed over, and how many files they brought. This is the other half of the picture: two conversations can take the same effort and be worlds apart in complexity, and the pair together is far more telling than either alone.
how many times the person wrote (each counts 1.5, up to 20 turns)
+ how much they wrote or handed over (up to 30 — the largest term,
maxed out at ~12,000 input tokens)
+ attachments (each one counts 3, up to 5)
- Quick
- Under 6 · asked once, gave little context.
- Light
- 6–17 · a few rounds of polishing.
- Substantial
- 18–34 · a real working session with genuine revisions.
- Deep
- 35 and up · a long collaboration, lots of back-and-forth, often with files.
Tokens
Tokens are the units AI providers bill in — roughly, pieces of words. Every conversation is counted piece by piece with a real tokenizer over the actual text, not a "characters ÷ 4" rule of thumb, and split by who produced it, because the two directions are priced very differently.
- Input tokens
- Everything the model reads on a turn: the conversation so far, the new message, and anything a tool handed back. These pile up as a conversation gets longer, which is why long sessions cost more than their word count suggests.
- Output tokens
- Everything the model writes: its replies, the tool calls it makes, and its extended thinking. There are usually far fewer of these, but each one costs several times more.
- Agentic tokens
- The slice of the above that is machinery rather than conversation — tool calls, what tools returned, and extended thinking. Counting it separately is how the Cost tab can show what share of a bill went to autonomous work versus plain writing and answering.
- What isn't counted
- Images have no text to count, so they add nothing here — a picture-heavy conversation reads cheaper than it was. A single very large tool result is also capped, so one enormous file read can't dominate a month. Both make totals conservative rather than inflated.
Cost model
Published per-million-token rates × the tokens your team used. Reading and writing are priced separately because they differ by three to five times.
How sure we are — the three levels
A dollar figure is only as good as what we know about which model did the work. The Cost tab labels every provider with one of three levels, and each connection you add can only move it up. Wherever a number still leans on an assumption it carries a ~ and an est. — those marks are the honest part, not decoration.
- ○ Estimated
- We can see the conversations but not which model ran them, so a mid-range rate for that provider is applied. Good enough to compare topics, people and months against each other; treat the absolute total as a ballpark. This is where an uploaded export starts.
- ◐ Model mix
- Your provider's usage reporting tells us which models your team actually ran each day, so each day is priced on its own real blend. Token counts are calibrated against the provider's own reported totals too, which removes most of the remaining drift.
- ● Actual
- Every conversation carries the model that handled it, so nothing is blended or assumed. The highest level available, and what a Compliance API connection unlocks.
Levels are per provider, not per company — Claude can sit at Actual while a second provider is still Estimated, and the tab says so. Adding a source only ever improves the level; it never quietly downgrades one you already earned. Setting these up: the data sources guide ↗.
Rates we price with (per million tokens)
Published list prices, kept current with each provider's pricing page. If your company has negotiated rates these will read high — the comparisons between topics, teams and months hold either way.
These dollars are measured AI spend. They are a different thing entirely from the labor dollars on the Action plan and Toolbox, which value people's time at the hourly rate you set.
Token budget review
A focused, per-person read for one decision: should this employee get more tokens? Opened from the Cost tab's by-user table when someone hits a cap. SightLift measures the spend and offers an advisory read — it never approves anything; that happens in your own budget tool, which is why you type the “since” date in yourself (we can't see the approval).
The read
It reads efficiency, not return on investment: whether the spend looks like real work done efficiently — not what that work produced, which we can't yet trace from a session to the thing it shipped. Every threshold comes from your own team's numbers, never a fixed industry figure.
- Supports request
- Aligned, efficient, converging work.
- Supports, with notes
- Efficient, but some orphan spend or right-sizing worth noting.
- Needs your call
- Signals conflict — telemetry can't classify it; you decide.
- Worth a look
- Spend looks like spinning, not progress.
Alignment — is it real work?
The share of a person's tokens that went to work matching one of your company objectives (you write these in Settings → Company profile). A conversation on a topic an objective names specifically gets full credit; one that only falls under a goal you set for the whole company gets partial credit — otherwise one company-wide goal would make everything read as perfectly aligned. Personal conversations never count as aligned, whatever the goal says. Whatever's left over is unclaimed spend. Topics come from what the conversation was actually about, so nobody can improve their own number by labeling their work differently. No company profile, no alignment figure. This shows where the money went; it never pushes anything up or down your action plan. The Cost tab's Aligned column and its “Share going to priority work” tile credit each conversation exactly the same way — they only differ in what they divide by (the Cost tab compares against work spend, this review against every dollar in the window, so personal use lowers the number here but not there).
Efficiency
Three stand-ins for “done efficiently,” all measured against your own team rather than a fixed cut-off. Where there's enough history each one also shows the same person's previous stretch of equal length (“was …”), so improvement is judged against themselves and not only against colleagues. Fewer than five earlier sessions is never shown as a trend.
- Cost vs peers on similar work
- Their tokens compared with what work of the same complexity typically costs across your company. About the same is normal; less is better.
- Repeated or duplicate work
- The share of tokens spent going over near-identical ground again. Less is better.
- Overspent on simple tasks
- Simple conversations that cost more than almost all other simple conversations — expensive effort on cheap tasks. Less is better. (Spotting the other kind of mismatch — a top-end model doing a rename — needs per-conversation model data we don't have on every source yet.)
Agentic share
Of everything the AI wrote for this person, the share that was machinery — tool calls and extended thinking — rather than plain answer text. High means tool-heavy, autonomous work; low means mostly writing and answering. Neither is good or bad on its own; it only becomes a warning sign when someone's share sits in the top quarter for your team, which is what feeds the “spinning” read below. (The Cost tab shows the same idea across all tokens, weighted by dollars.)
Why the spend is high
Everyone reviewed is already a heavy spender, so raw size separates no one. The read looks at the shape of the spend on two axes — signs of progress vs signs of spinning — and routes into three outcomes.
- Signs of progress
- Depth (genuinely complex work), novelty (new ground, not re-tread), convergence (closing in over time).
- Signs of spinning
- Repetition (the same ground again), tool loops (machinery overhead in the top quarter for your team), no new surface, and simple tasks costing too much.
- Genuinely hard problem
- Strong progress signals, little spinning → supports the request.
- Unclear — your call
- Signals conflict → routed to you. The default when it's genuinely unclear.
- Rework without progress
- Three or more signs of spinning line up with little sign of progress → worth a look. Deliberately hard to trigger: wrongly flagging someone doing genuinely hard work is worse than funding some waste, so a short history is never flagged at all.
What this can't see: whether quiet work finished or was given up on — both look the same from here. Telling them apart needs a link from a session to what it produced, which doesn't exist yet. Until it does, steady progress toward a dead end reads as a hard problem.
Getting around
- The tabs, in order
- Dashboard is the summary. Action plan is what to do about it. Toolbox is every AI skill your team uses; Capabilities is where you publish and manage your own. Team holds the score and the per-person views, Explore the raw data and the detailed signals, Cost the spend. Revenue appears for people cleared to see customer and deal information.
- Clicking a number
- Almost every figure here opens the conversations behind it. Click a bar, a chip, a table row, or a card and you get a filtered, sortable list — click again from inside it to narrow further. The line across the top always says which filters are applied.
- Bars split by effort
- On the Dashboard, the "By topic", "By complexity" and "By output type" panels use bar length for volume and the shading inside for the effort mix. Hover any segment for the numbers.
- Hovering a person
- Hovering a person's row pops up their last 30 days as a calendar — one cell per day, darker on busier days. Clicking the row expands it into their topics, outputs, trend, tools and hardest conversations.
- MTD
- Month-to-date — everything since the 1st, compared with last month at the same point rather than with a full month (which would always look like a drop).
- ⓘ and ↗
- An ⓘ opens this glossary at the matching definition. An arrow like How your Toolbox works ↗ opens a longer guide in a new tab.
- Switching organizations
- If your account belongs to more than one organization, the header shows which one you're looking at and lets you switch. Each one is completely separate — its own data, and your own role and permissions within it, which may differ from organization to organization.
- Connecting a data source
- In Settings → Data sources, each provider connects one of three ways. A live API connection uses a read-only key to pull new activity on a schedule, going back as far as that provider keeps history. Live telemetry has the tools report activity as it happens — the only way to see Cowork, and it starts from the day you turn it on rather than filling in the past. A manual upload takes an export file and brings in everything in it, which is how you backfill history or work around a blocked outbound key. You can change how often we sync under Settings → Notifications & schedules. Step-by-step setup and a table of what each product can and can't report: the data sources guide ↗.
- What SightLift can see
- What you connect decides what we can measure, and some things simply aren't visible without a particular connection. Which level unlocks what, what we store, and how it's kept separate from every other company: data access & privacy ↗.
Agent opportunities
Work your team does over and over that a custom agent or skill could take on. We find them in two passes: first by grouping conversations that cover the same ground, then by having a model read each group and judge how well it would suit automation. Every candidate lands in the Detected column of the Capabilities tab with its full write-up; the strongest also become Automation plays on your Action plan. Both come from the same detection. This answers "what could we build?" — the Data Signals under Explore → Signals answer "what should we do about the people?"
- Ranking
- How often the work comes up, how many people across the team do it, and how well it would suit an agent — multiplied together, so something frequent but only one person's job doesn't beat something frequent and shared.
- Stars
- A one-to-five rating of how strong a candidate this is. More stars, better fit for automation.
- Tagline & problem
- One line saying what it is, plus the underlying problem this repeated work represents.
- What goes in, what comes out
- Roughly what you'd hand the agent and what it would hand back — enough to tell whether it's worth building, shown on the candidate's detail panel.
- Why it's strong
- The reasoning behind the rating: what specifically makes this pattern a good fit.
- Risks
- What to weigh before building — how accurate it needs to be, awkward cases, and where a person should stay in the loop.
- Savings estimate
- A rough figure for the time or cost an agent could save, shown only where the data supports one. Shown in dollars, that's the estimated time valued at the hourly rate you set.
- Who does this today
- The people doing this work now, plus real examples behind the candidate — click "Show spec" to read them.
Capabilities
A capability is one repeatable piece of work, packaged so your team can reuse it instead of redoing it. The Capabilities tab is where they're proposed, proven and published; the Toolbox is where you see how the published ones are doing. You don't build or configure any of this yourself — SightLift proposes them from work your team already repeats and asks you to approve the ones worth keeping. The full walkthrough is How capabilities work ↗.
Nothing is ever proposed off one person's habit: the same work has to show up several times, across more than one person, before it becomes a candidate.
- Automation
- Work where the same input reliably gives the same result — pull a list, reconcile a file, format an export. It runs as plain code with no AI involved at the time it runs, which is where the largest savings come from.
- Skill
- Work that needs judgment — drafting, summarizing, deciding. The steps are consistent even though the thinking isn't mechanical, so it runs as a guided recipe rather than fixed code. SightLift sizes the model for each step; your own tools run it, on your own account.
The columns on the board
Capabilities move left to right, becoming more trustworthy at each step. Nothing publishes itself — a person signs off before anything is served. A column only appears when something is sitting in it, so you won't see all of these at once.
- Detected
- We spotted repeated work and proposed something. Nothing is committed. Your move: promote it, or leave it.
- Testing
- Being checked, not yet trusted — either replayed against your team's own past results, or running quietly alongside real work while it collects results people confirm. Usually nothing for you to do; you're told when it's ready.
- Verified
- It passed, and the evidence is attached. Your move: approve it to publish, or hold.
- Live
- Published into your tools and being measured — how often it's used, how often people keep the result, and what it's actually saving.
- Needs attention
- Something live has drifted away from what people do now. The work changed underneath it. Your move: re-check it against fresh examples, or retire it.
- Retired
- Unpublished and no longer served — retiring it removes the tool from your team's assistants. Nothing is deleted: it stays on the board with its proof intact, so you keep a record of what you used to run.
How strongly something is proven
The proof is the whole point — it's what lets you trust a capability without reading its internals. We say plainly which kind you're getting, and never dress one up as another.
- Verified on real work
- It reproduced your team's own past results exactly. The strongest proof there is, and only automations can earn it.
- Confirmed in live use
- It earned its proof from real results people confirmed, because there was nothing in the past to replay against. An honest notch below the above.
- Quality-checked
- A graded assessment, used for skills. Judgment work can't be certified by reproducing it, so a skill is graded and never called "proven" — and it says so wherever it's served.
- Failing forward
- When a candidate can't reproduce every example, it isn't thrown away — it's offered as a skill instead, with the differences shown. A "no" to automation is really "this part needs judgment."
Testing runs in a sealed sandbox: no data leaves it, and nothing runs against anything live. Every proof keeps a de-identified copy of the examples it passed, so it stays valid even if the original conversations are later deleted — and the person who approved it and the date are always attached.
Actions
The Action plan takes the many Data Signals about the same person, topic or group and bundles them into a short, ranked list of things to actually do — "coach this person", "spread this topic beyond one owner", "run an AI 101 for this group". Every action shows the signals and conversations it came from (click a card to read them), says which part of the AI-Use Score it would move, and how many points it would add. The list is ordered by those points alone, biggest first. Actions that name an individual are shown only to admins.
Each one carries a single category — the colored chip — which decides what kind of thing it's asking of you and what document you can generate from it:
- Coaching
- A 1:1 with one person — they've plateaued, gone quiet, or their results lag their effort. Generates a talking-points doc for the conversation.
- Structure
- Spread out work that sits with too few people, or agree on a way of working — one person holding a whole topic, a topic nobody owns, or a team all doing the same thing differently. Generates talking points.
- Automation
- Build or roll out a tool — work being done by hand over and over, the same prompt reinvented by different people, or a strong automation candidate. The candidate's own panel holds the write-up you'd build from.
- Cost
- Cut spend or waste — one or two people driving an outsized share of the bill, a top-end model doing routine work, or the same job being done twice on two different providers. Generates a talking-points brief.
- Enablement
- Run an "AI 101" for a group stuck on small, hands-on tasks in the same area. Admins see who's in the group; everyone else sees only how many. Generates a session outline.
- Governance
- Sit down with whoever owns the data and review how much sensitive work — legal, finance, personnel, security or strategy — is happening and who's doing it.
- Win
- Recognize a positive pattern — a comeback, a first adopter, or a winning workflow worth spreading. Visible to everyone.
- Priority
- The 0–100 pill adds up the signals underneath: more signals, and more serious ones, rank higher. The list you see is kept deliberately mixed, so one category can't crowd out everything else.
- Evidence
- How many signals and how many conversations sit behind the action. Click the card to read them.
- Generate a doc
- Coaching, structure, cost and enablement actions have a 📄 button that drafts a ready-to-use document — read it in place, then copy or download it.
- Who can see what
- Anything naming a person or the members of a group is limited to admins and to people given sensitive-data access. Those names never travel — not in an email, a Slack message, or a downloaded document. You only ever see them inside the app.
Dollars on an action — your loaded hourly rate
Some automation plays carry a dollar figure ("~$4,200/mo of manual time", "~$800 of avoidable rework"), as do the savings on a published capability. These are labor dollars, and the arithmetic is deliberately simple: estimated time × your loaded hourly rate. The time comes from your own data: how long this work is reckoned to take, multiplied by how often we actually saw it happen. The rate is a single number an admin sets in Settings → Monthly AI budget — an average hourly cost for the people doing this work, salary plus everything on top of it. Until someone sets it, a default is used, and every figure that depends on it says which rate it used when you hover.
This is money your team could redirect, not money that appears in a bank account — it values hours at what you told us an hour costs. It is also entirely separate from the dollars on the Cost tab and in an action's score impact, which are measured AI spend and carry no wage assumption at all.
Revenue & accounts
The Revenue tab asks where AI-assisted work goes across your customers, and what it changes. Accounts are picked up from conversations with no setup at all; uploading your account list or connecting your CRM firms up the matches and unlocks the deal views. It appears for admins and for people an admin has given revenue access (Settings → Users) — and only once there's account or deal data to show, so a new grant on an empty tenant isn't broken, there's just nothing to put in it yet.
Everything here that compares outcomes is a description of what went together, not proof that one caused the other. Rows are your customers — this tab never ranks, totals or compares your own staff. "Worked by" is a fact about the account, the same way your CRM shows who touched it.
The three views
- Accounts
- Every customer your team has done AI-assisted work for, how much, who worked it, and when. Available with no setup — this is the view you get on day one.
- Coverage
- The opposite question: which accounts are getting no AI-assisted work. This one needs a list of your accounts to compare against, because without one the question genuinely has no answer — so instead of showing you a misleading zero, the view says what it's missing.
- Impact
- Two views of "does it help", at different levels, neither one demoted to the other. Per deal: do deals with AI-assisted work behind them close better? Comparing deal to deal controls for who the seller is in a way that comparing sellers can't. Per person: who is turning AI use into results? Each needs its own data and says so plainly when it's missing.
How sure we are it's really that account
Three ways an account gets identified, and the difference is always visible on the row — a guess never gets styled like a certainty.
- From a web or email domain
- The conversation mentioned a domain that belongs to a known account. The strongest match, and it renders plainly.
- From your list
- The name matches an account you uploaded or that came from your CRM. Also renders plainly.
- Detected ~
- A company-shaped name that keeps coming up, with nothing to anchor it to. Shown with a ~ and a "detected" chip because it is a guess — it can pick up a partner, a competitor, or a name that just looks like a company. Confirm the ones that are real in Settings; until you do, they never write back anywhere.
Terms on this tab
- Single-threaded
- Only one person has done AI-assisted work on this account. Not a problem by itself, but worth knowing before that person takes leave or moves on.
- Shared sessions
- A conversation that covered more than one account. Its effort and dollars are split across them, so nothing gets double-counted.
- The 12-week strip
- Account work comes in bursts rather than a steady trickle, so recent weeks are drawn as separate cells instead of a line. A line would imply a trend that isn't really there.
- Untouched large deals
- Open deals above the tab’s large-deal threshold ($50k today, and the view states it) with no AI-assisted work behind them. Not a verdict on the deal — a list of the biggest places your team's proven approaches aren't being used.
- Handoff brief
- A one-page summary of everything the tab knows about an account, for handing over to a colleague. It's built only from facts already on screen, and nothing leaves until you've read it and said so.
- Small accounts still show
- Unlike the people-based views, there is no minimum here — a customer appears even if one person worked it once. These are companies, not employees, so the privacy floors that protect individuals don't apply. Sensitive conversations are still held back, and the tab says how many.
Data Signals
Under Explore → Signals you'll find these — patterns found by counting, the same way every time, across every conversation, and sorted by how much they matter. Each one is grouped into a family (the colored chips below). A signal points at a topic, a person or a habit, and always links to the exact conversations behind it; click any card to read them. Signals that name someone, or that touch sensitive work, are shown only to admins.
These are the raw findings. The Action plan is where many signals about the same person or topic get bundled into one thing to actually do.
Trend usage shifts over time
- Topic trend
- A topic's usage jumped or dropped sharply month-over-month.
- Person trend
- A person's overall usage jumped or dropped sharply month-over-month.
- New topic
- A brand-new topic appeared this month with real volume.
- Topic went quiet
- A previously-active topic dropped off — automated, moved, or abandoned?
- Consolidating
- A topic is shifting into fewer, longer sessions — ripe to formalize into a workflow.
- Output mix shift
- A topic's mix of output types changed notably (e.g. a surge in one artifact).
- Tool-use spike
- Agentic/tool use in a topic jumped — a pattern worth formalizing into a tool.
- Seasonal
- A topic recurs on a predictable cycle — pre-stage or automate ahead of the next peak.
Person individual trajectories
- Rising complexity
- A person's work is steadily getting more complex — showcase it or have them mentor.
- Stuck on simple
- A tenured person has stayed on simple tasks — a nudge could unlock higher-leverage use.
- Dormant
- A once-active person has gone idle — re-engage before they churn.
- High effort, few tools
- Heavy work done with few tools — agentic features could speed them up.
- Chat-only
- Real volume with zero tool use, on a team where peers do use tools — the capability is proven next to them.
- Prompt retyper
- One person retypes a near-identical prompt over and over — save it as a skill.
- Plateau
- A once-rising person's complexity rose then went flat — a stretch project could restart growth.
- Low artifact yield
- Heavy effort but few concrete artifacts — check for a blocker, tool, or template gap.
Cost token spend & model fit
- Topic cost concentration
- A topic is a large share of token spend — highest-ROI target for a shared template.
- Person cost concentration
- A person is a large share of token spend — confirm it's intentional heavy use.
- Token-efficiency outlier
- A person uses far more tokens per conversation than peers on the same topic.
- Budget review candidate
- Heavy recent spend that looks like circling, or that we can't call either way — open their budget review (Cost tab → By user) before approving more tokens. People doing genuinely hard work, and people with too little history to judge, are never flagged.
- Artifact cost outlier
- A team spends far more tokens per artifact (memo, deck, analysis…) than the org median for that output type — start with the two driving workflows. Teams doing measurably heavier work are never flagged. (Counted per team, never per person — the individual figure is admin-only.)
- Model downgrade
- Lots of simple-tier work on an expensive (Opus-class) model — a cheaper one would do. (Only where we can see which model ran each conversation.)
- Model underpowered
- Complex/deep work concentrated on the cheapest model available — a more capable one is likely worth it. (Only where we can see which model ran each conversation — a claude.ai export doesn't say.)
- Model-mix savings
- Org-level run-rate: routing routine (Simple/Medium) work to the cheapest in-family model would save a meaningful monthly amount. (Needs a live connection to your provider's usage reporting; the figure is an estimate and is marked as one.)
Risk concentration & provisioning
- Hidden expert
- Deep work in a topic is held by only one or two people — a single point of failure.
- Bus factor
- One person owns a large share of a topic — de-risk the silo.
- Hard-case outliers
- A mostly-simple topic with a few hard outliers — the hard cases may need a dedicated tool.
- Complexity variance
- Complexity varies widely across people on one topic — share the efficient approach.
- Over-provisioned thinking
- Extended thinking used on simple tasks — right-size the model to cut cost.
- Effort ↔ tier mismatch
- High effort spent on simple-tier tasks — a template or faster model could cut the grind.
- Abandoned long conversations
- Long conversations that produced nothing concrete — an unmet need, or a workflow that stalls.
- No agreed workflow
- No consistent output pattern across a team for a topic — align on a shared approach.
- Output anomaly
- A topic usually yields one output type but some conversations yield another — mislabeled or cross-functional?
- Automation candidate
- High-effort, zero-tool manual grind — a tool or skill could absorb it.
- Under-provisioned thinking
- Deep-tier work running without extended thinking — likely a quality lift.
- Knowledge hub
- Many people's work resembles one person's — a de-facto expert (and a dependency).
- Orphan topic
- A topic with real volume but no clear owner — assign one before it falls through the cracks.
Workload when & how hard people work
- Off-hours load
- A meaningful share of work happens outside the team's active hours — check on workload.
- Weekend load
- A meaningful share of activity lands on weekends — check workload and balance.
- Crunch
- Off-hours/weekend load jumped sharply vs baseline — possible crunch before it burns people out.
- Working-hours creep
- A person's active-hours window is widening over months — check workload.
Friction workflows that fight back
- Retry / rework
- Repeated short same-day attempts on a topic — the workflow or tool likely needs fixing.
Connect collaboration opportunities
- Cross-pollination
- Two people both heavy in a topic but with divergent approaches — introduce them.
- Reinvention
- Two people solving near-identical problems separately — point the later one at the earlier result.
Tools capabilities to spread
- Undiffused tool
- A useful capability is stuck with one person — spread it to the team.
- Shared prompt template
- Several people reuse a near-identical prompt — turn it into a shared template or command.
Roster membership hygiene
- Underutilized seat
- A person has a seat but almost no activity — onboard or reclaim it.
- Role adoption gap
- A person lags peers in the same role — a possible onboarding or enablement gap.
Multi-source provider spread
- Multi-provider overlap
- One person's work on a topic is split across multiple providers — consider consolidating.
Data quality trust the signals
- High uncategorized
- A large share of a person's work is unclassified — a blind spot to review before trusting other signals.
Win recognition
- First deep task
- A person just completed their first deep-tier task — recognize the milestone.
- Comeback
- A dormant person re-engaged after a long gap — welcome them back and learn what worked.
- First adopter
- The first person org-wide to use an output type new to the team — have them show others.
- Efficient producer
- Converts effort into concrete artifacts at an unusually high rate — worth learning from.
Governance sensitive activity admin-only
- Sensitive category exposure
- A sensitivity category (legal, finance, personnel, security, strategy) has enough volume to warrant a governance heads-up — confirm access controls.
- Sensitive activity by person
- An individual with a notable count of sensitive conversations — confirm handling and access. (The category comes from reading the conversation, with a keyword check as a backstop; only the category name is stored, never the text.)
- Do-not-paste patterns in chats
- Conversations contain a pattern your AI use policy says should stay out of chats (an email address, phone number, or card/ID number). (Matched on the shape of the pattern against the kinds you named in your policy; only the kind is stored, never what was matched.)
The signals below join AI usage with sales outcomes from your CRM (HubSpot) — they appear only when a CRM is connected and the data is matched. All are descriptive associations, not causal claims, and are admin/RevOps-only.
Outcomes plays to spread (team-level) admin-only · needs a connected CRM
- Winning workflow
- The AI output type top-revenue reps produce far more of than the lightest-usage quartile — a play to capture and spread.
- Winning play (subtopic)
- The specific sales subtopic top-revenue reps use AI for far more than the lightest quartile — the play to codify and coach.
- Agentic leverage
- Top-revenue reps run more tool/agentic AI workflows on sales work — automate the steps the winners automate.
- Cycle advantage
- The heaviest-AI-usage quartile closes deals materially faster — package their workflow as a team playbook.
- Usage without lift
- Top-usage reps use far more AI but show no win-rate edge — check whether it's going to the highest-leverage work.
Outcomes people to act on admin-only · needs a connected CRM
- Coaching candidate
- A strong closer sitting in the lightest AI-usage quartile — a quick adoption win to lift their output.
- Adoption win
- A rep who recently ramped AI on sales work and is closing well — a within-rep adoption story to replicate.
- Effort misallocation
- A rep pours high-effort AI into sales work but converts below the team median — the workflow may be wrong.
- Reinvention drag
- A below-median closer keeps re-deriving AI work others already did — hand them the team's proven templates.
- Expert, low yield
- A de-facto AI expert whose own deals lag — coach on execution, or lean on them to teach the winning play.
- Engagement decline
- A rep's sales-topic AI usage is falling while results stay weak — an early churn signal to act on.
- Denial loop
- The same tool got blocked for approval over and over — approve it up front, or publish it as a capability. (From live telemetry; counted per team, nobody named.)
- Error-retry loop
- A tool errors and gets retried in the same session over and over — a broken step burning model turns. (From live telemetry; the wasted spend is the real tokens burned between the error and the retry.)
- Repeated context sessions
- Long sessions send the whole conversation again every turn instead of reusing what was already sent — that repetition is the cost. (From live telemetry; counted per team.)
AI-Use Score
A number from 0 to 100 for how well your team uses AI — not just how much. It's the average of the five dimensions below (each also scored 0–100), using the weights you set. Every point traces back to something we counted in your own data — conversations, seats, dollars, shared-tool use — and the score page always shows the handful of numbers behind each dimension. Shown per team and for the whole org as a tier ribbon; the Weekly / Monthly / Quarterly toggle zooms the lookback (≈6 months / ≈18 months / ≈3 years). It's a steer, not a measurement instrument — read the band it lands in and which way it's moving, not the exact number. There is no individual score, and there never will be (see Privacy). The full walkthrough is How your AI-Use Score works ↗.
composite = weighted average of the dimensions we can measure
(weights equal by default — set them in Settings)
A dimension we can't measure yet is skipped, never guessed — the average is taken over the rest, so an unactivated Reuse or Governance dimension can't drag the number down.
The five dimensions
- Adoption is the whole team on board?
- Three ingredients, all counted from your usage: Activation (of the seats you pay for, how many were actually used this month), Regularity (how many days per month active people use AI — a daily habit beats a monthly binge), and Balance (whether usage is spread across the team or concentrated in one or two heavy users). Goes up when quiet seats come back to life, people build a habit, and usage stops depending on a single hero.
- Sophistication is AI doing real work?
- Two ingredients, graded on every conversation: Complexity (how big the task was — a quick lookup grades low, a multi-step build or analysis grades high) and Output (code, analysis, plans, and documents count for more than a quick email rewrite). We grade the work, not the typing — using a skill to do a complex job in one step counts as the complex job. Goes up when routine users graduate to bigger tasks and more conversations end in substantial output.
- Efficiency how much of your spend does real work?
- Six leaks, all counted in dollars and compared to your total spend: retry loops, abandoned sessions, error churn, re-sent context, model overkill (premium models on routine tasks — for most teams the biggest leak), and non-work spend (only if your AI use policy says it counts). We grade the leaks, not the bill: 100 means no measurable waste, and spending more never lowers this score — only wasteful spend does. Goes up when the waste dollars shrink — each Efficiency action names the exact leak and what it costs per month.
- Reuse does the team run on shared, proven tools?
- Two ingredients, counted across your shared tools (skills, templates, automations): Used, not redone (when work matches a shared tool the team already has, how often the tool was used instead of the work being done by hand) and Coverage (of the repeat work we can see, how much has a shared tool at all). We count the habit, not the logo — skills your team built before SightLift count from day one. Goes up when people adopt the standard tool instead of hand-rolling the work, and when repeated workflows get captured and published to your toolbox.
Not scored until we can see shared-tool activity or your first published capability — until then the score is simply the average of the other dimensions.
- Governance is sensitive work in the right hands?
- Two ingredients, measured only when you give us the context to judge fairly: Right hands (once we know each person's Function Sensitivity — set in Settings → Org & roles — the share of sensitive work in legal, finance, personnel, security, and strategy handled by people cleared for it) and Policy compliance (once we have your AI use policy — sensitive work staying on company accounts, AI work staying on approved tools, personal data staying out of chats). Your governance score is the average of the rules we're actively checking, and each violation pattern becomes an action on your plan. Without roles or a policy, the dimension isn't scored and sensitive findings appear as notes instead. Goes up when sensitive work lands with the right people and policy violations fade to zero.
Weights you control
The score uses your weights — equal by default (each dimension counts the same). Set them in Settings → Company profile → Score weights to match the behaviors you want to drive, and change the balance as your priorities evolve. Weights are a strategy lever, not a way to inflate the number: raising a weight makes that dimension matter more and makes its gaps cost more.
How actions connect to the score
Every action on your Action plan names the dimension it moves and shows the points it could add — computed from your own numbers, with the math on the card (e.g. "re-activating 4 quiet seats raises Activation from 67% to 100% — about +3 points"). The plan is ranked by those points, biggest first. Two kinds of projection, always labeled:
- Measured
- Worked out straight from things we counted — times a shared tool was skipped, wasted dollars, unused seats.
- Estimate
- Rests on one assumption, which is always written on the card — for example, "assumes this group's routine work moves up one level of complexity".
Some cards also show a dollar figure (like the monthly cost of manual work an automation would remove). Points rank the plan; dollars size the prize.
Tiers & movement
Banding is deliberate, so a 1-point wiggle can't read as a trend:
- Needs focus
- < 50
- Developing
- 50–59
- Solid
- 60–69
- Strong
- 70+
- ↑ / → / ↓
- Period-over-period direction (vs the previous bucket at the current grain). A change under 3 points shows as → (flat) on purpose — noise shouldn't look like movement.
Reading the ribbon
- Teams
- Grouped by the department you set in Settings → Org & roles where you've set one, and otherwise by what each person mostly works on. Someone's access level in SightLift — admin, member, viewer — is never used as a team.
- Hover a cell
- Shows what's behind that score — each of the five dimensions and what it contributed.
- Click a cell
- Opens the conversations behind that team-month.
- Not enough data
- A team with too few people or too few conversations to score reliably says so, instead of showing a number that would wobble on one person's week off.
Privacy
Scores exist at the team and company level only. There is no individual score, and there never will be — managers drill into examples of work patterns, never into a ranked list of people. CRM outcomes, where connected, only help validate the weights — they are never an input to the score.