Access metrics answer the wrong question

Most AI adoption reports open with the same numbers: licenses issued, accounts activated, weekly logins. These metrics rise reliably in the first months because what they measure is rollout logistics — provisioning, announcements, mandatory onboarding. Not one of them requires anyone's work to change. A team can log in every Monday, glance at the tool, and finish the job exactly the way it always has.

Here is a usable decision rule: a metric counts as adoption evidence only if it could not have been produced without the work itself changing. "Opened the assistant" fails that test. "Completed the proposal inside the assistant", "reviewed the sources and accepted the recommendation", "moved a first draft to approval largely intact" all pass it. Access tells you who could adopt. Behavior tells you who did.

Pair every leading indicator with the outcome it predicts

Behavior-change metrics come in two speeds. Leading indicators move within weeks: whether users start and finish the workflow inside the product instead of exporting halfway through and finishing in the old process; how human-approval interactions behave — acceptance rates, return-for-correction reasons, time to decision; and how much of a generated draft survives review, visible as edit and rework rates.

Lagging outcomes are what the business actually buys: cycle time, error cost, less rework in downstream systems. They usually need a quarter or more of clean data before they mean anything. Two failure modes follow from confusing the speeds. Judging a rollout on cycle time at week six means reading noise, and it can kill a product that was working. Celebrating leading indicators for a year without demanding outcome movement means funding a tool that changed clicks, not results.

The fix is structural: every leading indicator on the dashboard is written next to the lagging outcome it is supposed to predict, with a date by which that outcome must move. If the date passes and nothing has moved, either the causal story was wrong or the product needs to change.

Measure by role: the weakest handoff sets the ceiling

Company-level adoption rates hide the only structure that matters: the workflow. A drafting tool can be loved by the people who create documents and abandoned by the people who review them. The average looks healthy while the workflow is broken exactly at the handoff. Aggregate numbers cannot reveal this, because adoption failure is almost never uniform — it concentrates in the role that carries the effort while another role collects the benefit.

So segment by role in the workflow: who initiates, who reviews, who approves, who consumes the output. When one role's usage collapses, treat it as a product finding, not a training gap. The usual mechanism is asymmetry: the tool asks that role for new effort — checking, correcting, re-entering — while its benefit lands somewhere else. Repeating the training does not remove that effort; only a change in the product itself corrects the asymmetry. The decision rule: track adoption at every point where the work changes hands, and assume the weakest role sets the ceiling for the whole system.

No vanity precision: every number carries its assumption

Adoption reporting drifts toward false precision — a productivity percentage computed from self-reported estimates and a denominator nobody can reconstruct. The honest alternative is not fewer numbers; it is numbers that state their own limits. If cycle time was never measured before rollout, say so and start the baseline now rather than reconstructing one from memory. If the pilot has nine users, report counts, not percentages. If another process change shipped the same month, name the confound.

A practical test: would this number survive a skeptical CFO asking exactly how it was computed? If not, it does not belong on the dashboard. A defensible range beats a confident decimal, and "we do not know yet — measurement starts this month" beats both when it is the truth. Teams trust dashboards that admit uncertainty, and quietly discount the ones that never do.

A worked example: one workflow, one page

Consider a hypothetical: a mid-size company pilots a proposal-drafting agent with one sales team, with human approval on pricing and commitments. Its adoption dashboard is a single page — two rows plus footnotes. The behavior row holds the share of proposals started and completed inside the agent versus the old template path; the share of each first draft that survives review, tracked as an edit rate; and approval interactions — accepted, returned for correction, and the top three return reasons.

The outcome row holds first-draft turnaround, approval-cycle duration, and commercial errors caught after sending — each labeled with the date it becomes meaningful, because a few weeks of noise proves nothing. The footnotes state the baseline source, the user count, and the known confound. Notice what is absent: no satisfaction score, no company-wide usage percentage, nothing the team cannot act on. Every line implies a possible product change — which is exactly the point.

Adoption data must change the product, not the deck

The purpose of adoption measurement is iteration — the Compound step of the method, not the steering-committee slide. Recurring return-for-correction reasons are a template or input problem; fix them. A role completing its work outside the tool is an integration problem. A leading indicator that never converts into outcome movement is a prioritization problem. Every signal maps to a product decision, and a dashboard that produces no decisions is decoration.

That takes a mechanism, not good intentions: a fixed cadence where the people with the authority to change the product review the data, and every review ends either with a change or with an explicit, recorded "no change". Then usage data compounds into a better product instead of a thicker report.

Where to start: pick one workflow that is live or close to it. Before the next rollout phase, write down three behavior metrics and the lagging outcome each one predicts, name the role segments, state the baselines and assumptions, and put the review cadence on the calendar. If today's measurement cannot answer "whose work changed, and what did we change in the product because of it", that is the first gap to close.