Methodology Measurement spec · v3.23

How Depthdata measures AI at work.

The complete, open account of how Depthdata turns AI tool usage into governance and cost-per-outcome numbers a CFO, CIO, or auditor can trust. Every method here ships inside the product and with every export, because a number that cannot survive being checked should not be on the dashboard.

What Depthdata readsRead-only · metadata-scoped
Seats & sessions
Measured
Usage signals & spend
Measured
Completed-work metadata
Measured
Prompt / message / ticket text
Metadata-scoped APIs. No content, ever.
Never read
The pipeline / 01

Five stages. The pipeline never skips one.

Every number in Depthdata moves through the same five stages, in order: raw signals are ingested, normalized to a common shape, compared to a baseline, estimated only where estimation is unavoidable, and only then scored.

01INGEST
Raw signals, read-only
Pull tokens, spend, seats, sessions and completed-work metadata from vendor admin APIs. Incremental and cursor-based.
02NORMALIZE
Resolved to one person
Every signal is tied to a canonical person via directory identity (SCIM). Unmatched records are quarantined, never guessed.
03BASELINE
Graded against you
Each person and team is compared to their own pre-adoption history, per department. Never an industry benchmark.
04ESTIMATE
Modelled, with a band
Where a value cannot be measured directly, it is modelled from signals and samples and labelled, never disguised as measurement.
05SCORE
Published formula
Normalized, baselined and estimated inputs combine into scores with stated weights. The formula ships with every export.
The pipeline never inverts. Estimation is the fourth step, never the first.
Confidence labels / 02

Every number wears a label. The label is part of the number.

The fastest way to lose a finance team’s trust is to present an estimate as a fact. Every figure on every surface carries one of three labels, not as a footnote, but as part of the number itself.

Measured
Pulled directly from an API, unchanged. Tokens, spend, seat counts, raw completed-work counts. If the vendor returns it, we report it.
Derived
Computed from measured inputs, with the formula shown. Output-per-token, cost-per-outcome, weighted output, reproducible from the inputs.
Modelled
Estimated from signals and samples, always with a ± band. Baseline-relative lift, hours saved, never presented as proof.
A dashboard where measured and modelled numbers look identical is a dashboard you cannot take into an audit.
Cost & output / 03

A token is a cost, not a result.

Two people each spend a million tokens; one closes a hundred tickets, the other ten. On a usage report they look identical. Depthdata measures both halves of the equation so they never do. Output is weighted, not raw, tickets by story points and cycle time, pull requests by reviewed size, deals by amount. Both raw and weighted ship; weighted is labelled Derived with the formula shown.

Denominator · cost
Tokens, spend, adoption, per person, from vendor admin APIs.
Numerator · output
Resolved issues, merged PRs, closed deals, from the systems where work is recorded, joined on identity.
Weighted, not raw
Complexity weighting on both sides, so neither tokens nor output can be farmed. Cost-per-outcome is reported as a measured ratio, never a claim that AI did the work.
Honest gaps / 04

Where we cannot see, we say so.

Some roles have a clean, per-person unit of output. Some do not. Depthdata says which is which rather than manufacturing a number to fill the grid. A visible gap is more credible than an invented figure.

FunctionWhat is measuredWhere it comes from
EngineeringResolved issues, merged PRsJira, GitHub, Linear
SalesClosed-won deals, weighted by amountHubSpot, Salesforce
SupportSolved tickets, weighted by typeZendesk, Intercom
Product / Design / LegalNo clean per-person output unit, flagged, never estimatedNo signal
AI-console connectors
ChatGPT Enterprise, Claude, Copilot, Gemini, Cursor and others supply the cost side: usage, spend, seats.
Output connectors
Jira, GitHub, Linear, HubSpot, Salesforce, Zendesk supply the production side: completed-work metadata, joined on identity.
Content
Never read, by either class. Every scope is read-only; no agent is installed on any employee device. It is a property of the architecture, not a setting.
The principles / 05

The whole methodology, in one place.

Seven principles hold every number on every surface to the same standard.

Read-only, always.
Metadata-scoped API access. No writes, no device agents, no scraping, no content.
Prompt text is never read.
A property of the architecture, not a policy layered on top. Nothing in the pipeline ingests a prompt.
Every number is labelled.
Measured, Derived, or Modelled. Estimates are visible as estimates, with bands.
Baseline against yourself.
Lift is measured against your own pre-adoption history, per department, never a generic benchmark.
Weighted, not raw.
Complexity weighting on both sides of the ratio, so neither tokens nor output can be farmed.
Honest gaps.
Where a role has no clean output signal, we flag it. We never fabricate a number to fill a grid.
Open methodology.
Formula, weights and definitions ship in-app and with every export. This document is part of that commitment.
If a number cannot survive you checking it, it does not belong on the dashboard. That is the whole methodology in one sentence.
See it applied
The same numbers, end to end.

The case study walks this methodology through one company’s workspace, every figure carrying the label the app would give it.

Modelled workspace, not a customer. Every figure on this page matches the Depthdata product demo and carries the same confidence label the product would give it.