How it's used
One number. Four honest ways to use it.
The same measurement — outcome per dollar — reads four ways: a private mirror for one engineer, a name-free prompt for a team, a rollup for a manager or CFO, and a shared signal a group improves against together. Every use below is name-free or self-only. None of them ranks one person against another.
Four intents
Grade the work. Never the worker.
In its anonymized modes, TIER's served unit is the team over a quarter. A per-developer number exists in developer mode, but the policy TIER ships is that it stays a private self-diagnostic for an individual to raise their own yield — never a ranking input. That guardrail shapes all four uses.
See your own yield, privately.
tierd score --repo . runs on your laptop and shows where your own AI money went — no server, no token, nothing leaving your machine. In opt-in developer mode you see your own outcome-per-dollar over time. Read the dashboard's unattributed spend — mainline work not yet tied to a merged outcome, which says what TIER could not see, not that the work was wasted — and decide where to tighten: issue attribution, model choice for the task, less undirected wandering.
Move practices, not scores.
Run a standing retro on the name-free levers — cache-read share, premium-model share, unattributed spend, and yield within a work type. None names an individual, so they project safely on a shared screen. Agree one practice change, then test it with /scores/compare: two windows, before and after, with a CI-honest significant flag that is true only when the confidence intervals do not overlap.
A yield next to the line item.
Team and division rollups under a k-anonymity floor, so leadership sees efficiency without ever surfacing a name. For finance: Spend Leverage — what your teams' usage would cost at list price against what you actually paid (about 1× on pay-as-you-go, higher on a flat plan) — the figure that decides subscription vs. per-token, which nobody computes today. Plus unattributed spend: how much went to work not yet tied to a merged outcome.
Coach on the same transparent number.
Everyone reads the same number, in the open — so it guides how a group optimizes spend, not how it judges people. Rubric-version and price-version stamps keep comparisons honest as the scoring evolves; the CI-honest significance test stops a lucky month from being read as a win. The purpose is to spread high-yield practices and steer the token budget, never to appraise.
The CFO number
Spend Leverage decides the contract.
Renewals and per-seat vs. per-token decisions happen blind today, negotiated against a usage chart. Spend Leverage is the missing figure: what your teams' usage would cost at list price, divided by what finance actually paid. On pay-as-you-go it reads about 1×; above 1×, a flat plan is winning.
$10,000 of usage at list price on a $4,000 invoice = 2.5×
Costs are stored as integer micro-dollars, priced from a versioned reference table, and reconciled against the actual invoices finance posts — credit memos enter as negative rows, and the audit trail is row history, never an overwrite. Every dollar in the denominator traces to a table entry a CFO can check.
The hard line
What TIER is for, and what it is never for.
This is a policy, not a suggestion. A per-developer number is self-view or opt-in only; a manager may never request that an individual's number be shared or screenshotted. The request itself is out of bounds, regardless of the answer.
| Never do this | Do this instead |
|---|---|
| Rank teammates against each other by TIER | Compare the name-free levers and copy the practice behind the better one |
| Tie a tier, or any lever, to pay, promotion, or a review | Use it to coach token-spend habits; keep it out of appraisal entirely |
| Ask for, share, or screenshot another developer's number | Keep individual numbers self-view / opt-in; project only name-free aggregates |
| Read a lucky month, or a team-mode delta, as proof a change worked | Treat it as directional; re-check next cycle before claiming an effect |
| Compare yield across different work types | Compare within a single work-type segment (bugfix vs. bugfix) |
| Trust a short, recent after-window at face value | Account for the windowing skew before reading it |
Honest by design: yield reflects task mix and context, not raw talent — a senior untangling legacy code can score below a junior on greenfield. TIER is a lever for optimizing spend, and it ships its own guidance that it is not a performance-appraisal input. The anonymized team and division modes (k-anonymity floor, default 5, hard minimum 3) exist precisely so an organization can run every workflow above without ever naming an individual.
Read the peer-learning playbook → · why every metric pointed at individuals was gamed →