Anonymized portfolio copy — names and customers replaced; all data illustrative.
Token Cost Dashboard — Spec & Lo-fi Wireframe
v1 starting point for the a top-5 US bank MVP dashboards. Target: functional in ~5 days, function over polish (the CEO, 6/9 sync). Hover any tag for the source quote..[Red brackets] = unknown / decision pending — several depend on the CEO's ongoing research, deliberately left open. Nothing fabricated.
Feedback sources — hover any tag for the quote
CEO-n the CEO, leadership sync
DP a team member / a top-5 US bank pain points
MS a prior enterprise prospect
1PGR a top-5 US bank one-pager (sent 6/9)
WG a Senior Engineer
Decisions & status
TM Trent design direction
PHASE 2 explicitly deferred
FLAG open question
Review annotations — internal only, not product UI
✦ manager takeaway what a manager should leave with
⚠ weak takeaway fails the test: would a manager care?
Built for the engineering manager SINGLE PRIMARY USER
The transcript is explicit: "how much my team is spending on tokens" CEO-1 and "from a dashboard perspective for a manager". CEO-3Every panel answers a manager question: what's my team spending, against what budget, on what work, and what signals did my engineers get. Developers don't use this UI — their counters, warnings, and halts live in the IDE (AITrax extension) and roll up here in the Team Signals panel. A CIO/platform-owner view is a possible separate surface — see "Not in the 5 days" below.
Platform scope for MVP: Copilot first (a top-5 US bank is Copilot-only per the call); Codex and Claude Code already supported, Gemini CLI not needed. Runs on-prem, on the client's PostgreSQL, fed by the AITrax IDE extension + agentic monitor over the real-time bridge.
Team Token Dashboard — the wireframeMVP · 5 DAYS
Three interaction rules apply everywhereTM: every engineer name links (↗) to their existing developer profile page; every PR links to the pull request in the client’s SCM (GitHub today — Bitbucket when that integration lands WG2); and every KPI carries a trend vs last month. See something → investigate it; numbers are only impactful with context.
Spend right now
✦ Manager takeaway · internalIn one glance: am I okay on budget, who is running hot, and did anything get stopped.
Team Spend This Month CEO-1
[X.XM tokens · $X,XXX]
vs budget [$X — budget source TBD]FLAG
62% consumed · [X days] left in cycle
vs last month: [±X%]
✦ Manager takeaway · internalI know exactly what my team has spent, against what budget, and how much runway is left.
Instrumented at the IDE + agent level — not reconstructed from billing. That accuracy is the point: a team member’s homegrown tracking script "didn’t work… not accurate at all". DP Near-real-time; the manager view "doesn't have to be quite so real time" as the IDE counter. CEO-3 Refresh cadence: [TBD].
Burn & Forecast DP
[$X / day]
current burn · projected month-end: [$X — over/under budget]
vs last month’s pace: [±X%]
✦ Manager takeaway · internalI will know we are heading over budget before it happens — not after the bill.
The cliff problem is a trajectory problem — forecasting turns "who is at 90%" into "who will be, and when." Denominator waits on the budget-source decision. FLAG
Wasted Spend 1PGRFLAG
[X% · $X]
of this month’s tokens went to abandoned / discarded work
vs last month: [±X pts]
⚠ Weak takeaway · needs workWhy weak: the takeaway is right, but the panel cannot be trusted on day one — the token↔abandonment correlation is unconfirmed for the 5-day MVP. Fix: keep the panel; get a Senior Engineer’s feasibility call first. If it slips, ship without it and frame it to a top-5 US bank as the first fast-follow — never show an empty or estimated waste number.
Promised in the one-pager and already measured in the Engineering Health Report (AI Waste, abandonment telemetry) — so it belongs here, not in phase 2. FLAG: engineering to confirm the token↔abandonment correlation fits the 5 days (a Senior Engineer leaned "total use first" WG1). Advanced cuts stay deferred.
Engineers by Budget Runway DP
<50%[n]
50–75%[n]
75–90% ⚠[n]
>90% ⚠⚠[n]
warning bands (≥75%): [n] engineers vs [n] last month
✦ Manager takeaway · internalI know who is at risk of hitting their limit — before they do.
Kills the cliff problem at the team level: the manager sees who's running hot before the 90→100 jump with no signal.
Live Right Now CEO-4
[N]
active AI sessions · [n] approaching session cap ⚠
⚠ Weak takeaway · needs workWhy weak: a live session count is ambient awareness — no manager decision follows from the number alone. Fix: fold it into Team Signals as a “live now” header stat, or earn the slot by making it actionable: sessions approaching their caps, with engineer links to intervene.
Copilot’s billing delay is exactly why a top-5 US bank can’t intervene in time today DP — this panel exists because our data is instrumented live, not billed later. A session creeping toward its cap is the manager’s preview of an intervention.
Team Signals — in-IDE events, rolled up CEO-4DP
[N]
halts this week ([±n] vs last week) · [X tokens] consumed before halts · [X%] resumed
When
Engineer
Signal
State
[t]
[dev] ↗
⛔ halted — session cap
resumed ✓
[t]
[dev] ↗
⚠ 85% budget warning
—
[t]
[dev] ↗
⛔ halted — burn rate
halted
✦ Manager takeaway · internalNothing my engineers experienced — warnings, halts, resumes — is invisible to me.
No “spend avoided” figure — uncomputable once a session is halted (the CEO, 6/10). Every toast a developer sees in the IDE lands in this feed — the manager’s roll-up view of the developer surface. Halting ≠ lost work; sessions resume.
Where it goes
✦ Manager takeaway · internalI know what the spend bought, who spent it, and whether it produced delivered work.
Spend by Developer CEO-1
Engineer
Cohort
Tokens MTD
$
Δ vs last mo
Budget used
[Dev A] ↗
UNRESTRICTED
[n]
[$]
[±%]
—
[Dev B] ↗
GUARDRAILS
[n]
[$]
[±%]
—
[Dev C] ↗
GUARDRAILS
[n]
[$]
[±%]
—
✦ Manager takeaway · internalI can compare engineers fairly: spend in the context of cohort and delivery.
Engineer links open the developer profile TM. Cohort chip per row ties the manager view to the CIO cohort framework CEO-2 — over-budget + producing ≠ over-budget + flailing.
Spend by Deliverable / PR CEO-1MS
Deliverable / PR
Engineer
Tokens
$
Value read
[PR-123 — title] ↗
[dev] ↗
[n]
[$]
delivered ✓
[PR-124 — title] ↗
[dev] ↗
[n]
[$]
in progress
[PR-125 — title] ↗
[dev] ↗
[n]
[$]
[?]
✦ Manager takeaway · internalI know what each piece of work actually cost in AI spend.
Tokens-by-PR is exactly the a prior enterprise prospect ask — same table serves both clients. PR links open the pull request in the client’s SCM; the GitHub integration already exists. WG2 "You know how much goes into each PR" is the CEO’s two-week tracking bar. CEO-1
Value Interpretation CEO-1FLAG
Spend → delivered[%]
Spend → in flight[%]
Spend → unattributed[%]
⚠ Weak takeaway · needs workWhy weak: the “value” metric is undefined — today this is deliverable-state, and an “unattributed” bucket with no definition invites distrust. Fix: rename the panel “Spend by deliverable state” until the value metric is defined via the Copilot-backend analysis — or hold the panel until then.
The metric definition is the open item: requires the light value-analysis connection to the client’s own Copilot backend. Until defined, MVP ships this simpler cut (spend mapped to deliverable state) without inventing a value score.
Not in the 5 days — deferred or scope TBDPHASE 2
CIO / platform-owner view SCOPE TBDCEO-2
Separate surface: org-wide spend rollup + cohort administration (unrestricted / guardrails / progression). The pitch differentiator — but cohort enforcement can ship config-driven without an admin UI. Decide before building screens. FLAG
System vs developer prompt split DP1PGR
Promised in the one-pager ("breakdown of tokens billed to you but not authored by your developers") and positioned to a top-5 US bank as a second-month rollout per the sync — the data is already collected. Strong waste-selling point when it lands. P2
Waste differentiation — advanced cuts WG1
The basic wasted-spend KPI moved into the MVP wireframe (one-pager promise + consistency with the health report’s AI Waste). Deferred here: waste hidden inside shipped work, efficient-vs-inefficient over-budget spend, per-repo abandonment cuts.
Bitbucket integration WG2
Not an MVP blocker (repo setup is a single command); research/plan in parallel. Without it, PR-delivery attribution uses the GitHub-pattern foundation.
Model mix / model power tracking
the CEO: "we can also track model usage… all of that can come in, but I don't want to get too long" — keep MVP focused on their stated pain.
Coverage cross-check — every pain point, where it’s answered
Source ask
Answered by
Status
Token tracking "a total mess"; homegrown script failed, not accurate DP
The dashboard itself — instrumented at IDE + agent level, not billing reconstruction
MVP
No in-IDE tracking; cost visibility only post-hoc DP
AITrax in-IDE counters (developer surface) + Team Signals roll-up
Session caps, halts in Team Signals, resumable sessions
MVP
Same rules for every engineer; no power-user path DPCEO-2
Cohort chips on Spend by Developer; cohort admin view
MVP display / admin UI TBD
No system-vs-developer prompt breakdown DP
Deferred — second-month rollout, data already collected
PHASE 2
Tokens by PR MS / spend per deliverable with value interpretation CEO-1
Spend by Deliverable / PR table + Value Interpretation
MVP
One-pager: surface wasted & discarded tokens 1PGR
Wasted Spend KPI
MVP — feasibility FLAG
One-pager: alert before budgets are breached 1PGR
Runway bands, in-IDE warnings, Team Signals
MVP
One-pager: attribute every dollar to its work 1PGR
Spend by Developer + Spend by Deliverable / PR
MVP (prompt-split portion Phase 2)
One-pager: unlock proven power users 1PGR
Cohort framework (chips now, admin TBD)
MVP / TBD
Every a team member pain point and one-pager promise lands in MVP, Phase 2, or a flagged decision — nothing is unaccounted for.
Open questions — what we're waiting onFLAG
Budget source. What defines a developer's budget — Copilot plan limits, bank-set allocations, or ours? the CEO is researching Copilot token-limit semantics. Blocks the runway and counter panels' denominators.
Value interpretation metric. Needs the light value-analysis connection to a top-5 US bank's Copilot backend (their one extra approval). Until defined, ship spend-by-deliverable-state, don't invent a score.
CIO cohort admin in MVP? Cohort enforcement presumably yes; the admin UI could be config-driven first. Decide before building screens.
Refresh cadence for the manager view ("near-real-time" needs a number).
Alert channels. In-IDE confirmed; does the manager get dashboard-only, or email/Slack too?
Baselines in week one. Month-over-month context needs a month of data; the MVP launches with days. Show day-over-day at first, or a "baseline accruing — comparisons available [date]" state. Decide so empty trends don’t look broken.
Waste KPI feasibility. Promoted into the wireframe (the one-pager promises it; the health report already measures it). a Senior Engineer leaned "total use first" — engineering to confirm the token↔abandonment correlation fits the 5 days. If it does not, this is the first fast-follow and a top-5 US bank should hear that framing on day one.
Relationship to the Engineering Health Report: this dashboard is the operational tool (manager, daily); the report's §2 Cost & Spend is the executive narrative that consumes this same data quarterly. Build once, surface twice. Per the sync, Claratev's reports align to these same a top-5 US bank requirements (Claratev as rollout guinea pig) — one dashboard serves both, with Claratev-specific asks added later.