Anonymized portfolio copy — names and customers replaced; all data illustrative.
Engineering Health Report — Redesigned Wireframe
Lo-fi layout, v4 — v3 restructured to the 6/11 objectives doc; v4 closes the gaps from the full 6/11 transcript: tooling demoted to the back (§6), a print legend under every chart, cohorts formalized to the CEO’s bands with the top-10% baseline, waste concentration & comparative wastage added. Every metric checked against the atlas data inventory. RD11Hover any colored tag to read the actual feedback behind it. Chart forms mirror the original report 1:1. TR2v4 — static print/PDF brief, designed first for the CTO; the CFO and CEO get value from the same page. Tan ✦ strips = the takeaway each audience should leave with; red ⚠ strips = weak takeaways with fixes — internal, stripped before anything ships. Every metric checked against the atlas data inventory.[Red brackets] = data/decision needed, nothing fabricated.
Feedback sources — hover any tag for the quote
SA-n Stakeholder A item
SB-n Stakeholder B item
BANK a top-5 US bank / leadership sync (6/9)
RD-n Reports Discussion — the CEO + Trent, 6/11 (RD1–14 objectives doc · RD15–26 full transcript)
Decisions & status
TR-n Trent design decision (TR12–15: 6/11 persona audit)
NEW new content
FLAG open question / data needed
Review annotations — internal only, not report content
✦ exec takeaway what an executive should leave with
⚠ weak takeaway fails the test: would a CEO care?
Masthead & orientation — page head, not a report sectionFR1·2·10
✦ Exec takeaway · internalI know whose data this is, who this brief is for, and what it covers — before reading a single number.
AUDIENCE & FORM — Designed first for the CTO; structured so the CFO and CEO get value from the same page. RD2 Delivered as a static print/PDF brief, emailed standalone — a live in-product mode comes later, and nothing in the body relies on interaction. RD1RD14 Tone rule: actionable insights, never blame — no individuals or departments are named anywhere in the report. RD4
[Sub-brand logo]
▢▢▢
Branded under the AI module sub-brand, not the main CodeTogether mark FR10[sub-brand name TBD]
Engineering Executive Brief — Q2 2026
DATA BASIS:[One client of N engineers / aggregate across N clients / demonstration data — must be stated here]FR2FLAG
One line, top of page, answers "whose data is this?" before any number is read. Also fixes the $22M "doesn't excite me" problem: savings get normalized (e.g., "per 250 engineers") once the basis is declared.Savings figures are normalized (e.g., "per 250 engineers") once the basis is declared.
The Ask AI chip was removed from the masthead TR8: the brief is static-first RD1, so interactive affordances don't belong in the page body. "Ask AI" still ships in the app shell and answers the agent-connected ask in the future live mode FZ10; the print edition references it in the footer only.
What this report covers FR1
1 Executive Overview Where you stand, what to do
2 Cost & Spend Token spend, visibility & control
3 Output & Value What engineering produces per dollar
4 AI Adoption & Cohorts Who uses AI, how well, who’s earned more
5 Quality & Delivery Is the work holding up
6 Tooling Demographics The estate — last by design RD15Tools & models inventory
Fixes "jumps straight into data, no framework." These six cards mirror the six section headers exactly, so the table of contents is the structure.
Time conventions: snapshots are Q2 2026; trends are 9-month trailing (Oct–Jun). TR15
THE THREE QUESTIONS THIS BRIEF ANSWERSRD3 — How well am I spending money on AI? (§2) · How much is AI accelerating my work? (§3–4) · Where is spend I could optimize? (§1 cohorts & findings, §2 waste)
Print conventions RD12: every section header carries a one–two sentence narrative beside the title TR11, subsection headers are the orienting questions themselves RD17, and every chart carries a short two-line legend — print readers get no other context. In the live mode every number additionally carries provenance (source tables, quality band) from the platform's calculation dictionary.
Unit gloss RD13TR6
"Output in this brief is measured in vLOC — value lines of code, effort- and impact-weighted units. AI spend is token spend: tokens are the metered units AI vendors bill by, and every AI dollar figure is token cost [pricing basis: invoice vs blended rates — confirm]. Full methodology in the Definitions & glossary." TR15
Per 6/11: definitions and glossary live at the end, not inline at the top. Supersedes the placement half of TR3 (full methodology panel in the masthead) — the one-line gloss keeps the unit readable before §1 quotes vLOC figures; the credibility story (standard weighting, governance, client-tunable) moves to the glossary with FR3/FZ3 intact.
1Executive OverviewWe are asking you to approve five prioritized moves — added engineering capacity plus $1.8M in recovered waste. The evidence behind them: net engineering cost fell $373K this quarter while output rose 6%, and the §2 cost bridge attributes the gain to AI replacing costlier effort. The gap between our best AI users and everyone else — measured in value, not volume — is where the opportunity lives. RD12TR11TR12FR7·8SB-1,2BANKTR1
✦ Exec takeaway · internalIn 60 seconds: five prioritized moves — modeled capacity plus hard waste recovery — are waiting for approval; the first proof is in: costs fell while output rose, and the top cohort shows what the rest could return.
Where do we stand? FZ2TR5RD17
✦ Exec takeaway · internalI know exactly where we stand against every target — and cost is under control.
Net Engineering Cost · QoQ TR15
−$373K
AI spend rose, labor cost fell more (§2 bridge) · AI spend [$12.2M Q2 — state avg basis]
Target: [$ budget] · on track
AI as % of Eng Cost RD5
[X%]
of total engineering cost (payroll-anchored) · [Δ] QoQ
Ceiling: [X%]FLAG
Output · QoQ
+6%
3.2M vLOC/mo (Q2 average)
Target: [+X%/qtr]
Efficient Agentic Usage RD6
[X%]
of engineers clear the bar: ≥50% AI-driven share + high vLOC/$ + high output relative to cost — ≥50% is the proposed anchor RD23; 41% daily usage is only the volume input
Target: [X%] by [date]
AI Waste · QoQ
12%
of AI token spend on abandoned/discarded work · improving −1% QoQ
Target: ≤10% — on path by [date]
Every KPI carries a goal target (the AI Waste card treatment, applied everywhere). FZ2 Cost leads the row — the narrative is "control your AI token costs". BANK Two cards new in v3: AI as % of engineering cost answers the "AI ballooning relative to payroll" fear in one number RD5 (data: fact_cost_allocation + AI session cost — both captured); Efficient Agentic Usage replaces raw Daily AI Usage as the headline quality bar — usage alone is not the goal RD6TR9 (raw adoption detail stays in §4; composite formula is an open call). Regressions (+2%, the one metric that worsened) moves out of the first-10-seconds view into §5 with full context. FR6 Not deleted — relocated and explained.
What is the data telling us? TR1TR5
✦ Exec takeaway · internalAI work returns 2×+ the value per dollar, most output is still on the costliest path, and our fastest cohort is small — that is the case for the asks below.
Return on AI Spend
6.3 vLOC/$
agentic development
2.2× the return of traditional development (2.9 vLOC/$) · /$ basis [define: token vs loaded cost]
Savings Headroom TR15
62%
of output still runs on the costliest, non-AI workflow
non-AI runs $0.34/vLOC vs $0.16 agentic — the share is the headroom, not a savings %
Speed Upside
29%
faster delivery from power users
only [14% — cohort base unresolved: §4 funnel says 1.4%] of the team is in the cohort today FLAG
Was three paragraphs of narrative; now three stat callouts, cost-led. BANK Findings stay findings (the "why") — deliberately not a restatement of the recommendations (the "what"). TR1
Where does the improvement opportunity live? RD7RD19NEWTR5
✦ Exec takeaway · internalMy top 10% is the baseline: the gap from every other band to that bar, priced, is the improvement opportunity — and I can see exactly where waste concentrates.
Cohort
Engineers
Raw output (vLOC/mo)
Value output (vLOC/$)
Spend on low-value work RD20
Gap to top-10% baseline, priced
Top 10%
[n]
[X]
[X.X]
[$X]
— (baseline)
Top 11–25%
[n]
[X]
[X.X]
[$X]
[$X/yr]
Middle 50%
[n]
[X]
[X.X]
[$X]
[$X/yr]
Bottom 25%
[n]
[X]
[X.X]
[$X · X% of all low-value spend]
[$X/yr]
Legend:bands per the 6/11 formalization — bands: top 10 / top 25 / middle / bottom 25, with the bottom 10% called out inside the bottom band where waste concentrates. RD19 Raw vs value sit side by side — a high-volume, low-value band is a cost problem, not a productivity win. Baseline = the top-10% band, not the middle: “use my power users as the baseline to see opportunity across all other areas.”: power users set the bar for opportunity across all other bands.
Cohorts are aggregate bands only — no individuals or departments named. RD4 The low-value-spend column makes waste concentration visible — the CEO’s test case: “$10M token spend, $5M on low-value work in the bottom 25%” is the trigger for quotas or restricted access, priced in §4 follow-through. RD20RD9 Banding basis open: company-wide value output vs per-project — some projects inherently carry lower business value. RD22FLAG Data: per-engineer session/commit facts + Change AI value scores roll up to bands; the platform’s granularity policy serves anonymized aggregates by design.
Highest bang-for-buck pattern RD24NEW
✦ Exec takeaway · internalComparing raw output against value output tells me which workflow earns the best return — here in the summary, where I’ll actually see it.
[Workflow pattern with the best vLOC/$ — e.g. “supervised agentic + tool X”] · [supporting numbers]
Legend: the single best-returning workflow from the raw-vs-value comparison; the full pattern list is §3 Emerging Trends. Kept non-technical — workflows, not models, per the CEO’s own correction RD24.
What should we do? FZ1FR7·8TR1TR5
✦ Exec takeaway · internalFive ranked moves — I know what I am being asked to approve, why, and which dollars are bankable vs reinvested.
Five priorities ☨ · [$18.7–20.7M modeled capacity — frozen pending cohort-base fix] (reinvested, not banked) + $1.8M waste recovery (cash-equivalent) · ≈ [$X per 250 engineers]FZ1FR7TR13
#
Priority
Action
$ impact
Lever (one line)
1
HIGH
Expand power-user cohort [14% → 25% — base unresolved]
[$14.7M]
Unlock proven engineers via cohort token controls (§2) BANK
2
HIGH
Close adopted → daily-active gap
[$4–6M]
Prompt templates, in-IDE coaching, internal champions
How we rank ☨ Priority = modeled $ impact × confidence × strategic fit. Items without a dollar figure are ranked on enabling value and labeled as such. Modeled capacity = engineering time reinvested, not banked savings; waste recovery = cash-equivalent — the two are funded differently. TR13Fixes "why is #4 high but #3 medium?" — highs now sort above mediums, and unquantified items say why FR8
Full rationale and modeled-impact math live in Recommendations in detail at the end of the report — summary up front for the 60-second read, the full case after the evidence. Keeps Stakeholder B’s recs-after-snapshot placement while fixing text density. FZ1TR1
2Cost and SpendAI spend is [X%] of total engineering cost and every attributed dollar traces to a team and a work product. What spend changed, why, what it bought, and where control was applied this quarter. TR11NEWBANKFR4SB-4,7
✦ Exec takeaway · internalAI spend is visible, attributed to work products, and under enforced control — and every cost change has an explanation.
Is AI spend under control? NEWBANKTR5
The company's top priority gets the report's center of gravity. Exec framing: AI costs can spiral ($100M → $500M scenario); leadership needs visibility + enforcement, not just reporting. All figures below are placeholders — populate from MVP dashboards, never invent.
Spend Attribution Coverage Q2 2026TR7
✦ Exec takeaway · internal[X%] of AI dollars have a named owner and work product — attribution, not invoice reconstruction; the remainder is bucketed by reason.
[X%]
of Q2 AI spend tied to a specific team and deliverable
Team A[$X]
Team B[$X]
Team C[$X]
Legend: bars = attributed Q2 AI spend by team. The unattributed remainder is preserved and bucketed by reason — never silently assigned.
Was "Live Token Spend · real-time" — reframed as a quarterly outcome per the CEO’s altitude flag (resolved open call) and the static-brief decision. RD1 The live view stays on the Token Spend Dashboard.
Budget Cliff Outcomes Q2 2026TR7
✦ Exec takeaway · internalNo engineer fell off the budget cliff unwarned this quarter.
[N]
engineers hit a hard limit without prior warning · target 0
Quarter-end position — engineers by % of budget consumed:
<50%[n]
50–75%[n]
75–90% ⚠[n]
>90% ⚠⚠[n]
Legend: distribution at quarter close; ⚠ bands received early alerts before limits hit. Alerts fire before the cliff — no one bad prompt silently burns 10% of a budget.
Was a live "Budget Runway" view — reframed as the cliff-problem outcome RD1TR7; live runway bands stay on the dashboard.
Runaway Agent Outcomes Q2 2026TR7
✦ Exec takeaway · internalNo runaway session reached budget-impacting size this quarter — and halted work resumed without loss.
[N]
sessions halted before reaching budget-impacting size · [X%] resumed without lost work · [X tokens] consumed up to the halts
Sessions are resumable — no work lost. Stops the "agent dumps a million log lines into a model" failure mode. No "spend avoided" claim — uncomputable: once halted, the counterfactual burn is unknowable (the CEO, 6/10). Show halts, tokens consumed up to the halt, and resumes.
Cohort-Based Controls
✦ Exec takeaway · internalProven engineers run free; everyone else gets guardrails and a path up.
Unrestricted[n] proven power users
Guardrails[n]
In progression[n] on path to unrestricted
Not just cost control — a framework for building an effective AI-using org. Directly powers Recommendation #1 (expand the power-user cohort), and gives the cohort/authorization framing leadership asked for: authorize power users, guardrail the rest, with a visible path up. Bottom-cohort follow-through (quotas, restricted access) with dollar impact lives in §4. RD9
Wasted & Discarded Token Spend
✦ Exec takeaway · internalWaste has a number — even when it hides inside shipped work.
[$X]
tokens spent on abandoned/discarded work — including waste hidden inside shipped work
Wastage rate, bottom cohort vs effective users: [X×]RD21NEW
Legend: waste = token cost of later-abandoned work (session cost × the share of that work later abandoned). The multiple compares inefficient users’ wastage against the effective-user cohort — the contrast asked for on 6/11,, not an absolute number alone. RD16
Distinguishes efficient over-budget spend (producing value) from inefficient AI use. Reconciles with the 12% AI Waste KPI and the cohort low-value-spend column in §1.
Token Cost per Work Product Q2 2026BANKNEW
✦ Exec takeaway · internalI can see what a PR costs in tokens — spend is tied to what was actually produced, and to its value.
[$X]
median token cost per merged PR · the most expensive 10% of PRs start at [$Y] · per work item [$Z]
<$10/PR[n]
$10–50[n]
>$50 ⚑[n]
Legend: distribution of merged PRs by token cost; ⚑ tail = review candidates. Pair with §3 value-per-dollar to separate expensive-but-valuable from expensive-and-wasted.
Requested by a top-5 US bank and a prior enterprise prospect. B2 Data: fact_work_item (native branches + GitHub PRs) joined to session cost — "cost per outcome" is a measured spendAllocation metric in the inventory. RD11 Replaces the System-vs-Developer Prompt Split panel, dropped from the brief TR7 — prompt taxonomy stays Phase 2 on the Token Spend Dashboard.
Why did costs change, and what do they buy? Q2 2026FR4TR4TR5
Cost bridge callout FZ4
✦ Exec takeaway · internalNet cost fell because AI replaced costlier human effort — the trade is working.
"AI costs rose [$X], human labor costs fell [$Y], for a net QoQ gain of $373K."
Ready-made language for the CTO to take upstairs. Footnote must state whether −$373K includes AI tool spend and defect-remediation costs. FLAG
Where the Money Goes Q2 2026NOT CAPTURABLE TODAYTR15
✦ Exec takeaway · internalMost spend buys new features, not upkeep — once work-type classification is enabled.
New feature[52% · $19.0M · +3% QoQ]tgt [X%]
Maintenance[32% · $11.7M · −1% QoQ]tgt [X%]
Tech debt[16% · $5.9M · flat QoQ]tgt [X%]
Legend: attributed Q2 engineering spend by work type, target beside each — same period as the cost bridge. Values bracketed: work-type classification is not enabled today (see open calls). RD16
Target allocations added per category. FZ2 Paired with the cost bridge — same Q2 period, the numbers foot. FR4
RD11FLAG Data-feasibility check: work-type cost allocation exists structurally but every minute is 'unclassified' today (inventory §6). Change AI already classifies work types per commit — that allocation must be enabled before this chart is real. Open call.
Where are costs heading? 9-MO TRENDSTR5
Engineering Cost Composition — AI vs human capital 9-MO TRENDRD5
✦ Exec takeaway · internalAI’s share of engineering cost is growing on a steady, bounded path relative to payroll — substitution, not a spike.
Human vs AI spend over 9 months — trend context for the Q2 snapshot above; trends may run longer than cost snapshots FR4
AI as % of total engineering cost:[X%] → [Y%] over the trend · ceiling target [X%]FLAG
Legend: shaded area = AI spend; line = human engineering cost (loaded-cost policy). Read the gap: leadership’s fear is AI costs ballooning relative to payroll — this chart answers it directly.
Feeds the "AI as % of Eng Cost" KPI in §1 — same data, same period. Data: fact_cost_allocation (human cost from cost policy) + AI session cost, both captured. RD11
Cost per Value Unit 9-MO TREND
✦ Exec takeaway · internalA unit of value keeps getting cheaper — and the one exception is explained.
Dec spike: [reason — e.g. year-end slowdown] ⚑
$0.30/vLOC · −6.5% QoQ · target line at [$X]
Denominator [define: token cost only vs loaded engineering cost — must reconcile with the $12.2M AI-spend and $36.6M attributed totals]FLAG
Legend: blended cost of one value unit, trailing 9 months; every notable deviation annotated inline. RD16
Anomaly annotation inline on the chart. FZ7 Standing practice: every notable deviation gets a note — the report must be self-contained.
Cost & Spend leads as §2: token cost is the company's top priority and §1 already opens cost-first, so the story runs unbroken — position → cost control → the value it buys. Resolves open call #7. TR3BANK
3Output and ValueEngineering produces 3.2M vLOC a month (Q2 average), and half of delivered value — value-weighted; see the base note below — is already AI-touched. The economics run one way: the more agentic the work, the more value per dollar. TR11FER ×2SB-3
✦ Exec takeaway · internalThe unit economics favor AI: the more agentic the work, the more value per dollar.
How much work is produced? TR5RD17
Work Output 9-MO TREND
⚠ Weak takeaway · needs workWhy weak: growth alone has no so-what — every vendor chart goes up and to the right. Fix: pair with cost — “output rose 6% while net cost fell: more product for less money”.
3.2M vLOC/mo · +6% QoQ
Legend: total value-weighted output per month, trailing 9 months. RD16
Chart name unchanged — the duplication was fixed at the header layer: the subsection header is now the question itself, and the old "Work output" noun subhead is gone. RD17
Value-Weighted Output by Source Q2 2026RD18
✦ Exec takeaway · internalHalf of delivered value is already AI-touched.
Legend: value-weighted output split by author — human / AI-assisted / agentic — with absolute vLOC beside each share, not percentages alone. RD16RD18
Base note: shares here are value-weighted — they will not match raw-volume shares elsewhere (the 41% raw AI share in the trend, or the work-model split below). [one base must be declared and badged per chart]FLAGTR15
Absolutes bracketed until the value-weighted base is confirmed — % of value output ≠ % of the 3.2M raw figure; the two must not be conflated.
Is AI’s contribution growing — and sticking? TR5
AI-Attributed Output 9-MO TREND
✦ Exec takeaway · internalAI’s share of output nearly doubled in nine months.
0.7M → 1.3M vLOC/mo attributed to AI
Legend: monthly vLOC attributed to AI authorship, trailing 9 months. RD16
AI Work Product Retention 9-MO TREND
✦ Exec takeaway · internalAI-written work sticks — it is not throwaway code.
Retention rate (top) vs abandonment rate (bottom)
Legend: top line = AI-written code still in the codebase after 5 commits; bottom line = deleted within 5 (abandonment). RD16
Feeds the abandonment figure in §5 Quality Signals.
How impactful is recent work? TR5
Value Contribution by Work Model Q2 2026 ★
✦ Exec takeaway · internalThe more agentic the work, the more value per dollar.
Agentic6.3 vLOC/$ · $0.16 · 6% of output
AI-assisted4.2 vLOC/$ · $0.24 · 32%
Non-AI2.9 vLOC/$ · $0.34 · 62%
Legend: value per dollar by work model, with cost per vLOC and share of output beside each bar. Base: [raw vs value-weighted — must be badged; 6% agentic here vs 18% in the by-source split needs one declared base]FLAGRD16
Kept as the key metric — unchanged, both reviewers liked it.
Value by Platform NEWFR9
✦ Exec takeaway · internalWhich tools earn their seat — the standardization answer.
Copilot[vLOC/$]
Claude Code[vLOC/$]
Codex[vLOC/$]
Legend: value per dollar by AI platform — the standardization question answered in value terms. RD16
Answers "is Claude better than GitHub? Should we standardize on 1–2 tools?" Stakeholder A says we have this data. Feeds training/licensing decisions.
What’s emerging that deserves a CTO’s attention? RD8NEW
✦ Exec takeaway · internalI can see which tools and workflow patterns are earning the best return before the rest of the org gets there.
1.[Trend — e.g. tool/model pattern whose vLOC/$ is rising fastest QoQ] · [supporting numbers]
2.[Trend — e.g. workflow pattern (supervised vs autonomous mix) with best retention] · [supporting numbers]
3.[Trend — e.g. repo/setup pattern correlated with low abandonment] · [supporting numbers]
Legend: the 2–3 patterns with the best measured return this quarter, each with the numbers behind it. Computed, never anecdotal.
Data: per-tool efficiency, model mix, retention and value trends — all measured domains. RD11 The supervised/autonomous split is contracted but not yet in marts (inventory gap) — trend #2 stays bracketed until it lands. FLAG
4AI Adoption & CohortsA small proven cohort already works at the efficient-agentic bar; the bottleneck is adoption depth, not technology. Who uses AI and how well, who has earned more autonomy, where guardrails should tighten. TR11SB-5,9RD15TR10
✦ Exec takeaway · internalAdoption, not technology, is the bottleneck — a small proven cohort shows the upside of expanding.
Who is using AI, and how deeply? TR5
AI Adoption Funnel — anchored to total headcount FZ5
⚠ Weak takeaway · needs workWhy weak: the story is right but the number cannot be trusted — 41% (snapshot) vs 4.1% (funnel) is unresolved. Fix: keep the panel; reconcile the data first. Then the takeaway lands: “93% of the org has not adopted AI — the $14.7M opportunity is the gap”.
Org headcountEver used AIDaily activePower users
~1,000 · 100%66 · 6.6% of org41 · 4.1% of org (62% of adopters)14 · 1.4% of org (34% of daily)
Legend: each step as a share of total org headcount — ever used → daily active → power user. RD16
FLAG Snapshot says "41% daily" but the funnel implies 4.1% — same digits, different base. Must reconcile before this ships.
Adoption Trend 9-MO TREND + Cohort Delta insight
✦ Exec takeaway · internalAdoption climbs steadily; the power-user tier lags the most.
Legend: top line = daily-active share, bottom line = power-user share, trailing 9 months. RD16
Kept as-is. Insight box stays:Insight box: power users ship at 1.7d vs 2.4d — 29% delta.
Is the org ready for more? TR5
AI Readiness FZ9
✦ Exec takeaway · internalEnvironment readiness is the binding constraint on everything above.
READINESS ← [Q1: X%] two-quarter trend
Engineer Interactions83%
Session Steering82%
Prompt Specificity58%
Environment Readiness ⚑37%
Legend: composite readiness score and its four dimensions; ⚑ = binding constraint. RD16
One-line definition under each dimension [definitions TBD]. Callout: Environment Readiness (37%) is the foundational prerequisite — it multiplies Recommendations 1 & 2. Target 80+ by end of Q3. Quality band: proxy/estimated per the data inventory — badge it so the number isn't over-trusted. RD11
Who has earned more autonomy — and where should guardrails tighten? RD10RD9NEW
Recognition of excellence CONCEPT · CRITERIA TBDRD10
✦ Exec takeaway · internalExcellence has a bar — engineers and teams who clear it get recognized and unrestricted, and I can see hidden pockets of it.
[n] engineers and [n] teams above the certification thresholds [criteria TBD] this quarter
Hidden gems:[n] high-performing clusters outside the recognized cohort — surfaced as patterns, never by department or nameRD4
Legend: certification = sustained efficient-agentic usage (the §1 bar) over [period]. Clearing it feeds the unrestricted tier in §2 cohort controls — and can back recognition/bonus programs and knowledge-sharing from power teams ("this project team has a lot of power users — let’s share that knowledge"). RD10
Thresholds/certification are a concept, not yet platform machinery — inputs (value output, retention, efficiency per engineer) are all captured; the bar itself is an open call. FLAG
Bottom-cohort follow-through RD9
✦ Exec takeaway · internalWhere waste concentrates, the report names the action and its dollar impact — never the people.
Action
Applies to
Modeled $ impact
Usage quotas
bottom band · [n] engineers
[$X/yr]
Restricted model/tool access
highest-waste patterns
[$X/yr]
Coaching + progression path
guardrailed tier (§2)
enabler
Legend: concrete controls available where waste concentrates, each priced from measured waste cost (cost × abandoned share). Actions target patterns, not people. RD4
Pairs with §1 cohort analysis (the evidence) and §2 cohort controls (the mechanism); priced into recs #4–5. Note: per-cohort budgets need Work-Area-scoped cost policies — tenant-global today (inventory gap). FLAG
5Quality & DeliveryQuality is holding while AI’s share of the work grows: rework and abandonment are both improving, with the remediation dollars attached. Delivery is getting faster, not riskier. TR11SB-2,6FR6
✦ Exec takeaway · internalQuality and speed are holding while AI scales — every quality metric we can measure today is improving, and what can’t be measured yet is labeled, never asserted.
Is the work holding up — and are we still shipping fast? TR5
Quality Signals — with dollar impact ☨ FZ6
✦ Exec takeaway · internalEvery quality metric we can measure today is improving — each with a price tag; gaps are labeled.
Metric
Current
Δ QoQ
$ Impact
Status vs target
Rework rate
12%
−1%
[$X remediation]
improving, on path by [date]
Abandonment rate
14%
−1%
$1.8M recoverable (rec #4)
improving, on path by [date]
Review-risk exposure NEW
[n] high-risk items
[Δ]
—
measured (Change AI review-risk score)
Regression rate NOT CAPTURABLE TODAY
[5% — unverifiable]FLAGRD11
[Δ]
[$X remediation]
blocked — the platform doesn’t yet ingest build/test or bug-tracker data
Distinct measures: rework here (12% of merged code reverted) is not the §1 AI Waste KPI (12% of token spend) — the matching values are a coincidence of this period; abandonment (14%) is the deletion-rate input to waste. TR15
Regressions live here with context, not in the snapshot. FR6Footnote: does −$373K net cost already account for defect remediation? FLAG old snapshot said regressions = 8% of defects while this table says 5% — two different measures or an error; reconcile before shipping.
RD11 Data-feasibility check against the inventory: defect/CI provider facts are a listed gap — the platform cannot measure regression rate today, so the row is bracketed red, not asserted. Until that ingestion lands, the measured proxies are review-risk score and test confidence (Change AI), added above. Open call: drop, proxy, or wait.
Data basis ☨ Rework = reverted code / total merged (file-retention telemetry). Abandonment = AI-generated lines deleted within 5 commits. Review risk = Change AI semantic score per commit. Regressions = customer-flagged bugs traced to the originating change — no capture pipeline exists for this today. FLAGThe data-basis line answers "what is this data?" exactly the way our trust principle requires.
Delivery Speed
⚠ Weak takeaway · needs workWhy weak: faster shipping is unanchored without a cause or consequence. Fix: tie to the AI story — “cycle time fell 0.3d as AI’s share of work grew: AI is making delivery faster, not riskier”.
2.1d
cycle time · −0.3d QoQ
Target: [X.Xd]FZ2
Team Utilization
✦ Exec takeaway · internalThe team is properly loaded; a small tail needs attention.
ON TARGETTR15 12–20 hrs/wk active
Target: [X%]FZ2
Legend: share of engineers inside the configured 12–20 active hrs/week band; “active” = telemetry-measured hands-on engineering time, not total hours worked [exact definition TBD]. RD16
On-target 74% · Exceeding 15% · Below 11% — the donut shows on-target only (was 89% “hit target”, which silently counted the exceeding band — persona audit). TR15
6Tooling DemographicsThe tool and model inventory behind the numbers above. Deliberately last RD15: demographics describe the estate; the decisions are driven by value-per-platform (§3) and the cohorts (§1). TR11RD15TR10FR5FZ8
✦ Exec takeaway · internalReference inventory of the AI estate — read after the decisions, not before them.
What is the org running on? TR5
Platform Mix % OF DEVELOPERS USING TOOL · Q2
✦ Exec takeaway · internalCopilot is ubiquitous; the field behind it is fragmented.
Copilot71%
Claude Code38%
Codex24%
JetBrains AI17%
Internal / other9%
Legend: % of developers using each platform in Q2 — multi-select, so bars sum past 100%. RD16
Vendor Mix % OF USAGE VOLUME · Q2TR14
✦ Exec takeaway · internalThree-quarters of our AI usage flows through a single vendor — negotiating leverage and concentration risk in one number.
Anthropic · 74% · OpenAI · 16% · Other · 10%
Model detail: Sonnet 4.6 · 42% · Opus 4.7 · 18% · Haiku 4.5 · 14% (Anthropic) · GPT-4o-mini · 16% (OpenAI) · Other · 10%
Legend: share of Q2 usage volume by model vendor — concentration and dependency, not model taxonomy; model-level detail in the line above. RD16
Was “Model Mix” (model-name taxonomy) — regrouped by vendor per the panel’s own weak-takeaway fix, confirmed by the persona audit (CEO: “model taxonomy I can’t act on”). TR14
Bridge note under both charts:"Platforms and model vendors are different layers — one platform can serve several vendors' models, so these two charts measure different things and won't match." FR5 Distinct scope badges (developers vs usage volume) do the heavy lifting. FLAG confirm it isn't actually a mixed dataset.
Editor / IDE Breakdown → action FZ8
⚠ Weak takeaway · needs workWhy weak: tool choice is a manager decision; no CEO action follows from an editor chart. Fix (updated 6/11): demoted to back-of-report demographics with the rest of tooling RD15 —Fix: lives in back-of-report demographics — Value by Platform (§3) answers the standardization question at exec level; the conversion insight below is what earns this chart its place.
VS Code 50%
Cursor 30%
JetBrains 18%
Legend: % of developers by primary editor/IDE; the conversion insight below is the action layer. RD16
Insight strip added:Insight strip: power-user conversion rate and cycle time by editor — [e.g. "Cursor users convert to power users at X× the rate"] — surfaced as a standardization/training insight, not raw inventory.
Tooling moved out of §4 into this back-of-report section per the 6/11 transcript — "tooling to me is kind of like the last thing". RD15TR10 Value by Platform stays in §3: it answers a value question (FR9), not a demographic one.
Recommendations in detail — the full case, after the evidenceFZ1TR1
✦ Exec takeaway · internalEvery dollar figure in §1’s summary table must show its math here — bracketed figures stay frozen until the math exists; nothing is asserted without a basis.
The §1 table is the 60-second read; this section is for the reader who has seen the data and wants the argument. Same five items, same order — the two must stay in sync when numbers change.
HIGH1 · Expand power-user cohort [14% → 25% — base unresolved: §4 funnel says the cohort is 1.4% of org] · [$14.7M] annualized modeled capacity gain — frozen until the cohort base is resolved, then re-derived as population × Δ vLOC/$ × cost base, with the math shown here. Value Contribution shows agentic work at 6.3 vLOC/$ vs 2.9 for non-AI-assisted. Lever: convert [110] additional high-engagement AI-assisted users into the power-user cohort — enabled by cohort-based token controls in §2 (authorize proven engineers to run unrestricted). BANK
HIGH2 · Close the adopted → daily-active gap · [$4–6M] annualized depending on conversion rate — conversion model to be shown here; [25-point gap — §4 funnel shows 6.6% → 4.1%, a 2.5-point gap] today. Lever: prompt-template rollout, in-IDE coaching, internal champions in the highest-headcount underutilized departments.
HIGH3 · Raise AI readiness 65% → 80% · Enabler — multiplies items 1–2; no standalone $ modeled. Driven by Environment Readiness at 37% and Prompt Specificity at 58%. Lever: standardize repo setup checks, prompt specificity rubrics, ready-state gates for agentic workflows. FZ9
MED4 · Reduce AI abandonment 14% → 10% · $1.8M annualized waste recovery. Quality signals show 14% abandonment while the AI Waste KPI sits above the ≤10% target. Lever: prompt templates + tooling guidance targeted at the highest-abandonment repos surfaced by file-retention telemetry.
MED5 · Move below-target engineers into range 11% → 7% · No $ modeled — utilization hygiene. 110 engineers remain below the configured 12–20 active hrs/week range. Lever: manager-level workflow triage and repo-specific friction review.
Definitions & glossary — at the end by designRD13TR6FR3FZ3
✦ Exec takeaway · internalDefinitions are reference material: print readers flip back when they need a term — they don’t wade through methodology before the first number.
vLOC methodology — how to read every number in this report FR3FZ3TR3TR6
✦ Exec takeaway · internalThe measurement unit is credible: standard across clients, governed on a cadence, and tunable to our priorities.
Keeps the Strategic Value + Substance & Efficiency explanation, then replaces "we developed our own weighting" with:Strategic Value + Substance & Efficiency explanation, anchored by the weighting-governance statement:
"This report uses the standard weighting model — unchanged since [date], reviewed quarterly by [owner/committee]. Every CodeTogether client starts from this same model and can adjust the weights to fit their own priorities."
Kills the black-box read in a cold-call setting and answers the CFO's "what is a unit worth, and who decides?" Moved here from the masthead per the 6/11 requirement (definitions at the end, not inline at the top) RD13; the masthead keeps a one-line gloss so the unit is readable before §1 quotes vLOC figures.
Term definitions
Term
Definition (two lines max — print readers get no other context)
vLOC
Value lines of code — effort- and impact-weighted output units; full methodology above.
Token
The metered unit AI vendors bill by; every AI dollar figure in this brief is token cost [pricing basis: invoice vs blended rates — confirm]. TR15
Active hours
Telemetry-measured hands-on engineering time per week, not total hours worked [exact definition TBD].
Efficient agentic usage
The headline quality bar: high share of AI-driven work plus high vLOC/$ plus high output — not raw usage alone. Composite formula [TBD]. FLAG
Cohort bands
Engineers grouped by value output into top/middle/bottom aggregate bands [band definition TBD]; never resolved to individuals or departments in this report.
AI waste
Token spend on abandoned or discarded work, including waste hidden inside shipped work (cost × abandoned lineage share).
Abandonment
AI-generated lines deleted within 5 commits (file-retention telemetry).
AI readiness
Composite of capture/AI/prompt/semantic evidence — proxy/estimated quality band.
Quality band
Every platform number carries provenance: source tables, measures, and a quality band (derived / estimated / partial / unavailable). Missing data shows as unavailable — never fabricated.
FooterFR10FZ10
✦ Exec takeaway · internalThis report covers AI economics — one slice of CodeTogether’s value, not the whole platform.
[Sub-brand mark] + scope line
"This module covers AI economics; CodeTogether's full platform also measures [time on keyboard, time to code, …]" — so the report isn't mistaken for the entire value prop. FR10
The footer carries the report's only interactive reference: "Interactive version with Ask AI available in the app" — the brief itself is static-first. RD1FZ10TR8
Not in this brief / deferred RD25RD26NEW
• In-app Reports section — a reports button (QuickBooks-style, separate from dashboards), in-UI preview, PDF generation, Ask AI interaction, and a toggle to hide legends/descriptions once familiar. Build the report first; the reports UI manages the print↔live transition later. RD25
• Live dashboard mode of this report — future phase. RD1
• Insight cadence: report insights are computed month-over-month, not this-week — the backend logic differs from the dashboards. RD26
Process note (6/11): the report narrative is regenerated fresh each cycle — unbiased by the current draft — and iteratively reviewed from independent perspectives before finalizing.
Open calls & conflicts — decide before next sendFLAG
Data basis. One client vs aggregate vs demo changes the masthead line, the savings framing, and #2 below. Blocking decision.
Negative metrics vs trust (Stakeholder A ↔ report principles). Stakeholder A wants no worsened metrics in prospect settings; our trust principle says don't hide data. Proposed resolution: if demo data → curate freely and label it demo; if real client data → keep negatives but always annotated with cause + path to target. Consider a labeled prospect variant of the report rather than one doc doing both jobs.
Funnel base discrepancy — widened by the 6/11 persona audit. "41% daily" (snapshot) vs 41 engineers ≈ 4.1% of 1,000 (funnel); the same 10× error appears as "14% of the team" (§1 / Rec #1) vs "14 · 1.4% of org" (funnel), and Rec #2’s "25-point gap" vs the funnel’s 2.5 points. ~$19M of the recommendations ask is priced on this base — $14.7M and $4–6M are bracketed/frozen until one org base is declared and every dependent figure re-derived. FLAG
Regression rate. 8% (old snapshot) vs 5% (quality table). Reconcile or label as different measures.
−$373K definition. Does it include AI tool spend? Defect remediation? Needed for the cost bridge (FZ4) and quality $ column (FZ6).
Sub-brand name for masthead (AITrax is the agent monitor name in the a top-5 US bank one-pager — same brand or separate?).
§2 token panels — altitude re-scope — RESOLVED (Trent, 6/11). All three live/capability panels reframed as quarterly outcomes (Spend Attribution Coverage, Budget Cliff Outcomes, Runaway Agent Outcomes) TR7; Prompt Split dropped from the brief, stays Phase 2 on the Token Spend Dashboard. Confirmed by the 6/11 static-brief objective. RD1
Section order §2 vs §3 — RESOLVED. Cost & Spend leads as §2 (Trent, June 10). 6/11 amendment: full vLOC methodology moved from masthead to the back glossary TR6; masthead keeps a one-line gloss.
Cohort band definition. Bands formalized to the CEO’s 6/11 proposal — top 10 / top 25 / middle / bottom 25, bottom 10 called out within. RD19 Still open: exact cut points, and the banding basis — company-wide value output vs per-project, since some projects inherently carry lower business value. RD22 Blocks the §1 cohort table and §4 follow-through. RD7
Efficient-agentic-usage formula. the CEO proposed ≥50% AI-driven share as the anchor RD23; vLOC/$ and output are measured — the full composite, the output-relative-to-cost denominator, and thresholds remain undefined. RD6
Waste-by-cohort attribution. The §1 low-value-spend column and §2 wastage multiple need per-band waste cost (session cost × abandoned lineage share, rolled to bands). Inputs are captured; confirm the band rollup exists before populating. RD20RD21
Certification thresholds (RD10). Recognition-of-excellence criteria are a concept only — not yet platform machinery; define the bar and the period.
Regression/defect data gap. No CI or issue-tracker ingestion exists (inventory gap) — the §5 regression row is unverifiable. Decide: drop, run on the review-risk/test-confidence proxy, or wait for provider facts. RD11
Work-type cost allocation. "Where the Money Goes" requires cost-by-work-type, but all minutes are 'unclassified' today; Change AI per-commit work types are the path. Needs enabling before the chart is real. RD11
AI-as-%-of-eng-cost ceiling. The §1 KPI and §2 trend need a target ceiling [X%] — CFO input. RD5
vLOC/$ denominator & cross-layer reconciliation (persona audit, 6/11). $0.30/vLOC × 3.2M vLOC/mo ≈ $0.96M/mo, but stated AI spend is $12.2M (Q2) and attributed spend $36.6M — ~12× apart. Define the /$ denominator (token cost only vs loaded engineering cost), badge it on every vLOC/$ chart, and demonstrate one multiplication that closes. Until then the “2.2×” claim carries a basis bracket. FLAG
Output-split base (persona audit, 6/11). By-source: Human 50 / AI-assisted 32 / Agentic 18; by-work-model: Non-AI 62 / AI-assisted 32 / Agentic 6; the trend implies 41% raw AI share vs the §3 header’s “half of delivered value”. Likely raw vs value-weighted — declare one base, badge exceptions per chart, reconcile all three in one footnote. FLAG
Waste / rework / abandonment relationship (persona audit, 6/11). Three near-identical percentages (12% / 12% / 14%) now carry a distinct-measures note in §5 — confirm the definitions hold and the numeric coincidence survives real data.
Coverage: Stakeholder A FR1–FR10 · Stakeholder B FZ1–FZ10 · a top-5 US bank B1–B2 · Reports Discussion 6/11 RD1–RD26 (RD1–14 objectives doc, RD15–26 full transcript) — all tagged inline; hover any tag for the source feedback. Every metric checked against the atlas data inventory; gaps flagged red, never papered over. See Feedback-Traceability-Map.html for the side-by-side original→redesign mapping.