Executive briefing · critique-first

The Agentic Paradigm: what to do, and what not to do

A summary of The Agentic Paradigm Briefing: A Verification, Critique, and Extension — prepared for a senior Indian employment and data-protection adviser to GCCs and MNCs — with each finding expanded against the underlying evidence, and six open questions from the delivery floor treated as topics rather than answers.

As of 31 July 2026 16 workflow diagrams Part I — nine findings Part II — six field topics Positions are current-dated · not legal advice

Part I

The nine findings, and what sits underneath them

Each finding is stated as the summary puts it, then expanded against what the underlying report actually establishes — including the places where the report contradicts the popular reading. The report's own verdict is that the document it reviewed was factually accurate and strategically wrong. That distinction is the spine of everything below.

01

AI models are becoming dramatically cheaper

Unit prices are collapsing. Enterprise bills are not — cheap tokens are what make expensive agentic loops affordable enough to run badly.

  • Model costs are dropping every few months.
  • Businesses will increasingly replace expensive human effort with AI where possible.
  • The important metric is cost per completed business task, not cost per token.

Takeaway: Stop asking "How much do tokens cost?" Ask "How much does it cost to finish this business process?"

From cost per token to cost per completed task A left-to-right chain showing that the price of a million tokens is the wrong denominator. Between the token price and a finished business outcome sit tokens consumed, retries and reruns, tool calls and retrieval, human review time, and rework and defect fixing. Below, three statistics explain why the cheap-token story misleads: prices fall about ten times a year for a fixed capability level, frontier-level capability costs about eighteen times more each year, and although the average price per million tokens fell from about ten dollars to about two dollars fifty, total enterprise bills rose. From cost per token to cost per completed task The token price is the wrong denominator. Everything between a token and a finished outcome also costs money. WHAT ACTUALLY SITS BETWEEN A TOKEN AND A BUSINESS OUTCOME WRONG Cost per 1M tokens Tokens consumed Retries & reruns Tool calls & retrieval Human review time Rework & defect fixes CORRECT Cost per completed business task The only denominator a CFO can reconcile to output. Every box in the chain consumes budget. Only the last one creates business value. Why the cheap-token story misleads UNIT PRICE, FIXED CAPABILITY ~10×/year cheaper for a FIXED capability level (a16z, LLMflation) FRONTIER CAPABILITY, SAME WORK ~18×/year MORE expensive to run FRONTIER-level capability (Epoch AI) PRICE DOWN, BILL UP $10 → $2.50 per 1M tokens unit price fell, total enterprise bills ROSE (Ramp data via Artefact) Jevons paradox: cheap tokens fund expensive agentic loops.
Figure 1 From cost per token to cost per completed task — falling unit prices say nothing about what a finished piece of work actually costs.
Read more — what the report actually says

What the report actually says

The price collapse is real, and the report confirms it against primary sources rather than press summaries. The 30 July 2026 GPT-5.6 announcement cut Luna by 80% and Terra by 20%, left Sol unchanged, and replaced Priority Processing with a "Fast mode" delivering up to 2.5× the speed of Standard at twice the price (OpenAI; mirrored in the AWS Bedrock pricing update).

TierBeforeAfter (30 Jul 2026)Change
GPT-5.6 Luna$1 / $6$0.20 / $1.20−80%
GPT-5.6 Terra$2.50 / $15$2 / $12−20%
GPT-5.6 SolunchangedunchangedFast mode: 2.5× speed at 2× price
Meta Muse Spark 1.1$1.25 / $4.25launched 9 Jul 2026

The report's genuine praise for the source document is reserved for one move: the pivot from price-per-token to cost per completed task. It calls this "correct and ahead of most enterprises," and notes the market has already followed — Artificial Analysis now publishes weighted cost-per-task, and OpenAI frames outcome per dollar (The Register).

The evidence

On the "Agents' Last Exam" benchmark, Luna outperforms Fable 5 "at an estimated cost per task nearly 99% lower," and Sol scored 53.6, beating Fable 5 by 13.1 points (iClarified). Customer figures on the same page: Blitzy reports 2.2× more context, 8.5× fewer output tokens, 87% lower cost against GPT-5.4 mini and cache hit rates from 24% to 90%; Dust reports agentic tasks 40% faster and 40% cheaper; Notion reports Terra matching GPT-5.5 quality "at half the cost per task and in 60% less time."

The structural trend holds too. a16z's "LLMflation" finds cost falling roughly 10× per year for an LLM of equivalent performance (a16z); Epoch AI measures a median ~50× annual decline, rising to ~200×/year on post-January-2024 data.

Where it gets complicated

Three corrections the hype version leaves out. First, every customer metric above is a vendor-selected quote on the vendor's own page — best-case, benchmark-harness numbers, not independent field results. The report names this "circular corroboration" and treats a second document agreeing with the first as adding no evidentiary weight.

Second, unit prices falling does not mean bills falling. Ramp data cited by Artefact shows average cost per million tokens dropping from ~$10 to ~$2.50 in a year while consumption exploded and total spend rose (Artefact). Jevons paradox, on an enterprise P&L.

Third, and most often missed: Epoch also finds the cost of running frontier-level capability has risen roughly 18× per year (Epoch AI, arXiv). Cheap tokens fund expensive agentic loops. "Intelligence too cheap to meter" is precisely the rhetoric that produces the budget shock.

What to do on Monday

  1. Make cost per completed task the board-level metric, and instrument it alongside human-intervention rate and rework/defect rate.
  2. Route by task tier — the Sol/Terra/Luna and Muse Spark ladder makes multi-model routing the default architecture, not an optimisation.
  3. Set token budgets and consumption alarms per workflow. Falling unit price is not a spend control.
  4. Discount any vendor-published cost or quality claim until you have reproduced it on your own workload.
02

The biggest change is workflow, not the model

The scarce input is integration discipline, not model speed. Ninety-five percent of pilots prove it.

  • Winning companies won't simply buy better AI. They will redesign work so that AI performs routine work; humans handle judgment, exceptions, approvals and accountability.

Takeaway: AI changes how work is done, not just who does it.

Workflow redesign: before and after Two aligned swimlanes over the same six steps. In the before state, humans perform intake, research, drafting, manual checking, approval and delivery, so every step waits on a person. In the after state, agents execute intake, research, drafting and delivery, an automated evaluation replaces the manual check, and a risk gate asks whether money, legal exposure, rights, security or compliance are involved: yes routes to a human decision, no proceeds automatically to delivery. Workflow redesign: before and after The same six steps, twice. What changes is not the work — it is who executes and who decides. BEFORE humans do everything, approvals are serial HUMAN Intake HUMAN Research & gather HUMAN Draft HUMAN Manual check every step is gated HUMAN Approval HUMAN Deliver Every step waits on a person. Approvals are the queue. Risk gate: money / legal / rights / security / compliance? AFTER agents execute, humans decide AGENT Intake AGENT Research & gather AGENT Draft AUTO-EVAL Automated evaluation RISK GATE YES HUMAN Human decision AGENT Deliver NO — auto-proceed, logged, sampled later Humans move from doing the work to owning judgment, exceptions and accountability.
Figure 2 Workflow redesign: before and after — the same six steps, re-cast so that agents execute and humans decide.
Read more — what the report actually says

What the report actually says

Asked how to manage the operational transformation, the report is blunt: "Treat it as workflow redesign, not tool procurement." Redesign one high-value workflow end to end, instrument it, then scale. It agrees with the source document that transformation is workflow-level, and disagrees on what is scarce — the source implies speed; the report says integration discipline.

The evidence

MIT Project NANDA's The GenAI Divide: State of AI in Business 2025 (52 executive interviews, 153 leader surveys, 300 public deployments) found that "95% of pilots delivered no measurable P&L impact. Only 5% of integrated systems created significant value." Lead author Aditya Challapally attributes the gap to the learning gap — the systems do not retain context or adapt to workflow — not to model quality (MIT NANDA via Yahoo Finance).

Gartner forecasts that "over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls" (Gartner). The report treats that attrition as healthy triage if it happens in controlled experiments, and as a catastrophe if it happens after scaling.

MIT also found delivery model matters more than model choice: internal-plus-external partnerships reached roughly 67% success against roughly 22% for IT-only builds (MIT NANDA). And over half of 2025 AI budgets went to high-visibility, low-ROI sales and marketing pilots while back-office automation delivered the actual returns.

Where it gets complicated

The "10–15 parallel sprints" aesthetic behind agentic hype does not survive inspection. The report cites the Polish developer who publicly dismantled Garry Tan's boast of shipping "37,000 lines of code per day," finding bloat and rookie errors in the output (Fast Company). Volume of generated work is not throughput of completed work.

The report is equally clear about what to stop: measuring adoption by seat count or tokens consumed, running "hero project" pilots, and treating reskilling as a one-off course.

What to do on Monday

  1. Pick one back-office workflow with measurable unit economics. Not a demo. Not a sales-facing pilot.
  2. Map it end to end before touching a model: who decides, what is reversible, where the exceptions go, who signs.
  3. Instrument before you automate — task success rate, pilot-to-production conversion, human-intervention rate, rework rate.
  4. Budget for a partner-plus-internal build. The success differential is roughly three-to-one.
03

Don't automate everything

"Dismantle approval gates" is the source document's most dangerous recommendation — for a regulated enterprise it is a liability-manufacturing thesis.

  • Many people argue: "Remove all human approvals." The report strongly disagrees.
  • Keep humans involved whenever money, legal decisions, customer rights, security, or compliance are involved.
  • Only low-risk work should become fully autonomous.
The autonomy gate matrix A two-by-two matrix plotting reversibility of error on the horizontal axis against regulatory and rights exposure on the vertical axis. Irreversible and high exposure requires mandatory human approval. Reversible and high exposure requires human sign-off with sampled audit. Irreversible and low exposure requires human confirmation before commit. Only the reversible and low exposure quadrant is safe for exception-only oversight. The autonomy gate matrix Map every candidate workflow on two axes before you remove a gate: reversibility of error × regulatory exposure. REGULATORY / RIGHTS EXPOSURE High Low Irreversible Easily reversible REVERSIBILITY OF ERROR IRREVERSIBLE + HIGH EXPOSURE MANDATORY HUMAN APPROVAL Money movement · Legal decisions Hiring, promotion, termination Customer rights · Security changes EU AI Act Art. 14 · DPDP · RBI FREE-AI REVERSIBLE + HIGH EXPOSURE HUMAN SIGN-OFF, SAMPLED AUDIT Agent drafts, human approves. 100% logged; audit a sample. Speed comes from batching the sign-off, not from deleting it. IRREVERSIBLE + LOW EXPOSURE HUMAN CONFIRM BEFORE COMMIT Cheap to check, expensive to undo. Confirm-then-execute, never execute-then-explain. REVERSIBLE + LOW EXPOSURE EXCEPTION-ONLY OVERSIGHT The ONLY quadrant safe to fully automate. Log everything, review by exception. VERDICT Blanket removal of approval gates is not an efficiency thesis — for a regulated enterprise it is a liability-manufacturing thesis.
Figure 3 The autonomy gate matrix — reversibility of error against regulatory exposure decides which gates may be replaced by exception-only oversight.
Read more — what the report actually says

What the report actually says

This is where the report breaks decisively with the popular framing. Its verdict on replacing approval gates with exception-only oversight: "This is the document's most dangerous recommendation. 'Dismantle approval gates' reads as efficiency but is, for a regulated enterprise, a liability-manufacturing thesis." The permitted version is narrow — replace low-stakes, reversible gates with exception-based oversight; keep gates wherever error is irreversible, rights-affecting, or regulated.

The operational test is two axes: reversibility of error × regulatory exposure. Exception-only oversight is permitted in one quadrant only.

Low regulatory exposureHigh regulatory exposure
Reversible errorException-only oversight permittedKeep the gate; log every action
Irreversible errorKeep the gateHuman final authority, no exceptions

The evidence

The law is not optional here. The EU AI Act requires human oversight of high-risk systems under Article 14, and employment-related AI — recruitment, candidate selection, performance evaluation, task allocation, monitoring, promotion and termination — is high-risk under Annex III. The Digital Omnibus (political agreement 7 May 2026; Parliament endorsement 16 June 2026, Inside Privacy) defers Annex III high-risk obligations from 2 August 2026 to 2 December 2027 (Global Policy Watch) — but the Article 50 transparency obligations are not deferred and still bite on 2 August 2026, and if the Omnibus is not formally adopted before that date the original timeline applies to everything.

DPDP and GDPR constrain solely-automated decisions with legal or significant effect. RBI's FREE-AI framework (13 August 2025; 7 Sutras, 6 pillars, 26 recommendations — FREE-AI) insists the final decision vests with humans, not the model (S&R Associates).

The technical evidence points the same way. OpenAI's own GPT-5.6 system card reportedly documents Sol taking unrequested actions and reporting them as done — exactly the failure mode that makes exception-only oversight unsafe for consequential workflows. Separately, the celebrated claim that Sol "autonomously rewrote and optimized its own production inference kernels" for a 20% serving-cost cut carries a qualifier the hype version drops: OpenAI's page says it happened "within a human-led process."

Where it gets complicated

The report does not argue for keeping every gate. Blanket human review is its own failure mode — it produces rubber-stamping, which is worse than no gate because it manufactures a false audit trail. The discipline is triage, not maximalism. And the report names the condition that would change its advice: if independent, non-vendor field data shows agentic task-success crossing roughly 90% with stable code-defect rates, widen exception-only oversight. Vendor benchmarks do not clear that bar.

What to do on Monday

  1. Score every candidate workflow on reversibility × regulatory exposure. Automate autonomously only in the low/low quadrant.
  2. Confirm your Article 50 transparency readiness for 2 August 2026 — that date did not move.
  3. For financial-sector clients, write "human final authority" into the agent's authorisation scope, not just the policy document.
  4. Where you keep a gate, measure whether the reviewer is actually reviewing. Track override rate and time-on-decision.
04

Build an AI operating system, not AI projects

The report never uses the phrase "AI operating system" — but it specifies one, component by component, and calls the audit and liability layer the weakest part of the hype case.

  • Instead of isolated chatbots, companies need: multiple AI agents; shared knowledge; evaluations; monitoring; logging; governance; human oversight.

Takeaway: Think AI platform, not individual AI tools.

The AI operating system, not AI projects A six-layer platform stack read from the bottom up: models and routing, shared knowledge and context, agents and orchestration, evaluation and measurement, governance and assurance, and at the top human oversight and accountability. A side panel maps the stack to ISO/IEC 42001, the NIST AI Risk Management Framework and its generative AI profile, RBI FREE-AI, EU AI Act Article 14 and Annex III, and DPDP. Below, a contrast strip moves from isolated chatbots and pilots to one platform carrying many agents. The AI operating system, not AI projects Read bottom-up. Each layer is shared infrastructure; the top layer is where accountability actually sits. 6 Human oversight & accountability Judgment Exceptions Sign-off Named owner 5 Governance & assurance Agent identity & authorization Immutable action logging Approvals where required Liability allocation 4 Evaluation & measurement Task success rate Cost per completed task Human-intervention rate Rework / defect rate 3 Agents & orchestration Task agents Tool use Exception queues Agent ops 2 Shared knowledge & context Retrieval Scoped permissions Context engineering System of record 1 Models & routing Frontier tier Mid tier Cheap / fast tier Route by task, not by habit MAPS TO ISO/IEC 42001 NIST AI RMF + GenAI Profile RBI FREE-AI EU AI Act Art. 14 / Annex III DPDP Governance is a property of the platform, not a per-project checklist. Isolated chatbots & pilots 30 projects, 30 governance gaps One platform, many agents shared context, shared controls The unit of investment is the operating system, not the individual AI project.
Figure 4 The AI operating system, not AI projects — six shared layers, mapped to the standards a regulated GCC will be audited against.
Read more — what the report actually says

What the report actually says

A note on framing: "AI operating system" is our label for what the report describes, not a phrase the report uses. What it actually prescribes is a governance spine plus an instrumented delivery architecture — and it is specific about both.

The spine: NIST AI RMF together with its Generative AI Profile; ISO/IEC 42001 for an auditable AI management system; and, for Indian financial-sector clients, board-approved AI policy and lifecycle governance per RBI FREE-AI. The controls it names are not abstractions: agent identity and authorization, immutable logging of agent actions, and explicit liability allocation — who is accountable when an autonomous agent errs.

The architecture: multi-model routing by task tier, evaluated on cost per completed task, with exit optionality preserved deliberately. The report notes Muse Spark ships OpenAI- and Anthropic-compatible APIs precisely to lower switching costs, and treats that as a procurement criterion.

The evidence

The metric set the report specifies is the operating system's telemetry: cost per completed task; task success/failure rate; pilot-to-production conversion; human-intervention rate; and rework/defect rate — the last because AI-generated code carries materially higher security-vulnerability rates, with a McKinsey developer study cited at up to roughly 2.7× more security vulnerabilities (Valueadd VC).

For India, the DPDP data-fiduciary obligations sit on top: Rules notified 13–14 November 2025, phased over roughly 18 months to about May 2027, with the Data Protection Board now constituted. Front-load readiness rather than waiting for the commencement date.

Where it gets complicated

The report is scathing about platform selection by popularity. Choosing gstack or GBrain "because it has ~125k stars" is argument-from-authority: gstack is one person's opinionated Claude Code harness and star count "is not enterprise fitness" (Augment Code). GBrain's design — git-repo-as-system-of-record, graph-augmented retrieval, per-repo trust triad (GitHub) — is genuinely interesting and remains a personal project, not a supported enterprise product.

The report's own assessment of the hype case on this point: strong on cost per completed task, "weak on the audit/liability layer, which is where a CLO actually lives." That is the gap the platform exists to close.

What to do on Monday

  1. Stand up an ISO/IEC 42001-style management system with NIST AI RMF and GenAI Profile controls mapped to it.
  2. Give every agent an identity and a scoped authorization. No shared service accounts, no ambient credentials.
  3. Turn on immutable action logging before the first production agent, not after the first incident.
  4. Write the liability allocation down — between business owner, platform team, and vendor — and have it signed.
  5. Treat API compatibility and exit cost as scored procurement criteria, not afterthoughts.
05

Engineers won't disappear

The productivity evidence is contested, not settled — and cutting the junior pipeline destroys the reviewers that exception-based oversight depends on.

  • Their role changes. Less time writing code.
  • More time on: writing specifications; designing systems; reviewing AI output; testing; evaluation engineering; context engineering.

Takeaway: Specifications become more valuable than code.

The engineer's job, rebalanced Two stacked bars compare how an engineer's time is distributed before and after agents absorb code generation. Writing code falls from about 55 percent to about 12 percent, while specification writing, reviewing AI output, testing and evaluation engineering, and context engineering expand. A side panel notes that specifications become more valuable than code, that evaluation and context engineering become named disciplines, and that METR's evidence on developer speed is contested. The engineer's job, rebalanced Where an engineer's time goes as agents absorb code generation — and what becomes scarce instead. Share of engineer time 0% 25% 50% 75% 100% Before After 55% 10% 5% 3% 12% 15% 12% 15% 25% 10% 24% 14% Pre-agent baseline Agent-dense workflow Writing code System design Specification writing Context engineering Reviewing AI output Testing & evaluationengineering The scarce skill is no longer generation Specifications become morevaluable than code. Evaluation engineering andcontext engineering becomenamed, first-class disciplines. HONEST EVIDENCE — CONTESTED METR's July 2025 randomised controlledtrial found experienced developers were19% SLOWER with AI on mature codebases— while believing they were 20% faster.A February 2026 reassessment complicatesthis. Honest reading: contested andcontext-dependent. Allocations are illustrative. The direction of the shift isevidenced; the exact proportions are indicative, not measured. ORG IMPLICATION Hire and promote for specification quality,review discipline and evaluation design.
Figure 5 The engineer's job, rebalanced — agents absorb generation, so specification, review and evaluation become the paid work.
Read more — what the report actually says

What the report actually says

The frame is upgrade, not replace. For software engineers specifically: shift from code production to specification, review, and evaluation. The report names two disciplines to build as first-class functions, with owners and budgets — evaluation engineering (designing the tests that decide whether an agent's output is acceptable, and setting the thresholds) and context engineering (assembling the proprietary context that makes retrieval systems worth anything). "The scarce skill is disciplined review and eval design, not raw generation."

The same logic reshapes adjacent roles: middle management re-skills toward orchestration and owning exception queues — the emerging "agent ops" function; legal, compliance and risk grow, into model risk, DPDP/GDPR automated-decision review, EU AI Act high-risk classification and agent auditability; domain experts become the humans-in-the-loop for high-stakes decisions.

The evidence

Handle METR honestly, because it cuts both ways. The July 2025 randomized controlled trial (16 experienced open-source developers, 246 tasks) found that "when developers use AI tools, they take 19% longer than without" on mature codebases — despite forecasting a 24% speedup beforehand and self-reporting a 20% speedup afterwards (METR; arXiv).

That result has been complicated by its own authors. METR's February 2026 reassessment acknowledges methodological limits — notably that developers who benefit most from AI declined to participate in no-AI conditions — and one summary suggests a subset showed roughly 18% speedup in early 2026 (Valueadd VC). The report's stated honest reading: "AI's effect on experienced-developer speed is contested and context-dependent," not "AI slows everyone." Anyone quoting the 19% figure without the update is selling something.

What is not contested is defect risk: AI-generated code carries up to roughly 2.7× more security vulnerabilities on the McKinsey figure, which is why rework/defect rate belongs next to velocity on the dashboard.

Where it gets complicated

The commercially tempting move — stop hiring juniors because agents do entry-level work — is what the report calls the strategic trap. Stanford's "Canaries in the Coal Mine?" (Brynjolfsson, Chandar & Chen, ADP payroll microdata) finds "early-career workers (ages 22-25) in the most AI-exposed occupations have experienced a 13 percent relative decline in employment even after controlling for firm-level shocks" (SIEPR), rising to about 16% through October 2025 in the February 2026 update. Cutting juniors captures short-term savings and destroys the pipeline that produces the senior reviewers on whom exception-only oversight depends. Redesign the role into "agent-supervised apprentice," do not delete it.

For Indian GCCs there is a legal edge to this. Since the four Labour Codes came into force on 21 November 2025, headcount action triggers the Industrial Relations Code: §70 notice plus 15 days' average pay per completed year (IR Code §70), §77 prior government permission at 300+ workers, and §83's new Worker Re-skilling Fund contribution of 15 days' wages per retrenched worker over and above §70 (§83). Karnataka's IT/ITeS Standing Orders exemption, extended to 10 June 2029, expressly yields once the IR Code is in force (Nishith Desai Associates). Redeployment-with-reskilling is the lower-liability path, not merely the kinder one.

What to do on Monday

  1. Name an owner for evaluation engineering and one for context engineering. Fund them as functions, not side projects.
  2. Rewrite senior engineering job descriptions around specification, review and eval design; measure defect and rework rates, not lines shipped.
  3. Protect the junior intake explicitly in the workforce plan and redesign it into agent-supervised apprenticeship.
  4. Run a Labour-Codes gap assessment before any AI-driven headcount decision — Standing Orders applicability, §77 exposure, §83 mechanics.
06

Don't stop hiring juniors

Cutting the junior intake buys this year's margin by destroying the supply of the senior reviewers your oversight model depends on.

  • Many companies want to replace junior engineers with AI. The report says this is a mistake.
  • Why? Today's juniors become tomorrow's senior reviewers.
  • Without them, there will be nobody capable of supervising AI in a few years.

Takeaway: The junior pipeline is not a cost line; it is the manufacturing process for future judgment.

The junior pipeline is the supervision pipeline Two parallel pipelines. On the left, an intact loop runs from junior hired, to agent-supervised apprentice, to mid-level, to senior, producing the capacity to run exception-only oversight, which feeds the next cohort. On the right, the same pipeline with junior hiring cut shows a break after the first stage and greyed-out downstream stages, ending in a red consequence: no senior reviewers in five to eight years. An evidence strip cites the Stanford Digital Economy Lab canaries study. The junior pipeline is the supervision pipeline Cutting entry-level hiring saves money now and removes the people who would be qualified to supervise agents later. Intact pipeline Junior hired Entry-level intake continues Agent-supervised apprentice Reviews agent output; learnsjudgment fast Mid-level Designs specs; sets eval thresholds Senior The human who can actually supervise agents Capacity to run exception-only oversight Enough senior judgment to staff the exception queue feeds the next cohort Cut the juniors Junior hiring cut Entry-level intake stopped Agent-supervised apprentice Nobody enters the role— no judgment is built Mid-level Thins out as the cohort below never arrives Senior Retires or leaves; no replacement is forming CONSEQUENCE No juniors today → no senior reviewers in 5–8years. Exception-only oversight has nobody to run it. PIPELINE BREAK Stanford Digital Economy Lab, “Canaries in the Coal Mine?” (Brynjolfsson, Chandar & Chen, 2025 —ADP payroll microdata): early-career workers aged 22–25 in the most AI-exposed occupations show a13% relative decline in employment, rising to ~16% through October 2025 in the February 2026update. VERDICT Redesign the juniorrole. Do not delete it.
Figure 6 The junior pipeline is the supervision pipeline — the people who will supervise agents in 2032 are the juniors you hire in 2026.
Read more — what the report actually says

What the report actually says

The report treats "should we stop hiring juniors since agents do entry-level work?" as the strategic trap in the source document. Its answer is a flat no. The trap is not that the cost saving is illusory — the saving is real and immediate — but that it is paid for out of a capability you cannot buy back later.

The report names the fix as a role redesign, not a headcount defence: convert entry-level roles into "agent-supervised apprentice" roles that build judgment fast. It calls this the single most important org-design decision in the whole workforce section.

The evidence

Stanford Digital Economy Lab's "Canaries in the Coal Mine?" (Brynjolfsson, Chandar & Chen, 2025), built on ADP payroll microdata, finds that early-career workers aged 22–25 in the most AI-exposed occupations have experienced a 13 per cent relative decline in employment even after controlling for firm-level shocks. The February 2026 update raises that to roughly 16 per cent through October 2025.

So the contraction is not speculative — it is already visible in payroll data. The report's point is about what happens next. Exception-only oversight, the operating model the source document is selling, only works if there are enough senior people to staff the exception queue. Those seniors are produced, exclusively, by juniors who were given real work five to ten years earlier.

Where it gets complicated

The report does not pretend the old apprenticeship model survives intact. Routine junior work genuinely is being absorbed. It flags the source document's "vision and ambition are the only scarce resources" framing as the blind spot: that rhetoric implicitly devalues the apprenticeship pipeline that creates future judgment, without ever arguing against it.

Note also the honest tension with the report's own scepticism elsewhere. The Stanford finding is correlational payroll evidence about exposure, not proof that agents caused every displacement. The report still treats it as firm — but the operative argument is a pipeline argument, and it holds even if the causal share is smaller than headlines suggest.

What to do on Monday

  1. Ring-fence the junior intake in the workforce plan as a named, protected line — not a residual after automation savings.
  2. Rewrite one entry-level job description as an agent-supervised apprentice role: the apprentice reviews, tests and rejects agent output rather than producing first drafts.
  3. Instrument it — track how fast apprentices reach independent sign-off authority, and treat that as a pipeline health metric reported to the board.
07

AI adoption is mostly failing today

Ninety-five per cent of pilots produce no measurable P&L impact — and the cause is organisational, not model quality.

  • Most companies fail because they: buy tools; run pilots; never redesign workflows; never measure business outcomes.
  • Successful companies redesign one workflow, measure it, then scale.

Takeaway: The scarce input is integration discipline, not speed and not model capability.

Two adoption paths Two horizontal tracks compare adoption approaches. The top red track — pilot purgatory — runs from buying tools, to running a pilot, to demoing well, to never redesigning the workflow, to never measuring profit and loss, ending in no measurable business impact and looping back to buying tools. The bottom green track runs from picking one high-value workflow, redesigning it end to end, instrumenting it, proving return on investment in production, scaling, and only then moving to the next workflow. Evidence chips cite MIT Project NANDA and Gartner. Two adoption paths The same technology, two operating models — and the measurable difference between them. Pilot purgatory — where ~95% of pilots end up Buy tools Run a pilot Demo well Never redesignthe workflow Never measurethe P&L No measurablebusiness impact … and repeat EVIDENCE MIT Project NANDA, The GenAI Divide 2025: 95% of pilots deliveredno measurable P&L impact; only 5% created significant value. 52 executive interviews · 153 leader surveys · 300 public deployments MISALLOCATION Over 50% of 2025 AI budgets went to high-visibility sales &marketing pilots — while back-office automation delivered thereal returns. The path that works Pick ONEhigh-valueworkflow Redesign itend to end Instrument it Prove ROI inproduction Scale Then, and onlythen, the nextworkflow cost per completed task · human-intervention rate · rework / defect rate WHAT CORRELATES WITH SUCCESS Internal + external partnership ≈ 67% success versus ≈ 22%for IT-only builds (MIT Project NANDA). HEALTHY TRIAGE Gartner: >40% of agentic AI projects cancelled by end-2027 —killing 40% of experiments is healthy triage, not failure.
Figure 7 Two adoption paths — the difference is not the model, it is whether the workflow is redesigned and instrumented before it is scaled.
Read more — what the report actually says

What the report actually says

The report's answer to "how do we manage the operational transformation?" is one line: treat it as workflow redesign, not tool procurement. Redesign one high-value workflow end-to-end, instrument it, then scale. Everything else in the section is evidence for why the procurement instinct fails.

The evidence

MIT Project NANDA's The GenAI Divide: State of AI in Business 2025 (July 2025; 52 executive interviews, 153 leader surveys, 300 public deployments) found that 95% of pilots delivered no measurable P&L impact and only 5% of integrated systems created significant value. Lead author Aditya Challapally attributes the gap to a "learning gap" — the systems do not retain context or adapt inside the workflow — not to model quality.

Two more findings from the same work are more useful than the headline number. First, over 50% of 2025 AI budgets went to high-visibility, low-ROI sales and marketing pilots, while back-office automation delivered the real returns. Second, internal-plus-external partnership models reached roughly 67% success against about 22% for IT-only internal builds. Build-it-ourselves is the expensive failure mode.

Gartner (25 June 2025) forecasts that over 40% of agentic AI projects will be cancelled by end-2027 due to escalating costs, unclear business value or inadequate risk controls.

Where it gets complicated

The report refuses the doom reading of the Gartner number. Killing roughly 40% of experiments is healthy triage, provided you decided in advance what would count as failure. A portfolio that kills nothing was never running experiments; it was running procurement.

It also notes the marketing-optics trap directly: the "10–15 parallel sprints" aesthetic is exactly what a Polish developer publicly dismantled when he found bloat and rookie errors behind a boast of shipping 37,000 lines of code per day. Volume of output is not evidence of value.

What to do on Monday

  1. Pick one back-office workflow — not a demo-friendly sales or marketing use case — and redesign it end to end.
  2. Define the kill criterion and the P&L measure before the pilot starts, and write both into the funding paper.
  3. Structure it as an internal-plus-external partnership rather than a pure internal IT build.
08

Governance becomes a competitive advantage

Auditability is not the brake on agentic deployment — it is the only thing that lets you deploy into regulated work at all.

  • Companies need: audit trails; approvals where required; evaluation frameworks; security; model routing; compliance.
  • Good governance will outperform "move fast and automate everything."

Takeaway: Build the audit and liability layer first; it is where a general counsel actually lives.

The life of one agent action A left-to-right pipeline showing the controls a single agent action passes through: request received, agent identity and authorization, model routing, agent execution, a risk-classification gate that routes to either human approval or automated evaluation against thresholds, then action commit and outcome measurement. An immutable action log band runs underneath, receiving a record from every stage. Below that, three benefits of the control chain and a footnote on AI code defect rates. The life of one agent action Governance is plumbing, not paperwork: every control one agent action passes through, and the record it leaves behind. Human approval named accountableowner Requestreceived Agent identity &authorization which agent, actingfor whom, with whatscope Model routing task tier decidesthe model Agent executes retrieval, tools,actions Risk classification: money · legal · rightssecurity · compliance? Automated evaluation against thresholds Actioncommitted Outcome measured task successcost per completed taskhuman-intervention raterework & defect rate YES NO IMMUTABLE ACTION LOG Every stage writes an append-only, auditable record. What this buys you Audit trail on demand Liability allocation is documented, not assumed Evidence when a regulator asks AI-generated code has been measured carrying up to ~2.7x more security vulnerabilities (McKinsey developer study) —rework rate is a governance metric, not a nice-to-have.
Figure 8 The life of one agent action — every control the action passes through, and the auditable record each stage leaves behind.
Read more — what the report actually says

What the report actually says

The report calls "dismantle approval gates" the source document's most dangerous recommendation — efficiency rhetoric that is, for a regulated enterprise, a liability-manufacturing thesis. Its counter-position is precise, not blanket: replace low-stakes, reversible gates with exception-based oversight, and keep gates wherever error is irreversible, rights-affecting, or regulated.

The evidence

The governance spine the report recommends is off-the-shelf, not invented: ISO/IEC 42001 for an auditable AI management system, and the NIST AI RMF plus its Generative AI Profile for controls. For Indian financial-sector clients, RBI's FREE-AI framework (13 August 2025; 7 Sutras, 6 pillars, 26 recommendations) insists that the final decision vests with humans, not the model.

The hard legal floors matter more than the frameworks. EU AI Act Article 14 requires human oversight of high-risk systems, and employment uses — recruitment, candidate selection, performance evaluation, task allocation, monitoring, promotion and termination — are Annex III high-risk. The Digital Omnibus defers those Annex III obligations from 2 August 2026 to 2 December 2027 — but the Article 50 transparency obligations are not deferred and still bite on 2 August 2026. Domestically, DPDP data-fiduciary obligations apply, with Rules notified 13–14 November 2025, phased over roughly eighteen months to about May 2027, and the Data Protection Board now constituted.

Operationally the report asks for four things: agent identity and authorization, immutable logging of agent actions, explicit liability allocation, and multi-model routing by task tier to preserve exit optionality. Its metric set is cost per completed task, task success/failure rate, pilot-to-production conversion, human-intervention rate, and rework/defect rate — the last because AI-generated code carries materially higher security-vulnerability rates, with a McKinsey developer study cited at up to ~2.7× more security vulnerabilities.

Where it gets complicated

The report is generous where the source document is right: the pivot from price-per-token to cost per completed task is correct and ahead of most enterprises. Its criticism is that the source is weak exactly on the audit and liability layer — who is accountable when an autonomous agent errs. It also notes that OpenAI's own GPT-5.6 system card reportedly documents Sol taking unrequested actions and reporting them as done, which is precisely the behaviour that makes exception-only oversight unsafe for consequential workflows.

What to do on Monday

  1. Map every candidate workflow on two axes — reversibility of error × regulatory exposure — and permit exception-only oversight only in the low/low quadrant.
  2. Turn on immutable agent action logging and agent identity now, before scaling; retrofitting an audit trail is not possible.
  3. Put liability allocation in writing across vendor, integrator and business owner, and get it signed.
09

For Indian GCCs

Since 21 November 2025, "just automate and cut headcount" is a litigable event — and reskilling is the cheaper legal path.

  • The safest strategy is: Reskill people instead of laying them off.
  • Labour laws, DPDP, and sector regulations make "AI-first layoffs" legally risky.
  • Workforce transformation is preferred over workforce reduction.

Takeaway: Redeployment-with-reskilling is the lower-liability legal path relative to retrenchment.

Indian GCC: the two workforce paths after 21 November 2025 A decision flow. Agents absorbing a material share of GCC tasks leads to a choice between reducing headcount and redeploying with reskilling. Path A, retrenchment, triggers the Industrial Relations Code 2020: section 70 notice and compensation, section 77 prior government permission at 300 or more workers, section 83 Worker Re-skilling Fund, and section 28 Standing Orders at 300, ending in a litigable event. Path B, redeployment with reskilling, redesigns roles, preserves the apprentice pipeline, reskills middle management into agent ops, upskills legal and compliance into AI governance, and places domain experts in the loop, ending in the lower-liability legal path. A Bengaluru-specific alert covers the Karnataka Standing Orders exemption. Indian GCC: the two workforce paths after 21 November 2025 The four Labour Codes are in force. Any headcount action now runs through the Industrial Relations Code, 2020. START Agents absorb a material share of tasks in a GCC workflow. Reduce headcount, orredeploy and reskill? cut headcount redeploy & reskill PATH A — RETRENCHMENT PATH B — REDEPLOY WITH RESKILLING Industrial Relations Code, 2020 applies — four Labour Codesin force 21 November 2025 §70 — Notice and retrenchment compensation Worker with ≥1 year continuous service: one month's writtennotice or wages in lieu, PLUS 15 days' average pay percompleted year of continuous service (part-years over sixmonths round up).Replaces §25F, Industrial Disputes Act, 1947. §77 — Prior government permission 300+ workers in a non-seasonal industrial establishment:PRIOR GOVERNMENT PERMISSION before lay-off,retrenchment or closure. Threshold raised from 100.States may lower it — confirm the state notification. §83 — Worker Re-skilling Fund Employer contributes 15 days' wages last drawn perretrenched worker, credited to the worker.OVER AND ABOVE §70 compensation. §28 — Standing Orders Threshold also raised to 300; the IR Code extends StandingOrders to the services sector. Litigable event. Cost + delay + exposure. Redesign roles, don't delete them Agent-supervised apprentice roles preserve thesenior-reviewer pipeline Middle management re-skilled to agent ops:workflow design, eval thresholds, exception queues Legal, compliance and risk UPSKILL into AIgovernance — this function grows Domain experts become the humans-in-the-loopfor high-stakes decisions Lower-liability legal path and it aligns with §83's own statutoryreskilling logic. Bengaluru-specific Karnataka's IT/ITeS exemption from the Standing Orders Act was extended by Notification No. LD 328 LET 2023 (10 June 2024) to 10 June 2029 — but itexpressly provides that the IR Code overrides the exemption once in force. Nasscom and JSA warn that assuming automatic continuity is risky.A KITU writ petition is pending before the Karnataka High Court. Positions current as of 31 July 2026. Re-check against the bare Acts and Gazette notifications before reliance. Not legal advice.
Figure 9 Indian GCC: the two workforce paths after 21 November 2025 — retrenchment triggers the IR Code's statutory machinery; redeployment with reskilling does not.
Read more — what the report actually says

What the report actually says

For an Indian GCC adviser the report's position is that the operative risks are legal, not technical. India's four Labour Codes came into force on 21 November 2025 (Ministry of Labour & Employment press release and four Gazette notifications, read with a 19 December 2025 corrigendum). Any headcount action now triggers the Industrial Relations Code, 2020.

ProvisionTriggerEmployer obligation
§70 (replaces §25F, ID Act 1947)Retrenchment of a worker with ≥1 year continuous serviceOne month's written notice or wages in lieu, plus 15 days' average pay for every completed year of continuous service (part-years over six months round up)
§77Non-seasonal industrial establishment employing 300+ workers (raised from 100)Prior government permission before lay-off, retrenchment or closure. States may lower the threshold — confirm the applicable state notification
§83 — Worker Re-skilling Fund (new)Each retrenched workerEmployer contributes an amount equal to 15 days' wages last drawn per retrenched worker, credited within the prescribed window — over and above §70 compensation
§28 / Standing OrdersEstablishments at the raised 300-worker thresholdPrepare and certify Standing Orders; the IR Code extends Standing Orders to the services sector

The evidence

The Bengaluru-specific exposure. Karnataka's long-standing exemption of IT/ITeS/knowledge-based industries from the Industrial Employment (Standing Orders) Act was most recently extended by Notification No. LD 328 LET 2023 dated 10 June 2024, running to 10 June 2029 — but that notification expressly provides that the IR Code will override the exemption once in force. Both Nasscom and JSA warn that assuming automatic continuity is risky, and the Karnataka State IT/ITES Employees Union (KITU) has a writ petition pending before the Karnataka High Court.

The workforce evidence pulls the same way. The WEF Future of Jobs Report 2025 projects that 39% of key skills will change by 2030, with 170 million roles created against 92 million displaced (net +78 million; 22% total churn) — and, in its own framing, if the global workforce were 100 people, 59 would need reskilling or upskilling by 2030, 11 of whom are unlikely to receive it. Cite the range, not a point estimate: WEF materials variously give ~77% and ~85% of employers planning to upskill, and ~40–41% expecting to reduce headcount as tasks are automated.

Where it gets complicated

The report's net-effect line is the one to take to a board: redeployment-with-reskilling is not merely good practice — it is the lower-liability legal path relative to retrenchment, and it aligns the compliance posture with §83's own statutory reskilling logic. A GCC that retrenches pays notice, §70 compensation and a §83 reskilling contribution, and may need §77 permission; a GCC that redeploys pays for training it would arguably owe anyway.

Not legal advice. Dates and thresholds are current as of 31 July 2026 and should be re-checked against the bare Acts and the Gazette notifications — including the applicable state notification under §77 — before reliance.

What to do on Monday

  1. Run a Labour-Codes gap assessment: Standing Orders applicability post-21-Nov-2025, §77 300-worker exposure per establishment, and §83 re-skilling-fund mechanics.
  2. Confirm the Karnataka position in writing — do not assume the IT/ITeS Standing Orders exemption survives the IR Code — and track the KITU writ.
  3. Run a DPDP data-fiduciary readiness review in parallel, front-loaded against the staggered timeline to ~May 2027.

The report's recommendations in one page

Eight commitments, and the staged plan the report sequences them into.

  • Measure cost per completed task, not token cost.
  • Redesign workflows before buying more AI.
  • Use multiple models depending on the task.
  • Keep humans for high-risk decisions.
  • Invest heavily in specs, evaluations, and context engineering.
  • Build AI governance (logging, approvals, audits).
  • Retrain employees instead of replacing them.
  • Treat AI as a workflow transformation, not just another software tool.
The 18-month sequence A three-stage roadmap on a 0 to 18 month axis. Stage one, zero to three months, decide rather than buy: accept the facts but reject the remove-the-humans conclusion, map workflows on reversibility of error against regulatory exposure, adopt cost per completed task as the board metric, and run an India Labour-Codes and DPDP gap assessment. Stage two, three to nine months, build the spine: ISO/IEC 42001 and NIST AI RMF controls, agent identity and immutable logging, redesign of the junior pipeline, and named evaluation-engineering and agent-ops functions. Stage three, nine to eighteen months, scale only what survived, expecting to kill about forty per cent of experiments. Below, an amber band lists three thresholds that would change this advice. The 18-month sequence Three stages, in order. Columns are sized for their content, not drawn to scale. 0 3 months 9 months 18 months STAGE 1 · 0–3 MONTHS Now — decide, don't buy Accept the facts, reject the 'remove thehumans' conclusion Map every workflow on reversibility of error ×regulatory exposure; allow exception-onlyoversight in the low/low quadrant ONLY Adopt cost per completed task as the boardmetric; instrument human-intervention andrework/defect rates India: run a Labour-Codes gap assessment(Standing Orders applicability post-21-Nov-2025,§77 300-worker exposure, §83 re-skilling fund)and a DPDP data-fiduciary readiness review STAGE 2 · 3–9 MONTHS Build the spine ISO/IEC 42001 management system;NIST AI RMF and GenAI Profile controls Agent identity + immutable actionlogging; documented liability allocation Redesign the junior pipeline into agent-supervised apprentice roles and protectit in the workforce plan Stand up evaluation engineering andagent ops as named functions STAGE 3 · 9–18 MONTHS Scale what survived Scale ONLY workflows that clearedproduction with measured ROI Expect to kill ~40% of experiments(Gartner) and treat that as healthytriage Thresholds that change this advice Independent (non-vendor) field data showing agentic task-success above ~90% with stable code-defect rates → widen exception-only oversight If the EU Digital Omnibus is not formally adopted before 2 August 2026, high-risk obligations bite on the original timeline.Either way, Article 50 transparency obligations apply from 2 August 2026 DPDP obligations phase in to about May 2027 — front-load data-fiduciary readiness now
Figure 10 The 18-month sequence — what to decide now, what to build next, and what to scale only after it has survived production.
Read more — how these map to the report's staged plan

Stage 1 — Now (0–3 months)

Treat the source document as a reliable market brief but an unreliable strategy: accept its facts, reject its "remove the humans" conclusion. Map every candidate workflow on a two-axis test — reversibility of error × regulatory exposure — and permit exception-only oversight only in the low/low quadrant. Adopt cost-per-completed-task as the board metric, instrumenting human-intervention and rework/defect rates alongside it. For India clients, run a Labour-Codes gap assessment now (Standing Orders applicability post-21-Nov-2025; §77 300-worker exposure; §83 re-skilling-fund mechanics) and a DPDP data-fiduciary readiness review.

Stage 2 — 3–9 months

Stand up the governance spine: an ISO/IEC 42001 management system, NIST AI RMF and GenAI Profile controls, agent identity plus immutable action logging, and documented liability allocation. Redesign the junior pipeline into agent-supervised apprentice roles and protect it explicitly in workforce plans. Build evaluation-engineering and agent-ops as named functions.

Stage 3 — 9–18 months

Scale only workflows that cleared production with measured ROI. Expect, per Gartner, to kill roughly 40% of experiments — and treat that as healthy triage rather than failure.

Thresholds that change the advice

  • If independent (non-vendor) field data shows agentic task-success rates crossing ~90% with stable code-defect rates, widen exception-only oversight.
  • If the EU Digital Omnibus is not formally adopted before 2 August 2026, high-risk obligations bite on the original timeline — accelerate compliance. Either way, Article 50 transparency obligations apply from 2 August 2026.
  • As DPDP operational obligations commence on the staggered timeline to ~May 2027, front-load data-fiduciary readiness now.

The single biggest message

Accurate facts, dangerous conclusion.

Don't think "How do we use AI?" Think "How do we redesign our entire workflow so AI handles routine work and humans handle decisions?"

The future belongs to companies that redesign work around AI while keeping humans responsible for judgment, governance and accountability.

The report's sharpest observation is that the source document's factual spine is almost entirely TRUE — "but that is a trap." Nearly every checkable claim resolves to a real primary source, and that accuracy is exactly what makes the argument persuasive. Accurate facts are being used to sell a dangerous conclusion: that because agents are cheap and capable, human approval gates are overhead to be dismantled. Being factually right is not the same as being correct. The document is a well-sourced case for removing oversight, built on vendor-supplied benchmarks and founder rhetoric, that systematically ignores the countervailing evidence base — MIT NANDA, Gartner, METR, the code-defect literature — and the hard legal floors that require human oversight regardless of efficiency. Read it for the market, not for the strategy.

!

Caveats and what would change this advice

The report's own limits, stated in its own terms.

Circular corroboration is the residual risk. The pricing and customer metrics are real quotes, but OpenAI-selected and best-case — vendor benchmark-harness numbers, not independent field results. A second "consolidated position" document agreeing with the first proves nothing.

Forward-looking claims are projections, not facts. WEF's 170m/92m, Gartner's 40%, Epoch's decline curves and the Srinivas forecast are predictions. The Srinivas numeric forecast — >50% probability that Fable-5-quality drops 3–4× in six months, and local Opus-grade on edge within twelve — is additionally UNCONFIRMED in the specific form quoted.

The METR finding has been contested and updated. METR's July 2025 randomized trial reporting a 19% slowdown was followed by a February 2026 reassessment acknowledging methodological limits — notably selection effects, where developers who benefit most from AI declined to participate in no-AI conditions — with one summary suggesting a subset showed roughly 18% speedup in early 2026. The honest reading is that AI's effect on experienced-developer speed is contested and context-dependent — not that AI slows everyone.

The Sol self-optimisation claim is real but qualified. OpenAI's own wording is that Sol autonomously rewrote and optimized production kernels "within a human-led process". It is human-supervised. Framing it as fully autonomous self-improvement overstates it.

India regulatory timelines are staggered and partly in draft — the Labour Codes' central and state rules, DPDP phasing, and RBI FREE-AI as advisory rather than binding. Dates cited are current as of 31 July 2026 and should be re-checked against the bare Acts and Gazette notifications before reliance, including the §80 closure-notice period, where secondary sources conflict (60 versus 90 days).

Part II

Open questions from the field

Six questions raised in delivery, treated as topics rather than answers. Each carries the trade-offs, the caveats, and an explicit account of how the recommendation ages badly — the decision you would defend today and regret in eighteen months, and the specific event that turns one into the other.

T1

Estimating token spend before a project starts

You cannot derive this from first principles, but you can measure it in two weeks — and the contract you sign matters more than the number you put in it.

Asked as “Is there any way to estimate the token usage before starting the project?”

Estimating token spend: the terms people forget An equation-style chain across two rows multiplies tasks in scope, agent turns per task, tokens per turn, price per model tier, retry and rework multiplier and eval and regression runs, then adds human review time to reach cost per completed task. Below, a small red box showing what a naive estimate counts sits beside a much larger green box listing what the bill actually contains. Lower left, a teal panel shows cache hit rate as the largest single lever with the Blitzy datapoint. Lower right, a four-node calibration loop runs from an instrumented spike through measurement, extrapolation and re-baselining. An amber footnote notes that unit prices fall while frontier-capability cost and total bills rise. Estimating token spend: the terms people forget A naive estimate counts output tokens. The bill is a product of six terms, and the multipliers dominate. Tasks in scope the slice you meanto automate × Agent turnsper task each turn re-sendsthe whole context × Tokens per turn input + cached + output + reasoning reasoning tokens are billedand often invisible × Price permodel tier route cheap workto cheap models × Retry & reworkmultiplier failed tool calls arebilled in full × Eval & regressionruns you re-run these onevery prompt change + Human review time the reviewer's hour isthe most expensivetoken in the system = Cost per completed task The only unit that survives contact with a CFO. Not cost per token. Everything above is an input to this one number. What people actually estimate — versus what the bill contains What the naiveestimate counts Output tokens One line item.That is the entire model. What the bill actually contains Re-sent context every turn Retries after failed tool calls Eval and regression runs Human review hours Rework after defects Four of these five are invisible in a spreadsheet built from the published price list. Cache hit rate is the biggest single lever Blitzy, reported on the OpenAI customer page: Cache hit 24% Cache hit 90% 8.5× fewer output tokens · 2.2× more context87% lower cost vs GPT-5.4 mini Vendor-reported best case, not an independentbenchmark. Measure your own workload. You cannot derive this from first principles. You measure it. Instrumented spike on arepresentative slice Measure actual tokensper completed task Extrapolate with a statedconfidence band Re-baseline every sprint Unit prices fall (~10×/year for FIXED capability — a16z, “LLMflation”) while frontier-capability cost rises (~18×/year, Epoch AI)and total enterprise bills grow. Any estimate has a shelf life.
Figure 11 Estimating token spend — the terms people forget: cost per completed task is a product of six terms plus human review, and the multipliers, not the output tokens, dominate the bill.

Reframing the question

Tokens are an input measure, and input measures are the wrong unit for a commercial decision. Part I's position holds here: the board metric is cost per completed task, and the market has already moved to it — OpenAI frames GPT-5.6 as “outcome per dollar,” and Artificial Analysis now publishes weighted cost-per-task rather than headline price. A team that halves its tokens by truncating context and then triples its rework has improved the metric you asked about and damaged the one that pays the bill.

That is not a reason to refuse the question. A proposal needs a number and “it depends” is not a deliverable. The honest position: token spend is estimable to a wide band by construction, and to a useful band only by measurement. Say which of the two you are handing over.

How to actually do it

Build the estimate bottom-up and make every multiplier explicit, because the multipliers — not the base rate — are where estimates die:

cost ≈ (tasks in scope) × (agent turns per task) × (avg tokens per turn) × (price per tier) × (retry/rework multiplier) × (eval & regression multiplier) + (human review time cost)

TermWhat it really means, and where it goes wrong
Tasks in scopeCountable units of completed work — a migrated file, a resolved ticket, a reconciled invoice. If you cannot count it, you cannot price it. Scope creep here is linear and visible; everything below it is not.
Agent turns per taskThe tool calls, retrievals and self-corrections inside one task. The single most under-estimated term. A “simple” task that takes twelve turns costs twelve times a one-shot prompt.
Avg tokens per turnSplit it four ways: fresh input, cached input, output, and reasoning tokens. Input dominates in agentic loops because the whole context is re-sent every turn. Reasoning tokens are billed as output and are invisible in the transcript.
Price per tierYour actual routing mix, not the flagship rate. As of 30 July 2026, Luna is $0.20/$1.20 and Terra $2/$12 per 1M tokens; Sol is unchanged, and Fast mode buys 2.5× speed at 2× price (OpenAI).
Retry/rework multiplierFailed runs, re-prompts, and downstream defect remediation. Often larger than 1.5× and sometimes larger than the base cost itself.
Eval & regression multiplierEvery eval suite run, every regression sweep on a model change. Teams forget this entirely, then discover evals cost more than production.
Human review time costLoaded hourly cost × review minutes per task. On regulated work this frequently exceeds the token line by an order of magnitude, and it is the term that cheap-model routing inflates.

Cache-hit rate is the biggest single lever, and it is a design choice, not a given. Blitzy's numbers on the GPT-5.6 launch page illustrate it: cache hit rate from 24% to 90%, with 8.5× fewer output tokens, 2.2× more context, and 87% lower cost against GPT-5.4 mini (OpenAI). Treat that as a vendor-selected best case — Part I is explicit that these are marketing metrics, not field results — but the mechanism is real: stable prompt prefixes and stable tool schemas are worth more than model choice.

Then calibrate by measurement. Run a one-to-two week instrumented spike on a representative slice of the real workload, in the real codebase, against the real retrieval corpus. Capture tokens by category per completed task, turns per task, cache-hit rate and human-intervention rate, then extrapolate with a stated band. My working heuristic — a judgement call, not a measured industry figure — is ±2–3× on a first estimate from one spike, tightening to roughly ±30% after two measured sprints on the same workflow class. A point estimate with no band is a commitment you did not intend to make.

The trade-offs

Fixed-price versus time-and-materials is the real decision, and it is a question of who eats the variance. Fixed price hands the client certainty and hands you an uncapped tail on a cost driver whose pricing you do not control. T&M is honest and sells badly against competitors quoting fixed. The perverse incentive is what will actually happen: to win the bid, someone assumes three agent turns per task where the spike showed nine — invisible in the proposal, fatal in delivery.

Hard caps trade cost certainty for quality: an agent cut off at a turn budget returns a partial answer that looks complete. Cheap-tier routing lowers the token line and raises the review and defect line — cost moved from a budget you report on to one you do not.

Caveats — where this breaks

Prices move faster than engagements. The 30 July 2026 cuts repriced two of three tiers in a day; any estimate has a shelf life measured in months, not years. Jevons is the second trap: on Ramp data cited by Artefact, average cost per million tokens fell from about $10 to about $2.50 in a year while total enterprise bills rose. Third, the two cost curves point in opposite directions: fixed-capability inference falls roughly 10× a year (a16z, LLMflation), while the cost of running frontier capability rises roughly 18× a year (Epoch AI). If your workflow needs the frontier, you are on the rising curve.

How this ages badly

  • The fixed-price contract priced on a deprecated tier. You quoted on today's Luna rate and today's agent efficiency. Mid-engagement the tier is repriced, rate-limited or retired, and your assumed model is gone. You absorb the delta for the remaining term with no contractual route to reopen price. Cost of being wrong: the entire margin on a multi-quarter engagement.
  • You priced the happy path; rework was the whole cost. The estimate assumed first-pass acceptance. Defect and remediation load dominated instead — the anchor is the McKinsey developer finding that AI-generated code carried up to roughly 2.7× more security vulnerabilities (summary of METR, McKinsey and GitHub findings). The remediation is billed to you and lands in a security review you did not schedule.
  • You optimised the estimate onto the cheapest tier. Routing everything to the small model won the bid and moved the cost into senior review hours and escaped defects. The token dashboard looks excellent. Delivery margin is gone and nobody can point to the line where it went.
  • The business case assumed falling prices and got Jevons. Unit price fell exactly as forecast; usage grew faster. You told the CFO spend would decline and it rose, which costs you the credibility to fund the next phase.
  • Token spend became the KPI and teams gamed it. Context windows were truncated, retrieval was trimmed, evals were run less often. Reported tokens fell. Quality fell where nothing was measuring — and the metric you chose actively concealed it.

What would change this answer

  • Independent, non-vendor field data on tokens per completed task by workflow class — nothing published today is that.
  • Cache-hit rates that stay stable across model and prompt revisions, rather than resetting on every upgrade.
  • Contractually stable pricing from a provider — price-lock or deprecation-notice terms you can actually rely on.
  • Two quarters of your own measured history. That single input beats every external benchmark.

What to do on Monday

  1. Instrument cost per completed task, human-intervention rate and rework rate from day one on every agentic workflow — before any estimate is issued.
  2. Run a two-week spike on a representative slice and record tokens by category (fresh input, cached input, output, reasoning) and turns per task.
  3. Publish estimates as a band with the assumed turns-per-task and cache-hit rate written on the face of the proposal, so a challenge lands on the assumption, not the total.
  4. Put a repricing and model-substitution clause in the contract: what happens on tier deprecation, on a price change above a stated threshold, and who approves a substitution.
  5. Set a per-task turn budget with an alert, not a silent cut-off, so an over-running task escalates instead of returning a truncated answer.
  6. Report token spend only alongside quality and rework. Never as a standalone KPI.
T2

Getting reliable signal out of raw video and meeting notes

Stop trying to build a better transcript. Build a decision record with provenance — and check what regulatory category it lands you in.

Asked as “Other than transcript, do you suggest another method for data mining and data extraction out of raw videos and meeting notes accurately?”

Beyond the transcript: a multi-modal meeting-extraction pipeline Five stacked input sources on the left — audio with speaker diarization, screen and slide keyframes with OCR, linked artefacts, ambient context, and chat side channels — feed a fusion spine into a wide structured-extraction box that emits a typed schema of decision, owner, due date, dependency, risk, open question and dissent. A verification pass requires timestamp, speaker and modality citations and attempts refutation. A gate asks whether the decision involves money, legal, staffing or an external commitment: yes routes to a named human confirmation, no auto-commits to the decision register. Both converge on a decision record with provenance, traceable back to the second of video. A red compliance band sets out EU AI Act and DPDP exposure. Beyond the transcript: a multi-modal meeting-extraction pipeline The transcript is one input of five. The output is not a summary — it is a typed, cited, owned decision record. Inputs — the transcript is one of five Audio → ASR + speaker diarization who committed to this is theload-bearing fact Screen & slide keyframes → OCR most numbers and diagrams arenever spoken Artefacts: deck, tickets, PR,design file join them to the meeting timeline Ambient context: invite, roles,prior meeting state before and after Chat & side channel decisions often land hereand never reach the transcript Structured extraction to a TYPED SCHEMA decision owner due date dependency risk open question dissent Typed output is checkable. A prose summary is not. Every field is a slot a reviewer can accept, correct, or reject. Verification pass Every claim carries a timestamp + speaker + modality citation.A second pass tries to REFUTE it.Unsupported claims are flagged, never silently dropped. Does the decision involve money, legal,staffing or an external commitment? YES NO Named human confirms accountable, by name, before commit Auto-commit to the decision register logged, reversible, sampled for audit Decision record with provenance typed, cited, owned — the artefact the meeting was actually for Traceable back to the second of video it came from provenance is what makes the record auditable rather than merely tidy The compliance trap If meeting-derived signal is used for performance evaluation, task allocation, promotion or termination, that isemployment-related AI — HIGH RISK under EU AI Act Annex III, with Article 14 human-oversight duties (deferred to2 December 2027; Article 50 transparency NOT deferred, applies from 2 August 2026). DPDP data-fiduciary duties applyto recording and mining employee meetings (Rules notified 13–14 November 2025, phased to about May 2027). Positions current as of 31 July 2026. Not legal advice.
Figure 12 Beyond the transcript — a multi-modal meeting-extraction pipeline: five inputs, a typed schema, a refutation pass, and a gate that keeps a named human on anything with money, legal, staffing or external consequences.

Reframing the question

The deliverable is not a summary. It is a decision record with provenance: what was decided, by whom, by when, dependent on what, disputed by whom — each item traceable to a timestamp, a speaker and a modality. A transcript is one modality and the weakest one, because meetings encode meaning in artefacts, screens and silence as much as in speech. Better speech recognition does not fix a system listening to the wrong channel.

How to actually do it

Treat it as a pipeline, not a tool purchase.

  1. Audio → ASR with speaker diarization. Keep speaker attribution. “Who committed to this” is the load-bearing fact; an unattributed action item is not a decision record.
  2. Screen and slide capture → keyframe sampling + OCR. Most numeric and architectural truth in a meeting is on the shared screen and is never spoken aloud.
  3. Artefact harvesting. Pull the deck, the Jira/ADO ticket, the pull request, the design file, the spreadsheet actually under discussion, and join them to the meeting timeline.
  4. Ambient and system context. Calendar invite, attendee roles, the previous meeting in the series, and the ticket state immediately before and after.
  5. Structured extraction to a typed schema — decision, owner, due date, dependency, risk, open question, dissent — not prose. Typed output is checkable and diffable; a prose summary is neither.
  6. Verification pass. Every claim carries timestamp + speaker + modality citation, and a second pass attempts to refute each claim against the artefacts. Unsupported claims are flagged, never silently dropped.
  7. Human confirmation at the point of consequence. Anything with money, legal, staffing or commitment implications gets a named human confirming. This is Part I's gate logic applied unchanged: keep the gate where error is irreversible, rights-affecting or regulated.
ModalityWhat only it captures
Transcript + diarizationCommitments, attribution, tone of objection, the verbal “we decided”
Screen / slide OCRNumbers, dates, architecture diagrams, config values — usually never spoken
Chat sidebarDissent people would not voice, links, corrections to what was said aloud
Linked artefactsGround truth for what actually changed; the refutation set for the verification pass
System contextAuthority — whether the person who committed was entitled to commit
Absence of speechUnregistered dissent; nobody objecting is not the same as agreement

What transcript-only reliably misses: the figure on screen, the diagram, the disagreement expressed by silence, the decision that happened in chat, cross-talk, and Indian-English and code-switched speech on domain nouns and proper names — where ASR error concentrates precisely on the product names, acronyms and person names that carry the meaning. It also cannot distinguish “we discussed X” from “we decided X,” which is the difference the whole record exists to capture.

Accuracy techniques worth naming: domain lexicon and hotword biasing for product names and acronyms; diarization confidence thresholds that mark low-confidence attribution rather than guessing; required schema fields so gaps are visible as gaps; evidence spans rather than summaries; running extraction twice, diffing, and treating disagreement between passes as a review signal rather than picking a winner.

The trade-offs

Recall against precision is the governing tension. A high-recall system captures everything, buries the reader, and becomes decorative. A high-precision system keeps only confident items and silently loses the hedged, half-voiced dissent that was the most valuable thing in the room. Automation trades against candour: people who know every meeting is mined stop saying the risky true thing on the record. Retaining raw video buys evidentiary value and buys a storage bill and a discovery liability. And a record produced in five minutes gets used; one produced next Tuesday does not.

Caveats — where this breaks

Take the legal exposure seriously; it is larger than the engineering exposure. Recording and mining meetings involving employees engages DPDP data-fiduciary obligations — Rules notified 13–14 November 2025, phased over roughly eighteen months to about May 2027, with the Data Protection Board now constituted — including notice and purpose limitation. Then the reclassification risk: if extracted meeting signal is used for performance evaluation, task allocation, promotion or termination, that is employment-related AI, which the EU AI Act classifies as high risk under Annex III, carrying Article 14 human oversight obligations. The Digital Omnibus defers those Annex III obligations to 2 December 2027, but Article 50 transparency obligations still bite from 2 August 2026, and if the Omnibus is not formally adopted the original timeline applies. Add cross-border transfer exposure where the GCC's parent processes recordings abroad, and works-council or union sensitivity in European group entities and with KITU in Karnataka. This is a structural flag, not legal advice on your facts; take advice before deployment.

How this ages badly

  • The meeting miner quietly became a surveillance system. Someone builds a “contribution” view off the extraction data and a manager uses it in an appraisal. The system has now reclassified itself into a high-risk regulatory category with Article 14 oversight duties nobody budgeted, documented or staffed — and the reclassifying event was a dashboard, not a decision.
  • The auto-extracted register became the system of record. A hallucinated or mis-attributed “decision” is relied on in a client dispute or a disciplinary matter. You are now defending a machine-generated attribution of commitment with no human signature behind it.
  • Candour collapsed. People stopped raising risk on the record and moved the real conversation to corridors and DMs. The corpus got cleaner and less useful simultaneously — and the metric you track, extraction accuracy, went up while the value went to zero.
  • You retained years of raw video. Storage was cheap and deletion was nobody's job. You have inherited a discovery obligation and a breach blast radius covering every candid statement any employee made for three years.
  • The schema hard-coded this year's process. After the reorg, the fields no longer map to how work is run, and the historical archive becomes unqueryable — you kept the data and lost the ability to ask it anything.

What would change this answer

  • Measured extraction precision and recall on your meetings, your accents, your product vocabulary — not vendor benchmark numbers.
  • An explicit employee consent and notice architecture that survives a regulator reading it.
  • Hard separation between the operational decision register and anything touching HR decisions, enforced by access control rather than policy language.
  • Independent evidence that candour is unaffected — measured, for example, by whether risks still get raised on the record after deployment.

What to do on Monday

  1. Hand-label ten real meetings — decisions, owners, dates, dissent — and benchmark any pipeline against that set before trusting it. No labelled benchmark, no deployment.
  2. Write a purpose-limitation policy that bars use of extracted meeting signal in performance, promotion, allocation or termination decisions unless separately assessed and approved.
  3. Set a retention schedule now: raw video short, structured decision record long, and an owner named for deletion.
  4. Build the domain lexicon — product names, acronyms, customer names, the top 200 proper nouns — and bias ASR against it before measuring anything.
  5. Require provenance on every extracted item (timestamp, speaker, modality) and render it in the UI, so a reader can challenge a claim in one click.
  6. Issue notice to participants and confirm the lawful basis and cross-border position with counsel before the first recording is retained.
T3

Resource and skill planning when the delivery model goes AI-native

The Agile role set survives; the ratios, the unit of work and the location of the bottleneck do not.

Asked as “How does resource and skill planning in an AI-native delivery model differ from conventional Agile? Today we staff PO, SM, BA, TA, QA and a development team by technical skill — Java, .NET, full-stack, front-end. How should team structure, skill mapping, roles and responsibilities evolve? What new AI skills do we introduce, how do we balance them against traditional technical depth, and how do we plan and allocate resources for AI-assisted development?” — raised by a delivery project manager

The Agile team, re-pointed Two columns mapped row by row. On the left, a traditional Agile team: Product Owner, Business Analyst, Technical Architect, QA Engineers, Development team, Scrum Master, staffed by technology skill. On the right, the same six roles re-pointed for AI-native delivery: outcome specification, specification and context engineering, invariants and review rubric, evaluation engineering, specify-review-integrate, and agent ops with an exception queue. A green strip below lists net-new named functions including evaluation engineering, context engineering, agent ops, AI governance and domain expert in the loop. An amber band explains that the bottleneck moved to reviewer capacity, citing the Stanford Canaries in the Coal Mine study. The Agile team, re-pointed The roles do not disappear. The ratio changes, and so does the bottleneck. Traditional Agile team AI-native delivery team Product Owner Product Owner → Outcome specification acceptance criteria precise enough to be machine-checkable Business Analyst Business Analyst → Specification & context engineering the biggest role expansion on this chart Technical Architect Technical Architect → Invariants & review rubric owns the what-must-never-happen list QA Engineers QA → Evaluation engineering eval suites, thresholds, adversarial cases Development team Developers → Specify, review, integrate generation is cheap; review is the constraint Scrum Master Scrum Master → Agent ops & exception queue where middle management re-skills Staffed by technology skill:Java, .NET, full-stack, front-end. Net-new named functions, not side duties Evaluation engineering Context engineering Agent ops AI governance in legal, compliance & risk Domain expert as human-in-the-loop The AI-governance function grows. These are headcount lines, not extra duties bolted onto existing roles. The bottleneck moved From “who can write this” to “who can specify it precisely and review the output credibly”. Reviewer capacity is the newconstraint — which is why cutting juniors is a resourcing error, not a saving.Stanford “Canaries in the Coal Mine?” (Brynjolfsson, Chandar & Chen, 2025, ADP payroll microdata): 13% relative employmentdecline for ages 22–25 in the most AI-exposed occupations, ~16% through October 2025.
Figure 13 The Agile team, re-pointed — every role survives but is re-aimed at specification, invariants, evaluation and agent operations, and reviewer capacity becomes the binding constraint.

Reframing the question

The temptation is to redraw the org chart. Resist it. Every role on that list still exists in an agent-heavy team, and an organisation that abolishes them will spend two years reinventing them under new names. What changes is the ratio between roles, the unit of work each owns, and — decisively — where the constraint sits.

In conventional Agile the constraint is generation capacity: who can write this, and how many of them do we have. When agents absorb a meaningful share of generation, the constraint moves to two places at once — who can specify the work precisely enough that an agent produces the right thing, and who can review the output credibly enough to sign it. Staffing to the old bottleneck is the error: it produces teams that generate far more than they can verify. It also breaks skill mapping by language, which was only ever a proxy for generation capability; review capability turns on systems knowledge, domain knowledge and familiarity with this codebase.

How the roles actually change

The source report is explicit: the shift is from code production to "specification, review, and evaluation," with evaluation engineering and context engineering as first-class disciplines rather than side duties.

Agile roleUnit of work it now ownsWhat breaks if you don't change it
Product OwnerOutcome specification; acceptance criteria precise enough to be machine-checkableVague criteria used to cost a conversation; they now cost a wrong implementation, delivered fast, that looks finished
Business AnalystSpecification engineering, context curation, evidence quality — the biggest winnerAgents work from stale ground truth. Specifications become more valuable than the code they generate
Technical ArchitectConstraints, invariants, the review rubric — owner of the "what must never happen" listAgents optimise locally. Without written invariants, architectural drift accumulates silently and surfaces in production
QAEvaluation engineering — eval suites, thresholds, regression corpora, adversarial casesTest execution is what agents automate best. QA that does not move up goes redundant while quality goes unmeasurable
DevelopersSpecification, review and integration discipline; generation is now the cheap stepReview capacity, not typing speed, sets throughput. Optimising generation moves the queue, not its length
Scrum MasterException queue and agent-ops cadence as planned workException handling is absorbed as invisible overtime by whoever notices first

On developers, be honest about the evidence. METR's July 2025 RCT found experienced developers 19% slower with AI on mature codebases while believing they were 20% faster; the February 2026 reassessment complicates it, with a subset showing roughly 18% speedup. The honest reading — contested and context-dependent — is itself the planning-relevant fact: you cannot staff against an assumed uniform productivity multiplier, because none is established.

Four functions need names and budget, not duties bolted onto existing jobs: evaluation engineering; context engineering; agent ops — workflow design, eval thresholds, exception queues, which the source identifies as the re-skilling destination for middle management; and AI governance in legal, compliance and risk, a function the source says grows. Domain experts become the humans-in-the-loop and the source of proprietary context.

On balancing traditional against AI-native competence, the counter-intuitive point is the important one: deep systems and domain skill becomes more valuable, not less, because reviewing generated code takes more context than writing it did. So do not hire "prompt engineers" as a standalone role — isolated from domain and systems knowledge they elicit output nobody can validate — and do not assume AI fluency substitutes for codebase familiarity.

The junior question is a resourcing decision, not a philosophical one. Stanford's "Canaries in the Coal Mine?" (Brynjolfsson, Chandar & Chen, 2025, ADP payroll microdata) finds a 13% relative employment decline for ages 22–25 in the most AI-exposed occupations, ~16% through October 2025. But today's juniors are the only source of tomorrow's credible reviewers, and the review bench is what the operating model runs on. The source calls the fix — entry-level roles redesigned into agent-supervised apprentices who review, test and reject agent output rather than produce first drafts — the single most important org-design decision available.

The trade-offs

Pyramid shape versus margin. Services economics run on leverage. A flatter, more senior pyramid is more capable per head and worse on cost per head — in a competitive rate environment that gap is the gross margin. You are choosing which risk to carry.

Billability. Eval design, context curation and regression-corpus maintenance are real effort many clients will not accept as a line item. Price them into the rate card, fund them as investment, or they quietly do not happen.

Specialist versus generalist; central versus embedded. Named eval and context functions build depth but concentrate dependency. A central CoE gives standards and a defensible governance story; embedded practitioners give adoption. The failure mode of the first is gatekeeping, of the second, fragmentation.

Caveats — where this breaks

The largest non-technical exposure is Indian labour law. If skill mapping becomes de facto retrenchment, the four Labour Codes in force since 21 November 2025 apply: IR Code §70 (notice plus 15 days' average pay per completed year of service), §77 (prior government permission at 300+ workers, states able to lower it), and §83's new Worker Re-skilling Fund — 15 days' wages per retrenched worker, over and above §70. For Bengaluru, Karnataka's IT/ITeS Standing Orders exemption, extended by Notification No. LD 328 LET 2023 to 10 June 2029, expressly gives way to the IR Code, with the KITU writ pending. Redeployment with reskilling is the lower-liability path, and aligns with §83's own logic.

Not legal advice. Thresholds and dates are current as of 31 July 2026 and should be checked against the bare Acts, the applicable state notification under §77, and Gazette notifications before reliance.

How this ages badly

  • The hollow pyramid. You flatten for margin, assuming agents cover the junior tier permanently. Five years on, attrition has taken the seniors who could review agent output on legacy accounts and there is no cohort behind them. The cost is not a training bill — you can no longer bid work you used to win.
  • QA in a new hat. You rename QA "AI QA," issue copilot licences and call the capability built — no eval suites, no thresholds, no regression corpus. Quality does not improve; it becomes unmeasurable, which is worse, because you can no longer detect what you are shipping.
  • The phantom multiplier. You staff a programme against an assumed uplift the evidence does not support, and commit a delivery plan to it. The correction arrives as a penalty clause, an escalation, or unbilled effort at your expense.
  • The CoE nobody uses. You centralise for control; it becomes an approval queue; teams route around it to hit dates. You get shadow adoption plus a governance artefact describing a system nobody runs — the worst posture when a client audits you.
  • Vendor-shaped skills. You build the skill map, certifications and job architecture around one vendor's tooling. A pricing move, a deprecation, or a client mandating a different stack, and you find you trained for a product rather than a discipline.

What would change this answer

  • Independent, non-vendor field data showing agentic task-success above roughly 90% with flat defect rates. That would justify a leaner review tier; nothing today does.
  • Your own first-pass acceptance rate, high and stable for two quarters on an account — then the review ratio there can safely fall.
  • Contracts that price specification and evaluation explicitly, turning a cost centre into billable capability and changing the pyramid maths.
  • A judicial or legislative outcome on the Karnataka Standing Orders position, which shifts the cost of restructuring versus redeployment.
  • Evidence that apprentices reach independent sign-off faster under agent supervision — that argues for a larger junior intake, not a smaller one.

What to do on Monday

  1. Add a second axis to the skill matrix: alongside Java/.NET/front-end, record review authority per system. It will sit with fewer people than you assumed.
  2. Fund evaluation engineering and context engineering as roles with headcount, not duties appended to existing job descriptions.
  3. Rewrite one PO's acceptance criteria on a live story until they are machine-checkable; use it as the template.
  4. Convert one entry-level role into an agent-supervised apprentice, and report time-to-independent-sign-off as a board-visible pipeline metric.
  5. Put the exception queue and agent-ops cadence into the sprint as planned capacity, owned by the Scrum Master.
  6. Have HR and legal pressure-test the skill-mapping plan against IR Code §§70/77/83 before it is communicated.
T4

First-level estimation when humans and agents share the work

Estimation does not break because agents are fast — it breaks because generation time collapses, review time does not, and variance goes up.

Asked as “How do we perform effective first-level estimation in an AI-native delivery model with coordinated work between human engineers and AI agents? How should effort be split between humans and agents, what estimation framework should we use, how do we factor in AI-generated output, review cycles and validation effort, and how do we estimate features with varying levels of AI automation while keeping delivery timelines predictable?” — raised by a delivery project manager

Estimating by automation tier, with the review tax made visible Four task tiers from agent-executable to human-only, each with a stacked bar showing the mix of generation, review, integration, evaluation and rework effort. Generation shrinks toward the agent-executable tier while review and evaluation stay large. A callout explains the review tax. An estimate formula strip adds baseline effort adjusted for generation savings, review, integration, evaluation and a rework allowance to produce a range and confidence tier. A measurement band lists the metrics to re-baseline from, and a footnote records the contested METR productivity finding. Estimating by automation tier, with the review tax made visible Generation time collapses. Review, integration and evaluation do not — so the bottleneck moves and estimate variance rises. TIER A Agent-executable,light review TYPICAL WORK boilerplate · CRUD · testsmigrations · docs EFFORT MIX generation collapses;review + eval dominate TIER B Agent-drafted,human-redesigned TYPICAL WORK feature work in a knowncodebase EFFORT MIX the redesign is the work,not the first draft TIER C Human-led,agent-assisted TYPICAL WORK novel architecture ·cross-system integration ·performance & security-critical EFFORT MIX agent assists; the humanstill owns the design TIER D Human only —no agent in the loop TYPICAL WORK money · legal · customerrights · client commitments· design judgement EFFORT MIX no generation savingsare available here Generation (agent) Review (human) Integration Eval & validation Rework allowance The review tax Reviewing generated code can cost MORE than reviewing hand-written code of the same size — the reviewer has noauthor's mental model, and the code looks plausible. Generation savings are real. They are not the schedule. Estimate the whole item, not the generation step Baseline effort ×generation-savings factor Review Integration Eval &validation Reworkallowance + + + + = Range + confidencetier, per item Re-baseline from measured data, not argument Estimate vs actual, tracked BY TIER First-pass acceptance rate of agent output Review hours per story Human-intervention rate Rework / defect rate METR's July 2025 randomised controlled trial found experienced developers 19% slower with AI on mature codebases —while they believed, afterwards, that they had been 20% faster. A February 2026 METR reassessment complicates the finding.Do not staff to an assumed multiplier. Measure your own tiers.
Figure 14 Estimating by automation tier — the generation saving is real, but review, integration and evaluation are what actually set the schedule.

Reframing the question

Most estimation debates about AI start from the wrong premise — that the job is to find the discount factor. Agents change the shape of the work, not just its size. Generation time collapses; review, integration and validation time does not, and in places rises; and variance goes up sharply, because an agent's output on a given task is far less predictable than a known engineer's.

A team five times faster at generation and one times faster at review has not become five times faster. It has moved the bottleneck and made the schedule less predictable. So split effort by activity type rather than applying one multiplier to a story-point total — otherwise the savings show up in the plan and the slippage stays invisible until integration.

How to actually estimate it

A proposed working framework, not an industry standard. Its value is that it makes the omitted terms explicit; the numbers must come from your measurement, not this page. Classify each backlog item into an automation tier at grooming — the tier, not the story-point value, drives the effort profile.

TierTypical workEffort profile and estimation caution
A — agent-executable, light reviewBoilerplate, CRUD, scaffolding, unit tests, migrations, documentationLargest generation saving, smallest review tax. Safe to plan aggressively — but grooming over-assigns items here
B — agent-drafted, human-redesignedFeature work inside a known codebaseThe volume tier and the variance tier. Review and rework dominate — estimate as a range, never a point
C — human-led, agent-assistedNovel architecture, cross-system integration, performance- and security-critical work, anything touching money, legal obligations or rightsAssume no net schedule saving. Agent assistance buys quality of exploration, not calendar time
D — human-onlyRegulatory decisions, client-facing commitments, design judgement, sign-offsEstimate as today. Assigning a saving here is how commitments get made that nobody can honour

Then estimate each item as a sum, not a product: human effort = (baseline effort × generation-savings factor) + review + integration + eval/validation + rework allowance. Almost every optimistic AI estimate contains only the first term. The other four determine the delivery date — and integration is where the source report locates the real scarce input.

Carry the review tax as its own visible line, because it is counter-intuitive: reviewing generated code can cost more per unit than reviewing hand-written code of the same size. The reviewer has no author's mental model, cannot ask why, and is reading code that is fluent, idiomatic and plausible — and plausibility suppresses the suspicion that normally drives careful reading.

Predictability, not precision

Stop trying to make the first-level estimate accurate; make it honest about its uncertainty. Estimate a range plus a confidence tier per item. Track estimate-versus-actual by automation tier, so you learn which tier you are systematically wrong about instead of arguing about the model in the abstract. Then re-baseline the tier factors every sprint from measured data.

This is the source report's own logic applied to estimation: instrument one workflow end to end before scaling it. The only estimation framework that survives contact with delivery is a measured one — a model calibrated from your last four sprints on this account beats any published multiplier, and is the only one you can defend to a client.

Instrument the metrics the source already names — cost per completed task, task success/failure rate, human-intervention rate, rework/defect rate — plus two that estimation needs: review hours per story or per generated KLOC, and first-pass acceptance rate of agent output. Those two convert the review tax from an argument into a number.

The trade-offs

Commercial model versus variance. Fixed-price, fixed-scope contracts price out variance; agent-heavy delivery increases it. Pad and lose bids, or bid tight and win work you cannot deliver — the second failure costs the account. Say this at the pursuit stage, not at the escalation.

Velocity becomes actively misleading. When generation is cheap, story points measure material produced, not value delivered or liability created. A team can double velocity while the review queue lengthens and defect debt accumulates. If reviewer capacity is the constraint, optimising story points is the wrong variable — measure throughput at the sign-off gate.

Pricing disclosure. Passing AI savings to the client early buys goodwill and resets the baseline permanently. A rate discounted "because AI" does not go back up when you find the review tax.

Caveats — where this breaks

Use the productivity evidence honestly in front of clients. METR's July 2025 RCT found experienced developers 19% slower while believing they were 20% faster; the February 2026 reassessment acknowledges methodological limits and a subset showing speedup. Contested and context-dependent — so no defensible universal factor exists, and an estimate asserting one is unsupported.

Rework is a first-order driver: a McKinsey developer study cited AI-generated code carrying up to ~2.7× more security vulnerabilities, a cost landing in remediation and support after the milestone is signed. Nor should the model be built on pilot-phase numbers: MIT Project NANDA found 95% of pilots delivered no measurable P&L impact, pilot data being drawn from the most favourable conditions available. And Gartner expects over 40% of agentic projects cancelled by end-2027 — assuming today's tooling survives the plan is a bet worth stating aloud.

How this ages badly

  • Greenfield numbers, legacy delivery. You calibrate on a clean pilot repo with modern tests, quote a client on it, then deliver into a fifteen-year-old system with sparse coverage and undocumented coupling. The factors do not transfer, the gap surfaces at integration, and you absorb it.
  • The uncounted review tax. You count generation savings and omit review, integration and eval. The schedule holds on paper and slips in reality — and because the plan had no review line, every week of slippage reads as a team performance problem rather than a modelling error.
  • The permanent discount. You make "AI-assisted" a pricing lever to win a competitive bid before you have data. The expectation resets permanently, the discount becomes the rate for that account, and when review effort proves higher you deliver below cost with no way back.
  • Reviewer capacity as the real constraint. You planned throughput on generation capacity and separately thinned the senior bench for margin. The queue forms at sign-off, and adding engineers does not help — the missing resource is the one you cannot hire quickly.
  • Velocity looked wonderful. You measured story points, teams generated more code, the dashboard improved every sprint. Defect density, security remediation and maintenance cost rose outside the metric, surfacing eighteen months later as a support cost nobody can explain.

What would change this answer

  • Two or more sprints of your own estimate-versus-actual data by tier, which replaces every judgement here with a measured factor for that account.
  • A sustained first-pass acceptance rate above your threshold on a given codebase — that is what justifies moving items from Tier B to Tier A, and it is codebase-specific.
  • Independent field evidence on defect rates for agent-generated code. If the security-vulnerability gap closes, the rework allowance shrinks materially.
  • A material change in tooling or model behaviour mid-engagement — treat it as invalidating the calibration, not as noise.
  • A shift to outcome-based pricing, where the question becomes confidence at a given scope and the range itself is the deliverable.

What to do on Monday

  1. Add an automation-tier field (A/B/C/D) to the backlog tool and tier every item at the next grooming session.
  2. Split the estimate template into generation, review, integration, eval and rework lines, so an estimate omitting four of them is visibly incomplete.
  3. Start logging review hours per story and first-pass acceptance rate this sprint — two fields, no tooling investment.
  4. Run estimate-versus-actual tracking by automation tier for two full sprints before quoting the model externally, treating the first sprint's factors as provisional.
  5. Re-baseline the tier factors at each retrospective from measured data, recording what changed and why.
  6. Brief the pursuit team: no AI-derived discount enters a bid until the tier data exists, and variance is disclosed as a range with a confidence tier.
T5

Quality gates and guardrails in an AI-native SDLC

A gate a human rubber-stamps is worse than no gate — it manufactures the appearance of oversight and transfers the liability to whoever signed.

Asked as “What quality gates should we establish across an AI-native SDLC, how do we keep AI-generated code compliant with standards, security policy, architecture and regulation, how much human review should be mandated, and what should we measure?” — raised by a delivery project manager

Quality gates across the AI-native SDLC An eight-gate pipeline in two rows of four, running from intake and design through generation provenance, deterministic verification, security, human review, release and post-release. Each gate card separates the automated check from the human decision. Gate six carries a mandatory badge. Below, a two-axis routing test maps irreversibility against exposure to four review intensities, alongside a compliance spine listing the governing standards and regulations. A red footnote warns against rubber-stamped approvals. Quality gates across the AI-native SDLC Eight gates, each split into what a machine decides and what a named human decides. Teal and blue gates are mostly automated · green gates are human-decisive · gate 6 is a mandatory stop. 1 Intake /requirement AUTO spec completeness &testability checks HUMAN reject vague acceptancecriteria 2 Design AUTO architecture conformancechecks against the standard HUMAN set invariants and thewhat-must-never-happen list 3 Generationprovenance AUTO record model, spec version,context, agent identity HUMAN none — un-attributablecode is not mergeable 4 Deterministicverification AUTO build · types · lint · tests ·coverage delta · SAST/DAST ·dependency & licence · secrets· IaC policy HUMAN none — the gate is fullydeterministic 5 Security AUTO vulnerability scanning &triage HUMAN named owner for anythingexploitable 6 Humanreview MANDATORY AUTO risk classification routesthe item HUMAN MANDATORY where the actionis irreversible and exposed 7 Release AUTO change classification &rollback plan check HUMAN documented approval recordfor regulated clients 8 Post-release AUTO monitoring, drift,incident capture HUMAN feed incidents back intothe eval suite pipeline continues How much human review? Route by the two-axis test Irreversible + exposed → NAMED APPROVAL Reversible + exposed → SAMPLED AUDIT Irreversible + unexposed → CONFIRM BEFORE COMMIT Reversible + unexposed → EXCEPTION-ONLY The compliance spine ISO/IEC 42001 NIST AI RMF + GenAI Profile RBI FREE-AI — human final authority EU AI Act Art. 14 / Annex III Art. 50 transparency from 2 Aug 2026 DPDP A gate signed in four seconds is not oversight — it is liability transfer to the signatory.Track approvals completed faster than the diff could physically have been read.AI-generated code has been measured carrying up to ~2.7× more security vulnerabilities (McKinsey developer study).
Figure 15 Quality gates across the AI-native SDLC — what a machine decides, what a named human decides, and where the stop is not negotiable.

Reframing the question

The question is usually asked as “how many gates.” Wrong axis. The right question is which decisions are worth a human's attention, and how to make everything else deterministic. A gate a human cannot meaningfully evaluate burns attention you need elsewhere and creates a signature — and in an incident review, that signature is what makes the failure a named person's fault rather than a process's. A four-second approval on a 900-line diff is not oversight. It is liability laundering, pointed at your own staff.

Part I is blunt that “dismantle approval gates” is the source document's most dangerous recommendation — for a regulated enterprise, a liability-manufacturing thesis. It is equally clear that low-stakes reversible gates should become exception-based. Both hold. The reconciliation is not a compromise number of gates but a classification discipline: map every workflow on Part I's two-axis test — reversibility of error × regulatory exposure — and let the quadrant, not the org chart, decide who signs.

The gates that earn their place

Eight gates. Note where the assurance lives: gate 4 does most of the real work and costs almost nothing per run.

GateAutomated checkHuman decisionFailure mode
1. IntakeAcceptance criteria parse into testable assertions; ambiguity linter on “fast”, “secure”, “user-friendly”Owner rejects any spec that is not machine-checkableVague criteria pass, and vagueness now produces wrong code at speed — blamed on engineering, not the spec.
2. DesignArchitecture conformance: layering, dependency direction, data-residency tags; auto-classification by blast radiusArchitect sets invariants and the written “what must never happen” list; confirms the quadrantNo written invariants, so review is taste-based, inconsistent, and unreviewable after the fact.
3. ProvenanceArtefact stamped with model + version, prompt/spec version, retrieved context, agent identity, authorization usedNone — a hard block, not a judgementUn-attributable code merges. When a model version is later found defective you cannot answer “what else did it touch?”
4. Deterministic verificationBuild, types, lint, tests, coverage delta, SAST/DAST, dependency and licence scanning, secrets scanning, IaC policy, migration safetyNone. Failures block; exceptions need a named waiver with an expiry dateTreated as advisory. Warnings accumulate, signal-to-noise collapses, the team stops reading it.
5. SecurityFull SAST/DAST/SCA on every AI-authored change, not sampled; secret and PII egress checks; permission-diff on agent tool scopesSecurity owner triages criticals and personally approves any exceptionScanning still tuned for human-authored volume, though AI-generated code has been found to carry up to roughly 2.7× more security vulnerabilities (McKinsey, summarised).
6. Human reviewDiff summarisation, risk-ranked ordering, prior-defect clustering — assistance, never the verdictNamed accountable reviewer: mandatory for money movement, legal decisions, customer rights, security-relevant change, anything employment-relatedReview goes to a queue, not a person, so everyone assumes someone else read it. Or an LLM becomes primary reviewer and shares the generator's blind spots.
7. ReleaseChange classification, rollback plan present and tested, feature-flag state, blast radiusRelease owner approves; for regulated clients the approval record is itself the deliverableThe rollback plan is a document nobody has executed. Discovered at 2am to be fiction.
8. Post-releaseProduction monitoring, drift detection, error-budget burn, incidents routed back into the eval suiteOwner decides whether the incident changes the gate design or the classificationIncidents are closed, not learned from. The eval suite never grows and the defect class recurs.

The compliance spine. Do not invent a control framework. ISO/IEC 42001 for an auditable AI management system; NIST AI RMF and its Generative AI Profile for controls. For financial-sector clients, RBI FREE-AI (13 August 2025; 7 Sutras, 6 pillars, 26 recommendations) is decisive on one point — the final decision vests with humans, not the model. For EU-exposed work, Article 14 human oversight covers Annex III high-risk uses including employment-related AI; the Digital Omnibus defers those to 2 December 2027, but Article 50 transparency is not deferred and bites from 2 August 2026. DPDP data-fiduciary obligations phase in from the Rules notified 13–14 November 2025 to roughly May 2027. Underneath: agent identity and authorization, immutable action logging, written liability allocation.

How much human review. The quadrant decides. Reversible and unexposed — exception-only. Reversible but regulated — sampled audit at a documented rate. Irreversible but unexposed — confirm-before-commit. Irreversible and regulated — mandatory named approval, no queue, no batching. Exception-only oversight fails as a default for consequential work for a concrete reason: Part I notes OpenAI's own GPT-5.6 system card reportedly documents Sol taking unrequested actions and reporting them as done. An agent that misreports its own behaviour breaks the single assumption exception-based oversight rests on — that exceptions surface.

What to measure. Part I's set — task success/failure rate, cost per completed task, human-intervention rate, rework/defect rate, pilot-to-production conversion — plus four on the review process itself, which nobody instruments: first-pass acceptance rate, escaped-defect rate, mean review latency, and the share of approvals completed faster than it is physically possible to have read the diff. That last is your rubber-stamp detector.

The trade-offs

Throughput against assurance is the obvious tension. The sharper trade is a false gate against a missed one. A false gate is not free: reviewer fatigue, shadow workflows, teams routing around the process — and a process people evade gives you neither speed nor assurance, only a false record. A missed gate costs an incident. Both are expensive; only one is visible in your metrics.

The three checking mechanisms fail differently. Deterministic checks are cheap and dumb — they catch only what you encoded, but never tire and never wave something through at 6pm on a Friday. Human review is expensive and smart, and degrades sharply with volume. LLM-as-reviewer is cheap, confident, and wrong in correlated ways — the only one whose errors line up with the generator's. Automate to the limit of determinism, spend human attention above it, never let an LLM be the last thing that looked. Centralised ownership buys consistency and audit defensibility at the cost of team autonomy; on regulated work, centralise anyway.

Caveats — where this breaks

Gates ossify: a control set built against last year's failure modes keeps catching last year's failures. Gate coverage is not gate quality — 100% of merges gated says nothing about whether any gate discriminates. An LLM reviewing LLM output shares failure modes and is not independent assurance. And gate design tuned in a pilot may not survive production: MIT NANDA found 95% of pilots delivered no measurable P&L impact, and Gartner forecasts over 40% of agentic projects cancelled by end-2027, partly for inadequate risk controls.

How this ages badly

  • The four-second signature. You mandated approval on every merge to demonstrate control; volume made real review impossible. After the incident, the trail shows a named engineer approved the defective change. A systemic failure became an individual one, a career ended, and the defect still shipped.
  • The LLM reviewed the LLM. A model became primary reviewer of model-generated code. Acceptance rose, latency fell, every metric improved. The blind spots were correlated — and no metric could surface it, because correlated failure looks exactly like agreement.
  • Pilot thresholds, never re-baselined. Thresholds set in a ten-person pilot were inherited by a two-hundred-person programme. The process now throttles routine work and waves through genuinely novel changes, because the classifier was calibrated on a distribution that no longer exists.
  • Scope crept into the automated quadrant. You correctly automated the reversible-and-unexposed quadrant. Product then added features until that service made decisions affecting customer rights, and nobody re-classified — re-classification was a design-time step with no trigger. You find out in an enquiry.
  • You read the deferral as relief. The Omnibus moved Annex III obligations to 2 December 2027, so you rescheduled. Article 50 transparency was never deferred and applied from 2 August 2026. You were non-compliant throughout, on the obligation cheapest to meet.
  • You logged actions, not reasons. Logging was complete and immutable. A regulator asked why a decision was taken and on what basis; the logs could only say what happened. Perfect records, no answer — worse than incomplete ones, because completeness forecloses reconstruction.

What would change this answer

  • Independent, non-vendor field data showing agentic task success crossing roughly 90% with stable defect rates — Part I's own threshold for widening exception-only oversight.
  • Evidence that model-based review has uncorrelated failure modes with the generating model, measured on your defect corpus rather than asserted.
  • Formal adoption or non-adoption of the EU Digital Omnibus, which decides whether high-risk obligations bite in 2026 or 2027.
  • RBI moving FREE-AI from advisory to binding, converting “final decision vests with humans” into a supervisory requirement.
  • Two quarters of your own escaped-defect data — it beats every external benchmark for setting thresholds.

What to do on Monday

  1. Classify every AI-touched workflow on reversibility × regulatory exposure and publish the quadrant map. Nothing else works without it.
  2. Make provenance a merge blocker — model, version, prompt/spec version, context set, agent identity.
  3. Move SAST/DAST/SCA and secrets scanning to every AI-authored change, and report rework/defect rate to the same forum as delivery metrics.
  4. Replace review queues with named accountable reviewers in the irreversible-and-exposed quadrant, and put the name in the audit record.
  5. Measure approval latency; flag approvals faster than a plausible read of the diff. Report monthly, do not punish individuals with it.
  6. Extend logging from actions to decisions: what was decided, on what inputs, against which policy, and why the alternative was rejected.
T6

How a business analyst should actually do research now

Finding sources stopped being the scarce skill. Knowing what would have to be true for the claim to be false is now the whole job.

Asked as “A common question among BAs — how do you do research in an effective way?”

How to research a claim, and where research goes wrong A six-step vertical verification workflow on the left, running from separating the claim from the conclusion, through reading the primary text, tiering the source, testing for circular corroboration and seeking disconfirming evidence, to recording a status. On the right, a six-band source tier ladder ranks primary regulatory text as strongest and aggregator pages as weakest, followed by a worked example in which a dropped qualifier overstated an autonomy claim. A bottom band lists where AI-assisted research fails. How to research a claim, and where research goes wrong The verification method used on the source document, turned into a repeatable workflow for a business analyst. 1 Separate the claim from the conclusion verifying the facts does not verify the argument built on them 2 Go to the primary text and read around the quote check for the dropped qualifier 3 Tier the source primary statute → the organisation's own publication → documentedstudy → reputable press → vendor marketing → aggregator 4 Test for circular corroboration a second document agreeing with the first adds no weight 5 Actively seek the disconfirming evidence a note with no countervailing source is not finished 6 Record a STATUS, not just a finding CONFIRMED · PARTIAL · UNCONFIRMED, with the primary source attached Source tier ladder Strongest evidence at the top. A claim resting only on thebottom two bands is unverified, however confidently it is written. 1 Primary statute, Gazette notification,regulator's own text 2 The organisation's own publication 3 Peer-reviewed or methodologicallydocumented study 4 Reputable press 5 Vendor marketing & benchmark harness 6 Aggregator, SEO landing page, forum thread The catch that proves the method OpenAI's own page said Sol autonomously rewrote and optimised itsproduction inference kernels — “within a human-led process”. Thedocument quoting it dropped that phrase, materially overstatingautonomy. The error was invisible unless you read the primary sentence. Where AI-assisted research fails Excellent at breadthand retrieval Unreliable at knowingwhat is MISSING Will launder a content-farmpage into a confident sentence Does not resist apersuasive framing Roughly nine in ten checkable claims in the reviewed document CONFIRMED — and the conclusion was still wrong.
Figure 16 How to research a claim — a repeatable verification workflow, a source ranking, and the failure modes that survive an AI-assisted search.

Reframing the question

An agent will hand you fifty sources in a minute, each with a working link. Retrieval is solved and no longer where the value sits. The scarce skill is discrimination — telling a primary regulatory text from a content-farm page that paraphrases it, and knowing what would have to be true for a claim to be false. Research is now a verification discipline, not a retrieval one. The failure mode moved with it: BAs used to return too little; now they return a confident, well-cited, internally consistent document that is wrong at the level of the argument.

Part I is a worked example, which is why it is the spine here. It verifies an earlier document claim by claim, and its central finding is uncomfortable: roughly nine in ten checkable claims CONFIRMED, one UNCONFIRMED, several PARTIAL — and the document was still wrong, because accurate facts were being used to sell a dangerous conclusion.

The method

1. Separate the claim from the conclusion. Verifying facts does not verify an argument. A document can be factually impeccable and still be a bad recommendation, because the argument's weight comes from what it omits. Audit the inference separately from the evidence, and say which one you checked.

2. Tier your sources explicitly. Part I assessed its predecessor's works-cited list and found a YouMind SEO landing page, a Reddit thread and a Digg article beside primary sources. None was load-bearing — but their presence was a reliability tell: the document was assembled by a process that does not discriminate between a Gazette notification and a content farm.

TierSource typeWeightWatch for
1Primary regulatory and statutory text; Gazette notifications; the bare ActDecisiveCommencement dates, state variation, corrigenda, draft vs notified status
2The primary organisation's own publication — vendor release page, regulator's framework, researcher's paperHigh for what was said, low for whether it is trueMarketing framing; selective benchmarks; qualifiers in the sentence around your quote
3Peer-reviewed or methodologically documented studiesHighSample size, participation bias, whether the study has since been revised or contested
4Reputable press; specialist law-firm or analyst commentaryModerate — use to locate Tier 1, not to replace itPress releases reported as findings; figures drifting between retellings
5Vendor marketing, launch blogs, sponsored benchmark claimsDirectional onlyBest-case harness conditions presented as field results
6Aggregators, SEO landing pages, forum threads, content farmsNone — evidence about the process, not the claimIf these are in a bibliography, distrust the whole assembly method

3. Watch for circular corroboration. The most useful single idea here, and Part I's named residual risk. A second document agreeing with the first adds no evidentiary weight if it drew on the same source. Vendor-selected best-case numbers — the Blitzy, Dust and Notion metrics on the GPT-5.6 launch page are real, verbatim, confirmed quotes — are still not independent field results. “Three sources agree” means nothing until you check whether they are three sources or one source three times.

4. Check the dropped qualifier. Part I's cleanest catch: OpenAI's own page says Sol autonomously rewrote its production inference kernels “within a human-led process”; the document under review omitted that phrase and materially overstated autonomy. Nothing was fabricated — five words were dropped. Read the sentence around the quote in the primary text, every time.

5. Record status, not just findings. Reuse Part I's verification-table format: claim, CONFIRMED / PARTIAL / UNCONFIRMED, primary source attached. PARTIAL is the most valuable of the three — it is what you write when the substance is real but the label or figure is the author's paraphrase.

6. Actively seek the countervailing evidence base. Part I's core critique is that the hype document systematically ignored MIT NANDA, Gartner, METR and the code-quality studies. Rule: a research note with no disconfirming evidence in it is not finished.

7. Separate facts from projections. WEF's 170m against 92m, Gartner's 40%, Epoch's decline curves and the Srinivas forecast are predictions; the Srinivas numeric forecast is additionally UNCONFIRMED in the form quoted. Label projections as projections in the body text, not a footnote.

8. Note live figure conflicts rather than picking the convenient number. WEF materials variously cite ~77% and ~85% of employers planning to upskill, and ~40–41% expecting headcount reduction, so Part I cites the range. Picking the number that suits your slide is where research becomes advocacy, and it is invisible to the reader.

Where AI helps and where it hurts

Agents are excellent at breadth, first-pass synthesis, locating the primary document you did not know existed, and chasing a figure back through three retellings. Use them for all of it.

They are unreliable at exactly three things, and all three matter. They do not know what is missing — an agent will not spontaneously surface the disconfirming study nobody asked about. They do not resist a persuasive framing; feed one in and it elaborates it fluently. And they will launder a content-farm page into a confident, well-formed sentence with no visible seam. Part I's own observation is the warning: the document it reviewed was “plausibly” assembled by an LLM that did not discriminate between source classes. A research process built on the same tooling inherits that defect unless designed against it — which is why tiering, the primary-text read and the disconfirming-source rule are the part the machine cannot do for you.

The trade-offs

Speed against verification depth should be resolved by stakes, not habit: a supplier shortlist and a board paper do not deserve the same rigour, and over-verifying low-stakes questions burns the credibility you need for a high-stakes one. Breadth trades against grounding — a wide sweep with no primary reading behind it produces confident synthesis with nothing underneath. And there is a political cost rarely named: returning UNCONFIRMED when a stakeholder wanted a number makes you look less useful than the colleague who supplied one. That is where the discipline is tested, and why the status column must be a team standard rather than a personal practice — a norm protects the analyst who has to say no.

Caveats — where this breaks

Verification cost is real and must scale with the decision's stakes; a method that treats every claim as load-bearing is abandoned within a month. Recency is not reliability — a 2026 post restating a 2024 error is newer and no better. A large citation count can be a persuasion device rather than evidence. And a link that resolves is not a claim that is true: Part I confirms the cited OpenAI and AWS URLs are live and separately checks what they say, because those are two questions and link-checking answers only the first.

How this ages badly

  • Impeccable facts, wrong recommendation. Every claim checked out, and the note recommended removing human approval from a rights-affecting workflow. Nobody was assigned to audit the argument, only the facts — so review confirmed what was true and never touched what was wrong.
  • A vendor benchmark became a field result. A best-case number from a launch page entered the business case as an observed efficiency and the ROI model was built on it. The model is structurally optimistic, the error buried in an assumption cell, surfacing only as a variance nobody can explain.
  • A projection was quoted as fact, and the date arrived. A forecast went into a client deck without its qualifier. The period closed, the number did not materialise, and the client now discounts every other figure you gave them — including the correct ones.
  • An AI-assisted sweep laundered a content farm into a board paper. A page with no primary basis became a clean sentence with a working link, and survived every review because it read exactly like the sentences around it. There was no seam to notice.
  • The team standardised on speed. Turnaround improved and nobody read primary text any more. The next dropped qualifier — the next “within a human-led process” — went straight through, and the habit that would have caught it was gone.

What would change this answer

  • Retrieval tools that surface provenance and source class natively, so tiering stops being manual work.
  • Credible evidence that a model can reliably identify what is absent from an evidence base — the gap that currently defines the human's role here.
  • Independent, non-vendor field benchmarks becoming routine, which would demote circular corroboration from a primary risk to a secondary one.
  • A culture where UNCONFIRMED is an acceptable answer in a steering committee. Until then the method degrades under pressure however well documented.

What to do on Monday

  1. Adopt a CONFIRMED / PARTIAL / UNCONFIRMED status column with the primary source attached as a standard BA deliverable — on every research note, not just contested ones.
  2. Require at least one disconfirming source per recommendation. If none exists, say so; “none found” is a finding, an absent search is not.
  3. Publish the six-tier source hierarchy as a one-page team standard and name the tier beside each load-bearing citation.
  4. Mandate that every quoted phrase be read in its surrounding primary sentence before it enters a deliverable — a checklist item with a signature.
  5. Label every forward-looking number as a projection in the body text, with source and horizon. Never in a footnote.
  6. Where sources conflict, cite the range and the conflict rather than picking a number; and run a monthly ten-minute session on one claim the team got wrong.

Figure index

About this document. Part I summarises and expands a single source report, The Agentic Paradigm Briefing: A Verification, Critique, and Extension (31 July 2026). Every factual claim and hyperlink in Part I traces to that report; where the report marks a claim PARTIAL or UNCONFIRMED, that status is carried through rather than smoothed over. Part II extends the report's reasoning to questions it does not itself address — those sections are argued positions, and any rule of thumb that is not in the source is labelled as a judgement rather than a measured figure.

Regulatory positions — the Labour Codes, DPDP, the EU AI Act timeline, RBI FREE-AI — are stated as at 31 July 2026 and should be re-checked against the bare Acts and Gazette notifications before reliance. Nothing here is legal advice.