The research behind Radiq Case study

We read the research. Then we rebuilt the product.

We were building v1 - a contextual brain for product teams. Then we spent months with the most serious AI forecasts in the field. They taught us one thing: the bottleneck is not how capable the agent is. It is whether anyone can tell when it is wrong. So we rebuilt Radiq around verification. Every automation ships with receipts. When it cannot verify a number, it stops and asks.

For you The hard part of an AI doing your work is not making it capable - it is knowing when it is wrong. That is the part Radiq is built around.
01

Two papers, fifteen months apart.

The AI Futures Project publishes unusually specific, dated narratives about how AI development might go. Because they commit to detail, they are easy to hold to account. We studied them because they are the most serious forecasters in the field - and because what happened between their two papers explains why Radiq exists.

In April 2025 they published AI 20271 - a forecast. In July 2026 they published AI 2040: Plan A2, which opens by telling you it is not one.

April 2025 · a forecast

AI 2027

Kokotajlo, Alexander, Larsen, Lifland & Dean

A month-by-month scenario running to superhuman AI research agents. Its own first act does not stall on intelligence - it stalls on adoption. The agents in it are capable well before they are trusted.

What will happen. Read AI 2027 →
9 July 2026 · a recommendation

AI 2040: Plan A

Kokotajlo, Larsen, Lifland, Dean and two others

Ninety pages proposing that the US and China negotiate limits that push superintelligence from roughly 2030 out to 2040 - buying time to solve alignment and distribute benefits.

What should happen. Read Plan A →
15
months between them. The authors did not revise their capability estimates downward. They changed what kind of document they were writing.

“It's called Plan A because it's a recommendation, not a prediction.”

AI 2040: Plan A, opening section2

A prediction is a claim about capability. A recommendation is a claim about governance - about who is allowed to do what, when, and under what checks. The most aggressive published forecasters in the field spent fifteen months and moved from the first kind of document to the second.

They are arguing at the scale of treaties. We are not. But the observation underneath is the same one, and it holds at every scale down to a single laptop: the binding constraint is not how capable the agent is. It is whether anyone can tell when it is wrong.

The International AI Safety Report 2026 - Bengio as chair, ninety-two authors - exists to synthesise exactly this question for governments.3 An entire institutional apparatus is now pointed at verification. Almost none of it is pointed at the desk where the work actually happens. That is the gap Radiq is built for.

02

The error is reliable. So we built the check.

If the best forecasters in the world cannot verify a superhuman AI, what chance does a regular worker have verifying the automations that run their daily job? The researchers did not say AI will fail. They said we will not know when it fails - and that is a much harder problem.

METR measures how long a task an AI can complete - the "time horizon." It is the best quantitative handle anyone has on the capability curve, and the curve is steep: under their revised methodology, the frontier horizon has been doubling roughly every 131 days since 2023.4

But two things in the fine print change the picture. First: it is a median at a 50% success rate. A coin flip. METR also publishes an 80% threshold, which is materially shorter. Almost no work that matters runs at 50%. Second: the measurement itself is uncertain by a factor of four.

One model's time horizon, with its confidence interval

Claude Opus 4.5, at a 50% success rate. METR Time Horizon 1.1, January 2026.

The same numbers, as published.
QuantityValue
Point estimate, 50% success320 min
Lower bound170 min
Upper bound729 min
Width of the interval≈ 4.3×
Frontier doubling time, post-2023131 days
The interval spans 170 to 729 minutes. One working day sits comfortably inside it. Read at the lower bound the model is a useful assistant; read at the upper bound it runs your afternoon unattended. The published number does not distinguish between those two worlds - and this is the best-instrumented measurement the field has. Source: [4].

If the frontier's own capability is known only to within a factor of four, at a 50% success rate, then no amount of capability progress tells you whether a given action is safe to ship. That has to be established per action, at the moment of acting, against what actually happened. Which is a different piece of engineering than a better model - and nobody gets it for free when the next one lands.

Meanwhile the market is already crying out for this. The failures are well documented, and they cluster on verification, not capability.

40%
of agentic AI projects will be cancelled by the end of 2027 - on escalating cost, unclear business value and inadequate risk controls
Gartner, June 20255
95%
of enterprise GenAI pilots produce no measurable P&L return, against $30–40B of spend
MIT Project NANDA, The GenAI Divide, July 20256
87%
of developers worry about the accuracy of AI agents; 81% about the privacy of data handed to them. The top two concerns, from the same survey
Stack Overflow Developer Survey 2025, 49,000+ respondents7
19%
slower with AI tools - while the same developers believed they had been 20% faster. Nobody in the study could feel the difference
METR randomised controlled trial, July 2025 · 16 developers, 246 tasks. METR now labels this result historical8

The last one carries a caveat we are keeping in place: METR has since changed its experiment design and no longer presents that number as current.8 Cite it as a live measurement of today's tools and you are misusing it.

What it still demonstrates is the thing no revision touches: the participants could not perceive their own slowdown. Their self-reports were inverted relative to the stopwatch. Whatever intuition tells you about whether an agent is helping, it is not measurement.

The MIT finding points the same direction. Their diagnosis of the 95% is not that the models are bad - it is brittle workflows, weak contextual learning, and misalignment with the way people actually work day to day.6 The pilots don't fail at the model layer. They fail at the seam between the model and the real job.

And Gartner names inadequate risk controls as a top-three cause of cancellation, alongside cost and unclear value.5 Not inadequate capability.

So the demand is not in question. The trust to meet it is. That is exactly what Radiq is for.

Simulated data: Kelso Ltd and INV-4471 are fictional. The interface is real.

03

What we built instead.

An agent that runs your work has to know your work. That means watching. And watching, done the ordinary way, is how these products die - not in the market, but inside the company that deploys them. We chose a different architecture from the start.

In June 2026, Meta put software on US employees' laptops that captured mouse movements, clicks and keystrokes to train workplace AI. Staff circulated petitions, taped flyers in conference rooms, and Meta rolled parts of it back within weeks.10

The lesson we took from that is architectural, not moral. The sensitive asset in an observational product is the behavioural stream. A company with essentially unlimited engineering resources tried to centralise it and could not make it hold - not because the technology failed, but because the people being observed refused.

Which puts a hard constraint on anything built in this category: the observation layer cannot leave the device. Not as a privacy feature bolted on afterwards. As the precondition for the product being installable at all.

So this is what Radiq does, at the level of one worker's laptop:

Invoice · INV-4471
$18,400
Kelso Ltd · received today, 08:41
Purchase order · PO-2208
$16,900
Approved 14 Mar · signed by Marcus
Held at step 3 of 5

“I can't verify this.”

It stopped before step 4. Nothing was logged, nothing was sent. The two documents disagree by $1,500, and Radiq won't record a number it can't confirm, so it asked you which one is right.

That is the verification layer in action. Not a dashboard an admin checks later. Not a confidence score in a log file. A hard stop, at the moment of acting, because the number could not be confirmed. The same principle the AI Safety Report argues for at treaty scale - implemented at the scale of one invoice on one laptop.

The developer survey makes the shape of the demand explicit. The top two concerns about AI agents were accuracy at 87% and data privacy at 81%7 - roughly half of respondents said they had no plans to adopt agents, citing exactly those reasons. Those are not two separate objections to be traded off. They are one product requirement: check your work, and don't send my day anywhere.

That is what we built.

04

How these papers changed our build.

Four months ago we shipped Radiq v1 - a contextual brain for product teams. It had an ingestion pipeline across Slack, Jira, Confluence and meetings, a Neo4j knowledge graph, and an MCP server into Cursor and VS Code. We deployed it with one design partner and a second under signed LOI. Then we watched what users actually did.

They were lukewarm on the copilot layer. But they loved the few flows we automated end-to-end - the parts where a piece of finished work arrived without them asking for it. Across interviews one sentence kept repeating, even from people technical enough to build automations themselves: "I could build it, I just don't want to spend hours doing R&D on my own job."

That same summer, Ritvik was inside Neo4j watching well-resourced teams fail at automation for the identical reason. The bottleneck was not capability. It was that nobody can describe the long process they never wrote down.

Then we read these two papers. They gave us the language for what we were seeing. The researchers had moved from predicting capability to demanding verification - because capability was never the binding constraint. Neither is it for a knowledge worker's daily automations.

So we kept the engine - ingestion, knowledge graph, pattern detection - and killed the product around it. Then we rebuilt around three decisions:

01 Ambient discovery - never ask the user to describe their work 02 Verification-first - evidence before proposing, hold when uncertain 03 Local by default - the data stays on-device, five-minute install

The result is a desktop app that watches how you actually work, finds the loop you repeat without noticing, and proposes to take it over with the count attached. Approve it once and it runs on its own. When two numbers disagree, it stops and asks. Nothing it sees ever leaves your laptop.

Ninety early-access requests in the first three days told us the new aim is closer to true. We are onboarding our first design partners now.

05

What we learned, what we built, and how you check it.

Research that doesn't change the build is decoration. Here is the mapping - what we read, what it forced us to do, and how a user can confirm we actually did it rather than taking our word.

What we learned
What we built
How to see it
Horizons are 50% medians with 4× intervals[4]
Reliability is established per action, not inherited from the model. Every action is treated as a claim to be checked against what was observed.
The engine's verify step runs before anything ships, on every run.
Participants couldn't perceive their own slowdown[8]
Self-report is not evidence, so nothing relies on the user's sense of whether it helped. Counts are kept and shown.
Every discovery carries the occurrence count behind it - 14×, not "often."
Inadequate risk controls cancel projects[5]
It stops rather than guessing. When two sources disagree it holds at that step and returns the conflict instead of recording a number it can't confirm.
The hold at step 3 of 5, with both documents attached and nothing written.
Pilots fail on brittle workflows and misalignment with real work[6]
Nobody describes a process up front. It learns the loop from the work as performed, including the steps people skip.
The 13-of-14 step in your own loop. The gap is reported, not smoothed over.
Centralised behavioural capture provokes refusal[10]
The observation layer never leaves the device. No server holds your day, so there is no dashboard for anyone else to open.
Cut your network and watch observation, counting and discovery keep running.
Accuracy 87% and privacy 81% are one requirement[7]
Glass-box by default. What it saw, what it decided, and what it declined are all inspectable - including the things it chose not to raise.
The weekly digest lists what it refused to send, and what you said no to.
7.6 hours a week, €10.7M per thousand staff[9]
Target the undocumented repeated loop specifically - the part that never becomes a ticket, a spec or an automation because nobody ever wrote it down.
The first discovery it surfaces is one you never described to it.
06

References.

Every claim on this page was checked against its primary source on 27 July 2026. Where a source carries a caveat, the caveat is on the page too. Where we could not verify something, it isn't here.

  1. Kokotajlo, D., Alexander, S., Larsen, T., Lifland, E. & Dean, R. AI 2027. AI Futures Project, April 2025. ai-2027.com
  2. Kokotajlo, D., Larsen, T., Lifland, E., Dean, R. and others. AI 2040: Plan A. AI Futures Project, 9 July 2026. blog.aifutures.org/p/ai-2040-plan-a · also at ai-2040.comQuoted on this page: “It's called Plan A because it's a recommendation, not a prediction.”
  3. Bengio, Y. (chair) et al. International AI Safety Report 2026. 92 authors, submitted 24 February 2026. arXiv:2602.21012
  4. METR. Time Horizon 1.1, 29 January 2026. metr.orgPost-2023 frontier doubling time 131 days. Claude Opus 4.5 time horizon 320 min, 95% interval [170, 729], at a 50% success rate. Methodology introduced in Measuring AI Ability to Complete Long Tasks, 19 March 2025.The 80% threshold referred to in §02 is published in that paper's interactive chart and dataset rather than in its prose. We have therefore not put a specific 80% figure on this page.
  5. Gartner. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Press release, 25 June 2025. gartner.comCauses given: escalating costs, unclear business value, inadequate risk controls. The accompanying investment figures come from a January 2025 poll of 3,412 webinar attendees.
  6. MIT Project NANDA. The GenAI Divide: State of AI in Business 2025, July 2025.95% of pilots show no measurable P&L impact against $30–40B of enterprise spend; 5% capture the value. Based on 52 executive interviews, 153 leader surveys and 300 public deployments. Stated causes: brittle workflows, weak contextual learning, misalignment with day-to-day operations.
  7. Stack Overflow. 2025 Developer Survey - AI section. 49,000+ responses from 177 countries. survey.stackoverflow.co/2025/ai87% concerned about AI agent accuracy; 81% about security and privacy of their data. Roughly half report no plans to adopt agents; 23% use them at least weekly.
  8. METR. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. metr.org16 developers, 246 tasks. 19% slower with AI access; participants estimated they had been 20% faster, having expected 24%.Caveat, carried on the page: METR labels this result historical and has since revised its experiment design (update, 24 February 2026). It should not be cited as a measurement of current tools.
  9. Frends / Sapio Research. State of Integration & AI 2026. state-integration-ai.frends.com611 senior IT and business decision-makers at organisations of 201+ staff, across Denmark, Finland, Germany, the Netherlands, Norway and Sweden. 7.6 hours per employee per week lost to manual data work; €10.7M mean annual cost per 1,000-employee organisation.
  10. HR Grapevine. Meta scales back AI employee monitoring tool after staff backlash, June 2026. hrgrapevine.comSoftware on US employees' laptops captured mouse movements, clicks and keystrokes to help train AI systems. Staff petitions described the company as an “Employee Data Extraction Factory.” Meta rolled back parts of the programme; internal memo first reported by Reuters.
  11. SHRM / Maurer, R. AI Surveillance in the Workplace Linked to Employee Resistance, Turnover, 23 August 2024.Listed for transparency. An earlier version of our landing page attributed two figures (“54% would consider leaving”, “59% say tracking damages trust”) to this article. It contains neither; its only statistic is a 2023 i4cp finding that 6% of large companies use surveillance tools. Both figures have been removed rather than re-sourced.
  12. Radiq. Design spec: overview video and case study page, 27 July 2026.The citation ledger behind this page, including the two claims that failed verification and what replaced them. Held in the repository at docs/superpowers/specs/.

Capability is no longer the question. Verification is. Radiq is our answer.

Radiq watches the work you actually do, finds the loop you repeat without noticing, and proposes to take it over with the count attached. It stops when it can't verify. Nothing it watches leaves your machine.