Codex alone vs Codex + Rose for Agents¶
Date: 2026-10-01 · Ticket: IX-5364 · Data: 39 buyer questions on 15 web search and scraping APIs (backend/apps/r4a_mcp/bench/questions.json) · Transcript: every answer, side by side
TL;DR. After three rounds, Codex + Rose is more accurate than Codex with web search. On 50 held-out questions written after the last change and fact-checked blind on the vendors' pages, it is wrong on 1 answer against 3 (2% against 6%), fully right on 86% against 76%, and answers the whole question more often (1 partial answer against 6). The price is time and tokens: 46 s against 39 s median and 243k against 138k tokens, because Codex now checks the deciding facts on the vendor's page. What made the difference: audited sheets (434 corrections), pricing rules that say how a price was computed, and an instruction to verify the facts that decide the answer.
Question¶
Does a coding agent answer developer tool-buying questions better with the Rose for Agents connector than with web search alone? The questions cover pricing, features and compliance, getting started (MCP servers, SDKs), comparisons, listings and which vendor does what.
Verdict¶
Rose makes Codex more accurate on tool-buying questions, at a cost in time and tokens.
| Round (held-out questions) | Wrong: alone / with Rose | Fully right: alone / with Rose | Median time | Median tokens |
|---|---|---|---|---|
| 2 (40) | 3 / 5 | 78% / 75% | 38 s / 31 s | 136k / 124k |
| 3 (50) | 3 / 1 | 76% / 86% | 39 s / 46 s | 138k / 243k |
Round 2 had the richer catalog but no gain; round 3 added the audit pass, explicit pricing rules and verification of the deciding facts, and turned it into a gain. The samples are small: 3 against 1 wrong answers is a direction, not proof; the gap in fully right answers (43 against 38) and in partial answers (1 against 6) is the clearer signal.
Against the phase 0 gate: Rose now beats web search on correctness, but no longer on time and cost. The gate as written (match deep research on correctness, clearly beat it on time and cost) still needs the deep-research arm, which runs by hand, and real users.
Decisions it supports:
- Ship the instruction snippet (
bench/AGENTS.rose.md) with the connector's install steps; without it Codex never calls Rose, and without its "verify the deciding facts" line Rose was not more accurate. - Keep the two-pass sheet process (research, then audit) of
offers-research.md. - Decide the trade-off: verification costs about 7 s and 100k tokens per answer. A cheaper variant to test next: verify only prices and yes/no compliance facts.
- Keep measuring each change on a fresh held-out set, fact-checked blind.
- Next step, aiming at 99% fully right: IX-5395 (a knowledge base per vendor, answers that carry their caveats, cheaper verification, measured on at least 200 fresh questions).
Method¶
- Arms. Codex
gpt-5.6-sol, medium reasoning effort, live web search, in two isolatedCODEX_HOMEdirectories. The Rose arm adds the r4a-mcp stdio server and thebench/AGENTS.rose.mdsnippet a buyer installs with the connector. - Gold. One set of facts per question, read on the vendors' own pages. Where both arms contradicted a gold fact, the vendor page was checked again: four golds were wrong and fixed.
- Judge. Claude Sonnet 5 (a different model family from Codex): a 0-3 score per answer against the gold, and a blind preference between the two answers that counts only when it survives swapping their order.
- One answer per question and arm. The Codex-alone answers are from run r1; runs r2 and r3 re-ran only the Rose arm after each round of changes.
Results¶
| Run r3 | Codex alone | Codex + Rose |
|---|---|---|
| Score against gold (0-3) | 2.44 | 2.72 |
| Blind preference wins (39 questions, 12 ties) | 6 | 21 |
| Median time per answer | 28 s | 19 s |
| Median tokens per answer | 90k | 90k |
| Listing questions (l1-l4) | 70-130 s, 0.3-0.9M tokens | 21-45 s, 0.09-0.23M tokens |
| Category | Questions | Score alone | Score with Rose | Rose wins | Alone wins |
|---|---|---|---|---|---|
| Pricing | 12 | 2.50 | 2.75 | 10 | 0 |
| Features and compliance | 9 | 2.44 | 2.78 | 1 | 4 |
| Getting started | 4 | 2.50 | 2.50 | 2 | 1 |
| Comparison | 6 | 2.33 | 2.67 | 3 | 0 |
| Listing | 4 | 2.50 | 3.00 | 3 | 1 |
| Which vendor does what | 4 | 2.25 | 2.50 | 2 | 0 |
How the result moved across runs:
| Run | Change | Score alone / with Rose | Wins alone / with Rose |
|---|---|---|---|
| r1 | first sheets, connector installed without instruction | 2.30 / 2.24 | 16 / 14 |
| r2 | instruction snippet, reliable server start, audited sheets, get_offers for sheet vendors |
2.38 / 2.49 | 19 / 15 |
| r3 | complete-answer instructions, judge given today's date and order-swap check | 2.44 / 2.72 | 6 / 21 |
Round 2: richer knowledge, held-out questions¶
After round 1, the knowledge behind Rose was rebuilt without looking at the round 1 failures:
- Isolation. The 39 round 1 questions and their failures became a seen dev set. No fact from them was patched. A separate agent wrote 40 new questions (pricing, features, getting started, comparisons, listings and discovery, which does what) without seeing the sheets or round 1, and the sheets were frozen before anyone read them.
- Catalog. A market-mapping agent listed the 35 most relevant vendors from scratch (15
before). Research agents re-read every vendor from its own pages following
backend/apps/r4a_mcp/offers-research.md(docs index andllms.txt, raw pricing HTML, playground, trust and privacy pages, MCP docs and repository, changelog), without access to the old sheets: 303 plans, 297 products and 569 sourced notes. - Tools. Sheets gained products and free-text sourced notes;
compare_offerslists vendors without a volume, filters by product kind, names every vendor it covers, and says its list is not the whole market.
Results on the held-out set, fact-checked blind (bench/CHECKER.md, bench/factcheck.py):
| 40 held-out answers | Codex alone | Codex + Rose |
|---|---|---|
| Correct | 31 (78%) | 30 (75%) |
| Minor error | 6 | 5 |
| Wrong | 3 (8%) | 5 (12%) |
| False claims (material) | 14 (6) of 433 | 12 (5) of 448 |
| Median time | 38 s | 31 s |
| Median tokens | 136k | 124k |
| Answers that used Rose | 40 of 40 |
Why each Codex + Rose answer was wrong:
| Question | Wrong claim | Cause |
|---|---|---|
| Serper at 1M searches a month | two Standard packs, $750 | normalization: packs valid 6 months were priced as bought every month; the Scale pack comes to $500 a month |
| SerpApi at 300,000 searches | three $725 renewals, $2,175 | normalization: early renewal counted as whole extra plans; the Volume plan with renewal is about $1,770 |
| Bright Data JS rendering | must be switched on with render: true |
data: the sheet says so; the docs also say Web Unlocker detects JS sites and renders them itself |
| Browserbase hosted MCP | API key as a Bearer header | answer: the sheet says query parameter; Codex wrote a header |
| Exa content pricing | text or highlights cost extra on /search |
shared: Codex alone made the same misreading |
Codex alone was wrong on Exa (same misreading), a Tavily-versus-Exa winner (Tavily pay-as-you-go priced instead of its plan) and a ZenRows price (annual price used as monthly, then discounted again).
What round 2 shows:
- More data did not buy accuracy. Codex with live web search is already right on about nine answers in ten. Rose's data matched it but did not beat it.
- Normalization is Rose's main risk. Two of five errors come from turning prepaid packs and early renewal into a monthly price; the rule is in the code and the procedure, so it fails the same way for every vendor that sells packs. These rules need either a better model (validity windows, partial renewal) or to be shown as options, not as a single price.
- The speed gain shrank (19 s against 28 s in round 1, 31 s against 38 s here) because the richer sheets are larger tool results (15k characters median per vendor, 65k for SerpApi).
- Fixing these requires a new held-out set. These 40 questions are now seen; round 3 must be measured on questions written after its changes.
Round 3: audited sheets, pricing rules, verified answers¶
Changes, all generic, decided before the round 3 questions existed:
- Audit pass. A second agent re-read every fact of the 35 sheets on the vendors' pages and logged 434 corrections and additions with a source and a quote: defaults (what an API does when a parameter is not set), exact auth parameter names and commands, prepaid packs, annual-versus-monthly prices, search counts (Decodo's were about a third too high).
- Pricing rules. Prepaid packs carry their validity and are spread over the months they
last; plans renewed early are prorated; every computed price comes with a
price_rulethat says how it was computed. - Answer instructions. Copy commands, parameter names and prices exactly; never fill a
gap from memory; open the
source_urlof the one or two facts that decide the answer and trust the page when it differs; state the price rule next to the price.
Results on 50 new held-out questions (written by an agent that saw neither the sheets nor earlier rounds; sheets frozen before the questions were read), fact-checked blind:
| 50 held-out answers | Codex alone | Codex + Rose |
|---|---|---|
| Correct | 38 (76%) | 43 (86%) |
| Minor error | 9 | 6 |
| Wrong | 3 (6%) | 1 (2%) |
| False claims (material) | 13 (3) of 551 | 8 (1) of 593 |
| Answers the question only partly | 6 | 1 |
| Median time | 39 s | 46 s |
| Median tokens | 138k | 243k |
| Answers that used Rose / also searched the web | 48 / 50 |
Codex alone was wrong twice on Exa (charged page contents on top of a search, which Exa includes) and once on ScrapingBee (missed early renewal, so picked the wrong winner). Codex + Rose was wrong once: it gave Perplexity's general zero-retention default for the Search API, which Perplexity's Search API terms exclude.
Round 1: was the answer right?¶
Fact-check of every claim
The gold facts are one reading of the vendors' pages, not the truth: extra true information is fine, and a missing detail matters only when the question asks for it. So every answer was also fact-checked on its own. Eight checkers (Claude agents) received the 78 answers blind, labelled X and Y in random order, with no gold. They listed every claim that bears on the question (prices, allowances, yes/no features, commands, totals) and checked each on the vendor's own pages on 2026-10-01, recomputing the arithmetic. An answer is wrong when a false claim changes or misleads the answer to the question, minor error when a false claim does not.
| 39 answers (r3), fact-checked | Codex alone | Codex + Rose |
|---|---|---|
| Correct | 32 (82%) | 34 (87%) |
| Minor error | 3 (8%) | 2 (5%) |
| Wrong | 4 (10%) | 3 (8%) |
| Claims checked | 400 | 381 |
| False claims (of which material) | 11 (4) | 6 (4) |
| Answers the question only partly | 5 | 6 |
Codex alone was wrong on:
| Question | Claim | Truth on the vendor's page |
|---|---|---|
| p10 Linkup pricing | search costs $0.005-$0.006 "depending on depth" | deep search is $0.05-$0.055, 10 times more; the $0.005/$0.006 split is by output type |
| f4 GDPR DPA and EU processing | Jina's EU processing "not confirmed"; Firecrawl's DPA enterprise-only; Linkup "strict EU-only" | Jina offers EU residency (experimental); Firecrawl's DPA comes with paid plans; Linkup may route through the US, EU, Canada or APAC |
| f9 Bright Data MCP auth | API token, "not OAuth" | the hosted server also accepts OAuth 2.1 |
| l3 search only vs page content | Serper is search only | Serper now sells page fetching (scrape.serper.dev) |
Codex + Rose was wrong on:
| Question | Claim | Truth on the vendor's page | Cause |
|---|---|---|---|
| l2 free tiers without a card | "seven APIs" have one | many more do (ZenRows, ScrapingAnt, Scrape.do, Diffbot, Browserless...) | Rose's 15 vendors presented as the market |
| l3 search only vs page content | Serper is search only | Serper now sells page fetching | sheet misses Serper's new product |
| c2 SerpApi vs Serper at 50,000 searches | SerpApi $550/month on Big Data | Big Data covers 30,000; $550 assumes renewing it early within the month; the cheapest plan covering 50,000 is $725 | pricing rule stated as the plan price |
Minor errors: Rose's Firecrawl format list leaves out four formats, its Exa comparison ignores the cheaper Instant mode; Codex alone gives a wrong Context.dev allowance, a broken Oxylabs link and $40 for Tavily where $38 is possible.
What this changes:
- The gold-based count overstated Codex's errors. Against the gold, Codex alone looked wrong on 5 answers and Rose on none. The fact-check keeps only one of those (Linkup), and two of the gold's own figures were assumptions: Exa's $10 monthly credit applying to paid usage (Exa's docs do not say) and Tavily's free credits applying to pay-as-you-go.
- Rose fixes reading errors (Linkup's deep-search price, Jina's EU option, Firecrawl's DPA), the ones a web search makes on long or contradictory pages.
- Rose adds two error modes of its own: it answers market-wide questions from its 15 vendors as if they were the market, and a sheet misses a product the vendor sells (Serper's page fetching, visible only in its playground). The third error is a pricing rule (renewing a plan early) presented as a price.
The questions, answers and fact-check verdicts of every round are kept outside the repository (see backend/apps/r4a_mcp/bench/README.md for where, and how to recompute the tables).
Findings¶
- Codex must be told to use the connector. Codex loads MCP tools on demand. In r1 the Rose arm answered from the web on almost every question. The instruction alone was not enough: the model read "call Rose first", saw no Rose tools in its list and fell back to the web, until the snippet told it to search its tools for them.
- Where Rose wins: answers across the market, fast. Listing and comparison questions make Codex read page after page (up to 2 minutes and 0.9M tokens). Rose answers them from normalized data in one or two calls, at about the same accuracy. On completeness it is bounded by its catalog: asked for every scraping API with a free tier, it listed its own seven as if they were all.
- Where web search reads wrong. On Linkup (p10), Codex alone gave the deep-search price as $0.005 per request; the pricing page says $0.05 to $0.055. Once it called Rose (r2, r3), Codex gave the right price from the sheet. Both arms still score low on p10 because both leave out the page-fetch price the question's gold includes; the judge's 1/3 for Rose in r3 lists that omission as its only error, so it is harsh rather than a wrong figure. Long or contradictory pages are where the sheets help most: Linkup's marketing page shows the same misleading range as Codex's answer, and only the docs give the deep-search price.
- Where Codex alone still wins: extra detail. Its remaining wins add things the sheets do not hold, such as named community MCP servers for Serper, more Firecrawl output formats, and setup steps for several hosts.
- The sheets must be checked on first-party pages, not only pricing pages. A first pass missed facts on homepages, privacy policies and trust centers (ScrapingBee's SOC 2 Type II, Parallel's EU data residency, several DPAs). An audit fixed 25 facts across 11 vendors.
- Judge bias. The judge first marked Rose's "as of 2026-10-01" attributions as fabricated future dates; it now gets today's date. Its preference also changed with answer order; a preference now counts only when it survives swapping the answers.
Limits¶
- The gold facts and the sheets come from the same research, which favors Rose. Re-checking every gold fact that both arms contradicted reduces this but does not remove it.
- One answer per question and arm, and one judge model.
- Deep research modes and ChatGPT were not run: they have no headless runner.
Reproduce¶
From backend/apps/r4a_mcp, with Codex signed in and Secret Manager access:
just bench base ../.context/bench/runs/rN # Codex alone
just bench rose ../.context/bench/runs/rN # Codex + Rose
just bench-judge ../.context/bench/runs/rN
cd ../.. && poetry run python apps/r4a_mcp/bench/transcript.py ../.context/bench/runs/rN \
../docs/src/analysis/r4a-codex-vs-rose-transcript.md
Plan and decisions: docs/plans/2026-10-01-ix-5364-r4a-phase-0.md.