average retrievals
71.1s average latencyBetter-supported answers.
Less research work.
We asked an AI model to answer the same five decision fields for six AI coding Agents using official vendor pages alone, then using those pages plus Agenova's normalized data. A separate model scored both preserved result sets against accepted official sources.
Agenova reduced retrievals by 100%.
average retrievals
40.0s average latencyretrieval work
−44% latencySupport improved without hiding unknowns.
Identity errors were 0 in both cohorts, stale claims changed from 1 to 0, and unsupported claims from 15 to 0. Requirement completeness changed from 100% to 97% because one unsupported field stayed explicitly unknown. This supports additive value inside this bounded pilot; one run per cohort does not measure repeated-answer consistency, and this is not a universal performance claim.
- Agents
- 6
- Fields
- 5
- Runs per cohort
- 1
- Independent evaluator
- gpt-5.4
How to cite this result without overstating it.
Agenova AI Coding Agent data-value retrieval pilot (2026-08-28). https://agenova.io/benchmarks/ai-coding
Bounded post-repair comparison of six AI coding Agents, five decision fields, and one run per cohort; do not generalize this result to all Agents, fields, models, or retrieval conditions.
Open the exact machine-readable resultA real result with a deliberately narrow boundary.
- This is a bounded post-repair six-Agent, five-field, one-run-per-cohort pilot rather than the 100-prompt completion benchmark.
- The remaining missing pair is Aider enterprise controls, which stays unknown because no qualifying official product-level claim was found.
- Repeated-answer consistency is not assessed because this post-repair pilot has one run per cohort; the Agenova-assisted run had zero stale or unsupported claims.
- External search visibility and model citation remain separate observation metrics.