Real-world Pre-PoC decision · ACTIVE EVALUATION

~200 developers · ticket-to-PR · human review · company-wide evaluation

Governed ticket-to-PR workflow

A roughly 200-developer company is evaluating a governed ticket-to-pull-request workflow with agent implementation and human code review.

4Agents evaluated
4Ready for direct PoC
0Need vendor confirmation
0Blocked pre-PoC
0Disqualified
Recommended next stepRun the candidate-specific controlled tests before making a purchase decision.
Adapt this decision to my team
Candidate routing

Who moves forward—and why

Status is evidence-relative, not a universal product ranking.

QUALIFIED FOR POC

Claude Code

Why this status?

  • Autonomous task execution: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Issue-to-code workflow: UNKNOWN. P0 is designed for empirical verification; candidate may enter only the bounded PoC.
What happens next

RUN POC · Test Claude Code against the fixed scenario to measure issue-to-code-workflow; do not infer the outcome from documentation.

QUALIFIED FOR POC

OpenAI Codex

Why this status?

  • Autonomous task execution: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Issue-to-code workflow: UNKNOWN. P0 is designed for empirical verification; candidate may enter only the bounded PoC.
What happens next

RUN POC · Test OpenAI Codex against the fixed scenario to measure issue-to-code-workflow; do not infer the outcome from documentation.

RUN POC · Test OpenAI Codex against the fixed scenario to measure human-approval-boundaries; do not infer the outcome from documentation.

QUALIFIED FOR POC

Cursor

Why this status?

  • Autonomous task execution: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Issue-to-code workflow: UNKNOWN. P0 is designed for empirical verification; candidate may enter only the bounded PoC.
What happens next

RUN POC · Test Cursor against the fixed scenario to measure issue-to-code-workflow; do not infer the outcome from documentation.

QUALIFIED FOR POC

GitHub Copilot

Why this status?

  • Autonomous task execution: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Issue-to-code workflow: PASS. Accepted official evidence supports this criterion for the stated applicability.
What happens next

Keep the candidate in the controlled PoC and apply the stated pass/fail thresholds.

Unknown → action

What still needs to be resolved?

Claude Code

Issue-to-code workflow

No accepted official evidence in the Phase 1 source set establishes Issue-to-code workflow for Claude Code.

RUN POC
OpenAI Codex

Issue-to-code workflow

No accepted official evidence in the Phase 1 source set establishes Issue-to-code workflow for OpenAI Codex.

RUN POC
OpenAI Codex

Human approval boundaries

Codex environment controls constrain execution, but public evidence does not prove every material action requires approval.

RUN POC
Cursor

Issue-to-code workflow

No accepted official evidence in the Phase 1 source set establishes Issue-to-code workflow for Cursor.

RUN POC
Candidate-specific PoC

Test the remaining uncertainty

Claude Code · Can Claude Code turn the same fixed ticket into a reviewable pull request that passes tests and stays within the predeclared human-intervention and rework thresholds?

Why this candidate needs the test: gap-claude-code-issue-to-code-workflow

What to measure: task acceptance rate, elapsed minutes, human interventions, failed or invalid runs, estimated usage cost, policy exceptions

What counts as failure: Before execution, the buyer must declare a numeric acceptance threshold for task success, maximum elapsed time, maximum human interventions, maximum cost and zero tolerance for prohibited policy actions.

Repeat: 5 controlled runs.

Cannot prove: This test cannot prove contractual SLA, future vendor behavior, universal performance, or operation outside the tested plan, model, repository and environment.

OpenAI Codex · Can OpenAI Codex turn the same fixed ticket into a reviewable pull request that passes tests and stays within the predeclared human-intervention and rework thresholds?

Why this candidate needs the test: gap-codex-issue-to-code-workflow, gap-codex-human-approval-boundaries

What to measure: task acceptance rate, elapsed minutes, human interventions, failed or invalid runs, estimated usage cost, policy exceptions

What counts as failure: Before execution, the buyer must declare a numeric acceptance threshold for task success, maximum elapsed time, maximum human interventions, maximum cost and zero tolerance for prohibited policy actions.

Repeat: 5 controlled runs.

Cannot prove: This test cannot prove contractual SLA, future vendor behavior, universal performance, or operation outside the tested plan, model, repository and environment.

Cursor · Can Cursor turn the same fixed ticket into a reviewable pull request that passes tests and stays within the predeclared human-intervention and rework thresholds?

Why this candidate needs the test: gap-cursor-agent-issue-to-code-workflow

What to measure: task acceptance rate, elapsed minutes, human interventions, failed or invalid runs, estimated usage cost, policy exceptions

What counts as failure: Before execution, the buyer must declare a numeric acceptance threshold for task success, maximum elapsed time, maximum human interventions, maximum cost and zero tolerance for prohibited policy actions.

Repeat: 5 controlled runs.

Cannot prove: This test cannot prove contractual SLA, future vendor behavior, universal performance, or operation outside the tested plan, model, repository and environment.

GitHub Copilot · Can GitHub Copilot turn the same fixed ticket into a reviewable pull request that passes tests and stays within the predeclared human-intervention and rework thresholds?

Why this candidate needs the test: Buyer-specific measured performance

What to measure: task acceptance rate, elapsed minutes, human interventions, failed or invalid runs, estimated usage cost, policy exceptions

What counts as failure: Before execution, the buyer must declare a numeric acceptance threshold for task success, maximum elapsed time, maximum human interventions, maximum cost and zero tolerance for prohibited policy actions.

Repeat: 5 controlled runs.

Cannot prove: This test cannot prove contractual SLA, future vendor behavior, universal performance, or operation outside the tested plan, model, repository and environment.

Official evidence

Trace every positive qualification

The original public problem proves demand provenance. Only official sources support product qualification.

Claude Code · Autonomous task execution
  1. Requirementreq_v1_10 · autonomous-task-execution
  2. Factfact_claude-code_autonomous-task-execution
  3. Accepted evidenceev_claude-code_autonomous-task-execution
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceClaude Code overview →
Claude Code · Pull request workflow
  1. Requirementreq_v1_08 · pull-request-workflow
  2. Factfact_dcs_claude-code_pull-request-workflow
  3. Accepted evidenceev_dcs_claude-code_pull-request-workflow
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceClaude Code GitHub Actions →
Claude Code · Human approval boundaries
  1. Requirementreq_v1_23 · human-approval-boundaries
  2. Factfact_claude-code_human-approval-boundaries
  3. Accepted evidenceev_claude-code_human-approval-boundaries
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceClaude Code overview →
Claude Code · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_claude-code_audit-activity-logs
  3. Accepted evidenceev_p15_claude-code_audit-activity-logs
  4. ApplicabilityClaude Enterprise organizations · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceAccess audit logs →
OpenAI Codex · Autonomous task execution
  1. Requirementreq_v1_10 · autonomous-task-execution
  2. Factfact_codex_autonomous-task-execution
  3. Accepted evidenceev_codex_autonomous-task-execution
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCodex overview →
OpenAI Codex · Pull request workflow
  1. Requirementreq_v1_08 · pull-request-workflow
  2. Factfact_dcs_codex_pull-request-workflow
  3. Accepted evidenceev_dcs_codex_pull-request-workflow
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCodex code review →
OpenAI Codex · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_codex_audit-activity-logs
  3. Accepted evidenceev_p15_codex_audit-activity-logs
  4. ApplicabilityChatGPT Enterprise compliance and Codex usage logs · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceNew compliance and administrative tools for ChatGPT Enterprise →
Cursor · Autonomous task execution
  1. Requirementreq_v1_10 · autonomous-task-execution
  2. Factfact_cursor-agent_autonomous-task-execution
  3. Accepted evidenceev_cursor-agent_autonomous-task-execution
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor Agent documentation →
Cursor · Pull request workflow
  1. Requirementreq_v1_08 · pull-request-workflow
  2. Factfact_dcs_cursor-agent_pull-request-workflow
  3. Accepted evidenceev_dcs_cursor-agent_pull-request-workflow
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor Bugbot pull-request review →
Cursor · Human approval boundaries
  1. Requirementreq_v1_23 · human-approval-boundaries
  2. Factfact_dcs_cursor-agent_human-approval-boundaries
  3. Accepted evidenceev_dcs_cursor-agent_human-approval-boundaries
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor agent run modes →
Cursor · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_cursor-agent_audit-activity-logs
  3. Accepted evidenceev_p15_cursor-agent_audit-activity-logs
  4. ApplicabilityCursor Enterprise plan · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor Compliance and Monitoring →
GitHub Copilot · Autonomous task execution
  1. Requirementreq_v1_10 · autonomous-task-execution
  2. Factfact_depth_github-copilot-coding-agent_autonomous-task-execution
  3. Accepted evidenceev_depth_github-copilot-coding-agent_autonomous-task-execution
  4. ApplicabilityGitHub Copilot cloud agent on paid Copilot plans where enabled by policy · As offered by the provider; no broader regional availability inferred
  5. Verified2026-09-02 · CURRENT
  6. Official sourceGet started with Copilot agents on GitHub →
GitHub Copilot · Issue-to-code workflow
  1. Requirementreq_v1_09 · issue-to-code-workflow
  2. Factfact_github-copilot-coding-agent_issue-to-code-workflow
  3. Accepted evidenceev_github-copilot-coding-agent_issue-to-code-workflow
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceAbout GitHub Copilot coding agent →
GitHub Copilot · Pull request workflow
  1. Requirementreq_v1_08 · pull-request-workflow
  2. Factfact_github-copilot-coding-agent_pull-request-workflow
  3. Accepted evidenceev_github-copilot-coding-agent_pull-request-workflow
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceAbout GitHub Copilot coding agent →
GitHub Copilot · Human approval boundaries
  1. Requirementreq_v1_23 · human-approval-boundaries
  2. Factfact_depth_github-copilot-coding-agent_human-approval-boundaries
  3. Accepted evidenceev_depth_github-copilot-coding-agent_human-approval-boundaries
  4. ApplicabilityGitHub Copilot cloud agent on supported paid plans and repositories · As offered by the provider; no broader regional availability inferred
  5. Verified2026-09-02 · CURRENT
  6. Official sourceRisks and mitigations for GitHub Copilot cloud agent →
GitHub Copilot · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_github-copilot-coding-agent_audit-activity-logs
  3. Accepted evidenceev_p15_github-copilot-coding-agent_audit-activity-logs
  4. ApplicabilityGitHub Copilot Business and Enterprise on Enterprise Cloud · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceReviewing audit logs for GitHub Copilot →
Decision boundary

This narrows a PoC. It does not make the final purchase decision.

Public evidence cannot establish buyer-specific performance, negotiated terms, or future vendor behavior.