Real-world Pre-PoC decision · ACTIVE EVALUATION

Privileged workstation · secrets boundary · sandbox · approval before change

Secure coding-agent sandbox

A security-sensitive workstation with production-capable credentials needs a coding-agent setup that cannot read secrets, escape its sandbox, merge or deploy autonomously.

4Agents evaluated
2Ready for direct PoC
0Need vendor confirmation
2Blocked pre-PoC
0Disqualified
Recommended next stepResolve the blocking evidence questions, then test only the candidates that clear the gate.
Adapt this decision to my team
Candidate routing

Who moves forward—and why

Status is evidence-relative, not a universal product ranking.

BLOCKED PRE POC

Claude Code

Why this status?

  • Sandbox and isolation: UNKNOWN. P0 lacks the required proof and blocks pre-PoC qualification.
  • Network restrictions: PASS. Accepted official evidence supports this criterion for the stated applicability.
What happens next

SEARCH OFFICIAL EVIDENCE · Locate current official Claude Code documentation that explicitly establishes sandbox-isolation for the proposed product, plan, mode and region.

BLOCKED PRE POC

OpenAI Codex

Why this status?

  • Sandbox and isolation: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Network restrictions: UNKNOWN. P0 lacks the required proof and blocks pre-PoC qualification.
What happens next

SEARCH OFFICIAL EVIDENCE · Locate current official OpenAI Codex documentation that explicitly establishes network-restrictions for the proposed product, plan, mode and region.

RUN POC · Test OpenAI Codex against the fixed scenario to measure human-approval-boundaries; do not infer the outcome from documentation.

QUALIFIED FOR POC

Cursor

Why this status?

  • Sandbox and isolation: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Network restrictions: PASS. Accepted official evidence supports this criterion for the stated applicability.
What happens next

Keep the candidate in the controlled PoC and apply the stated pass/fail thresholds.

QUALIFIED FOR POC

GitHub Copilot

Why this status?

  • Sandbox and isolation: PASS. Accepted official evidence supports this criterion for the stated applicability.
  • Network restrictions: PASS. Accepted official evidence supports this criterion for the stated applicability.
What happens next

GET CUSTOMER CONTEXT · Obtain the buyer's required environment, policy and threshold for local-execution before qualifying GitHub Copilot.

Unknown → action

What still needs to be resolved?

Claude Code

Sandbox and isolation

No accepted official evidence in the Phase 1 source set establishes Sandbox and isolation for Claude Code.

SEARCH OFFICIAL EVIDENCE
OpenAI Codex

Network restrictions

No accepted official evidence in the Phase 1 source set establishes Network restrictions for OpenAI Codex.

SEARCH OFFICIAL EVIDENCE
OpenAI Codex

Human approval boundaries

Codex environment controls constrain execution, but public evidence does not prove every material action requires approval.

RUN POC
GitHub Copilot

Local execution

No accepted official evidence in the Phase 1 source set establishes Local execution for GitHub Copilot.

GET CUSTOMER CONTEXT
Candidate-specific PoC

Test the remaining uncertainty

Cursor · Can Cursor complete the fixed task while forbidden files, credentials, destructive shell commands and non-allowlisted network destinations remain inaccessible?

Why this candidate needs the test: Buyer-specific measured performance

What to measure: task acceptance rate, elapsed minutes, human interventions, failed or invalid runs, estimated usage cost, policy exceptions

What counts as failure: Before execution, the buyer must declare a numeric acceptance threshold for task success, maximum elapsed time, maximum human interventions, maximum cost and zero tolerance for prohibited policy actions.

Repeat: 5 controlled runs.

Cannot prove: This test cannot prove contractual SLA, future vendor behavior, universal performance, or operation outside the tested plan, model, repository and environment.

GitHub Copilot · Can GitHub Copilot complete the fixed task while forbidden files, credentials, destructive shell commands and non-allowlisted network destinations remain inaccessible?

Why this candidate needs the test: Buyer-specific measured performance

What to measure: task acceptance rate, elapsed minutes, human interventions, failed or invalid runs, estimated usage cost, policy exceptions

What counts as failure: Before execution, the buyer must declare a numeric acceptance threshold for task success, maximum elapsed time, maximum human interventions, maximum cost and zero tolerance for prohibited policy actions.

Repeat: 5 controlled runs.

Cannot prove: This test cannot prove contractual SLA, future vendor behavior, universal performance, or operation outside the tested plan, model, repository and environment.

Official evidence

Trace every positive qualification

The original public problem proves demand provenance. Only official sources support product qualification.

Claude Code · Network restrictions
  1. Requirementreq_v1_18 · network-restrictions
  2. Factfact_p15_claude-code_network-restrictions
  3. Accepted evidenceev_p15_claude-code_network-restrictions
  4. ApplicabilityClaude Code clients routed through corporate HTTP/HTTPS proxies · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCorporate proxy configuration →
Claude Code · Human approval boundaries
  1. Requirementreq_v1_23 · human-approval-boundaries
  2. Factfact_claude-code_human-approval-boundaries
  3. Accepted evidenceev_claude-code_human-approval-boundaries
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceClaude Code overview →
Claude Code · Local execution
  1. Requirementreq_v1_16 · local-execution
  2. Factfact_claude-code_local-execution
  3. Accepted evidenceev_claude-code_local-execution
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceClaude Code overview →
Claude Code · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_claude-code_audit-activity-logs
  3. Accepted evidenceev_p15_claude-code_audit-activity-logs
  4. ApplicabilityClaude Enterprise organizations · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceAccess audit logs →
OpenAI Codex · Sandbox and isolation
  1. Requirementreq_v1_22 · sandbox-isolation
  2. Factfact_codex_sandbox-isolation
  3. Accepted evidenceev_codex_sandbox-isolation
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCodex overview →
OpenAI Codex · Local execution
  1. Requirementreq_v1_16 · local-execution
  2. Factfact_dcs_codex_local-execution
  3. Accepted evidenceev_dcs_codex_local-execution
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCodex cloud environments →
OpenAI Codex · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_codex_audit-activity-logs
  3. Accepted evidenceev_p15_codex_audit-activity-logs
  4. ApplicabilityChatGPT Enterprise compliance and Codex usage logs · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceNew compliance and administrative tools for ChatGPT Enterprise →
Cursor · Sandbox and isolation
  1. Requirementreq_v1_22 · sandbox-isolation
  2. Factfact_depth_cursor-agent_sandbox-isolation
  3. Accepted evidenceev_depth_cursor-agent_sandbox-isolation
  4. ApplicabilityCursor 2.0+ on supported macOS and Linux configurations; Cloud Agents use a separate dedicated-machine boundary · As offered by the provider; no broader regional availability inferred
  5. Verified2026-09-02 · CURRENT
  6. Official sourceRun Modes →
Cursor · Network restrictions
  1. Requirementreq_v1_18 · network-restrictions
  2. Factfact_p15_cursor-agent_network-restrictions
  3. Accepted evidenceev_p15_cursor-agent_network-restrictions
  4. ApplicabilityCursor Enterprise security controls · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor Security and Privacy Hardening →
Cursor · Human approval boundaries
  1. Requirementreq_v1_23 · human-approval-boundaries
  2. Factfact_dcs_cursor-agent_human-approval-boundaries
  3. Accepted evidenceev_dcs_cursor-agent_human-approval-boundaries
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor agent run modes →
Cursor · Local execution
  1. Requirementreq_v1_16 · local-execution
  2. Factfact_dcs_cursor-agent_local-execution
  3. Accepted evidenceev_dcs_cursor-agent_local-execution
  4. ApplicabilityOnly the plan/scope named by the cited official source · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor agent run modes →
Cursor · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_cursor-agent_audit-activity-logs
  3. Accepted evidenceev_p15_cursor-agent_audit-activity-logs
  4. ApplicabilityCursor Enterprise plan · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceCursor Compliance and Monitoring →
GitHub Copilot · Sandbox and isolation
  1. Requirementreq_v1_22 · sandbox-isolation
  2. Factfact_depth_github-copilot-coding-agent_sandbox-isolation
  3. Accepted evidenceev_depth_github-copilot-coding-agent_sandbox-isolation
  4. ApplicabilityGitHub Copilot CLI and GitHub Copilot app sandbox modes; preview limitations apply · As offered by the provider; no broader regional availability inferred
  5. Verified2026-09-02 · CURRENT
  6. Official sourceAbout cloud and local sandboxes for GitHub Copilot →
GitHub Copilot · Network restrictions
  1. Requirementreq_v1_18 · network-restrictions
  2. Factfact_p15_github-copilot-coding-agent_network-restrictions
  3. Accepted evidenceev_p15_github-copilot-coding-agent_network-restrictions
  4. ApplicabilityGitHub Enterprise Cloud agent administration · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceAgent management for enterprises →
GitHub Copilot · Human approval boundaries
  1. Requirementreq_v1_23 · human-approval-boundaries
  2. Factfact_depth_github-copilot-coding-agent_human-approval-boundaries
  3. Accepted evidenceev_depth_github-copilot-coding-agent_human-approval-boundaries
  4. ApplicabilityGitHub Copilot cloud agent on supported paid plans and repositories · As offered by the provider; no broader regional availability inferred
  5. Verified2026-09-02 · CURRENT
  6. Official sourceRisks and mitigations for GitHub Copilot cloud agent →
GitHub Copilot · Audit and activity logs
  1. Requirementreq_v1_13 · audit-activity-logs
  2. Factfact_p15_github-copilot-coding-agent_audit-activity-logs
  3. Accepted evidenceev_p15_github-copilot-coding-agent_audit-activity-logs
  4. ApplicabilityGitHub Copilot Business and Enterprise on Enterprise Cloud · Only regions supported by the cited plan; otherwise UNKNOWN
  5. Verified2026-09-01 · CURRENT
  6. Official sourceReviewing audit logs for GitHub Copilot →
Decision boundary

This narrows a PoC. It does not make the final purchase decision.

Public evidence cannot establish buyer-specific performance, negotiated terms, or future vendor behavior.