# Shared starting conditions — producer only

Status: setup guide, not a completed experiment. Do not send these judging notes or prior outputs to a participant task.

Use identical assignment text, audience, deliverable limits, relevant instructions, skills, available tools, sign-in access, speed setting and starting folders. Use fresh independent tasks and isolated copies. Record actual effective settings; a task title is not evidence of configuration.

Core arm: Astra Low, Medium, High, Extra High and Max when available. Keep speed Standard and discretionary delegation disabled consistently in this arm. Ultra is a separate delegation-enabled condition; record children and effective settings. Sol High is an optional baseline specifically for Tibo’s claim.

Grok is an available research option, not a required action. Do not prescribe a tool sequence or grade usage of a particular tool as success. Identical access does not guarantee identical live research results. Different idea choices and live sources confound a pure effort-only interpretation. Report an open-ended project challenge.

Run a feasibility pilot first. Freeze one wall-clock ceiling and one rescue policy before collecting scored runs; the ceiling is not yet set. Save output when the agent ends or reaches that ceiling, before any help. Record failed/missing requirements and legitimate blockers separately. Then give one identical neutral continuation prompt where applicable, with a separately predeclared rescue ceiling. Do not silently grant one run additional hints or time.

Repeat the full comparison with fresh state before a strong recommendation. Randomize/interleave effort order. Isolate or serialize shared browser/computer-use sessions. Preserve all runs, ties and failures. Grade artifacts without effort labels where practical. Record elapsed time and attributable usage only when actually available; account-wide allowance movement is not a task bill.

Fixed acceptance: genuine relevant source evidence; explicit customer/opportunity/counterevidence; plausible but unproven defensibility path; editable connected planning canvas; working 3–5-step workflow with persistence; observed verification. Score each as passed, failed or missing with a receipt. Keep visual taste and product complexity separate.

## Exploratory batch R1, 7 September 2026

User authorized execution. Seven tasks launched together; model/effort verified in session records. 45-minute self-managed first attempts; then, if needed, one identical neutral continuation with a 15-minute self-managed allowance. No first-attempt coaching. Fresh folders, own browser tabs, no shared native-app control, local serving only. These concurrency restrictions are identical for all arms. Timing can reflect shared-machine contention and browser failures. Do not treat this batch as a controlled latency benchmark or derive a universal ranking. Archive before any rescue.
