curl -X POST https://app.selfbench.dev/api/orgs/mupt-ai/repos/mupt-ai/self-bench/evaluations/comparisons \
-H "Authorization: Bearer $SELFBENCH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"id":"7e0acb2b-3b1e-4df6-8afb-1b0a1cf2d6a8",
"tasks":[{"runId":"batch-…","taskId":"task-id"}],
"models":[{"catalogId":"<catalog-id>","credentialId":"<uuid>","harnesses":["codex"]}],
"sandbox":"e2b","sandboxCredentialId":"<uuid>","skipCompleted":true}'
Workspace API
Create a Comparison
Run model and harness settings on approved tasks.
POST
/
api
/
orgs
/
{org}
/
repos
/
{owner}
/
{name}
/
evaluations
/
comparisons
curl -X POST https://app.selfbench.dev/api/orgs/mupt-ai/repos/mupt-ai/self-bench/evaluations/comparisons \
-H "Authorization: Bearer $SELFBENCH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"id":"7e0acb2b-3b1e-4df6-8afb-1b0a1cf2d6a8",
"tasks":[{"runId":"batch-…","taskId":"task-id"}],
"models":[{"catalogId":"<catalog-id>","credentialId":"<uuid>","harnesses":["codex"]}],
"sandbox":"e2b","sandboxCredentialId":"<uuid>","skipCompleted":true}'
string
required
Workspace login.
string
required
GitHub owner.
string
required
Repository name.
string
required
Bearer plus a write API key.string
required
Client-generated UUID; reuse it when retrying.
array
required
Approved
{runId, taskId} pairs.array
required
Model catalog ID, credential ID, and harnesses (up to 12 model choices).
string
required
Sandbox from the evaluation catalog.
string
required
Matching credential ID.
boolean
Skip already completed trials.
curl -X POST https://app.selfbench.dev/api/orgs/mupt-ai/repos/mupt-ai/self-bench/evaluations/comparisons \
-H "Authorization: Bearer $SELFBENCH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"id":"7e0acb2b-3b1e-4df6-8afb-1b0a1cf2d6a8",
"tasks":[{"runId":"batch-…","taskId":"task-id"}],
"models":[{"catalogId":"<catalog-id>","credentialId":"<uuid>","harnesses":["codex"]}],
"sandbox":"e2b","sandboxCredentialId":"<uuid>","skipCompleted":true}'
202 with comparison status. Fetch tasks from GET …/evaluations/options, models and sandboxes from GET …/evaluations/catalog, and status from GET …/evaluations/comparisons/{id}. If submissionError is present, resume the same comparison at POST …/comparisons/{id}/resume.