Off-the-shelf agent training data

Long-horizon tool-use data for realistic AI agents.

Curated multi-environment scenario packs with prompts, verifiers, seeds, golden traces, and run evidence. Buy only the layer you need, from prompt-only tasks to full Harbor environments.

Marketplace previewHuman-reviewed
Scenario familyRecruiting ops across Email, Calendar, Teams, Asana

42 tool calls in golden trace. SQL + rubric verifier. Seeded actor persona. Optional environment.

Pack layers4
Verifier modesSQL + Rubric
DeliveryPrompt to runtime
Dataset catalog

Choose the artifact layer your team can use today.

Start with prompts, add verifiers when you need measurable evaluation, or license full executable packages when your team wants reproducible agent runs.

Coming soon50 Harbor CUA tasks + SpotHub desktop runtime + evidence

SpotHub CUA All-50 Pack

Fifty executable SpotHub computer-use tasks with the shared SpotHub desktop runtime and per-task run evidence.

Tasks50
RuntimeSpotHub desktop
  • 50 executable Harbor CUA tasks
  • Shared SpotHub desktop runtime kit
  • Per-task run evidence
  • Signed download access
Coming soon14 Harbor tasks + workplace four-gym runtime suite

Workplace Four-Gym Pack

Fourteen long-horizon workplace tasks spanning email, project, calendar, and chat surfaces on the shared four-gym runtime suite.

Tasks14
RuntimeFour-gym suite
  • 14 long-horizon workplace tasks
  • Email, project, calendar, and chat gyms
  • Golden trajectories + verifier evidence
  • Signed download access
Coming soon3 Harbor eval tasks + Default-DB runtime + seed dataset

Default-DB Eval Pack

Three multi-surface Default-DB evaluation tasks packaged with shared runtime gyms and the Default-DB seed dataset bundle.

Tasks1
DatasetDefault-DB v2
  • Multi-surface Default-DB eval task
  • Shared runtime gyms (Freshdesk, Jira, Slack, Zeta-SQL)
  • Default-DB seed dataset bundle
  • Signed download access
Why this is not commodity synthetic data

Proof-bearing datasets for teams that need more than rows.

Each pack is designed to support training and evaluation decisions: realistic workflow context, objective checks where possible, and evidence that the task stresses current agents.

01

Human-reviewed scenario design, not raw synthetic dumps.

02

Model-attempt evidence can be attached before purchase.

03

Buyer can choose prompts only, prompts plus verifiers, or full executable task package.

04

Harbor-compatible packaging keeps the data useful for RL, SFT, and eval workflows.

Scenario coverage

Realistic workplace workflows, packaged for agent teams.

Select from operational domains where long-horizon tool use matters: coordinating people, updating systems, reconciling context, and leaving an auditable final state.

Recruiting and HR workflowsCustomer support escalationsProject management operationsSales and CRM follow-throughIncident response coordinationFinance and compliance ops
Custom requests

Need a specific domain or package shape?

Tell us what your team is evaluating or training. We will use this to recommend an existing pack or scope a custom long-horizon workflow dataset.