TS/JS Core Logic
code-workbench · ~60m
Objective
Prove they can plan, implement, and verify TypeScript/JavaScript under fixed model and token caps—not just paste until it compiles.
AI-in-bounds assessment. Same caps. Exportable evidence.
One module, fixed caps—then the evidence a hiring manager actually reads.
TypeScript/JavaScript, React Product UI, agent-assisted delivery, and debugging under constraints—a full assessment for AI-native roles.
~2–3h total · sequential modules · same constraints per candidate
MODEL_LOCK · TOKEN_BUDGET · CONTEXT_CAP · GOALS_VERIFY
code-workbench · ~60m
Objective
Prove they can plan, implement, and verify TypeScript/JavaScript under fixed model and token caps—not just paste until it compiles.
ui-workbench · ~55m
Objective
Ship a product UI with AI in bounds: streaming interfaces, agent controls, and recovery when the first pass fails.
web_skill
Objective
Orchestrate an agent-assisted workflow end to end—scope, tool use, and a verifiable outcome.
code-workbench
Objective
Diagnose and fix a failing system under the same constraints: isolate the fault, recover, and leave evidence.
Clone a pack → send candidate links → review traces and output. One bar per opening.
This is the Forward Deployed Engineer pack—ready to run. Pilots can use it as-is or map another role as the catalog grows.
Annotation companies hire at volume. Speed tests and rubric quizzes still miss who can think under model pressure—and leave a trail your QA leads can actually review.
Reviewers · AI trainers · eval ops · technical annotators
Fixed caps and the same arena for every hire—so you compare judgment, not who got a longer prompt or a better tool tip.
Exportable traces: how they interpret instructions, recover from bad first passes, and whether the output would survive your review ladder.
Not another coding screen. Assess the people who decide label quality, preference ranking, and eval pass/fail—under real AI constraints.
Pilots scoped to your ladder · packs adapted to your domains.
Talk annotation pilot →Limited pilot slots. Scoped with your eng team.
Intro call & req scoping
Pick a pack (e.g. FDE) + adjust
Candidate links & trace review
Limited pilot slots · scoped in days, not quarters.
Yes—that's the point. AI is in-bounds under fixed constraints. You measure how they plan, prompt, recover, and ship—not whether they hid a tab.