Practice on real codebases
Every engineer getsa real codebase
Groundwork hands you a ticket and an isolated machine with the repository, a browser IDE and a terminal. You make the fix. Hidden checks evaluate it. Nothing to install.
How evaluation worksLive tasks · TypeScript and Python · hidden checks on every submit · runs in your browser
- TypeScript
- Node.js
- Express
- Jest
- Python
- pytest
- VS Code
- Linux sandbox
- Hidden checks
- Pair
- TypeScript
- Node.js
- Express
- Jest
- Python
- pytest
- VS Code
- Linux sandbox
- Hidden checks
- Pair
Where Groundwork fits
You do the engineering.We run the machine.
One realistic ticket, one isolated machine, one honest evaluation. In our cloud, from your browser.
[ The ticket ]01 / 04
A real task, not a puzzle.
Bug reports · failing suites · small features
- Repository
- an unfamiliar codebase, pinned
- Ticket
- what is broken and how to tell
- Budget
- 30 to 35 minutes on the clock
Where Groundwork fits
You do the engineering.We run the machine.
One realistic ticket, one isolated machine, one honest evaluation. In our cloud, from your browser.
[ The ticket ]01
A real task, not a puzzle.
- Repository
- an unfamiliar codebase, pinned
- Ticket
- what is broken and how to tell
- Budget
- 30 to 35 minutes on the clock
[ The machine ]02
A full machine per attempt.
- Editor
- VS Code in the browser
- Terminal
- a real shell in the sandbox
- Network
- no internet; dependencies baked in
- Pair
- explains the code, never writes the fix
[ Through the boundary ]03
Your diff out, hidden checks in.
- Hidden checks
- you see what they test, never the inputs
- Existing tests
- must keep passing
- Time limits
- slow code is reported as slow
[ The ledger ]04
Evidence, not a score.
- Diff
- exactly what you changed
- Checks
- each one passed or failed, by name
- Duration
- time against the budget
Tasks
Real ticketsfrom real industries
Retail stock, room scheduling, event logs, APIs. Bugs written up the way a team would hand them to you, in code you have never seen.
The workspace
Everything an engineerneeds at the desk
One isolated machine per attempt, with the tools around it: an editor, a terminal, three tiers of tests, a pair, and your diff.
❯ pytest tests/visible
src/ · tests/ · conftest.py
Editor and terminal
VS Code and a real shell in the browser, on an isolated machine with the repository.
- ✓Examples2 of 2
- +Your cases+3 added
- !Extended11 of 12 · 1 too slow
- Hiddenon submit
Tests in three tiers
Examples you can read, extended cases with time limits, hidden checks on submit.
Why does a retry pay out twice?
Find where a retry decides the first attempt failed. What does it check, and what happens if the first call was only slow?
GuardedPair
An AI pair that reads the code with you. In guarded mode it explains and points; it does not write the fix.
src/payouts/retry.ts
- if (!result.ok) await send(payout)
+ if (await alreadyPaid(payout.id)) return
+ if (!result.ok) await send(payout)
return result
Diff
See exactly what you changed before you submit it.
- Run
- Pausesaved
- Resume22:14 left
- Submit
Save and resume
Code autosaves. Leave, come back, and carry on with the time you had left.
Evaluation
Evidence,not a score.
Submitting is not the end of the attempt. See what passed, keep working, and submit again. Every result lands on your profile with its diff.
- 01
Hidden checks
Run when you submit. You see what each one tests and whether it passed, never its input.
- 02
Existing tests
The repository’s own tests run again. What passed before must still pass.
- 03
Time limits
Every check has one. Slow code is reported as too slow, not as wrong.
- 04
No ranking
No composite score, no leaderboard. Each row is something that happened.
Repair duplicate payout retries
Build
Compiled cleanlyPassed
Correctness
Hidden checks passed8 of 8
Regression avoidance
Existing tests still pass12 of 12
Duration
Inside the time budget31:04 of 45:00
Files changed
Source changed4
No composite score. Every row is something that happened.
FAQ
Questionsworth asking
LeetCode trains isolated puzzles. Groundwork trains the work around them: reading code you did not write, debugging a real service, repairing tests, keeping existing behavior intact. One repository, several realistic tickets.
Pair is built in. It reads the repository and explains: where something lives, what an error means, how to approach the ticket. By default it is guarded and will not hand you the fix; you can switch it to an ordinary assistant, and that choice is part of the record of the attempt.
Junior to mid-level web and backend engineers who want more time inside unfamiliar codebases, and anyone who wants to contribute to a large project but does not know where to start.
When you submit, hidden checks test the behavior the ticket asked for, each with a time limit, and the repository's existing tests run again and must still pass. You see what each check tests and whether it passed, never its input, and never a single grade. You can keep working and submit again.
Nothing to start. Create an account, pick a task, and the full loop is available: repository, workspace, Pair, and evaluation.
No. The editor, file tree, and a real terminal run in your browser, connected to an isolated environment for that task.
Your code, test runs, submissions and Pair conversations, plus events such as pasting into the editor or switching tabs (counts, never the text and never keystrokes). The privacy policy lists all of it.
Give yourselfa real ticket
Pick a task, open the workspace, and ship the fix. Free to start.