---
name: wack
description: Runs work on the user's wack boxes, persistent cloud Linux computers you drive through the wack MCP tools (box_*). Use it when a job needs a cloud computer or a clean Linux machine; should run in the cloud or split into parallel jobs; is too heavy for this machine (big builds, long test suites, Next.js or other dev servers); needs a headed browser, canvas or WebGL, a GPU (WebGPU, CUDA or ML jobs), or pixel-level visual testing and screenshots; should keep going while the user's laptop is closed; or whenever the user mentions wack, a box or boxes.
---

# wack

wack gives you your user's **boxes**: private Linux computers in the cloud. You are root in each,
files survive between sessions, and a box sleeps when idle and wakes in about a second on your
next call. Common tools are preinstalled: Python with `uv`, Node, Playwright and Chromium, office
and PDF tools, ffmpeg.

## Before the first job

1. **Check that the wack tools are available**: `box_docs`, `box_list`, `box_exec` and the other
   `box_*` tools (some agents list them by name and load them on demand). If they are not, stop
   and tell the user to connect wack to this agent at https://wack.sh/connect, then restart the
   agent. Do not guess a connect URL or make one up.
2. **Call `box_docs` once.** It returns the full, current manual: every tool, paths, limits and
   the rules of the box. It is the source of truth; this file is only the short version.

If `box_create` is missing, the connection reaches one box only: use that box.

## Which box for the job

| The job | Use |
| --- | --- |
| Scripts, data work, files, small builds, most tasks | An existing box from `box_list`, or `box_create` with the defaults (standard, no screen) |
| Next.js or other dev servers, big builds, test suites, several browsers at once | `size: "large"` (`"xl"` for the heaviest builds) |
| Anything that must run in a real window: headed Chrome, canvas/WebGL, clicking through an app, your human watching | `screen: true` (with `large` for heavy pages) |
| GPU work only: WebGL/WebGPU performance or pixel tests that need hardware rendering, CUDA or ML jobs | `size: "gpu"` (one RTX 4090 on spare capacity: it can be reclaimed at any time and is deleted when it stops). Download results as you go, and compare GPU box to GPU box |
| A quick picture of one page | No screen needed: `box_screenshot` on any box |
| The same environment every time, or A/B comparisons | `image`: the saved image from `box_list`, the same one for every box you compare. No image yet? Build one box with `setup`, then `box_image_save` |
| A private repo, a Chrome version, fonts | `setup` (`repos`, `chrome`, `fonts`) at create, or `box_setup` later |
| Many throwaway jobs in parallel | `ephemeral: true`, and delete them when done |

Start small: a bigger size or a screen costs more of your human's awake hours, so only ask for
them when the job needs them.

## Rules

- **Reuse before you create.** Check `box_list` first; a box the user already set up may hold the
  files and tools you need.
- **Verify artifacts, not exit codes.** Tools often exit 0 after writing nothing useful. Check the
  file itself (`test -s out.pdf && file out.pdf`, open the image) before you call it done.
- **Deliverables go in `/mnt/user-data/outputs`.** Files there show up on the user's dashboard.
- **Long jobs go in the background, and you poll.** One `box_exec` call is cut off at its timeout
  (900 s at most). Start longer work detached with its log in outputs, then check it:
  `nohup ./job.sh > /mnt/user-data/outputs/job.log 2>&1 &`, then
  `tail -n 20 /mnt/user-data/outputs/job.log`. A box sleeps after about 10 minutes without a
  call, which pauses background jobs, so poll at least every few minutes. For work that must run
  with nobody polling (the laptop closed, on a schedule), suggest a wack automation at
  https://wack.sh/automations.
- **GPU boxes are for GPU work only.** A `size: "gpu"` box never sleeps: stopping it, or 10
  minutes idle, deletes it with its disk, and it can be reclaimed at any time. `box_download`
  checkpoints, screenshots and results as you go. Run several at once with one `box_create` per
  box, in parallel. Compare GPU box to GPU box, never to a CPU box or a Mac; `box_docs` has the
  Chrome recipe (GPU boxes).
- **Delete what you made.** When you are done, `box_delete` the boxes you created (ephemeral ones
  delete themselves when they sleep; copy out anything you need first). Never delete a box you
  did not make without asking.
- **Keep secrets secret.** Never print the connect URL, put it in a file, or echo keys and tokens.
  If a job needs a key, ask the user for it.

## Example: A/B test a canvas change

The user wants to know whether a branch changes how a canvas renders.

1. `box_list`. If there is no saved image with the app's dependencies, make one box with
   `setup` (the repo, the Chrome version), check it works, then `box_image_save`.
2. `box_create` twice from that image, with `screen: true`, `ephemeral: true` and names like
   `ab-main` and `ab-branch`. Same image, so the only difference is the code.
3. In each, check out its revision, start the dev server in the background, and wait until
   `curl -sf http://localhost:3000` answers.
4. In each, open the page in headed Chrome on the screen (Playwright with `headless: false`),
   drive it to the same state and save a screenshot to `/mnt/user-data/outputs`. `box_screen`
   shows you the screen whenever you want to look.
5. Compare the two images pixel by pixel in one box (for example with `uv run --with pillow`),
   write a diff image to outputs, and verify every file is non-empty before reporting.
6. Report what changed, with the images, then `box_delete` both boxes.
