Which team should handle this?
Customer: I was charged twice for order 1182.
Real answers from decide on an M5 Max. Each took about 0.04 seconds.
Give it a question and your options. It picks one, on your Mac, in a blink.
curl -fsSL https://decide.run/install | sh
Needs an Apple Silicon Mac and about 9 GB of disk. It uses about 8 GB of memory, or 4.5 GB on 16 GB Macs. No API key, and nothing leaves your machine.
Use it wherever you'd write an if statement about text
- Route messagesSend each ticket or email to the right queue.
- Check rulesDoes this request meet the policy, yes or no?
- Steer an agentPick the next tool, or decide whether to hand over to a person.
- Label a backlogRun thousands of decisions from a file in one go.
$ decide ask "Policy: refunds need a receipt. Customer has none." \ "Is a refund allowed?" "allowed=refund is permitted" "denied=not permitted" denied denied 0.999 ████████████████████ allowed 0.001 ░░░░░░░░░░░░░░░░░░░░
# -q prints only the answer, for scripts $ team=$(decide ask -q "$(cat ticket.txt)" \ "Which team?" billing shipping tech) # read the context from a pipe $ cat email.txt | decide ask - \ "Does it ask for a refund?" yes no # many decisions from a JSONL file $ decide score decisions.jsonl
Keep Claude Code's subagents on the cheapest model that can do the job
A small hook asks decide what each new subagent's task is, before it starts, and picks the model to match. It adds about a tenth of a second, costs nothing, and never blocks an agent: when decide isn't sure, the task runs as Claude asked.
Simple lookups go to Haiku, everyday coding and anything written for people go to Sonnet, and hard reasoning goes to Opus. Every choice is logged, so you can check its judgement. Set it up.
| Count the test files | Haiku | simple, 0.997 |
| Find the handler for POST /api/orders | Haiku | simple, 0.987 |
| Add cursor pagination with tests | Sonnet | standard, 0.999 |
| Write a commit message for the rename | Sonnet | writing, 0.986 |
| Find the race that double-charges orders | Opus | hard, 0.986 |
As good as the big models on everyday decisions
Scored on the 231 public JevBench tasks. Short tasks are routing, policy checks, intent and extraction. Hard tasks are long documents and multi-step reasoning.
| Model | Short tasks | Hard tasks | Typical time |
|---|---|---|---|
| decide, 4B model on your Mac | 119 of 120 | 67 of 111 | 0.08 s |
| Jev 1.13, hosted API | 119 of 120 | 79 of 111 | 0.62 s |
| Open-Jev 27B, same Mac | 117 of 120 | 81 of 111 | 1.36 s |
Pick a size for your Mac
decide can load the model at lower precision to use less memory. It picks 8-bit on Macs with 24 GB of RAM or less; choose for yourself with decide start --bits 16, 8 or 4.
| Precision | Memory | Short tasks | Hard tasks | Typical time | Good for |
|---|---|---|---|---|---|
| Full (16-bit) | 8.0 GB | 119 of 120 | 67 of 111 | 0.08 s | Macs with 32 GB or more |
| 8-bit | 4.4 GB | 118 of 120 | 67 of 111 | 0.06 s | 16 and 24 GB Macs |
| 4-bit | 2.5 GB | 117 of 120 | 63 of 111 | 0.07 s | When memory is tight; weaker on hard tasks |
Times in the table were measured on an M5 Max. Smaller chips are slower: a short decision takes about 0.04 s on an M5 Max and about 0.35 s on a base M4 Mac mini (16 GB, 8-bit). The first decision after starting takes about twice as long. Accuracy is the same on any Mac.
Use a bigger model for the hard stuff. On long documents and multi-step reasoning, decide gets about 60% where the large models get 71–73%.
Trust the choice, check the confidence. The probabilities aren't calibrated, so test them on examples you've labelled before relying on them.
Name options clearly. A small model reads your wording closely. Describe options plainly and keep their order consistent.
What you're installing
- The installer
- Checks for Apple Silicon and Git, installs uv if needed, puts decide in its own Python environment, downloads the model once and runs a test decision. Read the script.
- The model
- Qwen3.5-4B, unmodified, pinned to one revision and run with Apple's MLX. decide reads the model's preference for each option directly; no text is generated.
- The background server
- Keeps the model loaded so answers come back instantly. It starts on first use, listens only on
127.0.0.1, anddecide stopfrees its memory. - For apps
- POST decisions to
localhost:8792/score, or send Jev-style requests to/v1/systemone. - Removing it
- Run
decide stopanduv tool uninstall decide-cli, then delete~/.decideand the model in~/.cache/huggingface.