Harnesses > Workflow methodologies
pstack
Draft. This entry has not been checked against its sources yet.
A Cursor plugin in which one command, /poteto-mode, routes each task to one of 23 playbooks that require reproducing problems first, verifying on the real surface, review by several models, and small stacked pull requests.
| Maintainer | Lauren Tan (poteto), published in Cursor's plugin repository |
|---|---|
| Category | Workflow methodologies |
| Hosts | Cursor |
| Components | 47 skills, 2 subagent definitions |
| License | MIT |
| Version | 0.15.5 |
| Source code | https://github.com/cursor/plugins/tree/main/pstack |
| Website | https://cursor.com/marketplace/cursor/pstack |
| Docs | https://github.com/cursor/plugins/tree/main/pstack/docs/guide |
| DeepWiki | https://deepwiki.com/cursor/plugins |
| Last researched | 2026-09-28, at commit ecc249f |
Features
Definitions are on the compare page.
| Feature | Supported | Notes |
|---|---|---|
| Stages of work it covers | ||
| Brainstorming | partial | No brainstorming skill, and the rules discourage asking what the agent can check itself. figure-it-out presents framing and tradeoffs before a long run. |
| Written plans | partial | Only the multi-phase-plan playbook writes a plan, to Cursor's agent store by default rather than the repository. README: 'the best spec is code.' |
| Test first | partial | tdd runs only when asked or when a cheap test exists. Bug fixes must commit the failing repro before the fix. |
| Debugging process | yes | bug-fix playbook: reproduce, binary-search hypotheses against runtime evidence, revert what a refuted hypothesis motivated. |
| Code review | yes | interrogate runs one read-only reviewer per configured model, three by default. Shipping needs a verdict from an agent that did not write the code. |
| Verification | yes | principle-prove-it-works: verify against the real artifact. Replies must label every claim as measured, inferred, or a guess. |
| Learning capture | partial | /reflect turns a transcript into proposed skill edits, applied only after you approve. Nothing is loaded at session start. |
| Branch to PR | yes | opening-a-pr, babysit, and shipping: worktree off main, small ordered commits, narrow PRs, independent verification before squash-merge. |
| How it runs | ||
| Single entry point | yes | /poteto-mode matches the task against 23 playbook descriptions and copies the chosen playbook's steps into the todo list. |
| Automatic activation | no | 46 of 47 skills are marked disable-model-invocation. You start with /poteto-mode, which then stays on across turns. |
| Session-start hook | no | No hooks. /setup-pstack writes an always-applied rule that holds model choices only. |
| Subagents | yes | |
| Parallel agents | yes | Parallel explorers, reviewers, and design candidates; some playbooks run one Cursor cloud agent per pull request. |
| Git worktrees | yes | Work happens in a worktree off main; parallel design candidates each get their own. |
| Adopting it | ||
| Multiple hosts | no | Cursor only. Community ports exist for Claude Code, Codex, OpenCode, Gemini CLI, and Pi. |
| Per-project setup | yes | Enable in the committed .cursor/settings.json. The model rule is always per user. |
| Team extensions | partial | /automate-me writes your own <handle>-mode skill to use alongside it. Adding playbooks to poteto-mode means forking. |
| No extra services | yes | Some steps use /deslop and control skills from the separate cursor-team-kit plugin, and some playbooks use Cursor cloud agents. |
How it works
The workflow
You start with /poteto-mode <goal>. The skill reads its own index of 23 principles, matches the task against 23 one-line playbook descriptions, and opens a todo list whose first items are the chosen playbook's steps "copied in verbatim". A step the agent decides to skip stays in the list with a one-line reason. Large or cross-cutting work, or work you plan to leave running, goes to figure-it-out, which writes a one-off playbook for the task. Once entered, poteto-mode stays on for later turns until you opt out.
The playbooks cover investigation, bug fixes, performance problems, hill-climbing on a metric, runtime and trace forensics, features, refactoring, prototypes, visual parity, authoring skills, evals, babysitting and shipping pull requests, autonomous runs, orchestrating multi-day programs, pausing and resuming, multi-phase plans, worktree cleanup, and opening a PR. Every code playbook ends with Opening a PR.
Two examples:
- Feature. Explain the affected subsystem (
how), explore designs in parallel (architect, which runs at least two structurally different candidates througharena), write a short throughput checkpoint, delegate the code to a subagent, verify on the surface the user will see, rebase into small ordered commits, runinterrogateif the design is contested, then open the PR. - Bug fix. Reproduce it yourself, binary-search the cause with hypotheses checked against runtime evidence, plan the fix, verify on the same surface, and commit the failing reproduction before the fix so the history shows both.
The author does not plan by default: "i don't believe in planning. the best spec is code." A plan document appears only in the multi-phase-plan playbook, and a script (check-plan.mjs) checks its structure.
Components
- 47 skills. 24 workflow and utility skills, including
poteto-mode,figure-it-out,how,why,architect,arena,swarm,interrogate,tdd,unslop,reflect,automate-me,show-me-your-work, andcreate-verification-skill. 23 principle skills, each one rule of 16 to 34 lines, in five groups: core, architecture, verification, delegation, and meta. - 23 playbooks, stored as Markdown files inside
poteto-moderather than as separate skills. - 2 subagents:
poteto-agent, the default for any subagent spawned inside a playbook, which reads poteto-mode in full before working; andComment Sicko, which removes needless code comments. - No commands and no hooks. Skills are invoked by name as slash commands.
- Scripts:
watch-pr(a Bun CLI that watches GitHub PR, check, and review state),orch(the store behind the Orchestrate playbook),check-plan.mjs,worktree-audit.sh, and a decision-log writer. - A dormant automation pack,
benny, for triaging Slack issue reports with Cursor Automations. It needs setup per repository.
How it steers the agent
All steering is text or tooling; there are no hooks.
- Playbook steps are copied into the todo list, so skipped steps are visible.
- A "Non-negotiables" section, and phrases such as "Mandatory: no skip-with-reason escape".
- Every reply must name the principles that shaped a decision, citing only principles whose skill file was read in the session, and must label each claim as measured, inferred, or a guess.
- Reviewers from several model families:
interrogategives the same prompt and rubric to each, on the view that "the adversarial signal comes from model diversity, not assigned personas." Merge verdicts are tied to a commit and are voided when the patch changes. - The autopilot playbooks re-read themselves from the main branch every 30 minutes and count "only side effects as progress".
- It favors acting over asking. The agent proceeds on reversible work and always pauses for irreversible writes, such as force-pushing shared branches, deploys, data deletion, and customer messages.
Files it writes
~/.cursor/rules/pstack-models.mdc: the model to use for each role and a reasoning budget, written by/setup-pstackand applied to every new session.- A decision log (
decisions.tsvor.audit/<task>.tsv), kept out of git by default. A later session picking up the work reads it. - For Orchestrate, a store in Cursor's agent directory with units, a ledger of verification verdicts, gates, and preferences.
.cursor/skills/verify-<app>/: a project skill that knows how to launch, drive, and check your app, created bycreate-verification-skill.- Your own
<handle>-modeskill, if you run/automate-me.
There is no standing memory file. Lessons go into skill edits through /reflect, and only with your approval.
Install
In Cursor, run /add-plugin pstack, then /setup-pstack to pick models per role and a reasoning budget, then start a new chat. For the full set of steps, also install cursor-team-kit, which provides /deslop, control-ui, and control-cli; pstack refers to them but does not include them. To enable it for a whole team, commit {"plugins": {"pstack": {"enabled": true}}} in .cursor/settings.json.
pstack only runs on Cursor. It depends on Cursor-specific tools: subagent options such as cloud execution, /loop, /goal, .mdc rules, and Cursor's transcript files. Community ports, none maintained by the author, include pstack-claude (Claude Code, Codex, OpenCode, Gemini), pstack-pi (Pi), and pstack-generic.
Scripts need git, gh, bun, and node. worktree-audit.sh assumes macOS versions of stat and date.
When to choose it
- Your team uses Cursor and wants one command that picks a rigorous playbook for each task.
- You want evidence over assertions: reproduce first, verify on the real surface, and review by more than one model.
- You run long or unattended work in parallel and land it as small stacked pull requests.
When not to
- You use a host other than Cursor. You would depend on community ports.
- You want a brainstorming step or written plans by default. The author rejects default planning.
- You want test-first on every change. TDD here is conditional.
- You want the agent to ask before acting. The defaults favor acting on anything reversible, including team chat and ticket updates.
- You need to keep token use low. The default setup fans out to three reviewers from different model families, and some playbooks run a cloud agent per pull request.
Sources
- cursor/plugins at ecc249f,
pstack/:README.md,docs/guide/,skills/,agents/,.cursor-plugin/plugin.json. The last commit touchingpstack/before that snapshot is12d587d(2026-09-23), per the commit history. - Cursor Marketplace listing and the author's post How I Use Cursor (linked, not read: blocked from the research environment)
- DeepWiki was not used: it was unreachable from the research environment.