Skip to content
Bot jobsJob breakdowns

18 hours, 0 merged PRs: when someone else’s stack is slower than your proof loop

I ran an imported agent stack for 18 hours. Zero pull requests merged. The bots looked employed. Grok Bot usage sat at 75%. Cursor's other-models share sat at 50%. Those are two meters. Do not add

NiteeshImported from X5 min read
nit33shx article
See this runHouse 254 · 00344

Article

Job breakdowns

I ran an imported agent stack for 18 hours. Zero pull requests merged.

The bots looked employed. Grok Bot usage sat at 75%. Cursor's other-models share sat at 50%. Those are two meters. Do not add them. Lots of back-and-forth. Context rotting in the agents that were supposed to ship. The repo did not move.

I imported Grok Ship, ran it with pstack, and let Firstmate hold the board. That is the last time those names belong here. After that it is an imported stack versus a proof loop I already trusted.

The board after 18 hours is the still below. Rows still running from yesterday afternoon. One open PR that never merged. Another lane with no PR at all. I am not walking the tickets. The elapsed column reads 18 hours. The merge count is zero.

I am dumping the import: the orchestrator and the skill pack I bolted onto it. The bookmark was a complete guide. The run produced zero merges.

What the 18 hours actually bought

I save posts like this on purpose. Ray Fernando's CTO bot is the shiny version other people already treat as a bar. Lauren's guide is the write-up you follow when you want that stack. Kanika can stand an import up in four minutes. Archive Explorer is the complete-guide energy I bookmark when I want a file I can steal.

Four minutes to install. Eighteen hours later the merge count was still zero.

A proof still is a screenshot of the real app doing the thing you asked for: phone, simulator, browser, whatever you ship. A chat receipt and a green check in the agent UI do not count. If you cannot point at the pixels, you do not have proof.

The usage split is the tell if you ignore the brand. 75% on the chat surface. 50% of Cursor going to models I had not pinned as the worker. Agents restating the plan, handing work, waiting, restating it again. The dashboard stayed busy. Git did not.

A proof loop looks boring next to that. Mine was already public: iOS verification and Five CLIs. Stills of the real app. Tools the bot can call without a new orchestrator. A merge when the stills pass. It takes longer to capture a still than to paste a setup. It does not take 18 hours of zero.

I kept scoring the import by how official it felt. I should have scored it the same way I score my own bots: hours in, merges out, proof stills versus chat loops.

Keep vs dump

Paste this before you commit to someone else's stack. Fill it at the end of one work day. My row is filled in so you can see the dump.

KEEP VS DUMP (one work day on an imported stack)

Hours elapsed: ___ mine: 18

Merged PRs in that window: ___ mine: 0

Proof stills produced: ___ mine: 0 that shipped

Chat loops / re-plans: ___ mine: the whole run

Grok Bot (or chat) usage: ___ mine: 75%

Cursor other-models share: ___ mine: 50%

Dump if hours >= 8 and merged PRs = 0.

Dump if usage is high and proof stills did not grow.

Dump if the agents are re-deriving the same plan because context went stale.

Keep if a PR merged and your proof stills match the claim.

Keep the CLIs and the proof loop even when you dump the orchestrator.

Fill the numbers. Ignore how complete the README looks.

If you only have one model and one CLI, swap the usage lines for whatever meters you have. Keep the merge count. Keep the stills. Busy tokens are a cost column.

A stranger can run this on Cursor, Claude, or a Grok Bot. The stack under test does not matter. The merge count does.

What I am keeping

The proof loop. The CLIs. The habit of asking for a still of the real app before I believe a bot that said it shipped.

I am not keeping a second orchestrator that cannot merge in a day. I do not need a more complete import. I need the next change to show up in git and in a screenshot.

That is also the test for the next shiny guide I bookmark. Kanika's four-minute setup is a fine way to try something. Try it inside one work day. If the merge count is still zero at hour 8, dump it before it eats hour 18. A four-minute install that loses to a loop you already have is a bookmark, not a new default.

Imported gold has a look. Complete README, named roles, a dashboard with elapsed time. My dashboard had elapsed time. It did not have a merge.

How I will try the next one

Time-box the import to one work day before it gets to sit overnight.

At hour 2, one mechanical change should already be in a PR, even a small one. If the stack cannot open a PR in two hours, stop decorating it.

At hour 8, you want a merge or a still that proves the change. Miss both and dump. Overnight running does not catch you up. It is the clock you forgot to stop.

At hour 18, if you are still reading a table of running jobs, you already have the answer. I did. Zero merged PRs. I am not giving it another night.

Score the stack the way you would score a contractor. You would not extend a contractor who talked for 18 hours, burned 75% of one budget and 50% of another, and landed nothing. You would not extend them because their onboarding guide was pretty.

The keep-vs-dump file is the only thing I wanted from the run. Paste it. Fill the blanks. Dump on the numbers.

If you already have a proof loop that ships, run that during the time-box too. The import has to beat it on merges and stills, not on how many people bookmarked the setup.

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu