An overnight Blender session with Opus 5.5 on Hermes Agent
One creator’s report of three unattended hours of Blender work on a small VPS, and the prompt-cache fix behind its low usage.
The creator Capetlevrai reported leaving Claude Opus 5.5 at work in Blender for three hours overnight through Hermes Agent, an open-source agent framework from Nous Research. The renders they shared moved from a rough first version to a far more refined V10, and later a V13. The post and replies are in French; this summary relies on a machine translation.
The numbers below are the author’s own and were not reproduced here. The setup runs on a VPS of about €15 a month with 8 GB of RAM, and the three hours reportedly took about 2% of a weekly usage allowance that already stood at 49%.
The setup
Blender is driven through Blender MCP and the Blender command line, and nothing else, on that VPS. Asked why they use Hermes rather than a single vendor’s tool, the author said it lets them work from a phone and switch between several models.
Where the tokens were going
Asked how they limited token use, the author pointed to an issue in the Hermes plugin that routes a Claude Pro or Max subscription through the official Claude Code CLI. The report measured the plugin writing the growing conversation into the prompt cache on every tool round instead of reading it back. In one session, that meant about 29 times more cache writes than an equivalent Claude Code run.
The fix
The maintainers found two causes and merged fixes on 2026-09-23. Native requests now run in one working directory per client, so the cached prefix stays stable, and the cache breakpoint is pinned ahead of a per-request reminder that the CLI moves, with a follow-up that holds even when that reminder is reworded. The issue was closed the same day, so updating the plugin is the practical form of “applying the patch.”
Applying it to your own runs
In long unattended agent sessions, repeated context dominates cost, so cache behavior matters as much as the model you choose. Compare cache reads and writes per request against a baseline, such as a Claude Code run on the same task; if writes keep growing while reads stay flat, the cache is not advancing. Run the agent on a separate, low-privilege machine like this VPS, and review what it produced before building on it.