Memory that survives a closed session, and context that costs less.
Persistent memory across sessions, smaller context, an official Anthropic plugin that audits your project, and a skill that learns from your corrections. What each one does, the install commands, and what it can see.
Free. No email required, nothing to download.
Every one of these is free and takes minutes to install, which is exactly why people end up with five of them and no idea what any of them can see.
So the rule first, then the tools. Install one at a time, for a problem you have already hit, and read what it can access before you point it at real work. A Claude Code add-on can read your files, your context and your credentials, and can talk to services outside your machine. That is what makes them useful and it is also the whole risk.
The add-ons are free and open source. Claude Code itself, and any third party model provider you wire in, still cost whatever they cost. Two of these route your work somewhere: Headroom through a proxy on your own machine, and Claude-Mem through a model to compress it. Both are covered below.
Start here, because it tells you which of the rest you actually need.
Claude Code Setup is published by Anthropic on its own plugin marketplace. It reads your codebase and recommends the automations that fit it: skills, MCP servers, hooks, subagents and slash commands. A React project might come back with a browser testing MCP. A data project will not.
It runs read only and does not modify your files, which makes it the safest thing on this page to try first.
Run it against a real project rather than an empty folder. On an empty folder it has nothing to reason about, which is how people conclude it does nothing.
Close a Claude Code session and the useful context goes with it. Next time you are re-explaining the project, the decisions you already made, and the three approaches that did not work.
Claude-Mem captures what happens during a session, compresses it, and makes the relevant parts available again in later sessions. So Claude keeps hold of your project, your previous decisions, the work already done and the problems you hit.
# the quick way npx claude-mem install # or from inside Claude Code /plugin marketplace add thedotmack/claude-mem /plugin install claude-mem
Restart Claude Code after installing, or it will not pick up the new configuration.
Before you point it at client work, know where the memory goes:
Worth knowing: Claude Code already reads a CLAUDE.md file in your project, and that is the right home for durable facts like commands, architecture and conventions. Claude-Mem is for the session history a file like that will never hold.
More context is not better context. Agents burn tokens on huge tool outputs, logs, files and JSON where most of it changes nothing about the answer.
Headroom compresses what your agent reads before it reaches the model. It ships as a library, a local proxy and an MCP server, and the proxy is the version that drops in front of Claude Code.
# install the CLI uv tool install --python 3.13 "headroom-ai[all]" # start the local proxy headroom proxy --port 8787 # point Claude Code at it ANTHROPIC_BASE_URL=http://localhost:8787 claude
There are pip, npm and Docker routes too. Check the README for the current invocation, because the project moves quickly.
The numbers it reports, by workload:
Your mileage depends entirely on what you do all day. Tool heavy and data heavy work saves most, and a conversation about copy saves almost nothing.
Compression runs on your machine. The project is explicit that no prompt or file content is sent anywhere to be compressed, which is the opposite trade-off to Claude-Mem. Think of the pair as remember what matters, and stop sending what does not.
The most interesting one, and the slowest to pay off.
Most people build an AI workflow and never touch it again, so they keep making the same corrections forever. Task Observer watches multi-step work and records what happened: the corrections you repeat, the gaps where no skill covers a task you keep doing by hand, what worked, and where it was wrong itself.
You review those observations and fold them into better skills. Claude does the task, you correct it, the correction gets observed, the skill improves, and you correct it less next time. It compounds when you are reusing the same skills, and does nothing much if every job is a one off.
npx skills add rebelytics/one-skill-to-rule-them-all --skill task-observer
Installing the files is not the same as switching it on, and this is where most people think it is broken. You also have to add the activation instruction from references/environments.md to your CLAUDE.md, or install the session start hook. To check it is actually running, do a few sessions of real work and look for a skill-observations/observation-log/ folder.
Observations land in skill-observations/observation-log/ and proposed changes in skill-updates/, both in your shared folder. On claude.ai it hands you a structured document at the end of the session instead.
The one you have seen in every thumbnail, and the one with conditions attached. Read the whole section before you install it.
OmniRoute is an open source AI gateway. It puts hundreds of providers behind a single endpoint, falls back automatically when one hits a quota, and pools the free tiers of the providers you connect. It is MIT licensed, has over 70,000 stars, and is developed daily. It is a real project, not a wrapper someone threw together.
The genuinely clever part is that a fresh install answers with no key at all, because a keyless provider is pre-wired into its auto routing.
# install and run npm install -g omniroute omniroute # dashboard: http://localhost:20128 # api: http://localhost:20128/v1
Then connect a provider in Dashboard, Providers, and point Claude Code at the endpoint. The base URL is http://localhost:20128/v1, the key comes from Dashboard, Endpoints, and the model auto lets it choose.
If you would rather run it in Docker, use their own command, because it binds to loopback rather than every interface:
docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \ -p 127.0.0.1:20128:20128 -v omniroute-data:/app/data \ diegosouzapw/omniroute:latest
One practical gotcha before the security one. The Docker image ships a 1024 MB heap, which is fine for the dashboard and light chat but not for coding agents. Claude Code sends long overlapping contexts and the process can die outright. Raise OMNIROUTE_MEMORY_MB and give the container more than 2 GB if you are actually routing Claude Code through it.
Now the part that decides whether you should run this at all.
OmniRoute has an open critical advisory, CVE-2026-88062, scored 9.5. Its custom agent endpoint takes a command from the request and runs it, and the only check is that two attacker-supplied fields agree with each other. Code then runs in the same process that holds every provider API key you connected.
The advisory covers every version up to and including 3.8.50, which is both the newest release and the newest published package, and it lists no patched version. We checked the unreleased 3.8.51 branch on 29 September 2026 and that endpoint is unchanged.
The unauthenticated path is narrower than the headline suggests. It needs one of two things: login switched off, or a brand new instance that has no management password yet. The default is login on, so the exposure is mostly self inflicted. That makes these steps non-optional rather than nice to have.
Before anything else, and before you connect a single provider. Until you do, the instance is in the bootstrap window where setup writes are accepted without credentials, which is one of the two ways the flaw is reachable.
It is on by default. Turning it off for convenience is the other way in, and it is exactly what people do on a local tool. Do not.
Bound to 127.0.0.1 and nowhere else. Not on a VPS, not behind a tunnel, not on the office network, not on your phone. If it is reachable, so is the endpoint.
Because it does. Connect throwaway provider accounts rather than keys that are attached to billing you care about.
And be clear about what the free token figure means. It adds up the free tiers of many separate providers, it is not extra Anthropic capacity, you create an account and a key with each one, and using their free tiers to drive a coding agent may breach their terms. It is real, it is just not what the thumbnails imply.
Our honest read: worth running on a machine where nothing matters if you want to see what pooled routing feels like. Not worth pointing at client work until a fixed release ships.
You do not need all four, and adding them all in one afternoon is how you end up unable to tell which one broke something.
First, against a real project. Read only, made by Anthropic, and it tells you what the project actually needs.
Once you are regularly picking work back up across several sessions. That is the point the re-explaining gets annoying enough to fix.
Once you have skills and workflows you reuse. Before that there is no pattern for it to observe.
When context size or token spend has become a real problem, not before. It is the most involved to set up and the easiest to skip.
Last, and only if you specifically want to experiment with routing through other providers. Read its section first, because it is the only one here with an open critical advisory and a setup you have to get right.
Stop at whichever one solved your problem.
None of this is paranoia about open source. It is that a Claude Code add-on runs with your access, to your files and your keys, and "it was free on GitHub" is not a security model.
Two guides that pair with this one:
We build content systems for small businesses that are not AI native. Tell us your content problem and we will come back within 48 hours with what yours would include.