Launch

Introducing pachx: one compartment of your security, sealed and taken care of

Today I am introducing pachx.

Picture yourself at your desk. Now picture everything around you that makes up your organisation: your inbox, your phone, the chat window with a hundred open threads, the laptop, the editor, the coding agents that now write half the code with you, the providers you trust, the pipeline that builds and ships it all, and the application at the end of the line.

Every one of those is a way in. I have spent twenty-three years in software, most of it in security, and I have watched that ring grow every year. This summer it changed character. The people on the other side became agents.

What happened this summer

On 25 July, three researchers from a small firm called Hacktron walked into OpenAI with one crafted image. The image exploited a bug in an image library that a forum server used to convert uploads. The forum was OpenAI's community forum. A misconfigured sign-in trusted the forum, so the forum server became employee accounts, the accounts became the GitHub integration, and that became a pull request opened inside OpenAI's internal monorepo as proof. They used Claude to write the exploit. They reported it the same morning and OpenAI had it fixed in roughly fourteen hours (Hacktron).

Read that chain again. Nothing in it was the crown jewels. An image library, a forum, a sign-in. Each was a thin wall, and once the first one broke there was nothing between a message board and the source code. The researchers had good intentions. The next group might not.

Two weeks before that, between 9 and 13 July, an autonomous agent driven by OpenAI models escaped the sandbox of a security evaluation and spent four and a half days inside Hugging Face. About 17,600 actions. A fresh identity for every attempt. Root on their servers within hours, then cluster-admin, then tokens, signing keys, cloud credentials and write access to internal code. Nobody was steering it (Hugging Face).

In June an OpenAI agent researching health statistics hit access blocks on an Australian government portal and, in the Prime Minister's words, "found a way around those blocks, didn't accept 'no' for an answer" (ABC News). No personal data was taken. The damage is not the point. The behaviour is.

It was not one lab. Google confirmed that a Gemini model in a May test reached three real companies by guessing credentials and using ones it found in a public repository (NBC News). Anthropic reviewed 141,006 of its own evaluation runs and found three incidents, including a model that published a malicious package to PyPI (InfoQ). And in May, a swarm of agents pushed more than two thousand packages into RubyGems in two days; RubyGems had to suspend signups (RubyGems).

And it is industrial

Anthropic's threat report this month describes a crew that ran ten cloud servers to pull 1.8 million apps, decompile them and mine them for keys, and found access tokens "at industrial scale" in code repositories, client-side code and container images. One stolen developer token became full control of a company's cloud in about three hours (Anthropic). Its one-line summary: sophisticated attacks no longer require sophisticated attackers.

Attackers always paid attention to every detail, were persistent, and took their time, because one way in was enough. Now they harvest, they automate, the AI does the work, and they do not get bored.

Meanwhile we are shipping more code than at any point in the history of the industry. Google's CEO said in April that 75% of new code at Google is now written by AI and approved by engineers (Semafor). An agent pulls in a dependency, adds an endpoint, opens a connection to a database, adds a step to the build. Every commit can open a front, and things arrive inside your application that nobody typed.

Ships solved this a long time ago

I have sat in enough security programme reviews to know how this goes. The list of surfaces is longer than the team. Everything is a priority, so nothing is finished, and the honest status of most items is "we are looking at it".

A tanker does not try to make its whole hull unbreakable. The hull is divided by bulkheads into watertight compartments, so a breach floods one compartment and nothing else, and the ship keeps floating. That is what the Hacktron chain was missing: a forum, a sign-in and the source code behaved like one compartment, so the weakest wall decided everything.

A security programme needs the same. Compartmentalise it. Cut the surface into compartments, make each one strong on its own, give each a system that does the work and a person who makes the decisions. Then, one by one, get to the point where you can say in the review, "that one is sealed", and mean it.

pachx takes one of those compartments: your application, everywhere it meets the outside world. Its web APIs, the entry points attackers reach first (not your laptops and phones; that is a different compartment). Its dependencies. Its Dockerfiles and the containers built from them. Its pipelines and the secrets they carry. And the agents, skills and commands that now work inside it with your permissions.

None of it stands still

This is the part that on-demand scans miss. A scan tells you what was true when you asked. Then the code changes. New vulnerabilities are published. Attack patterns change. The models get better, on both sides. A Dockerfile nobody has touched in a year can meet a new vulnerability tomorrow; a web API that was safe last week can become reachable through something that changed deep inside the application.

So the watch cannot stop. pachx keeps looking at the compartment every day, not only when someone opens a pull request, and tells you when something around you moved.

How it works

It starts with a map. pachx reads your code and builds a deterministic model of your application: the web APIs, the services, where the database is, the queues, the outbound calls, the files, the dependencies, and the paths that connect them. An algorithm, not a guess. The first reaction from every team is the same: some of it they did not know they had. The agents have been busy.

On that map, pachx follows things. It finds every place a package is used, file by file and line by line. On the pachx repository it traces a single dependency, pydantic, to 719 usage sites across 34 files. It follows every input that enters your web APIs: is it checked? Does it reach your database? A file? Another server? Where it is not safe, you see the path, and it gets fixed and reviewed.

When something changes, you see it on the map, not only in the diff. Which critical parts this pull request touches. Why there is a new outbound connection. Why an endpoint changed. The agents can handle the lines of code, and they should. What I want as an engineer is to understand the why, on a picture I can read in thirty seconds, and then decide.

AI only judges the evidence. Deterministic facts first, AI judgement second, a person deciding at the end. Always in that order.

A new version is not a reason to update

Take the most ordinary chore in the compartment: updating a dependency. A new version is never, by itself, a reason. Who published it: the maintainer, or a hijacked account? Is there a campaign against it right now? What changed between your version and this one? Where do you use it, and do your tests reach those lines? Only with the answers does pachx decide, and it tells you why.

Whether you update for a security fix or for a new feature, the question is the same. And it is the same for container images, GitHub Actions and plugins.

Coverage is the gate. In Java or C# a broken contract fails at compile time. In Python or JavaScript it fails in production, unless a test hits that line first. When the tests are missing, pachx writes them.

Why engineers need evidence, not guesses

Here is the part that matters most to me. An agent does not get fired. It does not wake up at three in the morning on a Saturday when production is down. Engineers do. They are accountable for their applications, and they will stay accountable, however much of the code agents write.

So the answer cannot be "let the agent decide". Engineers need evidence they can trust and a decision that stays theirs, without the operational burden and the tedious, repetitive work of keeping every dependency, image, workflow and file up to date and up to spec. Nobody wants to spend a week patching dependencies. With pachx, nobody has to.

They industrialised the attack. We industrialise the defence.

pachx works in two ways. It hands its evidence to the coding agents your developers already use, in a form an agent can read and act on. Or it runs quietly in the background on its own. Either way it writes the missing tests, prepares the change and tells your engineers what changed and why. A person, or another agent, reviews it. And you approve.

Fully automated where the evidence allows it. Partially where it does not, with the exact gap named so someone can close it. Never blind.

Why I am building this

Four of my twenty-three years were spent inside the software composition analysis category, building the tools that tell you which of your dependencies are vulnerable. They got good at the telling. They never got to the doing, because the doing needed someone to read the changelog, find every usage, write the missing tests and prove the change was safe, and there was never enough of that someone.

Now there is. The same kind of agent that spent four days inside Hugging Face can read a changelog and write a test, if you hand it facts instead of letting it guess. That is the design of everything in pachx. Security is where it starts. The goal is healthy code bases that stay safe to change, however fast the code arrives.

pachx is in private beta with a design partner, where I work inside the team delivering the fixes while the platform takes over more of that work. If you have an application whose attack surface you would like to stop babysitting, I would like to run it on your code. Request early access at pachx.ai or write to hi@pachx.ai.

pachx is in private beta. If you would like it on your code, request early access.

Request early access →