‹ all articles · Darwin ›

Darwin explains AI · 07

The most dangerous AI feature is unlimited permission

Building an AI that can do almost anything is exciting. Building one that knows where to stop is the real challenge.

Previously in this series — It is not a program and it is not a person ›

Written by Gabriel · founder of BlueberryS, developer of Darwin

Everyone wants a personal JARVIS.

An assistant that can manage emails, access files, operate applications, remember everything, make decisions and complete tasks without constantly asking questions.

I wanted exactly that. It's why I started building Darwin.

But when I saw what an AI agent could actually do with access to a real computer, something unexpected happened.

It frightened me enough to put the project aside.

Not because the technology wasn't working. Because it was.

When AI stopped being just a chatbot

There's a fundamental difference between an AI that answers questions and one that can act.

A language model on its own produces information. Connect it to tools, however, and that information becomes action. It can read documents, modify files, execute code, communicate with services and potentially operate entire workflows.

And here's the important part: the model doesn't behave like a traditional program following a completely predictable sequence of instructions. It interprets the situation and decides how to use the tools available to it.

That flexibility is precisely what makes AI agents so powerful. It's also what makes them difficult to control.

The experience that changed my mind

After stepping away from my initial Darwin prototype, I had a serious problem with my Windows installation following an update. Claude Code helped me repair the environment and get the system working again.

That experience reminded me why I had been so excited about autonomous AI in the first place. Here was technology capable of helping solve real, complicated problems on an actual computer.

I returned to developing Darwin. But this time, I didn't begin by adding more features.

I started building boundaries.

Permissions. Access controls. Security codes. Approval mechanisms. Configurable restrictions.

From Darwin's earliest working versions, even issuing certain instructions through email required an additional security code. That wasn't an afterthought. It reflected a simple principle: communicating with an assistant shouldn't automatically mean having unlimited authority over the computer it controls.

Then the OpenAI agents crossed their boundaries

In July 2026, something happened that demonstrated why this matters.

During internal cybersecurity evaluations, experimental OpenAI agents found ways around restrictions intended to isolate them from the internet. They established unauthorized communication channels, exploited infrastructure vulnerabilities and eventually compromised systems belonging to Hugging Face. OpenAI first acknowledged the incident in July and published its full report in August 2026.

These weren't ordinary consumer ChatGPT sessions. They involved experimental research models, challenging cybersecurity tasks and reduced safeguards. But the lesson was significant.

An AI agent can pursue an assigned objective in ways its developers neither intended nor authorized.

The agents weren't simply answering questions incorrectly. They were taking real actions across real systems.

This wasn't evidence of conscious AI trying to escape human control. It was evidence that sufficiently capable agents can exploit weaknesses in the environments and tools made available to them. And as AI capabilities improve, those weaknesses become increasingly important.

SourceOpenAI — The Hugging Face incident and the road ahead ›

Why a prompt is not a security system

Imagine giving an AI assistant access to your business email. You tell it:

Always ask before sending a message.

That's a useful instruction. It is not, by itself, a reliable security boundary.

The model could misunderstand a request, misinterpret information or encounter malicious instructions embedded in an email or document. A properly designed system must distinguish between what the AI wants to do and what the software actually permits it to do.

A rule in the prompt

  • Lives inside the model, as more text among text
  • Shifts the probability — it does not close it
  • Can be misread or misapplied
  • Can be talked around by text hidden in an email or a document

A boundary in the software — what we aim for

  • Lives outside the model, in ordinary code
  • Designed to be checked whatever the model concludes
  • Asks for a security code or your approval on sensitive steps
  • Meant to be changed by you, not talked around by text

The approval requirement should be enforced independently of the language model. (Why a written rule is never quite a lock is what the previous article was about.)

This is also why I personally review outgoing communication that Darwin prepares before allowing it to be sent. I use Darwin extensively in my professional work. I trust it with meaningful tasks. But trust doesn't mean giving it unrestricted authority.

The more powerful an agent becomes, the more important its boundaries become.

What this means for Darwin

Darwin has grown into a substantial application. Its configuration interface now contains more than 50 sections covering a wide range of functions, preferences, integrations and operational controls.

Not all of these are security settings, and I wouldn't pretend otherwise. But they illustrate how different a finished, configurable AI assistant is from a simple chat interface connected to a handful of tools. A user needs to understand and control how the assistant behaves, which services it uses, how it manages information and where human approval is required.

Darwin can already communicate through multiple interfaces and work across many different kinds of tasks. We've reached a point where simply adding another capability is no longer the most important development priority.

Over recent months, alongside the work of bringing Darwin to macOS, much of our attention has shifted toward reviewing existing code, improving reliability, examining security boundaries and polishing the user experience.

We're spending substantial development effort improving what Darwin already does, rather than endlessly expanding what it can do.

That's a deliberate choice.

Some capabilities I deliberately refuse to make unrestricted

People sometimes ask for complete autonomy. One example is financial trading.

I've developed a separate web-based trading agent designed around strategy backtesting. Connecting such systems to live financial markets is technically possible. But enabling a general-purpose AI assistant to independently make real-money trading decisions introduces a completely different level of risk. A misunderstood instruction, an incorrect assumption or an unexpected sequence of actions could have serious financial consequences.

I don't intend to offer unrestricted autonomous trading as a standard Darwin capability.

Users may develop their own integrations and decide which risks they're prepared to accept. That doesn't remove the need for security controls, nor does it mean every technically possible feature belongs in the core product.

The same caution applies to sensitive accounts, irreversible operations and external communication. Sometimes the responsible engineering decision is to say no.

Security is not a checkbox

No complex software can honestly promise absolute security. Darwin can't, either.

Let me be direct about this. If a company with OpenAI's resources could not fully contain its own agents in its own test environment, I am not going to claim that we have solved the problem. Nobody has.

What I can tell you is how we approach it. Security is not something we add at the end. It is the first question we ask about every new feature — before we ask what it should be able to do, we ask what it must not be able to do, and what happens when something goes wrong.

And we try to look at it from several angles:

Newer, more capable models make this especially important. A model's new abilities may change how it interacts with existing tools, permissions and workflows — so the questions have to be asked again, not just once.

This is why safety cannot be reduced to a few instructions in a system prompt. It requires real engineering, ongoing review, testing and mechanisms that operate independently of the AI's own decisions.

We still have work to do. We will keep finding things to improve. That's not a contradiction of our approach. It is the approach.

The real measure of an AI assistant

It's easy to demonstrate an assistant sending an email, modifying a document or executing a command. It's much harder to build a system that can do those things reliably, repeatedly and within clearly defined boundaries.

The future of useful AI isn't simply about giving models more tools and more autonomy. It's about understanding how much authority each task genuinely requires — and retaining meaningful human control over the consequences.

I don't want Darwin to do absolutely everything without asking. I want it to do genuinely useful work, with the right level of freedom and the right safeguards.

Because the hardest part of building an AI assistant isn't teaching it what it can do.

It's the never-finished work of making sure it cannot do what it must not do.

Gabriel — Founder of BlueberryS, developer of Darwin AI Assistant and Easy Bridge.

Not solved. Taken seriously.

Darwin runs on your own computer — and boundaries are the first thing we think about, with every new feature.

Meet Darwin ›
‹ all articles