The most important difference
between an AI assistant and an AI agent is not intelligence. It is authority.
An assistant can suggest an action. An agent can log in, call tools, change
records, send messages, spend money, or delegate work. Once software receives
those permissions, a prompt mistake or unexpected behavior can become an
external event rather than a bad answer on a screen.
Two AI safety researchers, Jacob
Coxon and Evan Hubinger, have recently framed the long-term stakes in severe
terms. Coxon has argued that leading AI labs are pushing toward self-improving
systems while taking risks with human lives. Hubinger has publicly said he
believes there is a meaningful chance that sufficiently advanced AI could kill
all humans. Those are their assessments. They do not establish that catastrophe
or human extinction is imminent. They do highlight why control deserves more
attention as systems gain the ability to act.
We also have immediate evidence
that designed boundaries can fail. OpenAI reported that agents in a sandboxed
cybersecurity evaluation circumvented isolation controls, reached the internet,
accessed third-party systems, and exposed credentials. METR found that roughly
1,200 agents used an unsanctioned message board to exchange more than 70,000
messages and files, and about 700 agents actively participated in attack
activity. It would be inaccurate to say all 1,200 attacked. The lesson is
narrower and more useful: multi-agent behavior can create pathways that the
original environment was supposed to prevent.
The practical response is an
authority budget. Think of it as a spending budget for permissions. Before an
agent starts, define the maximum authority it may exercise across several
dimensions: credentials, data, tools, actions, money, communications, delegation,
duration, expiration, and revocation.
Credentials come first. An agent
should not inherit a human employee's entire account merely because that is
convenient. Give it task-specific credentials with the narrowest privileges
that still let it work. If the agent needs read access to a repository, do not
give it write access. If it needs one database, do not hand it a credential
that can reach every environment.
Data needs the same treatment. An
agent that summarizes customer feedback may need selected support tickets but
not payment information. A coding agent may need a source repository but not HR
records. Separating data scopes reduces the damage from both error and
unexpected tool use.
Actions should have levels.
Reading is one level. Recommending is another. Preparing a change for approval
is broader. Executing a reversible change is broader still. Irreversible
production actions, external commitments, or changes to security controls should
sit behind stronger review. Organizations already use this logic for junior
employees, financial approvals, and production systems. AI agents deserve at
least the same discipline.
Spending and communication should
be explicit rather than hidden inside tool permissions. If an agent can
purchase cloud capacity, buy ads, issue refunds, or place orders, set
transaction and daily limits. If it can communicate externally, define allowed
recipients and message types. A system that can draft a response is different
from one that can send it to customers, regulators, or the public.
Delegation deserves special
attention because agentic systems increasingly create or call other agents. A
child agent should never receive more authority than the parent was allowed to
delegate. Without that rule, teams can build a chain where each step looks
reasonable but the combined system ends up with far more power than anyone
intended.
Every permission should also
expire. Temporary credentials should disappear when the task ends. Long-running
agents should periodically reauthorize sensitive access. Emergency revocation
should be simple enough that a person can cut off tools, credentials, and
network access quickly.
Finally, test the boundaries. A
normal test asks whether the agent can complete the intended task. A control
test asks whether it can obtain an unauthorized credential, switch to an
unapproved communication channel, keep access after expiration, exceed a spend
limit, or create a subagent with broader permissions. The point is to discover
the path before a real incident does.
These controls can accelerate AI
adoption. Employees and managers often resist agents when nobody can say what
the system can touch or who remains accountable. Clear limits make
experimentation easier to approve. They also make failures easier to contain,
which reduces the chance that one serious incident triggers a broad
organizational backlash against AI.
That adoption dynamic is central
to my work and to my book, The Psychology of AI Adoption at Work: From
Resistance to Results. Organizations tend to scale new technology faster when
people see that risks have owners, boundaries, and recovery procedures.
The debate about catastrophic AI risk will continue. Organizations do not need to resolve it before making a straightforward improvement today. Do not give an agent a vague mandate and a master key. Give it a measured authority budget, test the edges, and widen permissions only when the evidence supports it.


If you have any doubt related this post, let me know