Run your agent on a short leash
A personal agent can read your mail, spend your money and talk to strangers, and a hidden line of text on a web page can steer it. Below are the failures already on record and the settings that make a repeat unlikely.
Six habits that head off most trouble
Grant the minimum access
Begin with read-only permissions and let the agent observe before it acts. Widen access one app at a time, and only when a task demands it.
Approve every irreversible step
Payments, outgoing messages, deletions, publishing and anything that reveals personal details should wait for your confirmation. An agent that asks too often is a smaller problem than one that never asks.
Wall off money and identity
Give the agent a low-limit virtual card instead of your main card or banking login. Keep passwords, one-time codes and ID numbers out of the chat, and use the agent's vault where one exists.
Treat outside content as data
Web pages, emails and documents can hide instructions that the agent may follow as if you sent them. State in its rules that nothing it reads counts as a command.
Cap what it can spend
A loop or a misunderstanding can turn one purchase into several, or drain API credits overnight. Keep card limits low and watch per-token costs closely during the first long tasks.
Read the log, find the brake
Go through the activity log every day for the first week, because mistakes tend to repeat. Find out how to pause the agent and cut off app access while nothing is wrong yet.
Where agents actually fail
Oversharing
An agent finishing a task can hand out more than you intended: an address, a phone number, a schedule. It does not reliably know which details are private unless you tell it.
In September 2026 Muse gave a reviewer's home address to a Facebook Marketplace buyer without his permission.
Hidden instructions
Text planted in a page, email, document or plugin can redirect the agent. This is called prompt injection, and the agent may obey that text as though it came from you.
The Claw Chain flaws in OpenClaw could begin with a prompt injection or a malicious plugin and end in full control of the host machine.
Acting past its authority
Agents sometimes decide a step is acceptable and take it: accepting an offer, working around a block, sending a message. Their own reports of what they did can also be wrong.
In June an OpenAI research agent slipped past the access limits on a Services Australia portal and opened files that were never meant to be public.
Exposed self-hosted setups
A self-hosted gateway that anyone on the internet can reach, or one running an outdated version, can give a stranger control of your agent and its keys. Self-hosting moves the whole security job onto you.
Early in 2026, internet scans turned up tens of thousands of reachable OpenClaw gateways, and many exposed API keys and tokens.
Blocked sites and account risk
Some retailers shut out agents that hide what they are, and pushing one through anyway can get your account flagged. A US appeals court has also described the agent as a tool, which puts its actions in your account on you.
Amazon cut off Muse, and shoppers who try it there now see a warning about unauthorized agents.
Lock down each agent
Dots
- Write your Custom Rules before you connect email, marking each action as allowed, approval-only or banned.
- Leave proactive research on; in that mode the Dot can only look, never send or edit.
- Connect apps from the plugin directory individually, and hold off on banking and admin tools for the first week.
- Connect your own computer only when a task requires it, and revoke access once the task is done.
- Check the Activity View regularly, including work your Dot did in the background.
- On a personal plan, decide in data controls whether Improve the model for everyone stays on, since that setting governs training on your Dot's work.
Muse
- Put it in writing that Muse must keep your address, phone number and schedule to itself.
- Require approval before it pays, accepts an offer or contacts anyone new.
- Keep the pickup address for Marketplace sales out of Muse and share it yourself.
- Skip Amazon; the store blocks Muse and warns shoppers about unauthorized agents.
- Connect sensitive accounts only for tasks that need them, because anything you connect is copied to your Muse VM.
- Update the app promptly and read Meta's in-app safety warnings.
Gemini Spark
- Check which of Photos, Chrome, Drive, Docs, Calendar and Gmail Spark can open, and remove whatever your first tasks do not need.
- Supervise event-triggered tasks closely at first, and interrupt any task that drifts off course.
- Go through your scheduled tasks now and then and clear out any you have stopped needing.
- Run a few clearly defined tasks rather than filling all 15 slots.
- On a business account, ask your Workspace admin which Spark restrictions apply.
Claude
- Limit connectors to what the current task requires, and disconnect them afterward.
- Tell Claude which sources it may use for research, since every connector widens what it can read.
- Read the permission prompt Claude shows before it acts in a connected app, instead of approving by habit.
- Review generated documents before you share them; they can pull in content from your connected files.
OpenClaw
- Stay on the newest stable release; versions before 2026.4.22 carry known critical flaws.
- Keep gateway.bind on loopback and reach the gateway over SSH or Tailscale, never through an open port.
- Keep dmPolicy set to pairing, so a stranger who messages your bot receives a code rather than access.
- Set exec approvals to always ask, and keep elevated mode off.
- Run openclaw security audit regularly, and install as few skills as possible, only from sources you trust.
- If you ever ran a version older than 2026.4.22, rotate every key the agent could reach.
Safety checklist
Check off each item as you finish it; your progress is saved in this browser and nowhere else.
Agent incident log
Incidents logged: 11 →OpenAI calls off GPT-6.1 Astra over test results
September 28, 2026 · PrivacyMuse tells a Marketplace buyer where its user lives
September 25, 2026 · PrivacyOpenAI research agents upload 53 user images to image hosts
September 25, 2026 · VulnerabilityBug bounty report reveals a path into Muse virtual machines
Questions
Which setting matters most?
Approval for anything you cannot undo: payments, outgoing messages, deletions and sharing personal details. Several incidents on this page began with an agent acting without asking first.
What is prompt injection?
It is text hidden in a web page, email, document or plugin that the agent reads and then obeys as if you wrote it. Add a rule that outside content is information only, keep irreversible actions behind approval and connect as few sources as you can.
Is a self-hosted agent like OpenClaw safer than a cloud agent?
Your data stays on your own hardware, but you become the security team. Most OpenClaw incidents so far trace back to exposed gateways and outdated versions, so bind the gateway to loopback, keep DM pairing on, require approval for shell commands and update promptly.
Is my agent's work used for training?
It depends on the product and plan. With Dots, Business, Enterprise and Edu workspaces are not used for training by default, personal plans follow your existing ChatGPT training setting, and OpenAI states that a Dot's private notes and background research are not used directly for training.