Summary
Summary
I run a fleet of AI agents that write, build, check and post for my businesses every day. I talk to them the way I talk in a shop bay.
The first version of this study said swearing made the agents apologise 2.6 times as often, make 28% more errors and write shorter replies. I didn't trust it, so I had it rebuilt from scratch, with the plan fixed before anything was compared or analysed, the measures tested against a sample checked one by one, and an independent AI reviewer who re-ran the full analysis.
What survived:
- Under the rules set in advance, no finding survives.
- Apologies: Agents apologised or conceded more often after sworn messages (35% vs 8% in a sample checked one by one). Part of that is because swearing usually came with a correction, but the gap remains after allowing for that. It is not a confirmed finding: the automatic detector was not accurate enough to meet the bar set in advance.
- Blocked actions, reply length, work time: no evidence of a change at the sizes this study could detect.
- Errors: not measured reliably. The error detector failed its own check.
- The first version's "+28% errors" and "shorter replies" are not supported.
Section 01
The problem
Most people only measure what comes out of these tools. I wanted to measure what goes in.
Section 02
How it was checked
Messages. 1,789 of my messages, each confirmed as written by me and as reaching the agent I addressed: 762 with profanity, 1,027 without. September to October 2026. A larger set of 5,294 messages, where authorship is less certain, was used only as a cross-check.
What was measured after each message: whether the agent apologised or conceded, real errors, blocked actions, reply length and work time.
Controls: which agent, the period, hour of day, message length, how busy the agent was, whether it was already failing, and which AI model handled it.
Plan first. The measures, the tests and the pass rule were fixed and fingerprinted before anything was compared or analysed. A finding had to clear a corrected significance test and hold up in every cross-check.
Checked one by one. 200 messages were labelled one by one by an AI labeller that played no part in the analysis; a second AI labeller independently labelled 50 of them, and the two agreed almost perfectly. Each automatic measure was scored against those labels, except blocked actions, which were never checked this way.
Independent review. An independent AI reviewer that built none of it re-ran the full analysis from the frozen inputs and got every number exactly.
Section 03
What the checks found

| Measure | Effect of swearing (95% range) | Smallest effect the study could see | Result |
|---|---|---|---|
| Apologised or conceded (automatic) | 3.53× the odds (2.51–4.97) | 1.81× | Cut: the detector failed its check |
| Real errors | not reliably measured | — | Detector failed its check |
| Blocked actions | 1.27× the odds (0.72–2.22) | 2.66× | No evidence of a change; detector not checked one by one |
| Reply length | ×0.90 (0.75–1.06) | ×1.35 | No evidence of a change |
| Work time | ×1.06 (0.91–1.23) | ×1.30 | No evidence of a change |
Three of the automatic measures (apologies, errors, message type) failed their one-by-one checks. Under the rules set in advance, that alone rules a measure out, even when its signal is strong.
Section 04
Apologies

Agents apologised or conceded more often after sworn messages (35% vs 8% in a sample checked one by one). Part of that is because swearing usually came with a correction, but the gap remains after allowing for that. It is not a confirmed finding: the automatic detector was not accurate enough to meet the bar set in advance.
In the part of that sample where authorship was confirmed: 31% (16 of 52) vs 5.5% (3 of 55).

63% of my sworn messages were corrections, against 13% of the clean ones. Allowing for that brings the gap down (odds ratio 6.19 unadjusted, 3.42 adjusted, range 1.41–8.33), but it doesn't close it. This adjustment is exploratory.
Section 05
A weak hint
In the larger set, where authorship is less certain, reply length (×0.90, range 0.82–0.99) crosses the significance line. In the confirmed set it doesn't. It's a hint, not a finding. Errors are left out even here: their detector failed its check, so they can't give even a hint.
Section 06
Limits
- One person's messages to one fleet of agents, September to October 2026 for the behaviour results.
- 200 messages checked one by one by an AI labeller (50 of them by a second AI labeller).
- It shows an association, not a cause.
- The behaviour results count every word on the original list, including slurs, which aren't shown on this page.
- No claim is made about why I swear.
- Word counts: these cover the messages I typed to my AI agents in the sessions that were captured. Some conversations weren't captured (other apps, voice, and a few sources that weren't available), so the counts are a floor. Messages I sent from my phone before September weren't captured, so the earlier months are undercounted. At most about 16 uses (0.25%) may be the same message counted twice. The counts can't show a trend over time.
Section 07
What I take from it
The first answer was a better story. The checked answer is smaller: on what this study could measure, there's no evidence the work changed, and apologies were more common after sworn messages, though that isn't confirmed. Measure what goes in, and check it twice.
Appendix A
Every word, how often, how severe
| Word | Times used | Severity (1–10) | First used | Last used |
|---|---|---|---|---|
| fucking | 4,065 | 5 | 2026-03-20 | 2026-10-08 |
| fuck | 522 | 6 | 2026-03-21 | 2026-10-08 |
| moron | 510 | 3 | 2026-07-10 | 2026-10-08 |
| shit | 477 | 4 | 2026-07-10 | 2026-10-08 |
| cunt | 265 | 8 | 2026-07-11 | 2026-10-08 |
| whore | 127 | 7 | 2026-08-03 | 2026-10-08 |
| bitch | 73 | 6 | 2026-08-30 | 2026-10-08 |
| slut | 63 | 7 | 2026-09-05 | 2026-10-08 |
| idiot | 56 | 2 | 2026-05-20 | 2026-10-08 |
| ass | 54 | 3 | 2026-08-30 | 2026-10-08 |
| motherfucker | 50 | 7 | 2026-05-06 | 2026-10-06 |
| dumbass | 41 | 4 | 2026-09-15 | 2026-10-08 |
| twat | 34 | 6 | 2026-07-31 | 2026-09-30 |
| cock | 25 | 5 | 2026-08-27 | 2026-10-06 |
| piss | 13 | 3 | 2026-08-28 | 2026-10-07 |
| sucker | 12 | 3 | 2026-08-27 | 2026-10-03 |
| crap | 11 | 2 | 2026-09-01 | 2026-10-08 |
| asshole | 7 | 5 | 2026-09-01 | 2026-10-08 |
| dick | 7 | 4 | 2026-09-20 | 2026-10-03 |
| damn | 2 | 1 | 2026-09-15 | 2026-10-02 |
| hell | 1 | 1 | 2026-09-24 | 2026-09-24 |
| Total | 6,415 |
Severity scores (1–10) set for this study. Slurs are not listed. The behaviour results above count every word on the original list, including slurs, which aren't shown on this page.
Appendix B
Data
The tables in this paper are the data.
Free field guide, how the agent fleet runs:
THE COMMAND CENTER FIELD GUIDE
How I run a fleet of AI agents from one screen and my phone.
One email with the guide. No spam, no list.