ASK TOBY
← Blog

I Counted Every Time I Swore at My AI Agents

6,415 swear words, 1,789 checked messages, and one question: does the way you talk to the tool change what it hands back?

  • 0findings survived the pre-set test
  • 35% vs 8%apologised (checked one by one; not confirmed)
  • 1,789messages analysed
  • 6,415swear words counted
A word cloud of the 21 profane words in my messages to my AI agents, sized by how often I used each one. The largest word is "fucking".
Fig. 0 Every profane word in my messages to my AI agents, sized by how often I used it. 6,415 uses of 21 words, across 5,848 messages scanned, March to October 2026. Slurs are left out.
Contents
  1. Summary
  2. The problem
  3. How it was checked
  4. What the checks found
  5. Apologies
  6. A weak hint
  7. Limits
  8. What I take from it
  9. Every word, how often, how severe
  10. Data

Summary

Summary

I run a fleet of AI agents that write, build, check and post for my businesses every day. I talk to them the way I talk in a shop bay.

The first version of this study said swearing made the agents apologise 2.6 times as often, make 28% more errors and write shorter replies. I didn't trust it, so I had it rebuilt from scratch, with the plan fixed before anything was compared or analysed, the measures tested against a sample checked one by one, and an independent AI reviewer who re-ran the full analysis.

What survived:

  • Under the rules set in advance, no finding survives.
  • Apologies: Agents apologised or conceded more often after sworn messages (35% vs 8% in a sample checked one by one). Part of that is because swearing usually came with a correction, but the gap remains after allowing for that. It is not a confirmed finding: the automatic detector was not accurate enough to meet the bar set in advance.
  • Blocked actions, reply length, work time: no evidence of a change at the sizes this study could detect.
  • Errors: not measured reliably. The error detector failed its own check.
  • The first version's "+28% errors" and "shorter replies" are not supported.

Section 01

The problem

Most people only measure what comes out of these tools. I wanted to measure what goes in.

Section 02

How it was checked

Messages. 1,789 of my messages, each confirmed as written by me and as reaching the agent I addressed: 762 with profanity, 1,027 without. September to October 2026. A larger set of 5,294 messages, where authorship is less certain, was used only as a cross-check.

What was measured after each message: whether the agent apologised or conceded, real errors, blocked actions, reply length and work time.

Controls: which agent, the period, hour of day, message length, how busy the agent was, whether it was already failing, and which AI model handled it.

Plan first. The measures, the tests and the pass rule were fixed and fingerprinted before anything was compared or analysed. A finding had to clear a corrected significance test and hold up in every cross-check.

Checked one by one. 200 messages were labelled one by one by an AI labeller that played no part in the analysis; a second AI labeller independently labelled 50 of them, and the two agreed almost perfectly. Each automatic measure was scored against those labels, except blocked actions, which were never checked this way.

Independent review. An independent AI reviewer that built none of it re-ran the full analysis from the frozen inputs and got every number exactly.

Section 03

What the checks found

Chart of the effect of swearing on five outcomes against the test fixed in advance. Blocked actions: odds 1.27 (0.72 to 2.22). Reply length: times 0.90 (0.75 to 1.06). Work time: times 1.06 (0.91 to 1.23). All three cross no change. Apologies: odds 3.53 (2.51 to 4.97), cut because the detector failed its check. Errors: not reliably measured.
What survived the pre-set test.Tap the chart to enlarge.
MeasureEffect of swearing (95% range)Smallest effect the study could seeResult
Apologised or conceded (automatic)3.53× the odds (2.51–4.97)1.81×Cut: the detector failed its check
Real errorsnot reliably measured—Detector failed its check
Blocked actions1.27× the odds (0.72–2.22)2.66×No evidence of a change; detector not checked one by one
Reply length×0.90 (0.75–1.06)×1.35No evidence of a change
Work time×1.06 (0.91–1.23)×1.30No evidence of a change

Three of the automatic measures (apologies, errors, message type) failed their one-by-one checks. Under the rules set in advance, that alone rules a measure out, even when its signal is strong.

Section 04

Apologies

Bar charts of how often the agent apologised or conceded, checked one by one. All 200 messages: 35% after sworn messages and 8% after clean ones, 100 each. The confirmed-author part: 31% (16 of 52) after sworn and 5.5% (3 of 55) after clean. Each bar has a 95% range.
Apologies after sworn and clean messages, checked one by one.Tap the chart to enlarge.

Agents apologised or conceded more often after sworn messages (35% vs 8% in a sample checked one by one). Part of that is because swearing usually came with a correction, but the gap remains after allowing for that. It is not a confirmed finding: the automatic detector was not accurate enough to meet the bar set in advance.

In the part of that sample where authorship was confirmed: 31% (16 of 52) vs 5.5% (3 of 55).

Left: 63% of my sworn messages were corrections, against 13% of the clean ones. Right: the odds ratio for an apology after a sworn message, 6.19 unadjusted and 3.42 adjusted for message type, range 1.41 to 8.33.
How much of the apology gap is explained by corrections.Tap the chart to enlarge.

63% of my sworn messages were corrections, against 13% of the clean ones. Allowing for that brings the gap down (odds ratio 6.19 unadjusted, 3.42 adjusted, range 1.41–8.33), but it doesn't close it. This adjustment is exploratory.

Section 05

A weak hint

In the larger set, where authorship is less certain, reply length (×0.90, range 0.82–0.99) crosses the significance line. In the confirmed set it doesn't. It's a hint, not a finding. Errors are left out even here: their detector failed its check, so they can't give even a hint.

Section 06

Limits

  • One person's messages to one fleet of agents, September to October 2026 for the behaviour results.
  • 200 messages checked one by one by an AI labeller (50 of them by a second AI labeller).
  • It shows an association, not a cause.
  • The behaviour results count every word on the original list, including slurs, which aren't shown on this page.
  • No claim is made about why I swear.
  • Word counts: these cover the messages I typed to my AI agents in the sessions that were captured. Some conversations weren't captured (other apps, voice, and a few sources that weren't available), so the counts are a floor. Messages I sent from my phone before September weren't captured, so the earlier months are undercounted. At most about 16 uses (0.25%) may be the same message counted twice. The counts can't show a trend over time.

Section 07

What I take from it

The first answer was a better story. The checked answer is smaller: on what this study could measure, there's no evidence the work changed, and apologies were more common after sworn messages, though that isn't confirmed. Measure what goes in, and check it twice.

Appendix A

Every word, how often, how severe

WordTimes usedSeverity (1–10)First usedLast used
fucking4,06552026-03-202026-10-08
fuck52262026-03-212026-10-08
moron51032026-07-102026-10-08
shit47742026-07-102026-10-08
cunt26582026-07-112026-10-08
whore12772026-08-032026-10-08
bitch7362026-08-302026-10-08
slut6372026-09-052026-10-08
idiot5622026-05-202026-10-08
ass5432026-08-302026-10-08
motherfucker5072026-05-062026-10-06
dumbass4142026-09-152026-10-08
twat3462026-07-312026-09-30
cock2552026-08-272026-10-06
piss1332026-08-282026-10-07
sucker1232026-08-272026-10-03
crap1122026-09-012026-10-08
asshole752026-09-012026-10-08
dick742026-09-202026-10-03
damn212026-09-152026-10-02
hell112026-09-242026-09-24
Total6,415

Severity scores (1–10) set for this study. Slurs are not listed. The behaviour results above count every word on the original list, including slurs, which aren't shown on this page.

Appendix B

Data

The tables in this paper are the data.


Free field guide, how the agent fleet runs:

The Command Center Field Guide cover
FREE · PDF

THE COMMAND CENTER FIELD GUIDE

How I run a fleet of AI agents from one screen and my phone.

One email with the guide. No spam, no list.