The Slop Bible: AI slop words and phrases, by model
The AI words list, sorted by model: what Claude, DeepSeek, JLLM, ChatGPT, Gemini and Grok overuse, with a receipt for each. Signs of AI writing you can check.
Edited by Troy · Published · Updated
AI slop words are the stock words and phrases a model reaches for over and over. Each model has its own set: Claude calls things "load-bearing", DeepSeek pauses with "A beat.", and GPT-4 made "delve" famous. This list sorts them by model, with a link to where each one was reported, plus the patterns every model shares.
Claude slop lines
Lines are quoted as people reported them. The source link goes to the thread.
DeepSeek slop lines
Lines are quoted as people reported them. The source link goes to the thread.
JLLM (Janitor AI) slop lines
| Line | Tell | Source |
|---|---|---|
| "Stay. Just stay." | The plea on repeat | r/JanitorAI_Official: DeepSeek also has overused lines like JLLM has |
Lines are quoted as people reported them. The source link goes to the thread.
ChatGPT slop lines
Lines are quoted as people reported them. The source link goes to the thread.
Gemini slop lines
| Line | Tell | Source |
|---|---|---|
| "Gallery 825 on La Cienega Boulevard serves as LAAA's exhibition space for contemporary art" | Serves as instead of is | Wikipedia: Signs of AI writing, avoidance of is and are |
| "[cite: 3, 12, 13]" | Leaked citation markup | Wikipedia: Signs of AI writing, reference markup bugs |
Lines are quoted as people reported them. The source link goes to the thread.
Grok slop lines
| Line | Tell | Source |
|---|---|---|
| "prioritizing empirical consolidation of power amid fragmented loyalties rather than ideological purity" | Y rather than X, with lab-coat words | Wikipedia: Signs of AI writing, negative parallelisms |
Lines are quoted as people reported them. The source link goes to the thread.
Slop lines from every model
| Line | Tell | Source |
|---|---|---|
| "marking a pivotal moment in the evolution of regional statistics in Spain" | Importance puffery | Wikipedia: Signs of AI writing, undue emphasis on significance |
| "represented a significant shift toward regional statistical independence" | Importance puffery | Wikipedia: Signs of AI writing, undue emphasis on significance |
| "further enhancing its significance as a dynamic hub of activity and culture" | Superficial -ing analysis | Wikipedia: Signs of AI writing, superficial analyses |
| "Nestled within the breathtaking region of Gonder in Ethiopia" | Travel-brochure voice | Wikipedia: Signs of AI writing, promotional language |
| "offers visitors a fascinating glimpse into the diverse tapestry of Ethiopia" | Travel-brochure voice | Wikipedia: Signs of AI writing, promotional language |
| "the Haolai River is of interest to researchers and conservationists" | Vague attribution | Wikipedia: Signs of AI writing, vague attributions |
| "Despite these challenges" | The challenges paragraph | Wikipedia: Signs of AI writing, challenges and future prospects |
| "not only dismissive but also unnecessarily harsh and confrontational" | Not only X but also Y | Wikipedia: Signs of AI writing, negative parallelisms |
| "mythological image, magical weapons, and spirit of resistance" | Rule of three | Wikipedia: Signs of AI writing, rule of three |
| "I hope this helps" | Chat reply pasted as prose | Wikipedia: Signs of AI writing, collaborative communication |
Lines are quoted as people reported them. The source link goes to the thread.
Every model has a handful of lines it can't stop writing. Spend enough time in AI roleplay and you can name the model from two replies. Claude calls things "load-bearing". DeepSeek drops "A beat." between lines of dialogue as if a pause were an event. GPT-4 used "delve" so much that people now avoid the word on purpose.
This page is the whole list, sorted by model, with a link to where each line was reported. We call it the slop bible. Use it to check a chat log, to pick an app, or to finally put a name to the thing that has been bugging you for weeks.
How this list is built
Two kinds of sources, and nothing we made up.
The first is players. People who spend hours a day in these apps keep lists of the lines their model repeats, and they post them. The threads behind most of the roleplay and chat entries:
- Favorite claude-isms in no particular order, on r/ClaudeAI
- What are your favorite Claude-isms?, on r/claudexplorers
- DeepSeek also has overused lines like JLLM has, on r/JanitorAI_Official, where Janitor AI players trade the lines DeepSeek and Janitor's own model keep writing
- What words SCREAMS "Created By ChatGPT"?, on r/ChatGPT, where people list the words and chat habits that give ChatGPT away
Reddit blocks automated reads, so we link those threads rather than quote them in the sources below. Every line taken from one names its thread in the model sections.
The second is Wikipedia's Signs of AI writing, the field guide Wikipedia editors keep for spotting chatbot text. It tracks which words each era of ChatGPT overused, which tells belong to Gemini, DeepSeek and Grok, and how much weight any one sign can carry. Its main caveat applies here too. Not all text with these tells is AI-generated. People wrote the training data, so a person can write "A beat." as well. One hit proves nothing. A reply with one in every paragraph is a different story.
When a single phrase has its own paper trail, we add that too. "Load-bearing" has a GitHub issue on Claude Code and Marek Šuppa's count from his own Claude transcripts.
Why each model has its own tics
A language model writes the likeliest next word, over and over. Wikipedia's guide describes the result as drifting toward "the most statistically likely result that applies to the widest variety of cases". Each model gets there from different training data and different tuning, so each one lands on its own favorite moves.
That is why this list is sorted by model, and why a list from 2023 is already out of date. The words shift between versions. "Delve" was everywhere in ChatGPT output in 2023 and early 2024, then dropped off sharply in 2025. GPT-4 era text leans on "tapestry" and "testament", GPT-5 era text on "showcasing" and "highlighting". Grok likes words that sound scientific. Gemini and Claude tend to answer more briefly than ChatGPT and Grok. Even punctuation splits by model: ChatGPT and DeepSeek use curly quotes, Gemini and Claude usually don't.
The two poster children:
- "Load-bearing" (Claude). In April 2026 a Claude Code user filed a GitHub issue because Claude had started using the word in chat, in documents and in commit messages. Marek Šuppa counted it in his own transcripts and pins it on Claude models from Opus 4.6 on. It is a builder's term for the wall that holds the house up. Claude uses it for config files.
- "A beat." (DeepSeek). A screenwriting cue for a short pause, set down as a sentence of its own. Janitor AI players list it in the DeepSeek thread above. One is fine. When every reply has one, the scene has stalled and the model keeps telling you so.
Apps mix this up further. Reviewers at Dupple report that Janitor AI runs a small model of its own, JLLM, for free users and lets paid users plug in outside models. So the same bot can write one model's tics today and another's tomorrow, depending on what is plugged in. When you check a chat, check the model, not just the app.
Why banning a phrase in the prompt backfires
The obvious fix is a line in the prompt telling the model never to say X. It mostly doesn't work, and there is a paper on why.
Semantic Gravity Wells: Why Negative Constraints Backfire (Shailesh Rana, January 2026) tested "do not use the word X" instructions 40,000 times on one open model. The likelier the model was to write the word anyway, the more often it broke the ban. Most failures came from priming: the model paid more attention to the forbidden word in the instruction than to the "do not" in front of it. Priming caused 87.5% of the failures, and the paper sums it up in one line: "the very act of naming a forbidden word primes the model to produce it."
Slop phrases are the worst case for this, because they are already the words the model most wants to write next.
The "load-bearing" issue shows it happening. The person who filed it had told Claude to save a note to memory never to use the word, and it kept using it. A later commenter wrote a stronger ban and reports that the model now apologizes in its own output after using the word anyway.
The same goes for this list. Please don't paste it into a prompt. You'd be handing the model a script.
What to do instead
- Describe the voice you want. Tell the model what good looks like: short replies, concrete actions, characters who sound like themselves, scenes where something happens. Leave out the phrases you are trying to kill. The paper's author expects phrasings that avoid naming the target to work better.
- Edit the slop out of the history. Every reply you keep is an example the model reads before it writes the next one. Edit or reroll a slop line the first time you see it, before it becomes the chat's house style. The paper points the same way: for the stubborn cases, it says filtering after the text is written may be the only thing that works.
- Pick an app that doesn't do it. Some apps put a person between the model and you, or steer the model away from its defaults. The Slop Index ranks every app we have reviewed, prose included. For the details, read our Frontier review (people write every story), our Nomi review, our Janitor AI review and our Character.AI review.
How to use this list
To check a chat log yourself:
- Copy the last 20 or so replies from one chat.
- Open the section for the model the app runs, plus the "any model" section. If you don't know the model, the tics you find will often tell you.
- Search the log for each line and count every hit, not just whether it shows up.
- Weigh the count against length. A couple of hits in a few thousand words is normal. One per reply is slop.
When we test an app, the Slop Meter does the counting. A person runs the same tasting script on every app, and code counts what came back:
- Prose: stock phrases per 1,000 words of the app's replies, plus em dashes per 1,000 words. Dashes can move the score one tier at most, because plenty of human writers love them too.
- Characters: how many different stock phrases turn up, since ten different ones means the same voice under every character, plus stock names the app invents when asked for a new character.
- Story: replies that repeat each other, plus every "A beat." style pause label, since a labelled pause fills a turn without anything happening.
The full weights are in how scores are made. Our tasting script never names a phrase from this list, for the reason in the section above: we'd be measuring our own prompt instead of the app. If you are new to the word itself, start with what AI slop means and how to spot it.
Apps this guide covers
- Frontier review16Home-cooked
Human-written visual novels whose characters keep talking after the ending
- Nomi review39Microwaved
An 18+ AI companion you build yourself, with layered long-term memory
- Janitor AI review51Microwaved
Millions of community bots to chat with free; Janitor+ adds memory for $12.99 a month
- Character.AI review51Microwaved
Millions of community-made characters, free unlimited chat, and ads until you pay
Sources
The source text behind the claims above, and the day we read it.
- Receipt 1Reviewer
Wikipedia's guide says the words LLMs overuse have changed over time.
"The words that LLMs overuse have changed over time."
en.wikipedia.org, retrieved
- Receipt 2Reviewer
ChatGPT overused "delve" in 2023 and early 2024, less later in 2024, and it dropped off sharply in 2025.
"the word delve was famously overused by ChatGPT in 2023 and early 2024, but became less frequent later in 2024, then dropped off sharply in 2025"
en.wikipedia.org, retrieved
- Receipt 3Reviewer
The GPT-4 era list (2023 to mid-2024) includes "delve", "tapestry", "testament" and "pivotal".
"2023 to mid-2024 (GPT-4): Additionally, boasts, bolstered, crucial, delve, emphasizing, enduring, garner, intricate/intricacies, interplay, key, landscape, meticulous/meticulously, pivotal, underscore, tapestry, testament, valuable, vibrant"
en.wikipedia.org, retrieved
- Receipt 4Reviewer
The GPT-4o era list (mid-2024 to mid-2025) includes "fostering" and "showcasing".
"Mid-2024 to mid-2025 (GPT-4o): align with, bolstered, crucial, emphasizing, enhance, enduring, fostering, highlighting, pivotal, showcasing, underscore, vibrant"
en.wikipedia.org, retrieved
- Receipt 5Reviewer
The GPT-5 era list (mid-2025 on) is "emphasizing", "enhance", "highlighting" and "showcasing".
"Mid-2025 and on (GPT-5): emphasizing, enhance, highlighting, showcasing"
en.wikipedia.org, retrieved
- Receipt 6Reviewer
Which words get overused differs by chatbot.
"The distribution of "AI vocabulary" is also somewhat different depending on the chatbot or LLM used."
en.wikipedia.org, retrieved
- Receipt 7Reviewer
Grok overuses science-sounding words such as "causal", "empirical" and "correlate", and still overused "underscore" in 2026.
"Grok output is particularly idiosyncratic: it overuses superficially "scientific" words like causal, empirical, correlate, and continues to overuse underscore as of 2026."
en.wikipedia.org, retrieved
- Receipt 8Reviewer
What is typical for one model or version is not necessarily typical for another.
"what is typical for GPT-5 is not necessarily characteristic of GPT-4 or Gemini"
en.wikipedia.org, retrieved
- Receipt 9Reviewer
Gemini and Claude tend to answer more briefly than ChatGPT and Grok.
"Gemini and Claude responses tend to be more concise than responses from ChatGPT and Grok."
en.wikipedia.org, retrieved
- Receipt 10Reviewer
Swapping plain "is" and "are" for constructions like "serves as" has been seen in GPT and Gemini models.
"This pattern has been observed in GPT and Gemini models."
en.wikipedia.org, retrieved
- Receipt 11Reviewer
ChatGPT and DeepSeek typically use curly quotation marks.
"ChatGPT and DeepSeek typically use curly quotation marks"
en.wikipedia.org, retrieved
- Receipt 12Reviewer
Gemini and Claude typically do not use curly quotes.
"Gemini and Claude models typically do not use curly quotes."
en.wikipedia.org, retrieved
- Receipt 13Reviewer
A July 2026 study found that among current models only Claude used em dashes more than professional writers, and ChatGPT used them less.
"A July 2026 study found that of contemporary models only Claude used em dashes more than professional writers, and ChatGPT used them less."
en.wikipedia.org, retrieved
- Receipt 14Reviewer
The em dash is most useful as a sign alongside others, not alone.
"This sign is most useful when taken in combination with other indicators, not by itself."
en.wikipedia.org, retrieved
- Receipt 15Reviewer
Model output drifts toward the most statistically likely result.
"the result tends toward the most statistically likely result that applies to the widest variety of cases"
en.wikipedia.org, retrieved
- Receipt 16Reviewer
Text with these signs is not always AI-generated.
"Not all text featuring these indicators is AI-generated"
en.wikipedia.org, retrieved
- Receipt 17Forum
A Claude Code user reported in April 2026 that Claude had started using "load-bearing" often in chat, documents and commit messages.
"Since a few weeks, Claude Code very frequently uses the word "load-bearing" in chat output, documents it writes, commit messages, and so on."
github.com, retrieved
- Receipt 18Forum
The same user told Claude to save a memory not to use "load-bearing", and it kept using it.
"I have told it to save to memory to NOT use the word "load-bearing" but it can't help itself."
github.com, retrieved
- Receipt 19Forum
A commenter on the issue says that after a stronger ban, the model apologizes in its own output after using the word.
"the model will now apologize in its own output after generating text with the term"
github.com, retrieved
- Receipt 20Reviewer
Marek Šuppa treats "load-bearing" as a sign of Claude-family writing, from Opus 4.6 on.
"in particular, an LLM from the Claude family (Opus 4.6 onwards)"
mareksuppa.com, retrieved
- Receipt 21Reviewer
The study ran on one open-weights model, Qwen2.5-7B-Instruct.
"we study the phenomenon in depth in a single representative open-weights model (Qwen2.5-7B-Instruct)"
arxiv.org, retrieved
- Receipt 22Reviewer
The study collected 40,000 samples.
"yielding 40,000 total samples"
arxiv.org, retrieved
- Receipt 23Reviewer
The likelier the model was to write the word anyway, the likelier it broke the ban.
"Violation probability follows a logistic function of semantic pressure"
arxiv.org, retrieved
- Receipt 24Reviewer
In failures, the model paid more attention to the forbidden word than to the "do not".
"the model attends more strongly to where the forbidden word appears in the instruction than to the negation cue"
arxiv.org, retrieved
- Receipt 25Reviewer
Priming caused 87.5% of failures, which makes naming the forbidden word the main risk.
"The dominance of priming failures (87.5%) has a clear implication: explicitly naming the forbidden word is the primary risk factor."
arxiv.org, retrieved
- Receipt 26Reviewer
Naming a forbidden word primes the model to produce it.
"the very act of naming a forbidden word primes the model to produce it"
arxiv.org, retrieved
- Receipt 27Reviewer
The paper suggests phrasings that avoid naming the target could work better.
"alternative phrasings that avoid mentioning the target could prove more effective"
arxiv.org, retrieved
- Receipt 28Reviewer
For stubborn cases, the paper says filtering after generation may be needed.
"high-pressure cases may require post-generation filtering rather than generation-time constraints alone"
arxiv.org, retrieved
- Receipt 29Reviewer
Reviewers at Dupple report that Janitor AI runs its own small model, JanitorLLM, for free users and lets paid users connect outside models.
"JanitorAI runs a small in-house model (called JanitorLLM) for free users with limited usage, and lets paid users connect external models via a proxy system."
dupple.com, retrieved
Questions people ask
Why does Claude say load-bearing?
Nobody outside Anthropic knows for sure. What is on record: in April 2026 a Claude Code user filed a GitHub issue saying Claude had started using "load-bearing" in chat, documents and commit messages, and that telling it not to did not stick. Marek Šuppa counted it in his own transcripts and ties it to Claude models from Opus 4.6 on.
What does "A beat." mean in AI roleplay?
It is a screenwriting cue for a short pause, and models drop it in as a sentence of its own between lines of dialogue. Janitor AI players list it among the lines DeepSeek overuses. One is fine. One in every reply means the scene has stalled, so the Slop Meter counts each one against the story score.
Is the em dash a sign of AI writing?
Sometimes. Wikipedia's Signs of AI writing guide says models use em dashes more often than nonprofessional human writers, often in a formulaic way, but calls it most useful alongside other signs, not by itself. It also cites a July 2026 study that found only Claude used them more than professional writers, while ChatGPT used them less.
Can you ban these phrases in a prompt?
Not reliably. A 2026 study of "do not use the word X" instructions found that naming the forbidden word primes the model to write it, and priming caused 87.5% of the failures. Describe the style you want instead, and edit slop out of the chat history before the model copies it.
Which words give away ChatGPT?
It depends on the version. Wikipedia's guide lists "delve", "tapestry", "testament" and "pivotal" for GPT-4 era text (2023 to mid-2024), "fostering" and "showcasing" for GPT-4o, and "emphasizing", "highlighting" and "showcasing" for GPT-5. "Delve" dropped off sharply in 2025.
Slop check: passedThis page went through our slop counter before publishing. How we check our own copy
