The Bible Was AI Generated, Then It Wasn’t

Here's my best attempt to chronicle an everlasting battle against clients wanting to use dumb AI detectors

Today, I’m going to chronicle my always-on battle vs. clients using AI detectors that randomly flag my handwritten sentences as AI-generated, and then I have to rewrite them in convoluted ways just to “bypass” them because, otherwise, I won’t get paid. This is a struggle against AI detectors making inaccurate assumptions about a human’s original writing. And we’ll break it all down, including everything I learned in the process of reducing the “AI-ness” of an article.

But all of this began back in 2020. Shortly before we found out that the Bible was AI-generated.

The AI Before AI

OpenAI has been working on its GPT model for a while. Though ChatGPT was publicly released in 2022, the work on the actual models was finished way before that. In June 2018, OpenAI was done with GPT-1. GPT-2 was released in February 2019. GPT-3 was released in May 2020. Around this time, interested parties were already looking at how GPT can automate the process of writing.

Back in 2020, we got an order. A good couple of years before the public release of ChatGPT. The order was strange. The client wanted us to use a new tool to do our writing. We had a relationship with the client for bulk, monthly orders amounting to several tens of thousands of words. He wanted to create content at a much faster pace and possibly additional web properties as well, so he told us about this tool.

Copy.ai. It used GPT.

As you might guess, this was a tool to generate content back in 2020. It would create an outline, then write paragraphs under selected headings. It was worse than the worst ChatGPT content you’ll see, but stylistically speaking only. Theoretically, it did cover the topic. At that time, we did not have the fatigue and disdain for AI-generated text that we do today. At that time, it felt like you had a professor writing text for you. Robotic, dull, monotonous, but it was covering all aspects.

It’s a little hard to get this point across. Today, one would imagine if it was such plain, boring AI text, why did it even work? Well, at that time, there was no AI-generated text awareness among clients and website owners. Whatever it wrote (think early ChatGPT-like text) technically covered the topic wholly. And with no fatigue, it felt like it was truly a shortcut.

We said no to using that tool after a trial. Yes, it could create text faster. And yes, we didn’t have the same dislike toward AI-generated text that we do today for valid reasons. But it made some glaring mistakes here and there, hallucinated a lot (at that time we didn’t know this is what it would be called; we just liked to say it was “inventive”), and sometimes would write rubbish instead of proper sentences.

That caused a lot of wasted effort and extra work for our editors (editors are a more premium commodity vs. article writers). Add to that the fact that the niche was fairly simple (travel, hotels, flights, etc.), which we were already good at. It made little sense generating text via this tool and doing heavy editing. It just felt speedier and more natural to write it yourself. We told the client the same. He was okay with continuing on the previous arrangement. He likely offloaded this same work to some other service/agency because we saw a steady decline in orders from him.

All the way back in 2020, the ratio of AI slop to actual, well-researched human content had started to skew for the worse. And that brings us to our surprising discovery that the Bible was actually AI-generated!

When the Bible was AI-Generated

Fast forward a few years and everyone was generating exactly the same style of content it seemed. Turns out, ChatGPT was spitting out articles in full force. The year is 2023. It became so rampant that clients started becoming wary of ChatGPT text. There were also rumors that Google was penalizing AI content.

Side note on that: Google never said that it won’t rank AI content. The position has always been that AI or not, it must be valuable ultimately and follow EEAT principles. Somehow writing an excellent, to-the-point, and real-world article with AI will still rank you. A human writing objectively bad content laden with fillers and stuffed keywords will still not get ranked. But most people using AI to create not only articles but entire websites were, to say the least, quite lazy. The end result was content full of everything Google hates. And many websites got unranked due to that. Many more never even took off despite having thousands of articles in a month.

So, as clients started to become aware of this, they wanted articles that were human-written, not AI-generated. This created a niche category of tools: AI detectors. The detectors will tell you if it’s AI or not. Initially, they just looked at some patterns and words that AI liked to overuse. Pretty accurate. But over time, as AI models became more and more clever, and people started getting better at prompting them to follow specific styles, the AI detectors had to adapt.

And they adapted terribly.

Once during a long to-and-fro with a client trying to prove our content is genuinely human and not AI-generated, we came across something neat. Here’s how it went down.

The client alleged that we used AI because one of the detectors showed it as likely AI. We rewrote the piece. Still likely AI. So I went on the AI detectors and scanned myself. Yep, it was AI. According to these tools, at least. I tried to change certain things, and the score fluctuated wildly. On a whim, I took different samples from before AI:

Lo and behold, large sections of these were also AI-generated according to these AI detectors. This laid to rest our debate with that client.

The Motorola Droid review naturally used words like “boasts a gorgeous display,” “fully embraces the openness of the Android platform,” and had comma-separated values in sets of three. All pretty close to the signs that early AI detectors relied on. All pretty “ChatGPT-like” from 2022-2023.

But of course, the Bible wasn’t AI-generated. Tolkien didn’t have access to ChatGPT because it was freaking 1954. The tech reviewers at CNET couldn’t possibly have had access to OpenAI’s models a good decade before the models were created in the first place.

Today, things are different. Modern AI detectors still make false accusations. But they have had their own training on a good body of pre-AI text to tell it apart. Even if someone has written (by accident, let’s say) something that ChatGPT would write, and it’s popular, the AI detector will tell you it’s human. Change the topic and keep the structure the same; it’s suddenly AI.

And that brings us to our next point.

How Does a Modern AI Detector Work? Is There a Way to Reduce False Positives?

Modern AI detector tools from Originality, GPTZero, ZeroGPT, QuillBot, Sapling, Copyleaks, Scribbr, Winston, etc., work differently. And that gives clients a false sense of accuracy. Over the years, just to prove we’re indeed humans, we have had to use premium versions of many of these tools like Winston, Originality, GPTZero, etc., at different times. And we learned certain things.

Yes, 100% plain output from ChatGPT will get marked as AI accurately. 100% human content that has no formulaic patterns will get marked as human accurately. But anything in the middle is deceptively inaccurate.

Now, you’ll find a bunch of articles online telling you how perplexity and burstiness and maybe a bunch of other stuff help you avoid AI detectors. All of that is rubbish and went out of the window. Here’s what we did:

  • We planned an experiment. Dozens of human samples that came back as 0% AI and dozens of AI-generated samples that came back as 100% AI were taken.
  • Structural qualities were interchanged among them little by little. Word choices, sentence length variation, tone changes, transition words, and yes, the usual bunch of variables as well, like burstiness and perplexity.
  • The experiment’s objective was this: Tweak one thing at a time and see what makes human text AI and what makes AI text human in the eyes of these detectors.

We burned a ton of tokens on two leading AI detectors in doing this experiment. And here’s our list of variables that affect this “AI detection confidence” or score (most important to negligibly important):

  1. AI Model: Same brief, same topic, but different models. Claude’s Sonnet hit 75% consistently across samples, whereas GPT-5.5 hit 100% consistently.
  2. Token-Level Predictability (Perplexity): The per-token predictability is a measure of how AI tools generate text. This is extremely important. Understand that AI only has a handful of choices, as it’s essentially a glorified auto-complete in the hands of an amateur AI-based article writer. A human has more varied choices not just for the next word, but for the flow itself. Your topic is climate change. AI explains the ozone layer. Now it must follow it up with how we’re hurting it, why it is important to protect it, and maybe a few more possibilities. A human might choose to explain what the ozone layer is and then talk about that one obscure hoax about the ozone layer from 5 years ago because their memory happened to bring it up. A different human might follow up the explanation of the ozone layer by an example of a cell membrane to drive the point home. A different human writer might choose to go with a negative remark, a positive anecdote, historical data, further expansion, a specific region’s story, a personal remark, and so on.
  3. Expansion Ratio: More padding creates longer, more uniform stretches.
  4. Register or Contraction Density: Contracted text is likely to beat (sometimes, only marginally) formalized, uncontracted text.
  5. Summary: Wrap-ups, summaries, verdicts, etc., even when there’s no need, are known AI tells.
  6. Detector Noise: Shorter detection samples are noisier.
  7. Discourse Structure: Structure probably contributes a little, but it’s not a score-changing metric on its own. How information is ordered is different between human and AI text.
  8. Burstiness: This is sentence-length variance. Classic theory that it affects AI detection scores. Coupled with perplexity, might be significant.
  9. Mechanical Imperfections: Comma splices, inconsistent caps, number style, usage of backslashes, etc. might move the needle ever so slightly in edge cases. But asking an AI generator to introduce sloppiness deliberately is very different from human-originating writing mistakes.
  10. Signposting Phrases: Phrases like “another important feature” also affect the score.
  11. List Parallelism: AI loves clean, parallel lists. Minor tell.
  12. Residual Smoothness: Every sentence must land solidly. Have a premise, set-up, and an end. Be whole on its own. This is AI’s way of writing. A human often connects thoughts and transitions between ideas, trails off, connects random things beautifully, etc.
  13. Nested Sentences: How many ideas are nested within the same sentence? Sometimes, humans nest a lot.

And let’s talk about the myths:

  • Certain phrases/words give rise to AI confidence level. Nope. If you’re writing something completely on your own, it’s a short couple of paragraphs long piece, and you intentionally inject a few “AI words” like embrace, in the landscape of, etc., you will still get a human score.
  • The long ones—what we call em dashes—don’t increase AI likelihood at all. We injected em dashes into 100% human text; it was still 100% human. We removed all em dashes from an 80% AI text sample. The score remained 80%.
  • Constrained technical vocabulary might read more uniformly, but no evidence here either way. Speculative.

The best way to summarize all of my findings is this: AI detectors look at surrounding tissue, not exact words/phrases. If more variables (tells) are present, it’s more likely that the detectors will tune themselves to a lower threshold and capture more sentences as AI-like. If fewer tells are present, the exact same sentences will now get lower AI-like scoring.

We were able to get many parts of AI-generated text flagged as human. And fewer parts of human-written sections flagged as AI. No clear, objective knob you can turn here. But it’s clear as day that AI detectors are inaccurate in the middle ground. At least the Bible is no longer AI-generated according to Originality.

Formulaic vs. Thought-Led Writing

Okay, so what was the #1 thing I learned from this experiment and years of detection, arguments with clients, and rewriting to beat AI detectors? There are two main ways of writing: formulaic and thought-led.

Read passage A:

The first thing to understand about Swarm is that it’s not a separate software for you to install somewhere. It’s already built into Docker. You simply turn it on (activate it with docker swarm init). Docker’s Swarm mode has nodes, like bees. A node can be a manager, a worker, or both. Managers make scheduling decisions and maintain cluster state via the Raft consensus algorithm. The worker nodes run the actual containers.

Now, read passage B:

Docker Swarm is Docker’s built-in orchestration mode. It is not a separate product you need to install, and it does not require a separate orchestration stack to get started. If a server is already running Docker Engine, swarm mode is already available. You activate it with a simple command:

docker swarm init

Once enabled, Docker can manage a cluster of machines as one coordinated environment instead of leaving you to run containers manually on individual servers.

A Docker Swarm is made up of nodes. These nodes can be manager nodes, worker nodes, or both. Manager nodes handle the control plane. They make scheduling decisions, maintain the cluster’s state, and use the Raft consensus algorithm to keep that state consistent across the cluster. Worker nodes run the actual containers that your applications depend on.

A is human. B is GPT. Here’s the main difference. A is thought-led. B is formulaic. Why?

  • Thought-led means the human is thinking as they are forming the words. They add words based on that.
  • Formulaic means every sentence and section is self-contained (much like a Docker swarm worker node itself). It has a clear ending. No trailing off, bleeding meaning into the next sentence/paragraph, and elegant transition.

The passages are taken from our sample articles. Both articles are likely AI according to Originality AI, but 15% likely to be AI (human sample) and 80% likely to be AI (AI sample) according to GPTZero.

And the Joke

A side experiment I did was this: when Originality flags your content as AI/not AI, it assigns a confidence only. 40% doesn’t mean 40% is AI, it just means 40% confidence that AI generation was used.

A flawed system at best. And I suspect that’s why they have changed it more recently. Now you select a threshold of how much AI-likelihood is okay, and then it tells you the likelihood/confidence of that likelihood. When something needs to be so convoluted to interpret, it’s likely undependable by design.

Anyway, that’s not the joke. The real joke was this: a sample was roughly 40% AI. Originality assigns sentence-level color coding and tells you how likely that sentence is to be AI-generated or human-written. You can then pay extra tokens for a “deep scan” or “deep insights.” I have forgotten the exact term and I don’t care to top up my account again.

Once you do that, it tells you probably why that sentence is AI-like. So I took all sentences that scored over 50% AI-like and read their analysis. There were suggestions that Originality served to me. Do this, avoid that, this sentence does this, this sentence does that, this is why this is AI-like, etc. I followed 100% of those suggestions. I surgically rewrote all sentences having 50% or more likelihood of being AI, following exactly what Originality itself told me to do.

The overall score of “40% likely to be AI” actually jumped to “80% likely to be AI.”

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top