Free AI detector: what the 11 signals actually measure

By · · Inside Humanize 360

Every AI detector on the internet gives you a percentage. Very few tell you where the number came from. Ours does, because a number without an explanation is not something you can act on. This post walks through the eleven signals behind the Human Score in the Humanize 360 AI detector, what each one measures, and what the score does and does not mean.

First, what a detector is not

No AI detector can prove who wrote a text. Not ours, not the expensive ones sold to universities. All of them measure statistical properties that machine-generated prose tends to have more often than human prose. A careful human writer can trigger them; a carefully edited machine draft can avoid them. That is why we call our output a Human Score rather than a verdict, and why we show the signals instead of hiding them.

If you are a teacher, that also means a score alone is never grounds for an accusation. Treat it as a reason to have a conversation.

The eleven signals

1. Sentence length variance

Human writers produce sentences of wildly different lengths. Models cluster around a mean. We measure the standard deviation of words per sentence; low variance pushes the score down.

2. Sentence opener variety

How many of your sentences start with the same part of speech, or the same word? Machine drafts lean on "The", "This" and "It", and on adverbs like "Additionally". We count distinct openers across a window of ten sentences.

3. Tell-word density

The number of phrases per hundred words drawn from our list of expressions that occur far more often in generated text than in published writing: "delve", "tapestry", "it is important to note", "in today's fast-paced world" and several hundred more.

4. Lexical diversity

The ratio of unique words to total words, adjusted for length. Models repeat their favourite mid-frequency words. Humans repeat function words but range more widely elsewhere.

5. Contraction rate

"Do not" versus "don't". Formal human writing uses few contractions; casual human writing uses many. Machine drafts sit at an odd middle and often use none at all in a piece that otherwise reads casually. We compare the rate against the register the text seems to be aiming for.

6. Paragraph shape

Generated text loves the five-sentence paragraph with a topic sentence, three supports and a wrap-up. We look at the distribution of paragraph lengths and the presence of one-sentence paragraphs, which humans use for emphasis.

7. Specificity

The density of proper nouns, numbers, dates and quoted material. Vague text ("various stakeholders", "numerous benefits") scores low. This is the signal that most rewards real knowledge.

8. Hedging and filler

Phrases that add length without meaning: "it can be argued that", "in many ways", "to a certain extent", "plays a crucial role". Counted per hundred words.

9. Transition stacking

Whether consecutive paragraphs open with connective adverbs ("Furthermore", "Moreover", "In conclusion"). Three in a row is a strong machine pattern.

10. Punctuation range

Humans use dashes, parentheses, semicolons, questions and exclamation marks unevenly. Models default to full stops and commas. We measure the spread of punctuation types.

11. Predictability

A lightweight word-sequence model estimates how "expected" each word is given the previous few. Text that is uniformly predictable scores low. This is our closest cousin to the "perplexity" measure other detectors use, but it is one signal among eleven rather than the whole story.

How the signals combine

Each signal is scaled to a 0 to 100 range and weighted. Rhythm and tell-word signals carry the most weight because they are the most stable across topics; predictability carries less because it is sensitive to subject matter (a chemistry abstract is predictable whoever wrote it). The Human Score is the weighted mean, and the breakdown shows every signal so you can see which ones dragged the score down.

Reading a result

  • 80 and above: the text has the texture of human writing. Stop editing.
  • 60 to 79: readable, but one or two signals are weak. Usually rhythm or tell-words. Fix those.
  • Below 60: several machine patterns at once. Run it through the humanizer, then add your own specifics.

Short texts are noisy. Under 150 words, treat any score with suspicion; under 50, we do not show one at all.

Free limits

The detector is free for 300 words per check and three checks a day without an account. With a free account you get 1,500 words per check. Paid plans meter it at one fifth of a word credit per word, so a 1,000-word essay costs 200 credits from your wallet. The pricing page has the full table.

Why we built it this way

We could have trained a classifier and shown a single percentage. We chose measurable signals because they are explainable, because they do not need your text to leave our server, and because they give you a to-do list rather than a grade. When you then humanize the text you can watch the same eleven numbers move. That feedback loop is worth more than any verdict.

Students, read the honest-use notes before you lean on any of this. The detector is for checking your own drafts, not for beating a system.

Common questions

Can the score be wrong? Yes, in both directions. A formal report by a careful human can score in the sixties because its rhythm is even; a lightly edited machine draft can score in the eighties. Use the breakdown, not the headline number.

Does the detector store my text? It is processed on our server and discarded. Logged-in users can keep results in their private history; anonymous checks are not saved at all.

Why not show a percentage of "AI"? Because we cannot measure that. We can measure eleven properties of the text in front of us, and we would rather show you those honestly than dress a guess up as a fact.

Which signals should I fix first? Rhythm and tell-words. They carry the most weight and they are the quickest to change: split two sentences, delete three phrases, re-run.

ai detector human score writing