FreeToGenerate.com

1,277 impostors from 43 scripts — and most of them are not an alphabet.

Text to check

Checked in your browser against Unicode's own confusables table. Nothing is uploaded.

Lookalikes found, and a mixed-script check would not object to this text.

Either every character comes from one script, or the impostors belong to Common — which Unicode treats as compatible with everything. This is the case that gets through.

Reads as: paypal.com

The characters that are not what they look like

AtCharacterUnicode nameScriptImitates
0𝐩 U+1D429MATHEMATICAL BOLD SMALL PCommonp
1𝐚 U+1D41AMATHEMATICAL BOLD SMALL ACommona
2𝐲 U+1D432MATHEMATICAL BOLD SMALL YCommony
3𝐩 U+1D429MATHEMATICAL BOLD SMALL PCommonp
4𝐚 U+1D41AMATHEMATICAL BOLD SMALL ACommona
5𝐥 U+1D425MATHEMATICAL BOLD SMALL LCommonl
7𝐜 U+1D41CMATHEMATICAL BOLD SMALL CCommonc
8𝐨 U+1D428MATHEMATICAL BOLD SMALL OCommono

The table behind this

impostors
1,277
scripts
43
from Common
823
from Cyrillic
43
from Greek
38

The attack is named after Cyrillic, and most of the material is not Cyrillic. Common — mathematical alphanumerics, fullwidth forms, enclosed letters — supplies more impostors than every alphabet put together, and because Unicode treats it as script-neutral, a mixed-script check raises no objection to any of them.

Every Latin letter has at least one single-character impostor except two. Capital I is not a target at all, because Unicode folds I, l and 1 onto one representative. Lowercase m is imitated by the two-character sequence rn. The two gaps are exactly the two famous spoofs that work another way.

Our Cyrillic and Greek alphabet pages each carry a lookalike box built from a hand-written table for that one alphabet — 22 characters apiece. They stay useful as per-alphabet references; this page is the whole table.

Data from Unicode's confusables.txt, the table behind UTS 39, with script names from Scripts.txt. Positions are counted in code points, since several impostors sit above the basic plane.

Also available in: Español · Português · Français · العربية

Lookalike Characters: Check Text for Homoglyphs

Which characters in your text are pretending to be Latin letters, what script they come from, and whether the usual defence would catch them.

What is a lookalike character?

A lookalike character — a homoglyph — is a character from one script that is drawn almost identically to a character from another. The Cyrillic а and the Latin a are different code points that most fonts render the same way. Put one inside a domain name and you have a web address that reads correctly and goes somewhere else.

Unicode publishes the authoritative list of these as confusables.txt, the data behind its security report UTS 39. That file is what registrars and browsers consult when deciding whether two names look too alike to both exist. Reduced to single characters standing in for a single Latin letter, it holds 1,277 impostors drawn from 43 scripts.

This page checks your text against that table, names every impostor it finds, and tells you the one thing most tools do not: whether the standard defence would have caught it.

How to use it

  1. Paste the text. A domain, a username, a package name, an email address — anything you are about to trust. It is checked in your browser and never uploaded.
  2. Read the verdict. You get the Latin spelling the text actually resolves to, and a plain statement of whether a mixed-script check would object to it.
  3. Then look at the table. Every impostor is listed with its position, its code point, its official Unicode name, the script it belongs to and the letter it imitates.

The attack is named after Cyrillic and mostly is not Cyrillic

Cyrillic contributes 43 of the 1,277 impostors and Greek another 38. The largest single source, by a wide margin, is the script Unicode calls Common: 823 of them. Common is where the mathematical alphanumerics live, along with fullwidth forms, enclosed letters and a good deal else that is not part of any alphabet.

That is not a curiosity, because of how the usual defence works. The cheap and near-universal check is to ask whether a string mixes scripts, on the reasoning that a genuine Latin word has no business containing a Cyrillic letter. Unicode treats Common as compatible with every script — mathematical symbols are supposed to appear alongside any language — so a mathematical bold a sitting in the middle of Latin text raises no mixed-script objection at all.

There is a second way past the same check, and it is the one the Cyrillic reputation comes from: spell the whole word in one script. A word written entirely in Cyrillic characters is perfectly consistent, so nothing is mixed, and the check has nothing to say. The tool reports both cases explicitly rather than leaving you to work out which one you are looking at.

Two letters nobody can fake with one character

Every Latin letter has at least one single-character impostor except two, and the reasons are different — both recorded in the file itself.

Capital I has none because it is not a target at all. Unicode folds capital I, lowercase l and the digit 1 onto a single representative, so the table maps I to l rather than the other way round: 79 characters point at lowercase l, and none point at capital I. Lowercase m has none because its impostor is not one character. The table maps m to rn, the two-letter sequence that has been fooling readers since long before Unicode existed.

So the two gaps in the list are precisely the two famous spoofs that work by a different mechanism. That is a pleasing result, and the generator that builds this dataset refuses to emit it if the finding ever stops holding.

What this page cannot tell you

It cannot tell you that text is malicious. A great many of these characters have entirely legitimate uses: mathematical alphanumerics belong in mathematics, fullwidth forms belong in Japanese typesetting, and a Cyrillic а in a Russian word is simply the letter a. What the page reports is that a character is not the Latin letter it resembles, which is a fact rather than an accusation.

It also does not check the reverse direction. Unicode records 2,103 characters whose lookalike is a sequence rather than a single character — m for rn is the famous one — and those are a different shape of problem that a per-character scan cannot see.

And it is only about appearance. Text can deceive in ways that have nothing to do with how characters are drawn, most obviously by hiding characters that render as nothing at all. That is a separate check, and there is a separate page on this site for it.

Why is it free?

The table is a few kilobytes shipped with the page, and the check runs in your browser. Nothing you paste is uploaded, nothing is logged, and there is no account to make.

No sign-up, no limits, and no watermark on anything you copy out.