FreeToGenerate.com

Twenty-two of these letters can pass for Latin ones, and a word built from them is not mixed-script at all.

The 33 letters of the Russian alphabet

Ё is the seventh letter but sits at U+0401, outside the run that holds the other 32. Generating this alphabet as one range gives 32 letters and loses Ё without a word.

LetterNameCode pointsLatin lookalike
А аAU+0410 · U+0430A a
Б бBeU+0411 · U+0431
В вVeU+0412 · U+0432B
Г гGheU+0413 · U+0433 r
Д дDeU+0414 · U+0434
Е еIeU+0415 · U+0435E e
Ё ёIoU+0401 · U+0451
Ж жZheU+0416 · U+0436
З зZeU+0417 · U+0437
И иIU+0418 · U+0438
Й йShort IU+0419 · U+0439
К кKaU+041A · U+043AK
Л лElU+041B · U+043B
М мEmU+041C · U+043CM
Н нEnU+041D · U+043DH
О оOU+041E · U+043EO o
П пPeU+041F · U+043F
Р рErU+0420 · U+0440P p
С сEsU+0421 · U+0441C c
Т тTeU+0422 · U+0442T
У уUU+0423 · U+0443Y y
Ф фEfU+0424 · U+0444
Х хHaU+0425 · U+0445X x
Ц цTseU+0426 · U+0446
Ч чCheU+0427 · U+0447
Ш шShaU+0428 · U+0448 w
Щ щShchaU+0429 · U+0449
Ъ ъHard SignU+042A · U+044A
Ы ыYeruU+042B · U+044B
Ь ьSoft SignU+042C · U+044Cb
Э эEU+042D · U+044D
Ю юYuU+042E · U+044E
Я яYaU+042F · U+044F

Letters Russian does not use

There is no single Cyrillic alphabet. These sit just before the Russian block and belong to other languages written in the same script.

LetterNameUsed by
Ђ ђDjeSerbian
Ѓ ѓGjeMacedonian
Є єUkrainian IeUkrainian
Ѕ ѕDzeMacedonian
І іByelorussian-Ukrainian IUkrainian, Belarusian
Ї їYiUkrainian
Ј јJeSerbian, Macedonian
Љ љLjeSerbian, Macedonian
Њ њNjeSerbian, Macedonian
Ћ ћTsheSerbian
Ќ ќKjeMacedonian
Ў ўShort UBelarusian
Џ џDzheSerbian, Macedonian

Check text for lookalikes

Paste anything and see which characters are Cyrillic letters standing in for Latin ones.

Lookalikes found

Reads as
care
  • Position 1 · с · U+0441 · Imitates c
  • Position 2 · а · U+0430 · Imitates a
  • Position 3 · г · U+0433 · Imitates r
  • Position 4 · е · U+0435 · Imitates e

This text is entirely Cyrillic, so a mixed-script check sees nothing wrong with it. That is the whole problem.

Also available in: Español · Português · Français · العربية

The Cyrillic alphabet

All 33 Russian letters with their code points, the letters other Slavic languages add, and which ones imitate Latin.

What is the Cyrillic alphabet?

Cyrillic is the script used to write Russian, Ukrainian, Bulgarian, Serbian, Macedonian, Belarusian and dozens of languages across Eurasia. It descends from the Greek uncial alphabet by way of ninth-century Bulgaria, which is why several of its letters are Greek rather than Latin in origin and why a few look like neither.

There is no single Cyrillic alphabet, which is the first thing to get straight. Russian uses 33 letters and that is the set most people mean, but Ukrainian has і, ї, є and ґ, Serbian has ј, љ, њ, ћ and џ, and Bulgarian drops two letters Russian keeps. The table above lists the Russian 33 first and then the letters that belong to the other alphabets, because presenting one language's set as the alphabet is the mistake nearly every list makes.

The second thing is that the Russian alphabet is not a tidy range in Unicode. Thirty-two of its capitals run consecutively from U+0410 to U+042F, and the thirty-third, Ё, sits far away at U+0401 among letters Russian does not use at all. Anyone generating this alphabet from a range gets 32 letters and loses Ё without an error, which is exactly why the list here is built from the character database rather than counted off.

How to use this page

  1. Read the table for the letter you need. Each row gives the capital and lowercase forms, the letter's name, both code points, and the Latin letter it can be mistaken for if there is one. The code points are what you need when a letter has to survive a copy and paste into a system that stores bytes rather than glyphs.
  2. Check the second table before assuming a letter is Russian. The thirteen letters listed there are used by Ukrainian, Serbian, Macedonian or Belarusian and not by Russian. If you are transcribing a name or a place, that distinction is usually the thing that matters.
  3. Paste suspicious text into the checker. It flags every character that is a Cyrillic letter standing in for a Latin one, shows the position and code point of each, and prints what the text reads as once the imitations are resolved. It also tells you whether the string mixes scripts, which is the part worth understanding.

The lookalikes, counted rather than assumed

Cyrillic is the script that homograph attacks are named after, so the obvious expectation is that it has more Latin lookalikes than anything else. Measured against confusables.txt, the table Unicode publishes behind its security recommendations and the one registrars actually consult, it does not. Thirteen Russian capitals and nine lowercase letters are marked confusable with a Basic Latin letter, twenty-two in all. Greek, which this site measured the same way, has fourteen capitals and eight lowercase, also twenty-two, and is ahead on capitals.

So the count is a tie, and the reason Cyrillic dominates in practice is elsewhere. It is which letters the lookalikes cover. Cyrillic lowercase can stand in for a, c, e, o, p, r, w, x and y. Greek lowercase covers a, i, o, p, u, v and y. Those nine Cyrillic substitutes include c, e, r and w, and that is enough English to build whole words rather than to swap one letter inside them.

Run that against a real dictionary and the difference is stark. Of the 167,968 words in the ENABLE list this site uses for its word games, 394 can be written entirely in Cyrillic lookalikes and only 43 can be written entirely in Greek ones, a factor of about nine. Care, acre, carp, opera, reappear and preparer are all in the Cyrillic set. Score is not, because s has no Cyrillic imitator, and neither is apple, because l has none.

That is what makes the difference dangerous rather than merely interesting. The cheap and common defence against this kind of spoofing is a mixed-script check: if a string contains both Latin and Cyrillic characters, something is probably wrong. A word written entirely in Cyrillic contains no Latin characters at all, so it is perfectly consistent, single-script text and that check has nothing to say about it. The checker above demonstrates it: the word it opens with is not mixed-script, and it is not the word it appears to be.

Honest limits

The confusable pairs here come from Unicode's table, not from anyone's judgement of what looks alike, and that table is about shapes in isolation. Whether two characters are actually indistinguishable depends on the font you are reading in, and a few pairs that Unicode marks confusable are quite easy to tell apart in a typeface with a large x-height or distinctive terminals. Treat the column as the standard's answer rather than a promise about your screen.

Normalisation does not help, and this is worth stating plainly because it is the first thing people reach for. Unicode normal forms exist to unify characters that are canonically or compatibly the same, and a Cyrillic letter is not the same character as a Latin one that happens to look like it. Neither NFC nor NFKD folds them together, so the reflexive normalise-and-compare check catches none of this.

The page also covers only Basic Latin lookalikes. Cyrillic characters can imitate digits, punctuation and letters from other scripts, and confusables.txt records far more pairs than the ones shown here. Narrowing to A to Z keeps the table readable and matches what people are usually asking about, but it is a narrowing.

Finally, this is a reference and a checker, not a security control. Knowing that a string contains homoglyphs tells you a string is suspicious; it does not tell you a domain is hostile, and plenty of entirely legitimate text mixes scripts. Real defences live in browsers and registries, which apply script-mixing rules, per-registry allowed character sets and display policies that no single page can reproduce.

Why is it free?

Because it is a table and a small amount of arithmetic. Nothing you paste into the checker is uploaded, logged or stored: the character data ships with the page and the comparison happens in your browser, which matters more than usual here given that the natural thing to paste is a domain or a username you are already suspicious of.

So there is no account, no sign-up and nothing held back. The letters and their names come from the Unicode Character Database and the lookalikes from confusables.txt, both regenerated by a script committed beside the page, along with 70 checks that verify the result.