Typosquatting.ai
PATTERNS

Homoglyph Domains: Letters That Look Alike

The reader sees a shape, not a code point. Anything with the same shape passes. Illustrative lookalikes in this guide use the reserved .test ending, so that no real registration is named.

by Typosquatting.ai Research23 September 20265 min read

A homoglyph domain replaces one or more characters of a real name with characters that look the same or nearly the same, within the Latin script, such as rn for m and 1 for l, or across scripts, such as a Cyrillic a for a Latin a.

This guide covers what makes two characters confusable, the ASCII homoglyphs that need no special script, the mixed-script names that browsers now display as punycode, how many candidates each produces, and how the checker ranks them.

Family in a report
Lookalike characters
Sample
exarnple.test
Candidates for a six-letter domain
41 of 3,001, plus 2 Cyrillic

What a homoglyph is

Unicode Technical Standard #39 defines two strings as confusable when their skeletons are equal: strip each character down to the basic shape it is drawn with, and if the results match, the strings can be mistaken for one another. That definition covers the ASCII cases, where rn resembles m and l, 1 and I resemble one another, and the mixed-script cases, where a Cyrillic а is drawn exactly like a Latin a.

The generator produces both. Lookalike characters, 41 candidates for a six-letter label, substitutes one character at a time: an ASCII pair such as rn for m, or a single letter from another script such as a Cyrillic а. Cyrillic lookalike, two candidates, is the whole-label form, every letter replaced by its Cyrillic twin where one exists.

FamilyExampleCandidates for a six-letter label
Lookalike charactersexarnple.test (rn for m), exаmple.com (Cyrillic а)41
Cyrillic lookalikeехаmрlе.com, every letter Cyrillic2

The ASCII cases need no trick

rn for m is the classic, and it survives every defence that depends on scripts: it is plain ASCII, it registers anywhere, and in a proportional font at small size the two shapes are close to identical. l and 1 and I, 0 and o, vv and w, cl and d are the others the generator uses. These pass in an email, a chat message and a mobile address bar, and no browser converts them, because there is nothing to convert.

The mixed-script cases and what browsers do

In December 2001 Gabrilovich and Gontmakher registered a variant of microsoft.com with Cyrillic characters to show the attack was practical, and in 2005 a PayPal demonstration was disclosed at Shmoocon. Browsers responded by displaying a domain that mixes scripts, or that is confusable with a well-known name, in its punycode form, xn-- followed by an encoding, rather than as the lookalike letters. That is why the checker reports these names in punycode: the form a browser would show is the form you should recognise.

The blog guide on IDN homograph attacks works through the browser rules and the registries' own restrictions in more depth.

What a registered homoglyph usually is

Homoglyphs are not typed by accident. Nobody reaches exarnple by mistake, so a registered one was registered to be looked at rather than typed, in an email or a message where the reader glances. Defensive registration by the brand is the other explanation, and the Possibly yours tag will say so when the registrar and nameservers match. A homoglyph that is active, recent and carries a certificate is among the strongest signals a report can show, and it is still a review priority, not a finding of phishing.

How the checker finds and ranks it

A finding's review priority is a published rule, not a score. A name with an address or a mail record is active; active names registered under 90 days ago are Elevated, and so are active names under a year old that carry an impersonation keyword. Anything else that is active or has a registry record is Review, a name with no registry record and no DNS answer is Low signal, and a lookup the registry refused stays Unknown rather than being counted as a negative.

A single lookalike substitution is one edit away by distance, so it wears the nearest band; the whole-label Cyrillic form is several edits away in code points while being zero edits away to the eye, and the report lists it in punycode so the difference is visible.

What to do with one

Read the row. If it is active and recent, keep the report and file with the registrar; homoglyph evidence is usually the clearest case a registrar sees. Never open the name, and do not paste it into a browser to compare: the report's punycode already shows what it is.

Common questions

What is the difference between a homoglyph and a homograph?
A homoglyph is one character that looks like another. A homograph is a whole name built from them. An IDN homograph attack uses characters from another script; ASCII homoglyphs such as rn for m need no other script.
Do browsers protect against homoglyph domains?
Against mixed-script names, largely, by showing punycode. Against ASCII lookalikes such as rn for m, no, because there is nothing to convert.
Why does the report show xn-- names?
Because that is how a browser would display them. A lookalike in Cyrillic is shown in the encoded form so you can see it is not your name.
Can someone register a Cyrillic copy of my .com?
Only where the letters have Cyrillic twins, and many registries restrict mixed scripts. The generator produces single-letter substitutions in the lookalike family and the whole-label form in the Cyrillic family, two candidates for example.

Sources and further reading

  1. Unicode Technical Standard #39: Unicode Security Mechanisms
  2. Wikipedia: IDN homograph attack
  3. IDN homograph attacks: how lookalike letters fool browsers
  4. Types of typosquatting: 20 patterns with real examples
  5. Typosquatting.ai methodology: patterns, sources and review rules

Keep reading in Typosquatting patterns