Typosquatting.ai
TECHNICAL

IDN Homograph Attacks: How Lookalike Letters Fool Browsers

IDN homograph attacks: how Punycode and confusable characters make a lookalike render as the real domain, the 2005 and 2017 demonstrations, and what to do.

by Typosquatting.ai Research9 September 2026updated 23 September 202610 min read

Same look, different code points

An IDN homograph attack registers a domain whose characters look like another domain's but are different Unicode code points. Cyrillic a (U+0430) and Latin a (U+0061) render almost identically in many fonts. Internationalised domain names make this possible: RFC 5890 defines the U-label, the Unicode form a person reads, and the A-label, the ASCII form the registry stores, which always begins with the prefix xn-- followed by the output of the Punycode algorithm from RFC 3492. The browser decides which of the two forms to display.

Nothing is being exploited here in the sense of a flaw. The system is working: it was built so that people could register names in their own scripts, and a side effect of supporting every script is that some characters in different scripts look the same.

U-label
The Unicode form a person reads, such as the one containing a Cyrillic letter.
A-label
The ASCII form the registry stores, beginning with xn--.
Punycode
The encoding that converts between the two, defined in RFC 3492.
Confusable
A character that renders similarly enough to another to be mistaken for it. Unicode Technical Standard #39 defines two strings as confusable when their skeletons are equal.
Skeleton
A normalised form produced by mapping every confusable to one representative, so two names that look alike compare equal.

Two incidents that defined the problem

The attack was described in theory before browsers supported internationalised names, and twice it was demonstrated in a way that forced browsers to change.

The first was in February 2005. Wikipedia's account records that on 6 February 2005 Cory Doctorow reported a demonstration disclosed by 3ric Johanson of the Shmoo Group at the ShmooCon conference, in which browsers supporting IDNA appeared to direct a paypal.com address whose first a was a Cyrillic а to the real payment site, while it actually led to a spoofed page. The Mozilla bug opened for it describes the problem plainly: browsers incorrectly handled Punycode-encoded domain names, which allowed phishers to spoof the URLs of just about any domain name, including SSL certificates, and it gave the proof of concept on shmoo.com. Mozilla published a security advisory on 24 February 2005, noting that many supported characters are similar to others, if not identical in some fonts, so the mechanism could be used to construct perfect, indistinguishable phishing sites. The registered name behind the demonstration was xn--pypal-4ve.com.

The second was in April 2017. Xudong Zheng published a page on 14 April 2017 describing a domain written entirely in Cyrillic letters, xn--80ak6aa92e.com, which rendered as apple.com in the address bar. His write-up explains why the earlier defence did not catch it: the homograph protection in Chrome, Firefox and Opera failed when every character was replaced with a similar character from a single foreign script, because no script mixing was present to trigger it. Chrome's fix shipped in Chrome 58. Chromium's IDN policy document now uses the same string as its worked example of a whole-script confusable label shown in Punycode. Mozilla's bug on the 2017 case was first resolved as will not fix, and its reply to a request to show Punycode for everything was that it would make all non-Latin domain names show as gibberish.

What a rendered example looks like

The table shows five strings, their code points, the A-label the registry holds, and what the two main browser engines display. The two Cyrillic examples are the ones from the incidents above. The last row is an ASCII lookalike, included because no browser rule ever applies to it.

Unicode stringCode pointsA-labelChromium showsFirefox shows
pаypal (2005)Latin p, Cyrillic U+0430, then Latinxn--pypal-4vePunycode: mixed scriptsPunycode: mixed scripts
аppleU+0430 then Latin p p l exn--pple-43dPunycode: mixed scriptsPunycode: mixed scripts
аррӏе (2017)U+0430 U+0440 U+0440 U+04CF U+0435xn--80ak6aa92ePunycode: whole-script confusable under .comUnicode: single script is permitted
googléLatin with U+00E9xn--googl-fsaPunycode: skeleton matches a top domainUnicode: single script
exarnpleAll ASCII, rn in place of mexarnpleAs typedAs typed

How browsers limit the damage

Chromium and Firefox both show mixed-script labels in Punycode. Chromium goes further. Its published policy says that if all the letters in a label belong to a set of whole-script-confusable letters and the hostname does not have an allowed top-level domain for that script, it shows Punycode, and that if the skeleton of the registrable part of a hostname matches one of the top domains, it shows Punycode too; the policy has ignored the browser's language settings since Chrome 51. Firefox's published algorithm displays Unicode when all characters in each label come from a single script or an allowed combination; it states that the system permits whole-script confusables, that other browsers' solutions did too, and that in the end it is up to registries to make sure their customers cannot deceive each other. Mail clients differ again, and plain-text contexts such as SMS show whatever the sender wrote.

Where the address appearsWhat the reader is shownWhat that means for you
Chromium address barPunycode for mixed script, whole-script confusables and skeleton matchesThe strongest default protection, and the reason many attempts fail quietly
Firefox address barPunycode for mixed script, Unicode for whole-script confusablesA name blocked in one browser can render cleanly in another
Mail clientsVaries by client and by whether the link text was written by handDisplay rules cannot be assumed from the browser
Plain text, SMS, printed codesExactly what the sender wroteNo protection at all, because nothing is parsing it as a URL

Registries block some of this before it is ever registered

Most registries publish IDN tables that limit which scripts and which characters may appear in a name, and many forbid mixing scripts inside one label. A large share of the mixed-script names people worry about cannot be registered in the first place.

That protection has an edge. It applies per registry, it varies widely, and it does nothing about confusables that live inside a single permitted script. The 2017 demonstration was registered under .com precisely because an all-Cyrillic label is a valid single-script name there; Mozilla's first response to the bug was that any complaint belonged with the registry that let it be registered.

Confusables inside a single script

Not every lookalike crosses scripts. Within Latin, rn resembles m, a lowercase L resembles a capital i in many sans-serif fonts, and 0 resembles o. Unicode Technical Standard #39 publishes a confusables table that maps characters to a common skeleton, and it distinguishes mixed-script confusables, whose resolved script sets share nothing, from whole-script confusables, where each string is a single script. Security tooling compares skeletons to decide whether two names would be visually confused.

These are the ones that survive every layer above. The registry permits them because they are ordinary Latin characters. The browser displays them because there is no script mixing to report. The reader sees a name that looks right.

Skeleton matching answers whether two strings look alike. It does not answer whether anyone intended them to, which is a different question and not one a table can settle.

What a reader can do

Most of the defence is in the browser and the registry, but a reader has habits available that do not depend on either.

  • Copy the address as text and look for xn--. Paste it into a plain text editor rather than reading it in the address bar. If any label begins with xn--, the name you were shown was a Unicode rendering of something else.
  • Let a password manager do the matching. A manager fills credentials only on the exact domain it saved them for, and it compares the A-label, not the picture. A login page that looks right but gets no autofill is a page on a different domain.
  • In Firefox, set network.IDN_show_punycode to true in about:config. Zheng's 2017 write-up recommends this as the way Firefox users can limit their exposure; every internationalised name then displays in its xn-- form.
  • Do not trust the padlock to settle it. A certificate is issued for the A-label the registrant controls, so a homograph domain can carry a valid one. The padlock confirms the connection, not the identity you had in mind.
  • Type or bookmark the addresses that matter. A bookmark carries the A-label the browser resolved when you saved it, and it cannot be nudged by a link in a message.

How to investigate a suspected homograph

Copy the address from the message into a plain text editor and look for xn-- at the start of a label; that is the sign that the displayed form was Unicode. Run the Punycode form through a registration lookup and compare the registrar, creation date and nameservers with your own domain. Do not visit the site to see whether it looks like yours; passive evidence and registry data are enough to decide whether to escalate.

  1. Copy the address as text. Do not follow it, and do not retype it from a screenshot.
  2. Look for xn-- at the start of any label. If it is there, the name you were shown is not the name the registry holds.
  3. Look the ASCII form up in RDAP, the standardised replacement for WHOIS, and note the registrar, the registration date and the status codes.
  4. Compare the nameservers and registrar against your own domain, since a match frequently means the name is one of yours.
  5. Keep the address, the record and the time. If you escalate, that is the evidence pack.

What a checker does not enumerate

Typosquatting.ai generates ASCII variations from about twenty pattern families, including the rn-for-m substitution and a keyboard-neighbour map, across about 180 endings, and normalises any internationalised name you enter to its Punycode form so the registry lookup is exact. Its evidence comes from registry RDAP, DNS over HTTPS, certificate transparency logs and passive urlscan.io search, and it never connects to a suspect site. It does not enumerate Unicode homographs of your brand; that space is large, and while many mixed-script candidates are unregistrable under registry script rules, single-script lookalikes are not. Treat a clean result as covering spelling variants, not visual ones.

Stating that limit plainly is more useful than implying coverage that is not there. A report that quietly omits a category teaches you to trust it about that category.

Common questions

what is an IDN homograph attack
It is a lookalike domain built from characters that render like the characters of a real domain but are different Unicode code points, such as a Cyrillic а in place of a Latin a. The registry stores an ASCII form beginning xn--, and the browser decides whether to show that or the Unicode rendering.
what is punycode and why does it matter for phishing
Punycode, defined in RFC 3492, is the encoding that turns a Unicode domain label into ASCII so the DNS can carry it, and the encoded label starts with xn--. It matters because a name that looks identical to a familiar one on screen has a completely different xn-- form underneath.
do browsers still show homograph domains as the real thing
Mostly not for mixed-script names, which every major browser shows in Punycode. Chromium also shows Punycode for whole-script confusables such as an all-Cyrillic label under .com, while Firefox's published algorithm permits those and relies on registries.
how do I check if a URL is a homograph
Copy it as text and look for a label beginning xn--. If one is there, look the xn-- form up in RDAP and compare the registrar and nameservers with the domain you expected. Do not open the site to compare it visually.
how do I make Firefox show punycode
Open about:config, find network.IDN_show_punycode and set it to true. Every internationalised domain name will then display in its xn-- form.

Sources and further reading

  1. Unicode Technical Standard #39: Unicode Security Mechanisms (confusables)
  2. RFC 5890: Internationalized Domain Names for Applications (IDNA): definitions
  3. RFC 3492: Punycode
  4. Chromium: IDN display policy
  5. Mozilla: IDN display algorithm
  6. Mozilla Foundation Security Advisory 2005-29: IDN homograph spoofing
  7. Mozilla bug 279099: Protect against homograph attacks (spoofing using punycode IDNs)
  8. Mozilla bug 1332714: IDN phishing using whole-script confusables
  9. Xudong Zheng: Phishing with Unicode Domains (14 April 2017)
  10. Wikipedia: IDN homograph attack
  11. ICANN: Registration Data Access Protocol (RDAP)
  12. Typosquatting.ai: methodology and data sources