Emoji 17.0 Β· 3,944 characters Β· 827 curated senses
Most emoji references are dictionaries: they tell you what a picture is called. Emojisaurus tells you what it means, and lets you search from a meaning back to the emoji that carries it.
Emoji is Japanese: η΅΅ζε, from e (η΅΅, picture) and moji
(ζε, character). The resemblance to the English emotion and to
emoticon (the older :-) tradition) is a complete
coincidence. Emoji are picture-characters, and that is all the word claims.
The origin story is usually told as: Shigetaka Kurita designed 176 emoji for NTT DoCoMo's i-mode in 1999, and the Museum of Modern Art later acquired the set. That story is wrong, and Emojipedia formally corrected it in 2019. SoftBank shipped a 90-emoji set on the J-Phone SkyWalker DP-211SW in 1997 (two years earlier), and it included the ancestor of π©. Researchers have since pushed the date back further still, to picture-character sets on Japanese word processors such as the Toshiba JW 95F and Sharp PA-8500 in 1988. Kurita's set is the famous one, not the first one.
Emoji reached the rest of the world through Unicode. Version 6.0, in 2010,
was the first to encode them at scale, which is why an emoji is a character
in the same technical sense as A or ζΌ’. It has a
code point, an encoded byte length, and a name assigned by committee. This
project covers Emoji 17.0, released alongside Unicode
17.0 in September 2025: 3,944 fully-qualified emoji, and
5,225 rows once you count the partially-qualified spellings that
also occur in the wild.
Every emoji has an official name, assigned by Unicode and CLDR. Those names describe the picture, faithfully and uselessly:
skull, and means laughter.billed cap, and means someone is lying.face with steam from nose, a manga convention for triumph, and reads as indignation.love hotel, which really is the official name, for a specific Japanese institution most people outside Japan have never heard of.Three things make this worse. Emoji are polysemous: one glyph carries many unrelated senses at once. π has seven here, spanning six registers, and none of them is wrong. Meanings are regional: π π π€ βοΈ are obscene in some countries and neutral in others, and the anglophone consensus that π is "for old people" is not shared across Brazil, Indonesia or much of Africa. And meanings drift, quickly, in ways no committee ratifies.
The word is Latin thΔsaurus, from Ancient Greek thΔsauros: a treasure house or storehouse. Not a list of synonyms: a store of things worth keeping, arranged so you can find them.
Peter Mark Roget (a physician and secretary of the Royal Society) began compiling one for his own use in 1805 and published Thesaurus of English Words and Phrases on 29 April 1852. Its radical move was the arrangement. A dictionary is alphabetical: you arrive knowing the word and leave knowing the meaning. Roget inverted it, organising roughly 1,000 conceptual categories under six classes: Abstract Relations, Space, Matter, Intellect, Volition, Affections. This let you could arrive knowing only the idea and leave with the words that express it. The alphabetical index most people actually use was added afterwards, as a way back in.
It carries both layers, and they do different jobs.
The dictionary layer is mechanical and complete: the official CLDR name and keywords for all 3,944 fully-qualified emoji, plus the encoding facts (code points, UTF-8 byte size, UTF-16 units. That layer is authoritative and it is not mine; it comes straight from Unicode.
The thesaurus layer is the hand-written part, and it is what this project is actually for: 827 curated senses across 571 emoji, sorted into 17 registers. The registers are doing Roget's job (they are the conceptual classes).
Each groups senses by the world they belong to, so that π§ landing in five different registers tells you something real about the glyph rather than looking like a contradiction.
Every name and keyword in here is CLDR data: the Unicode Consortium's Common Locale Data Repository, and much less famous than the character standard itself.
Unicode proper says what characters exist. CLDR says how to present things to humans in a given language and place: date and time formats, number and currency formatting, sort order, pluralisation rules (Arabic has six plural forms, Japanese has one), translated names for languages, countries and timezones, measurement units, and emoji names and keywords.
You have almost certainly used it today without noticing. CLDR is the data
behind ICU, which Java, Android, iOS, macOS, Windows, PHP and JavaScript's
Intl API all rely on. When
toLocaleDateString() gives you 27/08/2026 in
Australia and 8/27/2026 in the United States, that difference
is CLDR. It is XML in a format called LDML, contributed and vetted by
native speakers through a public survey tool, and released a couple of
times a year.
This project reads two of its files, and takes three things from them:
kw: searches: 30,771 of them across the 3,944 fully-qualified emoji, averaging 7.8 each.U+0023.It would be tidy to say CLDR describes the picture and this project describes the meaning. That is too neat. CLDR keywords are search terms, and they are more sense-aware than the names suggest:
anger Β· angry Β· face Β· feels Β· fume Β· fuming Β· furious Β· fury Β· mad Β· nose Β· steam Β· triumph Β· unhappy Β· wonbody Β· dead Β· death Β· face Β· fairy Β· fairytale Β· i'm Β· lmao Β· monster Β· skull Β· tale Β· yolo
CLDR knows π€ covers both the manga triumph
it was drawn for and the anger people actually read. It even has
lmao on π.
What it does not do is adjudicate or explain. It is a flat bag of
terms with no indication that anger has displaced triumph, no note that
π became the standard laughter marker around 2020
as π aged out, and no grouping (you cannot ask
CLDR for "the finance vocabulary". It is also conservative by design:
π§’ gets baseball Β· bent Β· billed Β· cap Β· dad
Β· hat and nothing about lying;
π© gets construction Β· flag Β· golf Β·
post and nothing about dealbreakers. Those gaps are what the
curated layer is for.
An unexploited angle. CLDR ships these annotations in around a hundred languages. This project reads only the English file. Swapping in the German or Japanese one would localise the entire dictionary layer essentially for free. The curated senses would stay English, but the searchable names and keywords would not have to.
This is the part a plain emoji picker cannot do. Every example below opens the explorer with the query already run.
alt:lyingMeaning β emoji. You know the idea, not the picture. Finds π§’.
alt:"red flag"The dealbreaker sense, not the object. Finds π©.
π«‘Emoji β meanings. Paste a glyph to decode a message someone sent you.
reg:misreadA whole register: official names nobody recognises.
bytes:>25Neither. The encoding questions a thesaurus has no opinion on, but a byte does.
v:17Everything new in Emoji 17.0.
The curated layer is deliberately partial and openly subjective. There is no clean, current, licensable dataset of emoji sense. EmojiNet is a decade stale and its API is gone, the crowd-sourced Emoji Dictionary is unlicensed and noisy, and the academic corpora cover a couple of dozen glyphs. So every sense here was written by hand, tagged with a confidence of high, medium or low, and attributed in its source file to the class of evidence it rests on. Treat it as a curated reference, not a measurement. Where nothing is recorded, the literal name really is all there is.
Explicit senses are off by default. Three registers
sexual, kink and substances) sit
behind a toggle. They are documented for the same reason a slang
dictionary documents slang: moderation, safeguarding and plain
comprehension need them.
The gate is enforced in the search index rather than the display, so no query of any shape surfaces them while it is closed.