πŸ¦– Emojisaurus.me Explore β†’

Emoji 17.0 Β· 3,944 characters Β· 827 curated senses

A thesaurus for emoji

Most emoji references are dictionaries: they tell you what a picture is called. Emojisaurus tells you what it means, and lets you search from a meaning back to the emoji that carries it.

Dictionary  word β†’ meaning
Thesaurus  meaning β†’ word
Emoji need the second
3,944fully-qualified emoji
827curated senses
17sense registers
5,225total rows indexed
0dependencies, trackers, network calls

The origin of emoji

Emoji is Japanese: η΅΅ζ–‡ε­—, from e (η΅΅, picture) and moji (ζ–‡ε­—, character). The resemblance to the English emotion and to emoticon (the older :-) tradition) is a complete coincidence. Emoji are picture-characters, and that is all the word claims.

The origin story is usually told as: Shigetaka Kurita designed 176 emoji for NTT DoCoMo's i-mode in 1999, and the Museum of Modern Art later acquired the set. That story is wrong, and Emojipedia formally corrected it in 2019. SoftBank shipped a 90-emoji set on the J-Phone SkyWalker DP-211SW in 1997 (two years earlier), and it included the ancestor of πŸ’©. Researchers have since pushed the date back further still, to picture-character sets on Japanese word processors such as the Toshiba JW 95F and Sharp PA-8500 in 1988. Kurita's set is the famous one, not the first one.

Emoji reached the rest of the world through Unicode. Version 6.0, in 2010, was the first to encode them at scale, which is why an emoji is a character in the same technical sense as A or ζΌ’. It has a code point, an encoded byte length, and a name assigned by committee. This project covers Emoji 17.0, released alongside Unicode 17.0 in September 2025: 3,944 fully-qualified emoji, and 5,225 rows once you count the partially-qualified spellings that also occur in the wild.

Why emoji are hard to look up

Every emoji has an official name, assigned by Unicode and CLDR. Those names describe the picture, faithfully and uselessly:

Three things make this worse. Emoji are polysemous: one glyph carries many unrelated senses at once. πŸ™ has seven here, spanning six registers, and none of them is wrong. Meanings are regional: πŸ‘Œ πŸ‘ 🀘 ✌️ are obscene in some countries and neutral in others, and the anglophone consensus that πŸ˜‚ is "for old people" is not shared across Brazil, Indonesia or much of Africa. And meanings drift, quickly, in ways no committee ratifies.

So what is a thesaurus?

The word is Latin thΔ“saurus, from Ancient Greek thΔ“sauros: a treasure house or storehouse. Not a list of synonyms: a store of things worth keeping, arranged so you can find them.

Peter Mark Roget (a physician and secretary of the Royal Society) began compiling one for his own use in 1805 and published Thesaurus of English Words and Phrases on 29 April 1852. Its radical move was the arrangement. A dictionary is alphabetical: you arrive knowing the word and leave knowing the meaning. Roget inverted it, organising roughly 1,000 conceptual categories under six classes: Abstract Relations, Space, Matter, Intellect, Volition, Affections. This let you could arrive knowing only the idea and leave with the words that express it. The alphabetical index most people actually use was added afterwards, as a way back in.

Dictionaryword β†’ meaningYou have the symbol. You want the sense.
Thesaurusmeaning β†’ wordsYou have the sense. You want the symbol.

Why this one is a thesaurus

It carries both layers, and they do different jobs.

The dictionary layer is mechanical and complete: the official CLDR name and keywords for all 3,944 fully-qualified emoji, plus the encoding facts (code points, UTF-8 byte size, UTF-16 units. That layer is authoritative and it is not mine; it comes straight from Unicode.

The thesaurus layer is the hand-written part, and it is what this project is actually for: 827 curated senses across 571 emoji, sorted into 17 registers. The registers are doing Roget's job (they are the conceptual classes).

internetmisreadsexualironyworkplacesubculturenonenglisheverydaysportsreligiondatingregionalpoliticskinkfinancesubstancesplatform

Each groups senses by the world they belong to, so that 🧊 landing in five different registers tells you something real about the glyph rather than looking like a contradiction.

Where the dictionary layer comes from

Every name and keyword in here is CLDR data: the Unicode Consortium's Common Locale Data Repository, and much less famous than the character standard itself.

Unicode proper says what characters exist. CLDR says how to present things to humans in a given language and place: date and time formats, number and currency formatting, sort order, pluralisation rules (Arabic has six plural forms, Japanese has one), translated names for languages, countries and timezones, measurement units, and emoji names and keywords.

You have almost certainly used it today without noticing. CLDR is the data behind ICU, which Java, Android, iOS, macOS, Windows, PHP and JavaScript's Intl API all rely on. When toLocaleDateString() gives you 27/08/2026 in Australia and 8/27/2026 in the United States, that difference is CLDR. It is XML in a format called LDML, contributed and vetted by native speakers through a public survey tool, and released a couple of times a year.

This project reads two of its files, and takes three things from them:

A qualification worth making

It would be tidy to say CLDR describes the picture and this project describes the meaning. That is too neat. CLDR keywords are search terms, and they are more sense-aware than the names suggest:

CLDR knows 😀 covers both the manga triumph it was drawn for and the anger people actually read. It even has lmao on πŸ’€.

What it does not do is adjudicate or explain. It is a flat bag of terms with no indication that anger has displaced triumph, no note that πŸ’€ became the standard laughter marker around 2020 as πŸ˜‚ aged out, and no grouping (you cannot ask CLDR for "the finance vocabulary". It is also conservative by design: 🧒 gets baseball Β· bent Β· billed Β· cap Β· dad Β· hat and nothing about lying; 🚩 gets construction Β· flag Β· golf Β· post and nothing about dealbreakers. Those gaps are what the curated layer is for.

An unexploited angle. CLDR ships these annotations in around a hundred languages. This project reads only the English file. Swapping in the German or Japanese one would localise the entire dictionary layer essentially for free. The curated senses would stay English, but the searchable names and keywords would not have to.

It runs in both directions

This is the part a plain emoji picker cannot do. Every example below opens the explorer with the query already run.

What this is not

The curated layer is deliberately partial and openly subjective. There is no clean, current, licensable dataset of emoji sense. EmojiNet is a decade stale and its API is gone, the crowd-sourced Emoji Dictionary is unlicensed and noisy, and the academic corpora cover a couple of dozen glyphs. So every sense here was written by hand, tagged with a confidence of high, medium or low, and attributed in its source file to the class of evidence it rests on. Treat it as a curated reference, not a measurement. Where nothing is recorded, the literal name really is all there is.

Explicit senses are off by default. Three registers sexual, kink and substances) sit behind a toggle. They are documented for the same reason a slang dictionary documents slang: moderation, safeguarding and plain comprehension need them.

The gate is enforced in the search index rather than the display, so no query of any shape surfaces them while it is closed.