Tokenizer Playground

See exactly how text splits into tokens.

148 characters61 tokensvocabulary size 592

Tokens

Tokenizers␢split␢text␢into␢pieces␢before␢a␢model␢ever␢sees␢it.␢Try␢pasting␢your␢own␢prompt,␢code,␢or␢a␢sentence␢with␢contractions␢like␢don't␢or␢I'm.

This is an illustrative byte-level BPE tokenizer — the same family of algorithm GPT-2 and its descendants use (a regex pretokenizer plus a trained list of byte-pair merges). It was trained on a small local sample corpus, not any specific production model's full vocabulary, so exact token boundaries and ids won't match GPT, Claude, or Llama's real tokenizers — but the underlying merge behavior (common affixes and words collapsing into single tokens, unusual text staying fragmented) is real. ␢ marks a token that begins with a space.

Tokenized entirely in your browser as you type — nothing is sent anywhere.

What's next?