TL;DR: We introduce CoBPE, a tokenizer that attaches articles, prepositions, punctuation, capitalization and leading spaces to a neighboring word, so the model reads and predicts them together in one step. CoBPE represents text in 30% fewer positions than BPE. Models trained from scratch with CoBPE at 780M and 1.3B parameters outperform BPE by 1.2 and 1.3 points on a 30-task suite under matched training compute, and train and generate faster.
The idea
BPE gives every token its own sequence position. Determiners, prepositions and punctuation take up 29.8% of the positions in English web text under GPT-4's tokenizer, and each costs a full model step, the same as a content word.
“In a hole in the ground, there lived a hobbit. Not a dirty, wet hole, …”
_ marks a leading space.CoBPE represents these markers as modifiers attached to a lexical base token. In the example above, prepositions, determiners, capitalization, leading spaces and punctuation attach to their neighboring bases, reducing the sequence from 18 positions to 9. The original text can still be reconstructed exactly.
Multi-word tokens, as in SuperBPE, also shorten sequences, but each phrase needs its own vocabulary entry. CoBPE modifiers are shared: the same the attaches to any base, and a word keeps one base token across its surface forms.
| Text | BPE | CoBPE | |
|---|---|---|---|
| the ground | the_ground | 2 | grounddet: the |
| The Ground | _The_Ground | 2 | ground_capdet: thecap: det |
| In the ground | In_the_ground | 3 | groundprep: incap: prepdet: the |
| (in the ground) | (in_the_ground) | 5 | groundprefix: (prep: indet: thesuffix: ) |
ground. CoBPE always uses one compound token with the same base.Method
CoBPE changes only how tokens enter and leave the model. At the input, the model sums the embeddings of the base and its modifiers. At the output, it predicts the base token and then, conditioned on it, each modifier. The transformer is unchanged, and the roughly 100 English modifiers add under 0.1% to the parameter count.
-
(1) TokenizeOn the table.BPE ↓On_the_table.attach modifiers ↓table prep: On det: the suffix: .
-
(2) Embedtable+On+the+.↓one input position
-
(3) Transformerunchanged
-
(4) Predictbasemodifiers, given
table→ “On the table.” in one step
Results
We pretrain 780M and 1.3B models from scratch on English, changing only the tokenizer and keeping the architecture, training recipe, 32k vocabulary size and number of training tokens fixed. CoBPE shortens sequences by 30% (SuperBPE: 15%) and improves the 30-task average by 1.2 points at 780M and 1.3 points at 1.3B, while SuperBPE matches BPE. Because the token budget is matched, CoBPE models also see more raw text.
Relative sequence length
30-task average
Shorter sequences also make CoBPE faster end to end. On the nanochat speedrun, it reaches GPT-2-level performance in 6.45 hours instead of 7.92 (1.23× faster). When generating the same text, it takes 32.5% fewer steps and is 1.28× faster, even though each step is slightly slower.
Generation time (s)
Training time to GPT-2 level (h)
The paper also reports compression across vocabulary sizes, 18–32% shorter sequences in Spanish, German, Hindi and Indonesian, and per-task results.
Code and models
The code includes the tokenizer, the model changes, and scripts to reproduce the paper's experiments. Model weights are coming soon.
git clone https://github.com/schwartz-lab-NLP/cobpe
cd cobpe
uv sync --locked --extra cpu
uv run --locked --extra cpu python -m cobpe check-install
Related
Vocab Diet (Findings of ACL 2026) composes word forms such as walked from a base word and transformation vectors (walk + past tense). CoBPE uses the same input and output parameterization for grammatical markers instead of morphology.
Citation
@inproceedings{reif2026more,
title={More Than Words: Compositional Tokenization for Efficient Language Models},
author={Yuval Reif and Guy Kaplan and Roy Schwartz},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=Yw16kddexd}
}