More Than Words: Compositional Tokenization for Efficient Language Models

Yuval Reif Guy Kaplan Roy Schwartz

The Hebrew University of Jerusalem

COLM 2026

TL;DR: We introduce CoBPE, a tokenizer that attaches articles, prepositions, punctuation, capitalization and leading spaces to a neighboring word, so the model reads and predicts them together in one step. CoBPE represents text in 30% fewer positions than BPE. Models trained from scratch with CoBPE at 780M and 1.3B parameters outperform BPE by 1.2 and 1.3 points on a 30-task suite under matched training compute, and train and generate faster.

The idea

BPE gives every token its own sequence position. Determiners, prepositions and punctuation take up 29.8% of the positions in English web text under GPT-4's tokenizer, and each costs a full model step, the same as a content word.

“In a hole in the ground, there lived a hobbit. Not a dirty, wet hole, …”

BPE (18 tokens)
In_a_hole_in_the_ground,_there_lived_a_hobbit._Not_a_dirty,_wet
CoBPE (9 tokens)
holeprep: incap: prepdet: a
ground_prep: indet: thesuffix: ,
there_
lived_
hob_det: a
bitsuffix: .
not_cap
dirty_det: asuffix: ,
wet_
Modifiers: leading spacecapitalizationdeterminerprepositionpunctuation
BPE and CoBPE tokenizations of the same text. Highlighted BPE tokens are grammatical markers; _ marks a leading space.

CoBPE represents these markers as modifiers attached to a lexical base token. In the example above, prepositions, determiners, capitalization, leading spaces and punctuation attach to their neighboring bases, reducing the sequence from 18 positions to 9. The original text can still be reconstructed exactly.

Multi-word tokens, as in SuperBPE, also shorten sequences, but each phrase needs its own vocabulary entry. CoBPE modifiers are shared: the same the attaches to any base, and a word keeps one base token across its surface forms.

TextBPECoBPE
the ground
the_ground
2
grounddet: the
The Ground
_The_Ground
2
ground_capdet: thecap: det
In the ground
In_the_ground
3
groundprep: incap: prepdet: the
(in the ground)
(in_the_ground)
5
groundprefix: (prep: indet: thesuffix: )
BPE needs 2–5 tokens and several spellings of ground. CoBPE always uses one compound token with the same base.

Method

CoBPE changes only how tokens enter and leave the model. At the input, the model sums the embeddings of the base and its modifiers. At the output, it predicts the base token and then, conditioned on it, each modifier. The transformer is unchanged, and the roughly 100 English modifiers add under 0.1% to the parameter count.

  1. (1) Tokenize
    On the table.
    BPE ↓
    On_the_table.
    attach modifiers ↓
    table prep: On det: the suffix: .
  2. (2) Embed
    table
    +
    On
    +
    the
    +
    .
    ↓
    one input position
  3. (3) Transformer
    unchanged
  4. (4) Predict
    base
    tablechairfloor
    modifiers, given table
    Onnonethea.none
    → “On the table.” in one step
(1) BPE tokens are grouped into a compound token. (2) Its input embedding is the sum of the base and modifier embeddings. (3) A standard transformer processes the sequence. (4) The model predicts the next base, then each modifier group conditioned on that base.

Results

We pretrain 780M and 1.3B models from scratch on English, changing only the tokenizer and keeping the architecture, training recipe, 32k vocabulary size and number of training tokens fixed. CoBPE shortens sequences by 30% (SuperBPE: 15%) and improves the 30-task average by 1.2 points at 780M and 1.3 points at 1.3B, while SuperBPE matches BPE. Because the token budget is matched, CoBPE models also see more raw text.

Relative sequence length

BPE 1.00 SuperBPE 0.85 CoBPE 0.70

30-task average

29303132333435 780M BPE 30.4 CoBPE 31.6 SuperBPE 30.4 1.3B BPE 33.3 CoBPE 34.6

Shorter sequences also make CoBPE faster end to end. On the nanochat speedrun, it reaches GPT-2-level performance in 6.45 hours instead of 7.92 (1.23× faster). When generating the same text, it takes 32.5% fewer steps and is 1.28× faster, even though each step is slightly slower.

Generation time (s)

BPE 1.94 CoBPE 1.51

Training time to GPT-2 level (h)

BPE 7.92 CoBPE 6.45

The paper also reports compression across vocabulary sizes, 18–32% shorter sequences in Spanish, German, Hindi and Indonesian, and per-task results.

Code and models

The code includes the tokenizer, the model changes, and scripts to reproduce the paper's experiments. Model weights are coming soon.

git clone https://github.com/schwartz-lab-NLP/cobpe
cd cobpe
uv sync --locked --extra cpu
uv run --locked --extra cpu python -m cobpe check-install

Related

Vocab Diet (Findings of ACL 2026) composes word forms such as walked from a base word and transformation vectors (walk + past tense). CoBPE uses the same input and output parameterization for grammatical markers instead of morphology.

Citation

@inproceedings{reif2026more,
  title={More Than Words: Compositional Tokenization for Efficient Language Models},
  author={Yuval Reif and Guy Kaplan and Roy Schwartz},
  booktitle={Third Conference on Language Modeling},
  year={2026},
  url={https://openreview.net/forum?id=Yw16kddexd}
}