Sillage

A frozen language model that remembers what it reads — 4.2 MB, no gradients, no fine-tuning, no vector database.

This Space runs GPT-2 124M on a free CPU. Its prose is weak in absolute terms — that is not what is on display. What is on display is the difference between the two columns, and it comes from a 4.2 MB matrix written while reading, with no gradient anywhere.

The memory loaded here has read Sillage (paper 1), 8 969 tokens — the paper that describes this very mechanism. GPT-2 has never seen that text. Complete a sentence from it and watch the right-hand column recall what the left one cannot know.

Try one of these