Working Memory vs Long-Term Memory: The Difference That Decides What You Keep
Almost everything that goes wrong in self-study comes down to confusing two systems that behave nothing alike. One is tiny, fast and wiped clean within seconds. The other is vast and durable, and getting anything into it is slow work that no amount of rereading will speed up.
Read a phone number, hold it while you cross the room, dial it, forget it. That is working memory doing its job. Now think of your own phone number, or the chorus of a song you have not heard since school: no effort, no rehearsal, and it is simply there. Same brain, two systems that share almost no properties.
The distinction sounds academic until you notice that most study advice quietly assumes the wrong one. Rereading a chapter, highlighting it, listening to a lesson twice: all of that keeps material moving through the small fast system while doing very little to load it into the durable one.
Working memory: small, quick, and gone in seconds
Working memory is the workspace where you hold and manipulate whatever you are dealing with right now. George Miller's 1956 paper put its capacity at seven items give or take two, a number that escaped into popular culture and never left. Later work has been less generous: Nelson Cowan argued the realistic figure is closer to four chunks once you stop people from silently rehearsing.
Alan Baddeley and Graham Hitch reframed it in 1974 as something more structured than a holding pen, with a loop for verbal material, a separate store for visual and spatial material, and an executive that decides where attention goes. The useful takeaway for a learner is not the architecture; it is that the space is small, it is separate for words and images, and anything sitting in it decays within seconds unless you keep refreshing it.
Long-term memory: the opposite in almost every respect
Nobody has found the ceiling on long-term memory, and nobody expects to. It is durable, it survives sleep and distraction, and its failures are usually failures of retrieval rather than storage: the word is in there, you just cannot get at it right now, which is why it surfaces ten minutes after the conversation moved on.
The cost is that writing to it is slow and does not happen automatically. New material has to be encoded, then consolidated over hours and days, with sleep doing a good deal of that work. This is the part people skip, because it cannot be done faster by trying harder in the moment.
The bottleneck everything has to pass through
John Sweller built cognitive load theory on exactly this asymmetry: working memory is severely limited when handling anything new, while long-term memory is effectively unlimited. Every fact you will ever know has to squeeze through the narrow end first.
Which is why sitting down with fifty new words is a worse plan than it looks. The list does not fail because you lack discipline; it fails because you overran a four-item workspace and the surplus never got encoded at all. The case for a small daily number of new words rests on this and not on motivation.
Chunking, or how expertise widens the gate
The limit is on chunks rather than raw items, and what counts as one chunk depends on what you already know. William Chase and Herbert Simon showed this with chess in 1973: masters glanced at a position from a real game and reproduced it far better than weaker players, but when the pieces were scattered randomly the advantage collapsed. They were not holding more pieces; they were holding fewer, larger units built from years of stored patterns.
Language works the same way. A beginner reading a sentence in a new language handles it letter by letter, then word by word, and the workspace fills before the sentence ends. A fluent reader takes whole phrases as single units, which is why the same sentence costs them almost nothing. Long-term memory is what buys you room in working memory, and that is the real argument for knowing a lot of vocabulary cold rather than half-knowing three times as much.
Why cramming feels like it worked
Cram for four hours and the material is genuinely available at the end of it, sitting in working memory and freshly retrievable, which feels exactly like knowing something. The exam next morning may even confirm it. A week later most of it has gone, and the honest description is that it was never written to the durable store in the first place.
Hermann Ebbinghaus measured the decay in the 1880s and the forgetting curve has been replicated often enough since that its shape is not in dispute: steep at first, then flattening. Fluency during a session tells you almost nothing about what survives it, which is the single most expensive illusion in self-study.
The two systems side by side
The differences are stark enough that treating them as one thing guarantees wasted hours.
| Working memory | Long-term memory | |
|---|---|---|
| Capacity | Roughly four chunks | No known limit |
| Duration | Seconds without rehearsal | Years, potentially permanent |
| Getting things in | Instant, automatic | Slow, needs encoding and consolidation |
| Typical failure | Overload: the surplus vanishes | Retrieval: it is there but will not come |
| What helps | Fewer new items, less distraction | Retrieval practice, spacing, sleep |
What follows for how you study
Two practical consequences, and they point in different directions. To protect working memory, cut the number of genuinely new items per session and remove the noise around them. To build long-term memory, do the opposite of comfortable: pull material out of your head rather than putting it back in front of your eyes, which is what active recall means, and let time pass between attempts so spaced repetition can do its work.
Sleep belongs on the list too, and not as a wellness aside. Consolidation happens largely while you are asleep, so a short night after a long study session costs you part of what the session was for.
- Keep new items per session low; the ceiling is your workspace, not your willpower.
- Test yourself instead of rereading, even when rereading feels more productive.
- Space the attempts out rather than stacking them into one sitting.
- Build enough automatic vocabulary that sentences cost you fewer chunks to process.
Putting it to work
None of this is new science. Miller is seventy years old, Ebbinghaus older than that, and the practical advice has barely changed in decades. What has changed is that keeping track of when each item should come back is no longer something you have to do by hand, which used to be the reason people abandoned the method that works.
MindDory is built around the second system rather than the first. You capture words as you meet them, each session is assembled from the ones closest to slipping instead of whatever you happened to add last, and the sessions stay short enough that you are not trying to force twenty new items through a four-item gate. The card-level details are in how to use flashcards effectively, and how many words per day covers the intake side.
The short version: stop judging a study session by how much you understood while it was happening. That is working memory reporting on itself, and it is a flatterer.