Building MemDecks
Why a translation app is not a weekend project
You can build the demo in a day, and it is not fake. A translator is one of the standard demonstrations that anybody can now ship an app in a day — alongside the habit tracker and the note app — because its core really does look like one function: a string goes in, a string comes out. I started MemDecks on 8 December 2025 as an exercise of exactly that kind. As of 4 September 2026 it is 760 commits and 233 tracked issues, 101 of them closed. This post is about the decisions that consumed them: not the bugs, which are endless and boring, but the structural choices a demo gets to skip and a product does not.
What the one-day version genuinely gets right
An afternoon with a model and a translation API produces something that works: you type bread, you get the word in your study language, an image, an audio clip, and a saved card. Add a multiple-choice game over the saved cards and it demos well.
That is not an illusion and it is not worthless. It is the correct first step, it answers the question "is this pleasant to use", and the parts an LLM wrote in minutes would genuinely have taken a week a few years ago. The starting line moved. What did not move is everything that comes after the second hard word.
A word is not a string, and this is the decision everything else hangs off
The demo stores the word the user typed. That holds until the second thing you build and never again, because the string is not the identity.
wave is a noun and a verb with two unrelated translations, so it is two cards, and a learner shown the bare word cannot know which one is being asked for. March is a month and a procession, and a lookup keyed on the letters will return the wrong one to a calendar pack that already knew it wanted the month. Identity therefore has to carry at least the lemma and the part of speech, and where two senses share both, a sense key on top of that.
The tempting shortcut is the database id, and it is worth naming as a trap. The MemDecks catalogue is SQLite and words.id is an autoincrement, which makes it environment-local: test and production hold different ids for the same word. Anything that crosses an environment carrying ids — an exported deck, a content import, a fixture — is carrying numbers that mean something else on the other side. The rule the deck platform settled on in August 2026 is blunt: portable identity is the lemma and part of speech plus a stable per-deck card key, and the numeric word id may exist only as a nullable per-environment cache. Never a foreign key, never part of a comparison.
Every pair runs through one language, and that is a decision, not a default
A translator appears to support any language to any language. Do that honestly and the content problem is quadratic: sixteen study languages is 240 ordered pairs, each with its own translations, audio, examples and failure modes — and nobody can inspect the quality of a pair between two languages they do not speak.
MemDecks removed cross-translation before its first release and made English the pivot: a card is always English on one side and the study language on the other. Sixteen pairs instead of 240, and one side of every card in a language the author can actually check.
It costs something real, and the docs state it rather than hiding it: a learner whose own language is not English gets an English hop in the middle. That is the trade. It is also the kind of decision that does not exist on day one, because on day one there is one language pair and it is yours.
One card is four independent services, not one call
The demo makes a card with one request. A usable card needs a canonical translation, a romanisation if the script is not Latin, audio, and lexical detail — part of speech, gender, senses, examples. Those are four providers with four latencies, four failure modes and four reasons to be re-run later, so they are separated by contract: a base translator, a romanisation chain, a TTS provider, and an enrichment provider that may add detail and is never allowed to silently overwrite the canonical translation.
Two things forced that separation, and neither was visible from outside:
- The cloud does not cover your language. Google's romanisation supports neither of the two languages MemDecks needed it for most, so Armenian and Greek run on a deterministic local implementation instead. Every general-purpose service has a coverage edge, and a small product tends to live exactly on it.
- Cached output needs provenance. Audio is stored, so the provider, model, voice and prompt that produced it are stored with it. Without that you cannot answer "which files need regenerating" after a voice changes — you either regenerate everything or trust stale files forever.
Grammatical gender is the miniature of the whole problem. Four study languages — Greek, Russian, Spanish and German — inflect for it, so a noun card without it teaches the noun wrong. One small field, arriving from four places depending on the word: a lexicon, the head noun of a phrase, a rule for the language, or a paid provider call for what is left. Asking a model about every noun fails on cost and on accuracy at once, because rows that entered through runtime translation carry no part of speech, and a naive loop pays a provider to think about words the catalogue already knows are verbs.
Content and progress are different things with different lifetimes
In the demo a card is one row. In a product it is two things wearing one name: the content — word, translation, picture, audio — which somebody authors, versions and republishes, and the learner's progress on it, which belongs to the learner and has to survive all of that.
Every rule that matters here follows from taking the split seriously.
- Re-importing a deck may rewrite content and may never write progress. Same rule for a followed release, an Anki re-import, an admin repair.
- Deleting a deck takes nothing away from someone who already has its cards. Membership survives, pinned to the last release that user saw.
- De-duplication needs one rule in one place. This one was learned late: the database guarded saved cards with a unique index on the user and the resolved word, while the deck importer de-duplicated on the source word and the language pair. Two keys for one invariant means the two paths disagree, and the disagreement surfaces as duplicate cards that no single component believes it created.
Release semantics belong here too. A published translation set has to reference exactly one published English revision, or you serve English from one version beside translations from another — which reads to everyone as a translation bug and is really a versioning bug.
You cannot argue about a catalogue from examples
Every report arrives as an anecdote: this word returns four rows, that word has blank senses, this example sentence is attached to the wrong translation. Anecdotes are how you find problems and a terrible way to decide what to fix, because they carry no size. So the counting became its own tool: run over the whole catalogue and count how many rows hold each structural shape, rather than arguing from the three words somebody happened to type.
| Shape | What it means for a learner |
|---|---|
| The same lemma and part of speech stored twice | Two rows a lookup has to choose between, and it may choose differently tomorrow. |
| An enrichment row with no senses at all | Only a flat translation. Nothing a card or a game can point at when the word has more than one meaning. |
| A flat translation that is not among the first sense's translations | Two answers to the question "what is this word", and different screens show different ones. |
| An example whose target token is not declared by the row | The sentence can surface under the wrong sense, teaching the wrong meaning convincingly. |
A number per shape tells you which of them is a migration and which is a rounding error. That distinction does not exist in the demo, because in the demo there are eleven words and you know all of them.
And then the half that has nothing to do with language
The lexical problems are the interesting ones. They are not the majority. A partial list of what a working version needed and a demo does not: audio generated by a worker that must not take the server down with it when a voice fails; cards that learn the path to their audio file after the worker finishes rather than staying silent forever; an import pipeline that does not quietly rewrite the English it was given; a test environment that does not share a database with production; account deletion; guest sessions that survive a refresh; per-tier limits that mean the same thing on the server as on the pricing page.
None of that is hard in the sense of being clever. All of it is mandatory, and all of it is invisible until the day it is the only thing anyone notices.
If you are about to build one anyway
You should. Five things I would decide on day one rather than month four:
- Decide what a word is before you store any. Not the string — the identity, and a portable one that survives leaving the database it was born in. Everything else keys on it, and changing it later means migrating every card, score and example that pointed at the old shape.
- Separate what you author from what the learner earns. Content gets rewritten by imports and releases; progress must never be touched by either. Two lifetimes, one row, is the bug you will find last.
- Write the invariants down and check them in bulk. "Every sense has an identity", "the flat translation agrees with the first sense". A rule you cannot count violations of is a preference, not an invariant.
- Treat every model call as a source that can be wrong, not as an answer. Ask the data you already have first; spend the call on what is genuinely unknown; record which source each field came from.
- Pick one language you actually speak. Armenian is why half of MemDecks exists in the shape it does. You cannot see these bugs in a language you cannot read.
So is the one-day version worth building
Yes, as long as you know what it proved. It proved the idea is pleasant and that the plumbing connects. It did not prove that the domain is small, and in vocabulary the domain is where all the work is — every language you add brings its own grammar, its own script, its own idea of what counts as one word.
The tools got dramatically better at the first day. They did not shorten the ninth month, and being able to start in an afternoon mostly means arriving sooner at the part that was always going to be the whole job.