Guide · Updated August 2026
Audio spaced repetition, explained.
Audio spaced repetition is a full SRS review loop — prompt, recall, grading, rescheduling — carried out entirely in sound. The app speaks the question, you say the answer, it registers how you did and decides when that item comes back. It keeps everything that makes Anki work while removing the screen, which makes real studying possible while driving, walking, or doing chores. Kotoba implements it for Japanese vocabulary, free.
Why spaced repetition normally needs a screen
Spaced repetition has two moving parts: active recall (you're asked, you retrieve) and per-item scheduling (each card returns just before you'd forget it, based on your history with that card). Classic SRS apps deliver both through a screen — read the card, flip it, tap a grade. The grade is the crucial data point: without knowing whether you got the card right, the scheduler is blind. That single tap is why "just listen to your deck" never worked as studying — export your cards to audio and you keep the content but lose the recall and the scheduling. A playlist plays word #47 on Tuesday whether you've known it for a year or never learned it.
Moving the loop into sound
Every screen interaction has an audio equivalent:
Recall — a pause. You retrieve the word and say it: 「りんご」
Grade — speech recognition checks what you said; when it can't be sure, you self-grade by voice ("got it" / "missed it")
Schedule — the result feeds the SRS algorithm; the word's next appearance moves out (or snaps back)
That grading step is what separates audio spaced repetition from audio playback. Once the system hears your answer, the scheduler has the same information a screen tap would give it — so the intervals stay honest, reviews stay short, and hard words get the extra reps automatically. Kotoba runs this loop with FSRS, the same modern scheduling algorithm recent versions of Anki use.
Does it work as well as visual flashcards?
For spoken vocabulary, often better. A text flashcard trains you to produce a word when you see English text; an audio flashcard trains you to produce it in real time from sound, with your mouth — which is the skill conversation actually demands. Every rep doubles as listening and pronunciation practice, and hearing each word inside a native example sentence builds comprehension alongside recall. The honest limits: anything inherently visual — kanji recognition, spelling — still needs your eyes, and speech recognition occasionally mishears, which is why a voice self-grade fallback matters. Ears for the vocabulary, eyes for the writing system, is the pairing that works.
Tools that implement it
Disclosure: this page is by the developer of Kotoba.
| Tool | Recall step | Grading | Scheduling | Cost |
|---|---|---|---|---|
| Kotoba (web) | Spoken answer | Speech recognition + voice self-grade | FSRS per word | Free |
| Danki (iOS) | Recall pause | Voice commands ("good"/"again") | FSRS | $9.99 once |
| Audio Flash | Recall pause | Honor system (none registered) | Playback-level, not per-recall | Freemium |
| Anki auto-advance | Recall pause | Auto-grades every card the same | Degrades — scheduler gets no real data | Free |
The pattern: the further right you go on grading fidelity, the more the scheduling actually means something. Auto-advancing Anki keeps the audio but feeds the scheduler fiction; a timed pause without grading is closer to a playlist with gaps. If your deck is Japanese vocabulary, the full loop already exists — see Japanese audio flashcards for the hands-on comparison, or Anki while driving if you're coming from an existing deck.
Free · JLPT N5–N3 · The first session takes two minutes to try
Frequently asked questions
Is listening to vocabulary audio on repeat spaced repetition?
No — repetition alone isn't spaced repetition. Without a recall attempt and a registered result, there's nothing to space. That's recognition practice, useful but much weaker.
What algorithm does Kotoba use?
FSRS — the open spaced-repetition scheduler that modern Anki versions ship. Each word's interval adapts to your actual recall history with it.
Can audio SRS handle sentences, not just words?
Yes — as words mature in Kotoba, prompts graduate from single words to cloze sentences: you hear the native example sentence with a beep where the word belongs and supply it. Same scheduling, higher difficulty rung.
What about languages other than Japanese?
The concept is language-agnostic, but Kotoba currently implements it for Japanese only. For other languages the closest options are passive audio courses, which drop the grading half of the loop.