Guide · Updated August 2026

Audio spaced repetition, explained.

Audio spaced repetition is a full SRS review loop — prompt, recall, grading, rescheduling — carried out entirely in sound. The app speaks the question, you say the answer, it registers how you did and decides when that item comes back. It keeps everything that makes Anki work while removing the screen, which makes real studying possible while driving, walking, or doing chores. Kotoba implements it for Japanese vocabulary, free.

Why spaced repetition normally needs a screen

Spaced repetition has two moving parts: active recall (you're asked, you retrieve) and per-item scheduling (each card returns just before you'd forget it, based on your history with that card). Classic SRS apps deliver both through a screen — read the card, flip it, tap a grade. The grade is the crucial data point: without knowing whether you got the card right, the scheduler is blind. That single tap is why "just listen to your deck" never worked as studying — export your cards to audio and you keep the content but lose the recall and the scheduling. A playlist plays word #47 on Tuesday whether you've known it for a year or never learned it.

Moving the loop into sound

Every screen interaction has an audio equivalent:

Prompt — the app speaks the cue: "apple … in Japanese?"
Recall — a pause. You retrieve the word and say it: 「りんご」
Grade — speech recognition checks what you said; when it can't be sure, you self-grade by voice ("got it" / "missed it")
Schedule — the result feeds the SRS algorithm; the word's next appearance moves out (or snaps back)

That grading step is what separates audio spaced repetition from audio playback. Once the system hears your answer, the scheduler has the same information a screen tap would give it — so the intervals stay honest, reviews stay short, and hard words get the extra reps automatically. Kotoba runs this loop with FSRS, the same modern scheduling algorithm recent versions of Anki use.

Does it work as well as visual flashcards?

For spoken vocabulary, often better. A text flashcard trains you to produce a word when you see English text; an audio flashcard trains you to produce it in real time from sound, with your mouth — which is the skill conversation actually demands. Every rep doubles as listening and pronunciation practice, and hearing each word inside a native example sentence builds comprehension alongside recall. The honest limits: anything inherently visual — kanji recognition, spelling — still needs your eyes, and speech recognition occasionally mishears, which is why a voice self-grade fallback matters. Ears for the vocabulary, eyes for the writing system, is the pairing that works.

Tools that implement it

Disclosure: this page is by the developer of Kotoba.

ToolRecall stepGradingSchedulingCost
Kotoba (web)Spoken answerSpeech recognition + voice self-gradeFSRS per wordFree
Danki (iOS)Recall pauseVoice commands ("good"/"again")FSRS$9.99 once
Audio FlashRecall pauseHonor system (none registered)Playback-level, not per-recallFreemium
Anki auto-advanceRecall pauseAuto-grades every card the sameDegrades — scheduler gets no real dataFree

The pattern: the further right you go on grading fidelity, the more the scheduling actually means something. Auto-advancing Anki keeps the audio but feeds the scheduler fiction; a timed pause without grading is closer to a playlist with gaps. If your deck is Japanese vocabulary, the full loop already exists — see Japanese audio flashcards for the hands-on comparison, or Anki while driving if you're coming from an existing deck.

Hear the loop yourself →

Free · JLPT N5–N3 · The first session takes two minutes to try

Frequently asked questions

Is listening to vocabulary audio on repeat spaced repetition?

No — repetition alone isn't spaced repetition. Without a recall attempt and a registered result, there's nothing to space. That's recognition practice, useful but much weaker.

What algorithm does Kotoba use?

FSRS — the open spaced-repetition scheduler that modern Anki versions ship. Each word's interval adapts to your actual recall history with it.

Can audio SRS handle sentences, not just words?

Yes — as words mature in Kotoba, prompts graduate from single words to cloze sentences: you hear the native example sentence with a beep where the word belongs and supply it. Same scheduling, higher difficulty rung.

What about languages other than Japanese?

The concept is language-agnostic, but Kotoba currently implements it for Japanese only. For other languages the closest options are passive audio courses, which drop the grading half of the loop.