Audio courses are audio. Textbooks are text. Almost every language product picks one channel and commits to it.
That split is a production artefact, not a teaching decision. Recording audio was one job and typesetting a book was another, and getting them to line up — the words on screen and the words in your ear locked together, across a whole library — was work nobody could justify. It is cheap now. Which is why the version of this that should always have existed is only being built at this point.
Both choices are wrong, and the reason is one of the better-established findings in the psychology of learning.
Two Channels, Not One
Your brain processes verbal and visual information through separate systems that run simultaneously without competing for the same capacity.
Which means material encoded through both channels at once isn't stored twice as strongly. It's stored twice — two independent routes to the same thing, so a failure of one doesn't lose it.
Audio alone gives you one route. Text alone gives you one route. Both together give you two, at almost no additional cost in time or effort, because the channels were never competing.
What Audio Alone Loses
There's a specific problem with audio-only in a language you don't yet speak. You can't hear where the words end.
Fluent speech has no gaps in it. Native speakers don't pause between words — the gaps you perceive in your own language are supplied by your brain, which already knows where the boundaries are. In a new language they simply aren't there.
So you hear one long continuous sound, learn it as one long continuous sound, and can reproduce the noise without knowing which part means what. That's the state most people are in with a foreign song they can sing.
Seeing the line while you hear it puts the boundaries in. Suddenly it's four things instead of one noise.
What Text Alone Loses
The opposite failure, and it's the worse one.
Read a phrase and your brain supplies a pronunciation. It will be wrong — assembled from the sound system of your own language, because that's the only one your reading has ever mapped to. Then you rehearse the wrong version until it's automatic.
Learners who read heavily before listening spend years unpicking pronunciations they invented in silence.
Why the Lyric Video Specifically
Text and audio arriving separately doesn't do it. A transcript you read afterwards is a different event from the sound.
They have to be simultaneous and synchronised — the line on screen as the line is sung, the boundary visible at the moment the sound arrives. That's when both channels encode the same item together, and it's the whole reason for a lyric video rather than a track with a PDF.
Then Take the Text Away
Once the boundaries are in and the pronunciation is right, the text has done its job and becomes a crutch.
Which is why the video comes first and the audio-only relisten comes after. Learn it with both channels. Reactivate it with one, in the car, with your eyes on the road.
About Outputly
Every Outputly song is learned first as a lyric video — the line on screen exactly as it's sung — then reactivated in the app with audio alone. Both channels to encode it, one channel to keep it.
