All guides

How to shadow with podcasts (and which episodes actually work)

· 8 min read

Podcasts are the largest free supply of natural speech that has ever existed, in almost every language, on almost every subject. They are also, episode for episode, mostly bad shadowing material — and the reasons why are specific enough to work around.

What makes an episode usable

One speaker, or two who take turns

Crosstalk is the killer. Two friends talking over each other is delightful listening and impossible to shadow: there is no single line of speech to trail. Solo-narrated shows and formally structured interviews work; panel shows and comedy banter do not.

Scripted or semi-scripted delivery

Read narration has consistent pace, clean sentence boundaries and few false starts. Unscripted conversation is full of restarts, filler and abandoned clauses — genuinely useful to hear, but frustrating to reproduce, because you end up drilling someone’s stumble.

A good compromise is a show that is planned but spoken: the host knows what they are going to say and says it naturally.

Recorded in a studio

Background music under the speech, phone-quality guest audio and room echo all make phrases harder to hear precisely and harder to split automatically, because the pauses stop being quiet. If a show has a music bed running under the whole episode, skip it.

Slightly above your comfort

You should follow the meaning without a transcript while being unable to produce the sentences yourself. Shows made for learners are the reliable starting point; native-audience shows are the goal.

Quick test: play ninety seconds. If you can summarise it but could not say any of it, that is your episode.

Learner shows versus native shows

Podcasts made for language learners have real advantages for this exercise. They are slower, cleanly recorded, usually single-voiced, and very often come with a transcript — which matters, because the first pass of a shadowing session is much easier with the words in front of you.

Their disadvantage is that the speech is not quite real. Deliberate, over-articulated delivery has none of the compression and reduction that makes native speech hard, so a learner who only ever shadows learner content stays fluent at a speed nobody actually talks at.

The practical answer is to use both: learner material to build the habit and native material to keep raising the ceiling. News bulletins are a useful middle rung — professionally read, clearly recorded, naturally paced, and free everywhere.

Getting the audio

Shadowing needs a file, not a stream, because you have to cut it. Most podcasts publish a plain MP3 — the download button in a podcast app, or the enclosure link in the show’s RSS feed, will get you one.

Use shows that offer downloads and respect what the publisher allows. Practising with a downloaded episode privately is ordinary personal use; redistributing the audio, or the clips you cut from it, is not.

Shadowly never uploads any of this. The file is decoded in your browser and stays on your device, which means the copyright question stops at your own machine — nothing is copied to a server.

Cutting an episode down

A forty-minute episode is not a session; it is a source of about twenty sessions. Find one stretch of sixty to ninety seconds that is worth drilling and work only on that.

Pick the stretch for its language, not its content. The most interesting ninety seconds is often the worst material — that is where people interrupt each other and get excited. Look for a passage where one person explains something at an even pace.

From there the workflow is the same as any other file:

  1. Load the episode into the editor and zoom the waveform in on the stretch you chose.
  2. Drop boundaries in the silences between phrases, or let auto-split place them and nudge the few it gets wrong. Two to six seconds per phrase is the target.
  3. Delete nothing — a phrase is just the span between two boundaries, so you simply practise the range you care about and leave the rest.
  4. Set repeats and a gap, then work through the phrases as described in the step-by-step method.

Why the pauses in a podcast are so convenient

Podcast speech is breath-driven, and speakers breathe at grammatical boundaries. That means the silences in the waveform almost always fall exactly where a phrase ends — which is why automatic splitting works far better on a podcast than on, say, a song or a lecture recorded in a hall.

It is also why relative silence detection matters. A fixed volume threshold splits a loudly mastered show into fragments and a quiet one not at all. Judging silence against the recording’s own speech level and noise floor is what lets the same two settings work across wildly different shows.

Practising on the move

Podcasts are commute material by nature, and shadowing survives the commute better than most study. Once a clip is split and configured, you can export the entire session as a single MP3 — every phrase, with its repetitions and gaps baked in — and put it on your phone next to the original episode.

Walking while you shadow is not incidental, either. It was part of the method as originally taught, and it works: it is very hard to slide into passive listening while you are moving and speaking at volume.

A realistic first week

  • Day 1. Find an episode, cut ninety seconds, split it into phrases. Listen once. Shadow the first three phrases.
  • Days 2–3. Same clip. All phrases, five repetitions each. Record two of them and compare.
  • Day 4. Same clip, no transcript at all.
  • Day 5. Full clip end to end, twice, at full speed.
  • Days 6–7. New clip from the same show — the voice is familiar, so only the content is new.

Five days on one clip feels excessive and is where the results come from. Open the editor and load an episode to start.

Keep reading

Try it on your own audio

Split a recording into phrases, loop each one, and record yourself in the gap. Free, no account, and your audio never leaves your device.

Open the editor