After writing this piece, I started wondering about something: why am I (or are we) so demanding about stumbles and pauses when listening to audio podcasts, yet so much more forgiving when the same content is presented as video? 🤔
The reason is simple.
Without visuals, all of a listener’s attention is focused on the sound. The information in the sound (or, you could say, its features) gets magnified without limit, and stumbles, pauses, and the like become impossible to ignore. This is also why so many podcasters are so “picky” about how audio content is presented (I don’t mean that negatively here; I mean it neutrally).
Once visual information is added to the sound, the viewer’s attention is divided. We may start noticing details in the set, what people look like, their facial expressions and body language, the recording equipment, and changes in the visuals...
As a result, the way something feels to watch is generally different from the way it feels to listen~
P.S. Whether stumbles/pauses in audio production “must be cut” or should be “handled case by case,” as well as the common video-editing technique of “using facial expressions, body language, or shots of other people instead”... all of this has to do with the medium. Of course, it also has to do with the standards of the creator and producer~
I saw someone ask what the difference is between video podcasts and the video interviews that simply weren’t called podcasts before. That’s actually pretty interesting. If we’re comparing TV interviews vs. video podcasts, the former starts with a need for visuals and then fills in the content, while the latter starts with content and treats the visuals as optional. 🧐
Another P.S. These thoughts are also connected to some media theories. I won’t go into them here, but if you’ve read this far and are interested, you can look into McLuhan’s hot and cool media, as well as redundancy and fault tolerance in information theory~