#13. Video-calls in public are mathematically rude
The physics that makes video-calls fundamentally annoying
Back in 2010, during the Apple’s reveal of iPhone 4, Steve Jobs proudly presented FaceTime as the new way of talking to each other, which seemed to come straight from the future. Yet more than 15 years later, video calls are not exactly as common as he must have expected, even though they are much more common than I wish they were.

The problem with video calls is that they create a bad user experience for pretty much everyone within a 3-5 meter radius, unless the perfect conditions are met – people on both ends of the call are in quiet private spaces with no people around them. Otherwise, someone for sure will be disturbed, even if they don’t explicitly show it.
There is a straightforward mathematical mechanism that makes video calls more annoying than other types of communication, such as texting or audio calls.
In the following, I’ll explain exactly how it works.
Building blocks of a Video call
A video call relies on the following four components:
a camera 📷 – to record your image;
a screen 📱 – to show you the person you’re talking to;
a microphone 🎤 – to record your voice;
a speaker 🔈 – to hear the voice of the other person.
To make the illustration more simple, let’s consider just a bare phone, without any earphones, since that’s how the most annoying video calls are made. That’s what I observe regularly in regional trains and public transport in Italy, Switzerland and the UK. Things are a little different with earphones, but just a little. I might discuss those aspects in another post in the future.
Video vs Audio
There is a fundamental difference between the video and audio components of a video call, and this difference is what defines the level of discomfort that a video call creates.
▶ Video is transmitted by photons of light that travel in straight lines and get absorbed as soon as they hit any surface. So when you’re watching a video, it can hardly disturb anyone else, unless you’re watching it at extreme brightness in a dark place. And even in that case they can simply turn their head away and isolate themselves from the light if it disturbs them.
▶ Audio, on the other hand, travels in all directions, like concentric circles around you, and gets reflected from hard surfaces, bouncing back and forth many times, traveling at much longer distances. Therefore, audio is able to disturb way more people around you than video. And unlike photons, sound waves have long enough wavelength to propagate through materials, so it’s almost impossible to completely isolate yourself from unwanted sounds, even when you close your ears. The best you can do is reduce its intensity by moving further away from it or putting something soundproof into your ears. Both require much more effort than simply turning your head away from someone’s screen.
This is why being mindful of the acoustic noise that you create is so important. Unless you actually want to be rude to the people around, of course.
The inverse-square law
The thing that makes video calls fundamentally annoying is the way sound behaves. Like I said earlier, it propagates as concentric waves in all directions. Since we live in 3D space, all directions means that the wavefront of the sound emitted by your phone has a shape of a sphere that is expanding from its centre.
According to the law of conservation of energy, as that sphere is expanding, the original energy of the soundwave gets stretched over its surface, making the amplitude of the sound smaller as it expands. This is why sound gets quieter as you move away from the source. But the crucial detail is that this decrease of loudness is not linear. It’s quadratic, speaking in mathematical terms. And here is why…
If you studied geometry at school, you might remember the formula for the surface area of a sphere:
Its most relevant aspect is that the surface area (A) is proportional to the distance from the centre (r) squared. As the acoustic energy is stretched over that sphere, the perceived loudness of the sound is proportional to the inverse of that surface area, hence the inverse-square law1:

What it means in practice is that as you move away from the source of the sound, its loudness will decrease as the square of the distance from it. If you double the distance from your phone, the sound will become 4 times quieter (=2^2). If you move 10 times further away, it will sound 100 times quieter.
💡 In reality it’s a simplification, which works well for small distances and infinitely-large open spaces. If you consider sound propagation at large distances in a room or a train, the sound will bounce off the walls and other hard surfaces, fading out slower than you’d expect from the inverse-square law alone.
It’s all about distance
The inverse-square law by itself is actually great, because it provides a very simple way to improve the Signal-to-Noise ratio (S/N) – the typical way of describing how loud the signal is compared to the noise. For example, if you’re sitting between two people talking at the same volume, simply moving twice closer to one of them will make their voice sound ⨉4 times louder, while the sound of the other person will become ⨉2.25 times quieter, leading to ⨉9 (=4 x 2.25) times better S/N ratio.

S/N = 1 means that signal and noise are equally loud;
S/N = 4 means that signal is x4 times louder than the noise;
S/N = 0.25 = 1/4 means that signal is x4 times quieter than the noise;
This is what makes regular phone calls so clear and effortless, even when you’re in a noisy environment. By keeping the microphone close to your mouth and the speaker close to your ear, the signal becomes tens or hundreds of times closer to you than the noise. And when you square it, you get an enormous signal-to-noise ratio that allows people on both ends of the call to hear each other perfectly fine, even when whispering, without disturbing anyone else.
The problems start as soon as you move away from the signal, which is exactly what happens in a video call. Given that the phone camera needs to capture your face, it must be at a distance of about 30-40 cm. This dramatically reduces the signal-to-noise ratio, forcing you to be much louder than necessary.
Audio-call vs Video-call
Let’s compare an audio call and a video call from the microphone perspective, where your mouth is the signal, and a talking person sitting 1 metre away is the noise. In an audio call your mouth is roughly 5 cm away from the microphone, which gives a distance ratio between the signal and the noise of ⨉20 (=100cm/5cm), translating to a loudness ratio of ⨉400 (=20^2).

That ⨉400 signal-to-noise ratio is exactly why people can hear you so well when you’re holding your phone next to your mouth, even when you’re in a noisy pub or on a train station. But when you switch to a video call, the microphone moves away from your mouth, let’s say to 30 cm. That brings the distance ratio down to 3.3 (=100/33), and the signal-to-noise ratio becomes only ⨉10 (=3.3^2). Not 400!

Let me repeat this again. When you and another person 1 meter away are talking at the same volume, your voice will appear 400 times louder in an audio call, but only 10 times louder in a video call. That’s a massive difference!
And how do most people work around that? They raise their voice. But even if you start speaking twice louder, the signal-to-noise ratio will only become ⨉20 instead of ⨉10, which is still nowhere close to the original ⨉400 that you would get in an audio call.
So you make a mild improvement for the person on the other end of the line, but you are now speaking twice louder than everyone else, becoming an annoying source of noise yourself. Eventually, this forces other people in the room to talk louder as well, simply for staying above the noise that you’ve created. Ultimately, this chain reaction can turn the whole space into a mess, with everyone shouting, all started by a single video call.
Changing direction
Exactly the same principle applies to the sound of your interlocutor that you’re trying to hear. In this case, your ear plays the role of the microphone, the sound coming from your phone is the signal, and the person 1 metre away is the noise. Like before, the distance between the signal and your ear is super small, less than 1 cm, making the signal-to-noise ratio huge – around ⨉10000 (=100^2). In fact, that tiny speaker above your phone’s display is called “earpiece“, exactly because it’s meant to be right next to your ear.
When you switch to a video call, that distance goes from 1 cm to about 30 cm, at which point there is not that much difference any more between the signal and the noise. The perceived loudness ratio drops from ⨉10000 to just ⨉10, and the only way for you to still hear the signal is to increase the volume again. At that point, your phone switches from the earpiece to the main speaker, which is more powerful.
Besides, the increased volume in your phone creates more distortions. Even though the main speaker is slightly bigger than the earpiece, it’s still quite limited in power. It can’t reproduce the full acoustic spectrum of someone’s voice in good quality. When you push the volume to the maximum, it starts sounding dry and scratchy, adding even more to the irritation of the people around.
The Voice Notes
Another mistake I notice people making quite often is when they record Voice Notes.
From the engineering perspective, a voice note should provide exactly the same sound quality and signal-to-noise ratio as an audio call, because it only requires sound and no video. But from the human-behaviour perspective, that usually doesn’t work because people choose to hold their phone differently. I’ve seen many variations, but the most common one looks like in the photo below – a phone slightly tilted, with the top of the phone roughly levelled with the mouth, at a distance of 10-15 cm.
Judging from the distance alone, we can tell that its signal-to-noise ratio must be somewhere between a voice call and a video call. Yet a slight adjustment in how you’re holding your phone can bring it much closer to the audio-call quality. And to understand it, we have to look at the anatomy of a smartphone. I’ll use my iPhone 13 mini as an example.
At the top it has the earpiece, and at the bottom it has the main speaker and the microphone. As you can imagine, the microphone will be right next to your mouth during a phone call, if you’re holding it in your right hand. If you’re left-handed though, it will be the speaker that is the closest, while the microphone will be a few centimeters away from your mouth, unless you rotate it a bit.
But the more important part is the second microphone. It is placed on the opposite side of the phone to receive more sound coming towards you, for example when recording a video on the main camera. When recording a selfie video or just an audio, like you do in a Voice Note, that microphone is used for noise-cancellation. Being exposed more to the background noise than the main microphone, a dedicated algorithm uses the two audio streams to suppress the part of the sound that is common between the two, which normally represents the background noise.

And again, it is designed with the same inverse-square law in mind. Having the main microphone at the bottom, it is normally a bit closer to your mouth than the second one at the top, which improves the S/N ratio, letting the algorithm better separate the signal from the noise. But when you talk into the top of your phone, your voice is treated by the algorithm more like noise rather than a signal, because both microphones will hear your voice at a similar intensity. Therefore, noise cancellation will not work as well as if you talked right into the main microphone.
All you need to fix that is to slightly tilt your phone horizontally, so that the 2nd microphone stays significantly further from your mouth than the main one. This way the inverse-square law will work for you rather than against you.
Epilogue
I’ve tried my best to explain the simple consequence of the inverse-square law – changing distance between the signal and the noise has greater effect than changing the volume of your voice. And more often than not, it doesn’t mean that you have to physically get up and walk somewhere. Simply moving the phone closer to your face is enough to improve experience for everyone.
I believe that if more people do this, our public spaces will become quieter, more peaceful, and more inclusive for people who are sensitive to sounds.








Well, this was an interesting read. The diagrams made the physics much easier to visualise, which I appreciate. I'm not sure I'd agree with every conclusion, but it definitely got me thinking about how some mobile phone features are great in private but become everyone else's unwanted business in public. Loudspeaker calls and people blasting music... top of my list! 😂