From Sound Waves to Spectrograms: The Physics & Signal Processing of Speech AI

When a clinician speaks during a telehealth consultation, the microphone diaphragm converts vibrating air molecules into a continuous electrical voltage signal. Yet neural networks (like Whisper or Azure Speech Services) do not operate on raw 1D audio waveforms. Instead, they “look” at audio as 2D images called Mel Spectrograms. How does physical sound transform into a frequency visual? Here is the journey through acoustic physics and signal processing. 1. The Physics of Sound: Pressure Perturbations Sound is a mechanical longitudinal wave. When vocal cords vibrate, they compress and rarify surrounding air molecules, producing periodic fluctuations in atmospheric pressure: ...

September 12, 2026 · 3 min · Siva Madhavan