Attempts at Understanding Listening, Conversation, and Health Through Audio

Speaker: Neeraj Kumar Sharma, Assistant Professor, Mehta Family School of Data Science and Artificial Intelligence, Indian Institute of Technology Guwahati

Date: 07 August 2026

Listening, conversing, and breathing are among the most natural human behaviours, yet each poses surprisingly difficult scientific questions when examined through audio. At the SPIN (Sensing, Perception, Intelligence) Lab at IIT Guwahati, Neeraj Kumar Sharma and his team revisit these problems by asking not only whether models work, but also what they actually learn and what they fail to capture. In his talk, Neeraj Kumar Sharma presented three recent directions. First, he discussed a reaction-time study that reveals how listeners integrate interaural time and level cues during sound localisation, with implications for spatial audio rendering. Second, he introduced the speaker-switch test, a control experiment designed to determine whether conversational AI representations capture genuine dyadic adaptation or merely encode speaker identity. Third, he presented his team’s ongoing effort to build a carefully controlled, gender-balanced dataset of healthy breathing and cough sounds as a foundation for respiratory acoustic analysis. Although these projects span perception, conversation, and health, they share a common philosophy: progress often comes from looking beyond aggregate performance measures and examining the finer structure of audio, behaviour, and representation.