Audio Ducking and Erasing for Voice Assistant Noise Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Background noise from nearby devices interferes with accurate speech recognition in voice-interaction modes, affecting user experience in smart home devices and voice assistant systems.
Innovation Solution
A primary computing device detects nearby audio devices and transmits audio control signals to reduce or eliminate background noise by adjusting volume levels or erasing audio streams, improving voice command recognition through audio ducking and erasing techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If nearby devices continue playing media audio streams, then users can enjoy background music and media content, but speech recognition accuracy deteriorates due to background noise interference
Solution Approach 1:
The system extracts and removes audio streams from nearby devices that are causing background noise interference. The primary computing device identifies audio streams from secondary devices and selectively removes them from the audio input processed for speech recognition, thereby eliminating the harmful background noise while preserving the user's ability to speak clearly to the device.
Solution Approach 2:
The system applies preliminary anti-action by proactively reducing or muting audio output from nearby devices before they can interfere with speech recognition. When the primary computing device detects that speech recognition is needed, it sends control signals to secondary devices to reduce their volume or pause playback in advance, preventing background noise interference before it degrades speech recognition accuracy.
2Measurement precision
If the primary computing device requests audio stream data from nearby devices for processing, then background noise can be removed, but network bandwidth and device complexity increase
Solution Approach 1:
The primary computing device acts as an intermediary that coordinates audio control between multiple devices. It receives audio stream data from nearby secondary devices, processes this data to identify and remove interfering audio streams, and then applies the cleaned audio data for speech recognition. This intermediary approach centralizes the complex audio processing logic in one device rather than requiring each device to independently analyze and process audio streams.
3Measurement precision
If audio streams from nearby devices are completely eliminated, then speech recognition accuracy improves, but user experience deteriorates due to loss of background media playback
Solution Approach 1:
The system dynamically adjusts the audio output of nearby devices based on the operational state of the primary computing device. When the primary device is in voice-interaction mode, nearby devices reduce or pause their audio playback to improve speech recognition. When the primary device is not actively listening, nearby devices resume normal playback. This dynamic adjustment allows the system to optimize speech recognition accuracy only when needed, while preserving background media playback during other times, thus maintaining overall user experience quality.
Data Source
AI summary
A smart home device (e.g., a voice assistant device) includes an audio control system that determines a set of one or more audio devices to include nearby devices that are capable of providing audio streams that are audibly detected by a microphone of the smart home device. The audio control system initiates a voice-interaction mode for operating the smart home device to receive voice commands from a user and provide audio output in response to the voice commands. The audio control system transmits an audio control signal to nearby devices that configures each nearby device to implement one or more of: reducing a volume level associated with the audio streams generated by the nearby devices while the smart home device is operating in the voice-interaction mode; and transmitting, to the smart home device, audio stream data associated with a current audio stream generated for audible output by the nearby device.


