Wake Phrase Detection in Muted Audio Sessions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems in communication devices are unable to provide natural language processing services when the sound sensor is muted during a communication session, limiting user interaction and functionality.
Innovation Solution
The system detects and processes spoken commands by refraining from transmitting audio data during a muted sound sensor condition and monitoring for a wake phrase, initiating natural language processing when the phrase is detected, either locally or through a remote service.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If the sound sensor is muted during a communication session, then the audio data transmission is prevented and privacy is maintained, but the speech recognition service becomes unavailable
Solution Approach 1:
The patent extracts the wake phrase detection function from the main audio transmission path. When the sound sensor is muted, the system separates the audio processing into two paths: one for communication (blocked) and one for wake phrase detection (active). This allows the device to detect wake phrases locally without transmitting audio data during muted states, resolving the contradiction between maintaining privacy and enabling speech recognition.
Solution Approach 2:
The system performs preliminary action by continuously monitoring for wake phrases even when the sound sensor is muted for communication. The wake phrase detection is prepared in advance and operates independently, so when a wake phrase is detected, the system can immediately transition from muted state to active speech recognition mode without delay.
2Adaptability or versatility
If the sound sensor remains unmuted to enable speech recognition, then natural language processing is available, but audio data is transmitted during muted periods compromising session privacy
Solution Approach 1:
The patent introduces an intermediary mechanism - the wake phrase detection system - that acts as a mediator between the muted sound sensor and the speech recognition service. This intermediary continuously monitors audio input locally without transmitting data, and only triggers audio transmission when a wake phrase is detected, thus enabling speech recognition while preventing unwanted audio transmission during muted periods.
Solution Approach 2:
The system dynamically adjusts the audio transmission state based on wake phrase detection. Rather than maintaining a static unmuted state for speech recognition, the system transitions from muted to unmuted only when necessary - specifically when a wake phrase is detected. This dynamic state change allows speech recognition service availability while minimizing unwanted audio transmission.
3Ease of operation
If wake phrase detection is implemented during muted state, then speech commands can be issued without disturbing others, but the system complexity increases
Solution Approach 1:
The patent segments the audio processing architecture into distinct functional modules: a wake phrase detection module that operates during muted states, and a full speech recognition module that activates upon wake phrase detection. This segmentation allows wake phrase detection to be implemented with minimal additional complexity, as it uses a simplified processing path that only activates when needed, rather than requiring the entire speech recognition system to run continuously.
Data Source
AI summary
A method includes generating first audio data based on sound detected by a sound sensor at a first time. The method further includes transmitting, via one or more communication interfaces, the first audio data to another device during a communication session based on a determination that the sound sensor is unmuted with respect to the communication session at the first time. The method further includes generating second audio data based on sound detected by the sound sensor at a second time. The method further includes refraining from transmitting the second audio data to the other device during the communication session based on a determination that the sound sensor is muted with respect to the communication session at the second time. The method further includes initiating a natural language processing operation on the second audio data based on detecting a wake phrase in the second audio data.


