End of Query Detection for Voice Assistant Endpointing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Natural language processing systems face inaccuracies and resource wastage due to incorrect endpointing of user utterances, particularly on slow networks where speech decoders struggle to determine the completion of queries, leading to delayed actions and unnecessary resource usage.
Innovation Solution
An end of query detector using machine learning and neural networks is employed to quickly determine whether a user has finished speaking by analyzing acoustic characteristics and confidence scores, allowing for timely deactivation of microphones and efficient resource management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a speech decoder is used to determine when a user has finished speaking, then the system can process audio data, but on slow networks the determination is delayed causing the microphone to remain open longer than necessary
Solution Approach 1:
The system performs preliminary endpointing analysis using acoustic characteristics and confidence scores before the speech decoder completes its full processing. This allows the system to make an early determination about whether the user has finished speaking, enabling timely microphone deactivation even while waiting for the speech decoder to process the audio data.
2Reliability
If the microphone remains open to capture complete utterances, then no speech is missed, but computational resources and power are wasted processing unnecessary audio
Solution Approach 1:
The system continuously monitors acoustic characteristics and confidence scores during audio capture, providing real-time feedback about the likelihood that the user has finished speaking. When the confidence score indicates the user has finished, the system provides feedback to deactivate the microphone, thereby capturing complete utterances while avoiding unnecessary processing of additional audio that would waste power.
3Productivity
If audio data is transmitted in larger packets on slow networks, then network bandwidth is utilized more efficiently, but the speech decoder cannot determine utterance completion in a timely fashion
Solution Approach 1:
The system segments the endpointing determination process from the full speech decoding process. While audio data is transmitted in larger packets for network efficiency, the endpointing function is segmented out and performed independently using acoustic characteristics and confidence scores. This allows timely determination of utterance completion without waiting for the full speech decoder to process the large audio packets.
4Productivity
If traditional endpointing based on pause duration is used, then the system can identify potential utterance boundaries, but it frequently segments incomplete phrases and activates unnecessary processing
Solution Approach 1:
The system changes the parameters used for endpointing from simple pause duration to a combination of acoustic characteristics (pitch, loudness, intonation, sharpness, articulation, roughness, instability, speech rate) and confidence scores. This parameter change enables more accurate determination of whether an utterance is complete, reducing false segmentation of incomplete phrases while maintaining efficient processing by avoiding unnecessary activation of processing components.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for detecting an end of a query are disclosed. In one aspect, a method includes the actions of receiving audio data that corresponds to an utterance spoken by a user. The actions further include applying, to the audio data, an end of query model. The actions further include determining the confidence score that reflects a likelihood that the utterance is a complete utterance. The actions further include comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold. The actions further include determining whether the utterance is likely complete or likely incomplete. The actions further include providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.


