End of Query Detection for Voice Assistant Endpointing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing systems face inaccuracies and resource wastage due to incorrect endpointing of user utterances, particularly on slow networks where speech decoders struggle to determine the completion of queries, leading to delayed actions and unnecessary resource usage.

Innovation Solution

An end of query detector using machine learning and neural networks is employed to quickly determine whether a user has finished speaking by analyzing acoustic characteristics and confidence scores, allowing for timely deactivation of microphones and efficient resource management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a speech decoder is used to determine when a user has finished speaking, then the system can process audio data, but on slow networks the determination is delayed causing the microphone to remain open longer than necessary

Engineering Contradiction:
Improveaccuracy of endpointingVSAvoiddelay in determining utterance completion
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary endpointing analysis using acoustic characteristics and confidence scores before the speech decoder completes its full processing. This allows the system to make an early determination about whether the user has finished speaking, enabling timely microphone deactivation even while waiting for the speech decoder to process the audio data.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the microphone remains open to capture complete utterances, then no speech is missed, but computational resources and power are wasted processing unnecessary audio

Engineering Contradiction:
Improvecompleteness of utterance captureVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system continuously monitors acoustic characteristics and confidence scores during audio capture, providing real-time feedback about the likelihood that the user has finished speaking. When the confidence score indicates the user has finished, the system provides feedback to deactivate the microphone, thereby capturing complete utterances while avoiding unnecessary processing of additional audio that would waste power.

Inventive Principle:
Principle #23Feedback

3Productivity

If audio data is transmitted in larger packets on slow networks, then network bandwidth is utilized more efficiently, but the speech decoder cannot determine utterance completion in a timely fashion

Engineering Contradiction:
Improvenetwork transmission efficiencyVSAvoiddelay in processing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system segments the endpointing determination process from the full speech decoding process. While audio data is transmitted in larger packets for network efficiency, the endpointing function is segmented out and performed independently using acoustic characteristics and confidence scores. This allows timely determination of utterance completion without waiting for the full speech decoder to process the large audio packets.

Inventive Principle:
Principle #1Segmentation

4Productivity

If traditional endpointing based on pause duration is used, then the system can identify potential utterance boundaries, but it frequently segments incomplete phrases and activates unnecessary processing

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidaccuracy of utterance segmentation
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system changes the parameters used for endpointing from simple pause duration to a combination of acoustic characteristics (pitch, loudness, intonation, sharpness, articulation, roughness, instability, speech rate) and confidence scores. This parameter change enables more accurate determination of whether an utterance is complete, reducing false segmentation of incomplete phrases while maintaining efficient processing by avoiding unnecessary activation of processing components.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11551709B2End of query detection
Publication Date: 2023.01.10 GOOGLE LLC
  • US11551709B2 patent drawing
  • US11551709B2 patent drawing
  • US11551709B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for detecting an end of a query are disclosed. In one aspect, a method includes the actions of receiving audio data that corresponds to an utterance spoken by a user. The actions further include applying, to the audio data, an end of query model. The actions further include determining the confidence score that reflects a likelihood that the utterance is a complete utterance. The actions further include comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold. The actions further include determining whether the utterance is likely complete or likely incomplete. The actions further include providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.