Endpoint Detection via Speech Encoder Parameters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional endpoint detection techniques for automatic speech recognition systems are computationally intensive and unsuitable for mobile devices with limited resources, leading to inefficient power consumption and delayed user responses, as they require analyzing the entire audio signal to detect speech endpoints.
Innovation Solution
Performing endpoint detection on a mobile device by analyzing information available from a speech encoder, such as internal state information or encoding parameters, without decoding the audio signal, to estimate speech endpoints and reduce computational intensity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional endpoint detection techniques are used that analyze the entire audio signal, then speech endpoint detection accuracy is improved, but computational intensity and power consumption increase significantly
Solution Approach 1:
The patent extracts only the necessary encoding parameters from the speech encoder that are sufficient for endpoint detection, rather than analyzing the entire decoded audio signal. This extraction approach maintains adequate detection accuracy while significantly reducing computational intensity and power consumption on mobile devices.
Solution Approach 2:
Instead of the conventional approach of decoding the audio signal and then performing endpoint detection, the patent inverts the process by performing endpoint detection directly on the encoded parameters before decoding occurs. This reversal eliminates the need for full signal decoding, reducing computational load while maintaining detection capability.
2Measurement precision
If conventional endpoint detection techniques are used that analyze the entire audio signal, then speech endpoint detection accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs endpoint detection on the encoded parameters before the audio signal is fully decoded. This preliminary action allows the system to identify speech endpoints earlier in the processing pipeline, reducing overall processing time while maintaining detection accuracy using the pre-decoded parameter information.
Solution Approach 2:
The patent extracts only the essential encoding parameters needed for endpoint detection, avoiding the computationally intensive process of fully decoding the audio signal. This extraction method maintains adequate detection accuracy while significantly reducing processing time on resource-constrained mobile devices.
3Adaptability or versatility
If a full ASR system is implemented on a mobile device, then speech recognition capability is improved, but device resource requirements increase
Solution Approach 1:
The patent segments the speech recognition system into distinct functional components: the mobile device performs encoding and endpoint detection using extracted parameters, while the server handles full decoding and speech recognition. This segmentation allows the mobile device to maintain speech recognition capability with reduced computational complexity by offloading intensive processing to the server.
Solution Approach 2:
The patent inverts the conventional ASR architecture by performing endpoint detection on encoded parameters at the client device before transmission, rather than waiting for full decoding at the server. This inversion enables the mobile device to contribute meaningfully to the speech recognition process with its limited resources, improving overall system adaptability without requiring full ASR capability on the device.
Data Source
AI summary
Systems, methods and apparatus for determining an estimated endpoint of human speech in a sound wave received by a mobile device having a speech encoder for encoding the sound wave to produce an encoded representation of the sound wave. The estimated endpoint may be determined by analyzing information available from the speech encoder, without analyzing the sound wave directly and without producing a decoded representation of the sound wave. The encoded representation of the sound wave may be transmitted to a remote server for speech recognition processing, along with an indication of the estimated endpoint.


