Display Voice Endpoint Detection Using Energy and Text Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition technologies fail to accurately recognize the endpoint of user speech input due to noise or other sounds, leading to undesired speech recognition results and user inconvenience.
Innovation Solution
A display device with a network interface and controller that obtains speech input, determines energy level and endpoint information, and uses text analysis to accurately recognize the end of speech input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech recognition technology uses only signal size information (amplitude, energy strength) to recognize speech endpoint, then the system operation is simple, but the speech endpoint recognition accuracy deteriorates due to noise or other sounds
Solution Approach 1:
The patent segments the speech recognition system into multiple independent modules: energy level detection module, text analysis module, and endpoint determination module. Each module performs a specific function (detecting energy levels, analyzing text for endpoint indicators, and综合 determining speech endpoint), allowing the system to achieve high accuracy through coordinated operation of specialized components rather than a single complex algorithm
Solution Approach 2:
The patent introduces text analysis as an intermediary layer between raw speech signal detection and endpoint determination. The speech signal is first converted to text, and the text analysis module identifies endpoint indicators within the text structure. This intermediary processing enables more accurate endpoint detection by leveraging linguistic patterns while filtering out noise interference
2Reliability
If the TV continues to receive speech input when the user has finished speaking, then no additional speech input is missed, but undesired speech recognition results are produced due to noise recognition
Solution Approach 1:
The patent implements feedback mechanisms where the text analysis module continuously monitors the recognized text for endpoint indicators and feeds this information back to the endpoint determination module. This feedback loop allows the system to dynamically adjust speech reception timing based on linguistic analysis, stopping exactly when the user finishes speaking rather than using fixed time thresholds, thus preventing noise recognition while avoiding missed input
Solution Approach 2:
The patent performs preliminary text analysis and endpoint indicator detection during the speech recognition process itself, rather than waiting for the speech to end. By analyzing the text structure and identifying endpoint indicators in advance, the system can proactively determine when to stop receiving speech input, improving both accuracy and efficiency
Data Source
AI summary
The present disclosure relates to a display device capable of accurately recognizing an end point of a speech input of a user, and the display device may comprise a network interface which communicates with a first server and a second server, and a controller which: acquires a speech input of a user; transmits, to the first server, a speech signal corresponding to the acquired speech input; receives, from the first server, the energy level of the speech signal, text corresponding to the speech input, and speech end point information for the speech input; and determines whether an utterance of the user has ended on the basis of the energy level and the speech end point information.


