Speech Recognition Terminal Intention Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Automatic Speech Recognition (ASR) systems face inefficiencies in responding to long audio signals, leading to delayed responses and potential loss of user intentions, which reduces accuracy and efficiency.
Innovation Solution
Implementing a continuous intention prediction method through real-time ASR processing, where the terminal determines and assembles answers in advance based on predicted intentions, and responds when a preset condition is met, allowing for immediate and accurate responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If ASR processing is performed only after the user finishes speaking, then the recognition accuracy is maintained, but the responding efficiency is seriously affected
Solution Approach 1:
The patent applies preliminary action by performing ASR processing and intention recognition in advance during the user's speech, rather than waiting for the speech to complete. The system continuously processes audio data and predicts intentions as the user speaks, assembling answers before the user finishes speaking, thereby reducing response time and improving responding efficiency.
Solution Approach 2:
The patent implements continuity of useful action through continuous ASR processing that operates throughout the user's speech without interruption. The system maintains continuous audio data processing, intention prediction, and answer assembly operations, ensuring that useful processing actions are performed continuously rather than in discrete batches after speech completion.
2Measurement precision
If single intention recognition is used, then the processing complexity is reduced, but the user intention is lost and responding accuracy decreases
Solution Approach 1:
The patent applies segmentation by dividing the intention recognition process into multiple segments. Instead of treating the entire speech as a single intention unit, the system segments the speech into multiple time points (first moment, second moment, etc.) and performs intention recognition on each segment separately, allowing for multiple predicted intentions to be captured and compared.
Solution Approach 2:
The patent utilizes parameter changes by dynamically adjusting the intention recognition parameters based on the speech progress. The system changes the recognition parameters and prediction models according to the current speech context at different time points, enabling more accurate intention identification while managing processing complexity through adaptive parameter adjustment.
Data Source
AI summary
A response method, a terminal, and a storage medium. The response method comprises: determining, at a first time point by means of speech recognition processing, a first target text corresponding to the first time point (1001); determining, according to the first target text, a first predicted intention and an answer to be pushed, wherein said answer is used for responding to speech information (1002); continuing to determine, by means of the speech recognition processing, a second target text corresponding to a second time point and a second predicted intention, wherein the second time point is the next successive time point of the first time point (1003); determining, according to the first predicted intention and the second predicted intention, whether a preset response condition is satisfied (1004); and responding according to said answer if the preset response condition is determined to be satisfied (1005).


