Speech Recognition Terminal Intention Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Automatic Speech Recognition (ASR) systems face inefficiencies in responding to long audio signals, leading to delayed responses and potential loss of user intentions, which reduces accuracy and efficiency.

Innovation Solution

Implementing a continuous intention prediction method through real-time ASR processing, where the terminal determines and assembles answers in advance based on predicted intentions, and responds when a preset condition is met, allowing for immediate and accurate responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If ASR processing is performed only after the user finishes speaking, then the recognition accuracy is maintained, but the responding efficiency is seriously affected

Engineering Contradiction:
Improveresponding efficiencyVSAvoidresponse time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing ASR processing and intention recognition in advance during the user's speech, rather than waiting for the speech to complete. The system continuously processes audio data and predicts intentions as the user speaks, assembling answers before the user finishes speaking, thereby reducing response time and improving responding efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuity of useful action through continuous ASR processing that operates throughout the user's speech without interruption. The system maintains continuous audio data processing, intention prediction, and answer assembly operations, ensuring that useful processing actions are performed continuously rather than in discrete batches after speech completion.

Inventive Principle:
Principle #20Continuity of useful action

2Measurement precision

If single intention recognition is used, then the processing complexity is reduced, but the user intention is lost and responding accuracy decreases

Engineering Contradiction:
Improveresponding accuracyVSAvoidintention recognition complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the intention recognition process into multiple segments. Instead of treating the entire speech as a single intention unit, the system segments the speech into multiple time points (first moment, second moment, etc.) and performs intention recognition on each segment separately, allowing for multiple predicted intentions to be captured and compared.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent utilizes parameter changes by dynamically adjusting the intention recognition parameters based on the speech progress. The system changes the recognition parameters and prediction models according to the current speech context at different time points, enabling more accurate intention identification while managing processing complexity through adaptive parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12165640B2Response method, terminal, and storage medium for speech response
Publication Date: 2024.12.10 BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
  • US12165640B2 patent drawing
  • US12165640B2 patent drawing
  • US12165640B2 patent drawing

AI summary

A response method, a terminal, and a storage medium. The response method comprises: determining, at a first time point by means of speech recognition processing, a first target text corresponding to the first time point (1001); determining, according to the first target text, a first predicted intention and an answer to be pushed, wherein said answer is used for responding to speech information (1002); continuing to determine, by means of the speech recognition processing, a second target text corresponding to a second time point and a second predicted intention, wherein the second time point is the next successive time point of the first time point (1003); determining, according to the first predicted intention and the second predicted intention, whether a preset response condition is satisfied (1004); and responding according to said answer if the preset response condition is determined to be satisfied (1005).