Scenario-Adaptive Voice Recognition Using Shared Terminal Text

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems require pre-acquisition of text materials and human intervention, limiting their usage scenarios and reducing accuracy in the face of sudden theme changes.

Innovation Solution

A method that involves acquiring text contents and time-associated information from multiple terminals, determining a shared text, and generating a customized language model and hot word list to improve speech recognition accuracy in real-time scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text materials and keywords are acquired in advance to update the speech recognition model, then speech recognition accuracy is improved, but human intervention is required and usage scenarios are limited

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidhuman intervention requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically collects text materials from multiple terminals during the actual usage scenario, performs word segmentation and classification to extract keywords, and updates the hot word list without human intervention. The speech recognition model self-adapts to the scenario by utilizing real-time data from the environment.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs word segmentation, classification, and keyword extraction in advance during the data collection phase, preparing the hot word list before speech recognition is executed. This preliminary processing enables the model to be ready for accurate recognition without requiring last-minute manual updates.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If text materials are acquired in advance to update the speech recognition model, then field recognition performance is improved, but the system cannot adapt to sudden or temporary theme changes

Engineering Contradiction:
Improvefield recognition performanceVSAvoidtheme change adaptation
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically updates the hot word list and speech recognition model based on real-time text materials collected from multiple terminals during the actual scenario. Instead of relying on static pre-acquired materials, the system continuously adapts to theme changes by incorporating new keywords and updating the language model on-the-fly.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system collects text materials from multiple terminals during the scenario, processes them to extract keywords, and uses this feedback to update the hot word list and refine the speech recognition model. This closed-loop feedback mechanism enables continuous improvement and adaptation to changing themes.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If a customized language model is generated based on shared text from multiple terminals, then speech recognition accuracy is improved, but data processing complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system merges text materials from multiple terminals to create a comprehensive shared text corpus. By combining data sources, the system generates a more robust and scenario-specific language model that leverages the collective information from all terminals, improving recognition accuracy through aggregated data.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs word segmentation on the shared text to break it down into individual words and phrases, then classifies them to extract relevant keywords. This segmentation approach simplifies the processing of large volumes of text by handling it in manageable units, reducing overall complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4083999B1Voice recognition method and related product
Publication Date: 2025.08.27 IFLYTEK CO LTD
  • EP4083999B1 patent drawingFigure 1~2
  • EP4083999B1 patent drawingFigure 3~5
  • EP4083999B1 patent drawingFigure 6~7

AI summary

A voice recognition method and a related product. The method comprises: obtaining text content and text time information sent by a plurality of terminals in a preset scene, and determining a shared text of a preset scene according to the text content and the text time information (S101); and obtaining a customized language model of the preset scene according to the shared text, and executing voice recognition of the preset scene by using the customized language model (S102). A terminal in a preset scene can be utilized to acquire text content and text time information of the preset scene to determine a shared text of the preset scene. Therefore, the customized language model is obtained according to the shared text, the correlation between the customized language model and the preset scene is higher, the voice recognition of the preset scene is executed by using the customized language model, and the accuracy of the voice recognition can be effectively improved.