Multi-Terminal Shared Text for Dynamic Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems require pre-acquisition of text materials and human intervention, limiting their usage scenarios and reducing accuracy in dynamic environments with sudden theme changes.

Innovation Solution

A method that involves acquiring text contents and time-associated information from multiple terminals, determining shared text, and generating a customized language model for real-time speech recognition, incorporating word segmentation, classification, and updating hot word lists to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pre-acquisition of text materials and human intervention are used to update speech recognition model, then speech recognition accuracy is improved, but usage scenario is limited and operation complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidusage scenario
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system automatically collects text materials from multiple terminals in the preset scenario, performs automated word segmentation, classification, and keyword extraction without human intervention. The speech recognition model is automatically updated based on the shared text from terminals, enabling the system to serve itself and eliminate the need for manual model updating while expanding usage scenarios.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If pre-acquisition of text materials and human intervention are used to update speech recognition model, then speech recognition accuracy is improved, but operation complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidoperation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system automatically collects text materials from multiple terminals in the preset scenario, performs automated word segmentation, classification, and keyword extraction without human intervention. The speech recognition model is automatically updated based on the shared text from terminals, enabling the system to serve itself and eliminate the need for manual model updating while expanding usage scenarios.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system merges text materials from multiple terminals in the preset scenario to create a comprehensive shared text corpus. By combining data sources from multiple devices participating in the same scenario, the system builds a more robust language model that reflects the actual usage context, thereby improving accuracy without requiring manual curation.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If original speech recognition model is used without updates, then operation simplicity is maintained, but speech recognition accuracy decreases in dynamic environments

Engineering Contradiction:
Improveoperation simplicityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system dynamically adapts the speech recognition model to changing scenarios by automatically detecting theme changes and updating the language model based on shared text from terminals. The model transitions from a static pre-trained state to a dynamic state that continuously adapts to the current preset scenario, maintaining both operational simplicity and high accuracy in dynamic environments.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system collects feedback from multiple terminals in the preset scenario through shared text materials, uses this feedback to identify theme changes, and automatically updates the speech recognition model accordingly. This feedback loop enables the system to maintain high accuracy in dynamic environments without requiring manual intervention, as the model continuously learns from actual usage data.

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If automated model updating from multiple terminals is implemented, then adaptability to dynamic scenarios is improved, but system complexity increases

Engineering Contradiction:
Improveadaptability to dynamic scenariosVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system merges text materials from multiple terminals in the preset scenario to create a comprehensive shared text corpus. By combining data sources from multiple devices participating in the same scenario, the system builds a more robust language model that reflects the actual usage context, thereby improving accuracy without requiring manual curation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system automatically collects text materials from multiple terminals in the preset scenario, performs automated word segmentation, classification, and keyword extraction without human intervention. The speech recognition model is automatically updated based on the shared text from terminals, enabling the system to serve itself and eliminate the need for manual model updating while expanding usage scenarios.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12451128B2Voice recognition method and related product
Publication Date: 2025.10.21 IFLYTEK CO LTD
  • US12451128B2 patent drawing
  • US12451128B2 patent drawing
  • US12451128B2 patent drawing

AI summary

A speech recognition method and related products are provided. The method includes acquiring text contents and text-associated time information transmitted by a plurality of terminals in a preset scenario and determining a shared text for the preset scenario based on the text contents and the text-associated time information, obtaining a customized language model for the preset scenario based on the shared text, and performing speech recognition for the preset scenario with the customized language model. The method provides improved speech recognition for the preset scenario due to the correlation between the customized language model and the preset scenario.