Multi-Terminal Shared Text for Dynamic Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems require pre-acquisition of text materials and human intervention, limiting their usage scenarios and reducing accuracy in dynamic environments with sudden theme changes.
Innovation Solution
A method that involves acquiring text contents and time-associated information from multiple terminals, determining shared text, and generating a customized language model for real-time speech recognition, incorporating word segmentation, classification, and updating hot word lists to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If pre-acquisition of text materials and human intervention are used to update speech recognition model, then speech recognition accuracy is improved, but usage scenario is limited and operation complexity increases
Solution Approach 1:
The system automatically collects text materials from multiple terminals in the preset scenario, performs automated word segmentation, classification, and keyword extraction without human intervention. The speech recognition model is automatically updated based on the shared text from terminals, enabling the system to serve itself and eliminate the need for manual model updating while expanding usage scenarios.
2Measurement precision
If pre-acquisition of text materials and human intervention are used to update speech recognition model, then speech recognition accuracy is improved, but operation complexity increases
Solution Approach 1:
The system automatically collects text materials from multiple terminals in the preset scenario, performs automated word segmentation, classification, and keyword extraction without human intervention. The speech recognition model is automatically updated based on the shared text from terminals, enabling the system to serve itself and eliminate the need for manual model updating while expanding usage scenarios.
Solution Approach 2:
The system merges text materials from multiple terminals in the preset scenario to create a comprehensive shared text corpus. By combining data sources from multiple devices participating in the same scenario, the system builds a more robust language model that reflects the actual usage context, thereby improving accuracy without requiring manual curation.
3Ease of operation
If original speech recognition model is used without updates, then operation simplicity is maintained, but speech recognition accuracy decreases in dynamic environments
Solution Approach 1:
The system dynamically adapts the speech recognition model to changing scenarios by automatically detecting theme changes and updating the language model based on shared text from terminals. The model transitions from a static pre-trained state to a dynamic state that continuously adapts to the current preset scenario, maintaining both operational simplicity and high accuracy in dynamic environments.
Solution Approach 2:
The system collects feedback from multiple terminals in the preset scenario through shared text materials, uses this feedback to identify theme changes, and automatically updates the speech recognition model accordingly. This feedback loop enables the system to maintain high accuracy in dynamic environments without requiring manual intervention, as the model continuously learns from actual usage data.
4Adaptability or versatility
If automated model updating from multiple terminals is implemented, then adaptability to dynamic scenarios is improved, but system complexity increases
Solution Approach 1:
The system merges text materials from multiple terminals in the preset scenario to create a comprehensive shared text corpus. By combining data sources from multiple devices participating in the same scenario, the system builds a more robust language model that reflects the actual usage context, thereby improving accuracy without requiring manual curation.
Solution Approach 2:
The system automatically collects text materials from multiple terminals in the preset scenario, performs automated word segmentation, classification, and keyword extraction without human intervention. The speech recognition model is automatically updated based on the shared text from terminals, enabling the system to serve itself and eliminate the need for manual model updating while expanding usage scenarios.
Data Source
AI summary
A speech recognition method and related products are provided. The method includes acquiring text contents and text-associated time information transmitted by a plurality of terminals in a preset scenario and determining a shared text for the preset scenario based on the text contents and the text-associated time information, obtaining a customized language model for the preset scenario based on the shared text, and performing speech recognition for the preset scenario with the customized language model. The method provides improved speech recognition for the preset scenario due to the correlation between the customized language model and the preset scenario.


