Filtering Model Training for Speech Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies face challenges in accurately recognizing speech due to variations in environment, user articulation habits, and inconsistencies between training data and actual speech signals, leading to inaccurate recognition results.
Innovation Solution
A filtering model training method that determines syllable distances between original and recognized syllables, using these distances to train a filtering model that improves the matching of sound signals with the speech recognition engine, thereby enhancing recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech recognition is performed using a preset database and speech recognition engine, then speech recognition can be implemented, but recognition accuracy deteriorates due to environmental variations, user articulation habits, and inconsistencies between training data and actual speech signals
Solution Approach 1:
The patent applies preliminary action by collecting actual speech data from users in their real environments and training a filtering model in advance. This pre-trained filtering model is then used to process new speech inputs, adapting to individual user characteristics and environmental conditions before actual recognition occurs, thereby improving accuracy without requiring changes to the core speech recognition engine
Solution Approach 2:
The patent introduces a filtering model as an intermediary component between the speech signal input and the speech recognition engine. This filtering model processes and adapts the speech signals based on user-specific characteristics and environmental conditions, serving as a mediator that bridges the gap between diverse real-world inputs and the fixed speech recognition engine, thereby improving recognition accuracy
2Reliability
If a filtering model is trained using syllable distances between original and recognized syllables, then speech recognition accuracy is improved, but system complexity increases due to additional training processes and data processing
Solution Approach 1:
The patent applies parameter changes by using syllable distance as a specific training parameter to optimize the filtering model. Instead of training on raw speech signals directly, the system calculates syllable distances between original and recognized syllables and uses these distance metrics as training parameters. This transforms the complex speech recognition problem into a more manageable parameter optimization problem, improving accuracy while controlling complexity
Solution Approach 2:
The system applies self-service by automatically collecting user speech data, calculating syllable distances, and training the filtering model without requiring manual intervention. The system serves itself by generating its own training data from actual usage and automatically optimizing its filtering model based on performance feedback, reducing the need for external configuration and maintenance
Data Source
AI summary
A filtering model training method includes obtaining N original syllables, obtaining N recognized syllables, and obtaining N syllable distances based on the N original syllables and the N recognized syllables, where the N syllable distances are in a one-to-one correspondence with N syllable pairs, the N original syllables and the N recognized syllables form the N syllable pairs, each syllable pair includes an original syllable and a recognized syllable that correspond to each other, and each syllable distance is used to indicate a similarity between an original syllable and a recognized syllable that are included in a corresponding syllable pair.


