Hybrid Frequency Acoustic Recognition Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition methods require separate acoustic models for different sampling frequencies, leading to cumbersome maintenance, insufficient training data, and limited robustness and generalization, especially when dealing with hybrid sampling frequencies and environmental noise.
Innovation Solution
A training method for a hybrid frequency acoustic recognition model that unifies recognition models for speech signals of different sampling frequencies by performing pre-training and supervised parameter training using both 16 kHz and 8 kHz speech signals, incorporating MFCC and fbank features, and employing a partially connected deep neural network to improve noise suppression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate acoustic models are established for different sampling frequencies, then the model can match the training environment better, but the update and maintenance becomes very cumbersome
Solution Approach 1:
The patent merges multiple sampling frequency data (8kHz, 16kHz, and other frequencies) into a unified acoustic recognition model. Instead of maintaining separate models for each sampling frequency, the system processes mixed-frequency speech data through a single model that learns frequency-invariant features, thereby reducing model maintenance complexity while preserving recognition accuracy across different frequencies
Solution Approach 2:
The unified acoustic recognition model is designed to handle multiple sampling frequencies universally. The model architecture and training methodology enable it to process speech data from various sources with different sampling rates without requiring frequency-specific configurations, making the system multi-functional and adaptable to diverse recording environments
2Measurement precision
If separate acoustic models are trained for different sampling frequencies, then each model can be optimized for its specific frequency, but the training data becomes insufficient and robustness is limited
Solution Approach 1:
The patent combines training data from multiple sampling frequencies (8kHz, 16kHz, and other frequencies) into a unified training set. This merging approach allows the acoustic recognition model to learn from a larger, more diverse dataset, improving both recognition precision and robustness simultaneously by exposing the model to varied frequency characteristics during training
Solution Approach 2:
The patent employs parameter changes in the form of data augmentation and frequency transformation techniques. By artificially generating or transforming speech data across different sampling frequencies, the system enriches the training dataset, enabling the model to learn frequency-invariant features that improve both precision and adaptability to different recording conditions
Data Source
AI summary
The invention discloses a training method and a speech recognition method for a mixed frequency acoustic recognition model, which belongs to the technical field of speech recognition. The method comprises: obtaining a first-type speech feature of the first speech signal, and processing the first speech data to obtain corresponding first speech training data (S1); obtaining the first-type speech feature of the second speech signal, and processing the second speech data to obtain corresponding second speech training data (S2); obtaining a second-type speech feature of the first speech signal according to a power spectrum of the first speech signal, and obtaining the second-type speech feature of the second speech signal according to a power spectrum of the second speech signal (S3); performing pre-training according to the first speech signal and the second speech signal, so as to form a preliminary recognition model of the hybrid frequency acoustic recognition model (S4); and performing supervised parameter training on the preliminary recognition model according to the first speech training data, the second speech training data and the second-type speech feature, so as to form the hybrid frequency acoustic recognition model (S5). The beneficial effects of the above technical solution are: the recognition model has better robustness and generalization.


