Hybrid Frequency Acoustic Recognition Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition methods require separate acoustic models for different sampling frequencies, leading to cumbersome maintenance, insufficient training data, and limited robustness and generalization, especially when dealing with hybrid sampling frequencies and environmental noise.

Innovation Solution

A training method for a hybrid frequency acoustic recognition model that unifies recognition models for speech signals of different sampling frequencies by performing pre-training and supervised parameter training using both 16 kHz and 8 kHz speech signals, incorporating MFCC and fbank features, and employing a partially connected deep neural network to improve noise suppression.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate acoustic models are established for different sampling frequencies, then the model can match the training environment better, but the update and maintenance becomes very cumbersome

Engineering Contradiction:
Improvemodel matching accuracyVSAvoidmodel maintenance complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple sampling frequency data (8kHz, 16kHz, and other frequencies) into a unified acoustic recognition model. Instead of maintaining separate models for each sampling frequency, the system processes mixed-frequency speech data through a single model that learns frequency-invariant features, thereby reducing model maintenance complexity while preserving recognition accuracy across different frequencies

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified acoustic recognition model is designed to handle multiple sampling frequencies universally. The model architecture and training methodology enable it to process speech data from various sources with different sampling rates without requiring frequency-specific configurations, making the system multi-functional and adaptable to diverse recording environments

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If separate acoustic models are trained for different sampling frequencies, then each model can be optimized for its specific frequency, but the training data becomes insufficient and robustness is limited

Engineering Contradiction:
Improverecognition precisionVSAvoidmodel robustness
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent combines training data from multiple sampling frequencies (8kHz, 16kHz, and other frequencies) into a unified training set. This merging approach allows the acoustic recognition model to learn from a larger, more diverse dataset, improving both recognition precision and robustness simultaneously by exposing the model to varied frequency characteristics during training

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent employs parameter changes in the form of data augmentation and frequency transformation techniques. By artificially generating or transforming speech data across different sampling frequencies, the system enriches the training dataset, enabling the model to learn frequency-invariant features that improve both precision and adaptability to different recording conditions

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11120789B2Training method of hybrid frequency acoustic recognition model, and speech recognition method
Publication Date: 2021.09.14 YUTOU TECH HANGZHOU
  • US11120789B2 patent drawing
  • US11120789B2 patent drawing
  • US11120789B2 patent drawing

AI summary

The invention discloses a training method and a speech recognition method for a mixed frequency acoustic recognition model, which belongs to the technical field of speech recognition. The method comprises: obtaining a first-type speech feature of the first speech signal, and processing the first speech data to obtain corresponding first speech training data (S1); obtaining the first-type speech feature of the second speech signal, and processing the second speech data to obtain corresponding second speech training data (S2); obtaining a second-type speech feature of the first speech signal according to a power spectrum of the first speech signal, and obtaining the second-type speech feature of the second speech signal according to a power spectrum of the second speech signal (S3); performing pre-training according to the first speech signal and the second speech signal, so as to form a preliminary recognition model of the hybrid frequency acoustic recognition model (S4); and performing supervised parameter training on the preliminary recognition model according to the first speech training data, the second speech training data and the second-type speech feature, so as to form the hybrid frequency acoustic recognition model (S5). The beneficial effects of the above technical solution are: the recognition model has better robustness and generalization.