Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2results about How to "Improve speech recognition performance" patented technology

Noise robust acoustic feature extraction method based on gamma channel scaled basis vector

ActiveCN115662408BGuaranteed carrying capacityImprove speech recognition performanceSpeech recognitionFrequency spectrumFeature extraction
The application discloses a noise-robust acoustic feature extraction method based on a gamma pass scaling basis vector. The noise-robust acoustic feature extraction method based on the gamma pass scaling basis vector comprises the following steps: performing pre-emphasis processing on a speech signal; performing frame processing on the speech signal after the pre-emphasis processing; performing Fourier transform on the speech signal after the frame processing; calculating a scaling coefficient according to the spectral distribution characteristics of a gamma pass filter bank, and optimizing a basis vector based on the scaling coefficient; and performing discrete cosine transform on the speech signal after the Fourier transform and the optimized basis vector to extract noise-robust acoustic features from the speech signal. According to the application, the scaling coefficient is calculated based on the frequency domain distribution characteristics of the gamma pass filter bank, and is directly applied to the basis vector for generating acoustic features, so that the original details of the speech signal are retained to the greatest extent, the information carrying capacity of the acoustic features under the interference of noise signals is ensured, and the speech recognition effect can be effectively improved.
Owner:INFORMATION SCI RES INST OF CETC +1

Training methods, devices, equipment, and storage media for speech recognition models

This application discloses a training method, apparatus, device, and storage medium for a speech recognition model, belonging to the field of artificial intelligence. The method includes: acquiring a sample audio set, the sample audio set including multiple sample audios; filtering candidate sample audios from the sample audio set based on an initial speech recognition model; extracting audio segments from the candidate sample audios; wherein the audio segments include audios aligned with consecutive identical text units from the candidate sample audios; and the initial speech recognition model, when performing speech recognition on the candidate sample audios, failed to correctly recognize the consecutive identical text units; and retraining the initial speech recognition model based on the audio segments to obtain a target speech recognition model. This application can improve speech recognition quality, particularly improving the accuracy of recognizing consecutive identical text units.
Owner:SOUNDAI TECH CO LTD