Speech enhancement and speech recognition cascade system and design method thereof

By introducing a bidirectional long and short-term memory network and a dynamic combination multi-head attention mechanism in the speech enhancement and speech recognition system, and combining gradient mapping technology and gradient scaling strategy, the problem of insufficient accuracy of noise speech recognition in existing systems is solved, achieving higher accuracy and robustness of speech recognition.

CN120048260APending Publication Date: 2025-05-27SHANGHAI JIAOTONG UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510274619.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the existing speech enhancement and speech recognition systems process noise-free speech, although the speech enhancement module is pre-enhanced speech, there is still a problem of insufficient recognition accuracy, mainly because the performance of the speech recognition system itself needs to be improved, especially in the modeling and analysis capabilities of speech acoustic features.

Method used

Short-time Fourier transform is used to convert the input voice signal time-frequency and input it into the voice enhancement module, which is built on a bidirectional long and short-term memory network. At the same time, a dynamic combination of multi-head attention mechanism is introduced into the Conformer module of the speech recognition module, and end-to-end training is combined with gradient mapping technology and gradient scaling strategies.

Benefits of technology

It effectively improves the recognition accuracy and robustness of the speech recognition system, and significantly improves the overall performance of the speech enhancement and speech recognition cascade system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048260A_ABST
    Figure CN120048260A_ABST
Patent Text Reader

Abstract

The invention provides a speech enhancement and speech recognition cascade system and a design method thereof, the system comprises a speech enhancement module and a speech recognition module, short-time Fourier transform is adopted to carry out time-frequency transformation on an input speech signal, and then the speech signal is input to the speech enhancement module; the speech enhancement module is constructed based on a bidirectional long short-term memory network, carries out time-frequency inverse transformation on the output of the speech enhancement module to obtain a pre-enhanced speech signal, inputs the pre-enhanced speech signal into the speech recognition module for speech recognition, and outputs a speech recognition result. The speech recognition module comprises a plurality of improved DC-Conformer modules serving as encoders and a plurality of Transformer modules serving as decoders, the improved DC-Conformer modules introduce a dynamic combined multi-head attention mechanism into the Conformer modules in a speech recognition task, and the dynamic combined multi-head attention mechanism is applied to the Conformer modules in the speech recognition task, so that the dynamic combined multi-head attention mechanism can be used for recognizing the speech. Performance improvement of a speech enhancement and speech recognition cascade system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular, to a speech enhancement and speech recognition cascade system and a design method thereof. Background Art

[0002] As one of the core algorithms for front-end processing of speech signals, the speech enhancement algorithm has received extensive attention in the academic and industrial fields. The purpose of speech enhancement is to suppress the noise component in the noisy speech signal and retain the target speech therein, thereby improving the perceptual quality and intelligibility of the speech signal. The speech enhancement module is often used as a pre-processing module for many downstream tasks, including automatic speech recognition, speaker recognition, etc. Among them, cascading the speech enhancement module and the speech recognition module to improve the recognition rate of the speech recognition module for noisy speech is a very important technical path and has attracted much attention from researchers.

[0003] Speech recognition technology is a technology that converts speech signals into text information. Traditional speech recognition algorithms mainly rely on the combination of acoustic models and language models, and realize the conversion from speech to text through a series of complex signal processing and pattern recognition technologies. In the existing research on cascaded speech enhancement and speech recognition systems and methods, although the speech pre-enhanced by the speech enhancement module has less interference noise and higher human ear perceptual quality, the speech enhancement process will also cause certain damage to the target speech and introduce some distortions that are insensitive to the human ear, resulting in further room for improvement in the recognition accuracy of the subsequent speech recognition system for the enhanced speech. An important reason for this phenomenon is that the performance of the existing speech recognition system itself needs to be improved, and there is room for optimization in its ability to model and analyze speech acoustic features. Summary of the Invention

[0004] Aiming at the defects in the prior art, the purpose of the present invention is to provide a speech enhancement and speech recognition cascade system and a design method thereof, which can effectively improve the speech recognition accuracy of the overall speech enhancement and speech recognition cascade system.

[0005] To solve the above problems, the technical solution of the present invention is as follows:

[0006] A cascaded system for speech enhancement and speech recognition, including a speech enhancement module and a speech recognition module. The input speech signal is subjected to time-frequency transformation using the short-time Fourier transform and then input into the speech enhancement module. The speech enhancement module is constructed based on a bidirectional long short-term memory network. The output of the speech enhancement module is subjected to inverse time-frequency transformation to obtain a pre-enhanced speech signal, which is input into the speech recognition module for speech recognition and the speech recognition result is output. The speech recognition module includes a number of improved DC-Conformer modules as encoders and a number of Transformer modules as decoders. The improved DC-Conformer module is a Conformer module with a dynamically combined multi-head attention mechanism introduced into the speech recognition task. By applying the dynamically combined multi-head attention mechanism, the performance of the cascaded system for speech enhancement and speech recognition is improved.

[0007] Preferably, the speech enhancement module includes three layers of bidirectional long short-term memory network layers for performing temporal analysis and modeling on the input short-time Fourier transform magnitude spectrum, and one fully connected layer is used to achieve feature mapping.

[0008] Preferably, the encoder of the speech recognition module includes 12 improved DC-Conformer modules, and the decoder of the speech recognition module includes 6 Transformer modules.

[0009] Preferably, the cascaded system for speech enhancement and speech recognition applies gradient mapping technology and gradient scaling strategy during training.

[0010] Furthermore, the present invention also provides a design method for a cascaded system for speech enhancement and speech recognition, including the following steps:

[0011] Construct a speech enhancement model for enhancing the input speech signal;

[0012] Construct a speech recognition model for recognizing the pre-enhanced speech signal;

[0013] Construct a data set to train and test the cascaded model for speech enhancement and speech recognition.

[0014] Preferably, in the step of constructing the data set to train and test the cascaded model for speech enhancement and speech recognition, the CHiME-4 data set is used for training and evaluating the speech enhancement and speech recognition system, and the gradient mapping technology and gradient scaling strategy are applied during the training process of the cascaded model for speech enhancement and speech recognition.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0016] 1. The present invention introduces a dynamic combined multi-head attention mechanism into a speech recognition model, effectively enhancing the ability of the neural network to model the long-distance context global dependence of input audio features, thereby improving the recognition accuracy and robustness of the speech recognition system, and ultimately significantly enhancing the performance of the proposed speech enhancement and speech recognition cascaded system.

[0017] 2. The present invention uses a speech enhancement module as a pre-preprocessing module of the speech recognition model, effectively improving the quality of the speech signal to be recognized; at the same time, applying gradient mapping technology and gradient scaling strategy, effectively ensuring that during end-to-end training, the gradient of the speech enhancement task in the backpropagation process of the neural network has a positive promoting effect on the optimization goal of the final speech recognition accuracy, thereby taking into account the performance of both speech enhancement and speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0019] Figure 1 is a schematic structural diagram of the speech enhancement and speech recognition cascaded system of the present invention;

[0020] Figure 2 is a schematic diagram of the calculation principle of the Compose function in the dynamic combined multi-head attention mechanism;

[0021] Figure 3 is a flowchart of the speech enhancement and speech recognition cascaded method of the present invention;

[0022] Figure 4 is a flowchart of the training and testing of the speech enhancement and speech recognition cascaded system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0024] Specifically, the present invention provides a speech enhancement and speech recognition cascaded system, as Figure 1As shown, the system includes a voice enhancement module and a speech recognition module. After receiving an input voice signal, it first performs time-frequency transformation on it to obtain a voice time-frequency spectrogram. In this embodiment, the specific time-frequency transformation method used is the Short Time Fourier Transform (STFT). After obtaining the STFT magnitude spectrum of the input voice, it is input into the voice enhancement neural network module. The voice enhancement neural network module is constructed based on the Bidirectional Long Short-Term Memory (BiLSTM), specifically including 3 layers of BiLSTM layers for performing temporal analysis and modeling on the input STFT magnitude spectrum, and using a fully connected layer to achieve feature mapping. Finally, the output of the voice enhancement module is the predicted STFT spectral mask, which is multiplied element-wise with the STFT spectrum of the input voice to filter out the interference noise, and finally the pre-enhanced voice spectrum is obtained. Performing inverse time-frequency transformation on the pre-enhanced voice spectrum, that is, inverse STFT, the pre-enhanced voice signal can be obtained. Calculate the mean square error loss between the STFT magnitude spectra of the pre-enhanced voice signal and the target clean voice signal as the loss of the voice enhancement task. Next, the pre-enhanced voice signal is input into the speech recognition module for speech recognition, and the cross-entropy loss is calculated between the speech transcription text information output by the speech recognition module and the corresponding label information as the loss of the speech recognition task.

[0025] In the present invention, the speech recognition module adopts an Encoder-Decoder (ED) architecture, including 12 improved DC-Conformer modules as the encoder and 6 Transformer modules as the decoder. Among them, the improved DC-Conformer module is a Conformer module that introduces the Dynamically Composable Multi-Head Attention (DCMHA) mechanism into the speech recognition task, that is, using the improved DC-Conformer module to replace the traditional Conformer module. For the improved DC-Conformer module, the calculation principle diagram of the Compose function in the included dynamically composable multi-head self-attention mechanism is as Figure 2 shown. The embedding dimension of the self-attention mechanism is set to 256, the number of attention heads is set to 4, and the dimension of the feed-forward layer is set to 2048.

[0026] In the improved DC-Conformer module, the mathematical representation of the dynamically composable multi-head self-attention modeling is:

[0027]

[0028] Among them, are the query matrix, key matrix, and value matrix of the input of the self-attention mechanism module (L represents the sequence length, and D m represents the sequence input feature dimension), is the feature mapping matrix of the i-th attention head (D h represents the feature dimension of each attention head), is the output mapping matrix (H represents the number of attention heads).

[0029] Compared with the traditional self-attention mechanism, the dynamic combination multi-head self-attention mechanism adds two combination calculation functions Compose(·) to process the attention score matrix A S and the attention weight matrix A W obtained in the self-attention feature calculation process, so as to realize the dynamic combination and adjustment of the attention scores / weights calculated by different self-attention heads. For the combination calculation function Compose(·), it can calculate the attention scores / weights of the H attention heads for each pair of query vectors Q i (i ∈ [1,..., L]) - key vectors K j (j ∈ [1,..., L]) and perform dynamic combination. The mathematical representation of this process is: as follows:

[0030]

[0031] Among them, corresponds to a linear mapping. represents element-wise multiplication. includes the following non-linear transformation:

[0032] ω q1 , ω q2 = Chunk(GELU(Q i W q1 ), W q2 , dim = 1)

[0033] ω q1 = Rmsnorm(Reshape(ω q1 , (H, R)), dim = 0)

[0034] ω q2 = Reshape(ω q2 , (R, H))

[0035] ω qg = tanh(Q i W qg)

[0036] ω k1 , ω k2 = Chunk(GELU(K j W k1 )W k2 , dim = 1)

[0037] ω k1 = Rmsnorm(Reshape(ω k1 , (H, R)), dim = 0)

[0038] ω k2 = Reshape(ω k2 , (R, H))

[0039] ω kg = tanh(K j W kg )

[0040] Among them,

[0041] By applying the above dynamic combined multi-head self-attention mechanism, the performance of the Conformer model in the field of speech recognition is optimized, thereby ensuring a further improvement in the performance of the final speech enhancement and speech recognition cascade system.

[0042] Meanwhile, further, based on the improved speech recognition model proposed above, the present invention also pre-sets a speech enhancement network to ensure higher speech quality input to the speech recognition system, thereby improving the speech recognition accuracy. However, in the backpropagation process of the actual end-to-end training of the cascade model, the optimization directions of the gradients corresponding to the speech enhancement task and the gradients corresponding to the speech recognition task may be contradictory, resulting in the gradient of the speech enhancement task hindering the effective optimization of the final speech recognition performance. In response, by referring to and applying the existing gradient compensation strategy, it is ensured that while using the gradient of the speech enhancement task to optimize the speech enhancement network, it will not have an obvious negative impact on the optimization of the speech recognition task, thereby ensuring that the entire system has excellent speech recognition accuracy. Specifically, in the process of parameter iteration update during the backpropagation gradient of the neural network, according to the respective loss functions of the speech enhancement task and the speech recognition task, the respective corresponding gradients G SE and G ASR . can be obtained. When the optimization directions of the gradients G SE and G ASR are opposite, that is, the angle between G SE and G ASR in the high-dimensional space is greater than 90 degrees, G SE is rotated through gradient mapping technology to ensure G SE and GASR The included angle between them is less than 90 degrees. At the same time, a gradient scaling strategy is also adopted to prevent the total gradient in the backpropagation process of the speech enhancement and speech recognition cascaded network from being dominated by G SE dominated.

[0043] Furthermore, the present invention also provides a design method for a speech enhancement and speech recognition cascaded system, as Figure 3 shown, the method includes the following steps:

[0044] S1: Construct a speech enhancement model for enhancing the input speech signal;

[0045] Specifically, the speech enhancement model is constructed based on the bidirectional long short-term memory network BiLSTM, and specifically includes 3 BiLSTM layers for performing temporal analysis and modeling on the input STFT magnitude spectrum, and one fully connected layer is used to achieve feature mapping.

[0046] S2: Construct a speech recognition model for recognizing the pre-enhanced speech signal;

[0047] Specifically, the speech recognition model includes 12 improved DC-Conformer modules as encoders and 6 Transformer modules as decoders. The improved DC-Conformer module is a Conformer module that introduces a dynamic combination of multi-head attention mechanisms into the speech recognition task, that is, the traditional Conformer module is replaced with the improved DC-Conformer module. By applying the dynamic combination of multi-head self-attention mechanisms, the performance of the Conformer model in the field of speech recognition is optimized, thereby ensuring the performance improvement of the speech enhancement and speech recognition cascaded system.

[0048] S3: Construct a data set to train and test the speech enhancement and speech recognition cascaded model.

[0049] Specifically, as Figure 4As shown, the CHiME-4 dataset is used for the training and evaluation of the speech enhancement and speech recognition system. This dataset is particularly suitable for studying speech recognition tasks in noisy environments and aims to promote the development of speech recognition technology in complex noisy backgrounds. The training set of the CHiME-4 dataset contains 8,738 noisy speech utterances (including 1,600 real-recorded noisy speech utterances and 7,138 synthetic simulated noisy speech utterances), the development set contains 3,280 noisy speech utterances (including 1,640 real-recorded noisy speech utterances and 1,640 synthetic simulated noisy speech utterances), and the test set contains 2,640 noisy speech utterances (including 1,320 real-recorded noisy speech utterances and 1,320 synthetic simulated noisy speech utterances). After preparing the CHiME-4 dataset, the cascaded model of speech enhancement and speech recognition built is trained on the CHiME-4 dataset, and the gradient compensation strategy is applied during the training process. Then, the performance of the trained cascaded model of speech enhancement and speech recognition is evaluated on the CHiME-4 test set. The performance evaluation metric used is the Word Error Rate (WER). The lower this metric is, the higher the speech recognition accuracy and the better the model performance. Taking the speech enhancement-speech recognition cascaded system without using the DC-Conformer module proposed in the present invention as the baseline model, it is compared whether the advanced system based on DC-Conformer proposed in the present invention brings performance gain compared to the baseline model. If no performance gain is observed, the hyperparameters of the dynamic combination multi-head self-attention mechanism in the DC-Conformer module are tried to be adjusted. The evaluation results of the word error rate metric on the CHiME-4 dataset are shown in Table 1 below.

[0050]

[0051] Table 1

[0052] It can be seen that the cascaded system of speech enhancement and speech recognition based on DC-Conformer proposed in the present invention has further improved the speech recognition accuracy compared to the traditional cascaded baseline system based on Conformer. In addition, if the preposed speech enhancement module is removed, the speech recognition accuracy will decrease, which also proves the performance gain and necessity of the preposed speech enhancement module for the entire system. For the speech enhancement module itself, its speech enhancement performance is also evaluated using the signal-to-noise ratio metric, and the results are shown in Table 2 below.

[0053] Development set Test set Original noisy speech 3.61 dB 1.53 dB Speech pre-enhanced by the pre-enhancement module of speech enhancement 4.92 dB 2.39 dB

[0054] Table 2

[0055] It can be seen that the speech enhancement module of the present invention has indeed improved the quality of the input original noisy speech, thus bringing a positive gain to the improvement of the recognition accuracy of the subsequent speech recognition system.

[0056] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. A speech enhancement and speech recognition cascade system, characterized in that: The system includes a speech enhancement module and a speech recognition module. The input speech signal is transformed into a time-frequency signal by using a short-time Fourier transform and then input into the speech enhancement module. The speech enhancement module is constructed based on a bidirectional long short-term memory network. The output of the speech enhancement module is transformed into a pre-enhanced speech signal by performing a time-frequency inverse transform. The pre-enhanced speech signal is input into the speech recognition module for speech recognition and outputting a speech recognition result. The speech recognition module includes a plurality of improved DC-Conformer modules as encoders and a plurality of Transformer modules as decoders. The improved DC-Conformer module introduces a dynamic combination multi-head attention mechanism into a Conformer module in a speech recognition task. By applying the dynamic combination multi-head attention mechanism, the performance of a speech enhancement and speech recognition cascade system is improved.

2. The speech enhancement and speech recognition cascade system according to claim 1, characterized in that: The speech enhancement module includes three layers of bidirectional long short-term memory network layers, which are used to perform time series analysis and modeling on the input short-time Fourier transform amplitude spectrum, and a fully connected layer is used to realize feature mapping.

3. The speech enhancement and speech recognition cascade system according to claim 1, characterized in that: The encoder of the speech recognition module includes 12 improved DC-Conformer modules, and the decoder of the speech recognition module includes 6 Transformer modules.

4. The speech enhancement and speech recognition cascade system according to claim 1, characterized in that: The speech enhancement and speech recognition cascade system applies gradient mapping technology and gradient scaling strategy during training.

5. A method for designing a cascade system of speech enhancement and speech recognition, characterized in that: The method comprises the following steps: Constructing a speech enhancement model for performing speech enhancement on an input speech signal; constructing a speech recognition model for performing speech recognition on the pre-enhanced speech signal; Build a dataset to train and test the cascade model of speech enhancement and speech recognition.

6. The method for designing a cascade system of speech enhancement and speech recognition according to claim 5, characterized in that: In the steps of constructing a data set and training and testing the speech enhancement and speech recognition cascade model, the CHiME-4 data set is used to train and evaluate the speech enhancement and speech recognition system, and the gradient mapping technology and gradient scaling strategy are applied in the training process of the speech enhancement and speech recognition cascade model.

Citation Information

Patent Citations

  • Model training method and device and device for model training

    CN113707134A

  • Speech recognition method fused with speech enhancement

    CN114495969A

  • Speech recognition method based on self-supervised pre-training and interactive fusion network

    CN116631383A

  • Cooling system using seawater, water supply and LNG cold heat

    KR102906331B1