A Fusion Method Based on End-to-End Speech Recognition Model and Language Model

By changing the attention mechanism in the end-to-end speech recognition model and training the fully connected network, an internal language model estimation model is formed, and the problem of insufficient adaptability and accuracy in the existing technology is solved, and a more efficient speech recognition effect is achieved.

CN114596843BActive Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210242872.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-07-08
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

The existing language model fusion technology lacks adaptability in the end-to-end speech recognition model, and the internal language model estimation accuracy is limited, resulting in limited improvement in fusion effect.

Method used

By changing the attention mechanism of the end-to-end speech recognition model, an independent decoder model is formed, and a fully connected network replaces the attention mechanism, training the estimation model of the internal language model, and combining the external language model for fractional fusion, it is suitable for end-to-end speech recognition models of various attention mechanisms.

Benefits of technology

It significantly improves the accuracy of the fusion between the end-to-end speech recognition model and the language model, and is suitable for Conformer, BLSTM and Transformer encoders, improving the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596843B_ABST
    Figure CN114596843B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of end-to-end speech recognition, and discloses a fusion method based on an end-to-end speech recognition model and a language model, including the following steps: S1. Use speech and text to train an end-to-end speech recognition model, and use text data to train an external language model; S2. Separate the decoder part of the trained speech recognition model and form an independent model; S3. Use the training data to text to train the independent model alone and obtain an estimation model of the internal language model after convergence; S4. Decode the score fusion of the speech recognition model, the external language model, and the estimation model of the internal language model to obtain a decoding result. This algorithm can improve the recognition accuracy after the fusion of the speech recognition model and the language model, and has a wide application prospect in the field of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech recognition technology, and particularly relates to a fusion method based on an end-to-end speech recognition model and a language model. Background Art

[0002] Currently, the most classic speech recognition method is a method combining a Hidden Markov Model (HMM) and a Deep Neural Network (DNN). Although this method makes good use of the short-term stationary characteristics of speech signals, it still has disadvantages such as multi-model cascading of an acoustic model, a pronunciation dictionary, and a language model, inconsistent model training objectives, and a large decoding space. The invention of end-to-end speech recognition simplifies the entire speech recognition process, and the training objectives are simple and consistent.

[0003] Currently, end-to-end speech recognition models can be mainly divided into three categories: Connectionist Temporal Classification (CTC), Recurrent Neural Network-Transducer (RNN-Transducer), and Attention-based End-to-End Model (A-E2E). Among them, the independence assumption is introduced in the CTC model. The RNN-Transducer is mainly applied to streaming speech recognition models. The sequence model based on the attention mechanism aligns frame-level speech signals with text sequences using the attention mechanism, and its accuracy is the highest among end-to-end speech recognition models. The end-to-end speech recognition framework mainly consists of three parts: an encoder, a decoder, and an attention mechanism. Currently, an auxiliary language model is also very important to achieve better recognition results. The mainstream fusion algorithm for language models and speech recognition models is the Shallow Fusion (SF) technique. This technique works very well for traditional speech recognition models, but its improvement for end-to-end speech recognition models is very limited. This is mainly because, different from traditional speech recognition models, end-to-end speech recognition models model the entire sentence, so they will inevitably learn an Internal Language Model (ILM). This internal language model will affect the fusion of the speech recognition model and the external language model. As end-to-end models are increasingly widely used, more and more solutions have been proposed. The most well-known one is the Density Ratio method proposed by Masashi Sugiyama. This method trains a small language model on the data used to train the speech recognition model to approximate the ILM, and subtracts this approximate ILM when fusing with the external language model to reduce the influence of the ILM. This work slightly improves the fusion performance, but since the accuracy of estimating the true ILM through the approximate ILM cannot be guaranteed, the system accuracy improvement is very limited. Based on the Density Ratio method, Microsoft proposed the Internal Language Model Estimation technique, which can directly and accurately estimate the internal language model of the speech recognition model, so as to subtract a more accurately estimated ILM in the fusion stage, thus obtaining a great performance improvement.However, the ILME method proposed by Microsoft can only be applied to end-to-end speech recognition models with bidirectional long short-term memory network encoders and cannot be used on the newly proposed Transformer encoders and Conformer encoders. Therefore, its application is greatly limited. At the same time, since the method proposed by Microsoft does not have an adaptive function, even when applied to end-to-end speech recognition models with BLSTM encoders, the optimal effect cannot be achieved (INTERNAL LANGUAGE MODEL ESTIMATION FOR DOMAIN-ADAPTIVE END-TO-END). Summary of the Invention

[0004] In view of the deficiencies of existing language model fusion technologies and internal language model estimation technologies, the present invention proposes a fusion method based on an end-to-end speech recognition model and a language model. The main problem to be solved is that the existing algorithms for internal language estimation do not have an adaptive ability. At the same time, the existing technologies have limited accuracy in estimating internal language models, and the improvement in accuracy during fusion is also limited. The main application scenario of the present invention is an end-to-end speech recognition model based on the attention mechanism, abbreviated as end-to-end speech recognition. By means of model training, the internal language model in the end-to-end speech recognition model is estimated, and this estimated internal language model is subtracted during the model inference and decoding stage. Compared with traditional language model fusion technologies, the present invention can greatly improve the recognition accuracy after the end-to-end speech recognition model is fused with an external language model. At the same time, it can be applied to all speech recognition models based on the attention mechanism, including models based on Conformer encoders, BLSTM encoders, and Transformer encoders, and can also be applied to models with LSTM decoders and Transformer decoders. Therefore, it has a wider scope of application.

[0005] The present invention is achieved by at least one of the following technical solutions.

[0006] A fusion method based on an end-to-end speech recognition model and a language model, comprising the following steps:

[0007] S1. Use speech and text pairs to train an end-to-end speech recognition model, and use text data to train an external language model; the end-to-end speech recognition model includes an encoder, a decoder, and an attention mechanism;

[0008] S2. Separate the decoder of the trained end-to-end speech recognition model and form an independent model;

[0009] S3. Use the training data to text to train the independent model alone, and obtain an estimation model of the internal language model after convergence;

[0010] S4. Decode the score fusion of the end-to-end speech recognition model, the external language model, and the estimation model of the internal language model to obtain the decoding result.

[0011] Further, taking out the decoder of the trained end-to-end speech recognition model alone to form an independent model specifically means: changing the topological structure of the decoder to form an independent model, and the way of changing is: replacing the attention mechanism with a fully connected network.

[0012] Further, separately train the independent model with the text part data originally used to train the end-to-end speech recognition model, and obtain the estimation model of the internal language model after convergence;

[0013] During training, fix the parameters originally belonging to the decoder and only update the parameters of the added fully connected network. The parameters to be updated include the weights and bias parameters in the newly added fully connected network, and obtain the estimation model of the internal language model after convergence.

[0014] Further, use the Beam Search algorithm for decoding, and the score calculation during decoding is: adding the score of the end-to-end speech recognition model to the score of the external language model and then subtracting the score of the estimation model of the internal language model; the calculation method of the score is to input the normalized probability distribution output by the corresponding model into the natural logarithm function to obtain.

[0015] Further, control the score weights of the end-to-end speech recognition model, the external language model, and the estimation model of the internal language model by setting two fusion weights.

[0016] Further, the end-to-end speech recognition model is a recurrent neural network language model or a Transformer language model.

[0017] Further, the end-to-end speech recognition model must be an end-to-end speech recognition model containing an attention mechanism.

[0018] Further, the encoder is a Conformer encoder, a bidirectional long short-term memory network encoder, or a Transformer encoder.

[0019] Further, the decoder is a long short-term memory network decoder or a Transformer decoder.

[0020] Further, the attention mechanism is not limited to an additive attention mechanism, a position-sensitive attention mechanism, or a monotonic attention mechanism.

[0021] Compared with the existing technologies, the beneficial effects of the present invention are as follows: By changing the attention mechanism of the end-to-end speech recognition model, the present invention can estimate the internal language model, and can be applied to all end-to-end speech recognition models with attention mechanisms. Compared with the traditional language model fusion algorithm, the present invention can greatly improve the effect after the fusion of the end-to-end speech recognition model and the language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 FIG. is the overall structural block diagram of a fusion method based on an end-to-end speech recognition model and a language model in an embodiment;

[0023] Figure 2 FIG. is the structural diagram of a speech recognition model in an embodiment;

[0024] Figure 3 FIG. is the structural diagram of an internal language model estimation model modified from the decoder part of the speech recognition model in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] Embodiment 1

[0027] As Figure 1 、 Figure 2 、 Figure 3A fusion method based on an end-to-end speech recognition model and a language model is shown. In this implementation, the selected speech recognition model consists of a Conformer encoder, an additive attention mechanism, and an LSTM decoder. The Conformer encoder consists of 12 layers, each with a width of 512 dimensions, and the number of self-attention mechanisms in the encoder is eight. During training, random dropout is used to prevent the model from overfitting. The decoder is composed of a double-layer long short-term memory network, each with a width of 2048. The language model uses an RNN language model, which consists of three LSTM networks. The width of the hidden layer of the LSTM network is 2048 dimensions, and random dropout is used during training to prevent the model from overfitting. In this example, a Chinese general dataset is used to train the speech recognition model, and a medical dialogue dataset is selected as the test set. To match the test set, a large amount of medical corpus text is used to train the external language model in this example. In this implementation, 40-dimensional MFCC features are extracted from the selected speech dataset as input audio features, with a feature extraction window length of 25 milliseconds and a window shift of 10 milliseconds. The specific steps are as follows:

[0028] S1. Use the speech and text pairs in the Chinese general dataset to train the end-to-end speech recognition model. When training the end-to-end speech recognition model, the learning rate adopts the exponential decay method. At the beginning, the learning rate is 0.015, and after 400,000 iterations, it decays exponentially to 0.0015. And use the medical corpus text data to train the external language model. When training the language model, the learning rate adopts the exponential decay method. At the beginning, the learning rate is 0.015, and after 100,000 iterations, it decays exponentially to 0.0015;

[0029] The language model used in this implementation is a long short-term memory network language model, and its formula is as follows:

[0030]

[0031] where represents the predicted output of the external language model at the current moment, y i represents the word input at the i-th moment, y i-1 ...y0 represents the text sequence composed of all words from 0 to the moment before the current moment. LSTM represents the long short-term memory network, and softmax represents the activation function. After passing through the softmax activation function, the output is the normalized probability.

[0032] As another embodiment, a Transformer language model is applicable to replace the long short-term memory network language model. The Transformer language model can be expressed as:

[0033]

[0034] Among them, Transformer represents the Transformer network, and softmax represents the activation function. After passing through the softmax activation function, the output is the normalized probability.

[0035] S2. Separate the decoder part of the trained speech recognition model to form an independent model. And the internal language model estimation model can be obtained by replacing the attention mechanism in the original speech recognition model with a fully connected network and removing the encoder part. The specific formula is as follows:

[0036]

[0037] Among them, all subscripts represent the decoding moments. is the hidden state of the autoregressive decoder at the i-th moment. The hidden state is composed of the hidden state at the previous decoding moment and the predicted output at the previous moment calculated together. The hidden state can be converted into the query vector at the current moment by a fully connected network The content vector is directly obtained from the query vector after passing through the fully connected network The fully connected network FNN ilme used is a two-layer fully connected network with a width of 512 dimensions, and the activation function is RELU; represents the concatenation of the hidden state and the content vector at the current moment. Similar to the end-to-end speech recognition model, after obtaining the content vector at the current moment, it is concatenated with the hidden state at the current moment and then sent into the fully connected network FNN2. After passing through the softmax activation function, the normalized probability output at the current moment can be obtained.

[0038] S3. Use the text in the Chinese general dataset to train this independent model alone. During training, the learning rate adopts the exponential decay method. At the beginning, the learning rate is 0.015, and it decays exponentially to 0.0015 after 10,000 iterations. During training, fix the parameters originally belonging to the decoder in the model and only update the parameters of the newly added fully connected network. The specific operation formula is as follows:

[0039]

[0040] Loss ilm = CE(Y ilm , GT)

[0041]

[0042] Among them, is the state of the autoregressive decoder, and the hidden state is composed of the hidden state at the previous decoding moment and the predicted output at the previous moment Content vector is directly obtained from the query vector after passing through a fully connected network, and the content vector output by the attention mechanism at the current moment is used and the hidden state at the current moment After splicing and passing through a fully connected network, the predicted output at the current moment can be obtained represents the predicted output of the internal language model estimation model, GT represents the standard output at all moments, CE represents the cross-entropy loss function, and Loss ilm is the loss function used to train the internal language model estimation model. During training, the gradient descent algorithm is used, and the partial derivative of the parameters in the newly added fully connected network is used to update the parameters in the newly added fully connected network. The parameters to be updated include the weight parameters and bias parameters in the newly added fully connected network. At the same time, the parameters in the encoder remain fixed, and the parameters of the decoder refer to the parameters in the decoder network Decoder and the weights and biases in FNN1 and FNN2.

[0043] The converged model is called the estimation model of the internal language model;

[0044] S4. Use the Beam Search algorithm to decode the medical dialogue dataset. The score calculation method during decoding is: the score of the speech recognition model plus the score of the external language model and then subtract the score of the estimation model of the internal language model. The calculation formula for its fusion method is as follows.

[0045]

[0046] Among them and respectively represent the output normalized probabilities of the speech recognition model, external language model, and estimated internal language model at the current moment, and log represents the natural logarithm. λ elm and λ ilm respectively represent the fusion weights of the external language model and the internal language model. These two data need to be set according to specific situations, and usually the effect is better when these two parameters are equal. score iRepresents the fused score at the current moment, and the subsequent beam search algorithm decodes based on this score and outputs the final result

[0047] S5. The decoding result is the decoding result of the fusion of the speech recognition model and the external language model.

[0048] Embodiment 2

[0049] The end-to-end speech recognition model used in this embodiment can be represented in the following structure. The speech to be recognized is input into the encoder of the speech recognition model after being extracted into a feature sequence through feature extraction, and after being encoded by the encoder, it is output to the attention mechanism module for later use. The decoder uses an autoregressive method and comprehensively calculates the predicted output at the current moment through the attention mechanism. The specific formula is as follows:

[0050] H = Conformer(X)

[0051]

[0052] q i = FNN1(s i )

[0053] c i = Attention(H, q i )

[0054]

[0055] where X = [x1, x2,..., x t ,..., x T is the audio feature sequence to be recognized, where x t represents the audio feature of the t-th frame, and X ∈ R T×d , T is the length of the audio sequence, and d is the feature dimension; H = [h1, h2,..., h t ,..., h T is the output after being encoded by the encoder, and h t is the encoded output corresponding to the acoustic feature at the t-th moment. s i is the state of the autoregressive decoder, and the hidden state is jointly calculated by the hidden state s i-1 of the previous decoding moment and the predicted output of the previous moment. The hidden state can be converted into the query vector q i at the current moment by a fully connected network. The attention mechanism calculates the corresponding content vector c i according to this query vector. After concatenating the content vector output by the attention mechanism at the current moment and the hidden state s i at the current moment and passing through a fully connected network, the predicted output at the current moment can be obtained Subsequently, repeat this process in an autoregressive manner until the end symbol is predicted and decoding stops.

[0056] The attention mechanism in the end-to-end speech recognition model has the following function: The decoder obtains the acoustic information processed by the encoder through the attention mechanism.

[0057] As another embodiment, a bidirectional long short-term memory network encoder or a Transformer encoder can be used to replace the Conformer encoder. These two encoders can be represented as:

[0058] H = BLSTM(X)

[0059] H = Transformer(X).

[0060] Embodiment 3

[0061] The present embodiment relates to a fusion method based on an end-to-end speech recognition model and a language model, including the following steps:

[0062] S1. Use speech and text pairs to train the end-to-end speech recognition model, and use text data to train an external language model; the end-to-end speech recognition model includes an encoder, a decoder, and an attention mechanism, and the decoder obtains the acoustic information processed by the encoder through the attention mechanism;

[0063] The encoder is a Conformer encoder, a BLSTM encoder, or a Transformer encoder;

[0064] The decoder is an LSTM decoder or a Transformer decoder;

[0065] The attention mechanism is an additive attention mechanism, a position-sensitive attention mechanism, or a monotonic attention mechanism. The end-to-end speech recognition model can also be composed of the above modules.

[0066] S2. Separate the decoder of the trained end-to-end speech recognition model, and replace the attention mechanism of the model with a fully connected network with a double layer width of 512 to form an independent model. The first two layers of this fully connected network use the RELU activation function, and the last output layer does not use an activation function;

[0067] S3. Use the training data to the text to train the independent model alone. When training, the parameters originally belonging to the decoder of the end-to-end speech recognition model need to be fixed, and only the parameters in the newly added fully connected layer are updated. The learning rate during training uses the exponential decay method. The initial learning rate is 0.015, and it decays exponentially to 0.0015 after 10,000 iterations. And an estimation model of the internal language model is obtained after convergence;

[0068] S4. Decode the score fusion of the end-to-end speech recognition model, the external language model, and the estimation model of the internal language model to obtain a decoding result. When performing BeamSearch decoding, a relatively large Beamsize should be used as the decoding parameter. In this example, the Beamsize is 60.

[0069] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A fusion method based on an end-to-end speech recognition model and a language model, characterized in that It includes the following steps: S1. Train an end-to-end speech recognition model using speech and text, and train an external language model using text data; the end-to-end speech recognition model includes an encoder, a decoder, and an attention mechanism; S2. Separate the decoder of the trained end-to-end speech recognition model and form an independent model. Specifically: change the topological structure of the decoder to form an independent model, and the way of changing is: use a fully connected network to replace the attention mechanism; S3. Train the independent model separately with the text of the training data, and obtain an estimation model of the internal language model after convergence; specifically, train the independent model separately with the text part data of the trained end-to-end speech recognition model. When training, fix the parameters originally belonging to the decoder and only update the parameters of the added fully connected network. The parameters to be updated include the weights and bias parameters in the newly added fully connected network; S4. Decode the score fusion of the end-to-end speech recognition model, the external language model, and the estimation model of the internal language model to obtain a decoding result.

2. The fusion method based on the end-to-end speech recognition model and the language model according to claim 1, wherein Use the Beam Search algorithm for decoding, and the score calculation during decoding is: add the score of the end-to-end speech recognition model to the score of the external language model and then subtract the score of the estimation model of the internal language model; the calculation method of the score is to input the normalized probability distribution output by the corresponding model into the natural logarithm function to obtain.

3. A fusion method based on an end-to-end speech recognition model and a language model according to claim 2, characterized in that, Control the score weights of the end-to-end speech recognition model, the external language model, and the estimation model of the internal language model by setting two fusion weights respectively.

4. A fusion method based on an end-to-end speech recognition model and a language model according to claim 1, characterized in that The end-to-end speech recognition model is a recurrent neural network language model.

5. A fusion method based on an end-to-end speech recognition model and a language model according to claim 1, characterized in that, The end-to-end speech recognition model is a Transformer language model.

6. A fusion method based on an end-to-end speech recognition model and a language model according to claim 1, characterized in that The encoder is a Conformer encoder, a bidirectional long short-term memory network encoder, or a Transformer encoder.

7. A method for fusing an end-to-end speech recognition model and a language model according to claim 1, characterized in that, The decoder is a long short-term memory network decoder or a Transformer decoder.

8. A method for fusing an end-to-end speech recognition model and a language model according to any one of claims 1 to 7, characterized in that The attention mechanism is not limited to an additive attention mechanism, a position-sensitive attention mechanism, or a monotonic attention mechanism.

Citation Information

Patent Citations

  • End-to-end long-time speech recognition method

    CN113516968A

  • Customizable speech recognition system

    US20200327884A1