An end-to-end Chinese speech recognition method

By combining the Conformer encoder with the LAS attention decoder, and adding the CTC decoder and the auxiliary training of the intermediate layer CTC loss, the problem of insufficient speech recognition accuracy and training speed in the prior art is solved, and higher speech recognition accuracy and faster training convergence on the Aishell-1 dataset are achieved.

CN114373451BActive Publication Date: 2025-06-24JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210077486.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-06-24
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

Existing end-to-end Chinese speech recognition technology has shortcomings in accuracy and training speed, especially when using the state-of-the-art Conformer model as listener, its effects have not been fully explored.

Method used

A hybrid CTC/Attention end-to-end Chinese speech recognition method based on Conformer is proposed, combining the Conformer encoder and the LAS attention decoder, and adding the CTC decoder and the intermediate layer CTC loss as subtask assisted training to improve the accuracy of speech recognition and the training convergence speed.

Benefits of technology

On the Aishell-1 dataset, the proposed Conformer-LAS-CTC model significantly improves the accuracy of speech recognition, and compared with other codec combination models, the word error rate is reduced by 19.52% to 46.74%, and achieves the best performance of CER4.54% in a 3-layer LSTM network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114373451B_ABST
    Figure CN114373451B_ABST
Patent Text Reader

Abstract

An end-to-end Chinese speech recognition method belongs to the field of speech recognition. First, the effect of the Transformer-LAS speech recognition model based on the Transformer encoder and the LAS decoder was explored. Aiming at the problem that the Transformer is not good at capturing local information, Conformer was used to replace the Transformer, and the Conformer-LAS model was proposed. Secondly, since the overly flexible alignment method of Attention will cause its performance to drop sharply in a noisy environment, connectionist temporal classification (CTC) was adopted in the research for auxiliary training to accelerate convergence, and a phoneme-level intermediate CTC loss was added for joint optimization, and the Conformer-LAS-CTC speech recognition model with better performance was proposed. Finally, the proposed model was verified on the open-source Chinese Mandarin Aishell-1 dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech recognition, and particularly relates to a hybrid CTC / Attention end-to-end Chinese speech recognition method based on Conformer Background Art

[0002] Automatic Speech Recognition (ASR) systems are widely used in many products to support various business applications, such as: mobile assistants, smart homes, customer service robots, meeting records, etc., and have become an indispensable part of life. Traditional ASR systems usually consist of three parts: an acoustic model, a pronunciation dictionary, and a language model. Building and adjusting these individual components is usually quite complex. In recent years, with the rapid development of computing power and the sharp growth of data resources, end-to-end (E2E) ASR systems that integrate the three traditional speech recognition modules have made remarkable progress. Different from the aforementioned hybrid architectures, E2E models only require audio and corresponding text labels, and can directly convert speech inputs into character sequence outputs by learning the mapping from speech to text in a single model, greatly simplifying the training process. Currently, popular E2E speech methods are mainly built based on the following three models: connectionist temporal classification (CTC), attention based encoder decoder (AED), and transducers. These deep learning models are easy to build and optimize, and their recognition rates in some application scenarios exceed those of models based on traditional speech recognition methods. They can also be flexibly combined to utilize the advantages of different basic models to achieve better results

[0003] Based on CTC, an end-to-end acoustic model is constructed without frame-level alignment labels in the time dimension, greatly simplifying the acoustic model training process. Graves [Graves A, Fernández S, Gómez F, et al. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks[C] / / Proceedings of the 23rd international conference on Machine learning. 2006:369-376.] first constructed the Neural network CTC (NN-CTC) acoustic model and verified its effectiveness for acoustic modeling; Hannun et al. [Hannun A, Case C, Casper J, et al. Deep Speech: Scaling up end-to-end speech recognition[J]. Computer Science, 2014.] adopted a 5-layer RNN with bidirectional recurrent layers, trained with CTC loss and corrected with a language model, achieving the best results at that time on the Switchboard dataset. At the same time, they also proposed some optimization schemes. Amodei et al. [Amodei D, Ananthanarayanan S, Anubhai R, et al. Deep speech 2: End-to-end speech recognition in english and mandarin[C] / / International conference on machine learning. PMLR, 2016:173-182.] achieved better results on this basis using a model with 13 hidden layers (including convolutional layers). Jaesong Lee [Lee J, Watanabe S. Intermediate loss regularization for ctc-based speech recognition[C] / / ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2021:6224-6228.] proposed an intermediate CTC loss to regularize CTC training and improve performance.

[0004] The self-attention-based Transformer architecture has been widely used in sequence modeling due to its ability to capture long-range interactions and high training efficiency. However, although Transformer is effective in extracting long-sequence dependencies, it is relatively weak in extracting fine-grained local feature patterns. In [Gulati A, Qin J, Chiu C C, et al. Conformer: Convolution-augmented transformer for speech recognition [J]. arXiv preprint arXiv:2005.08100, 2020.], it is assumed that both global and local interactions are important for parameter effectiveness. By combining CNN, which is good at extracting local features but requires more layers or parameters to capture global information, a new combination of self-attention and convolution, Conformer, is proposed, enabling self-attention to learn global interactions while convolution effectively captures local correlations based on relative offsets.

[0005] Chan proposed Listen, Attend and Spell (LAS) in [Chan W, Jaitly N, Le Q, et al. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition [C] / / 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2016.]. Different from previous methods, LAS does not make independence assumptions in the label sequence and does not rely on HMM. LAS is also based on a sequence-to-sequence learning framework with attention. It consists of an encoder recurrent neural network (RNN) as the listener and a decoder RNN as the speller. The listener uses a pyramidal RNN to convert low-level speech signals into high-level features. The speller uses an attention mechanism to specify the probability distribution of the character sequence and converts these higher-level features into output labels. However, previous work has not explored the effects brought by using the state-of-the-art Conformer model as the listener.

[0006] Based on the above content, the present invention first explores the performance of the LAS speech recognition system composed of different codec combinations, and compares the accuracy of speech recognition under different codec structures; secondly, a Conformer-based LAS speech recognition model (Conformer-LAS) is proposed by combining the Conformer encoder with the LAS model; in order to further improve the accuracy of speech recognition and accelerate the convergence speed of model training, CTC decoder joint training is added, and the intermediate layer CTC loss proposed in [Lee J, Watanabe S. Intermediate loss regularization for ctc-based speech recognition [C] / / ICASS 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021: 6224-6228.] is added as a sub-task auxiliary training, and a Conformer-LAS-CTC speech recognition model is proposed; finally, speech recognition research is carried out based on the Aishell-1 dataset, and the experimental effects of different models are compared. The experimental results verify the advanced nature of the Conformer-LAS-CTC speech recognition model proposed in the present invention. Summary of the invention

[0007] The present invention aims to solve the problems existing in the prior art and provide a hybrid CTC / Attention end-to-end Chinese speech recognition method based on Conformer.

[0008] The technical solution of the present invention:

[0009] An end-to-end Chinese speech recognition method, the steps are as follows:

[0010] 1. Data Preprocessing

[0011] The speech data is pre-emphasized, framed, and windowed, and fast Fourier transform is performed to calculate the spectral line energy, perform Mel filtering, and take the logarithm to obtain the Fbank feature; the pre-processed data is divided into a training set and a validation set;

[0012] 2. Building a Conformer-based hybrid CTC / Attention model

[0013] The Conformer-based hybrid CTC / Attention model consists of three parts: a shared Conformer encoder, a CTC decoder, and a LAS attention decoder.

[0014] The shared Conformer encoder first processes the input using a convolutional subsampling layer, and inputs the data processed by the convolutional subsampling layer into N Conformer encoder blocks. Each Conformer encoder block sequentially includes a feedforward module, a multi-head self-attention module MHSA (Multi-head self-attention Module), a convolution module (Convolution Module), a feedforward module (Feedforward module), and layer normalization. A residual unit is set after each module in the Conformer encoder. Among them, a half-step residual connection is adopted between the feedforward module and the multi-head self-attention module, and between the feedforward module and layer normalization; the multi-head self-attention module includes layer normalization, multi-head self-attention with integrated relative sine position encoding, and dropout; the convolution module contains a pointwise convolution with an expansion factor of 2, projects the number of channels through a GLU activation layer, followed by a one-dimensional depth convolution, and a Batchnorm and a swish activation layer are connected after the one-dimensional depth convolution. The shared Conformer encoder maps the input frame-level acoustic features x=(x1,...x T ) to the sequence high-level representation h=(h1,h2,...,h U ).

[0015] The described LAS attention decoder adopts a two-layer unidirectional LSTM structure and introduces an attention mechanism. The specific decoding process is as follows: using local attention to focus on the information output by the shared Conformer encoder, and using LSTM to decode the information. During the output process of each LSTM, the LAS attention decoder jointly performs attention decoding on the already generated text (y1,y2,...,y s-1 ) and the output features h=(h1,h2,...,h U ) of the shared Conformer encoder, and finally generates the target transcription sequence y=(y1,y2,...,y S ). Thus, the probability of the output sequence y is as follows:

[0016]

[0017] At each time step t, the conditional dependence of the output on the encoder feature h is calculated through the attention mechanism. The attention mechanism is a function of the current decoder hidden state and the encoder output feature, and compresses the encoder feature into a context vector u through the following mechanism it .

[0018]

[0019] where h i is the output feature of the shared Conformer encoder; vector b a , and matrix W h , W d are all learned parameters; d t represents the hidden state of the decoder at time step t. Then, softmax is applied to u it to obtain the attention distribution:

[0020] a t = softmax(u t ) (4)

[0021] Using α it , the corresponding context vector is obtained by weighted summation of h i :

[0022]

[0023] At each time step, the attention decoder hidden state d t for capturing the previous output context is obtained as follows:

[0024]

[0025] where d t-1 is the previous hidden state, is the embedding layer vector learned through y t-1 . At time step t, the posterior probability of the output y t is as follows:

[0026] P(y t |h, y<t) = softmax(W s [c t ; d t + b s ) (7)

[0027] where W s and b s are learnable parameters.

[0028] The CTC decoder described above decodes with the output feature h of the shared Conformer encoder as the input. After passing through the Softmax layer, the output of the CTC decoder is P(q t |h), where q t is the output at time step t. Then, the label sequence l is the sum of all path probabilities:

[0029]

[0030] where: Γ(qt ) is a many-to-one mapping of the label sequence. Since multiple paths correspond to the same label sequence, duplicate labels and blank labels in the path need to be removed. q t ∈A, t = 1, 2,..., T, where A is the set of labels with the blank label "-" added, and the labeled sequence l with the highest probability in the output sequence is * as follows:

[0031] l * = arg l maxP(l|h) (9)

[0032] The loss function of the CTC decoder is the sum of the negative logarithmic probabilities of all labels. The CTC network is trained through backpropagation:

[0033] CTC loss = -logP(l|h) (10)

[0034] In the training of the CTC decoder, skip all layers after the intermediate layer, and add the intermediate layer phoneme-level CTC loss, that is, InterCTC loss as an auxiliary task to induce a sub-model. By obtaining the intermediate representation of the CTC decoder to calculate the loss of the sub-model, similar to the complete model of the CTC decoder, the loss function of the sub-model is as follows:

[0035]

[0036] where, represents the output of the sub-model.

[0037] The hybrid CTC / Attention model based on Conformer jointly optimizes the model parameters using the CTC decoder and the LAS attention decoder, and at the same time adds the intermediate layer phoneme-level CTC decoder loss for regularizing the lower-level parameters. Therefore, the loss function is defined as follows during the training process:

[0038] T loss = λCTC loss + μInterCTC loss +(1 - λ - μ)Att loss (12)

[0039] where, CTC loss , InterCTC loss , Att loss are the CTC decoder loss, the intermediate layer phoneme-level CTC decoder loss, and the LAS attention decoder loss respectively. λ and μ are two hyperparameters used to measure the weights of the CTC decoder, the intermediate layer phoneme-level CTC decoder, and the LAS attention decoder.

[0040] During the training process, the loss decline curve is converged to a stable state, the training is ended, and the final model is obtained;

[0041] III. Train the Conformer-based hybrid CTC / Attention model, and use the trained model to verify the validation set to achieve end-to-end Chinese speech recognition.

[0042] The technical effects of the present invention: The present invention proposes a Conformer-LAS-CTC acoustic model for end-to-end speech recognition. We studied the recognition effects of different codec combinations, proved the combination of the Conformer encoder and the LAS decoder, added phoneme-level CTC assisted decoding, and introduced joint training of the intermediate CTC loss. The model shows the best performance on the Aishell-1 dataset. The present invention also compared the traditional speech recognition model and other end-to-end models, and verified the advancement of the Conformer-LAS-CTC acoustic model. The model achieved the best performance of CER 4.54% when the Conformer decoder has a 3-layer LSTM network. In future research, the effects of different hyperparameters on the model will be explored, and the robustness of the model will be studied by fusing external language model decoding. Description of the Drawings

[0043] Figure 1 is the Conformer encoder model architecture;

[0044] Figure 2 is the LAS model architecture;

[0045] Figure 3 is the Conformer-LAS-CTC speech recognition model;

[0046] Figure 4 is the loss during the training process;

[0047] Figure 5 is the word error rate on the validation set. Detailed Embodiments

[0048] 1 Related Work

[0049] 1.1 Conformer Encoder

[0050] The Conformer proposed by Anmol Gulati [Gulati A, Qin J, Chiu C C, et al. Conformer: Convolution-augmented transformer for speech recognition[J]. arXiv preprint arXiv:2005.08100, 2020.] compares [Dong L, Xu S, Xu B. Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition[C] / / 2018 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2018:5884-5888.] combines convolution and self-attention. Self-attention learns global interactions, while convolution effectively captures local correlations based on relative offsets, resulting in more effective results than using convolution or self-attention alone. The Conformer Encoder first processes the input using a convolutional subsampling layer and then uses a large number of conformer blocks instead of [Zhang Q, Lu H, Sak H, et al. Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss[C] / / ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2020:7829-7833.] [Karita S, Chen N, Hayashi T, et al. A comparative study on transformer vs rnn in speech applications[C] / / 2019 IEEE Automatic Speech Recognition and Understanding Workshop(ASRU). IEEE, 2019:449-456.] the Transformer blocks in to process the input, Figure 1The overall architecture of the Conformer encoder is shown on the left, and the specific structure of the Conformer block is shown on the right:

[0051] The Conformer block consists of three modules: a feedforward module, a multi-head self-attention module, and a convolution module. There is a feedforward layer before and after the Conformer block. The multi-head self-attention module and the convolution module are sandwiched in the middle. The feedforward layer uses a half-step residual connection. Layer normalization (Layernorm) follows each large module, and residual units are used in each module. Through this structure, convolution and attention are concatenated to achieve an enhanced effect.

[0052] In the adopted multi-head self-attention module (MHSA), an important technology of Transformer XL is also integrated, namely the relative sinusoidal position encoding scheme. Relative position encoding enables the self-attention module to have better generalization ability for different input lengths and makes the encoder more robust to changes in the utterance length.

[0053] The convolution module contains a pointwise convolution with an expansion factor of 2, projects the number of channels through a GLU activation layer, followed by a one-dimensional depth convolution, and then Batchnorm and a swish activation layer are connected after the convolution.

[0054] In the Conformer block, the same feedforward module is deployed before and after. Each FFN contributes half of the value, called a half-step FFN. Mathematically, for the input x of the i-th Conformer block i , the output h i The calculation formula is as follows:

[0055]

[0056] Among them, FFN refers to the feedforward module, MHSA refers to the multi-head self-attention module, Conv refers to the convolution module, Layernorm represents layer normalization, and residual connections are used between each module.

[0057] 1.2 LAS Decoder

[0058] The LAS model includes an encoder listener, a decoder speller, and an attention network. The general model architecture is as Figure 2 shown.

[0059] where listener is the encoder of the acoustic model that performs an encoding operation which converts the input acoustic sequence x = (x1,..., x T ) into a high-level representation h, where the length of the high-level feature sequence h can be the same as that of the input acoustic sequence x or a downsampled short sequence.

[0060] The present invention explores the influence of three different model structures, namely BLSTM, Transformer, and Conformer, as listeners on the overall speech recognition model.

[0061] The speller is an attention-based decoder. At each output step, the transducer generates a probability distribution for the next character based on all the characters seen previously, and thus the probability of the output sequence y is as follows:

[0062]

[0063] At each time step t, the conditional dependence of the output on the encoder feature h is calculated through the attention mechanism. The attention mechanism is a function of the current decoder hidden state and the encoder output features, and compresses the encoder features into a context vector u through the following mechanism it .

[0064]

[0065] where the vector b a , and the matrices W h , W d are all learned parameters; d t represents the hidden state of the decoder at time step t. Then, softmax is applied to u it to obtain the attention distribution:

[0066] α t = softmax(u t )(4) Using α it , the corresponding context vector is obtained by weighted summation of h i :

[0067]

[0068] At each moment, the decoder hidden state d t used to capture the previous output context is obtained as follows:

[0069]

[0070] where d t-1 is the previous hidden state. is the embedding layer vector learned through y t-1 At time t, the posterior probability of the output y t is as follows:

[0071] P(y t |h,y<t)=softmax(W s [c t ;d t +b s ) (7)

[0072] where W s and b s are learnable parameters. Finally, the model loss function is defined as:

[0073] Att loss =-log(P(y|x)) (8)

[0074] 1.3 Connectionist temporal classification (CTC)

[0075] CTC adds a blank symbol to the labeled symbol set, which means that there is no predicted value output for this frame. Therefore, there are many blank symbols in the predicted output of the model. Only one peak in a whole segment of speech corresponding to a phoneme is confirmed by the recognizer, and the others are recognized as blanks. As a result, it is equivalent to automatically segmenting the phoneme boundary, achieving the elimination of blank symbols and continuously occurring states, and finally obtaining the predicted character sequence.

[0076] Given the input sequence h, after the output of the Softmax layer, the output of the network is P(q t |h), where q t is the output at time t. The label sequence l is the sum of all path probabilities:

[0077]

[0078] where: Γ(q t ) is the many-to-one mapping of the label sequence. Since the same label sequence may correspond to multiple paths, it is necessary to remove duplicate labels and blank labels in the path. q t ∈A,t=1,2,...,T, A is the label set with the blank label "-" added. The most probable annotation sequence in the output sequence is:

[0079] l * =arg l maxP(l|h) (9)

[0080] The loss function of CTC is the sum of the negative logarithmic probabilities of all labels, and the CTC network can be trained through backpropagation:

[0081] CTC loss =-logP(l|h) (10)

[0082] 2 Model Architecture

[0083] To implement a better speech recognition model, the present invention uses the Conformer model as the encoder (listener), and the Attention and spell parts of the LAS model are jointly decoded with the CTC model to jointly construct an end-to-end Conformer-LAS-CTC speech recognition system. Figure 3 The model architecture is given.

[0084] It includes three parts, a shared encoder, a CTC decoder, and an attention decoder. The shared encoder consists of N Conformer encoder layers. The CTC decoder consists of a linear layer and a log softmax layer, and the CTC loss function is applied to the softmax output during training. The LAS decoder structure is introduced in detail in Section 1.2 above.

[0085] 2.1 Combination of Conformer and LAS

[0086] In the comparative experiments with other encoder models, Conformer achieved the best results in all cases. Among them, the convolutional block is the most important in terms of effect, and the effect of two half-step FFNs is also better than the structure with only one FFN. Integrating relative sinusoidal position encoding in the multi-head self-attention mechanism enables the self-attention module to have good generalization ability and stronger robustness even when the input lengths are different. Therefore, in the model proposed in the present invention, the Conformer encoder is used to map the input frame-level acoustic features x=(x1,...x M ) to a sequence of high-level representations (h1, h2,..., h U ).

[0087] The LAS decoder then specifies the probability distribution of the character sequence by using the attention mechanism. Compared with other end-to-end models, the LAS network generates the character sequence without making any independent assumptions between characters. This also determines that the decoding of this model will bring better accuracy. In the structure proposed in the present invention, the Conformer encoder and the LAS decoder are combined. The decoder performs attention decoding on the jointly hidden states (h1, h2,..., h s-1 ) of the already generated text (y1, y2,..., y U ), and converts and decodes these higher-level features, and finally generates the target transcription sequence (y1, y2,..., yS )。

[0088] 2.2 CTC-assisted training

[0089] Since CTC can be regarded as an objective function that can directly optimize the likelihood of the input sequence and the output target sequence, under this objective function, CTC automatically learns and optimizes the correspondence between the input and output sequences during training. Therefore, a phoneme-level CTC decoder is added to the structure of the present invention for assisted training.

[0090] In the residual network regularization technique, stochastic depth helps train very deep networks by randomly skipping some layers. However, due to its integration strategy, it is ineffective for regularizing the lower layers. Inspired by this, in CTC training, all layers after the intermediate layer are skipped, and an intermediate CTC loss (InterCTC loss ) is added as an auxiliary task to induce the submodel. Training a submodel that depends on the lower layers can regularize the lower part of the entire model, thereby further improving the performance of CTC.

[0091] We consider an N-layer encoder with a CTC loss function. Since the submodel and the complete model share the lower structure, by obtaining the intermediate representation of the model to calculate its corresponding CTC loss, the CTC loss is also used for the submodel in the same way as for the complete model:

[0092]

[0093] The output of the submodel is expressed as which is the intermediate representation of the complete model. Then, the original CTC loss and the intermediate CTC loss are used for training to regularize the lower layers with a very small computational overhead.

[0094] 2.3 Multi-task loss

[0095] CTC can learn the monotonic alignment between acoustic features and label sequences, which helps the encoder converge faster; the attention-based decoder can learn the dependencies between target sequences. Therefore, combining CTC and the attention loss not only helps the convergence of the attention-based decoder but also enables the hybrid model to utilize label dependencies.

[0096] The model of the present invention jointly optimizes the model parameters using CTC and the LAS decoder, and at the same time adds a phoneme-level CTC loss for the intermediate layer to regularize the lower-layer parameters to further improve the model performance. Therefore, the loss function is defined as follows during training:

[0097]

[0098] where CTC loss , InterCTC loss , Attloss They are the CTC loss, the intermediate CTC loss, and the attention loss respectively. λ and μ are two hyperparameters used to measure the CTC, the intermediate CTC, and the attention weights.

[0099] 3 Experimental Results and Analysis

[0100] 3.1 Experimental Data

[0101] The dataset used in the experiments of this invention is the 178h dataset (Aishell-1) open-sourced by Hill Shell, with a sampling rate of 16 kHz. It includes 400 speakers from different Chinese accent regions, and the corpus content covers finance, technology, sports, entertainment, and current affairs news. It is divided into a training set, a validation set, and a test set according to the non-overlapping principle. The training set has 120,418 audio files, the validation set has 14,331 audio files, and the test set has 7,176 audio files.

[0102] 3.2 Experimental Platform

[0103] The hardware configuration used in the experiments of this invention is an Intel(R) Core(TM) i7-5930K processor, 32GB of running memory, and the GPU graphics card is NVIDIA GeForce GTX TITAN X; the software environment is a Pytorch deep learning environment built on a 64-bit Ubuntu18.04 operating system.

[0104] 3.3 Experimental Steps

[0105] In the experiments of the present invention, 80-dimensional FBank (Filter Banks) is used as the input feature, where the frame length is 25 ms and the frame shift is 10 ms. During training, we use the Adam optimizer

Kingma D P, Ba J. Adam: A method for stochastic optimization[J]. arXiv preprint arXiv:1412.6980, 2014.

Zhang Q, Lu H, Sak H, et al. Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss[C] / / ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020:7829-7833.

[26] , and use SpecAugment proposed by Google [Park D S, Chan W, Zhang Y, et al. Specaugment: A simple data augmentation method for automatic speech recognition[J]. arXiv preprint arXiv:1904.08779, 2019.] to randomly mask a part of the information in the time domain and frequency domain, where the masking parameters are F = 27 and T = 100. At the audio feature input part, two 2-D convolutional neural network (CNN) modules are used. Each module has two convolutional neural networks, a batch normalization layer (BatchNorm2d), and a ReLu activation function. Each CNN has 32 filter banks, each filter kernel size is 3x3, and the stride is 1. Then it is followed by a 2D max pooling layer (2D-MaxPool) with a kernel size of 2x2 and a stride of 2. Then it passes through a linear layer (Linear) and outputs a dimension of 256. Finally, two 1D max pooling layers (1D-Maxpool) with a kernel size of 2 and a stride of 2 are used for downsampling to reduce redundant speech feature information. The main network structure is LAS. Listen uses the parameter configuration in the Conformer-based Encoder structure. The multi-head attention layer uses d_model = 256, h = 4, the feed-forward neural network layer d_ff = 1024. In the convolutional module, the input channel of the Pointwise CNN is 256 and the output is 512, and the kernel size is 1. The input channel of the depthwise CNN is 256, the output channel is 256, and the kernel size is 15. The Swish

[28] activation function is used. Before each module, Layernorm and residual connections are used to accelerate the model training convergence. The dropout ratio of each layer is 0.1 to improve the model robustness. In the middle layer of the encoder, a phoneme-level CTC loss (weight 0.1) is used to assist the training. In attend, local attention is used to focus on the information output by the encoder. Spell uses LSTM to decode the information, where the input dimension is 1024. Dropout is used during training with a ratio of 0.3. All the experimental results of the present invention are obtained without an external language model and hyperparameter optimization.

[0106] 3.4 Experimental Analysis

[0107] The present invention first verified the proposed Conformer-LAS on the Aishell-1 dataset, as well as the effect of Conformer-LAS-CTC assisted by phoneme-level intermediate CTC loss (weight 0.1), and compared its experimental effects with those of the baseline model and other codec combination models, as shown in Table 1. We used the character error rate CER as the evaluation criterion, and all evaluation results were rounded to two decimal places.

[0108] Table 1 Experimental results of different codecs on Aishell-1

[0109]

[0110] As can be seen from Table 1:

[0111] (1) When the LAS model is used in the decoder, the proposed Conformer-LAS-CTC model has a relative reduction of 19.52% in the character error rate compared to the model with BLSM as the encoder, and a relative reduction of 46.74% compared to the Transformer encoder model.

[0112] (2) The Conformer-LAS-CTC(+Inter CTC) model trained with phoneme-level intermediate CTC loss assistance achieved the best result, with a further 2.11% improvement on the test set compared to the Conformer-LAS-CTC model.

[0113] To better reflect the differences between the models, the present invention selects a loss value every 1000 steps in the training set, and the loss curves of each model on the training set are as Figure 4 shown; the first 80 epochs are selected in the validation set, and the recognition character error rate (CER) curve during the training process is as Figure 5 shown.

[0114] From Figure 4 the loss curve during the training process, it can be seen that in the initial 0 - 10k steps, Conformer-LAS-CTC already shows advantages. Compared with the loss curves of the Transformer-LAS and Conformer-LAS models, its slope is larger and it drops faster. After 10k steps, Conformer-LAS-CTC is more stable than the blstm-LAS model, which means that the Conformer-LAS-CTC model can train the loss value quickly and stably compared to other models. From Figure 5It can be seen from the word error rate curve on the validation set that as the number of iterations increases, the model gradually converges, and the word error rate finally stabilizes within a fixed range. The word error rates of Conformer-LAS and Conformer-LAS-CTC are significantly lower than those of the BLSTM-LAS model and the Transformer-LAS. Among them, Conformer-LAS-CTC uses the Conformer-LAS codec to learn the dependencies between target sequences and adopts CTC to assist in accelerating convergence, enabling it to learn more information on the training set, thus improving the model's generalization performance and accuracy.

[0115] The present invention also compared the proposed model with traditional speech recognition methods and the mainstream end-to-end models in the past two years on Aishell-1, and the results are shown in Table 2.

[0116] Table 2 Experimental results of different acoustic models on Aishell-1

[0117]

[0118] Compared with other end-to-end models in the table, the model proposed by the present invention also further reduces the word error rate, which clearly proves the effectiveness of the proposed Conformer-LAS-CTC model.

[0119] To further verify the performance of the proposed model, we also explored the influence of different decoding layers on the speech recognition effect. By controlling the number of LSTM layers used in the LAS decoder to be 1 layer, 2 layers, and 3 layers respectively, and comparing the obtained experimental results, the results are shown in Table 3.

[0120] Table 3 Experimental results of different decoding levels

[0121]

[0122] It can be seen from the table that as the number of spell layers increases, the word error rate of the speech recognition model on the test set gradually decreases. It can be concluded that more decoder layers will be beneficial to obtaining better recognition effects. The proposed model achieved an error rate of 4.54% when combining 3 decoding layers.

Claims

1. An end-to-end Chinese speech recognition method, characterized in that, The steps are as follows: I. Preprocessing of data For speech data, perform pre-emphasis, framing, windowing, conduct fast Fourier transform, calculate spectral line energy, perform Mel filtering, and take the logarithm to obtain Fbank features; Divide the preprocessed data into a training set and a validation set; II. Establish a hybrid CTC / Attention model based on Conformer The hybrid CTC / Attention model based on Conformer consists of three parts: a shared Conformer encoder, a CTC decoder, and a LAS attention decoder; The described shared Conformer encoder first processes the input using a convolutional subsampling layer, and inputs the data processed by the convolutional subsampling layer into N Conformer encoder blocks. Each Conformer encoder block sequentially includes a feed-forward module, a multi-head self-attention module MHSA, a convolutional module, a feed-forward module, and layer normalization. A residual unit is set after each module in the Conformer encoder. Among them, a half-step residual connection is adopted between the feed-forward module and the multi-head self-attention module, and between the feed-forward module and layer normalization; the multi-head self-attention module includes layer normalization, multi-head self-attention integrating relative sinusoidal position encoding, and dropout; the convolutional module contains a pointwise convolution with an expansion factor of 2, projects the number of channels through a GLU activation layer, and then is followed by a one-dimensional depth convolution, and a Batchnorm and a swish activation layer are connected after the one-dimensional depth convolution; the shared Conformer encoder maps the input frame-level acoustic features x = (x1,...x T ) to a sequence of high-level representations h = (h1, h2,..., h U ); The described LAS attention decoder adopts a two-layer unidirectional LSTM structure and introduces an attention mechanism. The specific decoding process is as follows: local attention is used to focus on the information output by the shared Conformer encoder, and LSTM is used to decode the information. During the output process of each LSTM, the LAS attention decoder combines the already generated text (y1, y2,..., y s-1 ) with the output features h = (h1, h2,..., h U ) of the shared Conformer encoder for attention decoding, and finally generates the target transcription sequence y = (y1, y2,..., y S ), so the probability of obtaining the output sequence y is as follows: At each time step t, the conditional dependence of the output on the encoder features h is computed by an attention mechanism; the attention mechanism is a function of the current decoder hidden state and the encoder output features, and compresses the encoder features into a context vector u through the following mechanism it ; where h i is the output feature of the shared Conformer encoder; vector b a , and matrix W h , W d are all learned parameters; d t represents the hidden state of the decoder at time step t; then softmax is performed on u it to obtain the attention distribution: α t = soft max(u t ) (4) Utilize α it By performing an operation on h i The corresponding context vector is obtained through weighted summation: At each moment, the attention decoder hidden state d for capturing the previous output context t is obtained by the following method: where d t-1 is the previous hidden state, and t-1 is the embedding layer vector learned through y; at time t, the posterior probability of the output y t is as follows: P(y t |h,y < t) = softmax(W s [c t ; d t + b s ) (7) Where W s and b s are learnable parameters; The CTC decoder decodes with the output feature h of the shared Conformer encoder as the input. After passing through the Softmax layer, the output of the CTC decoder is P(q t |h), where q t is the output at time t, and the label sequence l is the sum of all path probabilities: where: Γ(q t ) is a many-to-one mapping of the tag sequence; since there are multiple paths corresponding to the same tag sequence, it is necessary to remove duplicate tags and blank tags in the path; q t ∈A, t = 1, 2,..., T, A is the set of tags with the blank tag "-" added, and the labeled sequence l * with the highest probability in the output sequence is: l * = arg l maxP(l|h)(9) The loss function of the CTC decoder is the sum of the negative log probabilities of all labels, and the CTC network is trained through backpropagation: CTC loss = -logP(l|h) (10) Skip all layers after the intermediate layer during CTC decoder training, and add the intermediate layer phoneme-level CTC loss, i.e., InterCTC loss Induce a sub-model as an auxiliary task; calculate the loss of the sub-model by obtaining the intermediate representation of the CTC decoder. Similar to the complete model of the CTC decoder, the loss function of the sub-model is as follows: Among them, represents the output of the sub-model; The hybrid CTC / Attention model based on Conformer jointly optimizes the model parameters using the CTC decoder and the LAS attention decoder, and at the same time adds the loss of the intermediate layer phoneme-level CTC decoder for regularizing the lower-level parameters. Therefore, the loss function is defined as follows during the training process: T loss = λCTC loss + μInterCTC loss +(1 - λ - μ)Att loss (12) Among them, CTC loss , InterCTC loss , Att loss are the CTC decoder loss, the intermediate layer phoneme-level CTC decoder loss, and the LAS attention decoder loss respectively. λ and μ are two hyperparameters used to measure the weights of the CTC decoder, the intermediate layer phoneme-level CTC decoder, and the LAS attention decoder; During the training process, make the loss decline curve converge to a stable state, end the training, and obtain the final model; III. Train the hybrid CTC / Attention model based on Conformer, use the trained model to validate the validation set, and achieve end-to-end Chinese speech recognition.

Citation Information

Patent Citations

  • Mixed speech recognition system and method based on end-to-end model

    CN113763939A

  • Multi-task training architecture and strategy for attention-based speech recognition system

    US20200135174A1