Transformers dialect speech recognition method based on non-autoregression network
By adopting the Transformers model based on non-autoregressive networks in speech recognition technology, combining feature coding and location coding, and adding Bi-GRU network to the backend, the problem of high recognition error rate in dialect speech recognition is solved, and higher recognition accuracy and speed are achieved.
Patent Information
- Application Number
- CN202510072156.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-06
AI Technical Summary
Existing speech recognition technology is difficult to effectively handle the diversity of dialects, resulting in a high recognition error rate and a complex structure of traditional models, making it difficult to perform joint optimization.
The Transformers model based on non-autoregressive network is adopted, combining feature encoding and position encoding, and semantic signals of speech are extracted through Encoder and Decoder, and a non-autoregressive network (Bi-GRU) is added to the backend to correct identification errors.
It improves the accuracy and recognition rate of dialect speech recognition, reduces the error rate of text output, and the model can better understand context information, which is suitable for recognition of multiple dialects.
Smart Images

Figure CN119943029A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech intelligence technology, and in particular to a dialect speech recognition method based on transformers of a non-autoregressive network. Background Art
[0002] At present, voice intelligent devices are showing a booming trend and have become an indispensable part of people's daily production and life. Voice communication is fast and accurate, and is an important way of communication between people. Therefore, using voice to interact between people and intelligent devices has become an important research content. Traditional automatic speech recognition consists of an acoustic model and a language model. The traditional classic model structure DNN-HMM with a mixed model structure requires a separate training of the Gaussian mixture model in advance, and the pronunciation dictionary also requires the knowledge of linguistic experts to establish. The separate training of the acoustic model and the language model makes it difficult to jointly optimize the model. With the rise of Transformers and its great success in the field of NLP, the present invention introduces transformers. In order to improve the training and prediction rate and reduce resource consumption, the present invention proposes a transformer model based on non-autoregression to speed up the training and reasoning of the model.
[0003] (2) Dialects are diverse and vary greatly in pronunciation, intonation, and vocabulary. Even slight differences exist in the same region. For example, in Guiyang, Guizhou, there are obvious differences in dialects among its subordinate counties and districts. In order to make speech recognition universal, it is urgent to propose a model suitable for dialects.
[0004] (3) Since transformers use global information of the context, time series are added to the feature encoding to make the model have temporal characteristics. The above information can be used to infer the prediction of the current text, making the dialect recognition error rate lower and more in line with the characteristics of the dialect. Summary of the invention
[0005] The purpose of the present invention is to provide a dialect speech recognition method based on transformers of a non-autoregressive network, which solves the problems raised in the background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A dialect speech recognition method based on transformers of a non-autoregressive network includes dialect speech recognition. The dialect speech recognition model is mainly composed of feature coding, position coding, encoder, decoder, and non-autoregressive network. Feature coding is mainly responsible for converting speech information into digital coding information that can be understood by the model. Position coding encodes the position of the speech frame as a specific trigonometric function so that the model can understand the temporal characteristics of the speech. The encoder Encoder and decoder Decoder are used to extract and understand the semantic signals contained in the speech. The non-autoregressive network solves the problem of context understanding and matching of wrong words in the dialect.
[0008] Preferably, the feature coding converts the original waveform signal into a digital signal before the data is input into the model for processing, and processes it into an F-bank acoustic feature containing more original information; first, a pre-emphasis filter is used to amplify the high-frequency signal of the signal to eliminate the vocal cord and lip effects during the phonation process to compensate for the high-frequency components of the speech signal suppressed by the pronunciation system. The filter coefficient is usually within 0.95 to 0.97, and 0.97 is taken in the present invention; second, the pre-emphasized signal is framed, each frame is within 20 to 40 ms, and the present invention stipulates that each frame is 25 ms, 16 kHz The speech length will be divided into 400 sampling points, and the frame shift is set to 15ms, that is, there is a 15ms overlap between frames; third, a window function is used for the information after framing to make the left and right ends of the frame continuous and reduce spectrum leakage; fourth, a Fourier transform is performed on the result of the previous step to convert the time domain signal into the frequency domain; fifth, the frequency domain features are squared and Mel filtered; sixth, the logarithm of the filtered features is taken to obtain the F-bank features; F-bank is more in line with the nature of sound, and has a smaller amount of calculation, higher feature correlation, and more information than MFCC.
[0009] Preferably, due to the special attention mechanism of encoder-decoder, the input global information has no temporal features after entering the encoder-decoder, so the position information of the speech is added as input to the model, and the position information is encoded by trigonometric function position, and the formula is as follows:
[0010]
[0011]
[0012] Where i represents the frame number, D m represents the dimension of the model input, 2j and 2j+1 represent the different position codes corresponding to the feature sequence of the i-th frame input being an odd number or an even number, and 10000 is an empirical value; the position code vector of the i-th frame information is p(i) = [p(i,0), p(i,1), … p(i,D m-1)], trigonometric functions have the properties: sin(a+b)=sin(a)cos(b)+cos(a)sin(b) and con(a+b)=cos(a)cos(b)-sin(a)sin(b), the elements in p(i+n) can be represented by the elements in p(i), so the position codes of each frame of speech signal are related.
[0013] Preferably, the encoder structure passes the speech coding and position coding information into the encoder's Self-Attention, uses the attention mechanism to obtain the degree of attention between features, uses the residual mechanism to add the attention input and output, and its basic idea is that the worst result is no worse than the input; after layer normalization, it is input into the feed-forward neural network Feed-Forward for feature extraction; finally, the residual mechanism and layer normalization are used again for processing;
[0014] Among them, the three matrices Q (query), K (key), and V (value) in the Self-Attention structure are all generated from the same input. First, Q and K are dot-multiplied. To prevent the result from being too large, Scaling is performed, where d k is the dimension of vector K, and then the result is normalized by softmax. Finally, the result of softmax is multiplied by matrix V to get the representation of weighted sum:
[0015]
[0016] The specific calculation process is as follows: First, for each input feature vector, initialize three weight matrices W Q , W K , W V , multiply the feature vectors by the weight matrix respectively to obtain the new vector representation of the feature q1, q2…, k1, k2…, v1, v2…, and represent the new vectors of all features as matrices, namely Q, K, V. The calculation of the i-th feature in the sequence and its attention score is done by q i The dot product is calculated with k1, k2, etc. The attention of all features is QK T; Second, the results are scaled and normalized, which will help the training process to be more stable. Obviously, the current feature has the largest attention (score) with itself, and the largest score is the most relevant or important to itself. Similarly, the more relevant other features are to the current feature, the larger the score. Third, each vector in V is multiplied by the score after softmax in the second step. Its significance is to reduce the attention to irrelevant features while keeping the attention to the current feature unchanged.
[0017] The decoder is similar to the encoder. The Multi-head attention is composed of multiple self-attentions and has H taps:
[0018] Masked Multi-headattention means that mask operations are randomly used for splicing. The masked taps will not participate in the calculation of the current step, that is, their weights will not be updated. It means that when the model is trained at the current step, if the weight of a tap is not updated, it is equivalent to training a sub-network. During the entire training process, many sub-networks will be obtained. Different sub-networks perform well on different samples, that is, the trained network adapts to different types of samples, especially better performance on dialects. At the same time, the use of mask operations also speeds up the training process and saves resources.
[0019] Combining the Encoder with the Decoder is a transformer structure. In the present invention, a multi-layer transformer structure is stacked to achieve the best effect.
[0020] Preferably, a non-autoregressive network is added to the encoder-decoder backend. The non-autoregressive network is composed of a shallow Bi-GRU. The bidirectionally connected Bi-GRU structure can make the latent vector have contextual relevance, so as to better extract the feature vector of the text. Suppose the decoder output contains n vectors {a1, a2, ..., a n}, then the model output probabilities of the forward connection and the reverse connection are:
[0021]
[0022] Bidirectional propagation makes each feature vector have a contextual relationship. At the same time, multiple bidirectional connections can obtain grammatical information of different dimensions. The vector that integrates the contextual features can correct the above recognition errors.
[0023] GRU consists of an update gate and a reset gate. Compared with LSTM, this structure of GRU improves computational efficiency while maintaining the effect. It can retain or discard input information as needed. The function of the update gate is to determine how much information from the previous time step needs to be retained for the hidden state of the current time step. The output value of the update gate is between 0 and 1. The smaller the value, the greater the value of the current input information, and the larger the value, the more past information will be retained. The formula is:
[0024] z t =σ(W z ·[h t-1 , xt ]+b z )
[0025] Among them, W z and b z are weight parameters and bias vectors, h t-1 is the hidden state of the previous time step, x t Represents the current input, σ represents the sigmoid activation function. The main purpose of using the activation function is to prevent the gradient from disappearing or exploding;
[0026] The function of the reset gate is to determine to what extent the hidden state of the previous time step is ignored. The output value of the reset gate is between 0 and 1. When the output value tends to 0, the network forgets more information from the previous step and relies more on the current input. When the output value tends to 1, the network will retain more information from the previous step and rely less on the current input. The formula is:
[0027] r t =σ(W r ·[h t-1 , x t ]+b r )W r and b r are weight parameters and bias vectors;
[0028] The GRU forward propagation execution process is as follows: the first step is to calculate the update gate z t and reset gate r t , their output will control the update of hidden state and the degree of retention of historical information; the second step is to calculate the candidate hidden state, the formula is:
[0029]
[0030] Among them, r t ·h t-1 Indicates the influence of the hidden state of the previous time step and the reset gate control on the hidden state. W and b represent weight and bias respectively. In the third step, the current hidden state is calculated using the update gate. The formula is as follows:
[0031]
[0032] Final state h t Obtained by updating the memory unit output at the previous moment and reset at the current moment;
[0033] As shown above, the execution process of the sequential connection of Bi-GRU is similar to that of the sequential connection. It uses the latter state to update the current state, that is, update gate z. t and reset gate r tUsing the hidden state h of the next step t+1 Calculated, candidate hidden state The hidden state h in the next step t+1 Finally, the order is hidden and reverse hidden state After concatenation and activation function, the output of Bi-GRU is as follows:
[0034]
[0035] y t =σ(W o ·H t ) Among them, concat means concatenation, y t represents the output of Bi-GRU at time t, W o Represents weight.
[0036] Preferably, Layer-Normal means normalizing all features of the sample using the mean and standard deviation of the sample, Softmax is used for classification, and each output value represents the probability value of its corresponding character.
[0037] Preferably, when training the dialect speech recognition model, the model is first pre-trained using a general Chinese speech data set, the data set THCHS30 is 40 hours long, the data set ST-CMDS is 100 hours long, the data set AIShell-1 is 178 hours long, the data set Premewords is 100 hours long, and the data set MagicData is 755 hours long, a total of about 1173 hours, using 80% of the speech data as a training set and the remaining 20% of the data as a test set; secondly, based on the pre-trained model, the model is further trained and fine-tuned using a private data set, which includes Guizhou dialect, Mandarin with Guizhou dialect intonation, and standard Mandarin;
[0038] During the model optimization process, the mini-batch gradient descent method is used to update the model parameters.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. Use a general Chinese speech dataset to pre-train the model so that the model has the ability to recognize Mandarin. Mandarin with Guizhou pronunciation and intonation has the characteristics of standard Mandarin, so the recognition effect of Mandarin with Guizhou pronunciation and intonation is relatively good. Integrating Guizhou dialect into the pre-trained model and aligning the model's capabilities in dialect recognition is one of the strategies for good model performance.
[0041] 2. Considering the excellent performance of transformers in speech models, the present invention introduces the transformer structure. On the basis of giving priority to model effects, the complexity of the model is debugged. Too many layers will affect the recognition rate, and too few layers will result in poor recognition effect. In order to balance performance and rate, the present invention tests the number of stacked layers of transformers. The results show that when the number of stacked layers is 51, the performance and reasoning rate can be better balanced.
[0042] 3. The introduction of position information enables the model to understand the temporal information of the sequence, which is also one of the strategies for better model performance.
[0043] 4. The introduction of the non-autoregressive network greatly reduces the text output error rate of the model for similar speech, and at the same time greatly improves the accuracy of dialect text and dialect words. Since the non-autoregressive network can better understand the contextual information, the model can correct the current output value according to the above information during reasoning.
[0044] 5. The dialect speech recognition technology proposed in the present invention can be extended to other dialect recognition. Other dialects can achieve better results by simply training them on the model proposed in the present invention.
[0045] In summary, the Guizhou dialect speech recognition technology proposed in the present invention has a lower error rate and a faster recognition rate in the Guizhou dialect, and can quickly adapt to the speech recognition of other dialects without changing the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a structural diagram of the dialect speech recognition model of the present invention;
[0047] Figure 2 This is the structure diagram of the Encoder of the present invention;
[0048] Figure 3 This is the Self-Attention structure diagram of the present invention;
[0049] Figure 4 This is a structural diagram of the Decoder of the present invention;
[0050] Figure 5 This is the Bi-GRU structure diagram of the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] See also Figure 1-5 A dialect speech recognition method based on transformers of non-autoregressive networks includes dialect speech recognition. The dialect speech recognition model is mainly composed of feature coding, position coding, encoder, decoder, and non-autoregressive networks. Feature coding is mainly responsible for converting speech information into digital coding information that can be understood by the model. Position coding encodes the position of the speech frame as a specific trigonometric function so that the model can understand the temporal characteristics of the speech. The encoder Encoder and decoder Decoder are used to extract and understand the semantic signals contained in the speech. The non-autoregressive network solves the context understanding and matching of dialect typos.
[0053] Specifically, the feature coding converts the original waveform signal into a digital signal before the data is input into the model for processing, and processes it into an F-bank acoustic feature containing more original information; first, a pre-emphasis filter is used to amplify the high-frequency signal of the signal to eliminate the vocal cord and lip effects during the phonation process to compensate for the high-frequency components of the speech signal suppressed by the pronunciation system. The filter coefficient is usually within 0.95 to 0.97, and 0.97 is taken in the present invention; second, the pre-emphasized signal is framed, each frame is within 20 to 40 ms, and the present invention stipulates that each frame is 25 ms, 16 kHz The speech length will be divided into 400 sampling points, and the frame shift is set to 15ms, that is, there is a 15ms overlap between frames; third, a window function is used for the information after framing to make the left and right ends of the frame continuous and reduce spectrum leakage; fourth, a Fourier transform is performed on the result of the previous step to convert the time domain signal into the frequency domain; fifth, the frequency domain features are squared and Mel filtered; sixth, the logarithm of the filtered features is taken to obtain the F-bank features; F-bank is more in line with the nature of sound, and has a smaller amount of calculation, higher feature correlation, and more information than MFCC.
[0054] Specifically, due to the special attention mechanism of encoder-decoder, the global information input has no temporal features after entering the encoder-decoder. Therefore, the position information of the speech is added as input to the model, and the position information is encoded by trigonometric function. The formula is as follows:
[0055]
[0056] Where i represents the frame number, D m represents the dimension of the model input, 2j and 2j+1 represent the different position codes corresponding to the feature sequence of the i-th frame input being an odd number or an even number, and 10000 is an empirical value; the position code vector of the i-th frame information is p(i) = [p(i,0), p(i,1), … p(i,D m-1)], trigonometric functions have the properties: sin(a+b)=sin(a)cos(b)+cos(a)sin(b) and con(a+b)=cos(a)cos(b)-sin(a)sin(b), the elements in p(i+n) can be represented by the elements in p(i), so the position codes of each frame of speech signal are related.
[0057] Specifically, the encoder structure passes the speech coding and position coding information into the encoder's Self-Attention, uses the attention mechanism to obtain the degree of attention between features, and uses the residual mechanism to add the attention input and output. The basic idea is that the result is at least as good as the input; after layer normalization, it is input into the feed-forward neural network Feed-Forward for feature extraction; and finally, the residual mechanism and layer normalization are used again for processing;
[0058] Among them, the three matrices Q (query), K (key), and V (value) in the Self-Attention structure are all generated from the same input. First, Q and K are dot-multiplied. To prevent the result from being too large, Scaling is performed, where d k is the dimension of vector K, and then the result is normalized by softmax. Finally, the result of softmax is multiplied by matrix V to get the representation of weighted sum:
[0059]
[0060] The specific calculation process is as follows: First, for each input feature vector, initialize three weight matrices W Q , W K , W V , multiply the feature vectors by the weight matrix respectively to obtain the new vector representation of the feature q1, q2…, k1, k2…, v1, v2…, and represent the new vectors of all features as matrices, namely Q, K, V. The calculation of the i-th feature in the sequence and its attention score is done by q i The dot product is calculated with k1, k2, etc. The attention of all features is QK T; Second, the results are scaled and normalized, which will help the training process to be more stable. Obviously, the current feature has the largest attention (score) with itself, and the largest score is the most relevant or important to itself. Similarly, the more relevant other features are to the current feature, the larger the score. Third, each vector in V is multiplied by the score after softmax in the second step. Its significance is to reduce the attention to irrelevant features while keeping the attention to the current feature unchanged.
[0061] The decoder is similar to the encoder. The Multi-head attention is composed of multiple self-attentions and has H taps:
[0062] Masked Multi-headattention means that mask operations are randomly used for splicing. The masked taps will not participate in the calculation of the current step, that is, their weights will not be updated. It means that when the model is trained at the current step, if the weight of a tap is not updated, it is equivalent to training a sub-network. During the entire training process, many sub-networks will be obtained. Different sub-networks perform well on different samples, that is, the trained network adapts to different types of samples, especially better performance on dialects. At the same time, the use of mask operations also speeds up the training process and saves resources.
[0063] Combining the Encoder with the Decoder is a transformer structure. In the present invention, a multi-layer transformer structure is stacked to achieve the best effect.
[0064] Specifically, a non-autoregressive network is added to the encoder-decoder backend. The non-autoregressive network consists of a shallow Bi-GRU. The bidirectionally connected Bi-GRU structure can make the latent vector have contextual relevance, so as to better extract the feature vector of the text. Suppose the decoder output contains n vectors {a1, a2, ..., a n}, then the model output probabilities of the forward connection and the reverse connection are:
[0065]
[0066] Bidirectional propagation makes each feature vector have a contextual relationship. At the same time, multiple bidirectional connections can obtain grammatical information of different dimensions. The vector that integrates the contextual features can correct the above recognition errors.
[0067] GRU consists of an update gate and a reset gate. Compared with LSTM, this structure of GRU improves computational efficiency while maintaining the effect. It can retain or discard input information as needed. The function of the update gate is to determine how much information from the previous time step needs to be retained for the hidden state of the current time step. The output value of the update gate is between 0 and 1. The smaller the value, the greater the value of the current input information, and the larger the value, the more past information will be retained. The formula is:
[0068] z t =σ(W z ·[h t-1 , xt ]+b z )
[0069] Among them, W z and b z are weight parameters and bias vectors, h t-1 is the hidden state of the previous time step, x t Represents the current input, σ represents the sigmoid activation function. The main purpose of using the activation function is to prevent the gradient from disappearing or exploding;
[0070] The function of the reset gate is to determine to what extent the hidden state of the previous time step is ignored. The output value of the reset gate is between 0 and 1. When the output value tends to 0, the network forgets more information from the previous step and relies more on the current input. When the output value tends to 1, the network will retain more information from the previous step and rely less on the current input. The formula is:
[0071] r t =σ(W r ·[h t-1 , x t ]+b r )
[0072] Among them, W r and b r are weight parameters and bias vectors;
[0073] The GRU forward propagation execution process is as follows: the first step is to calculate the update gate z t and reset gate r t , their output will control the update of hidden state and the degree of retention of historical information; the second step is to calculate the candidate hidden state, the formula is:
[0074]
[0075] Among them, r t ·h t-1 Indicates the influence of the hidden state of the previous time step and the reset gate control on the hidden state. W and b represent weight and bias respectively. In the third step, the current hidden state is calculated using the update gate. The formula is as follows:
[0076]
[0077] Final state h t Obtained by updating the memory unit output at the previous moment and reset at the current moment;
[0078] As shown above, the execution process of the sequential connection of Bi-GRU is similar to that of the sequential connection. It uses the latter state to update the current state, that is, update gate z. tand reset gate r t Using the hidden state h of the next step t+1 Calculated, candidate hidden state The hidden state h in the next step t+1 Finally, the order is hidden and reverse hidden state After concatenation and activation function, the output of Bi-GRU is as follows:
[0079]
[0080] y t =σ(W o ·H t )
[0081] Among them, concat means splicing, y t represents the output of Bi-GRU at time t, W o Represents weight.
[0082] Specifically, Layer-Normal means normalizing all features of the sample using the mean and standard deviation of the sample. Softmax is used for classification, and each output value represents the probability value of its corresponding character.
[0083] Specifically, when training the dialect speech recognition model, the model is first pre-trained using a general Chinese speech dataset, with a duration of 40 hours for the THCHS30 dataset, 100 hours for the ST-CMDS dataset, 178 hours for the AIShell-1 dataset, 100 hours for the Premewords dataset, and 755 hours for the MagicData dataset, for a total of about 1,173 hours. 80% of the speech data is used as the training set, and the remaining 20% of the data is used as the test set. Secondly, based on the pre-trained model, the model is further trained and fine-tuned using a private dataset, which includes Guizhou dialect, Mandarin with Guizhou dialect intonation, and standard Mandarin.
[0084] During the model optimization process, the mini-batch gradient descent method is used to update the model parameters.
[0085] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A dialect speech recognition method based on transformers of non-autoregressive networks, characterized by: Including dialect speech recognition, the dialect speech recognition model is mainly composed of feature encoding, position encoding, encoder, decoder, and non-autoregressive network. Feature encoding is mainly responsible for converting speech information into digital coding information that the model can understand. Position encoding encodes the position of the speech frame as a specific trigonometric function so that the model can understand the temporal characteristics of the speech. The encoder Encoder and decoder Decoder are used to extract and understand the semantic signals contained in the speech. The non-autoregressive network solves the problem of context understanding and matching of dialect typos.
2. The dialect speech recognition method based on transformers of non-autoregressive networks according to claim 1, characterized in that: Feature coding converts the original waveform signal into a digital signal before inputting the data into the model for processing, and processes it into an F-bank acoustic feature containing more original information; first, a pre-emphasis filter is used to amplify the high-frequency signal of the signal to eliminate the vocal cord and lip effects during the phonation process to compensate for the high-frequency components of the speech signal suppressed by the pronunciation system. The filter coefficient is usually within 0.95 to 0.97, and is taken as 0.97 in the present invention; second, the pre-emphasized signal is framed, and each frame is within 20 to 40 ms. The present invention stipulates that each frame is 25 ms, and the 16 kHz speech The length will be divided into 400 sampling points, and the frame shift is set to 15ms, that is, there is a 15ms overlap between frames; third, use the window function for the information after framing to make the left and right ends of the frame continuous and reduce spectrum leakage; fourth, perform Fourier transform on the result of the previous step to convert the time domain signal into the frequency domain; fifth, square the frequency domain features and perform Mel filtering; sixth, take the logarithm of the filtered features to obtain F-bank features; F-bank is more in line with the nature of sound, and has smaller calculation amount, higher feature correlation, and more information than MFCC.
3. The dialect speech recognition method based on transformers of non-autoregressive networks according to claim 1, characterized in that: Due to the special attention mechanism of encoder-decoder, the global information input has no temporal features after entering encoder-decoder. Therefore, the position information of the speech is added as input to the model, and the position information is encoded by trigonometric function. The formula is as follows: Where i represents the frame number, D m represents the dimension of the model input, 2j, 2j+1 represent the different position codes corresponding to the feature sequence of the i-th frame input being an odd number or an even number, and 10000 is an empirical value; the position code vector of the i-th frame information is p(i) = [p(i,0), p(i,1), … p(i,D m -1)], trigonometric functions have the properties: sin(a+b)=sin(a)cos(b)+cos(a)sin(b) and con(a+b)=cos(a)cos(b)-sin(a)sin(b), the elements in p(i+n) can be represented by the elements in p(i), so the position codes of each frame of speech signal are related.
4. The dialect speech recognition method based on transformers of non-autoregressive networks according to claim 1, characterized in that: The encoder structure passes the speech coding and position coding information into the encoder's Self-Attention, uses the attention mechanism to obtain the degree of attention between features, and uses the residual mechanism to add the attention input and output. The basic idea is that the result is at least as good as the input; after layer normalization, it is input into the feed-forward neural network Feed-Forward for feature extraction; finally, the residual mechanism and layer normalization are used again for processing; Among them, the three matrices Q (query), K (key), and V (value) in the Self-Attention structure are all generated from the same input. First, Q and K are dot-multiplied. To prevent the result from being too large, Scaling is performed, where d k is the dimension of vector K, and then the result is normalized by softmax. Finally, the result of softmax is multiplied by matrix V to get the representation of weighted sum: The specific calculation process is as follows: First, for each input feature vector, initialize three weight matrices W Q , W K , W V , multiply the feature vectors by the weight matrix respectively to obtain the new vector representation of the feature q1, q2…, k1, k2…, v1, v2…, and represent the new vectors of all features as matrices, namely Q, K, V. The calculation of the i-th feature in the sequence and its attention score is done by q i The dot product is calculated with k1, k2, etc. The attention of all features is QK T; Second, the results are scaled and normalized, which will help the training process to be more stable. Obviously, the current feature has the largest attention (score) with itself, and the largest score is the most relevant or important to itself. Similarly, the more relevant other features are to the current feature, the larger the score. Third, each vector in V is multiplied by the score after softmax in the second step. Its significance is to reduce the attention to irrelevant features while keeping the attention to the current feature unchanged. The decoder is similar to the encoder. The Multi-head attention is composed of multiple self-attentions and has H taps: Masked Multi-head attention means that mask operation is randomly used for splicing. The masked taps will not participate in the calculation of the current step, that is, their weights will not be updated. It means that when the model is trained at the current step, if the weight of a tap is not updated, it is equivalent to training a sub-network. During the whole training process, many sub-networks will be obtained. Different sub-networks perform well on different samples. That is, the trained network adapts to different types of samples, especially better in dialects. At the same time, the use of mask operation also speeds up the training process and saves resources. Combining the Encoder with the Decoder is a transformer structure. In the present invention, a multi-layer transformer structure is stacked to achieve the best effect.
5. The dialect speech recognition method based on transformers of non-autoregressive networks according to claim 1, characterized in that: A non-autoregressive network is added to the encoder-decoder backend. The non-autoregressive network consists of a shallow Bi-GRU. The bidirectionally connected Bi-GRU structure can make the latent vector have contextual relevance, so as to better extract the feature vector of the text. Suppose the decoder output contains n vectors {a1, a2, ..., a n }, then the model output probabilities of the forward connection and the reverse connection are: Bidirectional propagation makes each feature vector have a contextual relationship. At the same time, multiple bidirectional connections can obtain grammatical information of different dimensions. The vector that integrates the contextual features can correct the above recognition errors. GRU consists of an update gate and a reset gate. Compared with LSTM, this structure of GRU improves computational efficiency while maintaining the effect. It can retain or discard input information as needed. The function of the update gate is to determine how much information from the previous time step needs to be retained for the hidden state of the current time step. The output value of the update gate is between 0 and 1. The smaller the value, the greater the value of the current input information, and the larger the value, the more past information will be retained. The formula is: z t =σ(W z ·[h t-1 ,x t ]+b z ) Among them, W z and b z are weight parameters and bias vectors, h t-1 is the hidden state of the previous time step, x t Represents the current input, σ represents the sigmoid activation function. The main purpose of using the activation function is to prevent the gradient from disappearing or exploding; The function of the reset gate is to determine to what extent the hidden state of the previous time step is ignored. The output value of the reset gate is between 0 and 1. When the output value tends to 0, the network forgets more information from the previous step and relies more on the current input. When the output value tends to 1, the network will retain more information from the previous step and rely less on the current input. The formula is: r t =σ(W r ·[h t-1 ,x t ]+b r ) Among them, W r and b r are weight parameters and bias vectors; The GRU forward propagation execution process is as follows: the first step is to calculate the update gate z t and reset gate r t , their output will control the update of hidden state and the degree of retention of historical information; the second step is to calculate the candidate hidden state, the formula is: Among them, r t ·h t-1 Indicates the influence of the hidden state of the previous time step and the reset gate control on the hidden state. W and b represent weight and bias respectively. In the third step, the current hidden state is calculated using the update gate. The formula is as follows: Final state h t Obtained by updating the memory unit output at the previous moment and reset at the current moment; As shown above, the execution process of the sequential connection of Bi-GRU is similar to that of the sequential connection. It uses the latter state to update the current state, that is, update gate z. t and reset gate r t Using the hidden state h of the next step t+1 Calculated, candidate hidden state The hidden state h in the next step t+1 Finally, the order is hidden and reverse hidden state After concatenation and activation function, the output of Bi-GRU is as follows: y t =σ(W o ·H t ) Among them, concat means splicing, y t represents the output of Bi-GRU at time t, W o Represents weight.
6. The dialect speech recognition method based on transformers of non-autoregressive networks according to claim 1, characterized in that: Layer-Normal means normalizing all features of the sample using the mean and standard deviation of the sample. Softmax is used for classification, and each output value represents the probability value of its corresponding character.
7. The dialect speech recognition method based on transformers of non-autoregressive networks according to claim 1, characterized in that: When training the dialect speech recognition model, we first pre-trained the model using a general Chinese speech dataset, including 40 hours of THCHS30, 100 hours of ST-CMDS, 178 hours of AIShell-1, 100 hours of Premewords, and 755 hours of MagicData, for a total of about 1,173 hours. We used 80% of the speech data as the training set and the remaining 20% of the data as the test set. We then used a private dataset based on the pre-trained model to continue training and fine-tuning the model. The private dataset included Guizhou dialect, Mandarin with Guizhou dialect intonation, and standard Mandarin. During the model optimization process, the mini-batch gradient descent method is used to update the model parameters.