Speech Adaptive Recognition Method, System, Device and Storage Medium
By introducing attention-based gating scaling adaptive layer in the end-to-end voice recognition system based on CTC, the problem of data sparsity and low adaptability of the speaker adaptive and complex structure network in the end-to-end voice recognition system is solved, and a high-accuracy speech recognition and simplified adaptive process are achieved.
Patent Information
- Application Number
- CN202111482314.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-06
AI Technical Summary
The prior art realizes speaker adaptation in end-to-end speech recognition systems with challenges of data sparsity, low adaptability of complex structure networks, and acoustic and language model adaptation.
By introducing an attention-based gated scaling adaptive layer in the CTC-based end-to-end speech recognition system, the auxiliary network learns distinctive speaker characteristics through the self-attention mechanism and reweights the hidden layer activation output of the main network to achieve online speaker adaptive.
This method effectively reduces the mismatch between the tester and the training model, improves the accuracy of speech recognition, simplifies the adaptive process, and realizes effective adaptation to complex structural networks.
Smart Images

Figure CN114187900B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing technology, and in particular to a speech recognition method, system, device and storage medium. Background Art
[0002] In recent years, with the widespread application of neural networks in the field of speech recognition, the performance of speech recognition systems has been significantly improved. At present, there are two main mainstream speech recognition systems, one is the HMM-based speech recognition system (Graves A, Fernández S, Gomez F, et al. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks [C] / / Proceedings of the 23rd international conference on Machine learning. ACM, 2006: 369-376.), and the other is the end-to-end speech recognition system (Maas A, Xie Z, Jurafsky D, et al. Lexicon-free conversational speech recognition with neural networks [C] / / Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015: 345-354.). Compared with the HMM-based speech recognition system, the end-to-end speech recognition system has a simpler structure. It directly converts the input speech feature sequence into a text sequence through a neural network. It does not require a pronunciation dictionary, decision tree, or word-level annotation alignment information of the HMM system. Due to its simple implementation and excellent performance, it has become a hot topic in current research.
[0003] The first implementation of end-to-end speech recognition was when Alex Graves of Google and Navdeep Jaitly of the University of Toronto introduced the Connectionist Temporal Classification (CTC) criterion into the speech recognition system (GRAVES A, JAITLY N. Towards end-to-end speech recognition with recurrent neural networks [C] / / International conference on machine learning. PMLR, 2014: 1764-1772.). CTC is essentially a loss function, but it solves the hard alignment problem when calculating the loss. It was originally proposed to solve sequence-to-sequence prediction tasks (GRAVES A, S, GOMEZ F, et al. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks [C] / / Proceedings of the 23rd international conference on Machine learning. 2006: 369-376.). Speech recognition is a typical prediction task from speech sequence to text sequence. The introduction of CTC criterion successfully realizes the process of directly mapping input speech to text label. In combination with RNN or Convolutional Neural Network (CNN) to model temporal information, CTC criterion is widely used in speech recognition system (LI J, YE G, DAS A, et al. Advancing acoustic-to-word CTC model [C] / / 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018: 5794-5798.).
[0004] Speaker adaptation technology in speech recognition is used to solve the mismatch problem caused by speaker differences in training and test environments. Since different speakers have different pronunciation characteristics such as accent, speaking speed, volume, and intonation, when the speech recognition model encounters a speaker that has not appeared in the training set, the performance of the recognition system will decline. The greater the difference between the speakers in the test environment and the training environment, the more serious the decline in recognition performance. Speaker adaptation technology can reduce this difference to a certain extent and improve speech recognition performance in mismatched environments. In principle, speaker adaptation technology can be divided into two categories: model space adaptation and feature space adaptation. Model space adaptation transforms the acoustic model to match the test environment, while feature space adaptation transforms the model input features to match the training model.
[0005] In many cases, there is no strict distinction between model domain adaptation and feature domain adaptation, because the network used to transform features can also be regarded as part of the entire acoustic model. The most fundamental difference between different speaker adaptation technologies lies in the specific technical implementation. From the perspective of technical implementation, speaker adaptation can be divided into re-estimation adaptation and adaptive training. Figure 1 As shown, there are two different adaptation strategies, the left part is re-evaluation adaptation, and the right part is adaptive training.
[0006] 1) Re-estimation adaptation uses a small amount of target speaker adaptation data to estimate the adaptation parameters from the trained speaker-independent (SI) model, thereby converting the speaker-independent model into a speaker-dependent (SD) model. This method is a two-step training, that is, first train the SI model, and then retrain the SD model with the adaptation data. The training data can be labeled or unlabeled, corresponding to supervised adaptation and unsupervised adaptation, respectively. For unsupervised adaptation, it is necessary to first decode the adaptation data to obtain pseudo labels, and then use the pseudo labels to retrain the SI model.
[0007] 2) Adaptive training is a one-time training, which directly models the speaker acoustic variables during model training. Usually, the environment contained in the training data is more complex, in order to make the model more generalizable. Therefore, there are many differences within the training data itself. Adaptive training technology uses training data to directly model acoustic variables such as speakers during model training, so that the features learned by the network show stronger robustness to different acoustic variables, thereby reducing the impact of the adaptation of the training and test environments on performance.
[0008] Research on adaptive training can be summarized as follows:
[0009] 1) Adaptive methods based on speaker auxiliary features. Such methods provide features with the ability to represent speaker information to the SI model during training and testing, enabling the model to learn supplementary information related to the speaker, thereby eliminating the differences caused by speaker mismatch. Adaptive training methods based on i-vector or bottleneck vectors are typical representatives of such methods. The i-vector or bottleneck vector of the speaker is obtained using a pre-trained speaker recognition model, and then concatenated with the acoustic features of the corresponding speaker and sent to the input of the network (MIAO Y, ZHANG H, METZE F. Speaker adaptive training of deep neural network acoustic models using i-vectors[J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2015, 23(11): 1938-1949.). Another type of method uses a control network to generate SD parameters with speaker embeddings as the input. The control network is shared among all speakers (CUI X, GOEL V, SAON G. Embedding-based speaker adaptive training of deep neural networks[J]. arXiv preprint arXiv:1710.06937, 2017.), thereby enhancing the role of speaker auxiliary features in the network. The method based on multi-basis fusion provides several groups of normalization transformations for each layer of the network, and each group of normalization transformations corresponds to a speaker. Then, a set of difference weights is used to fuse these transformations as the SD parameters, and the difference weights and transformation parameters are learned together (TAN T, QIAN Y, YIN M, et al. Cluster adaptive training for deep neural network[C] / / 2015 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2015: 4325-4329.).
[0010] 2) Adaptive methods based on multi-task learning. Similarly, multi-task learning enables the model to learn supplementary information related to the speaker, thereby eliminating the differences caused by speaker mismatch. Similar to the work in domain adaptation (MENG Z, LI J, GONG Y, et al. Adversarial teacher-student learning for unsupervised domain adaptation [C] / / 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018: 5949-5953.), an intuitive method is to use multi-task learning to put the acoustic network and the speaker classification network together for joint optimization (MENG Z, LI J, CHEN Z, et al. Speaker-invariant training via adversarial learning [C] / / 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018: 5969-5973.).
[0011] 3) Online adaptive methods based on auxiliary networks. This type of method uses auxiliary networks to directly model speaker acoustic variables during acoustic model training without the need for any additional adaptive data or speaker auxiliary features. Unlike adaptive training methods based on multi-task learning, the auxiliary network in this type of method is retained as part of the acoustic model during the test phase. The auxiliary network can dynamically generate adaptive parameters based on the input, so that the model can adapt to different test environments. The online adaptive method based on the auxiliary network only requires a single training to obtain an adaptive model, which greatly simplifies the adaptive process and has gradually attracted the attention of researchers in recent years. Typical work includes sequence summary neural network (SSNN), dynamic layer normalization (DLN) and learning hidden unit contribution (LHUC) to re-weight the activation of each hidden unit, and then use adaptive data to directly learn the weight coefficient of each unit; the above three parts of the technology correspond to the literature 1 ( K,WATANABE S, K, et al. Sequence summarizing neural network for speaker adaptation[C] / / 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2016: 5315-5319.), Document 2 (KIM T, SONG I, BENGIO Y. Dynamic layer normalization for adaptive neural acoustic modeling in speech recognition[J]. arXiv preprint arXiv:1707.06065, 2017.), Document 3 (SWIETOJANSKI P, LI J, RENALS S. Learning hidden unit contributions for unsupervised acoustic model adaptation[J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2016, 24(8): 1450-1463.).
[0012] Although so much progress has been made in speaker adaptation research, the above adaptation methods are all for HMM-based hybrid recognition models, and there is still a lack of research on speaker adaptation in end-to-end speech recognition systems. Speaker adaptation in end-to-end speech recognition systems faces three main challenges. First, end-to-end recognition systems usually directly use characters or words as modeling units, and the number of output units is usually in the tens of thousands. In this case, speaker adaptation faces a serious data sparsity problem. Because in the limited adaptation data, most modeling units do not appear. Therefore, it is difficult to ensure that the model learns the differences of speakers during the adaptation process instead of overfitting the limited word targets. Second, when the acoustic modeling network is transformed from DNN to recurrent neural network (RNN) or long short-term memory network (LSTM) that is more suitable for time series modeling in end-to-end systems, the difficulty of improving the adaptation performance of such a complex structure network is further increased. For example, some researchers only reported a relative performance improvement of about 4% (MIAO Y, METZE F. On speaker adaptation of long short-term memory recurrent neural networks [C] / / Sixteenth Annual Conference of the International Speech Communication Association. 2015.). This may be partly because the recurrent topology of LSTM makes it more difficult to effectively capture and integrate speaker-related features than DNN. In addition, in hybrid systems, only the adaptation of the acoustic model needs to be considered, because the language model is a separate model and is usually based on probabilistic statistics. In end-to-end speech recognition systems, the acoustic model and the language model are integrated together. Therefore, the adaptation of the end-to-end model is expected to alleviate the mismatch in acoustic and language conditions at the same time, making it more challenging than the adaptation of the hybrid model. Summary of the invention
[0013] The purpose of the present invention is to provide a speech adaptive recognition method, system, device and storage medium, which aims at improving the recognition rate by studying speaker adaptation technology in an end-to-end recognition framework based on CTC.
[0014] The objective of the present invention is achieved through the following technical solutions:
[0015] A speech adaptive recognition method, comprising:
[0016] The input speech sequence is encoded into a deep feature sequence using an acoustic model, and blank labels are added to the relevant dictionary during the encoding process. The acoustic model is an end-to-end model based on CTC;
[0017] Applying the CTC criterion to the deep feature sequence, converting the deep feature sequence into a probability distribution sequence, during which each deep feature is activated through several hidden layers of the acoustic model, and at least one hidden layer generates a scaling transformation vector of the corresponding deep feature through a corresponding attention-based gated scaling adaptive layer, and reweights the activation output of the corresponding hidden layer using the scaling transformation vector; wherein the attention-based gated scaling adaptive layer is used to capture relevant information in the deep features of sentence-level acoustic modeling, so that each frame of the speech sequence contains information of all other frames;
[0018] During speech recognition, the probability distribution sequence is used to calculate the conditional probability of any path formed in the dictionary elements after adding blank labels, and the paths are aggregated according to the conditional probabilities, the blank labels in the paths are deleted, and the label sequence obtained by merging consecutive repeated labels is the recognition result.
[0019] A speech adaptive recognition system comprises: an acoustic model, and an attention-based gated scaling adaptive network composed of multiple attention-based gated scaling adaptive layers; the acoustic model and the attention-based gated scaling adaptive network realize speech recognition through the aforementioned method.
[0020] A processing device, comprising: one or more processors; a memory for storing one or more programs;
[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0022] A readable storage medium stores a computer program, which implements the above method when the computer program is executed by a processor.
[0023] It can be seen from the technical solution provided by the present invention that the problem of online speaker adaptation in speech recognition acoustic modeling is solved, and adaptive learning of the main network (i.e., the acoustic model) is achieved by introducing an auxiliary network (i.e., a network composed of multiple attention-based gated scaling adaptive layers). At the same time, the auxiliary network adopts a self-attention mechanism to learn more distinctive speaker personality characteristics, thereby improving the recognition rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0025] Figure 1Schematic diagram of two different adaptive strategies provided as background technology of the present invention.
[0026] Figure 2 A flowchart of a speech adaptive recognition method provided by an embodiment of the present invention;
[0027] Figure 3 A schematic diagram of path aggregation provided by an embodiment of the present invention;
[0028] Figure 4 AGS adaptive network schematic diagram provided by an embodiment of the present invention;
[0029] Figure 5 A schematic diagram of the structure of the AGS adaptive layer and the generation of a scaling transformation vector provided by an embodiment of the present invention;
[0030] Figure 6 A schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0032] First, the terms that may be used in this article are explained as follows:
[0033] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.
[0034] The following is a detailed description of the speech adaptive recognition solution provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professionals in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer shall be followed.
[0035] Embodiment 1
[0036] like Figure 2 As shown, a speech adaptive recognition method mainly includes the following steps:
[0037] Step 1: Use the acoustic model to encode the input speech sequence into a deep feature sequence.
[0038] In the embodiment of the present invention, the acoustic model is an end-to-end model based on CTC.
[0039] Step 2: Apply the CTC criterion to the deep feature sequence and convert the deep feature sequence into a probability distribution sequence. During the conversion process, each deep feature is activated through several hidden layers of the acoustic model, and at least one hidden layer generates a scaling transformation vector of the corresponding deep feature through the corresponding attention-based gated scaling adaptive layer, and the scaling transformation vector is used to re-weight the activation output of the corresponding hidden layer.
[0040] In an embodiment of the present invention, the attention-based gated scaling adaptive layer is used to capture relevant information in the deep features of sentence-level acoustic modeling, so that each frame of the speech sequence contains information of all other frames.
[0041] Step 3: During speech recognition, the probability distribution sequence is used to calculate the conditional probability of any path formed in the dictionary elements after adding blank labels, and the paths are aggregated according to the conditional probabilities (that is, the most likely path is selected for aggregation), the blank labels in the paths are deleted, and the label sequence obtained by merging consecutive repeated labels is the recognition result.
[0042] The above scheme of the embodiment of the present invention is a speech recognition method with gated scaling adaptive using an attention mechanism, which is used in an end-to-end acoustic model based on CTC. It is an improvement on the existing acoustic model, which can reduce the mismatch between the test person and the training model and improve the accuracy of speech recognition. For ease of understanding, the following is a detailed introduction to the network consisting of an end-to-end acoustic model based on CTC and a gated scaling adaptive layer based on attention.
[0043] 1. End-to-end acoustic model based on CTC.
[0044] Linked Temporal Classification (CTC) is an end-to-end technology for sequence labeling problems, which directly converts input sequences with time information into shorter label sequences by removing time and alignment information. CTC is mainly used to handle time series classification tasks, especially when the alignment results of the input signal and the target label are unknown. When applied to acoustic modeling, CTC can automatically learn the alignment between the input speech frame sequence and its label sequence (such as phonemes or words) without using frame-level alignment information. The main principle is to introduce additional blank labels during alignment, delete blank labels during decoding, and merge consecutive repeated labels to obtain a unique corresponding sequence. The criterion for neural network training through CTC technology is called the CTC criterion. The CTC criterion avoids a series of operations such as forced alignment and state binding in traditional HMM-based acoustic models, greatly simplifying acoustic modeling in speech recognition.
[0045] The main recognition process of the acoustic model is as follows:
[0046] The input speech sequence is recorded as X = {x 1 ,x 2 ,...,x T}, where x t represents the t-th frame of speech data, each frame corresponds to a time, t = 1, 2, ..., T, T represents the total number of frames, that is, the length of the input speech sequence; the acoustic model encodes the input speech sequence X into a deep feature sequence F = {f 1 ,f 2 ,...,f T}, deep feature f t is a vector of dimension |V|+1, where |V| represents the number of modeling units in the dictionary V, which is equal to the number of labels, and 1 represents an additional dimension, which corresponds to the blank label.
[0047] The general conventional criterion is that one frame feature corresponds to only one probability, while the CTC criterion considers all frames together and then forms a probability for a single frame. In the embodiment of the present invention, the CTC criterion is applied to the deep feature sequence F, and the deep feature sequence F is converted into a probability distribution sequence Y={y 1 ,y 2 ,...,y T}, where y t is a vector of dimension |V|+1, Represents the t-th frame of speech data x t The probability of belonging to label i, Represents the t-th frame of speech data x tThe probability of belonging to a blank label, i = 1, 2, ..., |V| + 1. The above conversion process will re-weight the activation output of the hidden layer in combination with the scaling transformation vector calculated by the attention-based gated scaling adaptation layer, and the relevant operations will be introduced later.
[0048] After adding the blank tag, the dictionary is recorded as Indicates that the definition is in the dictionary The set of all sequences of length T on The elements in constitute the path. Specifically, each moment corresponds to an element. Any combination of elements constitutes a path. Definition: For a given speech sequence X, the set The conditional probability of any path π in is calculated as follows:
[0049]
[0050] Among them, π t represents the tth label in the path π, that is, the label corresponding to time t.
[0051] Through the above calculation, the speech sequence {x 1 ,x 2 ,...,x T} is mapped to a path π of the same length. During the mapping process, each frame input feature x t are mapped to a specific label π t . Therefore, it can be considered that the mapping process from the input sequence to the path is an alignment process. In addition, from the calculation of conditional probability, it can be seen that there is an important assumption: the independence assumption, that is, each element in the output sequence is independent of each other. The label selected as the output at any time step will not affect the label distribution of other time steps. However, in the process of encoding features in the acoustic model, each moment is affected by both past and future information.
[0052] In the process of calculating the path probability, it can be found that the length of the path is the same as the length of the input sequence, which is inconsistent with the actual situation. Usually, the length of the transcription is much smaller than the length of the corresponding speech sequence. Therefore, in the CTC criterion, a many-to-one, long-to-short mapping is used to aggregate multiple paths to obtain a final short label sequence.
[0053] Let V ≤TIt represents the set of all label sequences with length less than or equal to T defined on the dictionary V. It can be obtained by calculating the probability, because each frame of T frames corresponds to a word in the dictionary (i.e., label). According to the probability, a frame can have multiple words, such as 90% corresponding to a, 5% corresponding to b... and so on. Path aggregation is defined as a mapping function B, which converts The paths in V are mapped to ≤T In the embodiment of the present invention, the path corresponding to the maximum conditional probability is selected for path aggregation according to the conditional probability calculated above. Path aggregation mainly includes two operations: first, merging the same continuous labels. If there are continuous identical labels in the path, they are merged and only one of them is retained.
[0054] like Figure 3 As shown, it is an example of path aggregation, which shows the process of merging multiple possible paths to generate recognition results. The acoustic model in the embodiment of the present invention is an end-to-end model, that is, a model directly from speech data to text. The input of the neural network is the characteristics of the sound, and the output node is the text.
[0055] 2. Attention-based gated scaling adaptive network.
[0056] In the embodiment of the present invention, the attention-based gated scaling (AGS) adaptive network is a series of attention-based gated scaling adaptive layers (AGS adaptive layers) attached to the main network, which belongs to an auxiliary network (or additional network). The auxiliary network generates control weights and adjusts the output parameters of the main network. The main network is any network for acoustic modeling in the general sense (i.e., the acoustic model introduced above), which can be DNN, CNN, RNN or LSTM. In the embodiment of the present invention, the CTC modeling structure using LSTM is used as an example for introduction.
[0057] The auxiliary network provides scaling transformation operations. As mentioned above, each deep feature is activated through several hidden layers of the main network. In an embodiment of the present invention, an AGS adaptive layer can be arranged for at least one hidden layer to generate a corresponding scaling transformation vector for the corresponding hidden layer, and the activation output of the corresponding hidden layer can be re-weighted.
[0058] Those skilled in the art can understand that deploying an AGS adaptive layer for at least one hidden layer can be understood as deploying an AGS adaptive layer for one or more hidden layers respectively; of course, the preferred solution is to deploy an AGS adaptive layer separately for all hidden layers; the specific deployment scheme can be set according to actual conditions, and the relationship between the number of deployed AGS adaptive layers and the recognition effect is explained through experiments later.
[0059] like Figure 4As shown, each hidden layer selected for activation scaling corresponds to a separate AGS adaptive layer, and the scaling transformation vector re-weights the activation of each hidden layer unit. This hierarchical structure provides distributed linear transformations for different levels of deep features, with the expectation of generating deep features that are insensitive to acoustic condition changes or features that are more discriminative for content information.
[0060] In an embodiment of the present invention, in order to capture the most relevant information in the deep features of sentence-level acoustic modeling, the self-attention mechanism is introduced into the AGS adaptive layer. As Figure 5 shown, it is a schematic structural diagram of the AGS adaptive layer. For the speech data x t corresponding to the t-th frame, the corresponding deep feature f t , the corresponding attention weight α t is calculated through the standard self-attention mechanism, and the deep feature f t is weighted, expressed as:
[0061] K t = W k * f t
[0062] Q t = W q * f t
[0063] V t = W v * f t
[0064]
[0065] C t = α t * V t
[0066] Among them, K t , Q t , V t successively represent the K, Q, and V matrices in the attention mechanism; represents the matrix obtained by transposing the matrix K t ; W k , W q and W v are d a ×d f dimensional mapping matrices, d a represents the dimension of the mapped K t , Q t and V t , d f represents the number of network nodes when the acoustic model encodes the speech sequence; C tRepresents the weighted features; considering the efficiency in space and time, calculate the attention weight α t When K t and Q t The similarity score between them is calculated using the point-wise attention mechanism.
[0067] The weighted features are nonlinearly activated, and the activation function uses a 2x sigmoid function to constrain the value of the scaling transformation vector to be in the interval [0,2]. For the AGS adaptive layer corresponding to the lth hidden layer, the scaling transformation vector is obtained by the following formula: It is expressed as:
[0068]
[0069] in, They represent the weight parameters and bias parameters used for the nonlinear activation of the AGS adaptive layer corresponding to the lth hidden layer respectively; is a h dimensional vector, d h is the dimension of hidden layer activation, that is, the number of hidden layer units.
[0070] As mentioned before, AGS adaptive layers can be arranged for one or more hidden layers, and K used by different AGS adaptive layers to calculate the scaling transformation vector t , Q t 、V t are the same, the main difference is that the weight parameters and bias parameters used in the nonlinear activation are different.
[0071] In the embodiment of the present invention, for the t-th frame of voice data x t The corresponding depth feature f t , the activation output of the l-1th hidden layer of the acoustic model is recorded as at the same time As the input of the lth hidden layer; generate deep features f through the attention-based gated scaling adaptation layer corresponding to the lth hidden layer t The scaling transformation vector And re-weight the activation output of the lth hidden layer, expressed as:
[0072]
[0073] Where ⊙ represents the Hadamard product, D l represents the transformation operation of the lth hidden layer, as mentioned above, it can be DNN, CNN or RNN, and the present invention selects LSTM; represents the activation output of the lth hidden layer; It represents the result of re-weighting the activation output of the lth hidden layer, which is used as the input of the next hidden layer. If the lth layer is the last layer, the corresponding probability distribution y is obtained through the softmax function. t .
[0074] In the above scheme, the auxiliary network generates a scaling transformation vector for each frame feature, and the self-attention mechanism reweights the input features at the frame level so that each frame of speech contains information from all other frames in the current sentence. This can also be seen as a modeling of temporal information, thereby providing complementary information for the modeling of LSTM in the main network. Moreover, the attention mechanism has the characteristic of selective weighting, which enables the auxiliary network to focus on providing contextually discriminative features. Since the auxiliary network operates on the entire sentence and all feature dimensions, the transformation of the hidden layer features of the main network can be seen as being performed in both time and feature dimensions. In addition, similar to SSNN and DLN, the auxiliary network is dynamically trained and optimized together with the main network, and the adaptation is achieved by dynamically generating scaling transformation parameters based on the input sequence without the need to learn them using additional adaptive data. This approach makes the entire adaptation process simple and fast. Since deeper networks can achieve better recognition accuracy, the strategy of generating different scaling transformation parameters for each layer of the main network can also further improve model performance.
[0075] The above solution of the embodiment of the present invention mainly achieves the following beneficial effects compared with the prior art:
[0076] 1) Compared with the traditional adaptive speech recognition training scheme, the auxiliary network used in the present invention is an online adaptive method. This type of method uses the auxiliary network to directly model the speaker acoustic variables during acoustic model training. The auxiliary network is trained while the main network is trained. No additional adaptive data or speaker auxiliary features are required, which greatly simplifies the adaptive process.
[0077] 2) Self-attention mechanism is widely used in speech recognition. We introduce the self-attention mechanism in the auxiliary network to learn more discriminative features.
[0078] In order to verify the effectiveness of the above solution of the present invention, an experiment is given to illustrate it.
[0079] 1. Select the data set and evaluation indicators.
[0080] The experiments of the present invention are conducted on the public AI-SHELL1 Chinese dataset. The dataset contains a total of 180 hours of speech, including 150 hours of training set, a total of 120,098 sentences, recorded by 340 speakers; 20 hours of development set, a total of 14,326 sentences, recorded by 40 speakers; 10 hours of test set, a total of 7,176 sentences, recorded by 20 speakers. There is no overlap among all speakers in the training set, development set, and test set. During the experiment, all the training set data were used to train the acoustic model, the development set data were used to control the model training process, and the performance of the model was verified on the test set. Table 1 shows the distribution of the dataset.
[0081]
[0082] Table 1 Distribution of AISHELL1 dataset
[0083] For Chinese, a typical language consisting of single words, the character error rate (CER) is usually used as an evaluation indicator for speech recognition systems. The lower the CER, the higher the corresponding recognition accuracy and the better the performance of the recognition system.
[0084] 2. Recognition results of different models.
[0085] In addition to the baseline system (SI) without an adaptive model and the solution of the present invention, three other different adaptive models were established for comparison with the extracted methods, namely SSNN, DLN and LHUC. For the SSNN adaptive model, the SSNN network is added to the input of the SI model. In the configuration of the SSNN network structure, three fully connected layers are used, each with a dimension of 512, and tanh is used as the nonlinear activation function. The output layer uses a linear layer with the same dimension as the input feature dimension of the SI model, i.e. 108. In addition, in order to reduce the possible overfitting of the network, a dropout regularization of 0.1 is used for each layer of the SSNN network. For the LHUC adaptive model, the scaling vector is added after each BLSTM layer. Dropout regularization is also used for the scaling vector, with a size of 0.3. For the DLN adaptive model, the dimension of the sentence-level feature vector is set to 256, and the DLN adaptation is added before the input of all BLSTM layers.
[0086] Table 2 lists the experimental results of all adaptive models used for comparison. The content in brackets is the improvement relative to the baseline. For the LHUC and DLN adaptive models, all LSTM layers are also adapted, and this configuration achieves the best results. It can be seen that the three models have achieved performance improvements to varying degrees. The LHUC algorithm performs best among the three comparison models, and the solution of the present invention has further achieved significant performance improvements relative to the LHUC algorithm.
[0087]
[0088] Table 2 CER of different adaptive models on the test set (CER / (relative improvement%))
[0089] 3. Adaptive recognition results at different layers.
[0090] The effect of adapting different LSTM layers is explored through experiments. Table 3 shows the effects of different combinations of adaptive LSTM layers on the development set and test machine.
[0091]
[0092]
[0093] Table 3 CER of different adaptive LSTM layers on the development set and test set
[0094] It can be seen that when only one LSTM layer is adapted, the deep layer, i.e. the hidden layer close to the output, achieves better results than the shallow layer: on the development set, the deep and shallow layers achieve 7.60% and 8.34% CER respectively; on the test set, the deep and shallow layers achieve 8.68% and 9.46% CER respectively. Furthermore, the CER decreases steadily with the increase in the number of adaptive layers. When all LSTM layers are adapted, the solution provided by the present invention achieves a CER of 7.94% on the test set, which is a 20.28% improvement over the baseline.
[0095] Embodiment 2
[0096] The present invention also provides a speech recognition system, which mainly includes: an acoustic model, and an attention-based gated scaling adaptive network composed of multiple attention-based gated scaling adaptive layers. Based on the acoustic model and the attention-based gated scaling adaptive network, speech recognition is achieved by adopting the scheme introduced in the aforementioned method embodiment. The specific recognition process is described in detail in the aforementioned method embodiment, so it will not be repeated here.
[0097] Embodiment 3
[0098] The present invention also provides a processing device, such as Figure 6 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the aforementioned embodiments.
[0099] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0100] In the embodiment of the present invention, the specific types of the memory, input device and output device are not limited; for example:
[0101] The input device may be a touch screen, an image acquisition device, a physical button or a mouse, etc.;
[0102] The output device may be a display terminal;
[0103] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0104] Embodiment 4
[0105] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0106] In the embodiment of the present invention, the readable storage medium is a computer-readable storage medium and can be set in the aforementioned processing device, for example, as a memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk, etc., which can store program codes.
[0107] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A voice adaptive recognition method, characterized in that, it includes: Encoding the input voice sequence into a deep feature sequence by using an acoustic model, and a blank label is added to the relevant dictionary during the encoding process. The acoustic model is an end-to-end model based on CTC; Applying the CTC criterion to the deep feature sequence to convert the deep feature sequence into a probability distribution sequence. During the conversion process, each deep feature is activated through several hidden layers of the acoustic model, and at least one hidden layer generates a scaling transformation vector for the corresponding deep feature through a corresponding attention-based gated scaling adaptive layer, and uses the scaling transformation vector to re-weight the activation output of the corresponding hidden layer; wherein, the attention-based gated scaling adaptive layer is used to capture relevant information in the deep features of sentence-level acoustic modeling, so that each frame of the voice sequence contains the information of all other frames; During voice recognition, use the probability distribution sequence to calculate the conditional probability of any path formed by the dictionary elements after adding the blank label, perform path aggregation according to the conditional probability, delete the blank label in the path, and the label sequence obtained by merging consecutive repeated labels is the recognition result.
2. A voice adaptive recognition method according to claim 1, characterized in that, The encoding of the input voice sequence into a deep feature sequence by using the acoustic model includes: Denote the input speech sequence as X = {x 1 , x 2 ,..., x T}, where x t represents the speech data of the t-th frame, and each frame corresponds to a moment, t = 1, 2,..., T, where T represents the total number of frames; Encode the input speech sequence X into a deep feature sequence F = {f 1 , f 2 ,..., f T} using an acoustic model, where the deep feature f t is a vector of dimension |V| + 1, where |V| represents the number of modeling units in the vocabulary V, which is equal to the number of labels, and 1 represents an additional dimension corresponding to the blank label.
3. A voice adaptive recognition method according to claim 1, characterized in that, The at least one hidden layer generates a scaling transformation vector for the corresponding deep feature through a corresponding attention-based gated scaling adaptive layer, and uses the scaling transformation vector to re-weight the activation output of the corresponding hidden layer, including: For the t-th frame of speech data x t The corresponding depth feature f t , the activation output of the hidden layer of the (l - 1)-th layer of the acoustic model is denoted as Meanwhile As the input of the l-th hidden layer; Generate the depth feature f through the attention-based gated scaling adaptive layer corresponding to the l-th hidden layer t The scaling transformation vector And re-weight the activation output of the l-th hidden layer, expressed as: where ⊙ represents the Hadamard product, and D l represents the transformation operation of the l-th hidden layer, represents the activation output of the l-th hidden layer, represents the result obtained by re-weighting the activation output of the l-th hidden layer, which is used as the input of the next hidden layer. If the l-th layer is the last layer, the corresponding probability distribution y is obtained through the softmax function t .
4. A voice adaptive recognition method according to claim 1 or 3, characterized in that, The steps for the attention-based gated scaling adaptive layer to generate a scaling transformation vector for the corresponding deep feature include: For the t-th frame of speech data x t the corresponding depth feature f t , calculate the corresponding attention weight α through the self-attention mechanism t , and weight the depth feature f t as follows: K t = W k * f t Q t = W q * f t V t = W v * f t C t = α t * V t Among them, K t , Q t , V t respectively represent the K value, Q value, and V value in the attention mechanism; represents the matrix obtained after transposing the matrix K t ; W k , W q and W v are d a ×d f dimensional mapping matrices, where d a represents the dimension of the mapped K t , Q t and V t , and d f represents the number of network nodes when the acoustic model encodes the speech sequence; C t represents the weighted feature; Perform non-linear activation on the weighted features. For the attention-based gated scaling adaptive layer corresponding to the l-th hidden layer, obtain the scaling transformation vector through the following formula Denoted as: where sigmoid represents the sigmoid function used for non-linear activation, respectively represent the weight parameter and bias parameter used for non-linear activation of the attention-based gated scaling adaptive layer corresponding to the l-th hidden layer.
5. A voice adaptive recognition method according to claim 1, characterized in that, The probability distribution sequence is represented as Y = {y 1 , y 2 ,..., y T}, where y t is a vector of dimension |V| + 1, |V| represents the number of modeling units in the dictionary V, and 1 represents an additional dimension corresponding to the blank label; represents the probability that the t-th frame of speech data x t belongs to label i, represents the probability that the t-th frame of speech data x t belongs to the blank label, i = 1, 2,..., |V| + 1, t = 1, 2,..., T, and T represents the total number of frames.
6. A voice adaptive recognition method according to claim 5, characterized in that, The calculation of the conditional probability of any path formed by the dictionary elements after adding the blank label by using the probability distribution sequence includes: Denote the dictionary after adding blank tags as which represents the set of all sequences of length T defined on the dictionary . The elements in the set are called paths. The conditional probability of any path π in the set is calculated by the following formula: Among them, X represents the input speech sequence, X = {x 1 , x 2 ,..., x T}, T represents the total number of frames, that is, the length of the input speech sequence, and π t represents the t-th label in the path π.
7. A voice adaptive recognition method according to claim 1, characterized in that, During the training process, the acoustic model and all attention-based gated scaling adaptive layers are dynamically trained and optimized together.
8. A voice adaptive recognition system, characterized in that, it includes: An acoustic model, and an attention-based gated scaling adaptive network composed of multiple attention-based gated scaling adaptive layers; The acoustic model and the attention-based gated scaling adaptive network implement voice recognition through the method described in any one of claims 1 to 7.
9. A processing device, characterized in that, it includes: One or more processors; A memory for storing one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1 to 7.
10. A readable storage medium stores a computer program, characterized in that, when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Self-adapting method of DNN acoustic model based on personal identity characteristics
CN109637526A
Language learner voiceprint recognition method based on multi-task self-attention mechanism
CN112908341A