A method, apparatus and processing device for training a speech recognition model
By introducing MoCov2 and Wav2vec2.0 models into the speech recognition model, and combining cross-modal and temporal attention mechanisms, lip and audio features are fused, and the training process is optimized, the problem of low speech recognition accuracy under environmental noise interference is solved, and higher-precision speech recognition is achieved.
Patent Information
- Application Number
- CN202211392542.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-11-08
AI Technical Summary
Existing speech recognition technologies have low accuracy rates under environmental noise interference.
The MoCov2 and Wav2vec2.0 models were used for pre-training, and cross-modal attention and temporal attention mechanisms were combined to train the speech recognition model through lip and audio feature fusion. The model was then optimized using a Transformer decoder and a CTC model.
It improves the accuracy of speech recognition models, reduces the interference of environmental noise on speech recognition, and obtains more accurate recognition results.
Smart Images

Figure CN115881101B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, specifically to a training method, apparatus, and processing device for a speech recognition model. Background Technology
[0002] In the application of Artificial Intelligence (AI), speech recognition is a major application scenario. Speech recognition technology can be understood as converting the lexical content of human speech into computer-readable input. In this way, data can be directly input from audio and video based on sound. It can be applied to various fields such as industry, home appliances, communications, automotive electronics, medical care, home services, and consumer electronics products, and has a broad application prospect.
[0003] During the recording of audio and video data, it is virtually impossible for users to be in an absolutely quiet place. In reality, there is often some ambient noise around the user, which will be recorded into the audio and video data, and this will interfere with the subsequent speech recognition processing.
[0004] During the research process of existing related technologies, the inventors discovered that the accuracy of existing speech recognition technologies needs to be improved when affected by environmental noise. Summary of the Invention
[0005] This application provides a training method, apparatus, and processing device for a speech recognition model, which is used to train a speech recognition model with higher accuracy. In practical applications, this can greatly reduce the interference of environmental noise on speech recognition and obtain more accurate speech recognition results.
[0006] Firstly, this application provides a method for training a speech recognition model, the method comprising:
[0007] Obtain the sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2;
[0008] The audio data in the LRW audio and video data is input into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model.
[0009] The video data L in the audio and video data LRS2 v The input is a video encoding module configured with a pre-trained MoCov2 model and a first Transformer encoder, which encodes the lip features X. v And the audio data L in the audio and video data LRS2 aThe input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a ;
[0010] Through lip features X v and audio feature X a The input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f.
[0011] The fused feature f is input to a speech recognition model consisting of a Transformer decoder and a CTC model for training. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the LRS2 audio and video data. w Calculated.
[0012] Secondly, this application provides a training device for a speech recognition model, the device comprising:
[0013] The sample acquisition unit is used to acquire a sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2.
[0014] The pre-training unit is used to input the audio data in the LRW audio and video data into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model.
[0015] Feature coding unit, used to convert video data L in audio and video data LRS2 v The input is a video encoding module configured with a pre-trained MoCov2 model and a first Transformer encoder, which encodes the lip features X. v And the audio data L in the audio and video data LRS2 a The input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a ;
[0016] Feature fusion unit, used to fuse lip features X v and audio feature X a The input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f.
[0017] The training unit is used to train a speech recognition model consisting of a Transformer decoder and a CTC model by inputting the fused features f. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the LRS2 audio and video data. w Calculated.
[0018] Thirdly, this application provides a processing device, including a processor and a memory, wherein a computer program is stored in the memory, and when the processor invokes the computer program in the memory, it executes the method provided by the first aspect of this application or any possible implementation of the first aspect of this application.
[0019] Fourthly, this application provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform the method provided in the first aspect of this application or any possible implementation thereof.
[0020] From the above, it can be concluded that this application has the following beneficial effects:
[0021] For training the speech recognition model, this application incorporates the MoCov2 and wav2vec2.0 models into the model training architecture to obtain more robust and stable audio and video features. Then, cross-modal attention and temporal attention mechanisms are introduced to correct and align the audio and video feature information and obtain a fused representation. The speech recognition model is then trained to decode and output the speech recognition result. During this training process, video features are combined with text features to obtain video features with more textual information, thereby obtaining higher-quality sample data for better training of the speech recognition model. The resulting speech recognition model has higher speech recognition accuracy, which can significantly reduce the interference of environmental noise on speech recognition in practical applications, leading to more accurate speech recognition results. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating one method for training the speech recognition model of this application.
[0024] Figure 2 This is a schematic diagram of one possible model training architecture for this application;
[0025] Figure 3 This is a schematic diagram of a training device for the speech recognition model of this application;
[0026] Figure 4 This is a schematic diagram of one type of processing equipment used in this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved.
[0029] The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between modules shown or discussed may be through some interfaces, and the indirect coupling or communication connection between modules may be electrical or other similar forms, none of which are limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules may be selected to achieve the purpose of the solution in this application according to actual needs.
[0030] Before introducing the training method of the speech recognition model provided in this application, we will first introduce the background content involved in this application.
[0031] The speech recognition model training method, apparatus, and computer-readable storage medium provided in this application can be applied to processing devices to train speech recognition models with higher accuracy. This can greatly reduce the interference of environmental noise on speech recognition in specific applications and obtain more accurate speech recognition results.
[0032] The speech recognition model training method mentioned in this application can be implemented by a speech recognition model training device, or by different types of processing devices such as a server, physical host, or user equipment (UE) that integrates the speech recognition model training device. The speech recognition model training device can be implemented in hardware or software. The UE can specifically be a smartphone, tablet, laptop, desktop computer, or personal digital assistant (PDA) or other terminal device. The processing devices can be configured in a device cluster.
[0033] It is understood that the specific form of the processing device can be configured according to actual needs. Its main function is to serve as a carrier to realize the data processing involved in training the speech recognition model of this application. Furthermore, it can also load the trained speech recognition model to perform speech recognition application functions as needed.
[0034] In addition, a brief introduction will be given to the abbreviations involved in the following positions.
[0035] Connectionist Temporal Classification (CTC) based on neural networks;
[0036] Region of Interest (ROI);
[0037] Bidirectional Enoceder Representations from Transformers (BERT).
[0038] The training method for the speech recognition model provided in this application will now be introduced.
[0039] First, refer to Figure 1 , Figure 1 The diagram illustrates a flowchart of the training method for the speech recognition model of this application. The training method for the speech recognition model provided in this application may specifically include the following steps S101 to S105:
[0040] Step S101: Obtain the sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2;
[0041] Understandably, the initial stage of training a speech recognition model involves acquiring and processing its sample set.
[0042] The acquisition and processing of the sample set here can be either real-time collection of relevant data or capture and processing of thread data. This application does not specifically limit the acquisition method.
[0043] The sample set involved in this application specifically includes two types: one is character-level audio / video data LRW, and the other is sentence-level audio / video data LRS2. Specifically, the audio / video data LRW can include the audio data discussed below, and the audio / video data LRS2 can include the video data discussed below. v Audio data L a and text data L w The content can be extracted from the original video or configured directly with the original video.
[0044] For the character-level audio and video data LRW here, the dataset contains only a single character, which means that in audio and video speech recognition, only a single character or word needs to be recognized.
[0045] For the LRS2 audio and video data at the sentence level here, the dataset consists of a single sentence. The recognition should be performed on the entire sentence. In audio and video speech recognition, you can pay attention to distinguishing the spaces between English words and the sentence ending marks.
[0046] In practical applications, this audio and video data can be obtained from relevant audio and video libraries, or it may be a relevant audio and video library itself.
[0047] The sample set corresponds to the subsequent training process of the speech recognition model, and can be further divided into training set, validation set and test set, which correspond to the three different model training stages of training, validation and testing. This basic model training mechanism is not the focus of this application, so it will not be explained in detail.
[0048] As an example, the data in the sample set can be configured as follows: 98% as training data, 1% as validation data, and 1% as test data.
[0049] Furthermore, for the sample set, data preprocessing can also be applied to improve the data quality of the samples and / or to standardize the data. For example, the mouth region can be cropped from the video and a 112*112 ROI feature can be extracted and converted into a grayscale image, which can then be compared with the audio data extracted from the video. a Normalization is performed.
[0050] Step S102: Input the audio data in the audio and video data LRW into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model.
[0051] Regarding the overall model training architecture, this application combines a pre-trained model to assist in the training of the speech recognition model, thereby helping the speech recognition model to achieve better training results.
[0052] In this way, the audio data in the LRW audio and video data of the sample set can be used to train the initial state of the MoCov2 model (which can be referred to as the initial MoCov2 model). This completes the processing of a pre-trained model and obtains the pre-trained MoCov2 model.
[0053] It is understandable that the specific training method for the pre-training of the initial MoCov2 model is similar to that of existing technologies, so it is not explained in detail here.
[0054] For the MoCov2 model, the initial MoCov2 model will be used as an example to illustrate the characteristics of the model itself.
[0055] The initial MoCov2 model is a self-supervised model, which can specifically include an encoder module, a multilayer perceptron module, and a queue module. For these three components, we have:
[0056] The encoder module processes the tensor corresponding to the input image data to obtain the feature matrix and constructs the keys ("keys" in "key-value pairs") in the unsupervised learning dictionary data to retrieve the corresponding data;
[0057] Multi-layer sensing modules acquire image features;
[0058] The queue module stores and maintains dictionary data by setting queue rules.
[0059] Step S103, transfer the video data L from the audio / video data LRS2. v The input is a video encoding module configured with a pre-trained MoCov2 model and a first Transformer encoder, which encodes the lip features X. v And the audio data L in the audio and video data LRS2 a The input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a ;
[0060] After configuring the pre-trained MoCov2 and Wav2vec2.0 models, you can then begin training the subsequent speech recognition model.
[0061] In this application, the speech recognition model involves a Transformer decoder. The input to the speech recognition model is provided by both the pre-trained MoCov2 model and the Wav2vec2.0 model. Both models can also be equipped with Transformer encoders. For ease of explanation, the pre-trained MoCov2 model and the Wav2vec2.0 model are referred to as the first Transformer encoder and the second Transformer encoder.
[0062] In this application, the first convolutional layer of the general MoCov2 model can be truncated and replaced with a three-dimensional convolutional layer. This setting can be understood as the application intentionally configuring the output of the three-dimensional convolutional layer to be the same as the input of the first ResBlock in the MoCov2 model, thereby providing a compatible interface to extract deeper features in time and space.
[0063] At this point, during the training process of the speech recognition model, on the one hand, the video data L in the audio and video data LRS2 can be processed. v The input is a video encoding module configured with a pre-trained MoCov2 model and a first Transformer encoder, which encodes the lip features X. v On the other hand, it can also store audio data from LRS2 audio / video data. a The input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a It provides two aspects of features.
[0064] Among them, the Wav2vec2.0 model is a pre-trained existing model.
[0065] Step S104, using lip feature X v and audio feature X a The input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f.
[0066] Furthermore, to facilitate input into the speech recognition model and to represent more accurate feature content, the two features obtained earlier, namely lip features X... v And the audio feature Xa can be used for the feature fusion processing specially designed in this application to achieve the purpose of alignment correction and feature fusion.
[0067] The feature fusion processing involved here can be accomplished by a joint module consisting of the cross-modal attention module and the temporal attention module configured in this application.
[0068] As is easy to understand, this joint module includes two main modules: a cross-modal attention module and a temporal attention module. It achieves a better fusion effect between the two features through cross-modal attention mechanisms and temporal attention mechanisms.
[0069] For a more detailed understanding of the attention mechanisms of both, please refer to the following content.
[0070] 1) Feature processing for the cross-modal attention module includes the following:
[0071] X a X represents the audio features extracted from the audio modality. v This represents the video features extracted from the video modality. L represents the number of sequence segments in the given video and audio input sequences. and Let L represent the feature vectors of segments L = 1, 2, ..., L respectively.
[0072] Audio features X a and video feature X v The audio and video features are spliced together J = [X] a ;X v ]∈R d×L d = d a +d v d a d represents the feature dimension of mode A. v The feature dimension representing the V mode,
[0073] Audio Feature X a The correlation matrix C of splicing feature J a for:
[0074]
[0075] Video Feature X v And the correlation matrix C of the spliced feature J v for:
[0076]
[0077] Correlation matrix C a And correlation matrix C v It provides not only semantic relevance for the same modality, but also semantic relevance for different modalities, with the joint relevance matrix C. a Correlation matrix and C v The high correlation coefficient indicates that the corresponding samples are highly correlated in the same mode and in other modes.
[0078] In the correlation matrix C aand audio feature X a Based on this, combined with a learnable weight matrix and W a The attention weights for the audio modal are calculated as follows:
[0079]
[0080] Similarly, in the correlation matrix C v and video feature X v Based on this, combined with a learnable weight matrix and W v The attention weights for the video modal are calculated as follows:
[0081]
[0082] Attention features for the audio and video modalities are calculated using attention maps, and are represented as follows:
[0083] X att,a =W ha H a +X a ,
[0084] X att,v =W hv H v +X v ,
[0085] Feature X att,a and feature X att,v By concatenating the features, we obtain the features used for inputting the temporal attention modality:
[0086] X att =[X att,a ;X att,v ].
[0087] The specific setup of the cross-modal attention mechanism here is easy to understand. Audio features and video features belong to two different modalities, so they cannot be directly spliced together. It is necessary to fuse the features of the two modalities. Through the nested application of the above series of formulas, the features of the two modalities learn each other's features, which helps to better capture the semantic modal relationship between audio features and visual features.
[0088] 2) Feature processing for the temporal attention modality includes the following:
[0089] In the trainable weight matrix W T Based on this, create the corresponding representation matrix T = X att W T , T∈R L×dThe temporal attention features of the temporal attention modality are defined by the following formula:
[0090]
[0091] Here, Q is the query vector in the attention mechanism, K is the key vector in the attention mechanism, and V is the value vector in the attention mechanism. The three are obtained by multiplying the input feature X by their respective weight matrices.
[0092] This setting can be understood as follows: in order to calculate the attention score, the query matrix W is... T Multiply by the time matrix T, then multiply by its transpose to keep the dimensions unchanged, and then divide by the norm of the time matrix T to avoid getting too large a value.
[0093] The specific settings of the time attention mechanism here are easy to understand. It provides a specific implementation scheme for the aggregation and fusion features in the time dimension. This application takes into account that video data and audio data are time-series, so it can focus on the content information of the current moment and all previous moments.
[0094] Step S105: The fused feature f is input into the speech recognition model composed of the Transformer decoder and the CTC model for training. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the audio and video data LRS2. w Calculated.
[0095] As can be seen, the training of a speech recognition model involves not only the loss function that processes the speech recognition results output by the model itself, but also the standard speech recognition results, namely the text features X of the LRS2 audio and video data mentioned here. w .
[0096] For the text feature X of the LRS2 of this audio and video data w In addition to the acquisition and processing, the method of this application may also include corresponding extraction processing, namely:
[0097] From audio / video data LRS2 text data L w In this process, text features X are extracted using the BERT model. w .
[0098] For the BERT model, during the extraction process, token embedding, segment embedding, and position embedding are obtained through its input layer. The three are then added together to obtain the output vector of the input layer.
[0099] The BERT model can not only extract word vectors from text, but also obtain the semantics of the text and the relationship between words. BERT belongs to the language model. Other models only focus on the word vectors of each character or word when extracting text features, ignoring the context of the sentence. However, the BERT model uses the nature of its language model to adjust the text features in combination with the context.
[0100] Of course, in specific applications, for text feature X w It can also be implemented using other types of model algorithms, and can be adjusted according to actual needs.
[0101] Furthermore, at a more detailed level, the BERT model can further enhance text feature extraction by occluding random words before extracting text features X. w The prediction is used to obtain good word vector features.
[0102] It's important to understand that the BERT model itself can extract text features. Here, we've configured it with a task of randomly masking words. By replacing 15% of the words in the dataset with [mask], the goal is to identify these masked words. During the identification process, it needs to learn the possible masking values by considering the semantic relationships between the preceding and following text. By selecting the value with the highest probability as the final output, this task demonstrates that the BERT model can effectively learn the semantic relationships between sentences. Therefore, when extracting features from individual words, it also considers the contextual relationships, thus obtaining good word vector features.
[0103] As can be seen, the speech recognition model to be trained in this application is specifically composed of a Transformer decoder and a CTC model.
[0104] Understandably, the goal of Transformer is to project audio and video modalities into the same coding space through an encoder-decoder structure, thereby achieving the effect of fusing multimodal information. The decoder part consists of CTC and a Transformer decoder.
[0105] During the training of the Transformer decoder, the loss can specifically consist of L1 loss, CTC loss, and cross-entropy loss, and may include the following:
[0106] The following Smooth L1 loss function is used:
[0107]
[0108] in, d represents the text feature X of the LRS2 audio / video data. wLanguage features decoded by the Transformer decoder The difference between them
[0109] The CTC loss assumes conditional independence between each output prediction and takes the form:
[0110]
[0111] The autoregressive decoder eliminates assumptions by directly estimating the posterior of the chain rule, in the form of:
[0112]
[0113] x = [x1, ..., xT] represents the output sequence of the speech recognition model, y = [y1, ..., yL] represents the target, T represents the input length, and L represents the target length.
[0114] Total loss is defined as:
[0115] Loss=λlog p CTC (y|x)+(1-λ)log p CE (y|x)+L1(d),
[0116] Here, λ controls the relative weights of the hybrid CTC / attention mechanism between CTC loss and cross-entropy loss.
[0117] In addition, before training the Transformer decoder, it may involve setting related network settings, such as the number of hidden layers, the number of hidden layer nodes, and the learning rate.
[0118] During training, the model can be optimized using the calculation results of the loss function, such as the calculation results of Smooth L1 loss, CTC loss, and cross-entropy loss. Backpropagation is then performed to optimize the model parameters. Once the training requirements such as 1000 iterations, recognition accuracy, and training time are met after N rounds of propagation, the model training can be completed.
[0119] For a better understanding of the above exemplary embodiments, please refer to... Figure 2 The diagram shown represents one architecture of the model training architecture proposed in this application.
[0120] Comparison Figure 2 The training architecture shown can be understood as follows: In terms of visual modality, the visual front-end is first pre-trained using self-supervised learning, which is accomplished here using the MoCov2 model. Then, the sequence classification is modified and trained using word-level video segments in the audio and video data LRW.
[0121] After that, the visual front end is inherited by the pure video model, in which the visual back end and dedicated decoder are used. Finally, the audio and visual features are extracted through the trained audio model and lip reading model and fed into the fusion module. Considering the computational limitations, the audio and video back end outputs can be pre-calculated, and in the final stage, only the parameters of the fusion module and decoder are learned.
[0122] Once the model has been trained, it can be put into practical use for speech recognition applications.
[0123] Correspondingly, the method of this application may also include:
[0124] The speech data to be recognized is input into the speech recognition model, which performs speech recognition processing and outputs the speech recognition result.
[0125] In addition, this application can also verify the speech recognition accuracy of the trained model (the same verification method can also be used during training).
[0126] As an example, the speaker's video and audio can be fed into a trained model with signal-to-noise ratios of 0dB, 5dB, and 10dB, respectively. The model outputs the recognized text, and the word error rate (BER) is used to evaluate the recognition performance. BER is a metric used to evaluate speech recognition performance, specifically the error rate between the predicted text and the standard text. A lower BER is better. The formula for calculating BER is:
[0127]
[0128] Where S represents the number of substitutions that occur when converting a predicted sample into a real sample, D represents the number of substitutions that occur when converting a real sample into a predicted sample, I represents the number of insertions that occur when converting a test sample into a real sample, N represents the total number of characters or English words in the standard sample sentence, and C represents the number of correctly identified characters in the predicted sample sentence.
[0129] As can be seen from the above, for the training of the speech recognition model, this application incorporates the MoCov2 model and the wav2vec2.0 model into the model training architecture to obtain more robust and stable audio and video features. Then, it further introduces cross-modal attention and temporal attention mechanisms to correct and align the audio and video feature information and obtain a fused representation. The speech recognition model is then trained to decode and output the speech recognition result. During this training process, this application combines video features with text features to obtain video features with more textual information, thereby obtaining higher quality sample data, which can better train the speech recognition model. The speech recognition model trained in this way has higher speech recognition accuracy, which can greatly reduce the interference of environmental noise on speech recognition in specific applications and obtain more accurate speech recognition results.
[0130] The above is an introduction to the training method of the speech recognition model provided in this application. In order to facilitate better implementation of the training method of the speech recognition model provided in this application, this application also provides a training device for the speech recognition model from the perspective of functional modules.
[0131] See Figure 3 , Figure 3 This is a schematic diagram of a structure for a training device of the speech recognition model of this application. In this application, the training device 300 for the speech recognition model may specifically include the following structure:
[0132] The sample acquisition unit 301 is used to acquire a sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2.
[0133] The pre-training unit 302 is used to input the audio data in the audio-visual data LRW into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model.
[0134] Feature encoding unit 303 is used to encode video data L in audio and video data LRS2. v The input is a video encoding module configured with a pre-trained MoCov2 model and a first Transformer encoder, which encodes the lip features X. v And the audio data L in the audio and video data LRS2 a The input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a ;
[0135] Feature fusion unit 304 is used to fuse lip features X v and audio feature X aThe input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f.
[0136] Training unit 305 is used to train a speech recognition model consisting of a Transformer decoder and a CTC model by inputting the fused feature f. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the LRS2 audio and video data. w Calculated.
[0137] In one exemplary implementation, the initial MoCov2 model is a self-supervised model, which includes an encoder module, a multilayer perceptron module, and a queue module.
[0138] The encoder module processes the tensors corresponding to the input image data to obtain the feature matrix and constructs the keys in the unsupervised learning dictionary data to retrieve the corresponding data;
[0139] Multi-layer sensing modules acquire image features;
[0140] The queue module stores and maintains dictionary data by setting queue rules.
[0141] In yet another exemplary implementation, the apparatus further includes a feature extraction unit 306, for:
[0142] From audio / video data LRS2 text data L w In this process, text features X are extracted using the BERT model. w .
[0143] In another exemplary implementation, the BERT model, during text feature extraction, occludes random words before extracting text features X. w The prediction.
[0144] In yet another exemplary implementation, the feature processing of the cross-modal attention module includes the following:
[0145] X a X represents the audio features extracted from the audio modality. v This represents the video features extracted from the video modality. L represents the number of sequence segments in the given video and audio input sequences. and Let L represent the feature vectors of segments L = 1, 2, ..., L respectively.
[0146] Audio features X a and video feature X v The audio and video features are spliced together J = [X]a ;X v ]∈R d×L d = d a +d v d a d represents the feature dimension of mode A. v The feature dimension representing the V mode,
[0147] Audio Feature X a The correlation matrix C of splicing feature J a for:
[0148]
[0149] Video Feature X v And the correlation matrix C of the spliced feature J v for:
[0150]
[0151] In the correlation matrix C a and audio feature X a Based on this, combined with a learnable weight matrix and W a The attention weights for the audio modal are calculated as follows:
[0152]
[0153] In the correlation matrix C v and video feature X v Based on this, combined with a learnable weight matrix and W v The attention weights for the video modal are calculated as follows:
[0154]
[0155] Attention features for the audio and video modalities are calculated using attention maps, and are represented as follows:
[0156] X att,a =W ha H a +X a ,
[0157] X att,v =W hv H v +X v ,
[0158] Feature X att,a and feature X att,u By concatenating the features, we obtain the features used for inputting the temporal attention modality:
[0159] X att =[X att,a ;X att,v ].
[0160] In yet another exemplary implementation, the feature processing for the temporal attention modality includes the following:
[0161] In the trainable weight matrix W T Based on this, create the corresponding representation matrix T = X att W T , T∈R L×d The temporal attention features of the temporal attention modality are defined by the following formula:
[0162]
[0163] In yet another exemplary implementation, the training process of the Transformer decoder includes the following:
[0164] The following Smooth L1 loss function is used:
[0165]
[0166] in, d represents the text feature X of the LRS2 audio / video data. w Language features decoded by the Transformer decoder The difference between them
[0167] The CTC loss assumes conditional independence between each output prediction and takes the form:
[0168]
[0169] The autoregressive decoder eliminates assumptions by directly estimating the posterior of the chain rule, in the form of:
[0170]
[0171] x = [x1, ..., xT] represents the output sequence of the speech recognition model, y = [y1, ..., yL] represents the target, T represents the input length, and L represents the target length.
[0172] Total loss is defined as:
[0173] Loss=λlog p CTC (y|x)+(1-λ)log p CE (y|x)+L1(d),
[0174] Here, λ controls the relative weights of the hybrid CTC / attention mechanism between CTC loss and cross-entropy loss.
[0175] This application also provides a processing device from a hardware architecture perspective, see [link / reference]. Figure 4 , Figure 4 This diagram illustrates a structural schematic of the processing device of this application. Specifically, the processing device may include a processor 401, a memory 402, and an input / output device 403. The processor 401 executes the computer program stored in the memory 402 to implement, for example... Figure 1 The corresponding steps of the training method for the speech recognition model in the embodiment; or, when the processor 401 executes the computer program stored in the memory 402, it implements as follows: Figure 3 Corresponding to the functions of each unit in the embodiment, the memory 402 is used to store the functions executed by the processor 401 as described above. Figure 1 The computer program required for training the speech recognition model in the corresponding embodiment.
[0176] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 402 and executed by processor 401 to complete this application. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a computer device.
[0177] The processing device may include, but is not limited to, processor 401, memory 402, and input / output device 403. Those skilled in the art will understand that the illustrations are merely examples of the processing device and do not constitute a limitation on the processing device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the processing device may also include network access devices, buses, etc., and processor 401, memory 402, input / output device 403, etc., are connected via a bus.
[0178] Processor 401 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting various parts of the device through various interfaces and lines.
[0179] The memory 402 can be used to store computer programs and / or modules. The processor 401 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 402 and by calling data stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the processing device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0180] When processor 401 executes a computer program stored in memory 402, it can specifically perform the following functions:
[0181] Obtain the sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2;
[0182] The audio data in the LRW audio and video data is input into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model.
[0183] The video data Lv from the LRS2 audio / video data set was input into the video encoding module configured with a pre-trained MoCov2 model and a first Transformer encoder to encode the lip feature X. v And the audio data L in the audio and video data LRS2 aThe input is an audio encoding module configured with a Wav2vec2.0 model and a second-stage Transformer encoder, which encodes the audio features X. a ;
[0184] Through lip features X v and audio feature X a The input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f.
[0185] The fused feature f is input to a speech recognition model consisting of a Transformer decoder and a CTC model for training. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the LRS2 audio and video data. w Calculated.
[0186] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the training device, processing equipment, and corresponding units of the speech recognition model described above can be found in, for example... Figure 1 The training method of the speech recognition model in the corresponding embodiment will not be described in detail here.
[0187] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0188] Therefore, this application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the present application. Figure 1 The steps of the training method for the speech recognition model in the corresponding embodiment can be found in the following example. Figure 1 The training method of the speech recognition model in the corresponding embodiment will not be repeated here.
[0189] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0190] Because of the instructions stored in the computer-readable storage medium, the present application can be executed as described above. Figure 1 The steps of the training method for the speech recognition model in the corresponding embodiment can therefore achieve the results of this application. Figure 1 The beneficial effects that the training method of the speech recognition model in the corresponding embodiment can achieve are detailed in the preceding description and will not be repeated here.
[0191] The training method, apparatus, processing device, and computer-readable storage medium of the speech recognition model provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for training a speech recognition model, characterized in that, The method includes: Obtain a sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2; The video data in the LRW audio and video data is input into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model. The video data L in the audio and video data LRS2 v The input is a video encoding module configured with the pre-trained MoCov2 model and the first Transformer encoder, which encodes the lip features X. v And the audio data L in the audio and video data LRS2 a The input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a ; Through the lip feature X v and the audio feature X a The input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f. The fused feature f is input into a speech recognition model composed of a Transformer decoder and a CTC model for training. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the LRS2 audio and video data. w Calculated; The feature processing of the cross-modal attention module includes the following: X a X represents the audio features extracted from the audio modality. v This represents the video features extracted from the video modality. L represents the number of sequence segments in the given video and audio input sequences. and Let each represent a feature vector of a segment l = 1, 2, ..., L. Audio features X a and video feature X v The audio and video features are spliced together J = [X] a ;X v ]∈R d×L d = d a +d v d a d represents the feature dimension of the audio modality. v The feature dimension of the video modality is represented. The audio feature X a The correlation matrix C of splicing feature J a for: The video feature X v And the correlation matrix C of the splicing feature J v for: Wherein, the splicing feature J is the audio feature X. a and the video feature X v The concatenation result along the feature dimension, In the correlation matrix C a and the audio feature X a Based on this, combined with a learnable weight matrix and W a The attention weights for the audio modal are calculated using the following formula: In the correlation matrix C v and the video feature X v Based on this, combined with a learnable weight matrix and W v The attention weights of the video modal are calculated using the following formula: Attention features for the audio modality and the video modality are calculated using attention maps, and are expressed as follows: X att,a =W ha H a +X a , X att,v =W hv H v +X v , Feature X att,a and feature X att,v By concatenating the features, we obtain the features used as input to the temporal attention module: X att =[X att,a ;X att,v ]。 2. The method according to claim 1, characterized in that, The initial MoCov2 model is a self-supervised model, which includes an encoder module, a multilayer perceptron module, and a queue module. The encoder module processes the tensor corresponding to the input image data to obtain the feature matrix and constructs the key in the unsupervised learning dictionary data to retrieve the corresponding data. The multilayer perceptron module acquires image features; The queue module stores and maintains dictionary data by setting queue rules.
3. The method according to claim 1, characterized in that, The method further includes: From the text data L of the audio / video data LRS2 w In the process, the text features X are extracted using the BERT model. w .
4. The method according to claim 3, characterized in that, In the text feature extraction process, the BERT model occludes random words before extracting the text features X. w The prediction.
5. The method according to claim 1, characterized in that, The feature processing of the time attention module includes the following: In the trainable weight matrix W T Based on this, create the corresponding representation matrix T = X att W T ,T∈R L×d The time attention features of the time attention module are defined by the following formula:
6. The method according to claim 1, characterized in that, The training process of the Transformer decoder includes the following: The following Smooth L1 loss function is used: in, d represents the text feature X of the LRS2 audio / video data. w The language features decoded by the Transformer decoder The difference between them The CTC loss assumes conditional independence between each output prediction and takes the form: The autoregressive decoder eliminates assumptions by directly estimating the posterior of the chain rule, in the form of: x = [x1, ..., xT] is the output sequence of the speech recognition model, y = [y1, ..., yL] is the target, T represents the input length, and L represents the target length. Total loss is defined as: Loss=λlogp CTC (y|x)+(1-λ)logp CE (y|x)+L1(d), Wherein, λ controls the relative weights of the hybrid CTC / attention mechanism between the CTC loss and the cross-entropy loss.
7. A training device for a speech recognition model, characterized in that, The device includes: A sample acquisition unit is used to acquire a sample set, which includes character-level audio and video data LRW and sentence-level audio and video data LRS2. The pre-training unit is used to input the audio data in the LRW audio and video data into the initial MoCov2 model for pre-training to obtain the pre-trained MoCov2 model. Feature encoding unit, used to encode video data L in the audio and video data LRS2 v The input is a video encoding module configured with the pre-trained MoCov2 model and the first Transformer encoder, which encodes the lip features X. v And the audio data L in the audio and video data LRS2 a The input is an audio encoding module configured with a Wav2vec2.0 model and a second Transformer encoder, which encodes the audio features X. a ; Feature fusion unit, used to fuse the lip feature X v and the audio feature X a The input is a joint module consisting of a cross-modal attention module and a temporal attention module, which yields the fused feature f. The training unit is used to train a speech recognition model composed of a Transformer decoder and a CTC model by inputting the fused feature f into the model. The loss function during training is composed of the speech recognition result output by the speech recognition model and the text features X of the LRS2 audio and video data. w Calculated; Feature processing for the cross-modal attention module includes the following: X a X represents the audio features extracted from the audio modality. v This represents the video features extracted from the video modality. L represents the number of sequence segments in the given video and audio input sequences. and Let each represent a feature vector of a segment l = 1, 2, ..., L. Audio features X a and video feature X v The audio and video features are spliced together J = [X] a ;X v ]∈R d×L d = d a +d v d a d represents the feature dimension of the audio modality. v The feature dimension representing the video modality, Audio Feature X a The correlation matrix C of splicing feature J a for: Video Feature X v And the correlation matrix C of the spliced feature J v for: Wherein, the splicing feature J is the audio feature X. a and the video feature X v The concatenation result along the feature dimension is in the correlation matrix C. a and audio feature X a Based on this, combined with a learnable weight matrix and W a The attention weights for the audio modal are calculated as follows: In the correlation matrix C v and video feature X v Based on this, combined with a learnable weight matrix and W v The attention weights for the video modal are calculated as follows: Attention features for the audio and video modalities are calculated using attention maps, and are represented as follows: X att,a =W ha H a +X a , X att,v =W hv H v +X v , Feature X att,a and feature X att,v By concatenating the features, we obtain the features used as input to the temporal attention module: X att =[X att,a ;X att,v ]。 8. A processing apparatus, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the method as described in any one of claims 1 to 6 when it invokes the computer program in the memory.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Audio and video voice separation method and system
CN115171717A