Training Method, System and Electronic Device for Audio and Video Speech Recognition Model
By performing front-end processing and feature sequence splicing on audio and video data, combined with the encoder's attention module, the recognition problem of the audio and video voice recognition system in noisy environments and audio and video is solved, achieving higher speech recognition accuracy and stability, and adapting to single-modal input.
Patent Information
- Application Number
- CN202211695645.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-12-28
AI Technical Summary
In the prior art, single-mode speech recognition system is difficult to recognize speech in noisy environments. The audio-visual speech recognition system has poor recognition effect when the audio-visual video is not synchronized, and the lack of adaptive strategies leads to poor single-mode speech recognition effect.
By performing front-end processing of audio and video in audio and video, the acoustic feature sequence of sound mode and the lip feature sequence of video mode are extracted, and the time dimension is spliced to form a fusion feature sequence. The deep fusion feature encoding is determined using the encoder's attention module, and input it to the connection timing classification module and decoder to obtain prediction recognition results, and finally the model is trained based on the reference recognition results.
In the absence of noisy environments and audio and video, the accuracy and stability of speech recognition are improved, and the single-modal input can be adapted to single-modal input without loss when the visual mode is missing, and the performance is comparable to that of conventional single-modal systems.
Smart Images

Figure CN116013270B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent voice, and in particular, to a training method, system, and electronic device for an audio-visual speech recognition model. Background Art
[0002] With the development of intelligent voice technology, intelligent speech recognition technology has been gradually popularized in daily life. In speech recognition, when the speech signal is distorted by noise, the corresponding speech recognition effect will be affected. Considering that in daily speech interaction, in addition to the sound modality, the visual modality is sometimes necessary for accurate speech understanding, and the visual modality that is not affected by noise can be introduced to assist speech recognition. For example, AVSR (Audio-Visual Speech Recognition) integrates the visual stream into a mature automatic speech recognition model.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art:
[0004] Due to the interference of noise, a single-modal speech recognition system is difficult to recognize speech in a noisy environment.
[0005] Although an audio-visual speech recognition system can improve the performance of speech recognition in a noisy environment, when there is an audio-visual asynchrony, due to the limitation of the depth of the model, the recognition effect is often poor. And when an audio-visual speech recognition system recognizes single-modal speech, due to the lack of an adaptive strategy, the recognition effect is worse than that of a single-modal speech recognition system. Summary of the Invention
[0006] In order to at least solve the above recognition problems existing in audio-visual speech recognition in the prior art. In a first aspect, an embodiment of the present invention provides a training method for an audio-visual speech recognition model, including:
[0007] Perform front-end processing on the audio and video in the audio-visual respectively to obtain an acoustic feature sequence of the sound modality composed of audio modality encoding and audio position encoding, and a lip movement feature sequence of the video modality composed of video modality encoding and video position encoding;
[0008] Perform splicing in the time dimension based on the acoustic feature sequence of the sound modality and the lip movement feature sequence of the video modality to obtain a fused feature sequence, and use the fused feature sequence as the input to the encoder of the audio-visual speech recognition model;
[0009] Through the attention module of the encoder, determine the deep fusion feature encoding of each subsequence with cross-modal context in the fusion feature sequence, and input the partial feature encoding of the corresponding audio in the deep fusion feature encoding into the connectionist temporal classification module and decoder of the audio-visual speech recognition model to obtain the predicted recognition result;
[0010] Train the audio-visual speech recognition model according to the benchmark recognition result of the audio in the audio-visual and the predicted recognition result.
[0011] In a second aspect, an embodiment of the present invention provides a speech recognition method based on an audio-visual speech recognition model, including:
[0012] Input the received audio-visual file into the audio-visual speech recognition model, where the audio-visual file includes: audio-visual with synchronized audio and video, and audio-visual with asynchronous audio and video within a preset frame rate;
[0013] The audio-visual speech recognition model performs front-end processing on the audio and video in the audio-visual file respectively to obtain an acoustic feature sequence in the sound modality composed of an audio modality encoding and an audio position encoding, and a lip movement feature sequence in the video modality composed of a video modality encoding and a video position encoding;
[0014] Perform splicing in the time dimension on the acoustic feature sequence in the sound modality and the lip movement feature sequence in the video modality, and input the spliced fusion feature sequence into the encoder;
[0015] The deep fusion feature encoding of each subsequence with cross-modal context output by the encoder retains the partial feature encoding of the corresponding audio in the deep fusion feature encoding;
[0016] Input the partial feature encoding of the corresponding audio into the connectionist temporal classification module and decoder to obtain the speech recognition result of the audio-visual file.
[0017] In a third aspect, an embodiment of the present invention provides a training system for an audio-visual speech recognition model, including:
[0018] A front-end processing program module for performing front-end processing on the audio and video in the audio-visual respectively to obtain an acoustic feature sequence in the sound modality composed of an audio modality encoding and an audio position encoding, and a lip movement feature sequence in the video modality composed of a video modality encoding and a video position encoding;
[0019] A fusion program module for performing splicing in the time dimension based on the acoustic feature sequence in the sound modality and the lip movement feature sequence in the video modality to obtain a fusion feature sequence, and using the fusion feature sequence as the input to the encoder of the audio-visual speech recognition model;
[0020] An encoding program module, configured to determine, through the attention module of the encoder, the deep fusion feature encoding of each subsequence with cross-modal context in the fusion feature sequence, and input a partial feature encoding of the corresponding audio in the deep fusion feature encoding into the connectionist temporal classification module and the decoder of the audio-visual speech recognition model to obtain a predicted recognition result;
[0021] A training program module, configured to train the audio-visual speech recognition model according to the benchmark recognition result of the audio in the audio-visual and the predicted recognition result.
[0022] In a fourth aspect, an embodiment of the present invention provides a speech recognition system based on an audio-visual speech recognition model, including:
[0023] An audio-visual receiving program module, configured to input a received audio-visual file into the audio-visual speech recognition model, where the audio-visual file includes: an audio-visual with synchronized audio and video, and an audio-visual with asynchronous audio and video within a preset frame rate;
[0024] A front-end processing program module, configured to perform front-end processing on the audio and video in the audio-visual file by the audio-visual speech recognition model to obtain an acoustic feature sequence in the sound modality and a lip movement feature sequence in the video modality;
[0025] A fusion program module, configured to splice the acoustic feature sequence in the sound modality and the lip movement feature sequence in the video modality in the time dimension, and input the spliced fusion feature sequence into the encoder;
[0026] An encoding program module, configured to output the deep fusion feature encoding of each subsequence with cross-modal context by the encoder, and retain the partial feature encoding of the corresponding audio in the deep fusion feature encoding;
[0027] A speech recognition program module, configured to input the partial feature encoding of the corresponding audio into the connectionist temporal classification module and the decoder to obtain the speech recognition result of the audio-visual file.
[0028] In a fifth aspect, an electronic device is provided, including: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the training method of the audio-visual speech recognition model and the speech recognition method based on the audio-visual speech recognition model according to any embodiment of the present invention.
[0029] Sixth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the steps of the training method of the audio-visual speech recognition model and the speech recognition method based on the audio-visual speech recognition model according to any embodiment of the present invention.
[0030] The beneficial effects of the embodiments of the present invention are as follows: Since the encoder learns the context within a certain range of the feature sequence through the cross-modal attention mechanism during the training process, in addition to being able to obtain a lower error rate on pure or noisy audio-visual inputs, it can also resist the problem of audio-visual asynchrony within a certain range. For example, for audio-visual inputs with an asynchrony of 3 to 5 frames, a relatively stable and accurate result can still be obtained. And when the visual modality is missing, that is, it becomes single-modal speech recognition, the audio-visual speech recognition model trained by this method can adapt to single-modal inputs without loss, and its performance can be at a comparable level to that of a conventional single-modal system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 It is a flowchart of a training method of an audio-visual speech recognition model provided by an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of the existing baseline model of a training method of an audio-visual speech recognition model provided by an embodiment of the present invention;
[0034] Figure 3 It is a schematic diagram of the improved model of a training method of an audio-visual speech recognition model provided by an embodiment of the present invention;
[0035] Figure 4 It is a flowchart of a speech recognition method based on an audio-visual speech recognition model provided by an embodiment of the present invention;
[0036] Figure 5 It is a schematic diagram of the comparison of the word error rates of different models of a training method of an audio-visual speech recognition model under different conditions provided by an embodiment of the present invention;
[0037] Figure 6 It is a schematic diagram of the change of the word error rates of the baseline and the model of this method with different audio-visual offsets of a training method of an audio-visual speech recognition model provided by an embodiment of the present invention;
[0038] Figure 7 It is a schematic diagram of attention map visualization in unified cross-modal attention of a training method for an audio-visual speech recognition model provided by an embodiment of the present invention;
[0039] Figure 8 It is a schematic structural diagram of a training system for an audio-visual speech recognition model provided by an embodiment of the present invention;
[0040] Figure 9 It is a schematic structural diagram of a speech recognition system based on an audio-visual speech recognition model provided by an embodiment of the present invention;
[0041] Figure 10 It is a schematic structural diagram of an embodiment of an electronic device for training an audio-visual speech recognition model provided by an embodiment of the present invention. Detailed implementation manners
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] As Figure 1 shown is a flowchart of a training method for an audio-visual speech recognition model provided by an embodiment of the present invention, including the following steps:
[0044] S11: Perform front-end processing on the audio and video in the audio-visual respectively to obtain an acoustic feature sequence of the sound modality composed of an audio modality encoding and an audio position encoding, and a lip movement feature sequence of the video modality composed of a video modality encoding and a video position encoding;
[0045] S12: Perform splicing in the time dimension based on the acoustic feature sequence of the sound modality and the lip movement feature sequence of the video modality to obtain a fused feature sequence, and use the fused feature sequence as the input to the encoder of the audio-visual speech recognition model;
[0046] S13: Through the attention module of the encoder, determine the deep fusion feature encoding of each subsequence with cross-modal context in the fused feature sequence, and input the partial feature encoding of the corresponding audio in the deep fusion feature encoding into the connectionist temporal classification module and decoder of the audio-visual speech recognition model to obtain a predicted recognition result;
[0047] S14: Train the audio-visual speech recognition model according to the benchmark recognition result of the audio in the audio-visual and the predicted recognition result.
[0048] In this embodiment, in order to further improve and train the audio-visual speech recognition model, it is necessary to first describe the basic audio-visual speech recognition model in the prior art, such as Figure 2 As shown, the basic audio-visual speech recognition model in the prior art is a dual-encoding model of an encoder-decoder, which has an independent front end with a hybrid CTC (Connectionist Temporal Classification) / Attention (attention mechanism) structure and a constructed encoder.
[0049] In the acoustic front end, the audio input in the waveform is converted into filter bank features and downsampled to features by a 2D convolutional block The visual front end extracts lip movement features Extract lip movement features from video clips through a three-dimensional convolutional module ResNet (residual) network and a double-layer Bi-LSTM (Bi-directional Long Short-Term Memory) network. Among them, t a and t v are the number of frames in the auditory and visual features, and d is the dimension of the feature space. After adding positional encoding, the encoder encodes them into and representation forms. After upsampling the visual feature sequence to match the time length of the acoustic feature sequence, the representations of the two modules are concatenated along the channel dimension. An MLP (multi-layer perceptron) is connected to project multiple features into which can be input into the CTC and the transformer decoder. This intermediate fusion strategy can be calculated as:
[0050] f a = Encoder(PE(x a )),
[0051] f v = Encoder(PE(x v )),
[0052]
[0053] I = MLP(f av )
[0054] Among them, c represents the connection of the channel dimension. The above is the baseline audio-visual speech recognition model of the prior art. Based on this, this method is improved by adjusting the fusion strategy of audio-visual information. Different from the intermediate fusion strategy of the baseline audio-visual speech recognition model, the fusion is advanced to the beginning stage of the model and changed to concatenation in the time dimension.
[0055] For step S11, as Figure 3 shown, the audio-visual speech recognition model of this method includes two front-ends and an encoder for fusing audio-visual features. The input of the model is audio-visual, and the audio and video are respectively processed by the front-ends.
[0056] In the front-end processing, the acoustic feature sequences of the audio and video in the audio-visual are determined and the visual feature sequences are used to form a unified new sequence feature of the two modalities.
[0057] Specifically, the separate front-end processing of the audio and video in the audio-visual includes:
[0058] Performing short-time Fourier transform on the audio to obtain filter bank features, performing positional encoding on the filter bank features and concatenating the audio modality encoding corresponding to the audio to obtain the acoustic feature sequence of the sound modality; using a residual network to extract the lip movement features of the lip region in the video, performing positional encoding on the lip movement features and concatenating the video modality encoding corresponding to the video to obtain the lip movement feature sequence of the video modality.
[0059] A unified new sequence in two forms is formed For each sub-sequence of each modality, the corresponding PE (Positional Encoding) is calculated respectively. In addition to the positional encoding, learnable ME (Modality embedding) is also added to the acoustic features and visual features accordingly: ME a ∈R 1×d and ME v ∈R 1×d . The corresponding process can be expressed as:
[0060] x’ a =PE(x a )+ME a ,
[0061] x’ v =PE(x v )+ME v .
[0062] For step S12, after obtaining the acoustic feature sequence and the visual feature sequence, they are concatenated in the time dimension to match the size of the unified sequence. The corresponding process can be expressed as:
[0063]
[0064] where denotes concatenation along the time dimension. This fusion strategy requires only a few video stream parameters and does not require comparison between video sequences. The fused feature sequence x av is used as the input to the audio-visual speech recognition model of this method.
[0065] For step S13, using the attention blocks inside the encoder composed of Conformer, the different modalities of the fused feature sequence x av can be distinguished, and the relative order in each subsequence of the fused feature sequence can be noted. Then the encoder directly processes the unified sequence, determines the unrestricted context within a certain range through the cross-modal attention mechanism, and obtains the deep fusion feature encoding f av . The corresponding process can be expressed as: f av = Encoder(x av ). After obtaining the deep fusion feature encoding, at the end of the encoder, only the audio feature part corresponding to the audio is retained, and the feature part corresponding to the video is truncated. The corresponding process can be expressed as: I = f av [1,..., t a . Since the cross-modal information of the sound modality and the video modality has been exchanged in the encoder, considering the recognition of single-modal speech, only the audio feature part is input into the continuous CTC (Connectionist Temporal Classification) and the decoder to obtain the predicted recognition result of the corresponding speech.
[0066] For step S14, the obtained predicted recognition result and the existing benchmark recognition result are used for training. For example, the training data selects LRS3 (a large-scale dataset for visual speech recognition). The audio-visual speech recognition model of this method is trained in a mixed manner, and only the truncated audio feature part is used for model training. Since all the data in LRS3 contains audio and video modalities, the video data is randomly discarded during the training iteration. The input of the visual modality will be ignored once in the forward and backward training, that is, in the i-th iteration:
[0067]
[0068] Trained in this way, the training task will occasionally degrade to an ASR task that is more complex than the AVSR task. Therefore, the training style helps to provide a solid optimization foundation for the audiovisual model, especially the attention mechanism on the unified sequence.
[0069] As Figure 4 shown in the flowchart of a speech recognition method based on an audiovisual speech recognition model provided by an embodiment of the present invention, which includes the following steps:
[0070] S21: Input the received audiovisual file into the audiovisual speech recognition model, where the audiovisual file includes: audiovisual with synchronized sound and picture, and audiovisual with unsynchronized sound and picture within a preset frame rate;
[0071] S22: The audiovisual speech recognition model respectively performs front-end processing on the audio and video in the audiovisual file to obtain an acoustic feature sequence in the sound modality and a lip movement feature sequence in the video modality;
[0072] S23: Concatenate the acoustic feature sequence in the sound modality and the lip movement feature sequence in the video modality in the time dimension, and input the concatenated fusion feature sequence into the encoder;
[0073] S24: The depth fusion feature encoding of each subsequence with cross-modal context output by the encoder, and retain the partial feature encoding corresponding to the audio in the depth fusion feature encoding;
[0074] S25: Input the partial feature encoding corresponding to the audio into the connectionist temporal classification module and the decoder to obtain the speech recognition result of the audiovisual file.
[0075] For step S21, this method can be installed on various intelligent voice devices, such as smart TVs, smartphones, smart speakers, etc. For example, some current high-end smart TVs are equipped with cameras. When a user interacts with a smart TV, the smart TV not only collects the user's voice through a microphone, but also uses the equipped camera to collect video information corresponding to the voice. Thus, an audiovisual file is obtained. Due to network or other factors, an audiovisual with a slightly small error between the video and the audio may be collected, such as within 12 frames.
[0076] As an implementation manner, the audiovisual file further includes: single-modal pure audio lacking the visual modality. Specifically, for example, when this method is installed on a smart speaker or a recording pen and other smart devices without a camera, such smart devices without a camera are also commonly used in the prior art. Such devices will only collect speech files in the sound modality when in use.
[0077] For steps S22 - S25, the trained audio - video speech recognition model is used to perform front - end processing on the audio - video (the specific process of front - end processing has been described in step S11 and will not be elaborated here), obtaining the acoustic feature sequence of the audio modality and the lip - movement feature sequence of the video modality. The acoustic feature sequence of the audio modality and the lip - movement feature sequence of the video modality are concatenated in the time dimension (the concatenation process has been described in step S12 and will not be elaborated here). The fused feature sequence obtained by concatenation is input into the encoder, and the depth - fused feature encoding of each subsequence with cross - modal context is output. The features corresponding to the video in the depth - fused feature encoding are truncated, and the partial feature encoding corresponding to the audio is retained. The connectionist temporal classification module and the decoder are used to identify the partial feature encoding corresponding to the audio, obtaining the speech recognition result of the audio - video file. If there is only audio in the audio - video file, then the audio - video speech recognition model trained by this method can be directly used for single - modality speech recognition processing.
[0078] From the training of the audio - video speech recognition model and the implementation of speech recognition based on the audio - video speech recognition model, it can be seen that since the encoder learns the context within a certain range of the feature sequence through the cross - modal attention mechanism during the training process, in addition to being able to obtain a lower error rate on pure or noisy audio - video inputs, it can also resist the problem of audio - video asynchrony within a certain range. For example, for audio - video inputs with an asynchrony of 3 to 5 frames, a relatively stable and accurate result can still be obtained. And when the visual modality is missing, that is, it becomes single - modality speech recognition, the audio - video speech recognition model trained by this method can adapt to single - modality inputs without loss, and its performance can be comparable to that of conventional single - modality systems.
[0079] A specific experiment of this method is described. The model training and testing are carried out on LRS3. The LRS3 dataset contains 433 hours of utterances and their corresponding video clips and text transcripts. The dataset is officially divided into a pre - training set, a training set, and a test set. 10,000 samples are separated from the training set as validation data, and the remaining samples are combined with the samples in the pre - training set as training data. Any sample with a duration less than 1 second is removed from the dataset.
[0080] To simulate the real environment, the noise in the WHAM noise dataset is used, and noise samples are generated by mixing clean utterances and noise with SNRs uniformly sampled from [-6, 6]. These noisy samples are used for training together with the clean samples. Noise samples are also generated using validation / test utterances and unseen noise with SNRs of {-10, -5, 0, 5, 10, 15, 20}, and they are used for validation / testing together with the clean data.
[0081] Most hyperparameters are shared among all consistency-based models. In the acoustic front-end, the STFT (short-time Fourier transform) has n-fft = 512, window size = 400, and hop length = 160. The encoder consists of 12 conformer blocks, the hidden size d of the fully connected layer fc = 1024, the hidden size d of the attention module att = 256, the number of attention heads n head = 4, and the kernel size k of the convolutional module is 31. The decoder consists of 6 transformer blocks, d fc = 2048, n head = 6, and the size of the acoustic unit is 5000. Relative position encoding is used for the conformer blocks, and absolute position encoding is used for the transformer blocks. The pm of the hybrid training is empirically set to 0.35.
[0082] The visual front-end is pre-trained on the LRW dataset. 512-dimensional visual features are extracted in advance, and the visual front-end is excluded during the accelerated training. All trainable parameters are randomly initialized and optimized by the Adam optimizer using a warm-up learning rate scheduler. When the number of elements in the sound input is fixed at 45 million, the batch size is dynamically set. The model is trained for 80 epochs at a peak learning rate of 0.002 at the 15000th step. SpecAug data is used to augment the audio input. For visual input, a similar strategy is adopted to enhance the extracted feature vectors. To obtain a powerful language model, the text in the LibriSpeech-960h dataset and the text in the LRS3 training data are used as the training corpus. The language model of the baseline transformer is trained for 25 stages and gets 54 on the LRS3 test data. The weight of the language model for decoding is empirically set to 0.2.
[0083] In addition to testing the language model of this method and the existing baseline transformer, this method also tests the data results of other models, which are fusion models where the audio-visual features are concatenated along the channel dimension before a single encoder. It has 12 encoding blocks, but additional data from the VoxCeleb2 dataset is used for preprocessing and then fine-tuned on the LRS3 dataset.
[0084] As Figure 5As shown in the central part, the model of this method reduces the word error rate (WER) of clean samples from 2.3% to 2.1%. Although generally it is very difficult to improve this WER because lip movements are more confusing than clean speech. If hybrid training is eliminated, the performance on noisy speech will deteriorate slightly, while on clean speech, the WER will deteriorate severely. Due to the complementary effect of the visual modality, the performance of the baseline dual-encoder model under noisy conditions is much better than that of the audio-only baseline model. However, the model of this method significantly improves the performance by 23% again, demonstrating the advantage of the proposed unified cross-modal attention mechanism in utilizing visual information. The shared encoder model performs poorly on clean data but has good performance when the noise is extremely strong (SNR (SIGNAL-NOISE RATIO) is below 0 dB), which indicates that the early fusion strategy is naturally suitable for audiovisual fusion.
[0085] When the auxiliary visual modality is missing due to environmental interference, equipment failure, or other reasons (which is common in real scenarios), it is expected that the audiovisual model can continue to work and bring satisfactory performance through audio input. The model of this method is tested with other models on pure audio samples. The model of this method does not need to modify the input features, but for other models, the visual modality should be filled with zeros. As Figure 5 the bottom reports the experimental results. Compared with other models, the model trained by this method obtains the best score, and in this case, it is even much better than the pre-trained AVHubert. It is worth noting that in clean speech, the WER can still be maintained at 2.4%.
[0086] A natural advantage of the unified cross-modal attention mechanism of this method is that there is no need to manually enforce frame-level alignment of audiovisual sequences. To verify this hypothesis, after artificially shifting the visual input sequence by one frame offset in [-5, 5], that is, from about -200 to 200 milliseconds, the WER is measured. Figure 6 Shows the performance changes of the model of this method and the dual-encoder baseline model. For the bottom part of the clean test samples, the model of this method maintains stable performance even without this increase in misalignment in the training data, while the performance of the dual-encoder model jitters with the change of the offset. For the top of the noisy test samples, when the offset is large enough, the WER of both models drops to the same level, which means that severe misalignment will completely destroy the effectiveness of the visual modality. However, for an offset of [-3, 3] frames, the model of this method does not deteriorate as severely as the baseline model.
[0087] To conduct a qualitative evaluation of the unified cross-modal attention mechanism, in Figure 7The attention maps from the multi-head attention module are visualized. Four attention heads come from the last encoding block and they behave quite differently. Besides the normal attention within each modality, there is cross-modal attention to exchange information. It is worth mentioning that as the audio regions are more frequently attended to in the attention maps, the acoustic features are shown to be more prominent in the attention module. This is consistent with the hybrid training style that keeps the audio dominant and the video auxiliary. The interactions between and within modalities indicate that the model is able to learn alignment at the context level. Figure 1 Consistent. The interactions between and within modalities indicate that the model is able to learn alignment at the context level.
[0088] Overall, this method explores a new fusion mechanism for the audiovisual speech recognition task, where the input sequences from two modalities are concatenated along the time dimension in a unified space at the early stage of the model. The large and flexible cross-modal context with the hybrid training style contributes to the adaptive fusion of audiovisual information. On the large-scale LRS3 dataset, compared with the state-of-the-art, the proposed model with unified cross-modal attention significantly improves the performance in clean and noisy environments. The experiments also demonstrate the robustness of the model trained by this method to the possible absence of the visual modality or frame misalignment in the audiovisual stream.
[0089] As Figure 8 shown is a schematic structural diagram of a training system for an audiovisual speech recognition model provided by an embodiment of the present invention. The system can execute the training method of the audiovisual speech recognition model described in any of the above embodiments and is configured in a terminal.
[0090] A training system 10 for an audiovisual speech recognition model provided in this embodiment includes: a front-end processing program module 11, a fusion program module 12, an encoding program module 13, and a training program module 14.
[0091] Among them, the front-end processing program module 11 is used to perform front-end processing on the audio and video in the audiovisual respectively, to obtain an acoustic feature sequence of the sound modality composed of audio modality encoding and audio position encoding, and a lip movement feature sequence of the video modality composed of video modality encoding and video position encoding; the fusion program module 12 is used to perform splicing in the time dimension based on the acoustic feature sequence of the sound modality and the lip movement feature sequence of the video modality, to obtain a fusion feature sequence, and use the fusion feature sequence as the input to the encoder of the audiovisual speech recognition model; the encoding program module 13 is used to determine the deep fusion feature encoding of each subsequence with cross-modal context in the fusion feature sequence through the attention module of the encoder, and input the partial feature encoding corresponding to the audio in the deep fusion feature encoding into the connectionist temporal classification module and decoder of the audiovisual speech recognition model to obtain a predicted recognition result; the training program module 14 is used to train the audiovisual speech recognition model according to the reference recognition result of the audio in the audiovisual and the predicted recognition result.
[0092] An embodiment of the present invention further provides a non-volatile computer storage medium, which stores computer-executable instructions, and the computer-executable instructions can execute the training method of the audio-visual speech recognition model in any of the above method embodiments;
[0093] As an implementation, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as follows:
[0094] Perform front-end processing on the audio and video in the audio-visual respectively to obtain an acoustic feature sequence of the sound modality composed of audio modality encoding and audio position encoding, and a lip movement feature sequence of the video modality composed of video modality encoding and video position encoding;
[0095] Perform splicing in the time dimension based on the acoustic feature sequence of the sound modality and the lip movement feature sequence of the video modality to obtain a fusion feature sequence, and use the fusion feature sequence as the input of the encoder of the audio-visual speech recognition model;
[0096] Through the attention module of the encoder, determine the deep fusion feature encoding of each subsequence with cross-modal context in the fusion feature sequence, and input the partial feature encoding of the corresponding audio in the deep fusion feature encoding into the connectionist temporal classification module and decoder of the audio-visual speech recognition model to obtain a predicted recognition result;
[0097] Train the audio-visual speech recognition model according to the reference recognition result of the audio in the audio-visual and the predicted recognition result.
[0098] As Figure 9 shown is a schematic structural diagram of a speech recognition system based on an audio-visual speech recognition model provided by an embodiment of the present invention. The system can execute the speech recognition method based on the audio-visual speech recognition model described in any of the above embodiments and is configured in a terminal.
[0099] A speech recognition system 20 based on an audio-visual speech recognition model provided in this embodiment includes: an audio-visual receiving program module 21, a front-end processing program module 22, a fusion program module 23, an encoding program module 24, and a speech recognition program module 25.
[0100] Among them, the audio-video receiving program module 21 is used to input the received audio-video file into the audio-video speech recognition model. Among them, the audio-video file includes: audio-video with synchronized audio and video, and audio-video with unsynchronized audio and video within a preset frame rate; the front-end processing program module 22 is used to perform front-end processing on the audio and video in the audio-video file by the audio-video speech recognition model to obtain an acoustic feature sequence in the sound modality and a lip movement feature sequence in the video modality; the fusion program module 23 is used to splice the acoustic feature sequence in the sound modality and the lip movement feature sequence in the video modality in the time dimension, and input the spliced fusion feature sequence into the encoder; the encoding program module 24 is used to perform deep fusion feature encoding on each subsequence with cross-modal context output by the encoder, and retain partial feature encoding corresponding to the audio in the deep fusion feature encoding; the speech recognition program module 25 is used to input the partial feature encoding corresponding to the audio into the connectionist temporal classification module and the decoder to obtain the speech recognition result of the audio-video file.
[0101] An embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the speech recognition method based on the audio-video speech recognition model in any of the above method embodiments;
[0102] As an implementation, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:
[0103] Input the received audio-video file into the audio-video speech recognition model. Among them, the audio-video file includes: audio-video with synchronized audio and video, and audio-video with unsynchronized audio and video within a preset frame rate;
[0104] The audio-video speech recognition model performs front-end processing on the audio and video in the audio-video file respectively to obtain an acoustic feature sequence in the sound modality and a lip movement feature sequence in the video modality;
[0105] Splice the acoustic feature sequence in the sound modality and the lip movement feature sequence in the video modality in the time dimension, and input the spliced fusion feature sequence into the encoder;
[0106] Perform deep fusion feature encoding on each subsequence with cross-modal context output by the encoder, and retain partial feature encoding corresponding to the audio in the deep fusion feature encoding;
[0107] Input the partial feature encoding corresponding to the audio into the connectionist temporal classification module and the decoder to obtain the speech recognition result of the audio-video file.
[0108] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium and, when executed by a processor, execute the training method of the audio-visual speech recognition model and the speech recognition method based on the audio-visual speech recognition model in any of the above method embodiments.
[0109] Figure 10 It is a schematic diagram of the hardware structure of an electronic device for the training method of the audio-visual speech recognition model provided in another embodiment of the present application, as Figure 10 shown. The device includes:
[0110] One or more processors 1010 and a memory 1020, Figure 10 Taking one processor 1010 as an example. The device for the training method of the audio-visual speech recognition model may further include: an input device 1030 and an output device 1040.
[0111] The processor 1010, the memory 1020, the input device 1030, and the output device 1040 may be connected through a bus or other means, Figure 10 Taking connection through a bus as an example.
[0112] The memory 1020, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the training method of the audio-visual speech recognition model in the embodiments of the present application. The processor 1010 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 1020, that is, implements the training method of the audio-visual speech recognition model and the speech recognition method based on the audio-visual speech recognition model in the above method embodiments.
[0113] The memory 1020 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data, etc. In addition, the memory 1020 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 1020 may optionally include a memory remotely set relative to the processor 1010, and these remote memories can be connected to the mobile device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0114] The input device 1030 can receive input digital or character information. The output device 1040 can include display devices such as a display screen.
[0115] The one or more modules are stored in the memory 1020 and, when executed by the one or more processors 1010, perform the training method of the audio-video speech recognition model and the speech recognition method based on the audio-video speech recognition model in any of the above method embodiments.
[0116] The above product can execute the method provided in the embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.
[0117] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0118] An embodiment of the present invention also provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the training method of the audio-video speech recognition model and the speech recognition method based on the audio-video speech recognition model in any embodiment of the present invention.
[0119] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:
[0120] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0121] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as tablet computers.
[0122] (3) Portable entertainment devices: Such devices can display and play multimedia content. This type of device includes: audio and video players, handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.
[0123] (4) Other electronic devices with data processing functions.
[0124] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or also include elements inherent to such a process, method, article, or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article, or device that includes the said elements.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0127] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for an audio - video speech recognition model, comprising: Perform front-end processing on the audio and video in the audio-visual data respectively to obtain an acoustic feature sequence of the sound modality composed of an audio modality encoding and an audio position encoding, and a lip movement feature sequence of the video modality composed of a video modality encoding and a video position encoding; Based on the acoustic feature sequence of the sound modality and the lip movement feature sequence of the video modality, perform splicing in the time dimension to obtain a fused feature sequence, and use the fused feature sequence as the input to the encoder of the audio-visual speech recognition model; Through the attention module of the encoder, determine the deep fusion feature encoding of each subsequence with cross-modal context in the fused feature sequence, and input the partial feature encoding of the corresponding audio in the deep fusion feature encoding into the connectionist temporal classification module and the decoder of the audio-visual speech recognition model to obtain a predicted recognition result; Train the audio-visual speech recognition model according to the benchmark recognition result of the audio in the audio-visual data and the predicted recognition result; 2. The method according to claim 1, wherein, The front-end processing of the audio and video in the audio-visual data respectively includes: Perform short-time Fourier transform on the audio to obtain filter bank features, perform position encoding on the filter bank features, and splice the audio modality encoding corresponding to the audio to obtain the acoustic feature sequence of the sound modality; Use a residual network to extract the lip movement features of the lip region in the video, perform position encoding on the lip movement features, and splice the video modality encoding corresponding to the video to obtain the lip movement feature sequence of the video modality; 3. A speech recognition method based on an audio - video speech recognition model, comprising: Input the received audio-visual file into The audio-visual speech recognition model obtained by the training method according to any one of claims 1-2, wherein the audio-visual file includes: audio-visual data with audio-visual synchronization, audio-visual data with audio-visual asynchronization within a preset frame rate; The audio-visual speech recognition model performs front-end processing on the audio and video in the audio-visual file respectively to obtain an acoustic feature sequence of the sound modality and a lip movement feature sequence of the video modality; Perform splicing in the time dimension on the acoustic feature sequence of the sound modality and the lip movement feature sequence of the video modality, and input the spliced fused feature sequence into the encoder; For each subsequence of the deep fusion feature encoding with cross-modal context output by the encoder, retain the partial feature encoding of the corresponding audio in the deep fusion feature encoding; Input the partial feature encoding of the corresponding audio into the connectionist temporal classification module and the decoder to obtain the speech recognition result of the audio-visual file; 4. The method according to claim 3, wherein, The audio-visual file further includes: unimodal pure audio lacking the visual modality; The method further includes: performing unimodal speech recognition on the unimodal pure audio lacking the visual modality based on the audio-visual speech recognition model to obtain the speech recognition result of the unimodal pure audio; 5. A training system for an audio - video speech recognition model, comprising: A front-end processing program module for performing front-end processing on the audio and video in the audio-visual data respectively to obtain an acoustic feature sequence of the sound modality composed of an audio modality encoding and an audio position encoding, and a lip movement feature sequence of the video modality composed of a video modality encoding and a video position encoding; A fusion program module for performing splicing in the time dimension based on the acoustic feature sequence of the voice modality and the lip movement feature sequence of the video modality to obtain a fused feature sequence, and using the fused feature sequence as the input to the encoder of an audiovisual speech recognition model; An encoding program module for determining, through the attention module of the encoder, the deep fusion feature encoding of each subsequence with cross-modal context in the fused feature sequence, and inputting the partial feature encoding of the corresponding audio in the deep fusion feature encoding into the connectionist temporal classification module and decoder of the audiovisual speech recognition model to obtain a predicted recognition result; A training program module for training the audiovisual speech recognition model according to the benchmark recognition result of the audio in the audiovisual and the predicted recognition result.
6. The system according to claim 5, wherein, The front-end processing program module is used for: Performing short-time Fourier transform on the audio to obtain filter bank features, performing position encoding on the filter bank features, and splicing the audio modality encoding corresponding to the audio to obtain the acoustic feature sequence of the voice modality; Using a residual network to extract the lip movement features of the lip region in the video, performing position encoding on the lip movement features, and splicing the video modality encoding corresponding to the video to obtain the lip movement feature sequence of the video modality.
7. A speech recognition system based on an audio - video speech recognition model, comprising: An audiovisual receiving program module for inputting the received audiovisual file into the audiovisual speech recognition model obtained by the training system according to any one of claims 5-6, wherein the audiovisual file includes: audiovisual with synchronized audio and video, and audiovisual with asynchronous audio and video within a preset frame rate; A front-end processing program module for the audiovisual speech recognition model to perform front-end processing on the audio and video in the audiovisual file respectively to obtain the acoustic feature sequence of the voice modality and the lip movement feature sequence of the video modality; A fusion program module for splicing the acoustic feature sequence of the voice modality and the lip movement feature sequence of the video modality in the time dimension, and inputting the spliced fused feature sequence into the encoder; An encoding program module for the deep fusion feature encoding of each subsequence with cross-modal context output by the encoder, and retaining the partial feature encoding of the corresponding audio in the deep fusion feature encoding; A speech recognition program module for inputting the partial feature encoding of the corresponding audio into the connectionist temporal classification module and decoder to obtain the speech recognition result of the audiovisual file.
8. The system according to claim 7, wherein, The audiovisual file further includes: unimodal pure audio lacking the visual modality; The speech recognition program module is further used for: performing unimodal speech recognition on the unimodal pure audio lacking the visual modality based on the audiovisual speech recognition model to obtain the speech recognition result of the unimodal pure audio.
9. An electronic device, which comprises: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method according to any one of claims 1-4.
10. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Quantum, biological, computer vision, and neural network systems for industrial internet of things
CA3177620A1
Screen splicing synchronization method and system
CN111586453A