A cross-modal lip reading recognition method
Through the self-supervised cross-modal comparison learning and attribute learning module, better visual features are extracted using audio information, which solves the problem that traditional lip recognition methods rely on a large amount of labeled data and poor generalization capabilities, and achieves more efficient lip recognition performance and cross-speaker robustness.
Patent Information
- Application Number
- CN202110941080.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-08-17
AI Technical Summary
Traditional lip recognition methods rely on a large amount of accurate label data, and have poor generalization capabilities for different speaker data, making it difficult to learn better visually separable features without additional labeled data.
The cross-modal lip recognition method is adopted, through self-supervised cross-modal contrast learning, the audio information is used to help the lip recognition branch extract better distinguishable visual features from the video sequence, and the lip characteristics of different speakers are standardized in the attribute learning module.
Without the need for additional artificial annotation data, the performance of lip recognition is improved, and the lip video sequences with different pronunciations but similar lip shapes can be better distinguished, and has better cross-speaker robustness.
Smart Images

Figure CN113851131B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of recognition, and in particular to a cross-modal lip reading recognition method. Background Art
[0002] Lip reading recognition is a visual language recognition technology that mainly uses lip movement information in the video, combined with language prior knowledge and contextual information. Lip reading recognition plays an important role in language understanding and communication, and is often used when effective audio information is not available. It also has extremely high application value and can be used in the treatment of patients with speech disorders, security fields, military equipment and human-computer interaction.
[0003] The limitation of traditional lip reading recognition methods is that they only focus on video input information and cannot learn good visually separable features without additional experience and knowledge. Therefore, these methods usually rely on a large amount of accurately labeled data, but in real life, the cost of obtaining labeled data is extremely high. Summary of the invention
[0004] In view of the above problems, the present invention aims to provide a cross-modal lip reading recognition method, comprising:
[0005] S1, data preprocessing:
[0006] For video data, we first identify 68 key points of the face, normalize each face image to a frontal view through affine transformation, and finally crop the lip area;
[0007] For audio data, it is first downsampled to 16kHz and converted into Mel-frequency cepstral coefficient features. Then, the Mel-frequency cepstral coefficient vectors at all times are normalized and formed into a feature matrix in chronological order.
[0008] S2, model training:
[0009] S21, inputting the paired video data and audio data into the visual recognition branch and the speech recognition branch respectively, and performing speaker recognition task training in the attribute learning module of each branch;
[0010] S22, input the paired video data and audio data into the visual recognition branch and the speech recognition branch respectively, in the contrastive learning module shared by the two branches, use the representation obtained by the speaker recognition task to standardize the semantic features, and then perform audio and video cross-modal contrastive learning;
[0011] S23, only input the audio sequence, remove the speaker's timbre characteristics, standardize the speech features, and use the back propagation algorithm to update the model parameters of the speech recognition branch to ensure that the intermediate audio features S involved in contrastive learning are correct;
[0012] S24, inputting only the video sequence, removing the speaker's lip shape characteristics, standardizing the lip reading features, and updating the model parameters of the lip reading recognition branch using the back propagation algorithm;
[0013] Repeat S21-S24 above until the loss function value no longer decreases in multiple rounds of training after the learning rate decays, that is, the model converges; S3, model deployment:
[0014] Only non-training data video sequences to be recognized are input, and the visual recognition branch is used to remove the speaker's lip shape features and standardize the lip reading features. Finally, the lip reading features are mapped to text.
[0015] Preferably, the visual recognition branch includes a 3D convolution module, a first recursive neural network module, a first speaker feature extraction module, a first attribute learning module, a contrastive learning module, a second recursive neural network module, a first attention module and a first mapping module;
[0016] The 3D convolution module is used to obtain short-term features of lip movements;
[0017] The first recursive neural network module is used to establish a long-term dependency relationship of lip movements;
[0018] The first speaker feature extraction module is used to extract lip shape features of different speakers;
[0019] The first attribute learning module is used to eliminate the lip shape differences of different speakers by using the acquired speaker lip shape features;
[0020] The contrastive learning module is used to use a self-supervised contrastive learning method across audio and video data, so that the model obtains prior knowledge from another form of expression of the video data itself, audio, and guides the visual recognition branch to learn lip shape features;
[0021] The second recursive neural network module is used to strengthen the contextual relationship of the video intermediate feature S sequence after the contrast learning layer;
[0022] The first attention module is used to help the model ignore irrelevant video frames by assigning different weights to different time point features output by the second recurrent neural network module in the time domain;
[0023] The first mapping layer is used to map the final lip motion features output by the first attention module into the text domain.
[0024] Preferably, the speech recognition branch includes:
[0025] 2D convolution module, third recursive neural network module, second speaker feature extraction module, second attribute learning module, contrastive learning module, fourth recursive neural network module, second attention module and second mapping module;
[0026] The 2D convolution module is used to extract short-term speech features from the Mel-frequency cepstrum features;
[0027] The third recursive neural network module is used to establish a long-term dependency relationship of speech features;
[0028] The second speaker feature extraction module is used to extract timbre features of different speakers;
[0029] The second attribute learning module is used to eliminate the timbre differences between different speakers by using the acquired speaker timbre features;
[0030] The fourth recursive neural network module is used to strengthen the contextual relationship of the audio intermediate feature S sequence after the contrast learning module;
[0031] The second attention module is used to help the model ignore irrelevant audio segments by assigning different weights to different time point features output by the fourth recurrent neural network module in the time domain;
[0032] The second mapping module is used to map the final audio features output by the second attention module into the text domain.
[0033] Preferably, the first mapping layer comprises a classifier based on nonlinear mapping of a multilayer perceptron with a ReLU activation function.
[0034] Preferably, a connectionist temporal classification loss function is used to constrain the visual recognition branch and the speech recognition branch respectively.
[0035] This method uses a self-supervised cross-modal contrastive learning method, without the need for additional manually labeled data. It uses audio information to help the lip reading recognition branch extract visual features with better distinguishability from the input video sequence, and based on this, distinguish lip reading video sequences with different pronunciations but similar lip shapes.
[0036] Compared with the two-stage traditional lip reading recognition method, this method builds an end-to-end lip reading recognition system based on deep learning. The feature extraction has better generalization and robustness, can be used across speakers, and there is no need to train a set of model parameters separately for each category sample.
[0037] Traditional methods have poor generalization ability for data from different speakers, while this method uses attribute learning to standardize the lip reading features from different speakers, greatly improving the robustness of the algorithm in dealing with the lip shapes of different speakers.
[0038] This method basically does not require manual labeling. Instead, it uses audio modal information as a guide. Through an end-to-end cross-audio and video data self-supervised learning method, it helps the lip reading model obtain better visual features under the guidance of audio information, thereby improving the performance of the algorithm in lip reading recognition problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.
[0040] Figure 1 , which is a diagram of an exemplary embodiment of a model of a cross-modal lip reading recognition method of the present invention.
[0041] Figure 2 , which is a flow chart of the model training steps of the present invention. DETAILED DESCRIPTION
[0042] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0043] The present invention provides a cross-modal lip reading recognition method, comprising:
[0044] S1, data preprocessing:
[0045] For video data, we first identify 68 key points of the face, normalize each face image to a frontal view through affine transformation, and finally crop the lip area;
[0046] For audio data, it is first downsampled to 16kHz and converted into Mel-frequency cepstral coefficient features. Then, the Mel-frequency cepstral coefficient vectors at all times are normalized and formed into a feature matrix in chronological order.
[0047] S2, model training:
[0048] S21, inputting the paired video data and audio data into the visual recognition branch and the speech recognition branch respectively, and performing speaker recognition task training in the attribute learning module of each branch;
[0049] S22, input the paired video data and audio data into the visual recognition branch and the speech recognition branch respectively, and in the contrastive learning module shared by the two branches, use the representation obtained by the speaker recognition task to standardize the semantic features. Then perform cross-modal contrastive learning of audio and video;
[0050] S23, only input the audio sequence, remove the speaker's timbre characteristics, standardize the speech features, and use the back propagation algorithm to update the model parameters of the speech recognition branch to ensure that the intermediate audio features S involved in contrastive learning are correct;
[0051] S24, inputting only the video sequence, removing the speaker's lip shape characteristics, standardizing the lip reading features, and updating the model parameters of the lip reading recognition branch using the back propagation algorithm;
[0052] Repeat S21-S24 above until the loss function value no longer decreases in multiple rounds of training after the learning rate decays, that is, the model converges; S3, model deployment:
[0053] Only non-training data video sequences to be recognized are input, and the visual recognition branch is used to remove the speaker's lip shape features and standardize the lip reading features. Finally, the lip reading features are mapped to text.
[0054] This method basically does not require manual labeling. Instead, it uses audio modal information as a guide. Through an end-to-end cross-audio and video data self-supervised learning method, it helps the lip reading model obtain better visual features under the guidance of audio information, thereby improving the performance of the algorithm in lip reading recognition problems.
[0055] Preferably, the visual recognition branch includes a 3D convolution module, a first recursive neural network module, a first speaker feature extraction module, a first attribute learning module, a contrastive learning module, a second recursive neural network module, a first attention module and a first mapping module;
[0056] The 3D convolution module is used to obtain short-term features of lip movements;
[0057] The first recursive neural network module is used to establish a long-term dependency relationship of lip movements;
[0058] The first speaker feature extraction module is used to extract lip shape features of different speakers;
[0059] The first attribute learning module is used to eliminate the lip shape differences of different speakers by using the acquired speaker lip shape features;
[0060] The contrastive learning module is used to use a self-supervised contrastive learning method across audio and video data, so that the model obtains prior knowledge from another form of expression of the video data itself, audio, and guides the visual recognition branch to learn lip shape features;
[0061] The second recursive neural network module is used to strengthen the contextual relationship of the video intermediate feature S sequence after the contrast learning layer;
[0062] The first attention module is used to help the model ignore irrelevant video frames by assigning different weights to different time point features output by the second recurrent neural network module in the time domain;
[0063] The first mapping layer is used to map the final lip motion features output by the first attention module into the text domain.
[0064] Specifically, the data input and output relationship between each module is as follows:
[0065] Video sequence to be identified -> 3D convolution module -> short-term features of lip movements;
[0066] Short-term features of lip movements -> first recurrent neural network module -> long-term dependencies of lip movements, overall features of lip sequences;
[0067] Overall features of lip sequence -> first speaker feature extraction module -> lip shape features of different speakers;
[0068] Lip shape features and long-term dependencies of lip movements of different speakers -> First attribute learning module -> Long-term dependencies of lip movements that eliminate individual differences;
[0069] Eliminate individual differences in lip movement long-term dependencies, eliminate individual differences in audio long-term dependencies -> contrastive learning module -> more discriminative lip movement features, audio intermediate features;
[0070] More discriminative lip movement features -> Second recurrent neural network module -> More contextually relevant, highly discriminative lip movement features;
[0071] Highly discriminative lip movement features with closer contextual connections -> First attention module -> Ignore silent lip movement features;
[0072] Ignore unpronounced lip movement features -> first mapping module -> text.
[0073] Preferably, the speech recognition branch includes:
[0074] 2D convolution module, third recursive neural network module, second speaker feature extraction module, second attribute learning module, contrastive learning module, fourth recursive neural network module, second attention module and second mapping module;
[0075] The 2D convolution module is used to extract short-term speech features from the Mel-frequency cepstrum features;
[0076] The third recursive neural network module is used to establish a long-term dependency relationship of speech features;
[0077] The second speaker feature extraction module is used to extract timbre features of different speakers;
[0078] The second attribute learning module is used to eliminate the timbre differences between different speakers by using the acquired speaker timbre features;
[0079] The fourth recursive neural network module is used to strengthen the contextual relationship of the audio intermediate feature S sequence after the contrast learning module;
[0080] The second attention module is used to help the model ignore irrelevant audio segments by assigning different weights to different time point features output by the fourth recurrent neural network module in the time domain;
[0081] The second mapping module is used to map the final audio features output by the second attention module into the text domain.
[0082] Specifically, the data input and output relationship between each module is as follows:
[0083] Mel cepstral coefficient feature sequence of the audio to be identified -> 2D convolution module -> audio short-time features;
[0084] Audio short-term features -> the third recursive neural network module -> audio long-term dependencies, overall features of audio sequences;
[0085] Overall features of audio sequence -> second speaker feature extraction module -> timbre features of different speakers;
[0086] The timbre characteristics and long-term audio dependencies of different speakers -> second attribute learning module -> long-term audio dependencies that eliminate individual differences;
[0087] Eliminate individual differences in lip movement long-term dependencies, eliminate individual differences in audio long-term dependencies -> contrastive learning module -> more discriminative lip movement features, audio intermediate features;
[0088] Audio intermediate features -> fourth recurrent neural network module -> audio intermediate features with closer contextual connections;
[0089] Audio intermediate features with closer context -> second attention module -> ignore unpronounced audio intermediate features ignore unpronounced audio intermediate features -> second mapping module -> text.
[0090] Preferably, the first mapping layer comprises a classifier based on nonlinear mapping of a multilayer perceptron with a ReLU activation function.
[0091] Preferably, a connectionist temporal classification loss function is used to constrain the visual recognition branch and the speech recognition branch respectively.
[0092] This method uses an end-to-end trained neural network to implement lip reading recognition. Figure 1 As shown in the figure, the model as a whole consists of two independent branches. The right branch is responsible for lip reading recognition, and the left branch is responsible for speech recognition. The core idea of the algorithm is: based on the self-supervised contrastive learning method, the audio information with better discrimination is used to improve the model's ability to distinguish visual input signals, namely lip movements or lip shape features. Figure 2 , which is the model training step of the present invention.
[0093] In the visual recognition branch on the right,
[0094] We first apply a 3D convolution module to extract short-term dependency features of lip motion from video sequences, and apply ReLU activation function and max pooling layer after the convolution layer.
[0095] Since the 3D convolution module contains more parameters, it is very easy to overfit on a small-scale dataset. Therefore, we also apply a dropout layer to alleviate the overfitting problem.
[0096] As shown in the lip reading recognition branch on the right side of the overall structure diagram, after using the 3D convolution module to obtain the short-term features of the lip movements, we use a layer of bidirectional GRU, that is, the first recursive neural network module to establish the long-term dependencies of the lip movements.
[0097] Compared with unidirectional recurrent networks, bidirectional recurrent networks can model the forward and reverse order of sequences to obtain richer semantic information from the sequences. Compared with LSTM, the use of GRU reduces the number of parameters to a certain extent, further alleviating the overfitting problem.
[0098] In the speech recognition branch on the left side of the overall structure diagram,
[0099] The audio signal is converted into Mel-cepstral coefficients and input into the branch. Since the converted Mel-cepstral features are in the form of a two-dimensional matrix, we simplify the 3D convolution part used to extract short-term temporal features in the visual recognition branch into a 2D convolution and apply it to the speech recognition branch, while keeping the rest of the branch consistent with the visual recognition branch.
[0100] After modeling the long-term relationship between video and audio, as shown in the middle of the overall structure diagram (between the upper and lower GRU layers, where the contrast learning module CL and the attribute learning module AL are applied),
[0101] The obtained intermediate feature S is subjected to attribute learning to normalize the lip shape differences of different speakers so that the model can obtain robust features across speakers.
[0102] Then, the self-supervised contrastive learning method across audio and video data is used to enable the model to obtain a certain degree of prior knowledge from another form of expression (audio) of the video data itself, and guide the visual recognition branch to learn lip shape features with better distinguishability.
[0103] Then, as shown in the lower part of the structure diagram, we again use a layer of bidirectional GRU to strengthen the contextual relationship of the sequence, and use the attention module in the time domain to help the model ignore irrelevant video frames by assigning different weights to features at different time points.
[0104] Finally, we map the lip motion features learned by the model to the text domain. Because the mapping from lip shape to text does not satisfy the relationship of single injection and surjection, we use a multi-layer perceptron (MLP) with a ReLU activation function to design a nonlinear mapping classifier, as shown in the bottom rectangle of the figure.
[0105] Since the input and output text lengths of lip reading and speech recognition are different, there is an alignment problem. Therefore, we use the Connectionist Temporal Classification loss function to constrain the two network branches separately.
[0106] In order to obtain more robust cross-speaker lip reading features, we designed an attribute learning module (AL in the overall structure) in the algorithm to standardize lip features from different speakers. This attribute learning module is applied to both the video and speech recognition branches. Generally speaking, the final hidden layer features of the GRU contain the speaker's attribute information, emotional information, etc. As shown in the figure, we input the final output features of the GRU into the attribute learning module, and learn how to classify speakers through the overall features of the sequence under the supervision of the speaker label, that is, the speaker classification result is output below the AL module shown in the figure. In the case of this branch training journey, the intermediate features of the AL module are used as representations of the speaker information, and their transformation is used to standardize the lip reading features output by the GRU at each moment, as shown by the arrows in the figure.
[0107] The reason why lip reading recognition is so difficult is that there are few lip shapes that can be clearly distinguished. Lip shapes can only be represented by 14 visemes, while audio signals have 42 phonemes to represent speech. Therefore, compared with lip reading features, speech features are naturally more distinguishable, especially when the speaker says words with similar mouth shapes but different pronunciations. Therefore, using audio to guide video learning is an effective and feasible solution. In order to obtain more distinguishable lip features, we introduce audio features to improve the recognition ability of the video model branch for similar lip shapes. Using a self-supervised cross-audio-video modality contrastive learning method, the audio and video feature pairs from the same sample and the same time are constrained to be as similar as possible in the time dimension, and the video features at this time are made as different as possible from the audio or video features of other samples at the same time. Considering that different sentences at the same time are likely to have the same semantics, an optional way is to shuffle the feature sequence in the time dimension before performing contrastive learning constraints.
[0108] The present invention has the following advantages:
[0109] This method uses a self-supervised cross-modal contrastive learning method, without the need for additional manually labeled data. It uses audio information to help the lip reading recognition branch extract visual features with better distinguishability from the input video sequence, and based on this, distinguish lip reading video sequences with different pronunciations but similar lip shapes.
[0110] Compared with the two-stage traditional lip reading recognition method, this method builds an end-to-end lip reading recognition system based on deep learning. The feature extraction has better generalization and robustness, can be used across speakers, and there is no need to train a set of model parameters separately for each category sample.
[0111] Traditional methods have poor generalization ability for data from different speakers, while this method uses attribute learning to standardize the lip reading features from different speakers, greatly improving the robustness of the algorithm in dealing with the lip shapes of different speakers.
[0112] Although embodiments of the present invention have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A cross-modal lip-reading recognition method, characterized in that, it includes: S1, data preprocessing: For video data, first identify 68 key points on the face, and normalize each face image into a frontal view through affine transformation, and finally crop out the lip region; For audio data, first downsample it to 16 kHz and convert it into Mel cepstral coefficient features, then normalize the Mel cepstral coefficient vectors at all times and form a feature matrix in chronological order; S2, model training: S21, Input paired video data and audio data into the visual recognition branch and the speech recognition branch respectively, and perform training on the speaker recognition task in the attribute learning module of each branch; S22, Input paired video data and audio data into the visual recognition branch and the speech recognition branch respectively. In the contrast learning module shared by the two branches, use the representations obtained from the speaker recognition task to standardize the semantic features, and then perform cross-modal contrast learning between audio and video; S23, Only input the audio sequence, remove the speaker's timbre characteristics, standardize the speech features, and use the backpropagation algorithm to update the model parameters of the speech recognition branch to ensure that the intermediate audio feature S participating in the contrast learning is correct; S24, Only input the video sequence, remove the speaker's lip shape characteristics, standardize the lip-reading features, and use the backpropagation algorithm to update the model parameters of the lip-reading recognition branch; Repeat the above S21 - S24 until the value of the loss function no longer decreases within consecutive multiple rounds of training after the learning rate decays, that is, the model converges; S3, model deployment: Only input the non-training data video sequence to be recognized, use the visual recognition branch, remove the speaker's lip shape characteristics, standardize the lip-reading features, and finally perform the mapping from lip-reading features to text.
2. A cross-modal lip-reading recognition method according to claim 1, characterized in that, the visual recognition branch includes a 3D convolutional module, a first recurrent neural network module, a first speaker feature extraction module, a first attribute learning module, a contrast learning module, a second recurrent neural network module, a first attention module, and a first mapping module; the 3D convolutional module is used to obtain short-term features of lip movements; the first recurrent neural network module is used to establish long-term dependencies of lip movements; the first speaker feature extraction module is used to extract lip shape features of different speakers; the first attribute learning module is used to eliminate lip shape differences of different speakers by using the obtained speaker lip shape features; the contrast learning module is used to use the self-supervised contrast learning method across audio-visual data, so that the model obtains prior knowledge from another manifestation form of the video data, namely audio, and guides the visual recognition branch to learn lip shape features; the second recurrent neural network module is used to strengthen the context relationship of the video intermediate feature S sequence after passing through the contrast learning layer; the first attention module is used to help the model ignore irrelevant video frames by assigning different weights to the features at different time points output by the second recurrent neural network module in the time domain; the first mapping module is used to map the final lip movement features output by the first attention module into the text domain.
3. A cross-modal lip reading recognition method according to claim 2, It is characterized in that The speech recognition branch includes: 2D convolution module, third recursive neural network module, second speaker feature extraction module, second attribute learning module, contrastive learning module, fourth recursive neural network module, second attention module and second mapping module; The 2D convolution module is used to extract short-term speech features from the Mel-frequency cepstrum features; The third recursive neural network module is used to establish a long-term dependency relationship of speech features; The second speaker feature extraction module is used to extract timbre features of different speakers; The second attribute learning module is used to eliminate the timbre differences between different speakers by using the acquired speaker timbre features; The fourth recursive neural network module is used to strengthen the contextual relationship of the audio intermediate feature S sequence after the contrast learning module; The second attention module is used to help the model ignore irrelevant audio segments by assigning different weights to different time point features output by the fourth recurrent neural network module in the time domain; The second mapping module is used to map the final audio features output by the second attention module into the text domain.
4. The cross-modal lip reading recognition method according to claim 2, It is characterized in that The first mapping module includes a classifier based on nonlinear mapping of a multilayer perceptron with a ReLU activation function.
5. The cross-modal lip reading recognition method according to claim 3, It is characterized in that The connectionist temporal classification loss function is used to constrain the visual recognition branch and the speech recognition branch respectively.
Citation Information
Patent Citations
Mandarin lip language recognition method based on deep learning
CN109524006A
Lip language recognition method based on multi-granularity knowledge distillation
CN111223483A