Lip language recognition method, device and equipment, medium and product

By extracting intra and inter-frame features of video frames and fusion with text sequences, the shortcomings of existing lip recognition technology in terms of feature expression capabilities and discrimination are solved, and the accuracy of lip recognition is significantly improved.

CN120182771AActive Publication Date: 2025-06-20EAST CHINA JIAOTONG UNIVERSITY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510646116.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing lip recognition technology is poor in terms of feature expression ability and distinction, and has failed to effectively solve the problems of sequence-level lip recognition and homophones, resulting in low recognition accuracy.

Method used

By obtaining the intra-frame features and inter-frame features of each frame image, video features are obtained, and the multi-modal feature fusion method is used to fuse the text sequence with the video features to improve the model's feature expression ability and distinction.

Benefits of technology

It significantly improves the accuracy of lip recognition, enhances the model's feature expression ability and distinction, and can effectively deal with the problems of Chinese stress and insufficient feature points in the lip area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182771A_ABST
    Figure CN120182771A_ABST
Patent Text Reader

Abstract

The invention discloses a lip language recognition method and device, equipment, a medium and a product, and relates to the crossing field of machine vision and natural language processing, and the method comprises the steps: obtaining a to-be-recognized original video, initializing a text sequence, cutting the original video, obtaining a plurality of frame images, and extracting the intra-frame features and inter-frame features of each frame image. And obtaining video features according to the inter-frame features of all the frame images. And obtaining a lip language recognition result by adopting a lip language recognition model according to the video features and the text sequence. The lip language recognition model is a neural network multi-modal fusion model built in advance according to a training sample set. According to the method, the video features are obtained by obtaining the intra-frame features and the inter-frame features of each frame image, and multi-modal feature fusion is further performed on the text sequence and the video features, so that the feature expression capability and the discrimination degree of the model are improved, and the lip language recognition accuracy is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the cross - field of machine vision and natural language processing, and particularly to a lip - reading recognition method, device, equipment, medium and product. Background Art

[0002] Lip - reading recognition can translate a video - frame sequence of a speaker's lip region into text. Lip - reading recognition technology has broad application prospects. It can be applied to barrier - free communication systems for certain people who have lost their voices due to vocal cord or laryngeal injuries. It can also assist in collecting users' language information, such as interpreting the information of silent videos. In addition, lip - reading recognition can be combined with other biometric technologies to assist in enhancing the security of recognition.

[0003] During the lip - reading recognition process, there is no one - to - one correspondence between lip shapes and pronunciations, and between pronunciations and texts. There are often multiple possible alternative results, and it is necessary to calculate the most likely result in real time. Due to problems such as the stress of Chinese and the lack of feature points in the lip region, the feature representation ability and discrimination of existing lip - reading recognition systems are both poor. Although existing lip - reading recognition methods have achieved remarkable results in the recognition of some word - level datasets, there are still common defects: the sequence - level lip - reading recognition work is lacking, most lip - reading recognition models still stay at the word - level classification task, and do not consider the homophone problem from the perspective of sequences and contexts; the modeling of video and text sequences is insufficient, the attention to the frames within and between video sequences is not enough, and multi - modal data of text context and video frames is not combined for learning. Most of the work is two - stage and does not achieve end - to - end. This results in poor feature representation ability and discrimination of existing technologies, and low lip - reading recognition accuracy. Summary of the Invention

[0004] The purpose of the present application is to provide a lip - reading recognition method, device, equipment, medium and product, which can obtain video features by acquiring the intra - frame features and inter - frame features of each frame image, and further improve the feature representation ability and discrimination of the model and the accuracy of lip - reading recognition by performing multi - modal feature fusion on the text sequence and video features.

[0005] To achieve the above purpose, the present application provides the following solutions: In the first aspect, the present application provides a lip - reading recognition method, including: Obtain an original video to be recognized and initialize a text sequence; Cut the original video to obtain a plurality of frame images; each frame image includes the lip region of the speaker; Extract the intra - frame features of each frame image, and extract the inter - frame features of each frame image according to the intra - frame features of all frame images; Obtain video features according to the inter - frame features of all frame images; Based on the video features and the text sequence, a lip-reading recognition model is used to obtain a lip-reading recognition result; the lip-reading recognition model is a neural network multi-modal fusion model pre-built according to a training sample set; each training sample in the training sample set includes sample video features, sample text sequences, and target lip-reading recognition results.

[0006] Optionally, extract the intra-frame features of each frame image, and based on the intra-frame features of all frame images, extract the inter-frame features of each frame image, specifically including: For any frame image, cut the frame image to obtain a plurality of embedding vectors; Perform spatial position encoding on each embedding vector to obtain a plurality of spatial encoding vectors of the frame image; Based on all the spatial encoding vectors of the frame image, use a multi-head self-attention mechanism to extract the intra-frame features of the frame image; Perform temporal position encoding on the intra-frame features of the frame image to obtain the temporal encoding vector of the frame image; Based on the temporal encoding vectors of all frame images, use a multi-head self-attention mechanism to extract the inter-frame features of each frame image.

[0007] Optionally, the method for building the lip-reading recognition model includes: Obtain a plurality of sample videos and the corresponding target lip-reading recognition results for each sample video, and initialize the sample text sequence corresponding to each sample video; For any sample video, cut the sample video to obtain a plurality of sample frame images; each sample frame image includes the lip region of the speaker; Extract the intra-frame features of each sample frame image of the sample video, and based on the intra-frame features of all sample frame images of the sample video, extract the inter-frame features of each sample frame image of the sample video; Based on the inter-frame features of all sample frame images of the sample video, obtain the sample video features of the sample video; Based on the sample video features, sample text sequences, and target lip-reading recognition results corresponding to each sample video, obtain a training sample set; Use the training sample set to train a neural network multi-modal fusion model to obtain the lip-reading recognition model; the lip-reading recognition model is a trained neural network multi-modal fusion model.

[0008] Optionally, the text sequence includes a start character and an end character, and the neural network multi-modal fusion model includes a deep learning algorithm and a multi-modal feature fusion method; Based on the video features and the text sequence, a lip-reading recognition model is used to obtain a lip-reading recognition result, which specifically includes: For the i-th iteration, according to the first i labels of the text sequence, a deep learning algorithm is used to obtain the text features corresponding to the first i labels; i > 0; the first label of the text sequence is the start character; Based on the text features corresponding to the first i labels and the video features, a multi-modal feature fusion method is used to obtain the i-th predicted token, and the i-th predicted token is used as the (i + 1)-th label of the text sequence; if the i-th predicted token is the end character, the second to the i-th labels of the text sequence are the lip-reading recognition result, otherwise, the (i + 1)-th iteration is performed.

[0009] Optionally, for the i-th iteration, according to the first i labels of the text sequence, a deep learning algorithm is used to obtain the text features corresponding to the first i labels, which specifically includes: Respectively determine the word vectors corresponding to each label among the first i labels; Perform positional encoding on the word vectors corresponding to each label among the first i labels, and use a masked multi-head self-attention mechanism to obtain the text features corresponding to the first i labels.

[0010] Optionally, based on the text features corresponding to the first i labels and the video features, a multi-modal feature fusion method is used to obtain the i-th predicted token, and the i-th predicted token is used as the (i + 1)-th label of the text sequence, which specifically includes: Based on the text features corresponding to the first i labels and the video features, a cross-attention mechanism is used to obtain a fused feature; Based on the fused feature, a feed-forward neural network is used to obtain a strengthened feature; Based on the strengthened feature, the i-th predicted token is obtained, and the i-th predicted token is used as the (i + 1)-th label of the text sequence.

[0011] In a second aspect, the present application provides a lip-reading recognition device, which is applied to the above lip-reading recognition method, and includes: A data acquisition module, configured to acquire an original video to be recognized and initialize a text sequence; A video cutting module, configured to cut the original video to obtain a plurality of frame images; each frame image includes the lip region of the speaker; An image feature extraction module, configured to extract the intra-frame features of each frame image, and extract the inter-frame features of each frame image according to the intra-frame features of all the frame images; A video feature extraction module, configured to obtain video features according to the inter-frame features of all the frame images; A lip-reading recognition module, configured to obtain a lip-reading recognition result by using a lip-reading recognition model according to the video features and the text sequence; the lip-reading recognition model is a neural network multi-modal fusion model pre-built according to a training sample set; each training sample in the training sample set includes a sample video feature, a sample text sequence, and a target lip-reading recognition result.

[0012] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above lip-reading recognition method.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above lip-reading recognition method is implemented.

[0014] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above lip-reading recognition method is implemented.

[0015] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application provides a lip-reading recognition method, device, equipment, medium and product. By obtaining the intra-frame features and inter-frame features of each frame image to obtain video features, the feature representation ability and discrimination of the lip-reading recognition method are improved; a multi-modal feature fusion method is adopted to perform multi-modal feature fusion on the text sequence and the video features for lip-reading recognition, thereby improving the accuracy of lip-reading recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0017] Figure 1 It is an application environment diagram of a lip-reading recognition method in an embodiment of the present application.

[0018] Figure 2 It is a flowchart of a lip-reading recognition method provided in an embodiment of the present application.

[0019] Figure 3 It is an overall flowchart of implementing lip-reading recognition by means of a Transformer encoder provided in an embodiment of the present application.

[0020] Figure 4 ForFigure 2 Schematic diagram of the refined process of step 206 in

[0021] Figure 5 Schematic diagram of the functional modules of a lip-reading recognition device provided by an embodiment of the present application.

[0022] Figure 6 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] The present application proposes a lip-reading recognition method, device, equipment, medium and product. By obtaining the intra-frame features and inter-frame features of each frame image, video features are obtained, which improves the feature representation ability and discrimination of the lip-reading recognition method; a multi-modal feature fusion method is adopted to perform multi-modal feature fusion on the text sequence and video features for lip-reading recognition, which improves the accuracy of lip-reading recognition.

[0025] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0026] The lip-reading recognition method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the original video to be recognized to the server 104. After receiving the original video, the server 104 cuts the original video to obtain multiple frame images. For any frame image, its intra-frame features and inter-frame features are extracted. According to the inter-frame features of all frame images, video features are obtained. Finally, according to the text sequence and video features, a multi-modal feature fusion method is used to obtain the lip-reading recognition result. The server 104 can feedback the obtained lip-reading recognition result to the terminal 102. In addition, in some embodiments, the lip-reading recognition method can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly perform lip-reading recognition on the original video, or the server 104 can obtain the original video from the data storage system and perform lip-reading recognition on the original video.

[0027] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0028] In an exemplary embodiment, as Figure 2 shown, a lip-reading recognition method is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 therein as an example for illustration, it includes the following steps 201 to step 205.

[0029] Among them, in step 201, obtain the original video to be recognized and initialize the text sequence.

[0030] In step 202, cut the original video to obtain a plurality of frame images. Each frame image includes the lip region of the speaker.

[0031] In step 203, extract the intra-frame features of each frame image, and extract the inter-frame features of each frame image according to the intra-frame features of all the frame images.

[0032] In step 204, obtain the video features according to the inter-frame features of all the frame images.

[0033] In step 205, according to the video features and the text sequence, use a lip-reading recognition model to obtain the lip-reading recognition result. Among them, the lip-reading recognition model is a neural network multi-modal fusion model pre-built according to a training sample set. Each training sample in the training sample set includes sample video features, sample text sequences, and target lip-reading recognition results.

[0034] In an exemplary embodiment, the method for building a lip-reading recognition model includes the following steps.

[0035] (1) Obtain a plurality of sample videos and the corresponding target lip-reading recognition results for each sample video, and initialize the sample text sequence corresponding to each sample video.

[0036] (2) For any sample video, cut the sample video to obtain a plurality of sample frame images. Each sample frame image includes the lip region of the speaker.

[0037] Extract the intra-frame features of each sample frame image of the sample video, and extract the inter-frame features of each sample frame image of the sample video according to the intra-frame features of all sample frame images of the sample video.

[0038] (4)Obtain the sample video features of the sample video according to the inter-frame features of all sample frame images of the sample video.

[0039] (5)Obtain the training sample set according to the sample video features, sample text sequence, and target lip-reading recognition results corresponding to each sample video.

[0040] (6)Use the training sample set to train the neural network multi-modal fusion model to obtain the lip-reading recognition model. The lip-reading recognition model is the trained neural network multi-modal fusion model.

[0041] In this embodiment, the training effect of the neural network multi-modal fusion model is evaluated and optimized by calculating the cross-entropy loss and edit distance loss between the predicted sample text sequence and the target lip-reading recognition result.

[0042] In an exemplary embodiment, before Figure 2 step 203, crop each frame image according to the face key points to obtain the lip region image to reduce the interference of invalid pixels. Perform GABF (Gaussian-Adaptive Bilateral Filter) filtering and size normalization on the lip region image for subsequent image cutting. In this embodiment, the sizes of all lip region images are normalized to 64×64 pixels.

[0043] In an exemplary embodiment, Figure 2 step 203 in

[0044] Step 2031, for any frame image, cut the frame image to obtain multiple embedding vectors. In this embodiment, use the Frame Embedding module to cut the original video to be recognized into multiple frame images, extract the lip region from the frame images, and perform filtering and size normalization on the lip region image. Finally, evenly cut the processed 64×64 pixel lip region image into four parts to obtain four embedding vectors. Flatten the four embedding vectors into a one-dimensional sequence and map them to the low-dimensional feature space through a linear layer.

[0045] Step 2032, perform spatial position encoding on each embedding vector to obtain multiple spatial encoding vectors of the frame image.

[0046] Step 2033: According to all the spatial encoding vectors of the frame image, use the multi-head self-attention mechanism to extract the intra-frame features of the frame image.

[0047] Step 2034: Perform temporal position encoding on the intra-frame features of the frame image to obtain the temporal encoding vector of the frame image.

[0048] Step 2035: According to the temporal encoding vectors of all the frame images, use the multi-head self-attention mechanism to extract the inter-frame features of each frame image.

[0049] In this embodiment, a Transformer encoder is used to extract video features and text features, and feature fusion is performed to achieve lip reading recognition. The process of using the Transformer encoder to implement the above operations is as Figure 3 shown. When extracting video features, all the embedding vectors of all the frame images are input into the Transformer encoder. Taking the frame image as a unit, a learnable identifier "cls" is added to the zero-th position of each frame image. Spatial position encoding is performed on each embedding vector of each frame image. By adding each embedding vector of each frame image to the learnable spatial position information, a sequence with spatial position information is formed, thereby helping the Transformer encoder understand the spatial relationship and sequence order inside each frame image. The multi-head self-attention mechanism is used for spatial attention calculation to calculate the relationship between each position inside each frame image, so as to obtain the spatial feature representation of each frame image and extract the intra-frame features. The intra-frame features are lip shape features.

[0050] In this embodiment, as Figure 3 shown, based on the calculation result of spatial attention, temporal attention calculation is performed. A learnable identifier "cls" is added to the intra-frame features of each frame image. Temporal position encoding is performed on the intra-frame features of each frame image. The temporal attention module is used to calculate the temporal relationship between different frame images, so as to obtain the inter-frame features of the original video. The inter-frame features are the temporal order of lip movement.

[0051] Based on the calculation result of temporal attention, that is, the inter-frame features, a feed-forward neural network is used for feature fusion to obtain video features. The feed-forward neural network fuses the intra-frame features and inter-frame features after each layer, optimizes the feature representation through layer normalization and activation functions, and finally generates video features for lip reading recognition.

[0052] Text feature extraction is parallel to video feature extraction. Based on the trained neural network multi-modal fusion model, the lip reading recognition result can be obtained. In an exemplary embodiment, the text sequence includes a start character and an end character. The neural network multi-modal fusion model includes a deep learning algorithm and a multi-modal feature fusion method. As Figure 4 shown,Figure 2 Step 205 in can be replaced by the following Step 401 and Step 402.

[0053] Step 401, for the i-th iteration, according to the first i labels of the text sequence, using a deep learning algorithm, obtain the text features corresponding to the first i labels. Wherein, i > 0, and the first label of the text sequence is the start character.

[0054] Step 402, according to the text features corresponding to the first i labels and the video features, using a multi-modal feature fusion method, obtain the i-th predicted token, and use the i-th predicted token as the (i + 1)-th label of the text sequence. If the i-th predicted token is the end character, then the second to the i-th labels of the text sequence are the lip-reading recognition results; otherwise, perform the (i + 1)-th iteration.

[0055] In an exemplary embodiment, Figure 4 Step 401 in can be replaced by the following Step 4011 and Step 4012.

[0056] Step 4011, respectively determine the word vectors corresponding to each label among the first i labels.

[0057] Step 4012, perform positional encoding on the word vectors corresponding to each label among the first i labels, and use a masked multi-head self-attention mechanism to obtain the text features corresponding to the first i labels.

[0058] In an exemplary embodiment, Figure 4 Step 402 in can be replaced by the following Step 4021 to Step 4023.

[0059] Step 4021, according to the text features corresponding to the first i labels and the video features, use a cross-attention mechanism to obtain the fused features.

[0060] Step 4022, according to the fused features, use a feed-forward neural network to obtain the enhanced features.

[0061] Step 4023, according to the enhanced features, obtain the i-th predicted token, and use the i-th predicted token as the (i + 1)-th label of the text sequence.

[0062] In this embodiment, as Figure 3As shown, word vector embedding is performed based on the text sequence. Xavier is used for weight initialization, and the embedding weights are automatically updated through gradients and backpropagation. Position encoding is performed on the word vectors corresponding to each label in the text sequence. The word vectors are added to the learnable position encoding, enabling the Transformer decoder to consider the order information of the words in the text data. For the word vectors after position encoding, a masked multi-head attention mechanism is used for training to obtain text features. Specifically, the self-attention mechanism of the Transformer decoder can capture context information by calculating the relationships between each position in the text sequence. The text features are the sequential and semantic relationships of the text.

[0063] Based on the text features and video features obtained by parallel processing of text and video respectively, cross-attention calculation between text and video is performed to obtain fused features. The cross-attention calculation process simultaneously focuses on the spatio-temporal features of the video and the context relationship of the text, performs feature fusion between the video modality and the text modality, and enhances the multi-modal understanding ability of the lip-reading recognition method. The fused features are input into a feed-forward neural network for feature enhancement to obtain enhanced features. The enhanced features are decoded in the Transformer decoder to obtain the final embedding vector. Further processing is performed on the embedding vector to convert it into a predicted token, thereby obtaining the final result of lip-reading recognition.

[0064] In an exemplary embodiment, based on the trained neural network multi-modal fusion model, using the autoregressive mechanism, predicted tokens are sequentially obtained, and the obtained tokens are sequentially filled into the text sequence after initialization. As Figure 3 shown, the text sequence after initialization includes a start character and an end character. The start character is located at the first label position of the text sequence, and the end character is located at the last label position of the text sequence.

[0065] The lip-reading recognition method of the present application obtains video features at once based on the original video to be recognized. For the text part, based on the trained neural network multi-modal fusion model, the text features corresponding to each predicted token are sequentially obtained through iteration. In the first iteration, the text features corresponding to the first label are extracted, that is, the text features corresponding to the start character. According to the video features and the text features corresponding to the start character, a lip-reading recognition model is used to predict the first predicted token, which is used as the second label of the text sequence. In the second iteration, according to the first two labels of the text sequence, that is, the start character and the first predicted token, the text features corresponding to them are obtained. Further combined with the video features, a lip-reading recognition model is used to predict the second predicted token, which is used as the third label of the text sequence, and so on. When the iteration reaches the last label of the text sequence, that is, the end character, it is detected that the predicted token is the end character, the loop ends, and the label sequence obtained by removing the start character and the end character from the text sequence is used as the final lip-reading recognition result.

[0066] Based on the same inventive concept, an embodiment of the present application further provides a lip-reading recognition device for implementing the above-mentioned lip-reading recognition method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the lip-reading recognition device provided below can refer to the limitations on the lip-reading recognition method in the above text, and will not be repeated here.

[0067] In an exemplary embodiment, as Figure 5 shown, a lip-reading recognition device is provided, including: a data acquisition module 501, a video cutting module 502, an image feature extraction module 503, a video feature extraction module 504, and a lip-reading recognition module 505. Among them, the data acquisition module 501 is used to acquire the original video to be recognized and initialize the text sequence. The video cutting module 502 is used to cut the original video to obtain a plurality of frame images. Each frame image includes the lip region of the speaker. The image feature extraction module 503 is used to extract the intra-frame features of each frame image, and according to the intra-frame features of all frame images, extract the inter-frame features of each frame image. The video feature extraction module 504 is used to obtain video features according to the inter-frame features of all frame images. The lip-reading recognition module 505 is used to obtain the lip-reading recognition result according to the video features and the text sequence by using a lip-reading recognition model. Among them, the lip-reading recognition model is a neural network multi-modal fusion model pre-built according to a training sample set. Each training sample in the training sample set includes sample video features, sample text sequences, and target lip-reading recognition results.

[0068] In an exemplary embodiment, a computer device is provided. This computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the original video. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a lip-reading recognition method.

[0069] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0070] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0071] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0072] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0073] The beneficial effects of the present application are as follows: By parallel feature extraction of the original video and the text sequence, the present application successfully captures the spatio-temporal relationship inside and outside the video frame and the mutual connection of the text context, realizes the effective fusion of multi-modal features, and can still maintain a high recognition ability even in the absence of audio modal input, effectively reducing the character error rate.

[0074] The lip-reading recognition method provided by this application addresses the problems of insufficient Chinese stress and lack of feature points in the lip region in existing lip-reading recognition systems, proposes a practical and effective solution, and significantly improves the feature representation ability and discrimination of the model. By using facial key points to crop the lip region, the interference of invalid pixels is reduced, and spatio-temporal attention mechanisms are used to extract inter-frame and intra-frame features. Cross-attention is calculated by combining video features and text features, demonstrating excellent performance exceeding that of traditional convolutions.

[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0076] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0077] In each of the embodiments provided in the present application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on a blockchain, etc., and is not limited thereto. In each of the embodiments provided in the present application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.

[0078] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0079] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A lip reading recognition method, characterized in that: The lip reading recognition method comprises: Get the original video to be recognized and initialize the text sequence; Cutting the original video to obtain a plurality of frame images; each frame image includes a lip area of ​​the speaker; Extracting intra-frame features of each frame image, and extracting inter-frame features of each frame image based on the intra-frame features of all frame images; According to the inter-frame features of all frame images, video features are obtained; According to the video features and the text sequence, a lip reading recognition model is used to obtain a lip reading recognition result; the lip reading recognition model is a neural network multimodal fusion model pre-built according to a training sample set; each training sample in the training sample set includes a sample video feature, a sample text sequence and a target lip reading recognition result.

2. The lip reading recognition method according to claim 1, characterized in that: Extracting intra-frame features of each frame image, and extracting inter-frame features of each frame image based on the intra-frame features of all frame images, specifically including: For any frame image, the frame image is segmented to obtain multiple embedding vectors; Performing spatial position encoding on each embedded vector to obtain multiple spatial encoding vectors of the frame image; According to all spatial encoding vectors of the frame image, a multi-head self-attention mechanism is used to extract intra-frame features of the frame image; Performing temporal position coding on intra-frame features of the frame image to obtain a temporal coding vector of the frame image; According to the time coding vectors of all frame images, a multi-head self-attention mechanism is used to extract the inter-frame features of each frame image.

3. The lip reading recognition method according to claim 1, characterized in that: The method for building the lip reading recognition model includes: Obtain multiple sample videos and target lip reading recognition results corresponding to each sample video, and initialize a sample text sequence corresponding to each sample video; For any sample video, the sample video is cut to obtain a plurality of sample frame images; each sample frame image includes a lip area of ​​a speaker; Extracting intra-frame features of each sample frame image of the sample video, and extracting inter-frame features of each sample frame image of the sample video based on the intra-frame features of all sample frame images of the sample video; Obtaining sample video features of the sample video according to inter-frame features of all sample frame images of the sample video; A training sample set is obtained according to the sample video features, sample text sequence and target lip reading recognition results corresponding to each sample video; The training sample set is used to train the neural network multimodal fusion model to obtain the lip reading recognition model; the lip reading recognition model is the trained neural network multimodal fusion model.

4. The lip reading recognition method according to claim 1, characterized in that: The text sequence includes a start character and an end character, and the neural network multimodal fusion model includes a deep learning algorithm and a multimodal feature fusion method; According to the video features and the text sequence, a lip reading recognition model is used to obtain a lip reading recognition result, which specifically includes: For the i-th iteration, based on the first i labels of the text sequence, a deep learning algorithm is used to obtain text features corresponding to the first i labels; i>0; the first label of the text sequence is the start character; According to the text features corresponding to the first i labels and the video features, a multimodal feature fusion method is adopted to obtain the i-th predicted word element, and the i-th predicted word element is used as the i+1-th label of the text sequence; if the i-th predicted word element is the termination character, then the 2nd to i-th labels of the text sequence are lip reading recognition results, otherwise, the i+1-th iteration is performed.

5. The lip reading recognition method according to claim 4, characterized in that: For the i-th iteration, based on the first i labels of the text sequence, a deep learning algorithm is used to obtain text features corresponding to the first i labels, specifically including: Determine the word vector corresponding to each label in the first i labels respectively; The word vector corresponding to each of the first i tags is positionally encoded, and a masked multi-head self-attention mechanism is used to obtain the text features corresponding to the first i tags.

6. The lip reading recognition method according to claim 4, characterized in that: According to the text features corresponding to the first i labels and the video features, a multimodal feature fusion method is used to obtain an i-th predicted word, and the i-th predicted word is used as the i+1-th label of the text sequence, specifically including: According to the text features corresponding to the first i tags and the video features, a cross attention mechanism is used to obtain fusion features; According to the fusion features, a feed-forward neural network is used to obtain enhanced features; According to the enhanced features, an i-th predicted word is obtained, and the i-th predicted word is used as the i+1-th label of the text sequence.

7. A lip reading recognition device, applied to the lip reading recognition method according to any one of claims 1 to 6, characterized in that: The lip reading recognition device comprises: The data acquisition module is used to obtain the original video to be recognized and initialize the text sequence; A video cutting module, used for cutting the original video to obtain a plurality of frame images; each frame image contains the lip area of ​​the speaker; An image feature extraction module is used to extract intra-frame features of each frame image, and extract inter-frame features of each frame image based on the intra-frame features of all frame images; A video feature extraction module is used to obtain video features based on inter-frame features of all frame images; The lip reading recognition module is used to obtain the lip reading recognition result by adopting the lip reading recognition model according to the video features and the text sequence; the lip reading recognition model is a neural network multimodal fusion model pre-built according to the training sample set; each training sample in the training sample set includes a sample video feature, a sample text sequence and a target lip reading recognition result.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the lip reading recognition method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the lip reading recognition method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the lip reading recognition method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Lip language recognition method and device, computer equipment and storage medium

    CN110276259A

  • Lip language recognition method based on multi-granularity knowledge distillation

    CN111223483A

  • Visual language recognition method and device, electronic equipment and storage medium

    CN114581812A

  • Lip language recognition method based on deep learning

    CN116778577A

  • Lip language recognition method and device based on space-time network model, medium and product

    CN119851345A