Lip reading method, apparatus, device, medium and computer program product

CN122799472APending Publication Date: 2026-09-22CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332526.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]本发明提供一种唇语识别方法、装置、设备、介质和计算机程序产品,用以解决现有技术中唇部识别容易受光照条件、图像噪声的影响导致唇部图像分割不准确的缺陷,实现不受光照、图像噪声影响而捕捉唇部运动序列中的时间依赖性,从而提高唇部识别的准确性

Benefits of technology

[0016]本发明提供的唇语识别方法、装置、设备、介质和计算机程序产品,通过从视频片段中提取多个图像帧;将多个图像帧输入至三维唇部建模模型,得到三维唇部建模模型输出的唇部的三维特征图像;将唇部的三维特征图像输入至特征提取模型,得到特征提取模型输出的空间特征向量;将空间特征向量输入至门控循环单元,得到门控循环单元提取的唇部运动时间特征;将唇部运动的时间特征输入至分类层,得到分类层输出的唇语识别结果。本申请首先通过3D唇部建模可以更准确地捕捉唇部的三维结构和运动,使得重建的3D唇部模型在唇语识别过程中更接近真实。相比于传统的2D方法,这种方式能够更好地应对唇语识别任务中不同设备视角和光照条件的变化;进一步地,本申请还使用了GRU,GRU通过更新门和重置门,保留重要的序列信息,过滤掉不必要的噪音和冗余信息,从而提高识别准确性。本申请所提出的唇部识别方法容易受光照条件或图像噪声的影响,提升唇部分割效果,进而提高唇语识别效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799472A_ABST
    Figure CN122799472A_ABST
Patent Text Reader

Abstract

The application provides a lip-reading method, device, equipment, medium and computer program product, the method comprising: extracting a plurality of image frames from a video segment; inputting the plurality of image frames into a three-dimensional lip modeling model to obtain a three-dimensional feature image output by the three-dimensional lip modeling model; inputting the three-dimensional feature image of the lip into a feature extraction model to obtain a spatial feature vector output by the feature extraction model; inputting the spatial feature vector into a gated recurrent unit to obtain a lip motion time feature extracted by the gated recurrent unit; and inputting the lip motion time feature into a classification layer to obtain a lip-reading result output by the classification layer. The application can more accurately capture the three-dimensional structure and motion of the lip by 3D lip modeling, and compared with the traditional 2D method, this way can better cope with the changes of different device angles and light conditions in the lip-reading task, and improve the lip-reading effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a lip-reading method, apparatus, device, medium, and computer program product. Background Technology

[0002] When a smart terminal initiates a voice conversation, it generally requires a wake-up operation, such as button wake-up, touch wake-up, or voice wake-up. However, these wake-up operations are very rigid, requiring a wake-up every time a conversation begins. This raises the barrier to entry for users and leads to interactive blockage issues such as unsmooth conversations.

[0003] Currently, lip recognition is used for device wake-up, meaning the device enters a wake-up state when it detects a person's lips in an image. Existing lip recognition methods mainly include color thresholding and Active Contour Model (ACM). Color thresholding calculates the color difference between the lip region and its surrounding area, setting an appropriate threshold to segment the lip region. However, color thresholding is extremely sensitive to changes in lighting and performs poorly when the lip color is similar to the surrounding skin tone. ACM is based on curve evolution theory, using grayscale information and shape prior knowledge to extract the lip contour. By adaptively adjusting the contour shape, ACM can improve the segmentation accuracy of the lip region. However, this method is susceptible to image noise and has high computational complexity, limiting its applicability.

[0004] In summary, existing lip recognition methods are easily affected by lighting conditions or image noise, resulting in poor segmentation performance. Summary of the Invention

[0005] This invention provides a lip reading recognition method, apparatus, device, medium, and computer program product to address the shortcomings of existing technologies where lip recognition is easily affected by lighting conditions and image noise, leading to inaccurate lip image segmentation. It achieves the capture of the time dependence in the lip movement sequence without being affected by lighting or image noise, thereby improving the accuracy of lip recognition.

[0006] This invention provides a lip-reading recognition method, comprising: Extract multiple image frames from a video clip; The multiple image frames are input into the 3D lip modeling model to obtain the 3D feature image output by the 3D lip modeling model; The three-dimensional feature image of the lips is input into the feature extraction model to obtain the spatial feature vector output by the feature extraction model. The spatial feature vector is input into the gated loop unit to obtain the lip movement time features extracted by the gated loop unit; The temporal features of the lip movements are input into the classification layer to obtain the lip reading recognition results output by the classification layer.

[0007] According to a lip-reading recognition method provided by the present invention, the three-dimensional lip modeling model includes a fixed encoder, a perceptual encoder, a three-dimensional image contour model, and a differentiable renderer. The step of inputting the plurality of image frames into the three-dimensional lip modeling model to obtain a three-dimensional feature image output by the three-dimensional lip modeling model includes: The three-dimensional feature image is input into the perceptual encoder to obtain the chin parameters obtained by the perceptual encoder using spatial perception loss. The three-dimensional feature image is input into the fixed encoder to obtain the identity parameters and albedo parameters output by the fixed encoder. The chin parameter, the identity parameter, and the albedo parameter are input into the three-dimensional image contour model to obtain the lip movement time series output by the three-dimensional image contour model; The lip movement time series is input into the differentiable renderer to obtain the three-dimensional feature image output by the differentiable renderer.

[0008] According to a lip-reading recognition method provided by the present invention, the feature extraction model includes a two-dimensional convolutional module, multiple Ghost bottleneck modules, an average pooling module, and a fully connected layer connected in sequence; the step of inputting the three-dimensional feature image of the lips into the feature extraction model to obtain the spatial feature vector output by the feature extraction model includes: The three-dimensional feature image is input into the two-dimensional convolution module to obtain the basic feature map; The basic feature map is input to the plurality of Ghost bottleneck modules to obtain Ghost feature maps output by the plurality of Ghost bottleneck modules; wherein, the plurality of Ghost bottleneck modules are connected in sequence, and the output of the previous Ghost bottleneck module is used as the input of the next Ghost bottleneck module; The Ghost feature map is input into the average pooling module to obtain the spatial feature map output by the average pooling module; The spatial feature map is input into the fully connected layer to obtain the spatial feature vector output by the fully connected layer.

[0009] According to a lip-reading recognition method provided by the present invention, the gated recurrent unit includes an update gate, a reset gate, and a forget gate; the step of inputting the spatial feature vector into the gated recurrent unit to obtain the lip movement temporal features extracted by the gated recurrent unit includes: The spatial feature vector is input into the gated loop unit so that the update gate controls the information of the previous state and the reset gate controls the previous state, thereby obtaining the lip movement time features.

[0010] The present invention also provides a device wake-up method, comprising: Detect whether a video frame contains a complete lip image; If no complete lip image is detected in any frame of the video frame, then the steps in the above lip reading recognition method embodiment are applied to obtain the Chinese lip reading prediction result; The lip reading results are displayed.

[0011] According to a device wake-up method provided by the present invention, after detecting whether the video frame contains a complete lip image, the method further includes: If the video frame contains a complete lip image, the complete lip image is compared with an image in a preset image feature library to obtain the lip reading recognition result.

[0012] The present invention also provides a lip-reading recognition device, comprising the following modules: The image frame extraction module is used to extract multiple image frames from a video clip; A three-dimensional feature extraction module is used to input the multiple image frames into a three-dimensional lip modeling model to obtain a three-dimensional feature image output by the three-dimensional lip modeling model; The spatial feature vector output module is used to input the three-dimensional feature image of the lips into the feature extraction model to obtain the spatial feature vector output by the feature extraction model. The lip movement time feature extraction module is used to input the spatial feature vector into the gated loop unit to obtain the lip movement time features extracted by the gated loop unit; The lip reading result output module is used to input the temporal features of the lip movements into the classification layer to obtain the lip reading result output by the classification layer.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps as described in any embodiment of the lip reading method, or to implement the steps as described in any embodiment of the device wake-up method.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps as in any embodiment of a lip-reading method, or implements the steps as in any embodiment of a device wake-up method.

[0015] The present invention also provides a computer program product, comprising a computer program that, when executed by a processor, implements the steps as described in any of the above embodiments of the lip reading method, or implements the steps as described in any of the above embodiments of the device wake-up method.

[0016] The lip-reading recognition method, apparatus, device, medium, and computer program product provided by this invention extracts multiple image frames from a video clip; inputs these multiple image frames into a 3D lip modeling model to obtain a 3D feature image of the lips output by the 3D lip modeling model; inputs the 3D feature image of the lips into a feature extraction model to obtain a spatial feature vector output by the feature extraction model; inputs the spatial feature vector into a gated loop unit to obtain the temporal features of lip movement extracted by the gated loop unit; and inputs the temporal features of lip movement into a classification layer to obtain the lip-reading recognition result output by the classification layer. This application firstly uses 3D lip modeling to more accurately capture the 3D structure and movement of the lips, making the reconstructed 3D lip model closer to reality during lip-reading recognition. Compared to traditional 2D methods, this approach can better cope with changes in different device perspectives and lighting conditions in lip-reading recognition tasks. Furthermore, this application also uses a GRU (Generative Recognition Root), which retains important sequence information and filters out unnecessary noise and redundant information through update and reset gates, thereby improving recognition accuracy. The lip recognition method proposed in this application is easily affected by lighting conditions or image noise, which improves the lip segmentation effect and thus enhances the lip reading recognition effect. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the lip reading recognition method provided by the present invention.

[0019] Figure 2 This is a schematic diagram of the overall structure of the lip reading recognition model provided by the present invention.

[0020] Figure 3 This is a schematic diagram of the structure of the three-dimensional lip modeling model provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the processing flow of the three-dimensional lip modeling model provided by the present invention.

[0022] Figure 5 This is a schematic diagram of the feature extraction model provided by the present invention.

[0023] Figure 6 This is a schematic diagram of the interaction process of the device wake-up method provided by the present invention.

[0024] Figure 7 This is a schematic diagram of the lip reading recognition device provided by the present invention.

[0025] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] The following is combined with Figures 1-8 Specific embodiments of the present invention are described below.

[0028] Figure 1 This is a flowchart illustrating the lip-reading recognition method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following.

[0029] Step 101: Extract multiple image frames from the video clip.

[0030] The video clip refers to a video segment captured by a camera, which may contain facial images. This camera can be installed in smart home devices, such as televisions, smartphones, and voice assistants.

[0031] Specifically, when a user enters the field of view of a smart home device's camera, the camera begins to capture facial video clips. After capturing the video clips, they are sent to the processor, which then extracts a fixed number of image frames from them.

[0032] Step 102: Input the multiple image frames into the three-dimensional lip modeling model to obtain the three-dimensional feature image of the lips output by the three-dimensional lip modeling model.

[0033] like Figure 2 As shown, Figure 2 This paper presents a schematic diagram of the overall structure of the lip-reading recognition model used in this application. The lip-reading recognition model includes a three-dimensional lip modeling model and an improved GhostNet network (i.e.,...). Figure 2The network consists of Efficient-GhostNet, GRU (Gated Recurrent Unit), and Softmax layer.

[0034] Specifically, the aforementioned fixed number of image frames are input into the 3D lip modeling model to perform lip modeling, thereby enhancing the three-dimensional feature information of the lips, providing rich dimensional feature descriptions, and outputting a three-dimensional feature image of the lips.

[0035] Step 103: Input the three-dimensional feature image of the lips into the feature extraction model to obtain the spatial feature vector output by the feature extraction model.

[0036] Among them, the feature extraction model is Figure 2 The Efficient-GhostNet network is a lightweight convolutional neural network for feature extraction. It efficiently generates feature maps by introducing a Ghost Module, thus maintaining high feature extraction capabilities while reducing computational cost and parameter count. This approach is particularly suitable for applications on resource-constrained devices such as mobile or embedded devices.

[0037] Specifically, the three-dimensional feature image of the lips is input into the feature extraction model (i.e., the Efficient-GhostNet network) to extract efficient depth feature vectors, and the spatial feature vector of the lips is obtained from the output of the Efficient-GhostNet network.

[0038] Step 104: Input the spatial feature vector into the gated loop unit (GRU) to obtain the lip movement time features extracted by the gated loop unit.

[0039] GRU, short for Gated Recurrent Unit, is a variant of recurrent neural network (RNN) used to process sequential data. GRU uses a gating mechanism to control the flow of information, thus better capturing long-term dependencies in a sequence. It mainly includes two key gates: the reset gate and the update gate. GRU is suitable for handling complex dependencies in sequential data.

[0040] Specifically, the spatial feature vector of the lips is input into the gated recurrent unit, and the temporal features of the lip movement are extracted using the GRU network to capture its dynamic changes, thus obtaining the temporal features of the lip movement extracted by the gated recurrent unit.

[0041] Step 105: Input the temporal features of the lip movements into the classification layer to obtain the lip reading recognition result output by the classification layer.

[0042] The classification layer is the Softmax layer.

[0043] Specifically, the aforementioned lip movement time features are input into the SoftMax classification layer to generate the final prediction result of lip reading, which may be a Chinese character recognition result.

[0044] The above embodiments extract multiple image frames from a video clip; input these multiple image frames into a 3D lip modeling model to obtain a 3D feature image of the lips output by the 3D lip modeling model; input the 3D feature image of the lips into a feature extraction model to obtain a spatial feature vector output by the feature extraction model; input the spatial feature vector into a gated recurrent unit to obtain the temporal features of lip movement extracted by the gated recurrent unit; input the temporal features of lip movement into a classification layer to obtain the lip-reading recognition result output by the classification layer. This application firstly uses 3D lip modeling to more accurately capture the 3D structure and movement of the lips, making the reconstructed 3D lip model closer to reality during lip-reading recognition. Compared to traditional 2D methods, this approach can better cope with changes in different device perspectives and lighting conditions in lip-reading recognition tasks; furthermore, this application also uses a GRU, which retains important sequence information and filters out unnecessary noise and redundant information through update and reset gates, thereby improving recognition accuracy. The lip recognition method proposed in this application is easily affected by lighting conditions or image noise, improving lip segmentation and thus improving lip-reading recognition performance.

[0045] In one embodiment, lip reading using only 2D image frames performs poorly in handling complex lip shape changes, lighting conditions, and angle variations. Therefore, this proposal employs 3D lip modeling to enhance lip features in the extracted K-frame RGB image sequence (i.e., multiple image frames extracted from a video clip) to address complex real-world application scenarios. Figure 2 The 3D lip modeling model in the image includes a fixed encoder, a perceptual encoder, a 3D image contour model, and a differentiable renderer, such as... Figure 3 As shown, Figure 3 A schematic diagram of the structure of a 3D lip model is shown.

[0046] Figure 4 The diagram illustrates the processing flow of the 3D lip model, including step 102 mentioned above.

[0047] Step 401: Input the three-dimensional feature image of the lips into the perceptual encoder to obtain the chin parameters obtained by the perceptual encoder using spatial perception loss.

[0048] A perceptual encoder is an encoder based on human perception characteristics, typically used in image, audio, or video processing tasks. Its core idea is to simulate the characteristics of human perceptual systems (such as the visual or auditory systems), focusing on retaining the information most important to human perception while ignoring or compressing information with less impact on human perception. This approach can significantly reduce data volume or computational complexity while maintaining high-quality output.

[0049] Specifically, the three-dimensional feature image of the lips is input into a perceptual encoder, which is driven by a spatial perception loss function and can accurately estimate parameters closely related to lip movement, ultimately outputting chin parameters. The spatially aware loss function optimizes the model by comparing the differences between the encoder output and the original data in the feature space, rather than pixel-level differences.

[0050] Step 402: Input the three-dimensional feature image of the lips into the fixed encoder to obtain the identity parameters and albedo parameters output by the fixed encoder.

[0051] A fixed encoder is an encoder whose parameters remain unchanged and are not updated during model training. Fixed encoders typically use pre-trained models or manually designed feature extractors to extract features from the input data and pass these features to subsequent modules for processing. The fixed encoder in this application adopts the fixed encoder from the DECA (Detailed Expression Capture and Animation) framework.

[0052] Specifically, the three-dimensional feature image of the lips is input into a fixed encoder, which extracts the rigid transformation parameters and albedo parameters from the image. The rigid transformation parameters, specifically those for the human head, are mathematical parameters describing a rigid transformation. A rigid transformation is a geometric transformation that preserves the size and shape of an object while changing its position and orientation. Rigid transformations include translation and rotation, and are widely used in two-dimensional or three-dimensional space. Rigid transformation parameters are typically represented in matrix form. The albedo parameter represents the ratio of incident radiation (such as sunlight) reflected from the object's surface to the total incident radiation; it is a dimensionless value ranging from 0 to 1.

[0053] Optionally, the fixed encoder is also used to extract identity parameters, which are used to characterize the identity of the person in the image.

[0054] Step 403: Input the chin parameters, the rigidity transformation parameters, and the albedo parameters into the three-dimensional image contour model to obtain the lip movement time series output by the three-dimensional image contour model.

[0055] The 3D image contour model is built on the MobileNet architecture and is used to output a rough 3D image contour. The output of the MobileNet architecture in this application incorporates a temporal convolution kernel to capture the temporal changes in mouth movement, thereby achieving a more realistic 3D reconstruction.

[0056] Specifically, the above chin parameters The rigidity transformation parameters and albedo parameters are input into the three-dimensional image contour model to obtain the lip motion time series output by the three-dimensional image contour model.

[0057] Optionally, since there are differences between the subsequently rendered image and the original image, even though the spatial perception loss function can preserve high-level information, it relies on pre-trained task-specific CNNs. These networks cannot guarantee that the generated images are completely realistic. Therefore, it is necessary to improve the training process through geometric constraint functions in the 3D image contour model. These geometric constraint functions include chin parameters. of Norm and Relative loss function.

[0058] Among them, the penalty chin parameter is used. of norm This is used to control the deviation range between the output chin image and the actual chin image; additionally, the internal distance between mouth landmarks is... The relative loss function focuses on the relative distances between points inside the mouth, ensuring that the shape and structure of the mouth remain consistent during reconstruction without being excessively affected by deviations in the position of individual points.

[0059] Step 404: Input the lip movement time sequence into the differentiable renderer to obtain the three-dimensional feature image output by the differentiable renderer.

[0060] Among them, a differentiable renderer is a renderer that can calculate the gradient of the rendering process. Unlike traditional renderers, a differentiable renderer can not only generate images, but also calculate the gradient information between image pixels and rendering parameters (such as geometry, materials, lighting, etc.).

[0061] Specifically, the above-mentioned lip movement time series is rendered into a 2D image (i.e., a three-dimensional feature image) by a differentiable renderer, which reflects the shape, texture and lighting of the 3D lip image.

[0062] During model training, the process of rendering the 3D lip model into a 2D image is designed as a differentiable operation, and then backpropagation is used to optimize the parameters of the 3D lip model. The generated 2D image is compared with the real image, the loss is calculated, and the model parameters are updated through gradients, thereby achieving efficient 3D reconstruction and optimization.

[0063] The above embodiments use a fixed encoder in DECA to predict parameters such as identity, neck pose, and albedo for each frame, and extract chin parameters by designing a perceptual encoder. In particular, the use of perceptual lip motion loss better preserves the details of mouth movements. A coarse 3D model of the lips is estimated using a 3D image contour model, and a differentiable renderer is used to generate a textured 3D facial mesh, obtaining a more feature-rich 3D lip image sequence, providing accurate and effective images for subsequent lip-reading recognition. Through parameter learning and adjustment of the fixed encoder, perceptual encoder, and differentiable renderer, lip reconstruction from 2D to 3D is achieved, significantly enhancing feature capture capabilities. It effectively addresses changes in lighting and viewing angle, ensuring the accuracy of lip-reading recognition under different lighting conditions and camera perspectives.

[0064] In one embodiment, such as Figure 5 As shown, Figure 5 The diagram illustrates the structure of the feature extraction model (i.e., the Efficient-GhostNet network), which includes sequentially connected two-dimensional convolutional modules, multiple Ghost bottleneck modules, an average pooling module, and fully connected layers. Figure 5 The improved GhostNet proposed in this application mainly consists of a series of Ghost bottleneck modules: First, the first layer is a 3×3 convolutional kernel with 16 channels, and then Ghost bottleneck modules of different lengths are stacked to increase the number of channels and change the size of the feature maps. Finally, global average pooling and 2D convolution are used to transform the features into 1280-dimensional feature vectors. Compared with the original GhostNet, the improved model introduces a more efficient channel attention module, replacing the squeezing and activation modules, which effectively reduces the number of parameters while maintaining recognition accuracy.

[0065] Step 103 above includes: The 3D feature image of the lips is input into a 2D convolutional module to obtain a basic feature map. The basic feature map is then input into multiple Ghost bottleneck modules to obtain Ghost feature maps output by these modules. These Ghost bottleneck modules are connected sequentially, with the output of the previous Ghost bottleneck module serving as the input to the next. The Ghost feature maps are then input into an average pooling module to obtain a spatial feature map output by the average pooling module. Finally, the spatial feature maps are input into a fully connected layer to obtain a spatial feature vector output by the fully connected layer.

[0066] Specifically, let the three-dimensional feature image of the lips be... Where c represents the number of channels, height, and width of the 3D feature image X. First, a basic feature map is generated using ordinary 2D convolution (i.e., a two-dimensional convolution module) to complete the main convolution operation, as follows: ; (1) in, This represents the convolution operation. Represents m basic feature maps, It is a convolutional kernel used to generate basic feature maps.

[0067] The Ghost bottleneck module is designed based on the following two key ideas: (1) Ghost module: The Ghost module generates redundant feature maps through inexpensive operations (such as depthwise separable convolution or linear transformation), thereby reducing the computational cost of traditional convolution operations. It first generates a portion of the feature maps through a small number of convolution operations, and then generates the remaining feature maps through inexpensive operations.

[0068] (2) Bottleneck structure: The bottleneck structure reduces the amount of computation by first compressing the number of channels in the feature map, then performing convolution operations, and finally expanding the number of channels. It usually consists of three convolutional layers: 1x1 convolution (compression), 3x3 convolution (feature extraction), and 1x1 convolution (expansion).

[0069] In this embodiment, in order to further obtain the required n spatial feature vectors, the basic feature map is... Perform some ordinary convolution and depthwise separable convolution operations (i.e., input to the Ghost bottleneck module) to generate s Ghost feature maps. The specific operations are as follows: ; (2) in, yes The i-th basic feature map in the middle, This is the transformation operation that generates the j-th Ghost feature map. Each basic feature map... It can generate one or more Ghost feature maps. This represents the j-th Ghost feature map corresponding to the i-th basic feature map.

[0070] Finally, the Ghost feature map is input into the average pooling module to obtain the spatial feature map output by the average pooling module; the spatial feature map is then input into the fully connected layer to obtain the spatial feature vector output by the fully connected layer. The final output spatial feature vector Y consists of m×s feature maps. Since the linear operation of each channel consists of depthwise separable convolutions and ordinary convolutions, the computational cost of this structure is much lower than that of a channel composed entirely of ordinary convolutions.

[0071] The above embodiments, through the design of the Ghost bottleneck module, generate these redundant feature maps through simple linear operations, thereby reducing computational load. Compared with traditional convolution operations that generate a large number of feature maps, this reduces redundant feature maps, thus reducing computational load and achieving a lightweight design that maintains high-performance computation while reducing redundant features. Simultaneously, by using a combination of depthwise separable convolution and ordinary convolution, it achieves efficient channel self-attention computation, reducing computational overhead and enabling the model to process more data in a shorter time, thereby accelerating recognition speed and allowing it to run on resource-constrained devices. Furthermore, the Efficient-GhostNet model structure, used to extract spatial features of lip movement sequences, reduces the number of training layers, effectively improving training speed by 45%, making it suitable for home scenarios requiring efficient computation and storage. The Efficient-GhostNet design improves the Ghost module, reducing redundant feature maps while maintaining high-performance computation. By using a combination of depthwise separable convolution and ordinary convolution, it achieves efficient channel self-attention computation, significantly reducing computational overhead. This model can process more data in a shorter time, thereby accelerating recognition speed and running efficiently on resource-constrained devices, meeting the real-time processing requirements of smart home environments.

[0072] In one embodiment, the gated loop unit (GRU) includes an update gate, a reset gate, and a forget gate; step 104 includes: inputting a spatial feature vector into the gated loop unit so that the update gate controls the information of the previous state and the reset gate controls the previous state, thereby obtaining lip movement time features.

[0073] Specifically, for lip reading tasks, it is necessary to analyze the entire lip reading process and obtain time-related lip reading sequences. In deep learning, RNN (Recurrent Neural Network) is the most commonly used model for handling time-related problems, but RNN models suffer from long-term dependencies, vanishing gradients, and exploding gradients.

[0074] To address the vanishing and exploding gradient problems in RNNs, the GRU module in this application uses update and reset gates instead of traditional input, output, and forget gates. The update gate controls the transfer of information from the previous state to the current state, while the reset gate controls the writing of information from the previous state to the current candidate state, thus preserving much sequential information and removing irrelevant information. The formulas for each part of the GRU network are as follows: ; (3) in, The output of the Reset Gate indicates the hidden state at the previous moment. How much information needs to be ignored; Indicates input The weight matrix to the reset gate; Indicates the hidden state at the previous moment. The weight matrix to the reset gate; This indicates the offset of the door being reset; This indicates the sigmoid activation function, which restricts the output to the range [0, 1].

[0075] ; (4) in, The output of the update gate determines the hidden state at the current moment. How much of it comes from the hidden state of the previous moment? How many candidate hidden states are there from the current moment? ; Indicates input Update the weight matrix of the gate; Indicates the hidden state at the previous moment. Update the weight matrix of the gate; This indicates that the bias term of the updated gate is being updated.

[0076] ; (5) in, Represents the candidate hidden state, and represents the potential hidden state at the current moment; Indicates input The weight matrix to the candidate hidden state; Indicates the hidden state at the previous moment. The weight matrix to the candidate hidden state; The term represents the bias of the candidate hidden state; ⊙ represents element-wise multiplication, used to reset the gate. Applied to ; tanh is the hyperbolic tangent activation function, which restricts the output to the range [-1, 1].

[0077] ; (6) in, This represents the final hidden state at the current moment, which is a candidate hidden state. and the hidden state of the previous moment The weighted sum. and This indicates the output of the update gate, which controls the output of each gate. and The weight.

[0078] ; (7) in, This represents the output at the current moment, typically used for classification or prediction tasks. Indicates hidden state The weight matrix to be output; This indicates the sigmoid activation function, which restricts the output to the range [0, 1].

[0079] The spatial feature sequences extracted from GhostNet are input into the GRU. The GRU processes the sequence data through a gating mechanism to capture the temporal dynamics of lip movements. Simultaneously, the GRU retains important sequence information and filters out unnecessary noise and redundant information through update and reset gates, thereby improving recognition accuracy.

[0080] Finally, by combining the GRU and SoftMax layers, the lip reading model can efficiently process and classify lip movement sequences, thereby achieving high-accuracy lip reading recognition.

[0081] The above embodiment employs a variant of the recurrent neural network, GRU, designed to address the vanishing and exploding gradient problems inherent in recurrent neural networks. GRU introduces a gating mechanism to control the flow of information, thereby more effectively capturing long-short-term dependencies in sequence data. When lip occlusion occurs, this module captures the temporal dependencies in the lip movement sequence to achieve contextual consistency of the recognized content, thus improving the accuracy and robustness of lip reading. Furthermore, the GRU module has fewer parameters and faster training speed, further enhancing the model's lightweight nature. When faced with occlusion, GRU can utilize information from previous time steps to supplement the occluded information in the current time step, ensuring the accuracy of lip reading under occlusion conditions.

[0082] In one embodiment, this application also provides a device wake-up method, including the following steps: detecting whether a video frame contains a complete lip image; if no complete lip image is detected in each frame of the video frame, then applying the steps as described in any of the above lip reading recognition method embodiments to obtain the Chinese lip reading prediction result; and displaying the lip reading recognition result.

[0083] Optionally, if the video frame contains a complete lip image, the complete lip image is compared with an image in a preset image feature library to obtain the lip reading recognition result.

[0084] Specifically, such as Figure 6 As shown, Figure 6 The diagram illustrates the interaction flow of the aforementioned device wake-up method, applied to a voice assistant with a camera. During this process, the system process polls the face detection module to confirm the user's presence in the camera's field of view, collecting visual and lip information. If the lip image is available, the system completes recognition and provides feedback. If the lip image is unavailable, the system activates the lip movement recognition module, extracting lip features from video frames and checking again for lip image availability. If available, it performs offline image feature library comparison, completing recognition and providing feedback. If still unavailable, the system initiates active voice prompts and commands, ultimately completing the lip reading task.

[0085] In the above embodiments, it is detected whether a complete lip image is contained in the video frame; if no complete lip image is detected in any frame of the video frame, the steps as described in any of the above lip reading recognition method embodiments are applied to obtain the Chinese lip reading prediction result; and the lip reading recognition result is displayed. This method can effectively cope with changes in lighting, viewing angle, occlusion, and real-time requirements in real-world home application scenarios, significantly improving the accuracy of lip reading recognition and user experience.

[0086] The lip-reading recognition device provided by the present invention is described below. The lip-reading recognition device described below can be referred to in correspondence with the lip-reading recognition method described above.

[0087] like Figure 7 As shown, a lip-reading recognition device is provided, which includes the following modules: Image frame extraction module 701 is used to extract multiple image frames from a video clip; The three-dimensional feature extraction module 702 is used to input the multiple image frames into the three-dimensional lip modeling model to obtain the three-dimensional feature image output by the three-dimensional lip modeling model; The spatial feature vector output module 703 is used to input the three-dimensional feature image of the lip into the feature extraction model to obtain the spatial feature vector output by the feature extraction model. The lip movement time feature extraction module 704 is used to input the spatial feature vector into the gated loop unit to obtain the lip movement time features extracted by the gated loop unit; The lip reading result output module 705 is used to input the temporal features of the lip movements into the classification layer to obtain the lip reading result output by the classification layer.

[0088] In one embodiment, the three-dimensional lip modeling model includes a fixed encoder, a perceptual encoder, a three-dimensional image contour model, and a differentiable renderer. The aforementioned three-dimensional feature extraction module 702 is further used for: The three-dimensional feature image is input into the perceptual encoder to obtain the chin parameters obtained by the perceptual encoder using spatial perception loss. The three-dimensional feature image is input into the fixed encoder to obtain the identity parameters and albedo parameters output by the fixed encoder. The chin parameter, the identity parameter, and the albedo parameter are input into the three-dimensional image contour model to obtain the lip movement time series output by the three-dimensional image contour model; The lip movement time series is input into the differentiable renderer to obtain the three-dimensional feature image output by the differentiable renderer.

[0089] In one embodiment, the feature extraction model includes a two-dimensional convolutional module, multiple Ghost bottleneck modules, an average pooling module, and a fully connected layer connected in sequence; the aforementioned spatial feature vector output module 703 is further used for: The three-dimensional feature image is input into the two-dimensional convolution module to obtain the basic feature map; The basic feature map is input to the plurality of Ghost bottleneck modules to obtain Ghost feature maps output by the plurality of Ghost bottleneck modules; wherein, the plurality of Ghost bottleneck modules are connected in sequence, and the output of the previous Ghost bottleneck module is used as the input of the next Ghost bottleneck module; The Ghost feature map is input into the average pooling module to obtain the spatial feature map output by the average pooling module; The spatial feature map is input into the fully connected layer to obtain the spatial feature vector output by the fully connected layer.

[0090] In one embodiment, the gated loop unit includes an update gate, a reset gate, and a forget gate; the aforementioned lip movement time feature extraction module 704 is further used for: The spatial feature vector is input into the gated loop unit so that the update gate controls the information of the previous state and the reset gate controls the previous state, thereby obtaining the lip movement time features.

[0091] In one embodiment, a device wake-up device is also provided, comprising: The lip image detection module is used to detect whether a video frame contains a complete lip image; The lip reading module is used to obtain the lip reading result by applying the steps in any of the above embodiments of the lip reading method if no complete lip image is detected in any frame of the video frame. The lip reading result display module is used to display the lip reading results.

[0092] In one embodiment, after detecting whether the video frame contains a complete lip image, the method further includes: The aforementioned lip reading module is further configured to, if the video frame contains a complete lip image, compare the complete lip image with an image in a preset image feature library to obtain the lip reading result.

[0093] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a lip-reading recognition method. This method includes: extracting multiple image frames from a video clip; inputting the multiple image frames into a three-dimensional lip modeling model to obtain a three-dimensional feature image output by the three-dimensional lip modeling model; inputting the three-dimensional feature image of the lips into a feature extraction model to obtain a spatial feature vector output by the feature extraction model; inputting the spatial feature vector into a gated loop unit to obtain lip movement temporal features extracted by the gated loop unit; and inputting the lip movement temporal features into a classification layer to obtain a lip-reading recognition result output by the classification layer.

[0094] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0095] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the lip-reading recognition method provided by the above methods. The method includes: extracting multiple image frames from a video clip; inputting the multiple image frames into a three-dimensional lip modeling model to obtain a three-dimensional feature image output by the three-dimensional lip modeling model; inputting the three-dimensional feature image of the lips into a feature extraction model to obtain a spatial feature vector output by the feature extraction model; inputting the spatial feature vector into a gated loop unit to obtain lip movement temporal features extracted by the gated loop unit; and inputting the lip movement temporal features into a classification layer to obtain a lip-reading recognition result output by the classification layer.

[0096] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the lip-reading recognition method provided by the above methods. The method includes: extracting multiple image frames from a video clip; inputting the multiple image frames into a three-dimensional lip modeling model to obtain a three-dimensional feature image output by the three-dimensional lip modeling model; inputting the three-dimensional feature image of the lips into a feature extraction model to obtain a spatial feature vector output by the feature extraction model; inputting the spatial feature vector into a gated loop unit to obtain lip movement temporal features extracted by the gated loop unit; and inputting the lip movement temporal features into a classification layer to obtain a lip-reading recognition result output by the classification layer.

[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0098] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lip-reading recognition method, characterized in that, include: Extract multiple image frames from a video clip; The multiple image frames are input into the three-dimensional lip modeling model to obtain the three-dimensional feature image of the lips output by the three-dimensional lip modeling model; The three-dimensional feature image of the lips is input into the feature extraction model to obtain the spatial feature vector output by the feature extraction model. The spatial feature vector is input into the gated loop unit to obtain the lip movement time features extracted by the gated loop unit; The lip movement time features are input into the classification layer to obtain the lip reading recognition results output by the classification layer.

2. The lip-reading recognition method according to claim 1, characterized in that, The three-dimensional lip modeling model includes a fixed encoder, a perceptual encoder, a three-dimensional image contour model, and a differentiable renderer. The step of inputting the multiple image frames into the three-dimensional lip modeling model to obtain the three-dimensional feature image of the lips output by the three-dimensional lip modeling model includes: The three-dimensional feature image of the lips is input into the perceptual encoder to obtain the chin parameters obtained by the perceptual encoder using spatial perceptual loss. The three-dimensional feature image of the lips is input into the fixed encoder to obtain the rigid transformation parameters and albedo parameters output by the fixed encoder. The chin parameters, the rigidity transformation parameters, and the albedo parameters are input into the three-dimensional image contour model to obtain the lip movement time series output by the three-dimensional image contour model. The lip movement time series is input into the differentiable renderer to obtain the three-dimensional feature image output by the differentiable renderer.

3. The lip-reading recognition method according to claim 1, characterized in that, The feature extraction model includes a two-dimensional convolutional module, multiple Ghost bottleneck modules, an average pooling module, and a fully connected layer connected in sequence; the step of inputting the three-dimensional feature image of the lip into the feature extraction model to obtain the spatial feature vector output by the feature extraction model includes: The three-dimensional feature image of the lips is input into the two-dimensional convolution module to obtain the basic feature map; The basic feature map is input to the plurality of Ghost bottleneck modules to obtain Ghost feature maps output by the plurality of Ghost bottleneck modules; wherein, the plurality of Ghost bottleneck modules are connected in sequence, and the output of the previous Ghost bottleneck module is used as the input of the next Ghost bottleneck module; The Ghost feature map is input into the average pooling module to obtain the spatial feature map output by the average pooling module; The spatial feature map is input into the fully connected layer to obtain the spatial feature vector output by the fully connected layer.

4. The lip-reading recognition method according to claim 1, characterized in that, The gated loop unit includes an update gate, a reset gate, and a forget gate; the step of inputting the spatial feature vector into the gated loop unit to obtain the lip movement temporal features extracted by the gated loop unit includes: The spatial feature vector is input into the gated loop unit so that the update gate controls the information of the previous state and the reset gate controls the previous state, thereby obtaining the lip movement time features.

5. A device wake-up method, characterized in that, include: Detect whether a video frame contains a complete lip image; If no complete lip image is detected in any frame of the video frame, the lip reading recognition method as described in any one of claims 1 to 4 is applied to obtain the Chinese lip reading prediction result; The lip reading results are displayed.

6. The device wake-up method according to claim 5, characterized in that, After detecting whether the video frame contains a complete lip image, the process further includes: If the video frame contains a complete lip image, the complete lip image is compared with an image in a preset image feature library to obtain the lip reading recognition result.

7. A lip-reading recognition device, characterized in that, include: The image frame extraction module is used to extract multiple image frames from a video clip; The three-dimensional feature extraction module is used to input the multiple image frames into the three-dimensional lip modeling model to obtain the three-dimensional feature image of the lip output by the three-dimensional lip modeling model; The spatial feature vector output module is used to input the three-dimensional feature image of the lips into the feature extraction model to obtain the spatial feature vector output by the feature extraction model. The lip movement time feature extraction module is used to input the spatial feature vector into the gated loop unit to obtain the lip movement time features extracted by the gated loop unit. The lip reading result output module is used to input the lip movement time features into the classification layer to obtain the lip reading result output by the classification layer.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the lip reading method as described in any one of claims 1 to 4, or implements the device wake-up method as described in any one of claims 5 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the lip reading method as described in any one of claims 1 to 4, or implements the device wake-up method as described in any one of claims 5 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the lip reading method as described in any one of claims 1 to 4, or implements the device wake-up method as described in any one of claims 5 to 6.