Lip reading method and device based on event, equipment and storage medium
By collecting lip sequence images through an event camera and using voxel representation and lip reading model to extract spatial features and model time dependence, the problem of insufficient lip reading accuracy in existing methods is solved and stable lip reading recognition is achieved.
Patent Information
- Application Number
- CN202511122198.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing event-based lip reading methods cannot effectively capture lip structure information, resulting in inaccurate lip reading results.
An event-based lip reading model is adopted. The original event stream of lip sequence images is collected by an event camera, and converted into a frame-like event tensor using a voxel representation method. Spatial features are extracted by combining 3D convolutional layers and deep neural networks. Temporal dependency modeling is performed through a temporal semantic hypergraph and a recurrent neural network, and lip reading recognition is performed using a viseme supervision strategy.
The accuracy and stability of lip reading are improved, and stable recognition of lip movements is achieved.
Smart Images

Figure CN120635994A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an event-based lip reading method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With its microsecond temporal resolution and sparse visual encoding, event cameras offer a revolutionary paradigm for automatic lip reading (ALR). As an important complement to auditory speech recognition, ALR shows great potential in complex scenarios where speech signals are unclear or severely degraded.
[0003] Most existing event-based lip reading methods are based on 3D convolutional or recurrent neural networks, which focus on modeling temporal dynamics. These methods typically follow the design paradigm of video-based lip reading, implicitly assuming that event streams provide structurally coherent and semantically complete spatiotemporal representations. However, event data is fundamentally different from traditional frame-based images. They lack explicit spatial texture and continuous visual structure, failing to capture critical lip structural information, resulting in inaccurate lip reading results. Summary of the Invention
[0004] Based on this, it is necessary to provide an event-based lip-reading method, apparatus, computer device, computer-readable storage medium and computer program product that can improve lip-reading accuracy in order to address the above technical problems.
[0005] In a first aspect, the present application provides an event-based lip reading method, the method comprising:
[0006] The raw event stream of the lip sequence images is collected by the event camera;
[0007] The original event stream is converted into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume;
[0008] Extracting spatial features from the voxelized event volume through a front-end network in an event-based lip reading model to obtain multi-scale spatial features;
[0009] Modeling the time dependency of the multi-scale spatial features through a back-end sequence model in an event-based lip reading model to obtain a sequence encoding;
[0010] The content of lip reading recognition is determined according to the sequence code.
[0011] In one embodiment, the raw event stream of the lip sequence image captured by the event camera includes:
[0012] Asynchronously recording the brightness change of each pixel of the lip sequence image through an event camera, and triggering an event when the logarithmic intensity change of the pixel exceeds a preset contrast threshold;
[0013] Events occurring within a preset time window are acquired to obtain the original event stream.
[0014] In one embodiment, the voxel-based representation method converts the original event stream into a frame-like event tensor to obtain a voxelized event volume, including:
[0015] According to the events in the original event stream and the given time interval, a voxel grid method is used to scale the event timestamps to the given time interval to obtain a voxelized event volume.
[0016] In one embodiment, the front-end network includes: a 3D convolutional layer and a deep neural network. The front-end network in the event-based lip reading model extracts spatial features from the voxelized event volume to obtain multi-scale spatial features, including:
[0017] Performing preliminary spatiotemporal feature extraction processing on the voxelized event volume through the 3D convolution layer to obtain initial spatiotemporal features;
[0018] The initial spatiotemporal features are sequentially input into the convolutional layer and multiple residual blocks of the deep neural network, the outputs of the multiple residual blocks are subjected to global average pooling to obtain a high-level spatial feature map, and the outputs of the multiple residual blocks are subjected to shallow adaptive pooling to obtain multiple spatial feature maps;
[0019] Using the intermediate layer feature maps of multiple spatial feature maps as the reference resolution for alignment, the spatial dimensions of multi-level features are reshaped into a sequence of node representations for graph construction.
[0020] Inputting the node representation sequence in parallel into a frequency-aware modulation module for frequency enhancement processing to obtain enhanced features; the frequency enhancement processing includes: injecting noise into a first frequency band and enhancing discriminative edge dynamics of a second frequency band using adaptive frequency filtering; the lower limit frequency of the second frequency band is higher than the upper limit frequency of the first frequency band;
[0021] Inputting the enhanced features into the fully connected layer of the deep neural network to obtain multi-level spatial node features;
[0022] Inputting the multi-level spatial node features into the spatial region hypergraph respectively to extract high-order intra-frame dependencies between lip regions;
[0023] Multi-scale spatial features are obtained based on the high-order intra-frame dependencies between lip regions and high-level spatial feature maps.
[0024] In one embodiment, the time dependency of the multi-scale spatial features is modeled by a back-end sequence model in an event-based lip reading model to obtain a sequence encoding, including:
[0025] Converting the multi-scale spatial features into temporally structured features through a temporal semantic hypergraph;
[0026] A recurrent neural network is used to model the temporal structured features to obtain sequence encoding.
[0027] In one embodiment, determining the content of lip reading recognition based on the sequence code includes:
[0028] The target lip-reading model determines the content of the lip-reading according to the sequence encoding; wherein the target lip-reading model adopts a viseme-based supervision strategy, and the viseme-based supervision strategy is used to encode visual similarity into a label space.
[0029] In one embodiment, the viseme-based supervision strategy includes:
[0030] Associating the normalized word labels with phoneme sequences, and converting each phoneme sequence into a corresponding viseme sequence according to the preset phoneme-to-viseme mapping rules;
[0031] The similarity between viseme sequences of different words is quantified by the edit distance. The minimum viseme edit distance between any two different words is 1. The smaller the viseme edit distance, the more similar the words are in phonetics.
[0032] Construct a distance matrix by calculating pairwise viseme edit distances;
[0033] Converting the distance matrix into a probability distribution by row-wise normalization;
[0034] Build soft labels for each word based on the probability distribution.
[0035] In a second aspect, the present application further provides an event-based lip reading device, comprising:
[0036] An acquisition module, used for acquiring the original event stream of lip sequence images through an event camera;
[0037] A conversion module, configured to convert the original event stream into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume;
[0038] a feature extraction module, configured to extract spatial features from the voxelized event volume using a front-end network in an event-based lip reading model to obtain multi-scale spatial features;
[0039] A temporal modeling module, configured to perform temporal dependency modeling on the multi-scale spatial features using a back-end sequence model in an event-based lip reading model to obtain a sequence code;
[0040] The recognition module is used to determine the content of lip reading recognition according to the sequence code.
[0041] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0042] The raw event stream of the lip sequence images is collected by the event camera;
[0043] The original event stream is converted into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume;
[0044] Extracting spatial features from the voxelized event volume through a front-end network in an event-based lip reading model to obtain multi-scale spatial features;
[0045] Modeling the time dependency of the multi-scale spatial features through a back-end sequence model in an event-based lip reading model to obtain a sequence encoding;
[0046] The content of lip reading recognition is determined according to the sequence code.
[0047] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0048] The raw event stream of the lip sequence images is collected by the event camera;
[0049] The original event stream is converted into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume;
[0050] Extracting spatial features from the voxelized event volume through a front-end network in an event-based lip reading model to obtain multi-scale spatial features;
[0051] Modeling the time dependency of the multi-scale spatial features through a back-end sequence model in an event-based lip reading model to obtain a sequence encoding;
[0052] The content of lip reading recognition is determined according to the sequence code.
[0053] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0054] The raw event stream of the lip sequence images is collected by the event camera;
[0055] The original event stream is converted into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume;
[0056] Extracting spatial features from the voxelized event volume through a front-end network in an event-based lip reading model to obtain multi-scale spatial features;
[0057] Modeling the time dependency of the multi-scale spatial features through a back-end sequence model in an event-based lip reading model to obtain a sequence encoding;
[0058] The content of lip reading recognition is determined according to the sequence code.
[0059] The event-based lip-reading method, apparatus, computer device, computer-readable storage medium, and computer program product described above use an event camera to capture a raw event stream of sequential lip images. This allows for asynchronous recording of each pixel's brightness changes with microsecond temporal resolution, achieving ultra-low latency, high dynamic range, and sparse data representation. A voxel-based representation method converts the raw event stream into a frame-like event tensor, resulting in a voxelized event volume. This allows the event data to be converted into a video-like representation, facilitating subsequent feature extraction. The front-end network in the event-based lip-reading model extracts spatial features from the voxelized event volume, generating multi-scale spatial features. The back-end sequence model in the event-based lip-reading model models the temporal dependencies of these multi-scale spatial features, generating a sequence code. This allows for the fusion of spatial and temporal features, improving the accuracy and stability of the sequence code. The lip-reading content is determined based on the sequence code. This enables stable recognition of lip movements and improves lip-reading accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0061] Figure 1 is a flow chart of an event-based lip reading method according to one embodiment;
[0062] Figure 2 for Figure 1 Schematic diagram of the process of step 103 in the embodiment;
[0063] Figure 3 (a) is a standard diagram in one embodiment;
[0064] Figure 3(b) is a hypergraph in one embodiment;
[0065] Figure 4 FIG1 is a schematic diagram of the architecture of an event-based lip reading model in one embodiment;
[0066] Figure 5 Schematic diagram comparing the viseme-based supervision strategy with the traditional one-hot supervision strategy;
[0067] Figure 6 is a structural block diagram of an event-based lip reading device in one embodiment;
[0068] Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0070] It should be noted that the terms "including" and "having" and any variations thereof used in this application are intended to cover non-exclusive inclusions. The term "plurality" used in this application refers to two or more. The term "and / or" used in this application refers to one or more solutions.
[0071] To facilitate understanding of the technical solutions in various embodiments of the present application, the following explanations are first given of the professional terms that may appear in the embodiments of the present application:
[0072] An event-based camera (or simply an event camera) is a new type of sensor that detects changes in the brightness of a pixel and generates an event when the cumulative brightness change reaches a certain threshold. The output of an event camera is related to the brightness change, not the absolute value of the brightness. An event indicates when and at which pixel the brightness increased or decreased.
[0073] Voxel representation is a 3D modeling technique in computer graphics that breaks down an object into a series of voxels, or three-dimensional pixels, similar to pixels in a 2D image. Each voxel represents a small cube in space, and the shape and structure of a 3D object are represented by a collection of these voxels.
[0074] Multi-scale spatial features: First, image scale does not refer to the size of the image, but rather the degree of blur. For example, the degree of blur when viewing an object from close up differs from that when viewed from a distance. The increasing blurriness of an image from close to far is also a process of increasing image scale. Multi-scale spatial features are features extracted at different spatial scales. For example, by analyzing a signal using filters of different scales, feature information at different scales is obtained to comprehensively describe the image or signal.
[0075] Lip reading model: refers to a model used to recognize lip movements. It can process video and audio information, and by fusing the two modal data, it can identify the sound produced by the speaker to improve the accuracy of lip reading recognition.
[0076] Sequence coding: It is an important step for lip reading models to perform lip reading recognition. For example, the image sequence coding containing spatiotemporal features can be input into the lip reading model, and the lip reading model outputs the content corresponding to the image sequence coding.
[0077] Event cameras, with their microsecond temporal resolution and sparse visual encoding, offer a transformative paradigm for automatic lip reading. However, event data inherently lacks a clear spatial structure and exhibits significant frequency domain bias. Low-frequency components fail to capture critical lip structural information, fundamentally hindering the modeling of intra-frame topological dependencies and inter-frame semantic evolution, both of which are crucial for robust lip reading.
[0078] In response to the problems existing in the prior art, the present application aims to provide an event-based lip-reading method, which adopts an event-based lip-reading model (a frequency-aware spatiotemporal hypergraph framework) to improve the robustness of the model in capturing discriminative features and integrates adaptive high-frequency filtering to enhance edge-aware representation. Optionally, the present application also constructs a spatial region hypergraph and a temporal semantic hypergraph. The former is used to capture the intra-frame topological dependencies between lip regions, while the latter explicitly models the inter-frame structural associations during the entire lip movement process, thereby enabling the model to capture discriminative patterns in lip dynamics. In addition, the present application also provides a viseme-based label smoothing strategy, which can use viseme-level edit distance to quantify the visual similarity between categories and guide the construction of soft labels.
[0079] In an exemplary embodiment, Figure 1 As shown, an event-based lip reading method is provided, and the method may include the following steps 101 to 105. Among them:
[0080] Step 101: collect the original event stream of lip sequence images through an event camera.
[0081] It should be understood that the method in this embodiment can be applied to automatic lip reading scenarios, such as hearing aids, smart cockpits, human-computer interaction, and security monitoring.
[0082] Exemplarily, the brightness change of each pixel of the lip sequence image can be asynchronously recorded by an event camera, and an event is triggered when the logarithmic intensity change of the pixel exceeds a preset contrast threshold; the events occurring within the preset time window are acquired to obtain the original event stream.
[0083] The calculation formula of the preset contrast threshold is as follows:
[0084]
[0085] Where: Indicates the brightness intensity of the scene, Indicates the polarity of the brightness change (p=1 means increase, p=-1 means decrease). is a predefined contrast threshold, Represents the time elapsed since the last event was triggered at position (x, y). Therefore, the output of the event camera can be represented as a stream of events occurring within the time window [T0, T1]. The event stream is calculated as follows:
[0086]
[0087] Where: Represents a collection of event streams, represents the kth event, Represents the pixel horizontal coordinate of the event in space, Indicates the timestamp of the event. Indicates the pixel ordinate of the event in space, Indicates the polarity of the event (i.e., the direction of brightness change), and N represents the total number of events in the event stream, that is, the total number of events collected in the [T0, T1] time period.
[0088] Step 102 : Convert the original event stream into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume.
[0089] Exemplarily, based on the events in the original event stream and a given time interval, a voxel grid method may be used to scale the event timestamps to within a given time interval to obtain a voxelized event volume.
[0090] In this embodiment, a voxel-based representation method is used to convert the original event stream into a frame-like event tensor to preserve as much motion information as possible. Optionally, given a set of N input events:
[0091] and a time interval count T, the voxel grid method is used to first scale the event timestamp to the range [0, T−1], and then generate a T × H × W voxelized event volume :
[0092]
[0093]
[0094] Where: represents the normalized time index, T × H × W represents the time dimension, height and width in the spatial dimension of the event data respectively, Indicates the Nth timestamp, t1 indicates the 1st timestamp, represents the voxel intensity at position (x,y) and time t in the voxel grid.
[0095] After voxelization, event data is converted into a video-like representation X , where T is the number of event frames and the channel dimension C = 1 means that events of both polarities are accumulated together into each voxel frame.
[0096] Step 103 : extracting spatial features from the voxelized event volume through the front-end network in the event-based lip reading model to obtain multi-scale spatial features.
[0097] Among them, the front-end network may include: 3D convolutional layer and deep neural network.
[0098] Optionally, preliminary spatiotemporal feature extraction is performed on the input event frame X, and then deep semantic features are further extracted through Resnet-18.
[0099] For example, the front-end network can use a 3D convolutional layer with a kernel size of 5 × 7 × 7 and connect it to a ResNet-18 backbone network to extract spatial features from voxelized event frames.
[0100] Step 104 : Time dependency modeling of the multi-scale spatial features is performed using a back-end sequence model in the event-based lip reading model to obtain a sequence code.
[0101] Exemplarily, the multi-scale spatial features can be converted into time-structured features through a temporal semantic hypergraph; and a recurrent neural network is used to perform modeling based on the time-structured features to obtain sequence encoding.
[0102] In this embodiment, in the temporal modeling stage, the temporal structured feature representation is obtained through the temporal semantic hypergraph. :
[0103]
[0104] Where: Represented by Temporal Semantic Hypergraph Network HGNN TSH Extracted temporal features.
[0105] Optionally, a Bidirectional Gated Recurrent Unit (BiGRU) is used to model temporal dependencies and generate the final sequence encoding representation. :
[0106]
[0107] Where: Represents a bidirectional gated recurrent unit for modeling temporal dependencies.
[0108] In this embodiment, the temporal semantic hypergraph can clearly capture the semantic evolution between frames, thereby achieving holistic spatiotemporal representation learning.
[0109] Step 105: Determine the content of lip reading recognition according to the sequence code.
[0110] In this embodiment, the lip-reading content can be determined according to the sequence encoding by a target lip-reading model; wherein the target lip-reading model adopts a viseme-based supervision strategy, and the viseme-based supervision strategy is used to encode visual similarity into a label space.
[0111] Optionally, the viseme-based supervision strategy includes: associating normalized word labels with phoneme sequences, and converting each phoneme sequence into a corresponding viseme sequence according to a preset phoneme-to-viseme mapping rule; quantifying the similarity between viseme sequences of different words by editing distance; wherein the minimum viseme editing distance between any two different words is 1, and the smaller the viseme editing distance, the more phonetically similar the words are; constructing a distance matrix by calculating paired viseme editing distances; converting the distance matrix into a probability distribution by row-by-row normalization; and establishing soft labels for each word according to the probability distribution.
[0112] For example, Figure 5 As shown in Figure 3, a comparison between the viseme-based supervision strategy and the traditional one-hot supervision strategy is demonstrated. The viseme-based supervision strategy preserves the similarity of visemes between classes and achieves smoother label transition compared to the traditional one-hot supervision strategy.
[0113] In this example, lexical stress markers (e.g., "AH0" → and "ah") are removed, leaving only the basic phoneme units to obtain normalized word labels. Each word label is associated with a clean phoneme sequence. Each phoneme sequence is converted to its corresponding viseme sequence. Assume that the viseme sequences of two given words are represented as a = [a1, a2, . .. , a p ] and b = [b1, b2, . . . , b q ], where p and q represent the number of visemes in each sequence. Levenshtein distance is introduced to quantify the similarity between viseme sequences of different words. This metric is widely used in string matching and calculates the minimum number of editing operations required to transform one sequence into another, i.e., insertion, deletion, and substitution. Given two viseme sequences a = [a1, a2, . . . , a p ] and b = [b1, b2, . . ., b q ], the edit distance at the viseme level is defined as:
[0114]
[0115] where p = |a| and q = |b|.
[0116] Where: Denotes the minimum number of edit operations required to transform the first p items of sequence a into the first q items of b. represents the indicator function, if If true, it returns 1, otherwise it returns 0. To ensure that the model can effectively distinguish different lip pronunciation patterns, it is mandatory that the minimum viseme edit distance between any two different words is 1. The smaller the value, the more similar the two words are in phonetics and therefore more likely to be confused in the lip reading task. After calculating the pairwise viseme edit distance, we can get a distance matrix , where each element Dlev(i,j) represents the minimum number of edit operations required to transform the viseme sequence of word i into the viseme sequence of word j. Then, row-wise Softmax normalization is applied to convert these distances into probability distributions. Specifically, the soft label of the i-th word is given by Given, the smaller the distance, the higher the probability.
[0117] Where: Represents the soft label of the i-th category, softmax() represents converting the edit distance matrix into a probability distribution, It represents the temperature parameter that controls the sharpness of the soft label distribution. When is smaller, the resulting probability distribution becomes sharper, causing the model to focus more on categories that are highly similar to the target word. On the contrary, when When is larger, the distribution becomes smoother, which helps the model capture a wider range of inter-class visual similarities.
[0118] In this example, a viseme-based supervision strategy uses viseme-level edit distance to encode visual similarity into the label space, thereby constructing soft labels. Rather than treating all non-target categories as equally negative, the viseme-based supervision strategy introduces a structured supervisory signal, encouraging the model to learn visually coherent feature representations, thereby achieving clearer separation between categories.
[0119] In the aforementioned event-based lip-reading method, an event camera captures a raw event stream of sequential lip images. This allows for asynchronous recording of each pixel's brightness changes with microsecond temporal resolution, achieving ultra-low latency, high dynamic range, and sparse data representation. A voxel-based representation method converts the raw event stream into a frame-like event tensor, resulting in a voxelized event volume. This allows the event data to be converted into a video-like representation, facilitating subsequent feature extraction. The front-end network in the event-based lip-reading model extracts spatial features from the voxelized event volume, generating multi-scale spatial features. The back-end sequence model in the event-based lip-reading model models the temporal dependencies of these multi-scale spatial features, generating a sequence encoding. This allows for the fusion of spatial and temporal features, improving the accuracy and stability of the sequence encoding. The lip-reading content is determined based on the sequence encoding, enabling stable recognition of lip movements and improving lip-reading accuracy.
[0120] In an exemplary embodiment, Figure 2 As shown, step 103 includes steps 1031 to 1037. Among them:
[0121] Step 1031 : Perform preliminary spatiotemporal feature extraction processing on the voxelized event volume through the 3D convolution layer to obtain initial spatiotemporal features.
[0122] In this embodiment, a 3D convolutional layer with a kernel size of 5×7×7 can be used and connected to a ResNet-18 backbone network.
[0123] In step 1032, the initial spatiotemporal features are sequentially input into the convolutional layer and multiple residual blocks of the deep neural network. The outputs of the multiple residual blocks are subjected to global average pooling processing to obtain a high-level spatial feature map. The outputs of the multiple residual blocks are subjected to shallow adaptive pooling processing to obtain multiple spatial feature maps.
[0124] In this example, we take the input event frame X as an example to perform preliminary spatiotemporal feature extraction, and then further extract deep semantic features through Resnet-18. The deeper semantic representation captured by the Resnet-18 backbone network is as follows:
[0125]
[0126] in, Indicates the Residual Block The output, Represents multi-level spatiotemporal features. The final residual block output is globally average pooled to obtain the spatial feature map from ResNet The calculation formula is as follows:
[0127]
[0128] These four residual blocks encode hierarchical semantic information at gradually decreasing spatial resolutions. To align them spatially, adaptive pooling is applied to shallow features and bilinear interpolation is used to upsample deep features, ultimately resulting in a uniform spatial resolution. :
[0129]
[0130] Where: represents the alignment result of the spatial feature map of layer i, Indicates average pooling of shallow feature maps. Indicates bilinear interpolation of deep features.
[0131] In step 1033 , the intermediate layer feature maps of the multiple spatial feature maps are used as the reference resolution for alignment, and the spatial dimensions of the multi-level features are reshaped into a node representation sequence for graph construction.
[0132] For example, the third-level feature map can be selected as the reference resolution for alignment because it achieves a good balance between spatial granularity and computational efficiency. The spatial dimensions of the fourth-level features are reshaped into a sequence of node representations suitable for graph construction, denoted as ,in is the total number of spatial nodes.
[0133] Step 1034: Input the node representation sequence in parallel into a frequency-aware modulation module for frequency enhancement processing to obtain enhanced features.
[0134] The frequency enhancement processing includes: injecting noise into a first frequency band, and using adaptive frequency filtering to enhance the discriminative edge dynamics of a second frequency band; the lower limit frequency of the second frequency band is higher than the upper limit frequency of the first frequency band.
[0135] For example, spatial node feature sequences from four semantic levels are input in parallel to the frequency-aware modulation module. In the perturbation branch, noise is selectively injected into specific frequency bands to improve feature robustness, while the adaptive filtering branch emphasizes information-rich frequency components.
[0136] In this embodiment, the frequency-aware modulation module compensates for the missing spatial structure in the event representation by injecting low-frequency disturbances, and integrates an adaptive frequency filtering mechanism to enhance the dynamic encoding of the lip contour and motion trajectory.
[0137]
[0138] Where: Represents low-frequency features, represents the features after filtering, represents the low-frequency characteristics after noise disturbance, Represents the features after adaptive filtering.
[0139] Step 1035: input the enhanced features into the fully connected layer of the deep neural network to obtain multi-level spatial node features.
[0140] For example, all four layers of enhanced features are connected to form a multi-level spatial node representation, which is represented as and , the specific calculation formula is as follows:
[0141]
[0142] in, , Indicates the combined channel dimension.
[0143] Step 1036: Input the multi-level spatial node features into the spatial region hypergraph respectively to extract the high-order intra-frame dependencies between the lip regions.
[0144] and are input into the spatial region hypergraph to extract high-order intra-frame dependencies between lip regions. :
[0145]
[0146] Where: Represents the use of spatial region hypergraph neural network to analyze input features The output result after processing.
[0147] In this embodiment, the spatial region hypergraph can be used to model the high-order topological dependency between lip regions within a frame.
[0148] For example, Figures 3(a) and 3(b) show a comparison between standard graph representations and hypergraph representations. Specifically, at the viseme level, the standard graph and hypergraph representations of the word "methodology" are compared. The neural network in Figure 3(a) captures pairwise relationships, while the hypergraph neural network in Figure 3(b) can model high-order group interactions between multiple viseme subwords.
[0149] It should be understood that hypergraphs extend traditional graphs by allowing hyperedges to connect more than two nodes, making them well-suited for modeling higher-order relationships. As shown in Figure 3(b), compared to traditional neural networks that only capture pairwise relationships, hypergraph neural networks can model group interactions. In recommender systems, hypergraph structures are used to overcome oversmoothing and sparsity. Applying coarse-to-fine hypergraph modeling to action detection can capture multi-object temporal relationships.
[0150] Optionally, by constructing a spatial region hypergraph within each frame and a temporal semantic hypergraph between frames, the high-order structural dependencies of the lip region can be explicitly modeled in both spatial and temporal dimensions.
[0151] It should be understood that robust lip movement modeling can be achieved by combining frequency-aware feature refinement with hypergraph-based dependency reasoning in the event-based lip reading model.
[0152] Step 1037: Obtain multi-scale spatial features based on the high-order intra-frame dependencies between lip regions and the high-level spatial feature map.
[0153] In this embodiment, With advanced spatial features Connect them together to get a comprehensive multi-scale spatial representation :
[0154]
[0155] For example, Figure 4 FIG. 1 is a schematic diagram of an architecture of an event-based lip reading model in one embodiment. Figure 4As shown in the figure, the input event frame first passes through the 3D convolutional layer for preliminary spatiotemporal feature extraction. The extracted spatiotemporal features are then sequentially input into four residual blocks. The output of the first residual block is divided into two branches, one of which is transmitted to the second residual block, and the other is transmitted to the first frequency-aware multi-branch module. The output of the second residual block is divided into two branches, one of which is transmitted to the third residual block, and the other is transmitted to the second frequency-aware multi-branch module. The output of the third residual block is divided into two branches, one of which is transmitted to the fourth residual block, and the other is transmitted to the third frequency-aware multi-branch module. The output of the fourth residual block is divided into two branches, one of which is connected to the output of the temporal semantic hypergraph and transmitted to the bidirectional gated recurrent unit (BiGRU) before being sent to the classifier; the other branch is transmitted to the fourth frequency-aware multi-branch module.
[0156] like Figure 4 As shown, the outputs of the four frequency-aware multi-branch modules are each divided into two branches, denoted as the first branch and the second branch. The first branch of the first frequency-aware multi-branch module, the first branch of the second frequency-aware multi-branch module, the first branch of the third frequency-aware multi-branch module, and the first branch of the fourth frequency-aware multi-branch module are concatenated and input into the first spatial region hypergraph. The second branch of the first frequency-aware multi-branch module, the second branch of the second frequency-aware multi-branch module, the second branch of the third frequency-aware multi-branch module, and the second branch of the fourth frequency-aware multi-branch module are concatenated and input into the second spatial region hypergraph. The outputs of the two spatial region hypergraphs are concatenated and transmitted to the temporal semantic hypergraph.
[0157] like Figure 4 As shown, the frequency-aware multi-branch module can include: selective frequency routing and low-frequency perturbation branches. Data entering the selective frequency routing pair undergoes fast Fourier transform, learnable filter filtering, and inverse fast Fourier transform processing. Data entering the low-frequency perturbation branch undergoes fast Fourier transform, event low-frequency perturbation, and inverse fast Fourier transform processing.
[0158] Alternatively, a low-frequency perturbation branch injects controlled perturbations into low-frequency components to compensate for missing structural semantics, especially lip topology. Selective frequency routing adaptively emphasizes informative high-frequency features such as lip contours and motion boundaries while suppressing noisy responses.
[0159] For example, for an event sequence containing T frames, each frame contains Space nodes, the The node feature of the layer is represented as:
[0160]
[0161] Where: Represents the channel feature vector of the nth spatial node at time t.
[0162] To obtain the frequency domain representation, along each Apply a one-dimensional fast Fourier transform to the channel dimension of the node:
[0163]
[0164] Where: Represents the input features of node i under the nth branch in the tth frame , the discrete Fourier transform result of the c-th channel, Represents the value of the kth channel in the time domain signal, k represents the time domain channel index, and c represents the frequency domain channel index.
[0165] In the frequency enhancement branch, low-frequency components are extracted by applying a center-symmetric binary mask on the spectrum:
[0166]
[0167] in, Represents a low-pass filter mask, used to retain low-frequency components. Represents the proportion of the spectrum allocated to the low-frequency region. A binary mask is applied to the frequency representation to retain only the low-frequency components:
[0168]
[0169] in, Represents the low-frequency part after mask filtering, Represents an element-wise multiplication operation.
[0170] Next, the low-frequency spectrum of each node is modeled as a Gaussian distribution, and a controllable perturbation is introduced on this basis:
[0171]
[0172] in, represents the low-frequency part after disturbance, represents the perturbation intensity sampled from a uniform distribution, Indicates the corresponding frequency representation The standard deviation of represents uniform distribution, represents a hyperparameter that controls the perturbation range.
[0173] The perturbed frequency domain features are then transformed back to the original frequency domain through inverse fast Fourier transform, and then node-level feature extraction is performed by multi-layer perceptron (MLP):
[0174]
[0175] Where: IFFT stands for Inverse Fast Fourier Transform.
[0176] Different from the low-frequency band, the selective frequency routing branch directly performs adaptive filtering on the frequency domain features obtained by fast Fourier transform. A learnable filter W is introduced to modulate the spectral representation of each node in the frequency space:
[0177]
[0178] in: Represents the channel characteristics after weight screening in the frequency domain (high frequency selection), represents the learnable frequency channel weight, It represents the result of Fourier transform of feature z of node i, time frame t, and frequency branch n, that is, entering the frequency domain space.
[0179] Finally, the filtered frequency domain features are also transformed back to the spatial domain via IFFT and subsequently processed by MLP to obtain the node representation from the selective frequency routing branch:
[0180]
[0181] in: It represents the node features of node i in frame t and branch n, which are returned to the time domain after frequency channel selection.
[0182] After passing through the frequency-aware multi-branch module, each frame obtains two different spatial node representations:
[0183]
[0184]
[0185] It should be understood that Emphasizes stable low-frequency structural modes in node features, After suppressing background interference, it focuses on significant frequency domain information, thereby enhancing the model's sensitivity to key motion areas. These two representations are input in parallel to the spatiotemporal hypergraph module to capture richer spatiotemporal dependencies and high-order semantic relationships between nodes.
[0186] It should be understood that in order to explicitly model the spatiotemporal semantic structure in event data, we construct two types of hypergraphs: spatial region hypergraph and temporal semantic hypergraph. The spatial region hypergraph captures the high-order topological relationships between lip regions within each frame, while the temporal semantic hypergraph encodes the semantic evolution across frames.
[0187] For example, the spatial region hypergraph is defined as ,in represents the set of all spatial nodes in the frame, Represents a set of hyperedges used to model the high-order structural relationships between these nodes. For each node, the Euclidean distance to all other nodes is calculated, and its first k nearest neighbor nodes are selected. Then, hyperedges are formed by connecting the node with its k neighbors, effectively capturing local spatial dependencies. The resulting spatial region hypergraph can be encoded as a k nearest neighbor matrix:
[0188]
[0189] in, Represents a hypergraph adjacency matrix based on the construction, which applies hypergraph convolution to capture high-order structural dependencies between spatial nodes:
[0190]
[0191] in: Represents the node feature representation of the lth layer of the temporal semantic hypergraph, represents the activation function, represents the degree matrix of the node, represents the transposed adjacency matrix of the hypergraph, represents the hyperedge weight matrix, represents the degree matrix of the hyperedge, represents the node-hyperedge incidence matrix, represents the spatial feature representation of the l-th layer input, Represents the learnable spatial hypergraph convolution parameter matrix.
[0192] Average pooling is applied to the spatial region hypergraph node features from the low-frequency perturbation branch and the selective frequency routing branch. After T layers of spatial hypergraph convolution, the outputs of all layers are connected to obtain the spatial feature representation of the entire event sequence, which is expressed as:
[0193]
[0194]
[0195] in, represents the global spatial feature representation of the low-frequency branch after spatial hypergraph modeling, represents the spatial feature representation of the frequency selection branch after spatial hypergraph modeling, represents the output channel dimension of the spatial hypergraph module. These two representations are then combined with the baseline convolutional features Concatenated along the time axis T, we obtain a multi-scale spatial feature representation that integrates convolutional and hypergraph enhancement features.
[0196] After extracting spatial features through the spatial region hypergraph, a temporal semantic hypergraph is constructed as the backend temporal modeling module.
[0197] Among them, the temporal semantic hypergraph is defined as: .
[0198] in, Represents a set of nodes in a temporal hypergraph, where each node represents a frame-level spatial representation. represents the set of hyperedges in the temporal hypergraph, where each hyperedge connects a frame node with its temporal neighbors. The hyperedges in are constructed by computing the Euclidean distance between each frame and all other frames and connecting each frame to its top k nearest temporal neighbors.
[0199] Among them, the temporal hypergraph convolution formula is as follows:
[0200]
[0201] in, represents the frame node feature representation after the l+1th layer of temporal hypergraph convolution, represents the normalization term of the node degree matrix, represents the normalized term of the hyperedge degree matrix, Represents the time node features of the l-th layer input.
[0202] To further preserve the original temporal order between frames, a residual connection is introduced to add the output of the temporal hypergraph convolution to its input The obtained feature representation is then fed into a novel conditional image generation model (BiGR) module for global temporal modeling.
[0203] After testing on multiple challenging datasets, the event-based lip-reading model proposed in this application demonstrates significant advantages in identifying confusing words. This model outperforms existing technologies in both overall accuracy and the ability to discriminate between confusing subsets. This effectively improves the accuracy and robustness of lip-reading systems when processing visually similar word pairs, thereby enhancing the model's practicality and reliability.
[0204] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps. It is understandable that the various steps in different embodiments can be freely combined as needed, and the various non-contradictory schemes formed by the combination all fall within the scope of protection of this application.
[0205] Based on the same inventive concept, embodiments of the present application also provide an event-based lip-reading device for implementing the aforementioned event-based lip-reading method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more event-based lip-reading device embodiments provided below can be found in the aforementioned limitations of the event-based lip-reading method and will not be further elaborated here.
[0206] In an exemplary embodiment, Figure 6 As shown, an event-based lip reading device is provided, comprising: an acquisition module 601, a conversion module 602, a feature extraction module 603, a time modeling module 604 and a recognition module 605, wherein:
[0207] An acquisition module 601 is configured to acquire an original event stream of a lip sequence image through an event camera;
[0208] A conversion module 602 is configured to convert the original event stream into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume;
[0209] A feature extraction module 603 is configured to extract spatial features from the voxelized event volume using a front-end network in an event-based lip reading model to obtain multi-scale spatial features;
[0210] A time modeling module 604 is configured to perform time dependency modeling on the multi-scale spatial features using a back-end sequence model in an event-based lip reading model to obtain a sequence code;
[0211] The recognition module 605 is configured to determine the content of lip reading recognition according to the sequence code.
[0212] Exemplarily, the acquisition module 601 is specifically used to: asynchronously record the brightness change of each pixel of the lip sequence image through an event camera, and trigger an event when the logarithmic intensity change of the pixel exceeds a preset contrast threshold; obtain the events occurring within a preset time window to obtain the original event stream.
[0213] Exemplarily, the conversion module 602 is specifically configured to scale the event timestamps to within the given time interval using a voxel grid method according to the events in the original event stream and the given time interval, thereby obtaining a voxelized event volume.
[0214] Exemplarily, the front-end network includes: a 3D convolution layer and a deep neural network, a feature extraction module 603, specifically used to: perform preliminary spatiotemporal feature extraction processing on the voxelized event body through the 3D convolution layer to obtain initial spatiotemporal features; input the initial spatiotemporal features into the convolution layer and multiple residual blocks of the deep neural network in sequence, and obtain a high-level spatial feature map after the output of multiple residual blocks is processed by global average pooling, and obtain multiple spatial feature maps after the output of multiple residual blocks is processed by shallow adaptive pooling; use the intermediate layer feature map of multiple spatial feature maps as the reference resolution for alignment, and reshape the spatial dimensions of the multi-level features into a node representation constructed by the graph sequence; the node representation sequence is input in parallel into the frequency-aware modulation module for frequency enhancement processing to obtain enhanced features; the frequency enhancement processing includes: injecting noise into the first frequency band, and using adaptive frequency filtering to enhance the discriminative edge dynamics of the second frequency band; the lower limit frequency of the second frequency band is higher than the upper limit frequency of the first frequency band; the enhanced features are input into the fully connected layer of the deep neural network to obtain multi-level spatial node features; the multi-level spatial node features are respectively input into the spatial region hypergraph to extract high-order intra-frame dependencies between lip regions; multi-scale spatial features are obtained based on the high-order intra-frame dependencies between lip regions and the high-level spatial feature map.
[0215] Exemplarily, the time modeling module 604 is specifically used to: convert the multi-scale spatial features into time-structured features through a time semantic hypergraph; and perform modeling based on the time-structured features through a recurrent neural network to obtain sequence coding.
[0216] Exemplarily, the recognition module 605 is specifically configured to: determine the lip-reading content according to the sequence encoding through a target lip-reading model; wherein the target lip-reading model adopts a viseme-based supervision strategy, and the viseme-based supervision strategy is used to encode visual similarity into a label space.
[0217] Exemplarily, the viseme-based supervision strategy includes: associating normalized word labels with phoneme sequences, and converting each phoneme sequence into a corresponding viseme sequence according to a preset phoneme-to-viseme mapping rule; quantifying the similarity between viseme sequences of different words by editing distance; wherein the minimum viseme editing distance between any two different words is 1, and the smaller the viseme editing distance, the more phonetically similar the words are; constructing a distance matrix by calculating paired viseme editing distances; converting the distance matrix into a probability distribution by row-by-row normalization; and establishing soft labels for each word according to the probability distribution.
[0218] Each module in the event-based lip-reading device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0219] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means. The wireless means can be implemented via Wi-Fi, a mobile cellular network, near field communication (NFC), or other technologies. When executed by the processor, the computer program implements an event-based lip reading method. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0220] Those skilled in the art will understand that Figure 7The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0221] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0222] The method collects the original event stream of lip sequence images through an event camera; converts the original event stream into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume; extracts spatial features from the voxelized event volume through a front-end network in an event-based lip reading model to obtain multi-scale spatial features; models the time dependency of the multi-scale spatial features through a back-end sequence model in the event-based lip reading model to obtain a sequence code; and determines the content of lip reading recognition based on the sequence code.
[0223] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0224] The brightness change of each pixel of the lip sequence image is asynchronously recorded by an event camera, and an event is triggered when the logarithmic intensity change of the pixel exceeds a preset contrast threshold; the events occurring within a preset time window are acquired to obtain the original event stream.
[0225] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0226] According to the events in the original event stream and the given time interval, a voxel grid method is used to scale the event timestamps to the given time interval to obtain a voxelized event volume.
[0227] In one embodiment, the front-end network includes: a 3D convolutional layer and a deep neural network, and the processor further implements the following steps when executing the computer program:
[0228] The voxelized event volume is subjected to preliminary spatiotemporal feature extraction processing through the 3D convolution layer to obtain initial spatiotemporal features; the initial spatiotemporal features are sequentially input into the convolution layer and multiple residual blocks of the deep neural network, the outputs of multiple residual blocks are subjected to global average pooling processing to obtain high-level spatial feature maps, and the outputs of multiple residual blocks are subjected to shallow adaptive pooling processing to obtain multiple spatial feature maps; the intermediate layer feature maps of the multiple spatial feature maps are used as the reference resolution for alignment, and the spatial dimensions of the multi-level features are reshaped into a node representation sequence constructed by the graph; the node representation sequence is input in parallel into the frequency perception Frequency enhancement processing is performed in the modulation module to obtain enhanced features; the frequency enhancement processing includes: injecting noise into the first frequency band, and using adaptive frequency filtering to enhance the discriminative edge dynamics of the second frequency band; the lower limit frequency of the second frequency band is higher than the upper limit frequency of the first frequency band; the enhanced features are input into the fully connected layer of the deep neural network to obtain multi-level spatial node features; the multi-level spatial node features are respectively input into the spatial region hypergraph to extract high-order intra-frame dependencies between lip regions; multi-scale spatial features are obtained based on the high-order intra-frame dependencies between lip regions and the high-level spatial feature map.
[0229] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0230] Converting the multi-scale spatial features into temporally structured features through a temporal semantic hypergraph;
[0231] A recurrent neural network is used to model the temporal structured features to obtain sequence encoding.
[0232] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0233] The target lip-reading model determines the content of the lip-reading according to the sequence encoding; wherein the target lip-reading model adopts a viseme-based supervision strategy, and the viseme-based supervision strategy is used to encode visual similarity into a label space.
[0234] In one embodiment, the viseme-based supervision strategy includes: associating normalized word labels with phoneme sequences, and converting each phoneme sequence into a corresponding viseme sequence according to a preset phoneme-to-viseme mapping rule; quantifying the similarity between viseme sequences of different words by editing distance; wherein the minimum viseme editing distance between any two different words is 1, and the smaller the viseme editing distance, the more phonetically similar the words are; constructing a distance matrix by calculating paired viseme editing distances; converting the distance matrix into a probability distribution by row-by-row normalization; and establishing soft labels for each word according to the probability distribution.
[0235] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method steps in each of the above implementations are implemented.
[0236] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method steps in the above embodiments are implemented.
[0237] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0238] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0239] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0240] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An event-based lip reading method, characterized in that: The method comprises: The raw event stream of the lip sequence images is collected by the event camera; The original event stream is converted into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume; Extracting spatial features from the voxelized event volume through a front-end network in an event-based lip reading model to obtain multi-scale spatial features; Modeling the time dependency of the multi-scale spatial features through a back-end sequence model in an event-based lip reading model to obtain a sequence encoding; The content of lip reading recognition is determined according to the sequence code.
2. The method according to claim 1, characterized in that The original event stream of the lip sequence image collected by the event camera includes: Asynchronously recording the brightness change of each pixel of the lip sequence image through an event camera, and triggering an event when the logarithmic intensity change of the pixel exceeds a preset contrast threshold; Events occurring within a preset time window are acquired to obtain the original event stream.
3. The method according to claim 1, characterized in that The voxel-based representation method converts the original event stream into a frame-like event tensor to obtain a voxelized event volume, including: According to the events in the original event stream and the given time interval, a voxel grid method is used to scale the event timestamps to the given time interval to obtain a voxelized event volume.
4. The method according to claim 1, wherein The front-end network includes: a 3D convolutional layer and a deep neural network. The front-end network in the event-based lip reading model extracts spatial features from the voxelized event volume to obtain multi-scale spatial features, including: Performing preliminary spatiotemporal feature extraction processing on the voxelized event volume through the 3D convolution layer to obtain initial spatiotemporal features; The initial spatiotemporal features are sequentially input into the convolutional layer and multiple residual blocks of the deep neural network, the outputs of the multiple residual blocks are subjected to global average pooling to obtain a high-level spatial feature map, and the outputs of the multiple residual blocks are subjected to shallow adaptive pooling to obtain multiple spatial feature maps; Using the intermediate layer feature maps of multiple spatial feature maps as the reference resolution for alignment, the spatial dimensions of multi-level features are reshaped into a sequence of node representations for graph construction. Inputting the node representation sequence in parallel into a frequency-aware modulation module for frequency enhancement processing to obtain enhanced features; the frequency enhancement processing includes: injecting noise into a first frequency band and enhancing discriminative edge dynamics of a second frequency band using adaptive frequency filtering; the lower limit frequency of the second frequency band is higher than the upper limit frequency of the first frequency band; Inputting the enhanced features into the fully connected layer of the deep neural network to obtain multi-level spatial node features; Inputting the multi-level spatial node features into the spatial region hypergraph respectively to extract high-order intra-frame dependencies between lip regions; Multi-scale spatial features are obtained based on the high-order intra-frame dependencies between lip regions and high-level spatial feature maps.
5. The method according to claim 1, wherein The multi-scale spatial features are temporally dependently modeled by a back-end sequence model in an event-based lip reading model to obtain a sequence encoding, including: Converting the multi-scale spatial features into temporally structured features through a temporal semantic hypergraph; A recurrent neural network is used to model the temporal structured features to obtain sequence encoding.
6. The method according to any one of claims 1 to 5, characterized in that The determining of the content of lip reading recognition according to the sequence code includes: The target lip-reading model determines the content of the lip-reading according to the sequence encoding; wherein the target lip-reading model adopts a viseme-based supervision strategy, and the viseme-based supervision strategy is used to encode visual similarity into a label space.
7. The method according to claim 6, characterized in that The viseme-based supervision strategy includes: Associating the normalized word labels with phoneme sequences, and converting each phoneme sequence into a corresponding viseme sequence according to the preset phoneme-to-viseme mapping rules; The similarity between viseme sequences of different words is quantified by the edit distance. The minimum viseme edit distance between any two different words is 1. The smaller the viseme edit distance, the more similar the words are in phonetics. Construct a distance matrix by calculating pairwise viseme edit distances; Converting the distance matrix into a probability distribution by row-wise normalization; Build soft labels for each word based on the probability distribution.
8. An event-based lip-reading device, characterized in that: The device comprises: An acquisition module, used for acquiring the original event stream of lip sequence images through an event camera; A conversion module, configured to convert the original event stream into a frame-like event tensor based on a voxel representation method to obtain a voxelized event volume; a feature extraction module, configured to extract spatial features from the voxelized event volume using a front-end network in an event-based lip reading model to obtain multi-scale spatial features; A temporal modeling module, configured to perform temporal dependency modeling on the multi-scale spatial features using a back-end sequence model in an event-based lip reading model to obtain a sequence code; The recognition module is used to determine the content of lip reading recognition according to the sequence code.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Lip reading method based on adaptive semantic space-time diagram convolutional network
CN111259875A
Lip language recognition method and system, terminal equipment and medium
CN119851350A
Cited By
Three-channel event determination method and device based on action dynamic characteristics, equipment, medium and product
CN121585922A
Three-channel event determination method, device, equipment, medium and product based on motion dynamic characteristics
CN121585922B
Speaker recognition method based on lip movement
CN122049949A