Human body posture recognition method based on multi-modal neural network framework
Through the multimodal neural network framework, the multi-dimensional characteristics of event cameras are used to solve the problem of timing information loss and static area 'blind spot' in the prior art, and high-precision and low-power consumption human posture estimation is achieved.
Patent Information
- Application Number
- CN202510536638.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing human pose estimation method based on event cameras faces difficulties such as loss of timing information, static area 'blind spot problem', and calculation efficiency and accuracy trade-offs, and fails to effectively utilize the multi-dimensional characteristics of event cameras.
A human posture recognition method based on a multimodal neural network framework is proposed. Through technical means such as five-channel event representation, channel splitting and three-modal separation, multimodal feature extraction, dynamic feature enhancement and cross-modal fusion, framework constraints and space-time processing, the spatial, timing and structural information of event data are fully utilized.
It significantly improves the accuracy of pose estimation, solves the static area 'blind spot problem', enhances the robustness and dynamic adaptability of the model, reduces jitter and incoherence problems, and realizes high-precision and low-power consumption of human pose estimation.
Smart Images

Figure CN120048006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and neural network processing, and particularly relates to a human pose recognition method based on a multi-modal neural network framework. Background Art
[0002] An event camera is a bionic vision sensor. Different from traditional cameras that capture full-frame images at a fixed frequency, an event camera only asynchronously generates events when the pixel brightness changes exceed a preset threshold. Each event can be represented as a quadruple e=(x, y, t, p), where (x, y) is the pixel position, t is the timestamp in microseconds, and p is the polarity (indicating an increase or decrease in brightness). This working mode endows the event camera with advantages such as microsecond-level time resolution, high dynamic range (120dB vs 60dB), and low power consumption, and is particularly suitable for processing high-speed motion and complex lighting scenarios.
[0003] However, the sparse asynchronous characteristics of event data also bring processing challenges. Each independent event carries very little scene information and must be aggregated before it can be used for downstream tasks. Currently, event data processing methods are mainly divided into two categories: sparse representation and dense representation. Sparse representation preserves the original characteristics of event data, but is limited by the lack of dedicated hardware and efficient algorithms; dense representation converts events into a format compatible with traditional computer vision, facilitating the application of mature deep learning technologies.
[0004] In the field of human pose estimation, event cameras have received increasing attention due to their advantages. Some researchers have proposed the DHP19 dataset, which uses multiple synchronized event cameras to record human actions, estimates 2D key points from event frames through a CNN, and then obtains 3D poses through triangulation; some have proposed the Temporally Ordered Recent Events Volume (TORE) representation method, which improves the pose estimation accuracy by storing the original event temporal information; others have proposed LiftMonoHPE, which estimates 3D poses through a monocular event camera, first predicting 2D key points and then recovering 3D information through a deep learning model. Some scholars have also tried to process events as point cloud data and proposed the EventPointPose model, but have not fully utilized the temporal characteristics of events.
[0005] In terms of the combination of event data and neural networks, some researchers have proposed a hybrid ANN-SNN architecture, which initializes the state of a high-frequency running SNN through a low-frequency running ANN, solving the problems of slow SNN convergence and state decay. Others have designed a hybrid network combining SNN and ANN for event-driven optical flow estimation, where the SNN processes event encoding and the ANN is responsible for feature extraction and prediction.
[0006] However, existing event camera-based human pose estimation methods face multiple technical challenges, severely limiting the effectiveness of event cameras in practical application scenarios. Mainstream methods accumulate events into two-dimensional frame representations, such as event frames, temporal surfaces, or voxel grids. This processing inevitably loses the precise temporal information contained in event data, especially having a significant impact when dealing with fast and complex movements. At the same time, event cameras only generate events when there is a change in brightness, which leads to the "blind spot problem" in static regions: when a part of the human body remains stationary, no events are generated in that region and it disappears in the representation, resulting in incomplete pose estimation. In addition, existing methods often use a fixed time window to complete event stacking. When the event rate is too high or too low, it will lead to information redundancy or loss. Methods that directly process the event stream, such as spiking neural networks, can retain temporal characteristics but suffer from problems such as slow convergence and state decay, and have low computational efficiency. Methods that convert events into dense representations have higher computational efficiency but limited accuracy and cannot fully utilize the advantageous characteristics of event cameras. More importantly, most existing methods mix and process different characteristics of events (such as spatial position, timestamp, polarity) without effectively distinguishing and utilizing this heterogeneous information, resulting in low feature extraction efficiency. At the same time, most methods fail to effectively utilize the prior knowledge of the human skeleton and lack necessary skeleton structure constraints, making pose prediction in complex scenarios often not conform to human anatomical laws. In summary, the main limitations of existing methods can be roughly summarized as follows: First, methods that accumulate events into two-dimensional frame representations inevitably lose precise temporal information; second, methods that directly use event stream processing methods such as spiking neural networks (SNNs) retain temporal characteristics but are difficult to train, have limited accuracy, and face problems of slow convergence and state decay; third, there is a "blind spot problem" in static limb scenarios, that is, no events are generated in static regions and they disappear in the representation.
[0007] Currently, there are also some relevant solutions for event data stream representation methods at home and abroad. The literature "Exploring Event-based Human Pose Estimation with 3D Event Representations" published by author Jing Zhao in the journal Conference on Computer Vision and Pattern Recognition (CVPR) in 2021 proposed a voxel-based multi-scale Transformer network (VMST-Net), which improves information extraction efficiency by constructing event voxels in 3D space. This method attempts to capture both the spatial and temporal information of events, but has a high computational complexity and still does not solve the problem of missing information in static regions, and there are still stability problems in fast-changing movements and complex environments.
[0008] In summary, although existing methods have made progress in the field of event-driven human pose estimation, there is still a lack of a comprehensive solution that can effectively handle the spatio-temporal characteristics of events, solve the problem of missing information in static regions, and maintain a low computational complexity. Summary of the Invention
[0009] To solve the above problems, the present invention provides a human pose recognition method based on a multi-modal neural network framework.
[0010] The object of the present invention is to provide a human pose recognition method based on a multi-modal neural network framework, which specifically includes the following steps: S1. Data preprocessing: Collect event data and represent the event data using a five-channel event representation method; the five-channel event representation method is as follows: E' = (x, y, T, P, E); where x and y are pixel positions, T is the average timestamp, P is the cumulative positivity, and E is the event count; S2. Channel splitting and three-modal separation: Divide the five-channel event data preprocessed in step S1 into three channels, which are as follows: a spatial channel, including the normalized spatial coordinates (x, y); an event characteristic channel, including time and polarity features (T, P, E); a point cloud channel: regard the complete five-channel data as point cloud data and retain the original three-dimensional characteristics of the event; S3. Multi-modal feature extraction: After completing the channel splitting, perform feature extraction on the data of the spatial channel, event characteristic channel, and point cloud channel respectively; S4. Dynamic feature enhancement and cross-modal fusion: After feature extraction, optimize and integrate the features of different modalities through dynamic feature enhancement and cross-modal fusion; S5. Skeleton constraint and spatio-temporal processing: After feature fusion, adopt a skeleton constraint module and a spatio-temporal joint processing module to enhance the network's understanding of the human body structure and spatio-temporal dependence relationship; S6. Key point prediction and training: Adopt a prediction head module to output the final human key point coordinates; S7. Use the preprocessed data in step S1 to repeat steps S2~S6 for training. The training method includes a three-level loss function and uses the alternative gradient method for backpropagation to achieve end-to-end training, and obtain the pose recognition result.
[0011] Preferably, step S3 specifically includes: S301. Process the data of the spatial channel using a multi-gradient feature extraction module; S302. Process the data of the event characteristic channel using a pulse feature processor; S303. Use a hierarchical point cloud encoding module to perform point-level feature extraction, local region feature extraction, and global feature extraction on the point cloud data.
[0012] Preferably, the multi-gradient feature extraction module in step S301 includes four parallel branches, namely a small-scale feature extraction convolution kernel, a medium-scale feature extraction convolution kernel, a large-scale feature extraction convolution kernel, and an extra-large-scale feature extraction convolution kernel, and the sizes of the convolution kernels are 1×1, 3×3, 5×5, and 7×7 respectively; the outputs of each branch are concatenated and fused after being processed by batch normalization and activation functions.
[0013] Preferably, the point-level feature extraction in step S303 independently processes each point through a multi-layer perceptron; the local region feature extraction captures the relationship between a point and its k nearest neighbors; the global feature extraction obtains the overall structure information through a max pooling operation.
[0014] Preferably, the pulse feature processor in step S302 includes an adaptive LIF neuron or an Izhikevich neuron; the dynamic equation of the adaptive LIF neuron is as follows:
[0015] Membrane potential update (discretization):
[0016] Threshold update:
[0017] Pulse generation:
[0018] Among them, spike is the action potential fired by the neuron, v is the membrane potential, θ is the adaptive threshold, τ mem is the membrane potential time constant, τ thresh is the threshold time constant, θ rest is the resting threshold, I is the input current, v t represents the membrane potential at time t, θ t represents the adaptive threshold at time t, I t represents the input current at time t, and Heaviside is the step function; The dynamic equation of the Izhikevich neuron is:
[0019] Pulse triggering condition: If v≥threshold, then spike = 1, otherwise spike = 0; Reset after pulse: If spike = 1, then v = c, u = u + d; Among them, v is the membrane potential, u is the membrane potential recovery variable, I is the input current, and a, b, c, and d are parameters that control the behavior of neurons.
[0020] Preferably, the pulse feature processor in step S302 further includes a time gating unit and a bidirectional LSTM layer to further extract temporal features and long-term dependencies; the specific operations are as follows: perform convolutional preprocessing on the data of the event feature channels (T, P, E) through a convolutional pulse encoder; use an adaptive LIF neuron model to process the convolved features to enhance the processing ability of temporal information; perform temporal modeling on the features after adaptive processing through a bidirectional LSTM layer to capture the forward and backward dependencies of the time series; further enhance the processing ability of time information through the time gating unit to control the flow of information.
[0021] Preferably, step S4 specifically includes: S401. Dynamic feature enhancement: The dynamic feature enhancement module adaptively enhances the event features and spatial features respectively according to the statistical information of the input features through channel attention and spatial attention mechanisms to enhance key features; S402. Cross-modal fusion: Use a cross-modal adaptive fusion module to perform cross-modal fusion on the features.
[0022] Preferably, the mathematical representation of the channel attention mechanism in step S401 is as follows: ; The mathematical representation of the spatial attention mechanism is as follows:
[0023] Among them, F is the input feature map, AvgPool is global average pooling, MaxPool is global maximum pooling, W 1 and W 2 are weight matrices, δ is the ReLU activation function, σ is the Sigmoid activation function, and [;] represents concatenation in the channel dimension; The cross-modal adaptive fusion module in step S402 realizes adaptive fusion through the following steps: S4021. Modal-specific encoding: Process the spatial features and event features respectively through dedicated encoders to enhance their respective representation capabilities; S4022. Adaptive weight generation: Generate dynamic weights through global pooling and a multi-layer perceptron to achieve dynamic weighting between modalities; the weight calculation formula is as follows: ; S4023. Weighted fusion: Weight the features of different modalities according to the generated weights; ; S4024. Cross-modal attention interaction: Achieve deep interaction between modalities through the multi-head self-attention mechanism; ; Among them, F weighted represents the feature fusion result after weight weighting, F s and F e are spatial features and event features respectively, w = [w 1 , w 2 are the weights of the two modalities, Q, K, V are query, key, and value matrices, is the scaling factor.
[0024] Preferably, step S5 specifically includes the following operations: S501. Skeleton constraint module: Combine the spiking neural network with graph convolution, and introduce the spiking neuron mechanism in the standard graph convolution operation. The process is as follows: First, perform the standard graph convolution operation: ; Input the obtained result Z into the spiking neuron: ; Among them, X is the input feature, W is the weight matrix, b is the bias, Ã = A + I is the adjacency matrix with self-connection added, D̃ is the degree matrix of Ã, and SpikeNeuron( ) is the spiking activation function; To enhance the rationality of the human body structure, design the bone length consistency loss function: ; Among them, bone is the line segment connecting two key point joints in the human body skeleton model, bones are all the connecting line segments that make up the complete human body skeleton, length(bone) is the length of the bone, and std represents the standard deviation; S502. Adopt a spatio-temporal joint processing module to capture the spatio-temporal dependence relationship in the event data; the spatio-temporal joint processing module includes multi-scale temporal encoding and multi-head spatio-temporal attention; The multi-scale temporal encoding adopts a hybrid feature extraction method that combines a parallel convolution structure with frequency domain analysis; the parallel convolution structure includes convolution kernels of different sizes for processing the time dimension; the frequency domain analysis realizes frequency domain feature extraction through the fast Fourier transform; the expression is as follows:
[0025] Among them, F is the input feature, F time is the time domain feature, F freq is the frequency domain feature, F out is the output feature, Conv kDenotes a convolution with a kernel size of k, [;] is concatenation in the channel dimension, W fusion is the fused convolution, FFT is the Fast Fourier Transform, LN is layer normalization, and g is a learnable gating parameter; The multi-head spatio-temporal attention establishes long-range dependencies between different spatio-temporal positions through the self-attention mechanism, and the formula is as follows:
[0026] Among them, F attn is the attention output, G spatial is the spatial gating, G temporal is the temporal gating, and ⊙ is the element-wise multiplication.
[0027] Preferably, the prediction formula of the prediction head module in step S6 is as follows:
[0028] Among them, x pred 、y pred are the predicted joint coordinates, conf is the confidence, σ is the Sigmoid activation function, F global represents the global feature, Proj x 、 Proj y 、 Proj conf are the coordinate and confidence predictors respectively, and W and H are the width and height of the output feature map resolution; The loss function is expressed as follows: ; Among them, L pose is the key point position loss, which calculates the difference between the predicted key points and the true key points using the mean square error; λ bone represents the weight coefficient of the skeleton constraint loss, L bone is the skeleton constraint loss, which ensures that the predicted bone length conforms to the physical constraints; λ temp represents the weight coefficient of the temporal consistency loss, L temp is the temporal consistency loss, which encourages smooth transitions between consecutive frames.
[0029] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) Advantages of multi-dimensional information utilization: Through the channel splitting strategy and specially designed processors, the comprehensive utilization of spatial, temporal, and structural information is achieved, significantly improving the accuracy of pose estimation. (2) Static region representation ability: The dedicated pulse feature processor of the present invention effectively solves the "blind spot problem" existing in traditional event processing methods in static regions by maintaining the membrane potential state of neurons. (3) Robustness improvement: The skeleton constraint module of the present invention significantly improves the robustness of the model in scenarios with partial occlusion, complex backgrounds, and multiple people by introducing prior knowledge of human anatomy. (4) Dynamic adaptability: The cross-modal adaptive fusion module of the present invention can dynamically adjust the weights of different modalities according to the characteristics of the input data, enabling the network to adapt to various scenarios and motion patterns; this design performs excellently when dealing with rapidly changing speed and lighting conditions, improving stability compared to the fixed weight method. (5) Temporal consistency: Through multi-scale temporal encoding and spatio-temporal attention mechanisms, the temporal dependence relationships in continuous action sequences can be better captured, reducing jitter and incoherence problems.
[0030] In summary, the present invention proposes a novel hybrid pulse-point cloud neural network architecture that separates event data into spatial channels, event feature channels, and point cloud channels through an innovative channel splitting strategy. This method combines a multi-gradient feature extraction module, a dedicated pulse feature processor, and a hierarchical point cloud encoder to effectively capture multi-scale spatial features and temporal dynamic features. The system also introduces a cross-modal adaptive fusion mechanism to dynamically adjust the weights of different modalities, as well as a skeleton constraint module to constrain the prediction results using prior knowledge of human anatomy. Compared with the prior art, the present invention can more effectively process the sparse asynchronous characteristics of event data, effectively solve the "blind spot problem" of static limbs, and achieve high-precision human pose estimation while maintaining low power consumption, especially suitable for resource-constrained edge computing environments. Brief Description of the Drawings
[0031] Figure 1 It is a schematic diagram of an event camera.
[0032] Figure 2 It is a flowchart of a human pose recognition method based on a multi-modal neural network framework provided by an embodiment of the present invention.
[0033] Figure 3 It is an overall network structure diagram of a multi-modal neural network framework provided by an embodiment of the present invention.
[0034] Figure 4 It is a schematic diagram of five-channel representation and channel splitting provided by an embodiment of the present invention.
[0035] Figure 5 It is a schematic diagram of a multi-gradient feature extraction module, a pulse feature processor, and a hierarchical point cloud encoding module for multi-modal feature extraction according to an embodiment of the present invention. Detailed implementation manners
[0036] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the following description, the same modules are denoted by the same reference numerals. In the case of the same reference numerals, their names and functions are also the same. Therefore, their detailed descriptions will not be repeated.
[0037] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0038] The novel neural network architecture proposed by the present invention aims to solve the above challenges through an innovative channel splitting strategy and a multi-modal processing method, providing a more efficient and accurate solution for event-driven human pose estimation.
[0039] The present invention provides a human pose recognition method based on a multi-modal neural network framework, specifically including the following steps: S1. Data preprocessing: Collect event data and represent the event data using a five-channel event representation method; specifically including: Obtain the original event stream from the event camera, and each event is represented as a quadruple: e=(x, y, t, p), where: (x, y) is the pixel position, t is the timestamp in microseconds, and p is the polarity value (0 or 1); the schematic diagram of the event camera is as Figure 1 shown; Represent the event data using the five-channel event representation method, and the five-channel event representation method is as follows: E'=(x, y, T, P, E); where, x and y are the pixel positions, T is the average timestamp, P is the cumulative polarity, and E is the event count.
[0040] This step provides a unified data format for subsequent channel splitting and feature extraction; based on the characteristics of the event camera, an innovative five-channel event representation method is designed, and this representation method has the following advantages: effectively separating spatial information and event characteristic information; retaining the key statistical characteristics of the events; significantly reducing the data volume and improving the processing efficiency.
[0041] S2. Channel splitting and three-modal separation: Divide the five-channel event data preprocessed in step S1 into three channels, which are as follows: Spatial channel: Contains the normalized spatial coordinates (x, y), providing information about the position of the human body on the image plane; Event Feature Channel: It contains time and polarity features (T, P, E), providing information about the intensity and direction of actions; Point Cloud Channel: Regarding the complete five-channel event data as point cloud data, it preserves the original three-dimensional characteristics of the events.
[0042] In this step, the data is divided into three channels, providing clear data input for the subsequent feature extraction module, enabling the network to adopt specially designed processing modules for different types of information, and thus more effectively extracting and utilizing the multi-dimensional features contained in the event data.
[0043] S3. Multi-modal Feature Extraction: After completing the channel splitting, feature extraction is performed on the data of the spatial channel, event feature channel, and point cloud channel respectively; specifically including: S301. Processing the data of the spatial channel (x, y) using a multi-gradient feature extraction module; Specifically, the multi-gradient feature extraction module captures spatial features in different ranges through multi-scale parallel convolutional branches; the multi-gradient feature extraction module contains four parallel branches, namely a small-scale feature extraction convolutional kernel, a medium-scale feature extraction convolutional kernel, a large-scale feature extraction convolutional kernel, and an extra-large-scale feature extraction convolutional kernel, with the convolutional kernel sizes being 1×1, 3×3, 5×5, and 7×7 respectively; the outputs of each branch are concatenated and fused after being processed by batch normalization and activation functions; This design (feature pyramid multi-scale receptive field module) enables the network to simultaneously focus on local details and global structures, and is particularly suitable for processing the complex spatial relationships between human body bone joints.
[0044] S302. Processing the data of the event feature channel (T, P, E) using a spike feature processor; the spike feature processor includes an adaptive LIF (Leaky Integrate-and-Fire) neuron or an Izhikevich neuron; the specific operations are as follows: Performing convolutional preprocessing on the data of the event feature channel (T, P, E); specifically, the convolutional pulse encoder extracts signal features through convolutional operations; Processing the convolved features using an adaptive LIF neuron model to enhance the processing ability for temporal information; specifically, the adaptive LIF neuron dynamically adjusts its own threshold or other parameters according to the characteristics of the input signal to better adapt to different input patterns and capture the time dependence and dynamic characteristics of the signal; Performing temporal modeling on the features after adaptive processing through a bidirectional long short-term memory LSTM network (Bi-LSTM) to capture the forward and backward dependencies of the time series; further enhancing the processing ability for time information through a time gating mechanism to control the flow of information and avoid problems such as gradient disappearance or explosion.
[0045] Specifically, the dynamic equation of the adaptive LIF neuron is as follows:
[0046] Membrane potential update (discretization):
[0047] Threshold update:
[0048] Spike generation:
[0049] Among them, spike refers to the action potential generated by the neuron and is the basic signal unit in the spiking neural network (SNN); v is the membrane potential, θ is the adaptive threshold, τ mem is the membrane potential time constant, τ thresh is the threshold time constant (determining the speed of threshold decay), θ rest is the resting threshold, I is the input current, v t represents the membrane potential at time t, θ t represents the adaptive threshold at time t, I t represents the input current at time t, and Heaviside is the step function.
[0050] Specifically, the dynamic equation of the Izhikevich neuron is:
[0051] Spike triggering condition: If v ≥ threshold, then spike = 1, otherwise spike = 0; Reset after spike: If spike = 1, then v = c, u = u + d; Among them, v is the membrane potential, u is the membrane potential recovery variable, I is the input current, and a, b, c, and d are parameters controlling the neuron behavior.
[0052] In addition, the pulse feature processor also includes a time gating unit and a bidirectional LSTM layer to further extract temporal features and long-term dependencies; this design effectively solves the "blind spot problem" in the static region of traditional methods because even if no new events occur in a certain region, the membrane potential state of the spiking neuron will remain for a period of time, enabling the network to remember the features of the static region.
[0053] S303. Use a hierarchical point cloud encoding module to perform point-level feature extraction, local region feature extraction, and global feature extraction on the point cloud data; The point-level feature extraction processes each point independently through a multi-layer perceptron (MLP); the local region feature extraction captures the relationship between a point and its k nearest neighbors; the global feature extraction obtains the overall structural information through operations such as max pooling; The hierarchical point cloud encoding module regards the event data as a point cloud in the spatio-temporal domain and captures the local and global features of the point cloud through multi-level abstraction. This module processes the point cloud data in three stages: point-level feature extraction, local region feature extraction, and global feature extraction. This hierarchical design enables the network to comprehensively understand the structure of the event point cloud and provides rich three-dimensional information for human pose estimation.
[0054] S4. Dynamic Feature Enhancement and Cross-modal Fusion: After feature extraction, dynamic feature enhancement is performed through a dynamic feature enhancement module, and cross-modal fusion of features is carried out through a cross-modal adaptive fusion module to optimize and integrate features of different modalities; specifically including: S401. Dynamic Feature Enhancement: The dynamic feature enhancement module adaptively enhances the event features and spatial features respectively based on the statistical information of the input features through channel attention and spatial attention mechanisms to enhance the key features; The mathematical representation of the channel attention mechanism is as follows: ; The mathematical representation of the spatial attention mechanism is as follows:
[0055] where F is the input feature map, AvgPool is global average pooling, MaxPool is global max pooling, W 1 and W 2 are weight matrices, δ is the ReLU activation function, σ is the Sigmoid activation function, and [;] represents concatenation on the channel dimension.
[0056] This dynamic feature enhancement mechanism can significantly improve the network's attention to key features and enhance the representation ability.
[0057] S402. Cross-modal Fusion: A cross-modal adaptive fusion module is used to perform cross-modal fusion of features; this module realizes adaptive fusion through the following steps: S4021. Modality-specific Encoding: The spatial features and event features are respectively processed through dedicated encoders to enhance their respective representation abilities; S4022. Adaptive Weight Generation: Dynamic weights are generated through global pooling and a multi-layer perceptron to achieve dynamic weighting between modalities; the weight calculation formula is as follows: ; S4023. Weighted Fusion: Weight the features of different modalities according to the generated weights; the expression is as follows: ; S4024. Cross-modal Attention Interaction: Achieve deep interaction between modalities through the multi-head self-attention mechanism; ; Among them, F weighted represents the feature fusion result after weight weighting, F s and F e are spatial features and event features respectively, w = [w 1 , w 2 are the weights of the two modalities, Q, K, V are the query, key, and value matrices, is the scaling factor.
[0058] The cross-modal adaptive fusion module is one of the core innovations of the present invention, which is used to integrate features from different modalities; this adaptive fusion mechanism enables the network to dynamically adjust the attention to different modalities according to the characteristics of the input data, so as to adapt to various scenarios and motion patterns.
[0059] S5. Skeleton Constraint and Spatiotemporal Processing: After feature fusion, adopt the skeleton constraint module and the spatiotemporal joint processing module to enhance the network's understanding of the human body structure and the modeling of spatiotemporal dependence relationships; the specific operations include the following: S501. Skeleton Constraint Module: The skeleton constraint module represents the human skeleton as a graph structure and learns the dependence relationships between joints through the graph convolutional network; the mathematical representation of the standard graph convolution is as follows:
[0060] Among them, X' is the output feature, X is the input feature, W is the weight matrix, b is the bias, à = A + I is the adjacency matrix with self-connection added, and the degree matrix of à is D̃; In this step, the spiking neural network is combined with the graph convolution to design a spiking graph convolutional layer, and the spiking neuron mechanism is introduced into the standard graph convolution operation. The process is as follows: First, perform the standard graph convolution operation: ; Input the obtained result Z into the spiking neuron: ; Among them, X is the input feature, W is the weight matrix, b is the bias, à = A + I is the adjacency matrix with self-connections added, D̃ is the degree matrix of Ã, and SpikeNeuron( ) is the spiking activation function; To enhance the rationality of the human body structure, a bone length consistency loss function is designed: ; Among them, bone refers to the line segment connecting two key points (joints) in the human body skeleton model, bones refers to the set of all bones, that is, all the connecting line segments that make up the complete human body skeleton, length(bone) is the length of the bone, and std represents the standard deviation.
[0061] The skeleton constraint module also introduces bone consistency loss. By calculating the within-batch standard deviation of the predicted bone lengths, it encourages the consistency of bone lengths: This module significantly improves the robustness of the network in complex scenarios by introducing prior knowledge of human anatomy, especially performing excellently when dealing with partially occluded or blurred scenarios.
[0062] S502. Spatiotemporal joint processing module: It includes two key components, multi-scale temporal encoding and multi-head spatiotemporal attention, which are used to capture the spatiotemporal dependencies in event data; The multi-scale temporal encoding processes the time dimension using convolutional kernels of different sizes, and at the same time combines frequency domain feature extraction (implemented through FFT), that is, a hybrid feature extraction method combining a parallel convolutional structure and frequency domain analysis. The mathematical expression is as follows:
[0063] Among them, F is the input feature, F time is the time domain feature, F freq is the frequency domain feature, F out is the output feature, Conv k represents convolution with a kernel size of k, [;] is concatenation in the channel dimension, W fusion is the fusion convolution, FFT is the fast Fourier transform, LN is layer normalization, and g is a learnable gating parameter.
[0064] The multi-head spatiotemporal attention establishes long-range dependencies between different spatiotemporal positions through the self-attention mechanism. The formula is as follows:
[0065] Among them, F attn is the attention output, G spatial is the spatial gating, G temporal is the time gating, and ⊙ is the element-wise multiplication.
[0066] The above modules work together to enable the network to effectively capture the spatio-temporal dependencies in event data and improve the ability to understand complex action sequences.
[0067] S6. Key point prediction and training: Use a prediction head module to output the final human key point coordinates; the prediction head module includes separate X-axis and Y-axis prediction branches to predict the coordinates of key points in two directions respectively; the prediction formula is as follows:
[0068] where x pred , y pred are the predicted joint coordinates, conf is the confidence, σ is the Sigmoid activation function, F global represents the global feature, Proj x , Proj y , Proj conf are the horizontal and vertical coordinate and confidence predictors respectively, and W and H are the width and height of the output feature map resolution; S7. Use the preprocessed data in step S1 and repeat steps S2 - S6 for training. The training method includes a three-level loss function and uses the surrogate gradient method for backpropagation to achieve end-to-end training and obtain the pose recognition result; Specifically, the training of the multi-modal neural network (UltraHybridSpikeNet) adopts a multi-task learning framework, and the training method includes the collaborative optimization of a three-level loss function. The loss function is expressed as follows: ; where L pose is the key point position loss, and the mean square error (MSE) is used to calculate the difference between the predicted key points and the true key points; λ bone represents the weight coefficient of the skeleton constraint loss (controlling the importance of the bone length consistency loss in the total loss function), L bone is the skeleton constraint loss, ensuring that the predicted bone lengths conform to physical constraints; λ temp represents the weight coefficient of the temporal consistency loss (controlling the importance of the smooth transition constraint between consecutive frames in the total loss function), L temp is the temporal consistency loss, encouraging smooth transitions between consecutive frames; For the training of the SNN part, due to the non-differentiability of the impulse function, the surrogate gradient method is used for backpropagation: ; This enables the entire network to be trained end-to-end through standard backpropagation.
[0069] The flowchart is shown in Figure 2 , and the structural diagram of the multi-modal neural network framework is as shown in Figure 3 .
[0070] Through the above steps, the spatio-temporal characteristics of event camera data are fully utilized to achieve high-precision human pose estimation while maintaining low power consumption, effectively solving the problems of temporal information loss, "blind spot problem" in static regions, and trade-off between computational efficiency and accuracy in existing methods. Especially when dealing with fast movements and complex scenes, the present invention demonstrates significant advantages and provides a feasible solution for real-time human pose estimation in resource-constrained environments. The key technical points of the present invention are as follows: (1) A novel channel splitting strategy is proposed, which divides event data into spatial channels (x, y), event feature channels (T, P, E), and point cloud channels, and three dedicated feature extractors are designed to process different types of features respectively, maximizing the utilization of multi-dimensional information of events.
[0071] (2) A multi-gradient feature extraction module and a hierarchical point cloud encoding module are designed. Through multi-scale parallel convolution and hierarchical point cloud processing, spatial multi-scale features and point cloud structure characteristics are effectively captured, significantly enhancing the spatial representation ability.
[0072] (3) A dedicated spiking feature processor is developed, which combines advanced neuron models (adaptive LIF or Izhikevich) and a time gating mechanism to enhance the processing ability of event temporal characteristics and effectively solve the "blind spot problem" in static regions.
[0073] (4) A cross-modal adaptive fusion module is introduced. Through dynamic weight generation and a multi-head attention mechanism, adaptive fusion and complex interaction of spatial features and event features are achieved, enabling the network to dynamically adjust the processing strategy according to the characteristics of input data.
[0074] (5) A skeleton constraint module based on spiking graph convolution is designed, which integrates prior knowledge of human anatomy into the network and ensures the rationality of predicted poses through bone consistency loss, improving the robustness of the model in complex scenes.
[0075] (6) A multi-scale temporal encoding and a multi-head spatio-temporal attention mechanism are proposed. By capturing features at multiple time scales and establishing spatio-temporal dependencies, the network's ability to understand complex action sequences is enhanced.
[0076] (7) For the DHP19 dataset, an innovative five-channel input processing method (x, y, T, P, E) is proposed, which effectively distinguishes spatial and event features and provides richer spatio-temporal information for the model.
[0077] The advantages are as follows: (1) Advantage in multi-dimensional information utilization: Existing methods such as EventPointPose and LiftMonoHPE mainly focus on spatial features or simple temporal features, while ignoring the multi-dimensional characteristics of event data. Through the channel splitting strategy and specially designed processors, the present invention realizes the comprehensive utilization of spatial, temporal, and structural information, significantly improving the accuracy of pose estimation.
[0078] (2) Ability to represent static regions: Traditional event processing methods have a "blind spot problem" in static regions, that is, static regions do not generate events and disappear in the representation. The dedicated pulse feature processor of the present invention effectively solves this problem by maintaining the membrane potential state of neurons.
[0079] (3) Improvement in robustness: The skeleton constraint module of the present invention significantly improves the robustness of the model in scenarios with partial occlusion, complex backgrounds, and multiple people by introducing prior knowledge of human anatomy.
[0080] (4) Dynamic adaptability: The cross-modal adaptive fusion module of the present invention can dynamically adjust the weights of different modalities according to the characteristics of the input data, enabling the network to adapt to various scenarios and motion patterns. This design performs excellently when dealing with rapidly changing speed and lighting conditions, improving the stability compared to the fixed-weight method.
[0081] (5) Temporal consistency: Through multi-scale temporal encoding and spatio-temporal attention mechanisms, the present invention can better capture the temporal dependence relationships in continuous action sequences, reducing jitter and incoherence problems.
[0082] By means of an innovative hybrid pulse-point cloud neural network architecture, the key problems in existing event-driven human pose estimation methods are solved, achieving high accuracy and high robustness while maintaining low power consumption, providing a feasible solution for real-time human pose estimation in resource-constrained environments and having broad application prospects.
[0083] In summary, the present invention proposes a novel neural network architecture aimed at achieving efficient and accurate event-driven human pose estimation through innovative design. First, different from the existing methods that mix and process event characteristics, the present invention innovatively proposes a three-channel separation strategy, which divides event features into a spatial channel, an event characteristic channel, and a point cloud channel, and designs a dedicated processor to extract features of different dimensions specifically, maximizing the utilization of the multi-dimensional information of events. Second, the present invention uses a multi-gradient feature extraction module to process the spatial channel, effectively capturing spatial dependency relationships through multi-scale receptive fields; designs a dedicated spiking feature processor to process the event characteristic channel, combining an advanced neuron model and a time gating mechanism to efficiently process the event temporal characteristics and solve the "blind spot problem" in static regions; processes the point cloud representation through a hierarchical point cloud encoder to capture the three-dimensional structural characteristics of events. Third, the present invention develops a cross-modal adaptive fusion module, which realizes the intelligent integration of features of different modalities through dynamic weight generation and an attention mechanism, enabling the network to adaptively adjust the processing strategy according to the characteristics of the input data and enhancing the ability to understand complex scenes. Finally, the present invention designs a skeleton constraint module based on spiking graph convolution, integrating prior knowledge of human anatomy into the network to ensure the biophysical rationality of the predicted pose and improve the robustness under occlusion and complex backgrounds. In summary, the present invention aims to provide a hybrid architecture that takes into account both computational efficiency and accuracy, providing an efficient and feasible solution for real-time human pose estimation on edge devices.
[0084] It should be understood that the various forms of the processes shown above can be reordered, added, or deleted. For example, the steps recited in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitations are imposed herein.
[0085] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A human body posture recognition method based on a multimodal neural network framework, characterized in that: The specific steps include: S1. Data preprocessing: Collect event data and represent the event data using a five-channel event representation method; the five-channel event representation method is as follows: E'= (x, y, T, P, E); Where x and y are pixel positions, T is the average timestamp, P is the cumulative positive, and E is the event count; S2. Channel splitting and trimodal separation: The five-channel event data preprocessed in step S1 is divided into three channels, as follows: spatial channel, containing normalized spatial coordinates (x, y); event feature channel, containing time and polarity features (T, P, E); point cloud channel: the complete five-channel data is regarded as point cloud data, retaining the original three-dimensional characteristics of the event; S3. Multimodal feature extraction: After completing channel splitting, feature extraction is performed on the data of the spatial channel, event characteristic channel, and point cloud channel respectively; S4. Dynamic feature enhancement and cross-modal fusion: After feature extraction, the features of different modalities are optimized and integrated through dynamic feature enhancement and cross-modal fusion; S5. Skeleton constraint and spatiotemporal processing: After feature fusion, the skeleton constraint module and spatiotemporal joint processing module are used to enhance the network's understanding of human body structure and spatiotemporal dependencies; S6. Key point prediction and training: Use the prediction head module to output the final coordinates of the key points of the human body; S7. Using the preprocessed data of step S1, repeat steps S2 to S6 for training. The training method includes a three-level loss function, and uses the alternative gradient method for back propagation to achieve end-to-end training to obtain the posture recognition result.
2. A human body posture recognition method based on a multimodal neural network framework according to claim 1, characterized in that: The step S3 specifically includes: S301. Processing spatial channel data using a multi-gradient feature extraction module; S302. Processing event characteristic channel data using a pulse feature processor; S303. Use a hierarchical point cloud encoding module to perform point-level feature extraction, local area feature extraction and global feature extraction on the point cloud data.
3. A human body posture recognition method based on a multimodal neural network framework according to claim 2, characterized in that: The multi-gradient feature extraction module in step S301 includes four parallel branches, namely, a small-scale feature extraction convolution kernel, a medium-scale feature extraction convolution kernel, a large-scale feature extraction convolution kernel and an ultra-large-scale feature extraction convolution kernel, and the sizes of the convolution kernels are 1×1, 3×3, 5×5 and 7×7 respectively; the outputs of each branch are spliced and fused after batch normalization and activation function processing.
4. The method for human posture recognition based on a multimodal neural network framework according to claim 2, characterized in that: The point-level feature extraction in step S303 processes each point independently through a multi-layer perceptron; the local area feature extraction captures the relationship between a point and its k nearest neighbors; and the global feature extraction obtains the overall structural information through a maximum pooling operation.
5. A human posture recognition method based on a multimodal neural network framework according to any one of claims 2 to 4, characterized in that: The pulse feature processor in step S302 includes an adaptive LIF neuron or an Izhikevich neuron; the dynamic equation of the adaptive LIF neuron is as follows: Membrane potential update: Threshold Update: Pulse Generation: Where spike is the action potential fired by the neuron, v is the membrane potential, θ is the adaptive threshold, and τ mem is the membrane potential time constant, τ thresh is the threshold time constant, θ rest is the resting threshold, I is the input current, v t represents the membrane potential at time t, θ t represents the adaptive threshold at time t, I t represents the input current at time t, Heaviside is a step function; The dynamic equation of the Izhikevich neuron is: Pulse trigger condition: if v≥threshold, then spike=1, otherwise spike=0; reset after pulse: if spike=1, then v=c, u=u+d; Among them, v is the membrane potential, u is the membrane potential recovery variable, I is the input current, and a, b, c, and d are parameters that control the behavior of neurons.
6. The method for human posture recognition based on a multimodal neural network framework according to claim 5, characterized in that: The pulse feature processor in step S302 also includes a time gating unit and a bidirectional LSTM layer to further extract timing features and long-term dependencies; the specific operations are as follows: convolution preprocessing of event feature channel (T, P, E) data is performed through a convolution pulse encoder; the convolved features are processed using an adaptive LIF neuron model to enhance the processing capability of timing information; timing modeling is performed on the adaptively processed features through a bidirectional LSTM layer to capture the forward and backward dependencies of the time series; the processing capability of time information is further enhanced through a time gating unit to control the flow of information.
7. The method for human posture recognition based on a multimodal neural network framework according to claim 1, characterized in that: The step S4 specifically includes: S401. Dynamic feature enhancement: The dynamic feature enhancement module uses the channel attention and spatial attention mechanisms to adaptively enhance event features and spatial features based on the statistical information of the input features to enhance key features; S402. Cross-modal fusion: Use a cross-modal adaptive fusion module to perform cross-modal fusion of features.
8. The method for human posture recognition based on a multimodal neural network framework according to claim 7, characterized in that: The mathematical representation of the channel attention mechanism in step S401 is as follows: ; The mathematical representation of the spatial attention mechanism is as follows: Among them, F is the input feature map, AvgPool is the global average pooling, MaxPool is the global maximum pooling, W1 and W2 are weight matrices, δ is the ReLU activation function, σ is the Sigmoid activation function, and [;] represents concatenation in the channel dimension; In step S402, the cross-modality adaptive fusion module implements adaptive fusion through the following steps: S4021. Modality-specific encoding: spatial features and event features are processed separately through dedicated encoders to enhance their respective representation capabilities; S4022. Adaptive weight generation: Generate dynamic weights through global pooling and multi-layer perceptron to achieve dynamic weighting between modalities; the weight calculation formula is as follows: ; S4023. Weighted fusion: weight the features of different modalities according to the generated weights; ; S4024. Cross-modal attention interaction: Deep interaction between modalities through multi-head self-attention mechanism; ; in, F weighted It represents the feature fusion result after weighting. F s and F e are spatial features and event features respectively, w=[w1, w2] is the weight of the two modes, Q, K, V are query, key, and value matrices, is the scaling factor.
9. The method for human posture recognition based on a multimodal neural network framework according to claim 1, characterized in that: The step S5 specifically includes the following operations: S501. Skeleton constraint module: Combine the spiking neural network with graph convolution and introduce the spiking neuron mechanism into the standard graph convolution operation. The process is as follows: First perform the standard graph convolution operation: ; Input the result Z to the pulse neuron: ; Where X is the input feature, W is the weight matrix, b is the bias, Ã=A+I is the adjacency matrix with self-connection added, D̃ is the degree matrix of Ã, and SpikeNeuron( ) is the spike activation function; In order to enhance the rationality of human body structure, a bone length consistency loss function is designed: ; Among them, bone is the line segment connecting two key point joints in the human skeleton model, bones is all the connecting line segments that constitute the complete human skeleton, length(bone) is the length of the bone, and std represents the standard deviation; S502. Use a spatiotemporal joint processing module to capture the spatiotemporal dependencies in event data; the spatiotemporal joint processing module includes multi-scale temporal coding and multi-head spatiotemporal attention; The multi-scale time series coding adopts a hybrid feature extraction method combining a parallel convolution structure and frequency domain analysis; the parallel convolution structure includes convolution kernels of different sizes for processing the time dimension; the frequency domain analysis realizes frequency domain feature extraction through fast Fourier transform; the expression is as follows: Among them, F is the input feature, F time is the time domain feature, F freq is the frequency domain feature, F out is the output feature, Conv k represents a convolution with a kernel size of k, [;] is the concatenation in the channel dimension, W fusion is fused convolution, FFT is fast Fourier transform, LN is layer normalization, and g is a learnable gating parameter; Multi-head spatiotemporal attention establishes long-distance dependencies between different spatiotemporal positions through the self-attention mechanism. The formula is as follows: in, F attn is the attention output, G spatial It is spatial gating. G temporal is time gating and ⊙ is element-wise multiplication.
10. The method for human posture recognition based on a multimodal neural network framework according to claim 1, characterized in that: The prediction formula of the prediction head module in step S6 is as follows: Among them, x pred ,y pred is the predicted joint point coordinate, conf is the confidence, σ is the Sigmoid activation function, F global Represents global features, Project x , Project y , Project conf are the coordinates and confidence predictors, respectively, W and H are the width and height of the output feature map resolution; The loss function is expressed as follows: ; in, L pose is the key point location loss, using the mean square error to calculate the difference between the predicted key point and the true key point; λ bone represents the weight coefficient of the skeleton constraint loss, L bone is the skeleton constraint loss, ensuring that the predicted bone lengths meet physical constraints; λ temp represents the weight coefficient of the timing consistency loss, L temp is a temporal consistency loss that encourages smooth transitions between consecutive frames.
Citation Information
Patent Citations
Illumination-variable action recognition method based on event camera
CN115035597A
Gesture recognition method and electronic equipment
CN115661941A
Gesture recognition method based on global-local synaptic plasticity spiking neural network
CN118279716A
Multi-modal human body action recognition method
CN119418401A
Cited By
Intelligent ultra-deep gas well tubular column leakage identification method based on time sequence data and time-frequency fusion BiLSTM model
CN121997035A
Intelligent identification method for pipe string leakage of ultra-deep gas well based on time series data and time-frequency fusion BiLSTM model
CN121997035B
A Real-Time Drive System for Digital Humans Based on Multimodal Motion Capture
CN122550764A