Human Recognition Method and System Based on Variable Convolutional Spatiotemporal Attention Hybrid Architecture

By combining multimodal feature fusion and adaptive attention mechanism, the problems of poor environmental adaptability and mutual constraints between real-time performance and accuracy in existing technologies are solved, achieving high-precision fighting behavior recognition, reducing false alarm rate, and meeting real-time requirements.

CN120580724BActive Publication Date: 2025-10-31杭州半云科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511072231.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-31
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing technologies for human body recognition suffer from poor environmental adaptability, rely heavily on pure visual data leading to high false alarm rates, and face trade-offs between real-time performance and accuracy. They also struggle to distinguish between normal physical contact and violent behavior, and their recognition performance is particularly poor under conditions of changing lighting, occlusion, and complex scenarios.

Method used

A human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture is adopted. Through video analysis, human key point pose estimation and dynamic temporal modeling, combined with multimodal feature extraction and fusion, temporal modeling is performed using a long short-term memory network and a Transformer network to generate interpretable output to recognize fighting behavior.

Benefits of technology

It significantly improves the accuracy and recall of fight detection in complex scenarios such as changes in lighting, crowd obstruction, and rapid movements, reduces the false alarm rate, meets real-time requirements, and can effectively distinguish between normal physical contact and violent behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580724B_ABST
    Figure CN120580724B_ABST
Patent Text Reader

Abstract

This invention proposes a human body recognition method and system based on a variable convolutional spatiotemporal attention hybrid architecture. The method includes the following steps: S1, inputting raw video frames and preprocessing the input raw video frames, including denoising, resizing, and normalization; S2, employing a dual-branch heterogeneous architecture for multimodal feature extraction and fusion, achieving dynamic feature fusion through a cross-modal gating attention mechanism; S3, using a Long Short-Term Memory (LSTM) network-Transformer hybrid architecture to process short-time frame sliding window sequences, applying a Transformer encoder to the LSM network output, calculating inter-frame importance weights through a self-attention mechanism, and calculating the probability of fighting; S4, providing interpretable output, realizing attention heatmap generation, dynamic anchor box generation, alarm output, and visualization. This invention, through the combination of multimodal feature fusion and adaptive attention mechanisms, can effectively distinguish between normal physical contact and violent behavior, significantly reducing the false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human body recognition, specifically a human body recognition method and system based on a variable convolutional spatiotemporal attention hybrid architecture. Background Technology

[0002] Current two-person fight detection technologies based on computer vision and deep learning can be categorized into the following core methods:

[0003] (1) Action recognition based on single-frame images uses CNN (Convolutional Neural Network) to extract spatial features to judge limb actions (such as punching and kicking), but ignores temporal information and is prone to misjudging violent movements (such as running and dancing) as fighting.

[0004] (2) Based on human body key point pose analysis, real-time multi-person pose estimation models (such as the OpenPose model) are used to detect skeletal key points, and aggressive actions are identified through Euclidean space features of joint motion trajectories or angle changes. However, key points are easily occluded when multiple people are entangled, and the three-dimensional spatial relationship is difficult to accurately restore due to the influence of camera angle and distance.

[0005] It is evident that existing technologies face numerous problems:

[0006] (1) Poor environmental adaptability. Changes in lighting and shadow interference can cause visual feature distortion, rendering traditional image processing methods, such as background subtraction, ineffective.

[0007] (2) In complex scenes, visual models have difficulty separating fighting targets. Existing algorithms rely on interpolation to complete the processing of occluded targets, which reduces accuracy.

[0008] (3) Real-time performance and accuracy are mutually restrictive. Models that meet accuracy requirements need to process multiple frame sequences, with a delay of up to hundreds of milliseconds, which cannot meet the real-time early warning requirements of security scenarios.

[0009] (4) Insufficient joint modeling and spatiotemporal feature modeling, or deep dependence on device hardware.

[0010] (5) It relies heavily on pure visual data and lacks multimodal fusion, resulting in a high false alarm rate in low light or occluded scenes. It is difficult to adapt to the differences in fighting actions of different body types and cultural backgrounds, and it is easy to identify boxing, rugby and other sports as fighting behaviors.

[0011] Therefore, a high-precision human body recognition method is needed that can effectively distinguish between normal physical contact and violent behavior. Summary of the Invention

[0012] To overcome the aforementioned deficiencies in existing technologies, this invention provides a human body recognition method and system based on a variable convolutional spatiotemporal attention hybrid architecture. Through video analysis, human key point pose estimation, and dynamic temporal modeling, it achieves high-precision recognition of fighting behavior.

[0013] The primary objective of this invention is to propose a human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture, which includes the following steps:

[0014] S1. Input raw video frames. Perform data preprocessing on the input raw video frames. The data preprocessing includes noise reduction, size adjustment and normalization.

[0015] S2. A dual-branch heterogeneous architecture is adopted for multimodal feature extraction and fusion, and dynamic feature fusion is achieved through a cross-modal gating attention mechanism; the dual branches include a visual feature extraction branch and a pose feature extraction branch;

[0016] S3. Temporal modeling is performed using a hybrid architecture of Long Short-Term Memory (LSTM) network and Transformer network. A bidirectional LSTM network is used to process short-time frame sliding window sequences. A Transformer encoder is applied to the output of the LSTM network, and the inter-frame importance weights are calculated through a self-attention mechanism to process temporal blocks of various durations in parallel. The output collision probability is calculated.

[0017] S4. Interpretable output enables attention heatmap generation, dynamic anchor box generation, alarm output, and visualization.

[0018] Preferably, S2 includes:

[0019] S21. The visual feature branch uses a deformable convolutional network to construct a multi-scale feature pyramid, dynamically adjusts the sampling position of the convolutional kernel, and connects channel attention and spatial attention for feature filtering, outputting a visual feature map that includes high-resolution details and low-resolution global features.

[0020] S22. The pose feature branch constructs a fully connected human skeleton map and feature matrix through the coordinates of human key points; it aggregates neighbor node information through a graph attention network and obtains the pose feature map of limb motion dynamics after hierarchical pooling.

[0021] S23. Cross-modal fusion introduces a bidirectional cross-attention mechanism to achieve feature complementarity. It uses a gated fusion factor to dynamically weight visual feature maps and pose feature maps to complete feature fusion. The training optimization of the fusion model adopts a multi-task loss function.

[0022] Preferably, the channel attention employs a dual-path aggregation of global average pooling and global max pooling; the spatial attention generates a spatial weight map through convolution and outputs a subtle action feature map.

[0023] Preferably, the bidirectional cross-attention mechanism includes visual pose direction and pose-visual direction, and the multi-task loss function includes visual task loss, pose task loss and cross-modal task loss.

[0024] Preferably, S3 includes:

[0025] S31. A bidirectional long short-term memory network layer is used to process the input sequence and capture short-term action patterns of punching and shoving.

[0026] S32. The multi-scale spatiotemporal attention mechanism uses the Transformer encoder part to perform inter-frame importance weighting on the output of the bidirectional long short-term memory network, and calculates the inter-frame importance weights through a self-attention mechanism.

[0027] S33. Perform multi-scale analysis and process temporal blocks of different frame lengths in parallel.

[0028] S34, multi-head attention weighting, gated fusion, residual connection;

[0029] S35, a fully connected network combined with a sigmoid activation function, uses classification features to calculate the probability of a fight, optimizing sample balancing and temporal smoothing.

[0030] Preferably, S4 includes:

[0031] S41. Input the probability of fighting and the visual feature map, generate a color heatmap, use different colors to distinguish importance, and highlight the signs of physical contact conflict;

[0032] S42. Cluster the key points of the human body, identify dense conflict areas, and generate adaptive anchor boxes.

[0033] S43: Overlay the original image with a color-level heatmap, mark conflict areas with anchor points, and trigger an alarm when the probability of a conflict exceeds a preset threshold.

[0034] The second objective of this invention is to propose a human body recognition system based on a variable convolutional spatiotemporal attention hybrid architecture, which includes the following modules:

[0035] Data preprocessing module: performs noise reduction, resizing, and normalization on the input raw video frames;

[0036] Feature extraction and fusion module: Multimodal feature extraction and dynamic fusion are achieved through a dual-branch heterogeneous architecture and a cross-modal gated attention mechanism;

[0037] Temporal modeling module: Uses a hybrid architecture of long short-term memory network and transformer network for temporal modeling, and processes temporal blocks of various durations in parallel;

[0038] Interpretable output module: Enables attention heatmap generation, dynamic anchor box generation, alarm output and visualization.

[0039] Preferably, the feature extraction and fusion module includes:

[0040] The visual feature extraction submodule uses a deformable convolutional network to construct a multi-scale feature pyramid and performs feature filtering through cascaded convolutional block attention modules.

[0041] The pose feature extraction submodule is used to extract human pose features from video frames, normalize the coordinates of human key points into offset vectors relative to the torso, construct a fully connected human skeleton map and feature matrix, and aggregate neighbor node information through a graph attention network.

[0042] Cross-modal attention fusion submodule: This module introduces a bidirectional cross-attention mechanism, which uses a gated fusion factor to dynamically weight the outputs of the two modalities, and fuses and complements visual and pose features.

[0043] Preferably, the time series modeling module includes:

[0044] The Long Short-Term Memory (LSTM) network layer submodule employs a bidirectional LSM network layer to capture short-term actions.

[0045] The Transformer encoder submodule performs inter-frame importance weighting on the output of the bidirectional long short-term memory network layer and uses a multi-scale spatiotemporal attention mechanism to calculate cross-frame correlation.

[0046] The multi-scale analysis submodule processes temporal blocks of different frame lengths in parallel to capture motion features at different time scales.

[0047] Gated fusion submodule: used for multi-head attention weighting, gated fusion, and residual connections;

[0048] Fighting probability output submodule: used for sample balancing and temporal smoothing optimization, and to calculate the output fighting probability using classification features.

[0049] The technical solution of this invention has the following positive effects:

[0050] (1) By combining multimodal feature fusion and adaptive attention mechanism, background interference is effectively suppressed and target features are enhanced. Compared with mainstream fighting detection methods HOF, MoSIFT and 3DCNN, the accuracy and recall of recognition are improved.

[0051] (2) Through the fusion of visual and posture multimodal features and spatiotemporal attention mechanism, the system can accurately identify aggressive actions such as punching and kicking, and effectively distinguish between normal physical contact and violent behavior.

[0052] (3) In complex scenarios such as changes in lighting, crowd occlusion, and rapid movements, the accuracy of the fight detection in this solution is significantly higher than that of traditional methods. This solution can process up to 25 frames per second, meeting real-time requirements.

[0053] The method and system of this invention employ deep learning-driven multimodal data fusion technology to effectively distinguish between normal physical contact and violent behavior, significantly reducing the false alarm rate. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the human body recognition method according to an embodiment of the present invention;

[0055] Figure 2 This is a schematic diagram of the preprocessing flow of the human body recognition method according to an embodiment of the present invention;

[0056] Figure 3 This is a schematic diagram of the multimodal feature extraction process of the human body recognition method according to an embodiment of the present invention;

[0057] Figure 4 This is a schematic diagram of the timing modeling process of the human body recognition method according to an embodiment of the present invention;

[0058] Figure 5 This is a schematic diagram illustrating the output process of the human body recognition method according to an embodiment of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Example 1

[0061] This invention proposes a human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture, such as... Figure 1 As shown, it includes the following steps:

[0062] S1. Input raw video frames. Perform data preprocessing on the input raw video frames. The data preprocessing includes noise reduction, size adjustment and normalization.

[0063] like Figure 2As shown, the system receives raw video frames as input and then preprocesses the input data, performing denoising, resizing, and normalization operations on each frame to ensure data quality. Image denoising can employ a median filtering algorithm with a window size of 5×5, aiming to eliminate uneven lighting and sensor noise while preserving limb edge details. Normalization uses bilinear interpolation to scale the image to a uniform size to fit subsequent modules; for example, it can be scaled to 512×512 pixels. The value of each pixel is adjusted to the range [0,1] to improve image stability and consistency, ensuring comparability of features between different images and laying a solid data foundation for subsequent feature extraction and tracking.

[0064] S2. A dual-branch heterogeneous architecture is adopted for multimodal feature extraction and fusion, and dynamic feature fusion is achieved through a cross-modal gating attention mechanism; the dual branches include a visual feature extraction branch and a pose feature extraction branch.

[0065] like Figure 3 As shown, multimodal feature extraction and fusion adopts an improved dual-branch heterogeneous architecture, consisting of a visual feature extraction branch and a pose feature extraction branch, and achieves dynamic feature fusion through a cross-modal gated attention mechanism.

[0066] This step includes the following sub-steps:

[0067] S21. The visual feature branch uses a deformable convolutional network to construct a multi-scale feature pyramid, and then concatenates channel attention and spatial attention for feature selection.

[0068] To extract robust visual features against changes in lighting and occlusion, a three-layer deformable convolutional network is employed, dynamically adjusting the sampling position of the convolutional kernels to adapt to severe limb deformation. Channel attention and spatial attention are concatenated to suppress interference from complex backgrounds and focus on the contact area of ​​conflicting limbs.

[0069] In this embodiment of the invention, the visual feature branch input is a 512×512×3 RGB image, which is used to construct a multi-scale feature pyramid through three layers of deformable convolution. Each convolutional kernel is equipped with a learnable offset parameter Δp, and deformation perception is achieved through 3×3 grid sampling, dynamically adapting to local deformations caused by limb bending (such as 90° elbow flexion). The network uses multi-level dilation rates (dilation rates=[1,2,3]) to construct the pyramid receptive field, simultaneously capturing hand micro-movements (P3 level) and whole-body posture (P5 level) features. Subsequently, feature selection is performed through cascaded CBAM (Convolutional Block Attention Module). CBAM includes channel attention and spatial attention.

[0070] Channel attention employs a dual-path aggregation method using Global Average Pooling (GAP) and Global Max Pooling (GMP), generating a weight vector W through the Sigmoid function (a sigmoid function). c ∈[0,1] C It should be noted that in the definition of the weight vector, the quantity W... c ∈[0,1] C This indicates that Wc is a C-dimensional vector, with each component taking values ​​in the interval [0,1].

[0071] Spatial attention generates a spatial weight map W through 7×7 convolution. s ∈[0,1] H×W The focus is on enhancing the response in conflict areas such as limb intersections, ultimately outputting a P3-level feature map F. v ∈R {256×256×64} Its high resolution allows for precise positioning of pixel-level details.

[0072] S22. The pose feature branch constructs a fully connected human skeleton map and feature matrix through the coordinates of human key points; then, it obtains the spatial features of human key points through graph attention network and hierarchical pooling.

[0073] The pose feature branch quantifies the Euclidean physical relationships of limb interactions and the hidden connections in non-Euclidean space. A lightweight OpenPose model (a real-time multi-person pose estimation library) is used to output the coordinates of 17 human keypoints in real time, including the nose, elbow, and wrist, which are normalized to offset vectors relative to the torso. Subsequently, a fully connected human skeleton map and feature matrix are constructed. Edge features include relative distances between joint angles, such as the angle between the elbow-wrist line and the shoulder-elbow line. A Graph Attention Network (GAT) is used to aggregate neighbor node information, outputting a 256-dimensional spatial relationship vector, thus obtaining the keypoint spatial features fused from Euclidean and non-Euclidean spaces.

[0074] The input to the pose feature branch is a matrix of coordinates of 17 keypoints detected by OpenPose in two-dimensional space, where P ∈ R. {17×2} After normalization, a fully connected graph G=(V,E) based on human anatomy is constructed, where vertices V correspond to keypoints and edges E represent skeletal connections. A graph attention network (GAT) is used for hierarchical encoding, employing a 4-head mechanism, where each head calculates the attention coefficient a for joint i to j. ij :

[0075] ,

[0076] Where a ij Here, W is the learnable vector, Whi and Whj represent the feature vectors of joints i and j, respectively, W is the weight matrix, LeakyReLU is the activation function, and softmax is the normalization function.

[0077] Finally, the network aggregates local features step by step through hierarchical pooling, and the final output 256-dimensional vector contains limb motion dynamics features.

[0078] S23. Cross-modal fusion introduces a bidirectional cross-attention mechanism to achieve feature complementarity. In the feature fusion stage, a gated fusion factor is used to dynamically weight the outputs of the two modalities. The training and optimization of the fusion model adopts a multi-task loss function.

[0079] In this embodiment of the invention, cross-modal fusion introduces a bidirectional cross-attention mechanism to achieve feature complementarity. The 256×256×64 feature map is flattened into a patch sequence in the visual pose direction, which serves as the key / value pair. The pose vector serves as the query, and a similarity score is calculated. .

[0080] In the pose vision direction, the pose vector is used as the query, and the channel vectors of the visual feature map are used as the key / value pairs. A gated fusion factor is used in the feature fusion stage. ,in It is a weight matrix. It is a bias vector. This indicates that the concatenation comes from the visual feature branch and the pose feature branch.

[0081] The two modal outputs are dynamically weighted, and the visual and pose features are ultimately combined.

[0082] ,

[0083] Where F vis and F pose These represent visual features and posture features, respectively.

[0084] The training optimization of the fusion model adopts a multi-task loss function, in which the cross-modal contrastive loss Lcross promotes the similarity of positive sample pairs (v, p+) to be higher than that of negative sample pairs (v, p-) by at least a marginal value m. Modality-invariant representation learning is achieved through NT-Xent (Normalized Temperature-scaled Cross Entropy) loss, and the function formula is as follows:

[0085] .

[0086] Where L vis (BCE) is the loss for the visual task, using binary cross-entropy loss (BCE); L pose (MAE) is the loss for the pose task, using Mean Absolute Error Loss (MAE); L cross (InfoNCE) is the loss for cross-modal tasks, using information noise contrastive estimation (InfoNCE Loss). , and These are the weights of each loss function.

[0087] S3. Temporal Modeling: A hybrid architecture of Long Short-Term Memory (LSTM) network and Transformer network is used to process short-time frame sliding window sequences to capture short-time action continuity. A Transformer encoder is applied to the output of the LSM network, and the importance weights between frames are calculated through a self-attention mechanism. Three types of temporal blocks are processed in parallel; the probability of output collision is calculated.

[0088] like Figure 4As shown, time series modeling refers to the process of analyzing and modeling data points arranged in chronological order. Its purpose is to capture the temporal dependencies and patterns in the data for prediction, system description, system analysis, decision-making, and control. LSTM (Long Short-Term Memory) is a special type of recurrent neural network (RNN) that effectively handles long-term dependencies in time series data, solving the gradient vanishing or exploding problems faced by traditional RNNs when processing long-sequence data.

[0089] This step includes the following sub-steps:

[0090] S31. A bidirectional long short-term memory network layer is used to process the input sequence and capture short-term action patterns.

[0091] The bidirectional temporal dependency modeling uses a bidirectional LSTM layer (BiLSTM) to process the input sequence, with a hidden layer of 256 dimensions. The model solves the long-range gradient problem through a forget gate / input gate / output gate mechanism, focusing on capturing two types of short-term action patterns: (1) instantaneous burst actions such as sudden acceleration changes when punching; (2) continuous actions such as continuous pushing, which have periodic characteristics.

[0092] S32. The multi-scale spatiotemporal attention mechanism uses the Transformer encoder to perform inter-frame importance weighting on the BiLSTM output, and calculates the inter-frame importance weights through a self-attention mechanism to detect instantaneous burst actions and continuous behaviors.

[0093] This embodiment of the invention employs an 8-head attention mechanism with 2 stacked layers. The Q / K / V matrices are projected onto a 32-dimensional subspace (256 / 8) respectively, and cross-frame correlations are calculated:

[0094] ,

[0095] Where Q (Query), K (Key), and V (Value) are three matrices in the attention mechanism, d k It is the dimension of the key matrix.

[0096] S33. Perform multi-scale analysis and process time blocks of different frame lengths in parallel.

[0097] Multi-scale analysis is performed, with parallel processing of 9 / 16 / 25 time blocks (corresponding to 0.3s / 0.53s / 0.83s), captured through different windows. Short time blocks are used to detect instantaneous bursts and rapid movements, while long time blocks are used to identify posture entanglement patterns in sustained grappling.

[0098] S34, multi-head attention weighting, gated fusion, residual connection.

[0099] For the residual connection part, the model adds a cross-layer skip connection with learnable weight α=0.3 to alleviate the degradation problem of deep networks.

[0100] The model employs gated fusion for multi-scale attention outputs. Generate time-sensitive classification features. This indicates that features at different scales (9 frames, 16 frames, 25 frames) are concatenated. These features capture instantaneous motion and long-term pose patterns in the video, respectively.

[0101] S35 uses a two-layer fully connected network with a Sigmoid activation function to calculate the probability of fighting using classification features, and optimizes the design for sample balancing and temporal smoothing.

[0102] A two-layer fully connected network, combined with a Sigmoid activation function, calculates the output fighting probability p ∈ [0,1] using classification features. The optimization design includes sample balancing and temporal smoothing. Sample balancing uses Focal Loss (γ=2) to alleviate the imbalanced data problem, and temporal smoothing applies Temporal Gaussian smoothing (σ=3) to the output of 16 consecutive frames to eliminate instantaneous misclassification.

[0103] S4. Interpretable output enables attention heatmap generation, dynamic anchor box generation, alarm output, and visualization.

[0104] like Figure 5 As shown, this module is used to provide visual decision-making support. Based on the Grad-CAM++ algorithm, it backpropagates the classification gradient to the feature map and generates an attention heatmap. The red highlighted area indicates the key parts of the fighting behavior.

[0105] The Grad-CAM++ algorithm is an improved visualization algorithm used to generate Class Activation Maps (CAMs) to visually represent the image regions that deep learning models focus on when making classification decisions. This step includes the following sub-steps:

[0106] S41. Input the probability of fighting and the visual feature map, generate a color heatmap, use different colors to distinguish the importance, and highlight the signs of physical contact conflict.

[0107] The fighting probability score p ∈ [0, 1] from the received temporal modeling output is combined with the P3 level visual feature map (256×256×64) from feature extraction and fusion output. The feature map F is then calculated. ij Classification score y cHigher-order partial derivatives are used to generate pixel-level weighting coefficients α. c ij Then, ReLU activation is applied to the weighted feature map to preserve the positive contribution regions, generating a heatmap L. Grad-CAM++ The heatmap color mapping uses the Jet color scale, with red to blue indicating descending order of importance, highlighting signs of physical contact conflict.

[0108]

[0109] S42. Cluster the key points of the human body, identify dense conflict areas, and generate adaptive anchor boxes.

[0110] The model adaptively adjusts the shape of the anchor box based on the key point distribution of OpenPose to enhance the target enclosure accuracy. First, DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering was performed on the coordinates of 17 human key points to identify dense conflict areas. Then, the shape adaptation rules were used to generate the adapted anchor boxes. The adaptation rules are as follows: (1) a 1.2:1 rectangular box is used for a single person standing posture with a uniform key point distribution; (2) a 1:1 square box is used for a curled-up posture and a squatting posture with a high concentration of key points; (3) a 0.8:1.5 long strip box is used for a multi-person limb entanglement posture with a linear distribution of key points.

[0111] S43: Overlay the original image with a color-level heatmap, mark conflict areas with anchor points, and trigger an alarm when the probability of a fight exceeds a preset threshold.

[0112] The heatmap generated by the color gradation heatmap generation submodule is overlaid on the original video frame, and the identified conflict areas are marked with anchor boxes generated by the adaptation anchor box submodule. When the probability of conflict exceeds a preset threshold, an alarm mechanism is triggered; for example, the threshold can be set to 70%.

[0113] The steps in this embodiment of the invention, through the combination of multimodal feature fusion and adaptive attention mechanism, can effectively suppress background interference, enhance target features, and improve the accuracy of human body recognition.

[0114] Example 2

[0115] This invention proposes a human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture, comprising the following modules:

[0116] Data preprocessing module 10: preprocesses the input raw video frames to improve image stability and consistency. The preprocessing includes noise reduction, resizing, and normalization.

[0117] Noise reduction: A median filtering algorithm with a window size of 5×5 is used to eliminate uneven lighting and sensor noise while preserving limb edge details.

[0118] Size adjustment: The image is scaled to 512×512 pixels using bilinear interpolation to unify the input size and adapt to subsequent modules.

[0119] Normalization: Adjusts pixel values ​​to the range of [0,1] to ensure that features between different images are comparable.

[0120] The feature extraction and fusion module 20 includes a visual feature extraction submodule 21, a pose feature extraction submodule 22, and a cross-modal attention fusion submodule 23. It adopts an improved dual-branch heterogeneous architecture, consisting of a visual feature extraction branch and a pose feature extraction branch, and achieves dynamic feature fusion through a cross-modal gating attention mechanism.

[0121] The visual feature extraction submodule 21 is used to extract visual features from video frames. A multi-scale feature pyramid is constructed using a deformable convolutional network, and feature selection is performed through cascaded CBAM (convolutional block attention modules), including channel attention and spatial attention.

[0122] Channel attention: Employs dual-path aggregation using global average pooling and global max pooling, generating weight vectors through the Sigmoid function. Spatial attention: Generates a spatial weight map through 7×7 convolution, focusing on enhancing responses in conflict areas such as limb intersections.

[0123] The visual feature extraction submodule 21 captures hand micro-movements (P3 level) and whole-body posture (P5 level) features, corresponding to high-resolution limb details to low-resolution global movements, to enhance multi-scale detection capabilities.

[0124] The pose feature extraction submodule 22 is used to extract human pose features from video frames. A lightweight OpenPose model is employed to output the coordinates of 17 human keypoints in real time. The keypoint coordinates are normalized into offset vectors relative to the torso, and a fully connected human skeleton map and feature matrix are constructed. Neighbor node information is aggregated through a graph attention network (GAT), outputting a 256-dimensional spatial relationship vector, resulting in a fusion of Euclidean and non-Euclidean space keypoint spatial features.

[0125] The cross-modal attention fusion submodule 23 is used to achieve complementary fusion of visual features and pose features. A bidirectional cross-attention mechanism is introduced.

[0126] Visual pose orientation: The visual feature map is flattened into a patch sequence, which serves as the key / value pair, and the pose vector serves as the query. A similarity score is calculated. Visual pose orientation: The pose vector serves as the query, and the channel vectors of the visual feature map serve as the key / value pair. A gated fusion factor is used to dynamically weight the outputs from both modalities. A multi-task loss function is employed, including the cross-modal contrastive loss Lcross, and modality-invariant representation learning is achieved through NT-Xent loss.

[0127] The temporal modeling module 30 includes an LSTM layer submodule 31, a Transformer encoder submodule 32, a multi-scale analysis submodule 33, and an optimization design submodule 34. It processes the input sequence and captures motion features at different time scales through multi-scale analysis.

[0128] LSTM layer submodule 31 is used to process the input sequence and capture short-term action patterns. It adopts a bidirectional long short-term memory network layer (BiLSTM). The hidden layer has 256 dimensions and solves the long-range gradient problem through forget gate, input gate, and output gate mechanisms, focusing on capturing instantaneous burst actions and continuous actions.

[0129] Transformer encoder submodule 32 performs inter-frame importance weighting on the BiLSTM output to detect instantaneous bursts of action and sustained behavior. A multi-scale spatiotemporal attention mechanism is employed; in this embodiment, an 8-head attention mechanism with 2 stacked layers is used. First, the Q / K / V matrices are projected onto a 32-dimensional subspace, and cross-frame correlations are calculated.

[0130] The multi-scale analysis submodule 33 processes time blocks of different frame lengths in parallel to capture motion features at different time scales. The time blocks include 9 / 16 / 25 frame time blocks (corresponding to 0.3s / 0.53s / 0.83s). Short time blocks are used to detect instantaneous bursts of motion and rapid motion. Long time blocks are used to identify posture entanglement patterns in continuous grappling.

[0131] Gated fusion submodule 34 is used for residual connections: adding cross-layer skip connections with learnable weights α=0.3 to alleviate the degradation problem of deep networks; gated fusion: gating fusion is applied to multi-scale attention outputs to generate time-sensitive classification features.

[0132] The fighting probability output submodule 35 includes a classification network and optimization strategies. The classification network consists of a two-layer fully connected network with sigmoid activation, using classification features to calculate the output fighting probability. Optimization strategies include sample balancing and temporal smoothing.

[0133] The interpretable output module 40 includes a color-scale heatmap generation submodule 41, an adaptation anchor box submodule 42, and a heatmap overlay anchor box marking submodule 43. It is used to generate interpretable output results, including attention heatmaps and adaptation anchor boxes, so that users can understand the model's decision-making process.

[0134] The color-gradient heatmap generation submodule 41 generates an attention heatmap based on the Grad-CAM++ algorithm. It inputs the fighting probability score p∈[0,1] output by the temporal modeling module 30 and the P3-level visual feature map from the feature extraction and fusion module 20. It calculates the pixel-level weight coefficients of the feature map, applies ReLU activation to the weighted feature map, preserves positive contribution regions, and generates the color-gradient heatmap.

[0135] The adaptation anchor point box module 42 clusters key points, identifies dense conflict areas, and generates adaptation anchor point boxes to enhance target enclosing accuracy. Based on the OpenPose key point distribution, the DBSCAN clustering algorithm is used to identify dense conflict areas.

[0136] The heatmap overlay anchor box marking submodule 43 overlays the heatmap generated by the color-level heatmap generation submodule with the original video frame, and simultaneously marks the identified conflict areas with anchor boxes generated by the adaptation anchor box submodule. When the probability of conflict exceeds a preset threshold, an alarm mechanism is triggered.

[0137] The modules in this embodiment of the invention work together to realize a method for human body recognition, which can effectively distinguish between normal physical contact and violent behavior.

[0138] This invention was tested on the Violent-Flows public dataset for violent behavior recognition. This dataset contains 123 violent video clips and 123 non-violent video clips. Accuracy and recall were used as the core evaluation metrics in the experiments, and the results were compared with current mainstream fight detection methods such as HOF, MoSIFT, and 3DCNN.

[0139] Table 1. Comparison of different models on the Violet-Flows dataset

[0140] method accuracy Recall rate This embodiment 91.0% 91.8% HOF 82.5% 80.1% MoSIFT 88.3% 86.5% 3DCNN 90.1% 85.6%

[0141] Experimental results show that our system significantly outperforms the comparison methods in both accuracy and recall. Our system maintains a stable advantage on random test sets, especially in scenarios with crowd occlusion and fast-moving action.

[0142] It should be noted that the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0143] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture, characterized in that, Includes the following steps: S1. Input raw video frames. Perform data preprocessing on the input raw video frames. The data preprocessing includes noise reduction, size adjustment and normalization. S2. A dual-branch heterogeneous architecture is adopted for multimodal feature extraction and fusion, and dynamic feature fusion is achieved through a cross-modal gating attention mechanism; the dual branches include a visual feature extraction branch and a pose feature extraction branch; S3. A hybrid architecture of Long Short-Term Memory (LSTM) network and Transformer network is used for temporal modeling. The bidirectional LSTM network processes short-time frame sliding window sequences. A Transformer encoder is applied to the output of the LSTM network, and the inter-frame importance weights are calculated through a self-attention mechanism to process temporal blocks of various durations in parallel. Calculate and output the probability of a fight; S4. Interpretable output enables attention heatmap generation, dynamic anchor box generation, alarm output, and visualization. Step S3 includes: S31. A bidirectional long short-term memory network layer is used to process the input sequence and capture short-term action patterns of punching and shoving. S32. The multi-scale spatiotemporal attention mechanism uses the Transformer encoder part to perform inter-frame importance weighting on the output of the bidirectional long short-term memory network, and calculates the inter-frame importance weights through a self-attention mechanism. S33. Perform multi-scale analysis and process temporal blocks of different frame lengths in parallel. S34, multi-head attention weighting, gated fusion, residual connection; S35, a fully connected network combined with a sigmoid activation function, uses classification features to calculate the probability of a fight, optimizing sample balancing and temporal smoothing.

2. The human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture according to claim 1, characterized in that, S2 includes: S21. The visual feature branch uses a deformable convolutional network to construct a multi-scale feature pyramid, dynamically adjusts the sampling position of the convolutional kernel, and connects channel attention and spatial attention for feature filtering, outputting a visual feature map that includes high-resolution details and low-resolution global features. S22. The pose feature branch constructs a fully connected human skeleton map and feature matrix through the coordinates of human key points; it aggregates neighbor node information through a graph attention network and obtains the pose feature map of limb motion dynamics after hierarchical pooling. S23. Cross-modal fusion introduces a bidirectional cross-attention mechanism to achieve feature complementarity. It uses a gated fusion factor to dynamically weight visual feature maps and pose feature maps to complete feature fusion. The training optimization of the fusion model adopts a multi-task loss function.

3. The human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture according to claim 2, characterized in that, The channel attention uses a dual-path aggregation of global average pooling and global max pooling; the spatial attention generates a spatial weight map through convolution and outputs a subtle action feature map.

4. The human recognition method based on a variable convolutional spatiotemporal attention hybrid architecture according to claim 2, characterized in that, The bidirectional cross-attention mechanism includes visual pose direction and pose-visual direction, and the multi-task loss function includes visual task loss, pose task loss and cross-modal task loss.

5. The human body recognition method based on a variable convolutional spatiotemporal attention hybrid architecture according to claim 1, characterized in that, S4 includes: S41. Input the probability of fighting and the visual feature map, generate a color heatmap, use different colors to distinguish importance, and highlight the signs of physical contact conflict; S42. Cluster the key points of the human body, identify dense conflict areas, and generate adaptive anchor boxes. S43: Overlay the original image with a color-level heatmap, mark conflict areas with anchor points, and trigger an alarm when the probability of a conflict exceeds a preset threshold.

6. A human body recognition system based on a variable convolutional spatiotemporal attention hybrid architecture, characterized in that, Includes the following modules: Data preprocessing module: performs noise reduction, resizing, and normalization on the input raw video frames; Feature extraction and fusion module: Multimodal feature extraction and dynamic fusion are achieved through a dual-branch heterogeneous architecture and a cross-modal gated attention mechanism; Temporal modeling module: Uses a hybrid architecture of long short-term memory network and transformer network for temporal modeling, and processes temporal blocks of various durations in parallel; Interpretable output module: Enables attention heatmap generation, dynamic anchor box generation, alarm output, and visualization; The time series modeling module includes: The Long Short-Term Memory (LSTM) network layer submodule employs a bidirectional LSM network layer to capture short-term actions. The Transformer encoder submodule performs inter-frame importance weighting on the output of the bidirectional long short-term memory network layer and uses a multi-scale spatiotemporal attention mechanism to calculate cross-frame correlation. The multi-scale analysis submodule processes temporal blocks of different frame lengths in parallel to capture motion features at different time scales. Gated fusion submodule: used for multi-head attention weighting, gated fusion, and residual connections; Fighting probability output submodule: used for sample balancing and temporal smoothing optimization, and to calculate the output fighting probability using classification features.

7. The human recognition system based on a variable convolutional spatiotemporal attention hybrid architecture according to claim 6, characterized in that, The feature extraction and fusion module includes: The visual feature extraction submodule uses a deformable convolutional network to construct a multi-scale feature pyramid and performs feature filtering through cascaded convolutional block attention modules. The pose feature extraction submodule is used to extract human pose features from video frames, normalize the coordinates of human key points into offset vectors relative to the torso, construct a fully connected human skeleton map and feature matrix, and aggregate neighbor node information through a graph attention network. Cross-modal attention fusion submodule: This module introduces a bidirectional cross-attention mechanism, which uses a gated fusion factor to dynamically weight the outputs of the two modalities, and fuses and complements visual and pose features.

Citation Information

Patent Citations

  • Human body action recognition method based on single IMU (Inertial Measurement Unit) sensor

    CN119810927A

  • Group behavior identification method and system based on cross-feature interaction Transform

    CN120388335A