A Motion Behavior Recognition Method Based on Multimodal Collaborative Self-Attention Network

By introducing a multimodal collaborative self-attention network into behavior recognition technology, combining depth data and bone data to construct motion collaborative spatial characteristics, the shortcomings of the existing technology under the influence of light and background are solved, and more accurate and efficient recognition of human movement behaviors is achieved.

CN114944012BActive Publication Date: 2025-06-24CHANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210544862.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-06-24
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

The existing behavior recognition technology is not effective under the influence of light and background factors, and the traditional methods rely on artificial feature extraction, and the generalization performance is insufficient, which cannot effectively reflect the integrity and synergy of human movement.

Method used

The motor behavior recognition method based on multimodal coordinated self-attention network is adopted. By combining depth data and skeleton data, the movement coordinated spatial characteristics are constructed to reflect the integrity and coordination of human movement, and the multi-head attention mechanism and channel attention mechanism are used to achieve the fusion of multimodal data.

Benefits of technology

It improves the accuracy and generalization ability of human movement behavior recognition, can more effectively reflect the integrity and coordination of human movement, and reduces redundant information through multimodal fusion and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114944012B_ABST
    Figure CN114944012B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and particularly to a motion behavior recognition method based on a multi-modal collaborative self-attention network, including: constructing a motion collaborative spatial feature sequence set by using skeleton sequence data and a collaborative spatial vector algorithm; inputting the motion collaborative spatial feature sequence set into a skeleton self-attention sub-network based on a Transfomer network; training through a depth self-attention sub-network integrating an attention mechanism by using depth image data and classifying through a Softmax classifier; fusing the skeleton sequence classification and the depth image classification to obtain a fused classification result. The present invention uses depth data and skeleton data on the basis of a transformer model, comprehensively considers the motion collaborative spatial features of each joint point information, reflects the integrity and coordination of human motion, and at the same time proposes a quantization standard for the contribution degree of each joint point participating in motion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a motion behavior recognition method based on a multi-modal collaborative self-attention network. Background Art

[0002] Human behavior recognition is one of the research hotspots in the field of computer vision, and many research results have been widely applied in fields such as image analysis, human-computer interaction, intelligent monitoring, video retrieval, motion sensing games, and health detection.

[0003] Currently, the research on behavior recognition mainly focuses on depth image sequences, human skeleton sequences, and video sequences. The depth image sequences collected by structured light depth sensors are insensitive to light changes and provide depth data of human behaviors, but they are still restricted by factors such as light and background. Using skeleton sequences can well overcome the influence of appearance factors and has the advantages of clear and simple features and strong spatial information correlation. However, representing the human body as dozens of joint points inevitably loses a lot of human details. Video sequences contain the most complete spatio-temporal information, but there is a large amount of redundant information and the calculation time is too long.

[0004] Traditional behavior recognition mainly relies on manually extracting features for recognition, and there are certain defects in terms of generalization performance and the like.

[0005] Most of the existing behavior recognitions based on skeleton data separately utilize the information of each skeleton node, while each joint part in the human motion process is mutually cooperative and interactive, and human motion has the characteristics of integrity and cooperation. In order to better reflect these characteristics, the present invention proposes a motion collaborative spatial feature.

[0006] Multi-modal fusion refers to the technology that a machine obtains information from multiple fields, realizes information conversion and fusion, thereby improving the performance of the model. Since multi-modal data can make the data more complete, realize the complementarity of various heterogeneous information, and eliminate the redundancy between modalities, thus learning better feature representations, researchers have begun to focus on how to fuse data from multiple fields to improve the model performance and achieve better recognition effects. Summary of the Invention

[0007] Aiming at the deficiencies of the existing algorithms, the present invention uses depth data and skeleton data based on the transformer model, synthesizes the motion collaborative spatial features of the information of each joint point, reflects the integrity and cooperation of human motion, and at the same time proposes a quantization standard for the contribution degree of each joint point participating in the motion.

[0008] The technical solution adopted by the present invention is: a motion behavior recognition method based on a multi-modal collaborative self-attention network includes the following steps:

[0009] S1. Collect bone sequence data through an inertial sensor; construct a motion collaborative space feature sequence set using the bone sequence data and a motion collaborative space vector algorithm; input the motion collaborative space feature sequence set into a bone self-attention subnetwork based on a Transfomer network and classify it through a Softmax classifier;

[0010] Further, the bone sequence data includes the following actions: the right arm slides to the left, the right arm slides to the right, wave the right hand, clap hands in front, the right arm throws, cross arms in front of the chest, basketball shooting, draw an x with the right hand, draw a circle with the right hand (clockwise), draw a circle with the right hand (counterclockwise), draw a triangle, bowling (right hand), front punch, baseball right swing, tennis right forehand swing, spin arms (both arms), tennis serve, push with both hands, knock on the door with the right hand, grab an object with the right hand, pick up / throw with the right hand, jog in place, walk in place, sit-to-stand, stand-to-sit, forward lunge (left foot forward), squat (both arms straight);

[0011] Further, the collaborative space vector algorithm includes: when the human body is performing various actions, the movement conditions of each part of the body are different. Most of the existing behavior recognition based on bone data separately uses the information of each bone node, while the joint parts of the human body are coordinated and interact with each other during the movement process. The human movement has the characteristics of integrity and coordination. In order to better reflect these characteristics, constructing a motion collaborative space feature sequence is to construct a motion collaborative space vector, divide the human body into four regions: LeftArm (left arm region), RightArm (right arm region), LeftLeg (left leg region), and RightLeg (right leg region), calculate the collaborative space vector representing the limb in each region to describe the movement of the four limbs and reflect the integrity and coordination;

[0012] There are a total of four joint points in the motion collaborative space vector of the LeftArm region: LeftShoulder (LS left shoulder), LeftElbow (LE left elbow), LeftWrist (LW left wrist), and LeftHand (LH left hand). Taking the Spine point as the center, the states of each joint point are represented by vectors denoted as, and are concatenated to form a motion collaborative space vector

[0013] Further, respectively represent the collaborative space vectors describing the movement state of the left hand in the front and rear two frames of images. Taking the Spine point as the origin coordinate, the space formed by the two space vectors pointing to the same bone point between two adjacent frames of images is defined as the space vector amplitude S, and the included angle between the two space vectors is θ;

[0014] When θ ∈ (0, 90°], the formula is as follows:

[0015]

[0016] When the angle change is greater than 90°, it indicates that the movement amplitude of this part is larger. In order to keep S positively correlated with the movement amplitude, when θ ∈ (90°, 180°], the formula is as follows:

[0017]

[0018] Process the n-frame images of the LeftArm area in a complete skeletal diagram action sequence in turn to obtain The spatial vector amplitude value, and the formula is as follows:

[0019]

[0020] Among them, S XOY 、S YOZ 、S XOZ represent the spatial vector amplitude values projected onto three planes; similarly, calculate the spatial vector amplitude value S of LE 、S LW 、S LH ;

[0021] Then, after normalizing the values in the LeftArm area, the motion state measurement coefficient W is obtained, and the formula is as follows:

[0022]

[0023] Similarly, calculate the motion state measurement coefficients W of the spatial vectors of LE 、W LW and W LH ;

[0024] By multiplying the spatial vectors of each joint point by the corresponding motion state measurement coefficients of each joint point and then summing them up, the motion coordination spatial vector of the LeftArm area is obtained, and the formula is as follows:

[0025]

[0026] Similarly, calculate the motion coordination spatial vectors of the regions RightArm, LeftLeg, and RightLeg and and calculate the corresponding motion state measurement coefficients W LA 、W RA 、W LL and W RL .

[0027] Calculate the motion coordination space vector of the entire human skeleton The formula is as follows:

[0028]

[0029] Based on the motion coordination space vectors of n frames of bones, establish a motion coordination space feature sequence set (x1, x2, x i , … x n ), where X i is the motion coordination space vector of the i-th frame

[0030] Furthermore, the bone self-attention sub-network includes:

[0031] After the initial bone sequence extracts the motion coordination space feature sequence, the input received by the neural network is a sequence of vectors of various sizes, and there are spatio-temporal connections between different vectors. However, in actual training, the relationship between these inputs cannot be fully utilized, resulting in a poor model training result. To solve the problem that the fully connected neural network cannot establish a correlation for multiple related inputs, the self-attention mechanism is used. The self-attention mechanism updates each component of the sequence by aggregating global information from the complete input sequence;

[0032] Use to represent the motion coordination space feature sequence of n frames of bone movements (x1, x2, x i , … x n ), where d represents the embedding dimension of each vector; the goal of the self-attention mechanism is to capture the interaction between n vectors by encoding each vector according to the global information; first, define three learnable weight matrices Q and where d q = d k , to get Q = XW Q , K = XW K and V = XW V , and the output is:

[0033]

[0034] Secondly, in order to encapsulate multiple complex relationships between different vectors in the sequence, the multi-head attention mechanism includes h self-attention modules, each module has its own set of learnable weight matrices {W Qi , W Ki , W Vi}, i = 0, 1, …, h-1. For the input X, secondly, the outputs of the h self-attention modules in the multi-head attention mechanism are concatenated into a matrix and projected onto the weight matrix

[0035] Next, the motion coordination spatial feature sequence embeds positional encoding during input to retain the temporal information of the sequence. After performing the multi-Head Attention operation, it is connected to a fully-connected feed-forward network, and the same operation is performed on each vector respectively, including two linear transformations and a ReLU activation output:

[0036] FFN(x) = max(0, xW1 + b1)W2 + b2 (8)

[0037] where W1 and W2 are weights, and b1 and b2 are biases;

[0038] Finally, the skeleton self-attention sub-network is composed of multiple identical multi-head attention modules and feed-forward neural network modules stacked in combination, and finally connected to a softmax layer to obtain the classification result.

[0039] S2. Collect depth image data, extract DMM, train through a depth self-attention sub-network integrating channel attention mechanism, and classify through a Softmax classifier;

[0040] Yang et al. proposed using DMM (Depth Motion Maps) to describe the 3D structure and shape information of behaviors. To utilize the additional body shape and motion information in the depth map, each depth frame is projected onto three orthogonal Cartesian planes; then, the region of interest of each projected map is set as the bounding box of the foreground (non-zero) region and further normalized to a fixed size; this normalization can reduce the internal variations among different subjects when performing the same action, such as subject height and motion range. Therefore, each 3D depth frame generates three 2D maps according to the front view, side view, and top view, namely map f , map s , map t ; for each projected map, its motion energy is obtained by calculating the difference between two consecutive mapped maps and thresholding. The binary map of the motion energy represents the motion region or the position where the motion occurs in each time interval, and finally, the DMM is generated by superimposing the motion energy over the entire video sequence v ;

[0041]

[0042] where map v i and map v i-1 are two adjacent frames of images;

[0043] After extracting the DMM, it is input into the deep self-attention sub-network, which consists of a three-channel convolutional neural network and a channel attention mapping module;

[0044] Furthermore, the network model of the deep self-attention sub-network first is the input layer, and the input is the DMM reduced to 224×224. The three channels respectively input the front view, side view, and top view. In the middle is the training layer: after the convolutional layer, a pooling layer is connected, and this is stacked 4 times in total. The size and number of convolutional kernels are 7×7×32, 5×5×64, 3×3×128, and 3×3×256 in sequence. The stride of the convolutional layer is 1, and the convolutional kernel padding algorithm is adopted to avoid discarding the feature map information during the convolution process; a 2×2 pooling kernel is selected for max pooling, and the stride of the pooling layer is 2 to avoid the blurring effect of average pooling;

[0045] Furthermore, in order to measure the contribution of the features of the three channels to the key information, a channel attention mapping module is added to the network to improve the network's feature extraction ability;

[0046] The main components of the attention mapping module are dimension compression, excitation, and weighting. First, the global average pooling operation is used to turn each two-dimensional feature channel into a real number, and then the fully connected operation and activation functions (ReLU, Sigmoid) are used to obtain a more comprehensive channel-level weight relationship. Finally, element-wise multiplication is used to fuse the obtained weights with the original features.

[0047] The last two layers are fully connected layers. The first fully connected layer produces an output of 1024 neuron nodes (FC-1024), and the second fully connected layer produces an output of 20 neuron nodes (FC-20); at the same time, the rectified linear unit: ReLU activation function is adopted for both the convolutional layer and the fully connected layer to accelerate network training and improve the CNN feature learning ability; and Dropout is added after the ReLU function of the first fully connected layer to enhance the model's generalization ability and prevent overfitting. Finally, Softmax is connected to obtain the classification result.

[0048] S3. Fuse the skeleton sequence classification based on the skeleton self-attention sub-network and the depth image classification based on the deep self-attention sub-network to obtain the fused classification result;

[0049] The fusion method adopted here is the maximum value fusion method in the backend fusion method, that is, the classifier output scores of different modality data are fused, and the maximum value of the output scores of the two sub-networks is retained as the final classification result.

[0050] The beneficial effects of the present invention:

[0051] 1. A motion cooperation spatial feature that can synthesize information of each joint point is proposed, and a quantization standard for the contribution degree of each joint point participating in the motion is proposed, which can effectively reflect the integrity and cooperation of human motion;

[0052] 2. Based on the idea of multi-modal fusion, the bone self-attention sub-network based on bone data is combined with the depth self-attention sub-network based on depth data to make the data more complete, realize the complementarity of various heterogeneous information, eliminate the redundancy between modalities, and establish a new behavior recognition method system, providing new ideas and theoretical basis for the research and application of human behavior recognition methods, and achieving good results on the NTU dataset and the UTD-MHAD dataset. Brief Description of the Drawings

[0053] Figure 1 is the structural block diagram of the motion behavior recognition method based on the multi-modal cooperation self-attention network of the present invention;

[0054] Figure 2 is the schematic diagram of the human bone region of the present invention;

[0055] Figure 3 is the left upper limb cooperation space vector diagram of the present invention;

[0056] Figure 4 is the schematic diagram of the cooperation vector of two consecutive frames of images of a certain bone point of the left upper limb of the present invention;

[0057] Figure 5 is the schematic diagram of the SE Module attention module of the present invention. Detailed Embodiment

[0058] The present invention will be further described below with reference to the drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, so it only shows the components related to the present invention.

[0059] The present invention is tested on the multi-modal action datasets NTU RGB+D 60 and UTD-MHAD. All samples in UTD-MHAD are classified simultaneously. The samples of subject 1, 3, 5, 7 are used for training, and the samples of subject 2, 4, 6, 8 are used for testing.

[0060] Two experimental settings are adopted for the NTU RGB+D 60 dataset:

[0061] In cross-subject evaluation, 40 subjects were divided into a training group and a test group, with 20 subjects in each group; for this evaluation, the training set and the test set had 40320 and 16560 samples respectively. The IDs of the training subjects in this evaluation were: 1, 2, 4, 5, 8, 9, 13, 14, 15, 16, 17, 18, 19, 25, 27, 28, 31, 34, 35, 38; the remaining subjects were left for testing.

[0062] For cross-view evaluation, all samples of camera 1 were selected for testing, and samples of cameras 2 and 3 were selected for training. The training set included the front view and two side views of the action, while the test set included the left and right 45-degree views of the action performance. The training set and the test set had 37920 and 18960 samples respectively.

[0063] As Figure 1 shown, a motion behavior recognition method based on a multi-modal collaborative self-attention network includes the following steps:

[0064] S1. Collect bone infrared sequence data through an inertial sensor; use the infrared sequence data and a collaborative space vector algorithm to construct a motion collaborative space feature sequence set; input the motion collaborative space feature sequence set into the self-attention sub-network of the Transfomer network;

[0065] To evaluate the effects of different methods for extracting motion collaborative space features, three different models were established and compared on UTD-MHAD and NTU RGB+D: Model 1 was the motion collaborative space features after projection onto three Cartesian planes without using a motion state measurement coefficient; Model 2 was the motion collaborative space features that directly calculated the motion state measurement coefficient without projecting the motion collaborative space vector onto three Cartesian planes; Model 3 was the motion collaborative space features of the present invention after processing with a motion state measurement coefficient and projection onto three Cartesian planes; Model 1 and Model 3, Model 2 and Model 3 were two groups of control groups;

[0066] Table 1 Comparison of motion collaborative space feature models

[0067]

[0068] It can be seen from Table 1 that the recognition rate of Model 3 on UTD-MHAD increased by 8% compared to Model 1, and the recognition rates on NTU RGB+D increased by 4.2% and 7.1% respectively; the introduction of the motion state measurement coefficient enabled the model to have a clear discrimination criterion for the contributions of each part to the motion, provided a weight ratio discrimination basis for the extraction and splicing of the motion collaborative space vectors in each region, and improved the recognition effect.

[0069] The recognition rate of Model 3 on UTD-MHAD increased by 16.7% compared to Model 2, and the recognition rates on NTU RGB+D increased by 16.3% and 17.5% respectively; the processing of projecting onto three Cartesian planes strengthened the model's extraction of spatial information and significantly improved the recognition effect.

[0070] It can be seen that the introduction of the motion state measurement coefficient and the setting of projecting onto three Cartesian planes both effectively improved the recognition effect of the motion collaborative spatial features. Therefore, Model 3 will be used in the subsequent experiments of the present invention to extract the motion collaborative spatial features.

[0071] S2. Collect depth image data, extract DMM, train through the depth self-attention sub-network integrating the channel attention mechanism, and classify through the Softmax classifier;

[0072] In order to prove that the depth sub-attention mapping module effectively improves the performance of the depth self-attention sub-network, the present invention designed an ablation experiment to prove its effectiveness;

[0073] Table 2 Ablation experiment of the channel attention mapping module

[0074]

[0075]

[0076] As can be seen from Table 2, after adding the channel attention mapping module, the recognition rate on UTD-MHAD increased by 0.6%, and the recognition rates on NTU RGB+D increased by 0.3% and 1.4% respectively; it can be seen that the channel attention mapping module plays a good role in adjusting the attention to the data of each channel and effectively improves the recognition effect.

[0077] S3. Fuse the classification of the skeleton sequence based on the skeleton self-attention sub-network and the depth image classification based on the depth self-attention sub-network to obtain the fused classification result;

[0078] In order to evaluate the influence of modality fusion on the model effect, the results of using the skeleton self-attention sub-network alone, the depth self-attention sub-network alone, and after modality fusion were compared, as shown in the table:

[0079] Table 3 Comparison of single-modal and multi-modal methods (UTD-MHAD)

[0080]

[0081] Table 4 Comparison of single-modal and multi-modal methods (NTU RGB+D)

[0082]

[0083] As can be seen from Table 3, the recognition rates of the fused multi-modal collaborative self-attention network on UTD-MHAD are increased by 1.2% and 2.4% respectively compared with the two sub-networks, and the recognition rates on NTU RGB+D are increased by 3.8% and 4.5% on average, proving that the information provided by skeletal data and depth data is complementary, and the multi-modal data can describe human behaviors more accurately after fusion.

[0084] The present invention is also horizontally compared with mainstream methods on NTU RGB+D and UTD-MHAD, and the results are shown in Tables 5 and 6:

[0085] Table 5 Comparison of the accuracies of various methods on UTD-MHAD

[0086]

[0087] As can be seen from Table 5, the recognition rate of the method of the present invention on the UTD-MHAD dataset reaches 93.9%, which is higher than the other methods listed, and the recognition rate is increased by at least 1.1%. The evaluation results prove the advanced performance of this method on the UTD-MHAD dataset;

[0088] Table 6 Comparison of the accuracies of various methods on NTU RGB+D

[0089]

[0090] As can be seen from Table 6, the CS recognition rate of the method of the present invention on the NTU RGB+D dataset reaches 90.5%, which is higher than the other methods listed, and the CV recognition rate reaches 94.7%, which is equivalent to or better than the recognition rates of other methods. The evaluation results prove the superiority of this method.

[0091] Taking the ideal embodiments of the present invention described above as an inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A method for recognizing motion behaviors based on a multi-modal collaborative self-attention network, characterized in that, It includes the following steps: S1. Construct a motion cooperation spatial feature sequence by using bone sequence data and a motion cooperation spatial vector algorithm; input the motion cooperation spatial feature sequence into a bone self-attention sub-network based on the Transfomer network and classify it through a Softmax classifier; The motion cooperation spatial vector algorithm includes the following steps: Divide the human body bones into four regions: LeftArm, RightArm, LeftLeg, and RightLeg; The LeftArm region has four joint points: Left Shoulder (LS), Left Elbow (LE), Left Wrist (LW), and Left Hand (LH); The calculation of the spatial vector amplitude of LH includes: When θ ∈ (0, 90°]: When θ ∈ (90°, 180°]: Among them, respectively represent the collaborative spatial vectors describing the motion states of the left hand in the front and rear two frames of images; Process the n-frame images in the LeftArm area sequentially to obtain The magnitude of the spatial vector, and the formula is as follows: Among them, S XOY , S YOZ , S XOZ represent the magnitudes of the spatial vectors after projecting onto three planes; LS space vector The calculation formula for the motion state measurement coefficient is as follows: Among them, S LS , S LE , S LW , S LH are the spatial vector amplitudes of , and W LS , W LE , W LW and W LH are the corresponding motion state measurement coefficients; The calculation formula for the motion cooperation spatial vector of the LeftArm region is: According to the motion coordination space vectors of the four regions and and the motion state measurement coefficients W LA 、W RA 、W LL and W RL calculate the motion coordination space vectors of the entire human body skeleton, and establish a motion coordination space feature sequence (x1, x2, x i ,…x n ), where x i is the motion coordination space vector of the i-th frame; S2. Collect depth image data, extract DMM, train it through a depth self-attention sub-network integrating a channel attention mechanism, and classify it through a Softmax classifier; S3. Integrate the bone sequence classification based on the bone self-attention sub-network and the depth image classification based on the depth self-attention sub-network to obtain the fused classification result.

2. The motion behavior recognition method based on the multi-modal collaborative self-attention network according to claim 1, characterized in that The actions of the bone sequence data include: sliding the right arm to the left, sliding the right arm to the right, waving the right hand, clapping both hands in front, throwing the right arm, crossing the arms in front of the chest, basketball shooting, drawing an x with the right hand, drawing a circle with the right hand, drawing a triangle, bowling, front boxing, right swing in baseball, right forehand swing in tennis, spinning the arm, tennis serving, pushing with both hands, knocking on the door with the right hand, grasping an object with the right hand, picking up / throwing with the right hand, jogging in place, walking in place, sitting to stand, standing to sit, forward lunge, squatting.

3. The motion behavior recognition method based on the multi-modal collaborative self-attention network according to claim 1, characterized in that The construction of the bone self-attention sub-network includes the following steps: First, define three learnable weight matrices and to obtain Q = XW Q , K = XW K and V = XW V ; For the motion coordination spatial feature sequence (x1, x2, x i , … x n ), where d represents the embedding dimension of each vector; Secondly, connect the outputs of the h self-attention modules into a matrix Secondly, when the motion cooperation spatial feature sequence X is input, it is embedded with position encoding to retain the temporal information of the sequence. After performing the multi-HeadAttention operation, it is connected to a fully connected feed-forward network; Finally, stack multiple identical multi-head attention modules and feed-forward neural network modules, and finally connect the softmax layer to obtain the classification result.

4. The motion behavior recognition method based on a multi-modal collaborative self-attention network according to claim 1, wherein, The structure of the depth self-attention sub-network is: a convolutional layer is connected to a pooling layer, which is stacked 4 times in total. The size and number of convolutional kernels are 7×7×32, 5×5×64, 3×3×128, and 3×3×256 in sequence. The stride of the convolutional layer is 1, and the convolutional kernel padding algorithm is adopted; select a 2×2 pooling kernel for max pooling, and the stride of the pooling layer is 2.

5. The motion behavior recognition method based on the multi-modal collaborative self-attention network according to claim 1, characterized in that: The components of the channel attention mechanism are dimension compression, excitation, and weighting. First, use global average pooling operation to turn each two-dimensional feature channel into a real number, and then use a fully connected operation and the ReLU activation function to obtain the channel-level weight relationship; finally, use element-wise multiplication to fuse the obtained weight with the original feature.

Citation Information

Patent Citations

  • Human body behavior recognition method and system based on human body skeleton

    CN111950485A

  • Behavior recognition method and system based on human skeleton

    CN114373225A