Human action recognition method driven by multi-scale spatiotemporal cross attention

Through a multi-scale spatiotemporal cross-attention network, combined with multi-scale spatiotemporal cross-graph convolution and dynamic multi-scale time convolution, the problem of under-exploration of space-time dependence in the existing technology is solved, and higher accuracy and robustness of action recognition are achieved.

CN119810926BActive Publication Date: 2025-06-06JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510286585.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-06
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing skeleton-based human behavior recognition technology ignores the dependence between time and space, and it is difficult to effectively capture the timing changes of actions, resulting in insufficient recognition accuracy.

Method used

A multi-scale space-time cross-attention-driven human behavior recognition method is proposed. Through multi-scale space-time cross-graph convolution and dynamic multi-scale time convolution, the complex relationship between the time dimension and the spatial dimension is captured, and the spatial feature representation is fused to improve the recognition accuracy.

Benefits of technology

Through multi-scale spatiotemporal interaction, we can more effectively model spatiotemporal relationships, improve the accuracy and robustness of action recognition, and enhance the ability to recognize timing changes and subtle action changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810926B_ABST
    Figure CN119810926B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-scale spatiotemporal cross-attention driven human behavior recognition method, the method comprising: taking joint mode, skeletal mode, joint motion mode and skeletal motion mode as input modes, inputting multi-scale spatiotemporal cross graph convolution, and obtaining the final spatiotemporal feature representation; performing residual connection on the final spatiotemporal feature representation and the input mode, obtaining the residual connected features, and then inputting the residual connected features into dynamic multi-scale time convolution; in the dynamic multi-scale time convolution, the residual connected features are processed by dynamic multi-scale time convolution to obtain the output feature representation of dynamic multi-scale time convolution; based on the final spatiotemporal feature representation and the output feature representation of dynamic multi-scale time convolution, the prediction results of each modal branch are obtained; based on the prediction results of each modal branch, the final prediction category is obtained. The present invention proposes an efficient computing framework, which reduces the computing cost by optimizing the network structure and reasoning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a multi-scale spatiotemporal cross-attention driven human behavior recognition method. Background Art

[0002] As an important research topic in the field of computer vision, human action recognition has been widely used in intelligent monitoring, health detection, human-computer interaction and other fields. Compared with recognition methods based on RGB images or optical flow, skeleton data directly reflects human posture and motion information, which makes it more robust when dealing with changes in camera perspective and fluctuations in video appearance. At the same time, the development of low-cost depth sensors and efficient posture estimation algorithms has provided a solid foundation for skeleton-based action recognition applications.

[0003] However, due to the time dependency of actions, highly complex human postures, and similarities between different actions, existing skeleton-based human action recognition technologies have the following problems: First, existing methods ignore the dependency between time and space, and usually simply concatenate spatial convolution and temporal convolution, but fail to fully explore the complex relationship between the temporal and spatial dimensions. This processing method will lead to limitations in the model in capturing spatiotemporal dependencies. Second, most current methods focus on static feature extraction, which makes it difficult to effectively capture the temporal changes of actions, resulting in insufficient use of temporal information and affecting recognition accuracy. Summary of the invention

[0004] In view of the above situation, the main purpose of the present invention is to propose a multi-scale spatiotemporal cross-attention driven human behavior recognition method to solve the above technical problems.

[0005] The present invention proposes a multi-scale spatiotemporal cross-attention driven human behavior recognition method, which comprises the following steps:

[0006] Step 1: Take the joint mode, the skeleton mode, the joint motion mode and the skeleton motion mode as input modes and input them into the multi-scale spatiotemporal cross attention network framework, wherein the joint mode branch, the skeleton mode branch, the joint motion mode branch and the skeleton motion mode branch have the same working principle, and the multi-scale spatiotemporal cross attention network framework is composed of multi-scale spatiotemporal cross graph convolution (MSTC-GC) and dynamic multi-scale temporal convolution (DTCN);

[0007] Step 2: In the multi-scale spatiotemporal cross graph convolution, the input modality is processed through the multi-scale spatial axis convolution branch to obtain a spatial feature map;

[0008] The input modality is processed through a multi-scale time axis convolution branch to obtain a temporal feature map;

[0009] Based on the spatial feature map and the temporal feature map, through multi-head cross attention calculation, the fused feature map along the spatial axis and the fused feature map along the temporal axis are obtained respectively;

[0010] Apply graph convolution operations to the fused feature maps along the spatial axis to obtain high-level spatial feature representation;

[0011] Based on the high-level spatial feature representation and the fused feature map along the time axis, the final spatiotemporal feature representation is obtained;

[0012] Step 3: Perform residual connection on the final spatiotemporal feature representation and the input modality to obtain the residual connected features, and then input the residual connected features into dynamic multi-scale temporal convolution;

[0013] In the dynamic multi-scale temporal convolution, the features after residual connection are processed by the four branches of the dynamic multi-scale temporal convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch;

[0014] The output features of the first branch, the second branch, the third branch, and the fourth branch are fused to obtain the output feature representation of dynamic multi-scale temporal convolution;

[0015] Step 4: Fuse the final spatiotemporal feature representation with the output feature representation of the dynamic multi-scale temporal convolution to obtain a fused feature representation;

[0016] Obtain prediction results based on the fused feature representation;

[0017] Step 5, through steps 1 to 4, respectively obtain the prediction results of the joint modality branch, the prediction results of the bone modality branch, the prediction results of the joint motion modality branch, and the prediction results of the bone motion modality branch;

[0018] Integrate the prediction results of the joint modality branch, the prediction results of the bone modality branch, the prediction results of the joint motion modality branch, and the prediction results of the bone motion modality branch to obtain the final prediction category;

[0019] The target is identified based on the final predicted category.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. The present invention uses multi-scale spatiotemporal interaction to more effectively model the relationship between the time dimension and the space dimension, thereby better capturing the potential dependencies between time and space. It improves the accuracy of action recognition and the robustness when processing complex actions;

[0022] 2. The present invention adopts a new dynamic convolution method to capture the dynamic changes in each time period, improve the model's sensitivity to timing changes, and further enhance the ability to recognize subtle motion changes.

[0023] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A flowchart of the multi-scale spatiotemporal cross-attention driven human behavior recognition method proposed in the present invention;

[0025] Figure 2 The overall network framework diagram of the multi-scale spatiotemporal cross-attention driven human behavior recognition method proposed in the present invention;

[0026] Figure 3 A schematic diagram of a multi-scale spatiotemporal cross-graph convolution of a human behavior recognition method driven by a multi-scale spatiotemporal cross-attention proposed in the present invention;

[0027] Figure 4 Schematic diagram of dynamic multi-scale temporal convolution for the multi-scale spatiotemporal cross-attention driven human behavior recognition method proposed in the present invention. DETAILED DESCRIPTION

[0028] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0029] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0030] See also Figure 1 and Figure 2 The embodiment of the present invention proposes a multi-scale spatiotemporal cross-attention driven human behavior recognition method, which includes the following steps:

[0031] Step 1. Take the joint modality, bone modality, joint motion modality and bone motion modality as input modalities and input them into the multi-scale spatiotemporal cross-attention network framework, wherein the joint modality branch, the bone modality branch, the joint motion modality branch and the bone motion modality branch have the same working principles, and the multi-scale spatiotemporal cross-attention network framework is composed of multi-scale spatiotemporal cross graph convolution and dynamic multi-scale time convolution.

[0032] Specifically, multi-scale spatiotemporal cross-graph convolution is used as the main module of the network to capture the potential dependencies between feature time and space. Dynamic multi-scale temporal convolution is an auxiliary module of the network to dynamically model the action sequence and enhance the ability to recognize subtle actions.

[0033] Step 2: In the multi-scale spatiotemporal cross graph convolution, the input modality is processed through the multi-scale spatial axis convolution branch to obtain a spatial feature map;

[0034] The input modality is processed through a multi-scale time axis convolution branch to obtain a temporal feature map;

[0035] Based on the spatial feature map and the temporal feature map, through multi-head cross attention calculation, the fused feature map along the spatial axis and the fused feature map along the temporal axis are obtained respectively;

[0036] Apply graph convolution operations to the fused feature maps along the spatial axis to obtain high-level spatial feature representation;

[0037] Based on the high-level spatial feature representation and the fused feature maps along the time axis, the final spatiotemporal feature representation is obtained.

[0038] See also Figure 3 In step 2, in the multi-scale spatial axis convolution branch, 1×1, 3×1, and 5×1 convolution kernels are used to perform convolution operations on the input modality along the spatial dimension. After all convolution operations are completed, the convolution results of multiple scales are weighted summed through the 1×1 convolution layer to obtain and output the spatial feature map. The corresponding process has the following relationship:

[0039] ;

[0040] in, represents the spatial feature map, It means after 1×1 convolution processing, Indicates that after Scaled 1D convolution processing, It means that after layer normalization, Indicates input mode;

[0041] In the multi-scale time axis convolution branch, the input modality is convolved along the time dimension using 1×1, 3×1, and 5×1 convolution kernels, and then the convolution results of multiple scales are fused to obtain the time feature map. The relationship between the corresponding process is:

[0042] ;

[0043] in, Represents a time feature graph;

[0044] For the spatial axis, the temporal feature map is used as the query, and the spatial feature map is used as the key and value. The fused feature map along the spatial axis is obtained by multi-head cross attention calculation. For the spatial axis, the spatial feature map is used as the query, and the temporal feature map is used as the key and value. The fused feature map along the temporal axis is obtained by multi-head cross attention calculation. The relationship between the corresponding process is:

[0045] ;

[0046] in, represents the fused feature map along the spatial axis, represents the fused feature map along the time axis, represents the multi-head cross-attention computation operation along the spatial axis, Represents the multi-head cross-attention calculation operation along the time axis;

[0047] Based on the fused feature graph and adjacency matrix along the spatial axis, a graph convolutional network is used for feature propagation to obtain a high-level spatial feature representation. The corresponding process has the following relationship:

[0048] ;

[0049] in, represents a high-level spatial feature representation, It means that it has been processed by the activation function. represents the normalized adjacency matrix, represents the weight matrix of graph convolution, represents the degree matrix, represents the adjacency matrix;

[0050] Based on the high-level spatial feature representation and the fused feature map along the time axis, the final spatiotemporal feature representation is obtained. The relationship between the corresponding process is:

[0051] ;

[0052] in, represents the final spatiotemporal feature representation, Indicates that a splicing operation has been performed.

[0053] Step 3: Perform residual connection on the final spatiotemporal feature representation and the input modality to obtain the residual connected features, and then input the residual connected features into dynamic multi-scale temporal convolution;

[0054] In the dynamic multi-scale temporal convolution, the features after residual connection are processed by the four branches of the dynamic multi-scale temporal convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch;

[0055] The output features of the first branch, the second branch, the third branch, and the fourth branch are fused to obtain the output feature representation of dynamic multi-scale temporal convolution;

[0056] See also Figure 4 In step 3, the features after residual connection are processed by four branches of dynamic multi-scale time convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch. In the first branch, the time-varying weights are generated by the node internal feature representation processed by the one-dimensional convolutional neural network and the node inter-feature representation processed by the one-dimensional convolutional neural network. The steps for obtaining the node internal feature representation processed by the one-dimensional convolutional neural network are as follows:

[0057] Apply 3D adaptive average pooling to the skeleton embedding representation to obtain the description vector inside the node. The relationship between the corresponding process is:

[0058] ;

[0059] in, Represents the description vector inside the node, Indicates that after 3D adaptive average pooling processing, represents the skeleton embedding representation;

[0060] Applying a one-dimensional convolutional neural network to the description vector inside the node, we get the internal feature representation of the node after being processed by the one-dimensional convolutional neural network. The relationship between the corresponding process is:

[0061] ;

[0062] in, represents the internal feature representation of the node after being processed by a one-dimensional convolutional neural network. It means that it has been processed by the ReLU activation function. It means that after batch normalization, Indicates that it has been calculated by a 1D convolutional neural network;

[0063] The steps for obtaining the feature representation between nodes after processing by the one-dimensional convolutional neural network are as follows:

[0064] Apply 1D adaptive average pooling to the description vector inside the node for derivation, and obtain the inter-node feature vector after adaptive average pooling. The relationship between the corresponding processes is:

[0065] ;

[0066] in, represents the inter-node feature vector obtained after adaptive average pooling processing, Indicates the derivation process after 1D adaptive average pooling;

[0067] Based on the inter-node feature vector obtained after adaptive average pooling processing, the inter-node feature representation after one-dimensional convolutional neural network processing is obtained through 1D convolutional neural network calculation. The relationship between the corresponding processes is:

[0068] ;

[0069] in, Represents the feature representation between nodes after being processed by a one-dimensional convolutional neural network;

[0070] The relationship in the process of generating time-varying weights is:

[0071] ;

[0072] in, represents the time-varying weight.

[0073] Furthermore, in the second branch, a larger receptive field is achieved by setting the dilation rate instead of a larger convolution kernel;

[0074] In the third branch, a 3×1 max pooling operation is used to reduce the feature dimension to reduce the computational cost;

[0075] In the fourth branch, dynamic multi-scale temporal convolutions are introduced to facilitate training by introducing residual connections;

[0076] Further, in Figure 4 middle, Indicates the number of bone nodes, represents the number of time steps, Indicates the number of input channels.

[0077] Specifically, in this step, the dilation rate of the convolution kernel is dynamically adjusted in the form of adaptive dilated convolution to capture spatiotemporal patterns of different granularities;

[0078] In the spatial dimension, different expansion rates are assigned to different body parts according to the distance distribution between nodes, achieving a balance between local dense perception and global sparse perception;

[0079] In the time dimension, a learnable time gating mechanism is introduced to automatically select the optimal time window size. For example, a short window is used to capture instantaneous changes in fast movements (such as clapping), and a long window is used to model periodic patterns in slow movements (such as walking).

[0080] Step 4: Fuse the final spatiotemporal feature representation with the output feature representation of the dynamic multi-scale temporal convolution to obtain a fused feature representation;

[0081] The prediction results are obtained based on the fused feature representation.

[0082] Step 5, through steps 1 to 4, respectively obtain the prediction results of the joint modality branch, the prediction results of the bone modality branch, the prediction results of the joint motion modality branch, and the prediction results of the bone motion modality branch;

[0083] Integrate the prediction results of the joint modality branch, the prediction results of the bone modality branch, the prediction results of the joint motion modality branch, and the prediction results of the bone motion modality branch to obtain the final prediction category;

[0084] The target is identified based on the final predicted category.

[0085] Furthermore, in this step, in order to improve the adaptability to complex scenes, RGB video streams are introduced for collaborative modeling;

[0086] Specifically, the geometric correspondence between skeleton node coordinates and RGB pixel coordinates is used to generate a spatial alignment matrix to guide the fusion of visual features and skeleton features in key areas. The channel attention mechanism is used to dynamically adjust the contribution weights of different modal features to suppress redundant information. For example, for static postures, the skeleton modality weight is higher; for dynamic actions with rich textures, the RGB modality weight is increased.

[0087] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0088] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0089] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A multi-scale spatiotemporal cross-attention driven human behavior recognition method, characterized in that: The method comprises the following steps: Step 1, taking the joint mode, the bone mode, the joint motion mode and the bone motion mode as input modes, and inputting them into the multi-scale spatiotemporal cross attention network framework, wherein the joint mode branch, the bone mode branch, the joint motion mode branch and the bone motion mode branch have the same working principle, and the multi-scale spatiotemporal cross attention network framework is composed of multi-scale spatiotemporal cross graph convolution and dynamic multi-scale time convolution; Step 2: In the multi-scale spatiotemporal cross graph convolution, the input modality is processed through the multi-scale spatial axis convolution branch to obtain a spatial feature map; The input modality is processed through a multi-scale time axis convolution branch to obtain a temporal feature map; Based on the spatial feature map and the temporal feature map, through multi-head cross attention calculation, the fused feature map along the spatial axis and the fused feature map along the temporal axis are obtained respectively; Apply graph convolution operations to the fused feature maps along the spatial axis to obtain high-level spatial feature representation; Based on the high-level spatial feature representation and the fused feature map along the time axis, the final spatiotemporal feature representation is obtained; Step 3: Perform residual connection on the final spatiotemporal feature representation and the input modality to obtain the residual connected features, and then input the residual connected features into dynamic multi-scale temporal convolution; In the dynamic multi-scale temporal convolution, the features after residual connection are processed by the four branches of the dynamic multi-scale temporal convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch; The output features of the first branch, the second branch, the third branch, and the fourth branch are fused to obtain the output feature representation of dynamic multi-scale temporal convolution; Step 4: Fuse the final spatiotemporal feature representation with the output feature representation of the dynamic multi-scale temporal convolution to obtain a fused feature representation; Obtain prediction results based on the fused feature representation; Step 5, through steps 1 to 4, respectively obtain the prediction results of the joint modality branch, the prediction results of the bone modality branch, the prediction results of the joint motion modality branch, and the prediction results of the bone motion modality branch; Integrate the prediction results of the joint modality branch, the prediction results of the bone modality branch, the prediction results of the joint motion modality branch, and the prediction results of the bone motion modality branch to obtain the final prediction category; The target is identified based on the final predicted category.

2. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 1, characterized in that: In step 2, the input modality is processed by a multi-scale spatial axis convolution branch to obtain a spatial feature map, and the relationship between the corresponding process is: ; in, represents the spatial feature map, It means after 1×1 convolution processing, Indicates that after Scaled 1D convolution processing, It means that after layer normalization, Indicates input modality.

3. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 2, characterized in that: In step 2, the input modality is processed through a multi-scale time axis convolution branch to obtain a time feature map, and the relationship between the corresponding process is: ; in, Represents a time feature graph.

4. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 3, characterized in that: In step 2, based on the spatial feature map and the temporal feature map, a fused feature map along the spatial axis and a fused feature map along the temporal axis are obtained through multi-head cross attention calculation, and the relationship between the corresponding processes is: ; in, represents the fused feature map along the spatial axis, represents the fused feature map along the time axis, represents the multi-head cross-attention computation operation along the spatial axis, Represents the multi-head cross-attention calculation operation along the time axis.

5. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 4, characterized in that: In step 2, a graph convolution operation is performed on the fused feature map along the spatial axis to obtain a high-level spatial feature representation. The relationship between the corresponding process is: ; in, represents a high-level spatial feature representation, It means that it has been processed by the activation function. represents the normalized adjacency matrix, represents the weight matrix of graph convolution, represents the degree matrix, Represents the adjacency matrix.

6. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 5, characterized in that: In step 2, based on the high-level spatial feature representation and the fused feature map along the time axis, the final spatiotemporal feature representation is obtained, and the relationship between the corresponding process is: ; in, represents the final spatiotemporal feature representation, Indicates that a splicing operation has been performed.

7. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 6, characterized in that: In the step 3, the features after residual connection are processed by four branches of dynamic multi-scale time convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch. In the first branch, the time-varying weights are generated by the node internal feature representation processed by the one-dimensional convolutional neural network and the node inter-feature representation processed by the one-dimensional convolutional neural network. The steps for obtaining the node internal feature representation after the one-dimensional convolutional neural network are as follows: Apply 3D adaptive average pooling to the skeleton embedding representation to obtain the description vector inside the node. The relationship between the corresponding process is: ; in, Represents the description vector inside the node, Indicates that after 3D adaptive average pooling processing, represents the skeleton embedding representation; Applying a one-dimensional convolutional neural network to the description vector inside the node, we get the internal feature representation of the node after being processed by the one-dimensional convolutional neural network. The relationship between the corresponding process is: ; in, represents the internal feature representation of the node after being processed by a one-dimensional convolutional neural network. It means that it has been processed by the ReLU activation function. It means that after batch normalization, Indicates that it is calculated by 1D convolutional neural network.

8. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 7, characterized in that: In the step 3, the features after residual connection are processed by four branches of dynamic multi-scale time convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch. In the first branch, the time-varying weights are generated by the node internal feature representation processed by the one-dimensional convolutional neural network and the node inter-node feature representation processed by the one-dimensional convolutional neural network. The steps for obtaining the node inter-node feature representation processed by the one-dimensional convolutional neural network are as follows: Apply 1D adaptive average pooling to the description vector inside the node for derivation, and obtain the inter-node feature vector after adaptive average pooling. The relationship between the corresponding processes is: ; in, represents the inter-node feature vector obtained after adaptive average pooling processing, Indicates the derivation process after 1D adaptive average pooling; Based on the inter-node feature vector obtained after adaptive average pooling, the inter-node feature representation after one-dimensional convolutional neural network processing is obtained through 1D convolutional neural network calculation. The relationship between the corresponding processes is: ; in, Represents the node-to-node feature representation after processing by a one-dimensional convolutional neural network.

9. The multi-scale spatiotemporal cross-attention driven human behavior recognition method according to claim 8, characterized in that: In the step 3, the features after residual connection are processed by four branches of dynamic multi-scale time convolution respectively to obtain the output features of the first branch, the output features of the second branch, the output features of the third branch and the output features of the fourth branch. In the first branch, the time-varying weight is generated by the node internal feature representation after processing by the one-dimensional convolutional neural network and the node inter-feature representation after processing by the one-dimensional convolutional neural network. The relationship in the process of generating the time-varying weight is: ; in, represents the time-varying weight.

Citation Information

Patent Citations

  • Human action recognition method based on graph convolutional neural network

    CN112633209A

  • Skeleton action recognition method based on space-time adaptive feature fusion graph convolutional network

    CN116665300A