Human body action recognition method and system based on skeleton sequence

Through the improved CTR-GCN network, combining absolute and relative position coding, multi-scale adjacency matrix and multi-branch structure, the existing models have insufficient utilization and excessive smoothing of physiological laws, achieving higher action recognition accuracy and fine-grained classification effects.

CN120564253APending Publication Date: 2025-08-29NANJING MAX DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446355.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing human action recognition model based on bone sequences does not fully utilize the physiological prior knowledge of human skeletal structure, resulting in a lack of effective modeling of physiological laws when understanding complex actions, and there are excessive smoothing problems, affecting the recognition accuracy.

Method used

The space-time graph convolution neural network and contrast learning are used to improve the CTR-GCN network, and the adjacency matrix is ​​reshaped by combining absolute position coding and relative position coding with multi-scale spatial graphs, topology enhancement module is built, and a multi-branch structure is used for timing feature extraction, combining residual jump connection and contrast normalization layer to alleviate the problem of node oversmoothing.

Benefits of technology

It improves the recognition accuracy of complex actions, significantly improves the recognition accuracy of cross-view angles and cross-subject tests, and enhances the fine-grained classification ability of actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564253A_ABST
    Figure CN120564253A_ABST
Patent Text Reader

Abstract

The invention discloses a human body action recognition method and system based on a skeleton sequence, and the method comprises the steps: obtaining a human body action data set, and constructing a skeleton graph based on the data set; the CTR-GCN network is improved based on a space-time diagram convolutional neural network and comparative learning, and a preliminary recognition model is constructed; the improvement comprises performing absolute position coding and relative position coding of skeleton points in the backbone network, and adopting a multi-scale space graph to remodel an adjacent matrix; training the preliminary recognition model by using the skeleton graph, and updating model parameters to obtain a human body action recognition model; performing action recognition on a to-be-recognized target by using the human body action recognition model; according to the method provided by the invention, the performance of human motion recognition for the skeleton sequence data can be improved, and a relatively good recognition effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human motion recognition method and system, in particular to a human motion recognition method and system based on skeleton sequences, and belongs to the technical field of computer vision and artificial intelligence. Background Art

[0002] Human action recognition aims to accurately identify and classify different movements by analyzing the body's posture and movement changes over time. By capturing the motion trajectories and posture changes of key human parts (such as joints and limbs), combined with advanced algorithmic models, this technology can extract meaningful motion features from complex dynamic data, enabling the classification and understanding of movements. Due to its non-invasive nature, high accuracy, and wide applicability, human action recognition technology has demonstrated tremendous potential for application in a wide range of fields.

[0003] Currently, human action recognition primarily focuses on data from two modalities: RGB video and skeleton sequences. RGB video-based methods typically treat the video stream as a sequence of image frames and leverage optical flow information to build dynamic models of motion. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are widely used to process this type of data. However, while these methods have improved recognition accuracy to a certain extent, using RGB video as processing data still has many drawbacks, such as large data volumes and susceptibility to factors such as lighting conditions and complex backgrounds. In contrast, skeleton sequence data can overcome these shortcomings. Its sparsity and strong noise resistance make it suitable for scenarios with limited hardware or complex environments. Specifically, a skeleton graph is defined based on human actions, with human joints as nodes and bones as edges connecting two nodes. This representation effectively restores the spatial topological structure of the action and effectively expresses the temporal dependencies between actions in consecutive frames.

[0004] In recent years, with the outstanding performance of graph convolutional neural networks (GCNs) in graph data processing, researchers have attempted to apply them to the task of skeletal sequence-based action recognition, achieving state-of-the-art performance. By modeling skeletal sequences, the model not only learns the spatiotemporal semantics of actions but also captures their spatial topological information, significantly improving the model's expressiveness and robustness. Yan et al. first applied graph convolutional networks (GCNs) to skeletal action recognition and proposed the spatiotemporal graph convolutional network (ST-GCN), which significantly improved the performance of skeletal sequence-based action recognition. Subsequent research has further advanced this field. Li et al. proposed a spatiotemporal graph routing scheme for skeletal-based action recognition, aiming to adaptively learn the inherent high-order implicit relationships between physically separated nodes. This scheme consists of two core components: the spatial graph router (SGR) and the temporal graph router (TGR). The SGR focuses on discovering latent relationships between joints through subclustering in the spatial dimension, while the TGR aims to reveal structural information by measuring the correlation between the temporal joint node trajectories. On the other hand, to reduce model computational costs, Cheng et al. proposed the Shift-GCN framework, which abandons conventional graph convolution operations and adopts shifted graph convolution and lightweight point-by-point convolution operations. Compared with existing methods, the network complexity is reduced by more than 10 times. Song et al. proposed a multi-stream GCN model. This model integrates input features such as joint node positions and motion speed at an early stage. By introducing separable convolutional layers and a compound scaling strategy, it effectively reduces redundant parameters in the model and significantly improves its robustness. Unlike Song et al.'s approach, Li et al. innovatively proposed a symbiotic GCN network designed to simultaneously handle action recognition and motion prediction tasks. This model consists of a backbone network, an action recognition branch, and a motion prediction branch. The two branches work together to achieve mutual promotion and enhancement of action recognition and motion prediction. Recently, Chen et al. proposed a novel topology refinement graph convolutional network (CTR-GC) for modeling dynamic topology and multi-channel features of skeletal sequences. CTR-GC uses a shared topological matrix as common prior knowledge for all channels, and optimizes the topological structure by inferring the correlation of specific channels, thereby achieving refined modeling at the channel level. Based on the information bottleneck theory, Chi et al. proposed the InfoGCN network to learn the action representation with the maximum amount of information, and introduced the attention-based graph convolution operation to infer the skeletal topology related to the context. Pang et al. combined the Transformer with GCN and proposed the Interaction Graph Transformer Network. The interaction skeleton graph is constructed through the interaction semantics and distance correlation between body parts, and the action representation is enhanced by aggregating the information of interactive body parts on the learned graph to better capture the spatiotemporal information of the skeleton sequence. This series of methods provides a variety of novel technical means and research directions for human action recognition based on skeletal data.

[0005] Skeleton sequences are an important data modality for human action recognition. Research on skeletal sequence-based human action recognition has achieved remarkable results, but several challenges remain. The human skeleton contains a wealth of physiological prior knowledge (such as the kinematic properties of joints, skeletal symmetry, and constraints between joints). However, existing skeletal sequence models do not fully exploit this prior knowledge, resulting in a lack of effective modeling of physiological laws when understanding complex actions. Oversmoothing also leads to a loss of feature expression, which in turn affects the accurate recognition of human actions. Summary of the Invention

[0006] Purpose of the invention: The purpose of the present invention is to provide a human motion recognition method and system based on skeleton sequences that can improve recognition accuracy.

[0007] Technical solution: The human motion recognition method based on skeleton sequence described in the present invention comprises:

[0008] (1) Obtain a human motion dataset and construct a skeleton graph based on the dataset;

[0009] (2) Based on the spatiotemporal graph convolutional neural network and contrastive learning, the CTR-GCN network is improved to build a preliminary recognition model; the improvement includes encoding the absolute and relative positions of the skeleton points in the backbone network and reshaping the adjacency matrix using a multi-scale spatial graph;

[0010] (3) Use the skeleton graph to train the preliminary recognition model, update the model parameters, and obtain the human action recognition model;

[0011] (4) Use the human action recognition model to perform action recognition on the target to be identified.

[0012] Furthermore, the skeleton graph uses joints as vertices and bones as edges. Indicates that is a set of N vertices; is a set of edges, through the adjacency matrix Indicates that the element a in the adjacency matrix ij Used to represent the vertex v i and v j The correlation strength between them, i=1,2,…,N; j=1,2,…,N; vertex v i The neighborhood of express; is the feature set of N vertices, through the matrix Indicates that, where c is the number of channels, vertex v i The characteristic representation of

[0013] Furthermore, step (2) performs absolute position encoding and relative position encoding of the skeleton points in the backbone network, and reshapes the adjacency matrix using a multi-scale spatial graph, including: performing absolute position encoding of the skeleton points; using five topological enhancement modules with shared parameters in the backbone network to extract the correlation between human joints in parallel, and performing relative position encoding of the skeleton points at the same time;

[0014] The topology enhancement module consists of a multi-scale adjacency matrix, channel refinement features, and relative position encoding. The operation process is as follows:

[0015] First, the input features after absolute position encoding are linearly projected, and then the pairwise differences between feature vectors are calculated in the channel dimension to generate channel-refined features with high-order interactive information; the channel-refined features are element-wise added to the relative position encoding, and weighted aggregation is performed in combination with the multi-scale adjacency matrix to construct an enhanced topological structure representation; based on the enhanced topological relationship, graph convolution is used to extract high-order spatial features.

[0016] Furthermore, the improvement in step (2) also includes: adopting a multi-branch structure to extract temporal features of different scales, the multi-branch structure comprising four parallel branches, wherein the first branch realizes the dimensional transformation of features through 1×1 convolution and skips the connection to retain the original information; the second and third branches extract features of different temporal receptive fields through 1×1 convolution and 5×1 convolution respectively; the fourth branch combines 1×1 convolution with 3×1 maximum pooling operation to capture local short-term temporal information, and finally, the outputs of the four branches are spliced ​​in the channel dimension to form a fused output feature.

[0017] Based on the same inventive concept, the present invention also provides a human motion recognition system based on skeleton sequences, comprising:

[0018] Initialization module, obtains human action dataset, and builds skeleton graph based on the dataset;

[0019] The model building module improves the CTR-GCN network based on spatiotemporal graph convolutional neural networks and contrastive learning to build a preliminary recognition model. The improvements include encoding the absolute and relative positions of skeleton points in the backbone network and reshaping the adjacency matrix using a multi-scale spatial graph.

[0020] The model training module is used to train the preliminary recognition model using the skeleton graph, update the model parameters, and obtain the human action recognition model;

[0021] The recognition module is used to perform action recognition on the target to be recognized using a human action recognition model.

[0022] Furthermore, the skeleton graph uses joints as vertices and bones as edges. Indicates that is a set of N vertices; is a set of edges, through the adjacency matrix Indicates that the element a in the adjacency matrix ij Used to represent the vertex v i and v j The correlation strength between them, i=1,2,…,N; j=1,2,…,N; vertex v i The neighborhood of express; is the feature set of N vertices, through the matrix Indicates that, where c is the number of channels, vertex v i The characteristic representation of

[0023] Furthermore, the model construction module performs absolute position encoding and relative position encoding of skeletal points in the backbone network and reshapes the adjacency matrix using a multi-scale spatial graph, including: encoding the absolute position of skeletal points; using five topological enhancement modules with shared parameters in the backbone network to extract the correlation between human joints in parallel, while performing relative position encoding of skeletal points;

[0024] The topology enhancement module consists of a multi-scale adjacency matrix, channel refinement features, and relative position encoding. The operation process is as follows:

[0025] First, the input features after absolute position encoding are linearly projected, and then the pairwise differences between feature vectors are calculated in the channel dimension to generate channel-refined features with high-order interactive information; the channel-refined features are element-wise added to the relative position encoding, and weighted aggregation is performed in combination with the multi-scale adjacency matrix to construct an enhanced topological structure representation; based on the enhanced topological relationship, graph convolution is used to extract high-order spatial features.

[0026] Furthermore, the improvement of the model construction module also includes: adopting a multi-branch structure to extract temporal features of different scales, the multi-branch structure includes four parallel branches, wherein the first branch realizes the dimensional transformation of features through 1×1 convolution and skips the connection to retain the original information; the second and third branches extract features of different time receptive fields through 1×1 convolution and 5×1 convolution respectively; the fourth branch combines 1×1 convolution with 3×1 maximum pooling operation to capture local short-term temporal information. Finally, the outputs of the four branches are spliced ​​in the channel dimension to form a fused output feature.

[0027] Based on the same inventive concept, the present invention also provides a computing device, comprising: one or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the human motion recognition method based on skeleton sequences according to any one of the above items are implemented.

[0028] Based on the same inventive concept, the present invention also provides a storage medium, which stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the steps of the human motion recognition method based on skeleton sequences according to any one of the above items.

[0029] Beneficial effects: Compared with the existing technology, the present invention realizes multi-level perception of spatial domain features through the dynamic reshaping mechanism of the adjacency matrix combined with the absolute position encoding and relative position relationship encoding of joint nodes; secondly, it adopts a parallel multi-branch temporal convolution structure for time series modeling to effectively capture the multi-scale temporal features of action sequences; in addition, by introducing the residual jump connection mechanism, it realizes the efficient fusion of spatiotemporal features, and innovatively adopts the contrast normalization layer to alleviate the common node over-smoothing problem in graph neural networks; this method improves the recognition effect of complex actions, has significant advantages in the fine-grained classification of similar actions, improves the recognition accuracy of cross-view tests, and improves the recognition accuracy of cross-subject tests. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0031] Figure 2 A schematic diagram of spatial modeling according to an embodiment of the present invention;

[0032] Figure 3 A schematic diagram of time modeling according to an embodiment of the present invention;

[0033] Figure 4 Schematic diagram of a human motion recognition model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0035] As attached Figure 1 As shown, the human action recognition method based on skeleton sequence of this embodiment includes:

[0036] (1) Obtain a human motion dataset and construct a skeleton graph based on the dataset;

[0037] (2) Based on the spatiotemporal graph convolutional neural network and contrastive learning, the CTR-GCN network is improved to build a preliminary recognition model; the improvement includes encoding the absolute and relative positions of the skeleton points in the backbone network and reshaping the adjacency matrix using a multi-scale spatial graph;

[0038] (3) Use the skeleton graph to train the preliminary recognition model, update the model parameters, and obtain the human action recognition model;

[0039] (4) Use the human action recognition model to perform action recognition on the target to be identified.

[0040] Specifically, the dataset in step (1) is: NTU RGB+D 120, a large-scale dataset for RGB+D human action recognition. This dataset contains 120 different action categories, covering multiple categories such as daily activities, human interactions, and health-related actions. It contains a total of 114,480 video samples. They were performed by 106 different subjects. The dataset contains more than 8 million frames of video data.

[0041] The skeleton graph, specifically, a human skeleton can be represented as a graph consisting of joints as vertices and bones as edges. Indicates that is a set of N vertices. is a set of edges, which is represented by an adjacency matrix Indicates that the element a in the adjacency matrix ij Reflects the vertex v i and v j The strength of the correlation between the vertex v i The neighborhood of express. is the feature set of N vertices, which is represented by the matrix Represents that the vertex v i The characteristic representation of

[0042] Step (2) includes: Spatial Modeling: Figure 2 As shown in Figure 1, before inputting into the backbone network, we first perform absolute position encoding (APE) on the skeleton points. The backbone network uses five topology enhancement modules with shared parameters to extract the correlations between human joints in parallel. During this process, relative position encoding (RPE) of the skeleton points is also performed.

[0043] The topology enhancement module consists of three core components: Multi-scale Adjacency Matrix, Channel-refined Features, and Relative Position Encoding. The module's operation flow is as follows:

[0044] Feature transformation and channel refinement: First, linear projection is performed on the input features after absolute position encoding (Absolute Position Encoding), and then pairwise subtraction (Pairwise Subtraction) between feature vectors is calculated in the channel dimension to generate channel refinement features with high-order interaction information.

[0045] Topology enhancement calculation: The above channel refinement features and relative position encoding are added element by element, and weighted aggregation is performed in combination with the multi-scale adjacency matrix to finally construct an enhanced topological structure representation.

[0046] Feature extraction: Based on the enhanced topological relationship, graph convolution is used to further extract high-order spatial features to improve the representation ability of the model.

[0047] This module effectively enhances the discriminability of topological structures through multi-scale adjacency relationship modeling and fine-grained channel interaction, providing a more robust basic representation for subsequent spatiotemporal feature learning.

[0048] Improvements also include temporal modeling: Figure 3 As shown in the figure, a multi-branch structure is used to extract temporal features of different scales. The input tensor is (C, T, V), which represents the number of channels, time steps, and nodes, respectively. The model contains four parallel branches: the first branch uses 1×1 convolution to achieve feature dimensionality transformation and skip connections to retain the original information; the second and third branches respectively extract features of different temporal receptive fields through 1×1 convolution and 5×1 convolution with different dilation rates (dilation=1,2); the fourth branch combines 1×1 convolution with 3×1 maximum pooling operations to capture local short-term temporal information. Finally, the outputs of the four branches are spliced ​​in the channel dimension to form a fused output feature. This structure not only retains the input information, but also effectively improves the temporal modeling capability of the model.

[0049] In step (3), the preliminary recognition network is trained and the model parameters are continuously updated for fine-tuning to obtain the human action recognition model, such as Figure 4 As shown;

[0050] In step (4), the optimal model parameters are loaded, the image to be recognized captured by the Kinect camera is input, the model performs action recognition, and returns the corresponding action category.

[0051] The present invention's experiments employed four types of data streams for human action recognition: the first uses raw skeleton coordinates as input, known as "joint flow"; the second uses second-order information from joint points, known as "bone flow"; the third uses motion information from the joint flow, known as "joint motion flow"; and the fourth uses motion information from the bone flow, known as "bone motion flow." The final recognition result is obtained by summing the softmax scores of these four data streams. On the NTU RGB+D 120Xsub dataset, the present invention's various data streams achieved the following top-1 accuracy: 86.7% for joint flow, 86.3% for bone flow, 82.3% for joint-motion flow, and 82.2% for bone-motion flow. A recognition accuracy of 90.9% was achieved in a cross-view test, and 89.3% in a more challenging cross-subject test.

[0052] Based on the same inventive concept, this embodiment also provides a human motion recognition system based on skeleton sequences, comprising:

[0053] Initialization module, obtains human action dataset, and builds skeleton graph based on the dataset;

[0054] The model building module improves the CTR-GCN network based on spatiotemporal graph convolutional neural networks and contrastive learning to build a preliminary recognition model. The improvements include encoding the absolute and relative positions of skeleton points in the backbone network and reshaping the adjacency matrix using a multi-scale spatial graph.

[0055] The model training module is used to train the preliminary recognition model using the skeleton graph, update the model parameters, and obtain the human action recognition model;

[0056] The recognition module is used to perform action recognition on the target to be recognized using a human action recognition model.

[0057] Furthermore, the skeleton graph uses joints as vertices and bones as edges. Indicates that is a set of N vertices; is a set of edges, through the adjacency matrix Indicates that the element a in the adjacency matrix ij Used to represent the vertex v i and v jThe correlation strength between them, i=1,2,…,N; j=1,2,…,N; vertex v i The neighborhood of express; is the feature set of N vertices, through the matrix Indicates that, where c is the number of channels, vertex v i The characteristic representation of

[0058] Furthermore, the model construction module performs absolute position encoding and relative position encoding of skeletal points in the backbone network and reshapes the adjacency matrix using a multi-scale spatial graph, including: encoding the absolute position of skeletal points; using five topological enhancement modules with shared parameters in the backbone network to extract the correlation between human joints in parallel, while performing relative position encoding of skeletal points;

[0059] The topology enhancement module consists of a multi-scale adjacency matrix, channel refinement features, and relative position encoding. The operation process is as follows:

[0060] First, the input features after absolute position encoding are linearly projected, and then the pairwise differences between feature vectors are calculated in the channel dimension to generate channel-refined features with high-order interactive information; the channel-refined features are element-wise added to the relative position encoding, and weighted aggregation is performed in combination with the multi-scale adjacency matrix to construct an enhanced topological structure representation; based on the enhanced topological relationship, graph convolution is used to extract high-order spatial features.

[0061] Furthermore, the improvement of the model construction module also includes: adopting a multi-branch structure to extract temporal features of different scales, the multi-branch structure includes four parallel branches, wherein the first branch realizes the dimensional transformation of features through 1×1 convolution and skips the connection to retain the original information; the second and third branches extract features of different time receptive fields through 1×1 convolution and 5×1 convolution respectively; the fourth branch combines 1×1 convolution with 3×1 maximum pooling operation to capture local short-term temporal information. Finally, the outputs of the four branches are spliced ​​in the channel dimension to form a fused output feature.

[0062] Based on the same inventive concept, this embodiment also provides a computing device, including: one or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the human motion recognition method based on skeleton sequences according to any one of the above items are implemented.

[0063] Based on the same inventive concept, this embodiment also provides a storage medium, which stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the steps of the human motion recognition method based on skeleton sequences according to any one of the above items.

Claims

1. A human action recognition method based on skeleton sequence, characterized in that: include: (1) Obtain a human motion dataset and construct a skeleton graph based on the dataset; (2) Based on the spatiotemporal graph convolutional neural network and contrastive learning, the CTR-GCN network is improved to build a preliminary recognition model; the improvement includes encoding the absolute and relative positions of the skeleton points in the backbone network and reshaping the adjacency matrix using a multi-scale spatial graph; (3) Use the skeleton graph to train the preliminary recognition model, update the model parameters, and obtain the human action recognition model; (4) Use the human action recognition model to perform action recognition on the target to be identified.

2. The human motion recognition method based on skeleton sequence according to claim 1, characterized in that The skeleton graph uses joints as vertices and bones as edges. Indicates that is a set of N vertices; is a set of edges, through the adjacency matrix Indicates that the element a in the adjacency matrix ij Used to represent the vertex v i and v j The correlation strength between them, i=1,2,…,N; j=1,2,…,N; vertex v i The neighborhood of express; is the feature set of N vertices, through the matrix Indicates that, where c is the number of channels, vertex v i The characteristic representation of 3. The human motion recognition method based on skeleton sequence according to claim 1, characterized in that: Step (2) performs absolute position encoding and relative position encoding of the skeleton points in the backbone network and reshapes the adjacency matrix using a multi-scale spatial graph, including: performing absolute position encoding of the skeleton points; using five topological enhancement modules with shared parameters in the backbone network to extract the correlation between human joints in parallel, and performing relative position encoding of the skeleton points at the same time; The topology enhancement module consists of a multi-scale adjacency matrix, channel refinement features, and relative position encoding. The operation process is as follows: First, the input features after absolute position encoding are linearly projected, and then the pairwise differences between feature vectors are calculated in the channel dimension to generate channel-refined features with high-order interactive information; the channel-refined features are element-wise added to the relative position encoding, and weighted aggregation is performed in combination with the multi-scale adjacency matrix to construct an enhanced topological structure representation; based on the enhanced topological relationship, graph convolution is used to extract high-order spatial features.

4. The human motion recognition method based on skeleton sequence according to claim 1, characterized in that The improvement in step (2) further includes: adopting a multi-branch structure to extract temporal features of different scales, wherein the multi-branch structure includes four parallel branches, wherein the first branch realizes the dimensional transformation of features through 1×1 convolution and skips the connection to retain the original information; the second and third branches extract features of different time receptive fields through 1×1 convolution and 5×1 convolution respectively; the fourth branch combines 1×1 convolution with 3×1 maximum pooling operation to capture local short-term temporal information, and finally, the outputs of the four branches are spliced ​​in the channel dimension to form a fused output feature.

5. A human action recognition system based on skeleton sequence, characterized in that: include: Initialization module, obtains human action dataset, and builds skeleton graph based on the dataset; The model building module improves the CTR-GCN network based on spatiotemporal graph convolutional neural networks and contrastive learning to build a preliminary recognition model. The improvements include encoding the absolute and relative positions of skeleton points in the backbone network and reshaping the adjacency matrix using a multi-scale spatial graph. The model training module is used to train the preliminary recognition model using the skeleton graph, update the model parameters, and obtain the human action recognition model; The recognition module is used to perform action recognition on the target to be recognized using a human action recognition model.

6. The human motion recognition system based on skeleton sequence according to claim 5, characterized in that: The skeleton graph uses joints as vertices and bones as edges. Indicates that is a set of N vertices; is a set of edges, through the adjacency matrix Indicates that the element a in the adjacency matrix ij Used to represent the vertex v i and v j The correlation strength between them, i=1,2,…,N; j=1,2,…,N; vertex v i The neighborhood of express; is the feature set of N vertices, through the matrix Indicates that, where c is the number of channels, vertex v i The characteristic representation of 7. The human motion recognition system based on skeleton sequence according to claim 5, characterized in that: The model construction module performs absolute and relative position encoding of skeletal points in the backbone network and reshapes the adjacency matrix using a multi-scale spatial graph, including: encoding the absolute positions of skeletal points; using five topological enhancement modules with shared parameters in the backbone network to extract the correlation between human joints in parallel, while performing relative position encoding of skeletal points; The topology enhancement module consists of a multi-scale adjacency matrix, channel refinement features, and relative position encoding. The operation process is as follows: First, the input features after absolute position encoding are linearly projected, and then the pairwise differences between feature vectors are calculated in the channel dimension to generate channel-refined features with high-order interactive information; the channel-refined features are element-wise added to the relative position encoding, and weighted aggregation is performed in combination with the multi-scale adjacency matrix to construct an enhanced topological structure representation; based on the enhanced topological relationship, graph convolution is used to extract high-order spatial features.

8. The human motion recognition system based on skeleton sequence according to claim 5, characterized in that: The improvements in the model building module also include: using a multi-branch structure to extract temporal features of different scales, the multi-branch structure includes four parallel branches, wherein the first branch realizes the dimensional transformation of features through 1×1 convolution and skips the connection to retain the original information; the second and third branches extract features of different time receptive fields through 1×1 convolution and 5×1 convolution respectively; the fourth branch combines 1×1 convolution with 3×1 maximum pooling operations to capture local short-term temporal information. Finally, the outputs of the four branches are spliced ​​in the channel dimension to form a fused output feature.

9. A computing device, characterized in that include: One or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the human motion recognition method based on skeleton sequences according to any one of claims 1 to 4 are implemented.

10. A storage medium, characterized in that: The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor executes the steps of the human motion recognition method based on skeleton sequences according to any one of claims 1 to 4.