A 3D Skeletal Human Body Full-Body Action Recognition Method, System, Device and Medium

By adopting multi-style adjacency topology matrix and feature refinement processing in three-dimensional skeleton human movement recognition, the shortcomings of existing methods in capturing semantic associations between parts and paying attention to local features are solved, and more refined action feature learning and stronger robustness are achieved.

CN120014715BActive Publication Date: 2025-06-17NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510504014.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-17
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing three-dimensional bone human body movement recognition method based on graph convolution networks is difficult to effectively capture the complex semantic relationships between different functional parts, and it focuses too much on global features, neglecting the detailed information of local bone features and the differences in contributions of different parts, resulting in the limited understanding of complex actions of the model and insufficient robustness to occlusion and subtle movement changes.

Method used

A multi-style adjacency topology matrix is ​​adopted to decompose the traditional single matrix into three independent local matrices, namely the head, the trunk and the lower limbs, and fuse it with the global matrix through a spatial mapping aggregation mechanism to build a multi-level topology structure. At the same time, the output features are refined through deep convolution and spatial point convolution, local feature representation is generated, and a multi-part comparison loss function is introduced, combined with the label smooth cross entropy loss, to achieve complementary learning between global and local features.

Benefits of technology

The model's ability to capture local area context features is enhanced, the ability to identify movements of occluded parts is improved, and the model's ability to understand complex movements and its robustness to subtle movement changes is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014715B_ABST
    Figure CN120014715B_ABST
Patent Text Reader

Abstract

This application belongs to the field of computer vision and discloses a three-dimensional skeletal human body full-action recognition method, system, device, and medium based on a multi-style topological matrix and refined skeletal features. The method first obtains skeletal action sequence data containing the three-dimensional coordinates of multiple frames of human joint points and constructs a multi-style topological structure composed of three local adjacency matrices of the head, torso, and lower limbs and a global adjacency matrix. After spatio-temporal feature extraction through a graph convolutional network, the output features are refined to generate local and global feature representations, and the corresponding prediction probability distributions are calculated. A combined loss function that jointly optimizes the global and local prediction probability distributions is used to train the network, and finally, the action classification result is output. This method captures joint relationships through an effective skeletal topology representation, combines local and global feature fine-grained learning, and improves the performance of three-dimensional skeletal human body action recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a three-dimensional skeletal human body full-body action recognition method, system, device and medium based on a multi-style topological matrix and refined skeletal features. Background Art

[0002] Human body full-body action recognition is one of the core research directions in the field of computer vision, and has important application values in scenarios such as human-computer interaction, intelligent monitoring, virtual reality, etc. In recent years, with the development of deep learning technology and three-dimensional sensors, human action recognition based on three-dimensional skeletal data has attracted much attention due to its robustness to environmental factors such as light and background. Skeletal data accurately describes human postures and movements through joint point coordinates, providing effective information for action recognition.

[0003] Currently, graph convolutional networks (GCNs) have become the mainstream method for processing skeletal data. Existing GCN-based action recognition methods, such as spatio-temporal graph convolutional networks (ST-GCNs) and their improvements (such as DEGCN, STFGCN, etc.), usually rely on a predefined or adaptively learned single adjacency topological matrix to represent the connection relationships between skeletal joint points. Although this single matrix can capture the spatial relationships of physical connections or neighboring joints, it has limitations in characterizing the complex semantic associations within and between different functional regions of the human body (such as the upper limbs, torso, lower limbs). For example, hand actions and foot actions may have strong correlations in a specific behavior (such as "kicking a ball"), but this non-physically directly connected semantic relationship is difficult to effectively express through a single fixed topological structure.

[0004] In addition, in the feature learning and loss calculation stages, existing methods often focus on the representation and optimization of global skeletal features. For example, only the final global feature vector is used in contrastive learning or classification loss calculation. This approach ignores the contribution differences of different parts of the human body during action execution and the importance of local features, which may cause the model to be insensitive to local details (such as gestures, specific limb postures), or the performance to degrade when joint points are occluded. Although some methods attempt to introduce attention mechanisms or adaptive graph structures, they fail to fundamentally solve the limitations of single topological representation and global feature preference.

[0005] Therefore, a single fixed or globally adaptive adjacency topological matrix is difficult to effectively capture the complex semantic associations within and between different functional parts of the head, torso, and lower limbs, limiting the model's ability to understand complex actions; and, existing methods overly focus on global features in feature learning and loss calculation, ignoring the detailed information of local skeletal features and the contribution differences of different parts, resulting in insufficiently fine feature representation and low robustness to occlusion and subtle action changes. Summary of the Invention

[0006] The objective of the embodiments of this application is to provide a method, system, device, and medium for three-dimensional skeletal human full-body action recognition, aiming to design a more effective skeletal topology representation method to capture rich inter-joint relationships, and combine local and global features for more refined learning to improve the performance of three-dimensional skeletal human action recognition.

[0007] To solve the above technical problems, this application is implemented as follows:

[0008] In a first aspect, the embodiments of this application provide a method for three-dimensional skeletal human full-body action recognition, and the method includes:

[0009] Obtain three-dimensional skeletal action sequence data, where the three-dimensional skeletal action sequence data includes three-dimensional coordinates of human joint points on multiple time frames;

[0010] Construct a multi-style adjacency topology matrix, where the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, torso, and lower limb partitions respectively, and a global adjacency topology matrix;

[0011] Use the multi-style adjacency topology matrix to guide a graph convolutional network to perform spatio-temporal feature extraction on the three-dimensional skeletal action sequence data to obtain output features;

[0012] Refine the output features to generate three local feature representations and one global feature representation;

[0013] Based on the three local feature representations and the one global feature representation, calculate the corresponding predicted probability distribution;

[0014] Calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label, and the losses between each local predicted probability distribution and the true label;

[0015] Train the graph convolutional network according to the combined loss function;

[0016] Based on the trained graph convolutional network, finally output the predicted action classification result corresponding to the three-dimensional skeletal action sequence data.

[0017] As an optional implementation manner of the first aspect of this application, the step of constructing a multi-style adjacency topology matrix includes: dividing human skeletal joint points into three independent subsets including the head, torso, and lower limbs according to human anatomy; generating a local adjacency topology matrix for each independent subset, where the local adjacency topology matrix only contains the connection relationships of the joint points within the corresponding subset; and fusing the three local adjacency topology matrices and a global adjacency topology matrix through a spatial mapping aggregation mechanism to form a multi-level topology structure.

[0018] As an alternative implementation of the first aspect of the present application, the steps of using the multi-style adjacency topology matrix to guide the graph convolutional network to extract spatio-temporal features from the three-dimensional skeletal action sequence data and obtain output features include: in the graph convolutional network, using the Einstein summation law to perform matrix multiplication on the three-dimensional skeletal action sequence data to obtain output features.

[0019] As an alternative implementation of the first aspect of the present application, the steps of refining the output features to generate three local feature representations include: expanding the dimensions of the output features through deep convolution and spatial point convolution; decoupling the output features after dimension expansion by partitioning them into the head, torso, and lower limbs to generate three sets of local feature representations corresponding to three local adjacency topology matrices.

[0020] As an alternative implementation of the first aspect of the present application, the steps of decoupling the output features after dimension expansion by partitioning them into the head, torso, and lower limbs to generate three sets of local feature representations corresponding to three local adjacency topology matrices include: decomposing the original adjacency matrix in the output features into three independent sub-matrices for the head, torso, and lower limbs, which is expressed by the formula: , where respectively represent the independent skeletal parts of the head, torso, and lower limbs; refining the three independent sub-matrices into corresponding feature information to obtain three sets of local feature representations, which is expressed by the formula: , where respectively represent the independent skeletal features of the head, torso, and legs.

[0021] As an alternative implementation of the first aspect of the present application, the steps of calculating the corresponding prediction probability distribution based on the three local feature representations and the one global feature representation include: respectively performing regularization processing on the three local feature representations and the one global feature representation, and generating the corresponding prediction probability distribution through a fully connected layer and the Softmax function.

[0022] As an alternative implementation of the first aspect of the present application, in the steps of calculating the combined loss function, the combined loss function is calculated using label-smoothing cross-entropy loss, and its weight parameter is dynamically adjusted by learnable parameters. The combined loss function is: , where, represents the label-smoothing cross-entropy loss function, y represents the true label, represents the predicted action classification result, and respectively represent the global feature and the local features of different parts, represents the learnable weight parameter between different parts.

[0023] In a second aspect, an embodiment of the present application provides a three-dimensional skeletal human body full-body action recognition system, which includes:

[0024] A data acquisition module, configured to acquire three-dimensional skeletal action sequence data, where the three-dimensional skeletal action sequence data includes three-dimensional coordinates of human body joint points on multiple time frames;

[0025] A multi-style topological matrix construction module, configured to construct a multi-style adjacency topological matrix, where the multi-style adjacency topological matrix includes three local adjacency topological matrices corresponding to the head, torso, and lower limb partitions respectively, and a global adjacency topological matrix;

[0026] A spatio-temporal feature extraction module, configured to use the multi-style adjacency topological matrix to guide a graph convolutional network to perform spatio-temporal feature extraction on the three-dimensional skeletal action sequence data to obtain output features;

[0027] A feature refinement module, configured to perform refinement processing on the output features to generate three local feature representations and a global feature representation;

[0028] A probability calculation module, configured to calculate a corresponding predicted probability distribution based on the three local feature representations and the global feature representation;

[0029] A loss function calculation module, configured to calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label, and the losses between each local predicted probability distribution and the true label;

[0030] A model training module, configured to train the graph convolutional network according to the combined loss function;

[0031] An action recognition module, configured to finally output a predicted action classification result corresponding to the three-dimensional skeletal action sequence data based on the trained graph convolutional network.

[0032] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0033] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0034] Compared with the prior art, the present invention provides a three-dimensional skeletal human body full-body action recognition method, which has the following beneficial effects:

[0035] (1) Multi - style Adjacent Topology Matrix Module: Decompose the traditional single - adjacent topology matrix into independent sub - matrices for three functional regions: the head, the torso, and the lower limbs. Each sub - matrix independently processes local bone features and is fused with the global topology matrix through a spatial mapping aggregation mechanism to construct a multi - level topology structure, enhancing the model's ability to capture local - region context features.

[0036] (2) Fine - grained Bone Feature Module: Decompose the features after graph convolution processing to generate local feature representations corresponding to the partitioned topology matrix, and introduce a multi - part contrast loss function. Combining with the label - smoothed cross - entropy loss, it realizes complementary learning of global and local features, improving the model's ability to recognize actions of occluded parts.

[0037] (3) Overall Framework: Integrate the multi - style adjacent topology matrix module and the fine - grained bone feature module into the graph convolution network, extract sequence features through spatio - temporal convolution blocks, and optimize the loss function to enhance the model's generalization ability. Description of the Drawings

[0038] Figure 1 is a flowchart of a three - dimensional skeletal human body full - body action recognition method provided by the first embodiment of the present invention;

[0039] Figure 2 is the overall architecture diagram of the first embodiment of the present invention;

[0040] Figure 3 is the segmented multi - style bone structure diagram in the first embodiment of the present invention;

[0041] Figure 4 is the overall flowchart of the multi - style topology matrix (CRKC) method in the first embodiment of the present invention;

[0042] Figure 5 is the overall flowchart of the bone feature refinement (CRKC) method in the first embodiment of the present invention;

[0043] Figure 6 is a schematic structural diagram of a three - dimensional skeletal human body full - body action recognition system provided by the second embodiment of the present invention. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0045] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0046] In order to illustrate the technical solutions described in this application, the following will be described through specific embodiments.

[0047] Embodiment 1

[0048] Please refer to Figure 1 , which is a flowchart of a three-dimensional skeletal human full-body action recognition method proposed in the first embodiment of this application; please refer to Figure 2 , which is the overall architecture of this embodiment, mainly including a multi-style topology matrix (MTM) module and a refined skeletal feature (CRKC) module. In order to make the correlation degree of different connection points of the topology matrix between bones closer, the MTM module divides the original fixed topology matrix into upper, middle, and lower independent adjacency topology matrices on the bone parts, and uses the strong adaptability between local features and global features to aggregate into a new adjacency topology matrix through spatial mapping, helping the model obtain data and features in the entire graph convolution, so that the topological relationship between different bones can be fully learned. The CRKC module refines the data features processed by graph convolution into local modules corresponding to the adjacency topology matrix, proposes a multi-part contrast loss, and performs label smoothing loss on the global and local independent prediction distribution results and the true labels, thereby strengthening the training loss prediction results of the occluded part features, improving its robustness and adaptability, and enhancing the fine-grainedness between data features.

[0049] The steps of the proposed method are as follows:

[0050] S1: Obtain three-dimensional skeletal action sequence data, which includes the three-dimensional coordinates of human joint points on multiple time frames.

[0051] Specifically, in the data preprocessing stage, the format of the action sequence is unified as , where M, T, and V represent the number of people, frames, and joints respectively. For the three-dimensional skeletal coordinates, they are represented as , , , . Then is divided into three independent features:

[0052] Upper head ;

[0053] Middle torso ;

[0054] Lower legs .

[0055] The above represents the coordinate positions of the articulation points of the bones in different parts. The diverse bone structures are segmented as shown in Figure 3 . Each independent bone point is mapped to each other to enhance the fine-grainedness and adaptability of the feature information.

[0056] S2: Construct a diverse adjacency topology matrix, which includes three local adjacency topology matrices corresponding to the head, torso, and lower limb partitions respectively, and a global adjacency topology matrix.

[0057] In this embodiment, the human bone articulation points are divided into three independent subsets including the head, torso, and lower limbs according to human anatomy; a local adjacency topology matrix is generated for each independent subset, and the local adjacency topology matrix only contains the connection relationships of the articulation points within the corresponding subset; the three local adjacency topology matrices are fused with a global adjacency topology matrix through a spatial mapping aggregation mechanism to form a multi-level topological structure.

[0058] It should be noted that the distances between different bone key points reflect the changes in the human motion trajectory, and the position coordinates of the key bone points of different motions are also different, which poses a greater challenge to the generalization ability of the model for motions. However, all these important information is reflected through the adjacency topology matrix. Segmenting the diverse adjacency topology matrix (as shown in Figure 2 and Figure 3 , where each independent matrix corresponds to the color of the segmented bone structure) can improve the model's ability to obtain the key information of each motion, enhance the interaction between different bone articulation points in each frame sequence, and on this basis, consider strengthening the correspondence relationship between the global and local feature topological matrices, effectively improving the ability to recognize motions.

[0059] S3: Use the diverse adjacency topology matrix to guide the graph convolutional network to extract spatio-temporal features from the three-dimensional bone motion sequence data to obtain output features.

[0060] It should be noted that in the graph convolution stage, it depends on the adjacency topology matrix A and the self-learning topological weight matrix W , where k corresponds to different parts, represents the number of input channels, Let \(C\) represent the number of output channels and \(S\) represent the size of the kernel. For non-Euclidean structured data modeling of skeletal data, the relationships between skeletal joints represented by the topological matrix are utilized to guide the learning of the model.

[0061] Specifically, regarding the properties of the adjacency topological matrix, it is a graph with skeletal nodes: , where represents the set of nodes, represents the set of edges, and the graph represents a matrix. The elements in the matrix represent the proximity relationship between node and node . In the initial graph convolution, the output feature representation is:

[0062] ,

[0063] where represents the output feature, represents the activation function, represents the adjacency topological matrix, represents the input feature, and represents the learnable adjacency weight parameter matrix.

[0064] Existing human full-body action recognition methods based on 3D skeletons use the skeletal points corresponding to a single overall adjacency topological matrix for feature extraction, ignoring the topological relationships existing in the skeletal adjacency matrix, not comprehensively learning its topological relationships, affecting the training results of the overall model, and leading to a decrease in classification accuracy. Therefore, to solve this problem and enhance the connection between different skeletal joints, the MNTM method is proposed to solve the above problems.

[0065] Considering the relationships between the connection points corresponding to the adjacency topological matrices of different skeletal points, using the idea of a multi-graph adjacency matrix, the original adjacency matrix in the output feature is decomposed into three independent sub-matrices for the head, torso, and lower limbs, which is expressed by the formula:

[0066] ,

[0067] where respectively represent the independent skeletal parts of the head, torso, and lower limbs.

[0068] The three independent sub-matrices are refined into corresponding feature information to obtain three groups of local feature representations, which is expressed by the formula: , where respectively represent the independent skeletal features of the head, torso, and legs.

[0069] In this embodiment, the Einstein summation law is adopted in the graph convolutional network to perform matrix multiplication on the three-dimensional skeletal action sequence data to obtain the output features.

[0070] The most prominent feature in graph convolution is the matrix multiplication between adjacent topological matrices, and its method can be expressed as (Einstein summation law):

[0071] ,

[0072] where the matrix corresponds to dimensions, V represents the number of joint points, represents the dimension of the feature channels, represents an index of the spatial dimension; the feature corresponds to dimensions, represents the number of samples, represents the number of channels, represents the number of time frames, V represents the number of joint points, represents the same spatial dimension index as . The specific operation process is as follows: The elements on the tensor dimension are multiplied and added to the elements of the feature tensor dimension according to the index for item-by-item corresponding matching, resulting in a new tensor with the dimension of , where the meaning is similar to the above. This tensor effectively integrates the information of the two tensors in different dimensions.

[0073] Therefore, based on the architecture of graph convolution, the output features are represented as:

[0074] ,

[0075] By aggregating the global and local multiple non-linear adjacency matrices, it supports the linear relationship between different topological matrices, optimizes the coupling degree between the feature information of skeletal points in different parts, forms more specific skeletal features, and fully correlates the feature information of each part in graph convolution, promoting the pre-function of action recognition.

[0076] S4: Refine the output features to generate three local feature representations and one global feature representation.

[0077] In this embodiment, the output features are dimensionally expanded through deep convolution and spatial point convolution; the dimensionally expanded output features are decoupled by region into the head, torso, and lower limbs to generate three groups of local feature representations corresponding to three local adjacent topological matrices.

[0078] For example Figure 4As shown, an example of the MTM method is presented. Given the original skeletal adjacency matrix and skeletal features, according to the previous research methods for feature extraction, an independent spatio-temporal convolutional block is used to extract skeletal information. However, the MTM method adopts deep convolution and spatial point convolution. Here, k corresponds to the receptive field of skeletal information. Deep convolution interconnects the input and output channels of each convolutional kernel independently, increasing the receptive field of feature information, reducing the number of feature parameters, and enhancing the convolution efficiency. Spatial point convolution constructs the relationship between input channels point by point, extending the feature dimension. Subsequently, the global adjacency topology matrix and skeletal features are segmented into three parts of feature information: upper (head), middle (torso), and lower (legs). Each part is aggregated in the activation function, and the GELU method is used as the activation function to supplement the training effect. Finally, new global information is aggregated point by point through spatial point convolution.

[0079] S5: Based on three local feature representations and one global feature representation, calculate the corresponding predicted probability distribution.

[0080] In this embodiment, regularization processing is performed on the three local feature representations and one global feature representation respectively, and the corresponding predicted probability distribution is generated through a fully connected layer and a Softmax function.

[0081] Specifically, in the original action recognition method, the Softmax function and the cross-entropy loss function are often used to represent the function for predicting the probability distribution of each category and to measure the difference between the predicted value and the true label, respectively, to judge the result of action recognition.

[0082] The Softmax function is expressed as: ,

[0083] where . is the value of the th element of the input vector, is any element sum, represents the result obtained by exponentiating the th output feature vector and then dividing it by the sum of the exponentiations of all vectors, and it is the probability value of the th output.

[0084] The cross-entropy loss function is expressed as: ,

[0085] where is the true label data, is the skeletal feature information, represents the difference value between the predicted value and the true label.

[0086] For skeleton-based full-body human action recognition, discriminative features of skeleton information are particularly important. The original research methods cannot ensure that a certain action category is classified into the correct class. During the classification and comparison training, only global features are used for contrastive loss training, and the correlation between skeleton feature information is poor, resulting in inaccurate recognition of actions in occluded joint parts. To better represent discriminative features, the CRKC method refines the output features again, dividing them into three parts: the head, the torso, and the legs, and comparing them separately to represent the similarity between actions in each part.

[0087] S6: Calculate the combined loss function, which is the weighted sum of the loss between the global prediction probability distribution and the true label, and the losses between each local prediction probability distribution and the true label.

[0088] In this embodiment, the combined loss function is calculated using label-smoothing cross-entropy loss, and its weight parameter is dynamically adjusted by learnable parameters. The combined loss function is:

[0089] ,

[0090] where represents the label-smoothing cross-entropy loss function, y represents the true label, represents the predicted action classification result, and represent the global feature and the local features of different parts respectively, represents the learnable weight parameter between different parts. Integrate the global and local features and the contrastive loss with the true label, and optimize the parameters between models, so that the model can better fit the data and improve the generalization ability of the model to feature data.

[0091] Specifically, based on the original function, a multi-part contrastive loss method is proposed to distinguish the difference between predicting real actions and ambiguous actions, and improve the intra-class consistency of a certain action, as shown in the following formula:

[0092]

[0093]

[0094]

[0095] where is the parameter controlling the weight of the label-smoothing cross-entropy loss, , , represent the predicted values of the corresponding parts respectively. Indicates the difference between the real action and the predicted action

[0096] S7: Train the graph convolutional network according to the combined loss function.

[0097] S8: Based on the trained graph convolutional network, finally output the predicted action classification result corresponding to the three-dimensional skeletal action sequence data.

[0098] As Figure 5 shown, the overall flowchart of CRKC is shown. This method first processes a series of input skeletons with a shape of which represents the T-frame information of the joints in the three-dimensional skeletal space. The main chain of the graph convolution consists of 11 basic units. The graph convolution block is composed of TCN (temporal convolution) and GCN, which are collectively called the TGN block. This block comprehensively applies CNNs in the temporal dimension and the spatial dimension respectively to extract sequence features, and uses a learnable adjacency topology matrix to extract spatial skeletal point features. It should be noted that the cross-temporal block in the basic unit is a method to reduce the temporal dimension and increase the channel feature dimension. Refine the output features of the graph convolution, perform regularization operations on the global and local feature information to prevent overfitting of the data, then project them into a fully connected layer with a softmax activation function to predict the class distribution probability, and finally use the label smoothing loss to train with the true label to find the true action class. in the temporal dimension and the spatial dimension respectively to extract sequence features, and uses a learnable adjacency topology matrix to extract spatial skeletal point features. It should be noted that the cross-temporal block in the basic unit is a method to reduce the temporal dimension and increase the channel feature dimension. Refine the output features of the graph convolution, perform regularization operations on the global and local feature information to prevent overfitting of the data, then project them into a fully connected layer with a softmax activation function to predict the class distribution probability, and finally use the label smoothing loss to train with the true label to find the true action class.

[0099] Experimental results show that in the NTU RGB+D 60 dataset, the accuracies of X-sub (Top-1) and X-view (Top-1) reach 93.22% and 97.10% respectively, and in the NTU RGB+D 120 dataset, the accuracies of X-sub (Top-1) and X-set (Top-1) are 90.30% and 91.61% respectively, reaching the current advanced level.

[0100] Embodiment 2

[0101] Please refer to Figure 6 which shows the structural schematic diagram of a three-dimensional skeletal human body full-body action recognition system proposed in the second embodiment of the present application. The system includes:

[0102] A data acquisition module 100, configured to acquire three-dimensional skeletal action sequence data, where the three-dimensional skeletal action sequence data includes the three-dimensional coordinates of human joint points on multiple time frames;

[0103] A multi-style topology matrix construction module 200, configured to construct a multi-style adjacency topology matrix, where the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, torso, and lower limb partitions respectively and a global adjacency topology matrix;

[0104] A spatio-temporal feature extraction module 300, configured to use the multi-style adjacency topology matrix to guide a graph convolutional network to perform spatio-temporal feature extraction on the three-dimensional skeletal action sequence data, so as to obtain output features;

[0105] A feature refinement module 400, configured to refine the output features to generate three local feature representations and one global feature representation;

[0106] A probability calculation module 500, configured to calculate a corresponding predicted probability distribution based on the three local feature representations and the one global feature representation;

[0107] A loss function calculation module 600, configured to calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label, and the losses between each local predicted probability distribution and the true label;

[0108] A model training module 700, configured to train the graph convolutional network according to the combined loss function;

[0109] An action recognition module 800, configured to finally output a predicted action classification result corresponding to the three-dimensional skeletal action sequence data based on the trained graph convolutional network.

[0110] A three-dimensional skeletal human body full-body action recognition system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., and the embodiments of the present application do not make specific limitations.

[0111] A device of a three-dimensional skeletal human body full-body action recognition system in an embodiment of the present application. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application do not make specific limitations.

[0112] A three-dimensional skeletal human body full-body action recognition system provided by an embodiment of the present application can achieve Figure 1For the sake of avoiding repetition, the processes implemented in the method embodiments of a three-dimensional skeletal human body full-body action recognition method will not be elaborated here.

[0113] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned method embodiment of a three-dimensional skeletal human body full-body action recognition method and can achieve the same technical effects. For the sake of avoiding repetition, it will not be elaborated here.

[0114] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned method embodiment of a three-dimensional skeletal human body full-body action recognition method and can achieve the same technical effects. For the sake of avoiding repetition, it will not be elaborated here.

[0115] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.

[0116] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiment can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0118] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A three-dimensional skeleton human body action recognition method, characterized in that: The method comprises the following steps: Acquire three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; Constructing a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limbs partitions respectively and a global adjacency topology matrix; Using the multi-style adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features from the three-dimensional skeletal motion sequence data to obtain output features; Refining the output features to generate three local feature representations and one global feature representation; Based on the three local feature representations and the one global feature representation, calculating a corresponding predicted probability distribution; Calculating a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; Training the graph convolutional network according to the combined loss function; Based on the trained graph convolutional network, the predicted action classification result corresponding to the three-dimensional skeletal action sequence data is finally output.

2. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: The steps to construct a multi-style adjacency topology matrix include: The human skeleton joints are divided into three independent subsets including the head, trunk and lower limbs according to the human anatomy. Generate a local adjacency topology matrix for each independent subset, wherein the local adjacency topology matrix only contains the connection relationship of the joint points in the corresponding subset; The three local adjacency topology matrices are fused with a global adjacency topology matrix through a spatial mapping aggregation mechanism to form a multi-level topological structure.

3. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: The step of using the multi-style adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features from the three-dimensional skeletal motion sequence data to obtain output features includes: In the graph convolutional network, the Einstein summation law is used to perform matrix multiplication on the three-dimensional skeletal motion sequence data to obtain output features.

4. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: The step of refining the output features to generate three local feature representations includes: The output features are dimensionally expanded through deep convolution and spatial point convolution; The output features after dimension expansion are decoupled according to the head, torso and lower limb partitions to generate three groups of local feature representations corresponding to three local adjacency topology matrices.

5. A three-dimensional skeleton human body motion recognition method according to claim 4, characterized in that: The steps of decoupling the output features after dimension expansion according to the head, torso and lower limb partitions to generate three sets of local feature representations corresponding to three local adjacency topology matrices include: The original adjacency matrix in the output feature is decomposed into three independent sub-matrices of head, torso and lower limbs, which can be expressed as follows: ,in Representing separate skeletal parts of the head, trunk, and lower limbs; The three independent sub-matrices are refined into corresponding feature information to obtain three sets of local feature representations, which are expressed as follows: ,in Represents independent bone features of the head, torso, and legs respectively.

6. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: Based on the three local feature representations and the one global feature representation, the step of calculating the corresponding predicted probability distribution comprises: Regularization is performed on the three local feature representations and the one global feature representation, respectively, and corresponding prediction probability distributions are generated through a fully connected layer and a Softmax function.

7. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: In the step of calculating the combined loss function, the combined loss function is calculated using label smoothed cross entropy loss, and its weight parameters are dynamically adjusted through learnable parameters. The combined loss function is: , in, represents the label smoothed cross entropy loss function, y represents the true label, represents the predicted action classification result, and Respectively represent the global features and local features of different parts, Represents the learnable weight parameters between different parts.

8. A three-dimensional skeleton human body action recognition system, characterized in that: The system comprises: A data acquisition module, used for acquiring three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; A multi-style topology matrix construction module, used to construct a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limb partitions respectively and a global adjacency topology matrix; A spatiotemporal feature extraction module, used to extract spatiotemporal features from the three-dimensional skeletal motion sequence data using the multi-style adjacency topology matrix to guide the graph convolutional network to obtain output features; A feature refinement module, used for refining the output features to generate three local feature representations and one global feature representation; A probability calculation module, used to calculate a corresponding prediction probability distribution based on the three local feature representations and the one global feature representation; A loss function calculation module, used to calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; A model training module, used for training the graph convolutional network according to the combined loss function; The action recognition module is used to output the predicted action classification result corresponding to the three-dimensional skeletal action sequence data based on the trained graph convolutional network.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a three-dimensional skeletal human body whole body motion recognition method as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a three-dimensional skeletal human body whole body motion recognition method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • End-to-end human behavior recognition method and model based on skeleton nodes

    CN114613013A

  • Action recognition method based on dynamic local-global graph convolutional neural network

    CN114998525A