Three-dimensional skeleton human whole body action recognition method, system and equipment and medium
By adopting a multi-style adjacency topology matrix and refined feature processing method in three-dimensional skeleton human movement recognition, the problem of neglecting complex semantic associations and local features in the prior art is solved, and more refined feature learning and higher robustness are achieved.
Patent Information
- Application Number
- CN202510504014.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing three-dimensional bone human body movement recognition method based on graph convolution network is difficult to effectively capture the complex semantic relationships inside and between different functional areas of the human body. It focuses too much on global features, ignores the detailed information of local bone features and the contribution differences between different parts, resulting in insufficient feature representation of feature expression and insufficient robustness to occlusion and subtle movement changes.
A multi-style adjacency topology matrix is adopted to decompose the traditional single matrix into three independent local matrices, namely the head, the trunk and the lower limbs, and fuse it with the global matrix through a spatial mapping aggregation mechanism to build a multi-level topology structure. At the same time, the features after graph convolution are refined, local feature representations are generated, and a multi-part comparison loss function is introduced, combined with the label smooth cross entropy loss, to achieve complementary learning between global and local features.
The model's ability to capture local area context features is enhanced, the ability to identify movements of occluded parts is improved, and the model's ability to understand and robustness of complex actions is improved.
Smart Images

Figure CN120014715A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a three-dimensional skeleton human body motion recognition method, system, device and medium based on a multi-style topological matrix and refined skeleton features. Background Art
[0002] Human body motion recognition is one of the core research directions in the field of computer vision, and has important application value in scenarios such as human-computer interaction, intelligent monitoring, and virtual reality. In recent years, with the development of deep learning technology and three-dimensional sensors, human motion recognition based on three-dimensional skeleton data has attracted much attention due to its robustness to environmental factors such as lighting and background. Skeleton data accurately describes human posture and movement through joint point coordinates, providing effective information for motion recognition.
[0003] Currently, graph convolutional networks (GCNs) have become the mainstream method for processing skeletal data. Existing GCN-based action recognition methods, such as spatiotemporal graph convolutional networks (ST-GCNs) and their improvements (such as DEGCN, STFGCN, etc.), usually rely on a predefined or adaptively learned single adjacency topology matrix to represent the connection relationship between skeletal joints. Although this single matrix can capture the spatial relationship of physical connections or adjacent joints, it has limitations in representing the complex semantic associations within and between different functional areas of the human body (such as upper limbs, torso, and lower limbs). For example, hand movements and foot movements may have a strong correlation in a specific behavior (such as "kicking a ball"), but this non-physical direct semantic relationship is difficult to effectively express through a single fixed topological structure.
[0004] In addition, existing methods often focus on the representation and optimization of global skeletal features during feature learning and loss calculation stages, such as using only the final global feature vector in contrastive learning or classification loss calculation. This approach ignores the differences in contributions of different parts of the human body when performing actions and the importance of local features, which may cause the model to be insensitive to local details (such as gestures, specific limb postures), or performance degradation when joints are occluded. Although some methods have attempted to introduce attention mechanisms or adaptive graph structures, they have failed to fundamentally address the limitations of a single topological representation and global feature preference.
[0005] Therefore, a single fixed or globally adaptive adjacency topology matrix can hardly effectively capture the complex semantic associations within and between different functional parts of the head, torso, and lower limbs, limiting the model's ability to understand complex movements. In addition, existing methods place too much emphasis on global features in feature learning and loss calculation, ignoring the detailed information of local skeletal features and the differences in contributions of different parts, resulting in insufficiently refined feature representation and insufficient robustness to occlusion and subtle movement changes. Summary of the invention
[0006] The purpose of the embodiments of the present application is to provide a method, system, device and medium for three-dimensional skeletal human body motion recognition, aiming to design a more effective skeletal topology representation method to capture rich relationships between joints, and combine local and global features for more refined learning to improve the three-dimensional skeletal human body motion recognition performance.
[0007] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a method for recognizing full-body movements of a three-dimensional skeleton human body, the method comprising: Acquire three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; Constructing a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limbs partitions respectively and a global adjacency topology matrix; Using the multi-style adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features from the three-dimensional skeletal motion sequence data to obtain output features; Refining the output features to generate three local feature representations and one global feature representation; Based on the three local feature representations and the one global feature representation, calculating a corresponding predicted probability distribution; Calculating a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; Training the graph convolutional network according to the combined loss function; Based on the trained graph convolutional network, the predicted action classification result corresponding to the three-dimensional skeletal action sequence data is finally output.
[0008] As an optional implementation of the first aspect of the present application, the step of constructing a multi-style adjacency topology matrix includes: dividing the human skeletal joints into functional areas according to human anatomy, and dividing them into three independent subsets including the head, torso and lower limbs; generating a local adjacency topology matrix for each independent subset, and the local adjacency topology matrix only contains the connection relationship between the joints in the corresponding subset; and fusing the three local adjacency topology matrices with a global adjacency topology matrix through a spatial mapping aggregation mechanism to form a multi-level topological structure.
[0009] As an optional implementation of the first aspect of the present application, the step of using the diverse adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features of the three-dimensional skeletal motion sequence data to obtain output features includes: in the graph convolutional network, using Einstein's summation law to perform matrix multiplication on the three-dimensional skeletal motion sequence data to obtain output features.
[0010] As an optional implementation of the first aspect of the present application, the output features are refined to generate three local feature representations, including: dimensionality expansion of the output features through deep convolution and spatial point convolution; feature decoupling of the dimensionally expanded output features according to the head, torso and lower limb partitions, to generate three groups of local feature representations corresponding to three local adjacency topology matrices.
[0011] As an optional implementation of the first aspect of the present application, the output features after dimension expansion are partitioned into head, trunk and lower limbs for feature decoupling, and the step of generating three sets of local feature representations corresponding to three local adjacency topological matrices includes: decomposing the original adjacency matrix in the output features into three independent sub-matrices of head, trunk and lower limbs, which are expressed as follows: ,in Respectively represent the independent skeletal parts of the head, trunk and lower limbs; the three independent sub-matrices are refined into corresponding feature information to obtain three sets of local feature representations, which are expressed as follows: ,in Represents independent bone features of the head, torso, and legs respectively.
[0012] As an optional implementation of the first aspect of the present application, based on the three local feature representations and the one global feature representation, the step of calculating the corresponding prediction probability distribution includes: performing regularization processing on the three local feature representations and the one global feature representation, respectively, and generating the corresponding prediction probability distribution through a fully connected layer and a Softmax function.
[0013] As an optional implementation of the first aspect of the present application, in the step of calculating the combined loss function, the combined loss function is calculated using label smoothed cross entropy loss, and its weight parameters are dynamically adjusted through learnable parameters. The combined loss function is: ,in, represents the label smoothed cross entropy loss function, y represents the true label, Represents the predicted action classification result, and Respectively represent the global features and local features of different parts, Represents the learnable weight parameters between different parts.
[0014] In a second aspect, an embodiment of the present application provides a three-dimensional skeleton human body motion recognition system, the system comprising: A data acquisition module, used for acquiring three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; A multi-style topology matrix construction module, used to construct a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limb partitions respectively and a global adjacency topology matrix; A spatiotemporal feature extraction module, used to extract spatiotemporal features from the three-dimensional skeletal motion sequence data using the multi-style adjacency topology matrix to guide the graph convolutional network to obtain output features; A feature refinement module, used for refining the output features to generate three local feature representations and one global feature representation; A probability calculation module, used to calculate a corresponding prediction probability distribution based on the three local feature representations and the one global feature representation; A loss function calculation module, used to calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; A model training module, used for training the graph convolutional network according to the combined loss function; The action recognition module is used to output the predicted action classification result corresponding to the three-dimensional skeletal action sequence data based on the trained graph convolutional network.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] Compared with the prior art, the present invention proposes a three-dimensional skeleton human body action recognition method, which has the following beneficial effects: (1) Multi-style adjacency topology matrix module: The traditional single adjacency topology matrix is decomposed into independent sub-matrices for the head, torso, and lower limbs. Each sub-matrix independently processes local bone features and is fused with the global topology matrix through a spatial mapping aggregation mechanism to construct a multi-level topological structure, thereby enhancing the model's ability to capture local contextual features.
[0018] (2) Refined Skeleton Feature Module: Decompose the features after graph convolution processing to generate local feature representations corresponding to the partition topology matrix, introduce a multi-part contrast loss function, and combine it with label smoothed cross entropy loss to achieve complementary learning of global and local features, thereby improving the model's ability to recognize actions of occluded parts.
[0019] (3) Overall framework: The multi-style adjacency topology matrix module and the refined skeleton feature module are integrated into the graph convolutional network, sequence features are extracted through spatiotemporal convolution blocks, and the loss function is optimized to improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of a method for recognizing full-body movements of a three-dimensional skeleton human body provided by the first embodiment of the present invention; Figure 2 is an overall architecture diagram of the first embodiment of the present invention; Figure 3 is a segmented multi-style bone structure diagram in the first embodiment of the present invention; Figure 4 is an overall flow chart of the multi-style topology matrix (CRKC) method in the first embodiment of the present invention; Figure 5 is an overall flow chart of the skeleton feature refinement (CRKC) method in the first embodiment of the present invention; Figure 6 It is a structural schematic diagram of a three-dimensional skeleton human body full-body motion recognition system provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0022] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0023] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0024] Example 1 See also Figure 1 , which is a flow chart of a three-dimensional skeleton human body full body motion recognition method proposed in the first embodiment of the present application; please refer to Figure 2 , is the overall architecture of this embodiment, which mainly includes a multi-style topological matrix (MTM) module and a refined skeleton feature (CRKC) module. In order to make the correlation between different connection points of the topological matrix between bones closer, the MTM module divides the original fixed topological matrix into independent adjacency topological matrices at the top, middle and bottom of the skeleton parts, and uses the strong adaptability between local features and global features to aggregate them into a new adjacency topological matrix through spatial mapping, which helps the model obtain data and features in the entire graph convolution, so that the topological relationship between different bones can be fully learned. The CRKC module refines the data features processed by the graph convolution into local modules corresponding to the adjacency topological matrix, and proposes a multi-part contrast loss, which performs label smoothing loss on the global and local independent prediction distribution results and the true label, thereby strengthening the training loss prediction results of the occluded part features, improving its robustness and adaptability, and enhancing the fine-grainedness between data features.
[0025] The proposed method steps are as follows: S1: Acquire three-dimensional skeletal motion sequence data, where the three-dimensional skeletal motion sequence data includes three-dimensional coordinates of human joints in multiple time frames.
[0026] Specifically, in the data preprocessing stage, the format of the action sequence is unified as , where M, T, and V represent the number of people, frames, and joints, respectively. For the three-dimensional bone coordinates, it is expressed as , , , ] and then Divided into three separate features: Upper head ; Middle torso ; Lower Legs .
[0027] The above represents the coordinate positions of the bone joints in different parts. Segmentation of various bone structures such as Figure 3 As shown in Figure 2, each independent bone point is mapped to each other to enhance the fine-grainedness and adaptability of feature information.
[0028] S2: Construct a multi-style adjacency topology matrix, which includes three local adjacency topology matrices corresponding to the head, trunk and lower limb partitions and a global adjacency topology matrix.
[0029] In this embodiment, the human skeletal joints are divided into three independent subsets including the head, torso and lower limbs according to human anatomy to perform functional areas of the human body; a local adjacency topology matrix is generated for each independent subset, and the local adjacency topology matrix only contains the connection relationship between the joints in the corresponding subset; the three local adjacency topology matrices are fused with a global adjacency topology matrix through a spatial mapping aggregation mechanism to form a multi-level topological structure.
[0030] It should be noted that the distance between different skeleton key points reflects the changes in the trajectory of human body movements. The position coordinates of the key skeleton points of different movements are also different, which further tests the generalization ability of the model to the movement. However, these important information are all reflected through the adjacency topology matrix. Figure 2 and Figure 3 As shown in the figure, each independent matrix corresponds to the color of the segmented bone structure) can improve the model's ability to obtain key information of each action, enhance the interaction of different bone joints in each frame sequence, and on this basis consider strengthening the correspondence between the topological matrices of global and local features, effectively improving the ability of action recognition.
[0031] S3: Use a multi-style adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features from 3D skeletal motion sequence data to obtain output features.
[0032] It should be noted that in the graph convolution stage, the adjacency topology matrix A and the self-learning topology weight matrix W , where k corresponds to different parts, Indicates the number of input channels, Represents the number of output channels, S represents the size of the kernel, and the non-Euclidean structure data of the skeleton data is modeled. The relationship between the skeleton joints represented by the topological matrix is used to guide the learning of the model.
[0033] Specifically, for the adjacency topology matrix property, it has a Graph of bone nodes: ,in Represents a collection of nodes, Represents the edge set, graph Represents a The matrix of Representation Node and nodes The neighboring relationship between them. In the initial graph convolution, the output feature is represented as: , in, represents the output features, represents the activation function, represents the adjacency topology matrix, represents the input features, represents the learnable adjacency weight parameter matrix.
[0034] The existing human body action recognition method based on three-dimensional skeleton uses the bone points corresponding to the single whole adjacency topological matrix to extract features, ignoring the topological relationship of the bone adjacency matrix, and the topological relationship learning is not comprehensive enough, which affects the overall model training results and leads to a decrease in classification accuracy. Therefore, in order to solve this problem and improve the connection between different bone joints, the MNTM method is proposed to solve the above problem.
[0035] Considering the relationship between the corresponding connection points between the adjacency topological matrices of different bone points, the original adjacency matrix in the output feature is decomposed into three independent sub-matrices of head, trunk and lower limbs using the idea of multi-graph adjacency matrix, which can be expressed as: , in Represents separate skeletal parts of the head, trunk, and lower limbs.
[0036] The three independent sub-matrices are refined into corresponding feature information to obtain three sets of local feature representations, which can be expressed as follows: ,in Represents independent bone features of the head, torso, and legs respectively.
[0037] In this embodiment, Einstein's summation law is used in the graph convolutional network to perform matrix multiplication on the three-dimensional skeletal motion sequence data to obtain output features.
[0038] The most notable feature of graph convolution is the matrix multiplication between adjacency topology matrices, which can be used (Einstein's summation law): , in Matrix correspondence Dimension, V represents the number of joints, represents the dimension of the feature channel, An index representing a spatial dimension; Feature correspondence Dimensions, represents the number of samples, Indicates the number of channels, represents the number of time frames, V represents the number of joint points, Representation and The same spatial dimension index. Specific operation process: The elements of the tensor dimension are The feature tensor dimension elements are multiplied and added according to the corresponding index to obtain a new tensor with dimension , where the meaning is similar to the above, this tensor effectively integrates the information of two tensors in different dimensions.
[0039] To this end, based on the architecture of graph convolution, the output features are expressed as: , By aggregating global and local multiple nonlinear adjacency matrices, supporting linear relationships between different topological matrices, optimizing the coupling between the feature information of skeletal points in different parts, forming more specific skeletal features, and fully associating the feature information of each part in the graph convolution, it promotes the pre-function of action recognition.
[0040] S4: Refine the output features to generate three local feature representations and one global feature representation.
[0041] In this embodiment, the output features are dimensionally expanded through deep convolution and spatial point convolution; the output features after dimension expansion are decoupled according to the head, torso and lower limb partitions to generate three groups of local feature representations corresponding to three local adjacency topology matrices.
[0042] like Figure 4 As shown in Figure 1, an example of the MTM method is shown. Given the original bone adjacency matrix and bone features, according to previous research methods for extracting features, independent spatiotemporal convolution blocks are used to extract bone information, while the MTM method uses Deep convolution and In the spatial point convolution method, k corresponds to the receptive field of the skeleton information. The deep convolution interconnects the independent input and output channels of each convolution kernel, increases the receptive field of the feature information, reduces the number of feature parameters, and improves the convolution efficiency. The spatial point convolution constructs the relationship between the input channels point by point, extends the feature dimension, and then divides the global adjacency topology matrix and the skeleton features into three parts of feature information: upper (head), middle (trunk), and lower (legs). Each part is aggregated in the activation function. The activation function uses the GELU method to supplement the training effect, and finally aggregates the new global information point by point through spatial point convolution.
[0043] S5: Based on the three local feature representations and one global feature representation, calculate the corresponding prediction probability distribution.
[0044] In this embodiment, regularization is performed on three local feature representations and one global feature representation, and corresponding prediction probability distributions are generated through a fully connected layer and a Softmax function.
[0045] Specifically, in the original action recognition method, the Softmax function and the cross entropy loss function are often used to represent the function of predicting the probability distribution of each category and the difference between the measured predicted value and the true label to judge the result of action recognition.
[0046] The Softmax function is expressed as: , in . is the first The element value, Any element The sum of Indicates The output feature vector is indexed and then divided by the sum of all vector indexes to obtain the result. The probability value of the output.
[0047] The cross entropy loss function is expressed as: , in is the real label data, is the bone feature information, Represents the difference between the predicted value and the true label.
[0048] For skeleton-based human body action recognition, the discriminant features of skeleton information are particularly important. The original research method cannot guarantee that a certain action category is classified into the correct category. During the classification comparison training, only the global features are used for comparison loss training. The correlation between skeleton feature information is poor, and the recognition of occluded joint parts is not accurate. In order to better represent the discriminant features, the CRKC method further refines the output features and divides them into three parts: head, torso, and legs for comparison to represent the similarity between the actions of each part.
[0049] S6: Calculate the combined loss function, which is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label.
[0050] In this embodiment, the calculation of the combined loss function adopts label smoothed cross entropy loss, and its weight parameters are dynamically adjusted through learnable parameters. The combined loss function is: , in, represents the label smoothed cross entropy loss function,y represents the true label, Represents the predicted action classification result, and Respectively represent the global features and local features of different parts, Represents the learnable weight parameters between different parts. The global and local features and the true label contrast loss are integrated to optimize the parameters between models, so that the model can fit the data more fully and improve the model's generalization ability for feature data.
[0051] Specifically, based on the original function, a multi-part contrast loss method is proposed to distinguish the difference between predicting real actions and blurred actions, and improve the intra-class consistency of a certain action, as shown in the following formula: in, is a parameter that controls the weight of the label smoothed cross entropy loss, , , They represent the predicted values of the corresponding parts respectively. Represents the difference between the actual action and the predicted action S7: Training graph convolutional networks based on combined loss functions.
[0052] S8: Based on the trained graph convolutional network, the predicted action classification results corresponding to the 3D skeletal action sequence data are finally output.
[0053] like Figure 5 As shown in the figure, the overall flow chart of CRKC is shown. The method first inputs a series of shapes as The skeleton is represented as The T-frame information of the joint. The graph convolution main chain consists of 11 basic units. The graph convolution block consists of TCN (temporal convolution) and GCN, collectively referred to as the TGN block. This block applies CNN is used to extract sequence features, and a learnable adjacency topology matrix is used to extract spatial skeleton point features. Note that the cross-sequence block in the basic unit is a method used to reduce the time dimension and increase the channel feature dimension. The output features of the graph convolution are refined, and the global and local feature information are regularized to prevent overfitting of the data. Then, the data is projected into a fully connected layer with a softmax activation function to predict the category distribution probability. Finally, the label smoothing loss is used to train with the real label to find the real action category.
[0054] Experimental results show that in the NTU RGB+D 60 dataset, the accuracy of X-sub (Top-1) and X-view (Top-1) are 93.22% and 97.10% respectively, and in the NTU RGB+D 120 dataset, the accuracy of X-sub (Top-1) and X-set (Top-1) are 90.30% and 91.61% respectively, reaching the current advanced level.
[0055] Example 2 See also Figure 6 , which is a schematic diagram of the structure of a three-dimensional skeleton human body motion recognition system proposed in the second embodiment of the present application, and the system includes: The data acquisition module 100 is used to acquire three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; A multi-style topology matrix construction module 200, for constructing a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limbs, respectively, and a global adjacency topology matrix; A spatiotemporal feature extraction module 300 is used to extract spatiotemporal features from the three-dimensional skeletal motion sequence data using the multi-style adjacency topology matrix to guide the graph convolutional network to obtain output features; A feature refinement module 400 is used to refine the output features to generate three local feature representations and one global feature representation; A probability calculation module 500, configured to calculate a corresponding predicted probability distribution based on the three local feature representations and the one global feature representation; A loss function calculation module 600, used to calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; A model training module 700, configured to train the graph convolutional network according to the combined loss function; The action recognition module 800 is used to output the predicted action classification result corresponding to the three-dimensional skeleton action sequence data based on the trained graph convolutional network.
[0056] A three-dimensional skeleton human body whole body motion recognition system in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.
[0057] The device of a three-dimensional skeleton human body motion recognition system in the embodiment of the present application can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0058] The three-dimensional skeleton human body full body motion recognition system provided in the embodiment of the present application can realize Figure 1 The various processes implemented by a three-dimensional skeleton human body full body motion recognition method in the method embodiment will not be described here to avoid repetition.
[0059] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned three-dimensional skeletal human body whole body motion recognition method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0060] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the embodiment of the above-mentioned three-dimensional skeletal human body whole-body motion recognition method are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0061] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0062] It should be noted that, in this article, the term "comprises", "includes" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0063] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0064] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A three-dimensional skeleton human body action recognition method, characterized in that: The method comprises the following steps: Acquire three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; Constructing a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limbs partitions respectively and a global adjacency topology matrix; Using the multi-style adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features from the three-dimensional skeletal motion sequence data to obtain output features; Refining the output features to generate three local feature representations and one global feature representation; Based on the three local feature representations and the one global feature representation, calculating a corresponding predicted probability distribution; Calculating a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; Training the graph convolutional network according to the combined loss function; Based on the trained graph convolutional network, the predicted action classification result corresponding to the three-dimensional skeletal action sequence data is finally output.
2. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: The steps to construct a multi-style adjacency topology matrix include: The human skeleton joints are divided into three independent subsets including the head, trunk and lower limbs according to the human anatomy. Generate a local adjacency topology matrix for each independent subset, wherein the local adjacency topology matrix only contains the connection relationship of the joint points in the corresponding subset; The three local adjacency topology matrices are fused with a global adjacency topology matrix through a spatial mapping aggregation mechanism to form a multi-level topological structure.
3. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: The step of using the multi-style adjacency topology matrix to guide the graph convolutional network to extract spatiotemporal features from the three-dimensional skeletal motion sequence data to obtain output features includes: In the graph convolutional network, the Einstein summation law is used to perform matrix multiplication on the three-dimensional skeletal motion sequence data to obtain output features.
4. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: The step of refining the output features to generate three local feature representations includes: The output features are dimensionally expanded through deep convolution and spatial point convolution; The output features after dimension expansion are decoupled according to the head, torso and lower limb partitions to generate three groups of local feature representations corresponding to three local adjacency topology matrices.
5. A three-dimensional skeleton human body motion recognition method according to claim 4, characterized in that: The steps of decoupling the output features after dimension expansion according to the head, torso and lower limb partitions to generate three sets of local feature representations corresponding to three local adjacency topology matrices include: The original adjacency matrix in the output feature is decomposed into three independent sub-matrices of head, torso and lower limbs, which can be expressed as follows: ,in Representing separate skeletal parts of the head, trunk, and lower limbs; The three independent sub-matrices are refined into corresponding feature information to obtain three sets of local feature representations, which are expressed as follows: ,in Represents independent bone features of the head, torso, and legs respectively.
6. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: Based on the three local feature representations and the one global feature representation, the step of calculating the corresponding predicted probability distribution comprises: Regularization is performed on the three local feature representations and the one global feature representation, respectively, and corresponding prediction probability distributions are generated through a fully connected layer and a Softmax function.
7. A three-dimensional skeleton human body motion recognition method according to claim 1, characterized in that: In the step of calculating the combined loss function, the combined loss function is calculated using label smoothed cross entropy loss, and its weight parameters are dynamically adjusted through learnable parameters. The combined loss function is: , in, represents the label smoothed cross entropy loss function, y represents the true label, represents the predicted action classification result, and Respectively represent the global features and local features of different parts, Represents the learnable weight parameters between different parts.
8. A three-dimensional skeleton human body action recognition system, characterized in that: The system comprises: A data acquisition module, used for acquiring three-dimensional skeleton motion sequence data, wherein the three-dimensional skeleton motion sequence data includes three-dimensional coordinates of human joints in multiple time frames; A multi-style topology matrix construction module, used to construct a multi-style adjacency topology matrix, wherein the multi-style adjacency topology matrix includes three local adjacency topology matrices corresponding to the head, trunk and lower limb partitions respectively and a global adjacency topology matrix; A spatiotemporal feature extraction module, used to extract spatiotemporal features from the three-dimensional skeletal motion sequence data using the multi-style adjacency topology matrix to guide the graph convolutional network to obtain output features; A feature refinement module, used for refining the output features to generate three local feature representations and one global feature representation; A probability calculation module, used to calculate a corresponding prediction probability distribution based on the three local feature representations and the one global feature representation; A loss function calculation module, used to calculate a combined loss function, where the combined loss function is a weighted sum of the loss between the global predicted probability distribution and the true label and the loss between each local predicted probability distribution and the true label; A model training module, used for training the graph convolutional network according to the combined loss function; The action recognition module is used to output the predicted action classification result corresponding to the three-dimensional skeletal action sequence data based on the trained graph convolutional network.
9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a three-dimensional skeletal human body whole body motion recognition method as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a three-dimensional skeletal human body whole body motion recognition method as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Terminal passenger flow volume space-time distribution prediction method based on graph convolution network
CN112257614A
Human body 3D skeleton behavior recognition method based on graph attention convolutional neural network
CN114613011A
End-to-end human behavior recognition method and model based on skeleton nodes
CN114613013A
Action recognition method based on dynamic local-global graph convolutional neural network
CN114998525A
Multi-granularity human body action classification method based on graph convolutional network
CN115116139A
Cited By
Self-adaptive perception digital human skeleton binding method and system based on dynamic graph convolution
CN120726193A
Adaptive perceptual digital human skeleton binding method and system based on dynamic graph convolution
CN120726193B
Action recognition method and system based on human skeleton structure
CN121861729A