Human body behavior recognition method and system
By setting up a multi-headed spatial hypergraph convolution module and virtual feature expansion in the ST-GCN model, an adaptive hypergraph structure is built, which solves the problem of fixed bone topology accuracy and joint number in the existing technology, and improves the accuracy and generalization ability of human behavior recognition.
Patent Information
- Application Number
- CN202510512268.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-12
AI Technical Summary
The existing methods for modeling bone topology through hypergraphs are based on prior knowledge definitions or other fixed hypergraph definitions, resulting in poor bone topology accuracy and fixed number of bone joints in the input features of the behavior recognition model, making it difficult to accurately understand and recognize human behavior, resulting in poor human behavior recognition accuracy.
A multi-head spatial hypergraph convolution module is set up before the convolution of each time graph of the ST-GCN model. The multi-head spatial hypergraph convolution module divides the bone sequence data into multiple sub-features, and constructs a hypergraph. The virtual features are used to expand the joint dimensions, combine the physical adjacency matrix and the association matrix to generate modal data spatial features, and adopts an adaptive hyper-edge construction method and dense connection structure to optimize the loss function to improve the recognition accuracy.
The refined disassembly of complex human body movement characteristics can be realized, and the real relationship between human joints in behavior and actions can be expressed more comprehensively and accurately, the recognition accuracy and generalization ability of the human behavior recognition model are improved, and the limitation of the fixed joint number on the model capacity is broken.
Smart Images

Figure CN120472527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a human behavior recognition method and system. Background Art
[0002] Skeleton-based human action recognition has attracted considerable attention in the field of artificial intelligence, with widespread application in a variety of real-world scenarios, including video understanding, video surveillance, and human-computer interaction. Skeleton sequences, consisting of a series of 2D or 3D coordinates, can be collected using depth sensors or acquired using video-based pose estimation algorithms. Compared to RGB and optical flow images, skeletal data is highly popular in action recognition because it represents the fundamental physical structure of humans, offers low dimensionality, efficient information transmission, and robustness to lighting and scene variations.
[0003] To improve the accuracy of skeleton-based action recognition, both recurrent neural networks (RNNs) and convolutional neural networks (CNNs) have been extensively studied. However, RNNs themselves struggle to describe the inherent skeletal topology, while the learned filters of CNNs often ignore the spatiotemporal structure of the skeleton. Consequently, recent research has focused on directly modeling the skeletal topology. Because the physical topology of human joints and bones can be uniformly represented using a graph structure, graph convolutional networks (GCNs) are often introduced to aggregate the characteristic information of skeletal joints. Generally speaking, two adjacent joints can communicate information through their shared skeleton, which corresponds to the information exchange between two vertices along their connecting edges in a graph. Most existing methods assume binary connections between pairs of connected vertices, representing the topology of the adjacency matrix as a general graph. However, human actions are coordinated by multiple joints, involving not only binary relationships between node pairs but also multi-joint relationships. For example, the action of starting to run is represented by a person raising their left hand while simultaneously stepping forward with their right leg. Binary connections cannot fully capture this synergy. To overcome this problem, researchers have constructed hypergraphs to model the skeletal topology. Hypergraph topology involves multiple node connections, rather than the binary connections of ordinary graphs.
[0004] However, most of these studies constructing hypergraphs to model skeletal topology are based on prior knowledge definitions, such as using hyperedges to divide the human body's limbs and torso, or other fixed hypergraph definitions. Human movements are diverse, and the coordination patterns between joints within each action category are unique. Therefore, this construction approach is not conducive to the model's understanding of different action categories. Prior knowledge definitions and fixed hypergraph definitions focus on the relatively macroscopic structural divisions of the human body, but different action categories often differ in the subtle joint connections. For example, the two similar action categories of writing and drawing may have similar hyperedge configurations from the perspective of the overall limb and torso division. However, in reality, the coordination between finger joints in writing is detailed and regular, while the coordination between hand and arm joints in drawing differs in details from writing. The macroscopic structural division cannot delve into these subtle joint connections, resulting in a lack of sufficient information for the model to distinguish between these similar action categories, affecting accurate understanding of the movements and poor accuracy in human action recognition. Furthermore, the capacity of action recognition models is also limited by their input features. In various existing studies and datasets, the number of joints used to describe the human skeletal structure is often a fixed value, usually set at 25. The core of behavior recognition lies in accurately capturing the interactions between skeletal joints. If the model only analyzes based on the fixed 25 joints, when encountering complex movements, there may not be enough "joint data points" to fully present these key joint coordination relationships, resulting in the model being unable to accurately understand and recognize human behavior. Summary of the Invention
[0005] To this end, the technical problem to be solved by the present invention is to overcome the defects that the existing method of constructing a hypergraph to model skeletal topology adopts a priori knowledge definition or other fixed hypergraph definition methods, the constructed skeletal topology accuracy is poor, and the number of skeletal joints in the input features of the behavior recognition model is fixed, which makes it difficult for the model to accurately understand and recognize human behavior, resulting in poor accuracy of human behavior recognition.
[0006] To solve the above technical problems, the present invention provides a human behavior recognition method, comprising:
[0007] A multi-head spatial hypergraph convolution module is set before each temporal graph convolution in the ST-GCN model to build a human behavior recognition model;
[0008] Convert the to-be-identified skeletal sequence data into different modal data, input the physical adjacency matrix of the to-be-identified skeletal sequence data and each modal data into the human behavior recognition model respectively, output the predicted probability of each modality belonging to each type of human behavior, and fuse them to obtain the predicted probability of the to-be-identified skeletal sequence data belonging to each type of human behavior;
[0009] Among them, the modal data features input to the multi-head spatial hypergraph convolution module are divided into multiple sub-features along the channel dimension;
[0010] The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature;
[0011] According to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes, the nodes included in each hyperedge are determined, and the hyperedge of each sub-feature is constructed;
[0012] Based on the nodes included in each hyperedge, the association matrix of each sub-feature is obtained; based on the mapping features of all sub-features, the weight matrix of each sub-feature is obtained;
[0013] Construct a hypergraph for each sub-feature based on its nodes, hyperedges, association matrix, and weight matrix.
[0014] The modal data space features are generated based on the physical adjacency matrix, the normalized correlation matrix corresponding to each sub-feature and its hypergraph.
[0015] Preferably, the process of obtaining sub-features includes:
[0016] Constructing and feeding modal data features into the multi-head spatial hypergraph convolution module Virtual features with the same channel and time dimensions Among them, C in F in The channel dimension of F in The time dimension, V is F in The joint dimension, V h F h joint dimensions;
[0017] The virtual feature F h and the modal data features F input to the multi-head spatial hypergraph convolution module in After splicing along the joint dimension, the real-virtual modal data features are obtained The real-virtual modality data features are divided into multiple sub-features along the channel dimension.
[0018] Preferably, during the training of the human behavior recognition model, the loss function used is:
[0019]
[0020] in, is the loss function of the human action recognition model, is the cross entropy loss function between the predicted probability of the skeleton sequence data to be identified belonging to various human behaviors and the true label, is the difference loss of the virtual features corresponding to the l-th multi-head spatial hypergraph convolution module, is the joint dimension of the virtual feature corresponding to the l-th multi-head spatial hypergraph convolution module, C l is the cosine matrix of the virtual features corresponding to the l-th multi-head spatial hypergraph convolution module, is the virtual feature corresponding to the l-th multi-head spatial hypergraph convolution module, (·) T represents transpose, ‖·‖ 2 represents the Euclidean norm, a is the row index of the cosine matrix, b is the column index of the cosine matrix, C l For the element in row a and column b, ReLU(·) is the ReLU activation function.
[0021] Preferably, the human behavior recognition model includes a feature embedding layer, multiple hypergraph convolution layers, a global average pooling layer, and a category mapping layer connected in sequence along the forward propagation direction, and each hypergraph convolution layer includes a multi-head spatial hypergraph convolution module and a temporal graph convolution connected in sequence.
[0022] Preferably, every three supergraph convolution layers are regarded as a stage, and the output features of the first supergraph convolution layer and the output features of the second supergraph convolution layer in each stage are added element by element as the input of the third supergraph convolution layer.
[0023] Preferably, the mapping features based on all sub-features to obtain a weight matrix for each sub-feature includes:
[0024] The mapping features of each sub-feature are sequentially passed through the dimensionality reduction mapping function and the LeakeyReLU activation function to obtain the activation features of each sub-feature;
[0025] After connecting the activation feature channels of all sub-features, the weights are used to generate the mapping function and the Tanh activation function as the weight matrix corresponding to each sub-feature.
[0026] The formula for obtaining the weight matrix corresponding to each sub-feature is:
[0027]
[0028] Among them, W is the weight matrix corresponding to each sub-feature, is the mapping feature of the kth sub-feature, is the dimensionality reduction mapping function corresponding to the kth sub-feature, σ 1 (·) is the LeakeyReLU activation function, represents the channel connection operation, K is the number of sub-features, k is the sub-feature index, Ψ 2 (·) is the weight generation mapping function, σ 2 (·) is the Tanh activation function, Cin is the channel dimension of the modal data feature input to the multi-head spatial hypergraph convolution module, K is the number of sub-features, and C h is the hidden layer dimension of the mapping subspace, Ψ 2 are learnable parameters.
[0029] Preferably, determining the nodes included in each hyperedge according to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes includes:
[0030] A hyperedge is constructed with each node as the centroid. Based on the distance between each node and other nodes, the K nearest neighbor algorithm is used to obtain the K nodes closest to each node. m nodes, as the initial nodes included in the hyperedge corresponding to the node; where K m Hyperparameters for human action recognition models;
[0031] Count the number of times each node is included in a hyperedge. If the number of times the current node is included in a hyperedge exceeds the maximum number of times the node can be included in a hyperedge, then sort these hyperedges in descending order based on the distance between the current node and the node corresponding to the hyperedge containing the node. Delete the node in the hyperedge that exceeds the maximum number of times the node is included after sorting, so as to determine the nodes included in each hyperedge.
[0032] Preferably, the calculation formula for each element in the correlation matrix corresponding to each sub-feature is:
[0033]
[0034] Among them, set i is the node set contained in the hyperedge corresponding to the i-th node, i is the first node index, j is the second node index, i≠j, exp(·) is the exponential function with a natural constant as the base, m i,j is the distance between the i-th node and the j-th node, m i,k For the i-th node and set i The distance between the kth nodes in set i Midpoint index.
[0035] Preferably, the formula for generating the modal data spatial feature is:
[0036]
[0037] Among them, F out is the modal data space feature, Indicates a channel connection operation, is the physical adjacency matrix, α is the topological fusion learnable parameter, is the normalized association matrix corresponding to the hypergraph of the k-th sub-feature, is the kth sub-feature, K is the number of sub-features, P k The feature transformation corresponding to the kth sub-feature can learn the weight parameter, and k is the sub-feature index.
[0038] Preferably, the process of obtaining the mapping feature of each sub-feature includes:
[0039] After pooling each sub-feature in the time dimension, the mapping function is used to obtain the mapping features of each sub-feature.
[0040] The present invention also provides a human behavior recognition system, comprising:
[0041] The model construction module is used to set a multi-head spatial hypergraph convolution module before each temporal graph convolution of the ST-GCN model to build a human behavior recognition model;
[0042] The recognition module is used to convert the skeleton sequence data to be identified into different modal data, input the physical adjacency matrix of the skeleton sequence data to be identified and each modal data into the human behavior recognition model respectively, output the predicted probability of each modality belonging to each type of human behavior, and fuse the predicted probability of the skeleton sequence data to be identified belonging to each type of human behavior;
[0043] Among them, the modal data features input to the multi-head spatial hypergraph convolution module are divided into multiple sub-features along the channel dimension;
[0044] The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature;
[0045] According to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes, the nodes included in each hyperedge are determined, and the hyperedge of each sub-feature is constructed;
[0046] Based on the nodes included in each hyperedge, the association matrix of each sub-feature is obtained; based on the mapping features of all sub-features, the weight matrix of each sub-feature is obtained;
[0047] Construct a hypergraph for each sub-feature based on its nodes, hyperedges, association matrix, and weight matrix.
[0048] The modal data space features are generated based on the physical adjacency matrix, the normalized correlation matrix corresponding to each sub-feature and its hypergraph.
[0049] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0050] The present invention discloses a method and system for human behavior recognition. The present invention sets a multi-head spatial hypergraph convolution module before each time graph convolution of the ST-GCN model. The module divides the modal data features input to the multi-head spatial hypergraph convolution module into multiple sub-features along the channel dimension. Each sub-feature contains skeletal joint information of a specific dimension, thereby realizing a refined decomposition of complex human motion features and helping the human behavior recognition model to accurately capture the characteristics of each joint. The sub-features are subjected to a mapping function to obtain the mapping features of each sub-feature. While retaining the original features of the sub-features, spatial correlation information such as the relative position and connection relationship between joints is integrated into the mapping features. By independently constructing a hypergraph for each sub-feature under the multi-head structure, each branch can not only reduce the amount of calculation, but also mine action features from multiple perspectives. By integrating the normalized association matrix corresponding to each sub-feature and its hypergraph, the physical adjacency matrix reflecting the relationship between each joint is integrated to generate modal data space features. The joint relationship can be integrated from different levels and angles, so that the modal data space features can more comprehensively and accurately express the real relationship between human joints in behavioral actions, thereby improving the recognition accuracy of the human behavior recognition model.
[0051] The present invention determines the nodes included in each hyperedge based on two key parameters: the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes. Compared with the traditional uniform hypergraph, the number of joints included in each hyperedge in the present invention is not fixed, allowing each hyperedge to extract more differentiated associations. It can adaptively incorporate synergistic joints into the same hyperedge according to the characteristics of different actions, making the hypergraph structure more in line with the collaborative pattern between joints when the action actually occurs.
[0052] The present invention constructs virtual features with the same channel dimension and time dimension as the modal data features input to the multi-head spatial hypergraph convolution module; after splicing the virtual features with the modal data features input to the multi-head spatial hypergraph convolution module along the joint dimension, a real-virtual feature is obtained; that is, virtual nodes are added to the real bone nodes. The addition of virtual nodes expands the joint dimension of the modal data features, providing more data points for the model, so that the model can have richer "joint data points" when processing complex movements, thereby more comprehensively capturing the interaction and synergy between skeletal joints, breaking through the limitation of the fixed number of joints on the model capacity, and improving the model's recognition accuracy of complex human behaviors.
[0053] In addition, the present invention regards every three hypergraph convolution layers as a stage, and the output features of the first hypergraph convolution layer in each stage are added element by element to the output features of the second hypergraph convolution layer as the input of the third hypergraph convolution layer. Through this multi-stage dense connection, the feature transition between layers can be smoothed, so that the difference between the input of the deep network and the output of the shallow network can be effectively controlled, so that the human behavior recognition model can better comprehensively utilize features at different levels and improve the accuracy of human behavior recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0055] Figure 1 It is a flow chart of a human behavior recognition method of the present invention.
[0056] Figure 2 It is a structural diagram of a human behavior recognition method of the present invention.
[0057] Figure 3 This is a schematic diagram of the related structure of the multi-head spatial hypergraph convolution module. Figure 3 (a) is the overall structure diagram of the multi-head spatial hypergraph convolution module. Figure 3 (b) in the figure is a schematic diagram of the construction process of the hypergraph of each sub-feature. DETAILED DESCRIPTION
[0058] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0059] Reference Figure 1 As shown, this embodiment provides a human behavior recognition method, including the following steps:
[0060] Step S1: Set a multi-head spatial hypergraph convolution module before each temporal graph convolution of the ST-GCN model to build a human action recognition model (Hyper-GCN);
[0061] In this embodiment, specifically, the human behavior recognition model includes a feature embedding layer, multiple hypergraph convolution layers, a global average pooling layer, and a category mapping layer connected in sequence along the forward propagation direction. Each hypergraph convolution layer includes a multi-head spatial hypergraph convolution module and a temporal graph convolution connected in sequence.
[0062] In this embodiment, preferably, every three supergraph convolutional layers are regarded as a stage, and the output features of the first supergraph convolutional layer and the output features of the second supergraph convolutional layer in each stage are added element by element as the input of the third supergraph convolutional layer.
[0063] The present invention regards every three hypergraph convolution layers as a stage, and the output features of the first hypergraph convolution layer in each stage are added element-by-element to the output features of the second hypergraph convolution layer as the input of the third hypergraph convolution layer. Through this dense connection, the feature transition between layers can be smoothed, so that the difference between the input of the deep network and the output of the shallow network can be effectively controlled. The key information related to the basic structure of the human skeleton and the basic action units in the shallow network can be directly transmitted to the deep network, avoiding loss in the process of layer-by-layer feature transformation, so that the human behavior recognition model can better comprehensively utilize features at different levels, thereby improving the model's recognition accuracy and generalization ability of human behavior.
[0064] like Figure 3 As shown, Figure 3 This is a schematic diagram of the related structure of the multi-head spatial hypergraph convolution module. Figure 3 (a) is the overall structure diagram of the multi-head spatial hypergraph convolution module. Figure 3 (b) is a schematic diagram of the construction process of the hypergraph of each sub-feature. In the figure, A-NHG is the construction method of the hypergraph of each sub-feature constructed by the present invention.
[0065] In this embodiment, preferably, the modal data features input to the multi-head spatial hypergraph convolution module Divided into K sub-features along the channel dimension; among them, the F corresponding to the k-th sub-feature in The channel range is [(k-1)×C in / K+1,…,k×C in / K];
[0066] After pooling each sub-feature in the time dimension, each pooled sub-feature is obtained After each sub-feature corresponding mapping function Get the mapping features of each sub-feature Among them, Φ k is the mapping function corresponding to the k-th sub-feature, and the mapping feature of the k-th sub-feature Each row in represents the feature vector of a joint;
[0067] The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature;
[0068] According to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes, the nodes included in each hyperedge are determined, and the hyperedge corresponding to each sub-feature is constructed; the distance between each node and other nodes is calculated as follows:
[0069]
[0070] Among them, m i,j is the distance between the i-th node and the j-th node, m j,i is the distance between the jth node and the ith node, v i is the i-th node, v j is the jth node.
[0071] In this embodiment, preferably, determining the nodes included in each hyperedge based on the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes includes:
[0072] A hyperedge is constructed with each node as the center of mass. Based on the distance between each node and other nodes, the K nearest neighbor algorithm is used to obtain the K nodes closest to each joint. m nodes, as the initial nodes included in the hyperedge corresponding to the node; where K m Hyperparameters for human action recognition models;
[0073] Count the number of times each node is included in a hyperedge. If the number of times the current node is included in a hyperedge exceeds the maximum number of times the node can be included in a hyperedge, then sort these hyperedges in descending order based on the distance between the current node and the node corresponding to the hyperedge containing the node. Delete the node in the hyperedge that exceeds the maximum number of times the node is included after sorting, so as to determine the nodes included in each hyperedge.
[0074] The present invention determines the nodes included in each hyperedge based on two key parameters: the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes. Compared with the traditional uniform hypergraph, the number of joints included in each hyperedge in the present invention is not fixed, allowing each hyperedge to extract more differentiated associations. It can flexibly incorporate synergistic joints into the same hyperedge according to the characteristics of different actions, making the hypergraph structure more in line with the collaborative pattern between joints when the action actually occurs.
[0075] The mapping features of each sub-feature are sequentially passed through the dimensionality reduction mapping function and the LeakeyReLU activation function to obtain the activation features of each sub-feature;
[0076] After connecting the activation feature channels of all sub-features, the weights are used to generate the mapping function and the Tanh activation function as the weight matrix corresponding to each sub-feature.
[0077] In this embodiment, preferably, the formula for obtaining the weight matrix corresponding to each sub-feature is:
[0078]
[0079] Among them, W is the weight matrix corresponding to each sub-feature, is the mapping feature of the kth sub-feature, is the dimensionality reduction mapping function corresponding to the kth sub-feature, σ 1 (·) is the LeakeyReLU activation function, represents the channel connection operation, K is the number of sub-features, k is the sub-feature index, Ψ 2 (·) is the weight generation mapping function, σ 2 (·) is the Tanh activation function, C in is the channel dimension of the modal data feature input to the multi-head spatial hypergraph convolution module, K is the number of sub-features, C h is the hidden layer dimension of the mapping subspace, Ψ 2 are learnable parameters.
[0080] The conventional weight matrix uses a diagonal matrix Indicates that, the calculation formula for each element is:
[0081]
[0082] Among them, w ij is the element in the i-th row and j-th column of the diagonal matrix. When i=j, w ij represents the weight of the i-th hyperedge, and M is the number of hyperedges.
[0083] Passing the mapping features of each sub-feature through the dimensionality reduction mapping function and the LeakeyReLU activation function in sequence can effectively reduce the channel dimension and remove redundant information while retaining the most critical discriminant information of each node. After retaining the most basic discriminant information of each node, the activated feature channels of all sub-features are connected and then passed through the weight generation mapping function and the Tanh activation function in sequence as the weight matrix corresponding to each sub-feature to obtain the weight matrix of the hypergraph. In the human behavior recognition scenario, human motion data often contains a large amount of complex and redundant information, such as the human joint movement trajectory, posture changes, and other data with high dimensions and repeated features. Through this operation, we can focus on the core features that can reflect behavioral differences, avoid the computational burden and overfitting risk brought by high-dimensional data, and enable the model to more efficiently learn the essential characteristics of human behavior.
[0084] The Tanh activation function generates weights of [-1, 1], greatly enhancing the model's expressive power. It can more finely characterize the correlation and importance of various sub-features in human behavior. For example, when identifying running and walking behaviors, the weights of sub-features such as the movement amplitude and speed of different body parts will be precisely adjusted due to this activation function, enabling the model to accurately distinguish similar behaviors and improve recognition accuracy.
[0085] The application of the LeakeyReLU activation function effectively avoids the problem of neuron death. During the training of human action recognition tasks, the traditional ReLU function may cause some neurons to stop updating their parameters due to negative input data, thus becoming ineffective. The LeakeyReLU activation function, however, assigns a small non-zero slope to negative inputs, ensuring that all neurons participate in training. This allows the model to continuously learn various features from human action data, improving the effectiveness and reliability of model training.
[0086] The present invention sequentially generates a mapping function and a Tanh activation function through the activation features of all sub-features as the weight matrix corresponding to each sub-feature; and integrates these multi-perspective information as the weight matrix corresponding to each sub-feature, which can comprehensively capture the complex interactions and dependencies between joints in human behavior, so that each constructed sub-feature hypergraph can accurately match the coordinated changes of joints in different action stages, effectively improving the human behavior recognition model's ability to understand each action.
[0087] Construct a hypergraph for each sub-feature based on its nodes, hyperedges, association matrix, and weight matrix.
[0088] The calculation formula for each element in the conventional correlation matrix is
[0089]
[0090] Among them, e j is the hyperedge corresponding to the j-th node.
[0091] The discrete values of the conventional association matrix are difficult to adjust through optimization algorithms such as gradient descent, which limits the flexibility of model training. To ensure that the model is trainable, in this embodiment, preferably, a softmax operation is used to convert the distance relationship between nodes into the association probability between nodes. The calculation formula for each element in the association matrix corresponding to each sub-feature is:
[0092]
[0093] Among them, set i is the node set contained in the hyperedge corresponding to the i-th node, i is the first node index, j is the second node index, i≠j, exp(·) is the exponential function with a natural constant as the base, mi,j is the distance between the i-th node and the j-th node, m i,k For the i-th node and set i The distance between the kth nodes in set i Midpoint index.
[0094] Based on the physical adjacency matrix and the normalized correlation matrix of each sub-feature and its hypergraph, the output features of the multi-head spatial hypergraph convolution module are generated.
[0095] In this embodiment, preferably, the formula for generating the modal data spatial feature is:
[0096]
[0097] Among them, F out is the modal data space feature, Indicates a channel connection operation, is the physical adjacency matrix, α is the topological fusion learnable parameter, is the normalized association matrix corresponding to the hypergraph of the k-th sub-feature, is the kth sub-feature, K is the number of sub-features, P k The feature transformation corresponding to the kth sub-feature can learn the weight parameter, and k is the sub-feature index.
[0098] The calculation formula of the normalized correlation matrix corresponding to each sub-feature is:
[0099]
[0100] in, is the normalized association matrix corresponding to the hypergraph of the k-th sub-feature, D k,v is the node degree matrix corresponding to the hypergraph of the kth sub-feature, D k,e is the hyperedge degree matrix corresponding to the hypergraph of the kth sub-feature, H k is the association matrix corresponding to the hypergraph of the k-th sub-feature.
[0101] For the node degree matrix and hyperedge degree matrix Elements of and is defined as follows:
[0102]
[0103] Among them, w j is the weight of the j-th hyperedge.
[0104] In this embodiment, preferably, the process of obtaining the sub-features includes:
[0105] Constructing and feeding modal data features into the multi-head spatial hypergraph convolution module Virtual features with the same channel and time dimensions Among them, C in F in The channel dimension of F in The time dimension, V is F in The joint dimension, V h F h joint dimensions;
[0106] The virtual feature F h and the modal data features F input to the multi-head spatial hypergraph convolution module in After splicing along the joint dimension, the real-virtual modal data features are obtained The real-virtual modality data features are divided into multiple sub-features along the channel dimension.
[0107] The present invention constructs virtual features with the same channel dimension and time dimension as the modal data features input to the multi-head spatial hypergraph convolution module; after splicing the virtual features with the modal data features input to the multi-head spatial hypergraph convolution module along the joint dimension, a real-virtual feature is obtained; that is, virtual nodes are added to the real bone nodes. The addition of virtual nodes expands the joint dimension of the modal data features, providing more data points for the model, so that the model can have richer "joint data points" when processing complex movements, thereby more comprehensively capturing the interactions and collaborative relationships between skeletal joints, breaking through the limitation of the fixed number of joints on the model capacity, and improving the model's understanding and recognition capabilities of complex human behaviors.
[0108] In this embodiment, preferably, during the training process of the human behavior recognition model, the loss function used is:
[0109]
[0110] in, is the loss function of the human action recognition model, is the cross entropy loss function between the predicted probability of the skeleton sequence data to be identified belonging to various human behaviors and the true label, is the difference loss of the virtual features corresponding to the l-th multi-head spatial hypergraph convolution module, is the joint dimension of the virtual feature corresponding to the l-th multi-head spatial hypergraph convolution module, C l is the cosine matrix of the virtual features corresponding to the l-th multi-head spatial hypergraph convolution module, is the virtual feature corresponding to the l-th multi-head spatial hypergraph convolution module, (·) T represents transpose, ‖·‖ 2 represents the Euclidean norm, a is the row index of the cosine matrix, b is the column index of the cosine matrix, C l For the element in row a and column b, ReLU(·) is the ReLU activation function.
[0111] After adding virtual nodes to real skeleton nodes, virtual connected hyperjoints are formed. These virtual nodes and real nodes together form a more complex hypergraph structure. Supervised learning of superjoints, which guide these superjoints to support generalizable driving features embedded in large amounts of data, aims to coordinate features at different network depths by setting independent superjoints at each layer of the network. These virtually connected superjoints involve spatial hypergraph convolutions rather than temporal convolutions.
[0112] In order to diversify these super-joints, the present invention proposes a difference loss for super-joint optimization to alleviate their homogeneity. In the difference loss, we use the average value of the cosine matrix C of each layer of the network to measure the difference of the super-joint. Since the correlation of the super-joint itself cannot be optimized, the present invention subtracts V h to manually remove this part.
[0113] like Figure 2 As shown, Figure 2 This is a structural diagram of a human behavior recognition method of the present invention.
[0114] Step S2: Convert the bone sequence data to be identified into different modal data, input the physical adjacency matrix of the bone sequence data to be identified and each modal data into the human behavior recognition model respectively, output the predicted probability of each modality belonging to each type of human behavior, and fuse the predicted probability of the bone sequence data to be identified belonging to each type of human behavior.
[0115] In this embodiment, specifically, the process of converting the to-be-identified skeleton sequence data into different modal data includes: sequentially subjecting the to-be-identified skeleton sequence data to data denoising, normalization, and cropping to obtain preprocessed to-be-identified skeleton sequence data;
[0116] The pre-processed bone sequence data to be identified is modally converted to obtain joint mode, bone mode, joint motion mode, and bone motion mode.
[0117] The joint modality is obtained by sequentially arranging the three-dimensional coordinate information of the joint points in the preprocessed bone sequence data to be identified; the bone modality is obtained by calculating the coordinate difference of adjacent joint points in the preprocessed bone sequence data to be identified, obtaining the direction and length information of the bone vector and arranging it; the joint motion modality is obtained by calculating the coordinate change of the joint point at different time steps; the bone motion modality is obtained by first obtaining the bone vectors at different times, and then calculating the difference of the bone vectors at adjacent times, so as to obtain the bone motion information and construct it.
[0118] In this embodiment, specifically, the process of obtaining the physical adjacency matrix corresponding to the to-be-identified skeleton sequence data includes:
[0119] By obtaining the topological subset S corresponding to the sequence data of the skeleton to be identified, id ,s cf ,s cp}, where s id 、s cf and s cp Topological subsets representing self-connectivity, centrifugality, and centripetality, respectively;
[0120] The adjacency matrices corresponding to the three topological subsets are graph-normalized and then added together to form the physical adjacency matrix corresponding to the bone sequence data to be identified.
[0121] Here, we take the skeleton structure of X nodes as an example:
[0122] s id The corresponding matrix is an X×X diagonal matrix, where only the elements on the diagonal are 1, indicating that each node is connected only to itself and not to other nodes.
[0123] s cf It represents a centrifugal connection method. Once an origin is determined, the direction of all edges is from the origin to the outside of the human body. In the adjacency matrix, it reflects the connection relationship that diverges from a specific node to other nodes.
[0124] s cp It represents the centripetal connection mode. Once an origin is determined, the directions of all edges point to the origin. The adjacency matrix reflects the connection relationship of other nodes converging to a specific central node.
[0125] In this embodiment, specifically, the predicted probability of each modality belonging to each type of human behavior is fused to obtain the predicted probability of the to-be-identified skeleton sequence data belonging to each type of human behavior. The fusion method includes:
[0126] The predicted probability of each modality belonging to each type of human behavior is weighted and fused with its corresponding weight to obtain the predicted probability of the skeleton sequence data to be identified belonging to each type of human behavior.
[0127] The present invention is compared with several graph-based methods including MST-GCN, CTR-GCN, EfficientGCN-B4, InfoGCN, FRHead, HD-GCN, DS-GCN and BlockGCN, several hypergraph-based methods including Hyper-GNN, Selective-HCN and DST-HCN, and Transformer-based methods including DSTA-Net, IIP-Transformer and SkateFormer, as shown in Table 1. Table 1 gives a brief overview of different methods.
[0128] Table 1
[0129]
[0130]
[0131] The human action recognition model (Hyper-GCN) proposed in this paper was trained and tested on an NVIDIA RTX3090 GPU. During training, Hyper-GCN was optimized using stochastic gradient descent (SGD) with a Nesterov momentum of 0.9 and a weight decay of 0.0004. The model was implemented using a label-smoothed cross-entropy loss and the proposed discrepancy loss for a total of 140 epochs, starting with five warm-up epochs. The initial learning rate was 0.05, which was reduced to 0.005 at the 110th epoch and to 0.0005 at the 120th epoch.
[0132] As shown in Table 2, Table 2 compares the test results of the present invention with those of different methods.
[0133] Table 2
[0134]
[0135] Among them, J, B, JM, and BM represent the integration of joint mode, bone mode, joint motion mode, and bone motion mode. Different methods input different modes, integrate the predicted probabilities of different modes belonging to various human behaviors, and obtain the final predicted probabilities for comparison.
[0136] Increasing the number of channels can bring significant improvements, but it also increases the computational complexity of the model to a certain extent. Considering that the number of channels in state-of-the-art lightweight GCN methods is between 1M and 2M, for a fair comparison, we set the number of channels in the last three layers to 256, reducing the number of parameters to 1.06M, which is lighter than all related GCNs.
[0137] Compared with the GCN-based method, as shown in Table 3, Table 3 compares the performance of the lightweight version of the present invention with different methods.
[0138] Table 3
[0139]
[0140] In fact, on the most challenging dataset, NTU RGB+D 120, the performance of the human action recognition method proposed in this paper declined slightly, but still reached the top level with the smallest parameter size. Furthermore, while having parameters comparable to those of Transformer-based methods, the human action recognition method proposed in this paper also outperformed other algorithms.
[0141] Based on Example 1, this Example 2 sets the hypergraph convolution layer of the human behavior recognition model to 9, divided into 3 stages, each stage has dense connections, the channels of each stage are set to 128, 256 and 512 respectively, and the number of sub-feature divisions is set to 8. The skeleton sequence data to be identified is processed, including:
[0142] The bone sequence data to be identified is subjected to data denoising, normalization, and cropping in sequence to obtain the preprocessed bone sequence data to be identified;
[0143] Performing modality conversion on the pre-processed sequence data of bones to be identified to obtain joint modality, bone modality, joint motion modality, and bone motion modality;
[0144] Taking joint modalities as an example, the physical adjacency matrix corresponding to the joint modalities and the bone sequence data to be identified is input into the human behavior recognition model, and the predicted probability of the joint modalities belonging to various human behaviors is output, including:
[0145] Pass the joint modality through the feature embedding layer to extract the joint modality features
[0146] Pass the joint modal features through the first stage and extract the output features of the first stage
[0147] Among them, the joint modal feature F G Through the first hypergraph convolution layer of the first stage, the output features of the first stage are extracted, including:
[0148] The joint modal feature F G Through the first multi-head spatial hypergraph convolution module, the first joint modal space features are extracted, including:
[0149] The joint modal feature F G Divided into 8 sub-features along the channel dimension, the dimension of each sub-feature is Among them, the channel dimension range of the kth sub-feature corresponds to F G [(k-1)×C in / 8+1…,k×C in / 8];
[0150] After each sub-feature is pooled in the time dimension, it passes through the mapping function corresponding to each sub-feature Get the mapping features of each sub-feature
[0151] The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature; according to the maximum number of nodes contained in each hyperedge and the distance between each node and other nodes, the nodes contained in each hyperedge are determined to construct the hyperedge of each sub-feature;
[0152] Based on the distance between each node pair in each sub-feature and the nodes included in each hyperedge, the association matrix of each sub-feature is obtained. The calculation formula is:
[0153]
[0154] Mapping features of each sub-feature Pass the dimensionality reduction mapping function corresponding to each sub-feature in turn LeakeyReLU activation function to obtain the activation features of each sub-feature;
[0155] After connecting the activated feature channels of all sub-features, the mapping function is generated by weights in turn. Tanh activation function, as the weight matrix for each sub-feature;
[0156] Construct a hypergraph for each sub-feature based on its nodes, hyperedges, association matrix, and weight matrix.
[0157] The first joint modal space feature is generated based on the physical adjacency matrix, each sub-feature and its corresponding normalized correlation matrix of the hypergraph.
[0158] The first joint modal spatial feature is convolved through the first time graph to extract the first joint modal temporal feature. The first joint modal temporal feature is the output feature of the first hypergraph convolution layer in the first stage.
[0159] The temporal features of the first joint modality are passed through the second multi-head spatial hypergraph convolution module to extract the spatial features of the second joint modality;
[0160] The second joint modality spatial feature is convolved through the second time graph to extract the second joint modality temporal feature. The second joint modality temporal feature is the output feature of the second hypergraph convolution layer in the first stage.
[0161] The output features of the first hypergraph convolution layer in the first stage are added element by element with the output features of the second hypergraph convolution layer, and input into the third multi-head spatial hypergraph convolution module to extract the third joint modal space features;
[0162] The third joint modal spatial feature is convolved through the third time graph to extract the third joint modal temporal feature as the output feature of the first stage.
[0163] Pass the output features of the first stage through the second stage to extract the output features of the second stage
[0164] Pass the output features of the first stage through the third stage to extract the output features of the third stage
[0165] The output features of the third stage are sequentially passed through the global average pooling layer and the category mapping layer to obtain the predicted probability that the joint modality belongs to various types of human behaviors.
[0166] The nine hypergraph convolutional layers extract features at different levels. Lower-level convolutional layers capture basic features of human behavior, such as the local positions of joints and simple shapes. Deeper convolutional layers combine and abstract these basic features to produce higher-level, more representative features, such as the overall pattern and posture of complex movements. The nine hypergraph convolutional layers are divided into three stages, allowing for gradual information integration and transformation between stages. This allows the model to better adapt to the requirements of feature learning at different levels. Each stage can focus on features of different scales and complexities, improving the model's feature representation capabilities. The number of channels determines the number of features a convolutional layer can extract. As the network depth increases, gradually increasing the number of channels from 128 to 256 and 512 allows the model to have different feature representation capabilities at different stages. When divided into eight sub-features, the model achieves an optimal balance between training and validation sets, fully exploiting the feature information in the data while avoiding overfitting due to excessive sub-features or insufficient detail due to too few sub-features.
[0167] This third embodiment provides a human behavior recognition system, including:
[0168] The model construction module is used to set a multi-head spatial hypergraph convolution module before each temporal graph convolution of the ST-GCN model to build a human behavior recognition model;
[0169] The recognition module is used to convert the skeleton sequence data to be identified into different modal data, input the physical adjacency matrix of the skeleton sequence data to be identified and each modal data into the human behavior recognition model respectively, output the predicted probability of each modality belonging to each type of human behavior, and fuse the predicted probability of the skeleton sequence data to be identified belonging to each type of human behavior;
[0170] Among them, the modal data features input to the multi-head spatial hypergraph convolution module are divided into multiple sub-features along the channel dimension;
[0171] The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature;
[0172] According to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes, the nodes included in each hyperedge are determined, and the hyperedge of each sub-feature is constructed;
[0173] Based on the nodes included in each hyperedge, obtain the association matrix of each sub-feature; connect the mapping feature channels of all sub-features as the weight matrix of each sub-feature;
[0174] Construct a hypergraph for each sub-feature based on the nodes, hyperedges, association matrix, and weight matrix corresponding to each sub-feature;
[0175] The modal data space features are generated based on the physical adjacency matrix, the normalized correlation matrix corresponding to each sub-feature and its hypergraph.
[0176] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0177] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0178] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0180] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A human behavior recognition method, characterized in that: include: A multi-head spatial hypergraph convolution module is set before each temporal graph convolution in the ST-GCN model to build a human behavior recognition model; Convert the to-be-identified skeletal sequence data into different modal data, input the physical adjacency matrix of the to-be-identified skeletal sequence data and each modal data into the human behavior recognition model respectively, output the predicted probability of each modality belonging to each type of human behavior, and fuse them to obtain the predicted probability of the to-be-identified skeletal sequence data belonging to each type of human behavior; Among them, the modal data features input to the multi-head spatial hypergraph convolution module are divided into multiple sub-features along the channel dimension; The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature; According to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes, the nodes included in each hyperedge are determined, and the hyperedge of each sub-feature is constructed; Based on the nodes included in each hyperedge, the association matrix of each sub-feature is obtained; based on the mapping features of all sub-features, the weight matrix of each sub-feature is obtained; Construct a hypergraph for each sub-feature based on its nodes, hyperedges, association matrix, and weight matrix. The modal data space features are generated based on the physical adjacency matrix, the normalized correlation matrix corresponding to each sub-feature and its hypergraph.
2. A human behavior recognition method according to claim 1, characterized in that: The process of obtaining sub-features includes: Constructing and feeding modal data features into the multi-head spatial hypergraph convolution module Virtual features with the same channel and time dimensions Among them, C in F in The channel dimension of F in The time dimension, V is F in The joint dimension, V h F h joint dimensions; The virtual feature F h and the modal data features F input to the multi-head spatial hypergraph convolution module in After splicing along the joint dimension, the real-virtual modal data features are obtained The real-virtual modality data features are divided into multiple sub-features along the channel dimension.
3. A human behavior recognition method according to claim 2, characterized in that: During the training of the human behavior recognition model, the loss function used is: in, is the loss function of the human action recognition model, is the cross entropy loss function between the predicted probability of the skeleton sequence data to be identified belonging to various human behaviors and the true label, is the difference loss of the virtual features corresponding to the l-th multi-head spatial hypergraph convolution module, is the joint dimension of the virtual feature corresponding to the l-th multi-head spatial hypergraph convolution module, C l is the cosine matrix of the virtual features corresponding to the l-th multi-head spatial hypergraph convolution module, is the virtual feature corresponding to the l-th multi-head spatial hypergraph convolution module, (·) T represents transpose, ‖·‖ 2 represents the Euclidean norm, a is the row index of the cosine matrix, b is the column index of the cosine matrix, C l For the element in row a and column b, ReLU(·) is the ReLU activation function.
4. A human behavior recognition method according to claim 1, characterized in that: The human behavior recognition model includes a feature embedding layer, multiple hypergraph convolution layers, a global average pooling layer, and a category mapping layer connected sequentially along the forward propagation direction. Each hypergraph convolution layer includes a multi-head spatial hypergraph convolution module and a temporal graph convolution connected sequentially.
5. A human behavior recognition method according to claim 4, characterized in that: Every three hypergraph convolution layers are regarded as a stage. The output features of the first hypergraph convolution layer and the output features of the second hypergraph convolution layer in each stage are added element by element and serve as the input of the third hypergraph convolution layer.
6. A human behavior recognition method according to claim 1, characterized in that: The mapping features based on all sub-features are used to obtain the weight matrix of each sub-feature, including: The mapping features of each sub-feature are sequentially passed through the dimensionality reduction mapping function and the LeakeyReLU activation function to obtain the activation features of each sub-feature; After connecting the activation feature channels of all sub-features, the weights are used to generate the mapping function and the Tanh activation function as the weight matrix corresponding to each sub-feature. The formula for obtaining the weight matrix corresponding to each sub-feature is: Among them, W is the weight matrix corresponding to each sub-feature, is the mapping feature of the kth sub-feature, is the dimensionality reduction mapping function corresponding to the kth sub-feature, σ 1 (·) is the LeakeyReLU activation function, represents the channel connection operation, K is the number of sub-features, k is the sub-feature index, Ψ 2 (·) is the weight generation mapping function, σ 2 (·) is the Tanh activation function, C in is the channel dimension of the modal data feature input to the multi-head spatial hypergraph convolution module, K is the number of sub-features, and C h is the hidden layer dimension of the mapping subspace, Ψ 2 are learnable parameters.
7. A human behavior recognition method according to claim 1, characterized in that: The step of determining the nodes included in each hyperedge based on the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes includes: A hyperedge is constructed with each node as the centroid. Based on the distance between each node and other nodes, the K nearest neighbor algorithm is used to obtain the K nodes closest to each node. m nodes, as the initial nodes included in the hyperedge corresponding to the node; where K m Hyperparameters for human action recognition models; Count the number of times each node is included in a hyperedge. If the number of times the current node is included in a hyperedge exceeds the maximum number of times the node can be included in a hyperedge, then sort these hyperedges in descending order based on the distance between the current node and the node corresponding to the hyperedge containing the node. Delete the node in the hyperedge that exceeds the maximum number of times the node is included after sorting, so as to determine the nodes included in each hyperedge.
8. A human behavior recognition method according to claim 1, characterized in that: The calculation formula for each element in the correlation matrix corresponding to each sub-feature is: Among them, set i is the node set contained in the hyperedge corresponding to the i-th node, i is the first node index, j is the second node index, i≠j, exp(·) is the exponential function with a natural constant as the base, m i,j is the distance between the i-th node and the j-th node, m i,k For the i-th node and set i The distance between the kth nodes in set i Midpoint index.
9. A human behavior recognition method according to claim 1, characterized in that: The formula for generating the spatial features of modal data is: Among them, F out is the modal data space feature, Indicates a channel connection operation, is the physical adjacency matrix, α is the topological fusion learnable parameter, is the normalized association matrix corresponding to the hypergraph of the k-th sub-feature, is the kth sub-feature, K is the number of sub-features, P k The feature transformation corresponding to the kth sub-feature can learn the weight parameter, and k is the sub-feature index.
10. The human behavior recognition method according to claim 1, characterized in that: The process of obtaining the mapping features of each sub-feature includes: After pooling each sub-feature in the time dimension, the mapping function is used to obtain the mapping features of each sub-feature.
11. A human behavior recognition system, characterized in that: include: The model construction module is used to set a multi-head spatial hypergraph convolution module before each temporal graph convolution of the ST-GCN model to build a human behavior recognition model; The recognition module is used to convert the skeleton sequence data to be identified into different modal data, input the physical adjacency matrix of the skeleton sequence data to be identified and each modal data into the human behavior recognition model respectively, output the predicted probability of each modality belonging to each type of human behavior, and fuse the predicted probability of the skeleton sequence data to be identified belonging to each type of human behavior; Among them, the modal data features input to the multi-head spatial hypergraph convolution module are divided into multiple sub-features along the channel dimension; The feature vector of each joint in the mapping feature of the sub-feature is used as the node of each sub-feature; According to the maximum number of nodes that can be included in the hyperedge and the distance between each node and other nodes, the nodes included in each hyperedge are determined, and the hyperedge of each sub-feature is constructed; Based on the nodes included in each hyperedge, the association matrix of each sub-feature is obtained; based on the mapping features of all sub-features, the weight matrix of each sub-feature is obtained; Construct a hypergraph for each sub-feature based on its nodes, hyperedges, association matrix, and weight matrix. The modal data space features are generated based on the physical adjacency matrix, the normalized correlation matrix corresponding to each sub-feature and its hypergraph.
Citation Information
Patent Citations
Human skeleton action recognition method based on graph and hypergraph convolution composite neural network
CN118629089A
Cited By
Work clothes wearing identification method based on hypergraph calculation
CN121170505A
Fall behavior identification method based on hypergraph depth feature fusion
CN121236821A
Complex scene action classification method based on coupling of spatial semantics and pose features
CN122244561A