Human body action recognition method and device, electronic equipment and storage medium
By segmenting and processing the spatial and temporal features of the three-dimensional skeleton sequence in human body movement recognition, and generating target feature tensors, the problem of insufficient dynamic change capture of node dependencies in the prior art is solved, and accurate recognition of complex actions is achieved.
Patent Information
- Application Number
- CN202510026421.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-30
AI Technical Summary
When building a spatiotemporal graph model of the human body skeleton, existing human body movement recognition methods rely on a single data or a shared adjacency matrix, and cannot accurately capture the dynamic changes in the dependencies between the joint nodes, resulting in low recognition accuracy, especially when identifying complex actions.
By obtaining the human skeleton dataset, the three-dimensional skeleton sequence is determined, and segmented in the spatial and temporal dimensions, the action features are extracted, the spatial and temporal-related adjacency matrix is calculated, matrix coupling and splicing is performed, and the target feature tensors are generated to dynamically capture the dependencies between the joint nodes.
It realizes accurate recognition of simple or complex human body movements, improves the accuracy of human body movement recognition, and can dynamically capture changes in dependencies between nodes.
Smart Images

Figure CN120071431A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly relates to a human action recognition method, an apparatus, an electronic device, and a storage medium. Background Art
[0002] Human action recognition technology is widely used in many scenarios such as smart home, intelligent security, and human-computer interaction. In human action recognition technology, a spatio-temporal graph model of the human skeleton is usually constructed, and an adjacency matrix is used to represent the connection relationship between joint points in the human skeleton.
[0003] When constructing the spatio-temporal graph model of the human skeleton, some human action recognition methods often only rely on single skeleton data or image data, which limits the richness and accuracy of feature information and results in low accuracy of human action recognition. When processing dynamic human action data, some methods use a shared adjacency matrix. However, the dependence relationship and interaction between joint points change over time. When processing different time points in an action sequence, the same adjacency matrix is used, which cannot accurately capture the dynamic changes of the dependence relationship between joint points and leads to low accuracy of human action recognition, especially a significant decrease in accuracy when recognizing complex human actions. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention provides a human action recognition method, an apparatus, an electronic device, and a storage medium, which can accurately recognize simple or complex human actions and improve the accuracy of human action recognition.
[0005] A human action recognition method according to an embodiment of the first aspect of the present invention includes:
[0006] Obtaining a human skeleton data set containing human skeleton data of at least two objects;
[0007] Determining a three-dimensional skeleton sequence according to the human skeleton data set;
[0008] Performing combination processing on the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain a target skeleton sequence;
[0009] Segmenting the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, extracting action features from the skeleton sequence segments of multiple spatial scales to obtain spatial feature vectors, and performing optimization calculation based on the spatial feature vectors to obtain a spatial correlation adjacency matrix;
[0010] Segment the target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales, extract action features from the skeleton sequence segments of multiple time scales to obtain time feature vectors, and perform optimization calculations based on the time feature vectors to obtain a time-related adjacency matrix;
[0011] Multiply the space-related adjacency matrix and the space feature vector to obtain a first coupling matrix;
[0012] Multiply the time-related adjacency matrix and the time feature vector to obtain a second coupling matrix;
[0013] Concatenate the first coupling matrix and the second coupling matrix to obtain a target feature tensor;
[0014] Determine the human action recognition result according to the target feature tensor.
[0015] A human action recognition method according to an embodiment of the present invention has at least the following beneficial effects:
[0016] According to the technical solution of the embodiment of the present invention, first determine a three-dimensional human skeleton sequence from a human skeleton dataset, and then perform spatial and temporal segmentation, action feature processing and optimization calculation on the three-dimensional human skeleton sequence to obtain a space-related adjacency matrix and a time-related adjacency matrix; couple the space-related adjacency matrix and the time-related adjacency matrix with their respective feature vectors to obtain a first coupling matrix and a second coupling matrix; then concatenate the first coupling matrix and the second coupling matrix to obtain a target feature tensor; this target feature tensor is equivalent to directly modeling the temporal correlation and spatial correlation between human skeleton joint nodes, and can dynamically capture the changes in the dependence relationship between joint points for both simple and complex human skeleton joint nodes; finally, determine the human action recognition result according to the target feature tensor. It can be seen that the present invention can accurately recognize simple or complex human actions and improve the accuracy of human action recognition.
[0017] According to some embodiments of the present invention, the combination processing of the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence includes:
[0018] Divide the three-dimensional skeleton sequence into a first skeleton sequence and a second skeleton sequence; wherein, the first skeleton sequence and the second skeleton sequence together constitute the action three-dimensional skeleton sequence of a complete human body;
[0019] Select the first skeleton sequence of the first object as the first input data stream;
[0020] Select the second skeleton sequence of the second object as the second input data stream; wherein, the first object and the second object are different objects;
[0021] Combine the first input data stream and the second input data stream for processing to obtain a target skeleton sequence.
[0022] According to some embodiments of the present invention, the target skeleton sequence is segmented in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, action features are extracted from the skeleton sequence segments of multiple spatial scales to obtain spatial feature vectors, and an optimization calculation is performed based on the spatial feature vectors to obtain a spatial correlation adjacency matrix, including:
[0023] Perform multi-scale segmentation processing on the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales;
[0024] Extract features at a specific spatial scale from each skeleton sequence segment of the spatial scale to obtain corresponding spatial feature vectors;
[0025] Use the spatial feature vectors to calculate the spatial correlation between joints at different spatial scales to obtain a spatial correlation matrix;
[0026] According to a predefined spatial adjacency matrix, a preset learnable spatial parameter, and the spatial correlation matrix, perform a fusion calculation to obtain the spatial correlation adjacency matrix.
[0027] According to some embodiments of the present invention, the target skeleton sequence is segmented in the time dimension to obtain skeleton sequence segments of multiple time scales, action features are extracted from the skeleton sequence segments of multiple time scales to obtain time feature vectors, and an optimization calculation is performed based on the time feature vectors to obtain a time correlation adjacency matrix, including:
[0028] Perform multi-scale segmentation processing on the target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales;
[0029] Perform convolution processing on the skeleton sequence segments of different time scales through multiple parallel different convolution branches respectively to extract features at a specific time scale to obtain time feature vectors;
[0030] Use the time feature vectors to calculate the time correlation between joints at different time scales to obtain a time correlation matrix;
[0031] According to a predefined time adjacency matrix, a preset learnable time parameter, and the time correlation matrix, perform a fusion calculation to obtain the time correlation adjacency matrix.
[0032] According to some embodiments of the present invention, performing convolution processing on the skeleton sequence segments of different time scales through multiple parallel different convolution branches respectively includes:
[0033] For each parallel volume integration branch, the skeleton sequence segment corresponding to the time scale is first processed by a 1×1 convolutional layer for convolution to reduce the number of channels;
[0034] When no further feature extraction is required for the branch, the time-related adjacency matrix is directly constructed based on the skeleton sequence segment after convolution processing;
[0035] When further feature extraction is required for the convolution branch, a max pooling layer is selectively added for pooling operation or convolution operation is performed using convolutional kernels of size k 1 ×1 and k 2 ×1, and then the time-related adjacency matrix is constructed based on the skeleton sequence segment after pooling operation or convolution operation; where k 1 and k 2 are both positive integers.
[0036] According to some embodiments of the present invention, determining the human action recognition result based on the target feature tensor includes:
[0037] Inputting the target feature tensor into a long-tail classification model, and identifying the sample quantity level of the category to which the target feature tensor belongs through the long-tail classification model;
[0038] Determining a learning rate adjustment coefficient according to the sample quantity level, and updating the long-tail classification model according to the learning rate adjustment coefficient;
[0039] Processing the target feature tensor through the updated long-tail classification model to obtain the long-tail classification scores of each category;
[0040] Determining the human action recognition result according to the action labels corresponding to the long-tail classification scores of each category.
[0041] According to some embodiments of the present invention, determining the human action recognition result according to the action labels corresponding to the long-tail classification scores of each category includes:
[0042] Weighting the long-tail classification scores of each category to obtain target classification scores; where each target classification score corresponds to a specific action label;
[0043] Comparing the magnitudes of each target classification score to obtain the highest target classification score;
[0044] Obtaining the action label corresponding to the highest target classification score to obtain the human action recognition result.
[0045] According to an embodiment of the second aspect of the present invention, a human action recognition device, the device includes:
[0046] A detection unit, configured to obtain a human body skeleton data set including human body skeleton data of at least two objects;
[0047] A tracking unit, configured to determine a three-dimensional skeleton sequence according to the human body skeleton data set;
[0048] A spatial interaction unit, configured to perform combination processing on the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain a target skeleton sequence; segment the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, and perform action feature extraction on the skeleton sequence segments of multiple spatial scales to obtain a spatial feature vector and a spatial correlation adjacency matrix;
[0049] A temporal interaction unit, configured to segment the target skeleton sequence in the temporal dimension to obtain skeleton sequence segments of multiple temporal scales, and perform action feature extraction on the skeleton sequence segments of multiple temporal scales to obtain a temporal feature vector and a temporal correlation adjacency matrix;
[0050] An evaluation unit, configured to perform matrix multiplication on the spatial correlation adjacency matrix and the spatial feature vector to obtain a first coupling matrix; perform matrix multiplication on the temporal correlation adjacency matrix and the temporal feature vector to obtain a second coupling matrix; perform matrix splicing on the first coupling matrix and the second coupling matrix to obtain a target feature tensor; determine a human body action recognition result according to the target feature tensor.
[0051] An electronic device according to an embodiment of the third aspect of the present invention, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the human body action recognition method described in the first aspect above is implemented.
[0052] A computer-readable storage medium according to an embodiment of the fourth aspect of the present invention, the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the human body action recognition method described in the first aspect above.
[0053] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0054] Figure 1 is a flowchart of a human body action recognition method provided by an embodiment of the present invention;
[0055] Figure 2 is a schematic diagram of multimodal conversion provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic diagram of temporal and spatial convolution provided by an embodiment of the present invention;
[0057] Figure 4 It is a schematic diagram of multi-branch temporal convolution provided by an embodiment of the present invention;
[0058] Figure 5 It is a schematic diagram of the process of obtaining the target feature tensor provided by an embodiment of the present invention;
[0059] Figure 6 It is a schematic diagram of the multi-modal data fusion processing process provided by an embodiment of the present invention;
[0060] Figure 7 It is Figure 1 the flowchart of step S13 in
[0061] Figure 8 It is an example diagram of combining and processing the three-dimensional skeletons of different objects provided by an embodiment of the present invention;
[0062] Figure 9 It is Figure 1 the flowchart of step S14 in
[0063] Figure 10 It is Figure 1 the flowchart of step S15 in
[0064] Figure 11 It is Figure 10 the flowchart of step S42 in
[0065] Figure 12 It is Figure 1 the flowchart of step S19 in
[0066] Figure 13 It is a schematic diagram of the long-tail classification processing and multi-stream fusion process provided by an embodiment of the present invention;
[0067] Figure 14 It is Figure 12 the flowchart of step S64 in Detailed Embodiments
[0068] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0069] In the description of the present invention, it should be understood that regarding the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0070] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is two or more. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0071] In the description of the present invention, unless otherwise clearly defined, words such as "set", "install", "connect", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0072] The embodiments of the present invention provide a human action recognition method, device, electronic device, and storage medium. Through this human action recognition method, it is possible to directly model the temporal correlation and spatial correlation between the human skeleton joint nodes, and be able to dynamically capture the changes in the dependency relationships between the joint points; it also uses long-tail classification processing to identify low-frequency action categories, and performs weighted fusion on various data streams. Finally, it can accurately recognize simple or complex human actions, improving the accuracy of human action recognition.
[0073] Based on the drawings below, the human action recognition method of the embodiments of the present invention will be further elaborated.
[0074] Refer to Figure 1 , Figure 1 which is a flowchart of the human action recognition method provided by an embodiment of the present invention. The human action recognition method includes but is not limited to the following steps:
[0075] Step S11, obtain a human skeleton data set containing at least two objects;
[0076] Step S12, determine a three-dimensional skeleton sequence according to the human skeleton data set;
[0077] Step S13, perform combination processing on the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain a target skeleton sequence;
[0078] Step S14: Segment the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments at multiple spatial scales. Extract action features from the skeleton sequence segments at multiple spatial scales to obtain spatial feature vectors. Perform optimization calculations based on the spatial feature vectors to obtain a spatially related adjacency matrix;
[0079] Step S15: Segment the target skeleton sequence in the temporal dimension to obtain skeleton sequence segments at multiple temporal scales. Extract action features from the skeleton sequence segments at multiple temporal scales to obtain temporal feature vectors. Perform optimization calculations based on the temporal feature vectors to obtain a temporally related adjacency matrix;
[0080] Step S16: Multiply the spatially related adjacency matrix with the spatial feature vectors to obtain a first coupling matrix;
[0081] Step S17: Multiply the temporally related adjacency matrix with the temporal feature vectors to obtain a second coupling matrix;
[0082] Step S18: Concatenate the first coupling matrix and the second coupling matrix to obtain a target feature tensor;
[0083] Step S19: Determine the human action recognition result according to the target feature tensor.
[0084] In step S11, the object refers to a person performing certain actions. The object may perform some actions, including specific body actions such as gestures, walking, running, jumping, sitting, standing, etc. Human skeleton data is a digital representation of human posture and motion information. Human skeleton data usually involves the coordinate information of each key point (such as joints) of the human body. Human skeleton data can be used to represent the structure and actions of the human body. A human skeleton data set is a data set including human skeleton data of multiple objects. When performing a certain action, the human skeleton of the object will change. Therefore, by recognizing the human skeleton data of the object, the action made by the object can be obtained. In this embodiment, the number of targeted objects is at least two, such as two, three, ten, M, etc. Among them, M is a positive integer greater than 2.
[0085] In step S11, a human skeleton dataset can be obtained by performing human skeleton recognition on the collected initial data. The initial data mainly includes pictures or videos of human actions taken by a drone. During the process of human skeleton recognition: First, mark the pictures or videos of human actions to mark the key point positions of each human action, such as the head, shoulders, elbows, wrists, hips, knees, and ankles, etc.; then preprocess the pictures or videos of human actions, such as cropping, scaling, denoising, etc., and then use the pose estimation algorithm OpenPose to detect the key points of the preprocessed pictures or videos of human actions. The algorithm can automatically identify the key point positions of each human action in the image; according to the detected key point positions, construct the skeleton data of each human action to obtain the human skeleton dataset.
[0086] In step S12, according to the human skeleton dataset, determine the three-dimensional skeleton sequence; in one embodiment, step 12 includes extracting two-dimensional coordinates from the human skeleton dataset to obtain a two-dimensional skeleton sequence; then performing three-dimensional conversion on the two-dimensional skeleton sequence to obtain a three-dimensional skeleton sequence; wherein, the two-dimensional skeleton sequence is a sequence composed of the coordinates of human joint points in a two-dimensional plane and the connections between these coordinates, reflecting the two-dimensional position information of human joint points at a certain time point, including the two-dimensional coordinates of human joint points and the two-dimensional coordinates of human bone edges. The three-dimensional skeleton sequence is a sequence composed of the coordinates of human joint points in three-dimensional space and the connections between these coordinates. The coordinates of each joint point represent its position in three-dimensional space, including the three-dimensional coordinates of human joint points and the corresponding actions of human joint points.
[0087] Extract two-dimensional coordinates from the human skeleton dataset to obtain a two-dimensional skeleton sequence:
[0088] The two-dimensional coordinates of human joint points are represented as: ν i ={x i ,y i}; v j ={x j ,y j};
[0089] The two-dimensional coordinates of human bone edges are represented as:
[0090] where i is the index of the i-th joint point, v i is the i-th joint point, x i is the horizontal coordinate of the i-th joint point in the two-dimensional plane, y i is the vertical coordinate of the i-th joint point in the two-dimensional plane; j is the index of the j-th joint point, v j is the j-th joint point, x j is the horizontal coordinate of the j-th joint point in the two-dimensional plane, y jis the vertical coordinate of the j-th joint point in the two-dimensional plane; is the joint point v i and the joint point v j the bone edge between;
[0091] The set of human joints is represented as:
[0092] The set of human bones is represented as:
[0093] where N is the number of joints, and k i is the i-th relative index of the joint.
[0094] Perform three-dimensional conversion on the two-dimensional skeleton sequence to obtain a three-dimensional skeleton sequence;
[0095] The formula for converting the two-dimensional skeleton sequence into a three-dimensional skeleton sequence is: Pos 3d = Encoding(x i , y i ); Substitute the two-dimensional coordinates of the human joints into the above formula for conversion to obtain the three-dimensional coordinates of the human joints;
[0096] The three-dimensional coordinates of the human joints are represented as:
[0097] The corresponding actions of the human joint points are represented as:
[0098] where Pos 3d represents the position where the two-dimensional coordinates (x i , y i , y i ) of the joint point ν represents the three-dimensional coordinates of the human joints, x i represents the horizontal coordinate of the i-th joint point in the two-dimensional plane, C represents the dimension of the joint features, T represents the number of key frames in the skeleton sequence, V represents the number of joints, and N represents the total number of video sequences; y i represents the action category of a specific video sequence.
[0099] In steps S11 to S12, refer to Figure 2, the positions of joint points in the human skeleton dataset are extracted in a two-dimensional space. The connection line between two joint points forms a bone. Through the relationship between joint points and bones, the joint modality and bone modality are introduced in the two-dimensional space. In the three-dimensional space, the time scale is added. As time goes by, the human body movements change, and the positions of joint points and bones change. That is, the joint movement modality and bone movement modality are introduced in the three-dimensional space, realizing the conversion from the joint modality and bone modality to the joint movement modality and bone movement modality.
[0100] In step S13, in one embodiment, the three-dimensional skeleton sequence is segmented by using the method of average segmentation. First, by analyzing the joint points of the three-dimensional skeleton and their connection relationships, each part in the three-dimensional skeleton sequence is identified. After identifying N parts, they are evenly segmented into two parts, namely the first skeleton sequence and the second skeleton sequence. When N is an even number, the first skeleton sequence and the second skeleton sequence each contain N / 2 parts. When N is an odd number, the first skeleton sequence and the second skeleton sequence each contain as close as possible to N / 2 parts.
[0101] In step S13, in one embodiment, the three-dimensional skeleton sequence is segmented by using the segmentation method based on the characteristics or functions of parts. First, by analyzing the joint points of the three-dimensional skeleton and their connection relationships, each part in the three-dimensional skeleton sequence is identified. The parts related to movement (such as arms and legs) are divided into the first skeleton sequence, and the parts related to support or stability (such as the torso) are divided into the second skeleton sequence.
[0102] In step S13, there is no limit on the number of parts into which the three-dimensional skeleton sequence is divided. Only after the three-dimensional skeleton sequence is divided, different parts of different objects corresponding to the selected number are used as input streams for combined processing. In one embodiment, three different objects can be selected, and the human skeleton sequence of each object is divided into the first skeleton sequence, the second skeleton sequence, and the third skeleton sequence. The first skeleton sequence of the first object is used as the first input stream, the second skeleton sequence of the second object is used as the second input stream, and the third skeleton sequence of the third object is used as the third input stream. The first input stream, the second input stream, and the third input stream are combined and processed to obtain the target skeleton sequence. Among them, the first skeleton sequence, the second skeleton sequence, and the third skeleton sequence together constitute the complete three-dimensional skeleton sequence of the human body movement.
[0103] In step S13, the 3D skeletons representing different objects in the 3D skeleton sequence are combined to obtain the target skeleton sequence. The process of combination processing is that first, the 3D skeleton data of different objects need to be aligned to the same coordinate system, then a suitable skeleton fusion method is selected, and finally the fusion operation is performed; in one embodiment, first, the 3D skeleton data of the first skeleton sequence and the second skeleton sequence are aligned to the same coordinate system, and the aligned skeleton sequences are fused by weighted average, that is, the 3D coordinates of the corresponding joint points are weighted averaged to perform the fusion operation.
[0104] In step S14, the target skeleton sequence is segmented in the spatial dimension, which means that each part or joint in the target skeleton sequence is divided in space. First, the segmentation basis needs to be determined, and then the segmentation operation is performed. In one embodiment, the segmentation can be performed according to the natural structure of the human joints (such as head, torso, limbs, etc.), and then according to the segmentation basis, each joint point in the skeleton data is identified and divided; the action feature extraction for the skeleton sequence segments at multiple spatial scales means that each skeleton sequence segment at each spatial scale is processed separately, and static action features such as the position, angle, and distance of the joint points can be extracted, or dynamic action features such as the movement trajectory, speed, and acceleration of the joint points can be extracted. Finally, the extracted action features are encoded to construct a spatial feature vector.
[0105] In step S14, optimization calculations are performed based on the spatial feature vector to obtain a spatial correlation adjacency matrix; the process of performing optimization calculations means that an enhancement strategy is adopted in the process of constructing the spatial correlation adjacency matrix. By combining the decoupled multi-scale spatial features with a predefined spatial adjacency matrix A, the model's cognitive ability for the local connections between human joints can be strengthened. The specific process of constructing the spatial correlation adjacency matrix is as follows: the target skeleton sequence is segmented in the spatial dimension to obtain skeleton sequence segments at multiple spatial scales; for each skeleton sequence segment at each spatial scale, features at a specific spatial scale are extracted to obtain the corresponding spatial feature vector; using the spatial feature vector, the spatial correlation between joints at different spatial scales is calculated to obtain the spatial correlation matrix Q; according to the predefined spatial adjacency matrix A, the preset learnable spatial parameter α, and the spatial correlation matrix Q, fusion calculations are performed to obtain the spatial correlation adjacency matrix D. The expression of the spatial correlation adjacency matrix D is as follows:
[0106] D = Q·α + A
[0107] where Q is the spatial correlation matrix for capturing the static spatial relationship between joints; α is the preset learnable spatial parameter; and A is the predefined spatial adjacency matrix.
[0108] In step S15, segmenting the target skeleton sequence in the time dimension means cutting the continuous target skeleton sequence into multiple shorter and temporally continuous skeleton sequence segments in chronological order. Each skeleton sequence segment contains the motion information of the target skeleton over a period of time; in one embodiment, the skeleton sequence can be cut according to a fixed time length (such as per second, per half second, etc.).
[0109] In steps S14 to S15, referring to Figures 3 to 4 , in Spatial Decoupled Learning (SDL), for each skeleton sequence segment at each spatial scale, extract features at a specific spatial scale, and this process performs basic convolution operations; in Temporal Decoupled Learning (TDL), for skeleton sequence segments at multiple temporal scales, perform action feature extraction, and this process also performs basic convolution operations, that is, a 1×1 convolutional layer is used to reduce the number of channels; however, different from Spatial Decoupled Learning (SDL), the convolution in the Temporal Decoupled Learning (TDL) process can be adaptively adjusted according to whether further feature extraction is required for different branches; in one embodiment, input the target skeleton sequence, perform multi-scale segmentation processing on the target skeleton sequence in the time dimension to obtain skeleton sequence segments at multiple temporal scales; use a 1×1 convolutional layer for each of the skeleton sequence segments at different temporal scales through multiple parallel different convolutional branches to reduce the number of channels. When a branch does not require further feature extraction, directly proceed to subsequent processing; when a branch requires further feature extraction, it can adaptively choose to slide a window with a width of 3 and a height of 1 on the input feature map and select the maximum value within each window as the output; when a branch requires further feature extraction, use convolution kernels of size k 1 ×1 and k 2 ×1 for convolution operations; finally, splice the different skeleton sequences after the above different processes of multiple parallel different convolutional branches, and then construct a time-related adjacency matrix based on the spliced skeleton sequence.
[0110] In step S15, perform optimization calculations based on the time feature vector to obtain a time-related adjacency matrix; this process is similar to the process of constructing a space-related adjacency matrix, and also sets a predefined time adjacency matrix, preset learnable time parameters, and a time-related matrix, and performs fusion calculations to obtain a time-related adjacency matrix.
[0111] In steps S16 to S18, referring to Figure 5, multiply the space-related adjacency matrix with the space feature vector to obtain the first coupling matrix; multiply the time-related adjacency matrix with the time feature vector to obtain the second coupling matrix; in the spatio-temporal mixing module, the first coupling matrix and the second coupling matrix can be concatenated to obtain the target feature tensor; the process of fusing the first coupling matrix and the second coupling matrix reflects the comprehensive spatial and temporal information of the human skeleton sequence in three-dimensional space.
[0112] Through steps S11 to S19, referring to Figure 6 , first determine the three-dimensional human skeleton sequence from the human skeleton dataset, then perform spatial and temporal segmentation, action feature processing and optimization calculation on the three-dimensional human skeleton sequence to obtain the space-related adjacency matrix and the time-related adjacency matrix; couple the space-related adjacency matrix and the time-related adjacency matrix with their respective feature vectors to obtain the first coupling matrix and the second coupling matrix; then concatenate the first coupling matrix and the second coupling matrix to obtain the target feature tensor; this target feature tensor is equivalent to directly modeling the temporal and spatial correlations between the human skeleton joint nodes, and can dynamically capture the changes in the dependencies between the joint points for both simple and complex human skeleton joint nodes; finally, determine the human action recognition result according to the target feature tensor. This human action recognition method can accurately recognize simple or complex human actions and improve the accuracy of human action recognition.
[0113] According to some embodiments of the present invention, referring to Figure 7 , in Figure 1 In step S13 of the illustrated embodiment, perform combined processing on the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain the target skeleton sequence, including:
[0114] S21, divide the three-dimensional skeleton sequence into a first skeleton sequence and a second skeleton sequence; wherein, the first skeleton sequence and the second skeleton sequence together constitute the three-dimensional skeleton sequence of the actions of a complete human body;
[0115] S22, select the first skeleton sequence of the first object as the first input data stream;
[0116] S23, select the second skeleton sequence of the second object as the second input data stream; wherein, the first object and the second object are different objects;
[0117] S24, perform combined processing on the first input data stream and the second input data stream to obtain the target skeleton sequence.
[0118] In step S21, the first skeleton sequence and the second skeleton sequence may include the skeleton sequences of multiple human body parts; it only needs to meet the condition that in the same human action, the first skeleton sequence and the second skeleton sequence can jointly form a complete human action skeleton sequence; in one embodiment, a human action is divided into five core parts: the torso, the left arm, the left leg, the right arm, and the right leg; the first skeleton sequence includes: the torso, the left arm, and the left leg; the second skeleton sequence includes: the right arm and the right leg; the first skeleton sequence and the second skeleton sequence constitute a complete human action.
[0119] In steps S22 to S24, with reference to Figure 8 , in one embodiment, the torso, the left arm, and the left leg of the first object are used as the first input data stream, the right arm and the right leg of the second object are used as the second input data stream, and the first input data stream and the second input data stream are fused to obtain the target skeleton sequence.
[0120] Through steps S20 to S24, the three-dimensional skeleton sequence is divided into a first skeleton sequence and a second skeleton sequence; wherein, the first skeleton sequence and the second skeleton sequence jointly constitute the action three-dimensional skeleton sequence of a complete human body; the first skeleton sequence of the first object is selected as the first input data stream; the second skeleton sequence of the second object is selected as the second input data stream; wherein, the first object and the second object are different objects; the first input data stream and the second input data stream are combined and processed to obtain the target skeleton sequence. It can integrate the bone data from different perspectives and levels to construct a comprehensive multi-stream input data set.
[0121] According to some embodiments of the present invention, with reference to Figure 9 , in Figure 1 In step S14 of the illustrated embodiment, the target skeleton sequence is segmented in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, the action features of the skeleton sequence segments of multiple spatial scales are extracted to obtain spatial feature vectors, and based on the spatial feature vectors, an optimization calculation is performed to obtain a spatial correlation adjacency matrix, including:
[0122] S31, perform multi-scale segmentation processing on the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales;
[0123] S32, extract the features at a specific spatial scale for each skeleton sequence segment of the spatial scale to obtain the corresponding spatial feature vectors;
[0124] S33, use the spatial feature vectors to calculate the spatial correlation between joints at different spatial scales to obtain a spatial correlation matrix;
[0125] S34. Perform a fusion calculation based on a predefined spatial adjacency matrix, a preset learnable spatial parameter, and a spatial correlation matrix to obtain a spatial correlation adjacency matrix.
[0126] In step S31, perform a multi-scale segmentation process on the target skeleton sequence in the spatial dimension. Specifically, regard the entire skeleton sequence as an overall scale, and then segment the skeleton into smaller scales according to body parts. In one embodiment, perform an equal segmentation operation along the channel dimension on the input data. Execute an equal segmentation operation along the channel dimension.
[0127] In step S32, perform feature extraction on each spatial scale of the skeleton sequence segment. Through a basic 1x1 convolutional kernel, perform a convolution operation on each spatial scale of the skeleton sequence segment to extract the spatial features between joints. These features can be joint positions, joint angles, joint velocities, etc.
[0128] Through steps S31 to S34, perform a multi-scale segmentation process on the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales; extract features at a specific spatial scale for each spatial scale of the skeleton sequence segment to obtain corresponding spatial feature vectors; use the spatial feature vectors to calculate the spatial correlation between joints at different spatial scales to obtain a spatial correlation matrix; perform a fusion calculation based on the predefined spatial adjacency matrix, the preset learnable spatial parameter, and the spatial correlation matrix to obtain a spatial correlation adjacency matrix. This process constructs a spatial correlation adjacency matrix that can reflect the mutual relationship and dependence between joints in space.
[0129] According to some embodiments of the present invention, refer to Figure 10 , in Figure 1 In step S15 of the illustrated embodiment, segment the target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales, perform action feature extraction on the skeleton sequence segments of multiple time scales to obtain time feature vectors, and perform an optimization calculation based on the time feature vectors to obtain a time correlation adjacency matrix, including:
[0130] S41. Perform a multi-scale segmentation process on the target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales;
[0131] S42. Through multiple parallel different convolutional branches, perform convolution processing on the skeleton sequence segments of different time scales respectively to extract features at a specific time scale to obtain time feature vectors;
[0132] S43. Use the time feature vectors to calculate the time correlation between joints at different time scales to obtain a time correlation matrix;
[0133] S44. Perform a fusion calculation based on a predefined temporal adjacency matrix, a preset learnable temporal parameter, and a temporal correlation matrix to obtain a temporal correlation adjacency matrix.
[0134] In step S41, the target skeleton sequence is used as input data and evenly segmented in the temporal dimension to obtain skeleton sequence segments at multiple temporal scales.
[0135] In step S42, the skeleton sequence segments at multiple temporal scales are processed through multiple branches. When each branch processes the skeleton sequence segments at different temporal scales, it first reduces the number of channels through a 1×1 convolutional layer, and then provides different processing methods according to the specific requirements and characteristics of each branch: when the branch does not require any additional operations, it directly enters the subsequent processing stage; when the branch needs further feature extraction or dimensionality reduction, it will selectively add a max pooling layer to reduce the spatial dimension, or perform a convolution operation using convolutional kernels of size k 1 ×1 and k 2 ×1; where the values of k 1 and k 2 will be adaptively adjusted according to the current processed temporal scale.
[0136] Through steps S41 to S44, multi-scale segmentation processing of the target skeleton sequence is performed in the temporal dimension to obtain skeleton sequence segments at multiple temporal scales; through multiple parallel different convolutional branches, the skeleton sequence segments at different temporal scales are respectively convolved to extract features at specific temporal scales to obtain temporal feature vectors; using the temporal feature vectors, calculate the temporal correlation between joints at different temporal scales to obtain a temporal correlation matrix; perform a fusion calculation based on a predefined temporal adjacency matrix, a preset learnable temporal parameter, and the temporal correlation matrix to obtain a temporal correlation adjacency matrix; this process designs a specific temporal convolutional layer to analyze the temporal sequence of actions and capture the dynamic changes of actions; constructing the temporal correlation adjacency matrix can capture the features of the temporal sequence at different temporal scales and reflect the mutual relationship and dependence between joints in time.
[0137] According to some embodiments of the present invention, referring to Figure 11 , in Figure 10 the step S42 of the illustrated embodiment, convolving the skeleton sequence segments at different temporal scales through multiple parallel different convolutional branches respectively includes:
[0138] S51. First, perform a convolution operation on the skeleton sequence segments corresponding to each parallel convolutional branch at the corresponding temporal scale using a 1×1 convolutional layer to reduce the number of channels;
[0139] S52. When no further feature extraction is required for a branch, directly construct a time-related adjacency matrix based on the skeleton sequence segment after convolution processing;
[0140] S53. When further feature extraction is required for the convolution branch, selectively add a max pooling layer for pooling operation or use a convolution kernel of size k 1 ×1 and k 2 ×1 for convolution operation, and then construct a time-related adjacency matrix based on the skeleton sequence segment after pooling operation or convolution operation; where k 1 and k 2 values are adaptively adjusted according to the time scale processed by the current branch.
[0141] Through steps S51 to S53, for each parallel convolution branch to process the skeleton sequence segment corresponding to the time scale, first perform convolution processing using a 1×1 convolution layer, and then adaptively select convolution and size according to whether the branch needs further feature extraction; this process can more accurately adapt to the feature representation requirements in different branches and enhance the ability to understand complex human actions in long time series.
[0142] According to some embodiments of the present invention, referring to Figure 12 , in Figure 1 in step S19 of the illustrated embodiment, determine the human action recognition result according to the target feature tensor, including:
[0143] S61. Input the target feature tensor into the long-tail classification model to identify the sample quantity level of the category to which the target feature tensor belongs through the long-tail classification model;
[0144] S62. Determine the learning rate adjustment coefficient according to the sample quantity level, and update the long-tail classification model according to the learning rate adjustment coefficient;
[0145] S63. Process the target feature tensor through the updated long-tail classification model to obtain the long-tail classification scores for each category;
[0146] S64. Determine the human action recognition result according to the action labels corresponding to the long-tail classification scores of each category.
[0147] In step S61, joint modality and bone modality are introduced in the two-dimensional state; joint motion modality and bone motion modality are introduced in the three-dimensional state. Inputting the data streams of these four modalities into the long-tail classification model will increase two data streams of joint long-tail modality and bone long-tail modality.
[0148] In step S61, referring to Figure 13, in the long-tail classification model, the samples in the training set are carefully divided. According to the number of samples in each category, the samples are divided into three levels of sample numbers: high-frequency, medium-frequency, and low-frequency; a specific learning rate adjustment coefficient is assigned to each level of sample numbers.
[0149] In step S62, after each model training cycle is completed, the learning rate adjustment coefficient of each level of sample numbers can be automatically updated according to the performance of each level of sample numbers; the performance of each level of sample numbers includes indicators such as prediction accuracy and loss value; the learning rate adjustment coefficient can adjust the accuracy of unbalanced action categories, prevent the model from overfitting, and achieve the best effect.
[0150] In step S63, the long-tail classification scores of each category are obtained, that is, the long-tail classification scores of six modalities: output joint modality, bone modality, joint motion modality, bone motion modality, joint long-tail modality, and bone long-tail modality are obtained; the long-tail classification score of each modality includes the scores corresponding to different actions under the data of this modality; the higher the score, the more certain the model's classification result for this category; the lower the score, the greater the uncertainty of the model's classification result for this category.
[0151] Through steps S61 to S64, the target feature tensor is input into the long-tail classification model, and the level of the number of samples to which the target feature tensor belongs is identified through the long-tail classification model; the learning rate adjustment coefficient is determined according to the level of the number of samples, and the long-tail classification model is updated according to the learning rate adjustment coefficient; the target feature tensor is processed by the updated long-tail classification model to obtain the long-tail classification scores of each category; according to the action labels corresponding to the long-tail classification scores of each category, the human action recognition result is determined. This process can identify low-frequency actions, effectively balance the contributions of each human action category, and avoid overfitting to high-frequency human action categories.
[0152] According to some embodiments of the present invention, referring to Figure 14 , in Figure 12 In step S64 of the illustrated embodiment, according to the action labels corresponding to the long-tail classification scores of each category, determining the human action recognition result includes:
[0153] S71, weighting the long-tail classification scores of each category to obtain the target classification scores; wherein, each target classification score corresponds to a specific action label respectively;
[0154] S72, comparing the magnitudes of each target classification score to obtain the highest target classification score;
[0155] S73, obtaining the action label corresponding to the highest target classification score to obtain the human action recognition result.
[0156] In step S71, the long-tail classification scores of each category are weighted to obtain the target classification score. Specifically, the long-tail classification scores of the six modalities, namely the joint modality, the bone modality, the joint motion modality, the bone motion modality, the joint long-tail modality, and the bone long-tail modality, participate in multi-modal fusion according to the weight fusion formula. The weight fusion formula is:
[0157] where λ 1 represents the weight coefficient of the joint modality, λ 2 represents the weight coefficient of the bone modality, λ 3 represents the weight coefficient of the joint motion modality, λ 4 represents the weight coefficient of the bone motion modality, λ 5 represents the weight coefficient of the joint long-tail modality, λ 6 represents the weight coefficient of the bone long-tail modality.
[0158] In step S71, in some embodiments, the accuracies of the six modalities, namely the joint modality, the bone modality, the joint motion modality, the bone motion modality, the joint long-tail modality, and the bone long-tail modality, on the CSv1 benchmark set and the CSv2 benchmark set are calculated respectively.
[0159] Through steps S71 to S73, the long-tail classification scores of each category are weighted to obtain the target classification score. Among them, each target classification score corresponds to a specific action label. Compare the magnitudes of each target classification score to obtain the highest target classification score. Obtain the action label corresponding to the highest target classification score to obtain the human action recognition result. This process can perform weighted fusion on six different data streams, reducing the limitation of the single quantity in the human action feature extraction process of a single data stream.
[0160] The embodiment of the present application also provides a human action recognition device, including:
[0161] A detection unit, configured to obtain a human skeleton data set including human skeleton data of at least two objects;
[0162] A tracking unit, configured to determine a two-dimensional skeleton sequence and a three-dimensional skeleton sequence according to the human skeleton data set;
[0163] A spatial interaction unit, configured to perform combination processing on the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain a target skeleton sequence; segment the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, and perform action feature extraction on the skeleton sequence segments of multiple spatial scales to obtain a spatial feature vector and a spatial correlation adjacency matrix;
[0164] A timing interaction unit, configured to segment a target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales, extract action features from the skeleton sequence segments of multiple time scales, and obtain a time feature vector and a time-related adjacency matrix;
[0165] An evaluation unit, configured to perform matrix multiplication on the space-related adjacency matrix and the space feature vector to obtain a first coupling matrix; perform matrix multiplication on the time-related adjacency matrix and the time feature vector to obtain a second coupling matrix; splice the first coupling matrix and the second coupling matrix to obtain a target feature tensor; determine a human action recognition result according to the target feature tensor.
[0166] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned human action recognition method is implemented.
[0167] An embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for executing the human action recognition method as described above.
[0168] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] Those of ordinary skill in the art will appreciate that all or some of the steps and systems disclosed above in the methods can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0170] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A human motion recognition method, characterized in that: include: Obtain a human skeleton dataset including human skeleton data of at least two objects; Determining a three-dimensional skeleton sequence according to the human skeleton dataset; Combining the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain a target skeleton sequence; Segmenting the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, extracting motion features from the skeleton sequence segments of multiple spatial scales to obtain spatial feature vectors, and performing optimization calculation based on the spatial feature vectors to obtain a spatial correlation adjacency matrix; Segmenting the target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales, extracting action features from the skeleton sequence segments of multiple time scales to obtain a time feature vector, and performing optimization calculation based on the time feature vector to obtain a time-related adjacency matrix; Performing matrix multiplication of the spatial correlation adjacency matrix and the spatial eigenvector to obtain a first coupling matrix; Performing matrix multiplication of the time-dependent adjacency matrix and the time eigenvector to obtain a second coupling matrix; Performing matrix concatenation on the first coupling matrix and the second coupling matrix to obtain a target feature tensor; A human action recognition result is determined according to the target feature tensor.
2. The human motion recognition method according to claim 1, characterized in that: Combining the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence comprises: Dividing the three-dimensional skeleton sequence into a first skeleton sequence and a second skeleton sequence; wherein the first skeleton sequence and the second skeleton sequence together constitute a three-dimensional skeleton sequence of a complete human body action; Selecting a first skeleton sequence of a first object as a first input data stream; Selecting a second skeleton sequence of a second object as a second input data stream; wherein the first object and the second object are different objects; The first input data stream and the second input data stream are combined and processed to obtain a target skeleton sequence.
3. The human motion recognition method according to claim 1, characterized in that: The target skeleton sequence is segmented in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, action features are extracted from the skeleton sequence segments of multiple spatial scales to obtain spatial feature vectors, and optimization calculation is performed based on the spatial feature vectors to obtain a spatial correlation adjacency matrix, including: Performing multi-scale segmentation processing on the target skeleton sequence in spatial dimension to obtain skeleton sequence fragments of multiple spatial scales; Extracting features at a specific spatial scale from the skeleton sequence fragments at each of the spatial scales to obtain a corresponding spatial feature vector; Using the spatial eigenvectors, the spatial correlations between joints at different spatial scales are calculated to obtain a spatial correlation matrix; A fusion calculation is performed according to a predefined spatial adjacency matrix, a preset learnable spatial parameter and the spatial correlation matrix to obtain the spatial correlation adjacency matrix.
4. The human motion recognition method according to any one of claims 1 to 3, characterized in that: The target skeleton sequence is segmented in the time dimension to obtain skeleton sequence segments of multiple time scales, action features are extracted from the skeleton sequence segments of multiple time scales to obtain a time feature vector, and optimization calculation is performed based on the time feature vector to obtain a time-related adjacency matrix, including: Performing multi-scale segmentation processing on the target skeleton sequence in the time dimension to obtain skeleton sequence fragments at multiple time scales; Through multiple parallel convolution branches, skeleton sequence fragments of different time scales are convolved to extract features at a specific time scale and obtain a time feature vector. Using the time feature vector, calculate the time correlation between joints at different time scales to obtain a time correlation matrix; A fusion calculation is performed according to a predefined time adjacency matrix, a preset learnable time parameter and the time correlation matrix to obtain the time correlation adjacency matrix.
5. The human motion recognition method according to claim 4, characterized in that: The skeleton sequence fragments of different time scales are convolved through multiple parallel different convolution branches, including: Each parallel convolution branch processes the skeleton sequence fragment of the corresponding time scale by first using a 1×1 convolution layer to perform convolution processing to reduce the number of channels; When the branch does not need to further perform feature extraction, the time-dependent adjacency matrix is directly constructed according to the skeleton sequence fragments after convolution processing; When the convolution branch needs to further perform feature extraction, a maximum pooling layer is selectively added to perform a pooling operation or a convolution kernel of size k1×1 and k2×1 is used to perform a convolution operation, and then the time-dependent adjacency matrix is constructed based on the skeleton sequence fragments after the pooling operation or the convolution operation; wherein k1 and k2 are both positive integers.
6. The human motion recognition method according to any one of claims 1 to 3, characterized in that: Determining a human action recognition result according to the target feature tensor includes: Inputting the target feature tensor into a long-tail classification model, and identifying the sample quantity level of the category to which the target feature tensor belongs through the long-tail classification model; Determining a learning rate adjustment coefficient according to the sample quantity level, and updating the long-tail classification model according to the learning rate adjustment coefficient; Processing the target feature tensor through the updated long-tail classification model to obtain a long-tail classification score for each category; The human action recognition result is determined according to the action labels corresponding to the long-tail classification scores of each category.
7. The human motion recognition method according to claim 6, characterized in that: According to the action labels corresponding to the long-tail classification scores of each category, the human action recognition results are determined, including: Weighting the long-tail classification scores of each category to obtain a target classification score; wherein each target classification score corresponds to a specific action label; Comparing the sizes of each target classification score to obtain the highest target classification score; The action label corresponding to the highest target classification score is obtained to obtain the human action recognition result.
8. A human motion recognition device, characterized in that: The device comprises: A detection unit, configured to obtain a human skeleton data set including human skeleton data of at least two objects; A tracking unit, used for determining a three-dimensional skeleton sequence according to the human skeleton dataset; A spatial interaction unit is used to combine the three-dimensional skeletons representing different objects in the three-dimensional skeleton sequence to obtain a target skeleton sequence; segment the target skeleton sequence in the spatial dimension to obtain skeleton sequence segments of multiple spatial scales, and extract action features from the skeleton sequence segments of multiple spatial scales to obtain spatial feature vectors and spatial correlation adjacency matrices; A time sequence interaction unit, used to segment the target skeleton sequence in the time dimension to obtain skeleton sequence segments of multiple time scales, extract action features from the skeleton sequence segments of multiple time scales to obtain a time feature vector and a time-related adjacency matrix; An evaluation unit is used to perform matrix multiplication on the spatial correlation adjacency matrix and the spatial feature vector to obtain a first coupling matrix; perform matrix multiplication on the time correlation adjacency matrix and the time feature vector to obtain a second coupling matrix; perform matrix concatenation on the first coupling matrix and the second coupling matrix to obtain a target feature tensor; and determine a human motion recognition result based on the target feature tensor.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the human motion recognition method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the human motion recognition method as described in any one of claims 1 to 7.