A Digital Human Action Recognition Method, Device, Medium and Product
By training the digital human action recognition model, using the posture data of the action frame for encoding and multi-layer decoding calculation, the problem of low efficiency in digital human action sequence recognition is solved, and efficient recognition and efficiency improvement is achieved.
Patent Information
- Application Number
- CN202510232941.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The lack of action recognition schemes for digital human action sequences in the prior art, resulting in low efficiency in digital human generation tasks.
By training the digital human action recognition model, encode the posture data of the action frame for encoding, obtain the action sequence list, and obtain the predicted action through multi-layer decoding calculation, and combine the annotation action for matching calculation and loss optimization, optimize the model to improve the recognition efficiency.
It realizes efficient recognition of digital human action sequences, improves the efficiency of digital human generation tasks, and ensures that the model can contribute to action recognition at every level.
Smart Images

Figure CN119763196B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital humans, and particularly to a method, device, medium and product for digital human action recognition. Background Art
[0002] Virtual digital humans (referred to as "digital humans" for short) are one of the applications in the fields of 3D vision and generative artificial intelligence, and are widely used in multiple fields such as augmented reality, virtual reality, digital twins, film and television, and game production. To achieve the accuracy of digital human generation in specific scenarios, a large number of labeled data sets are required, and the numerous labeling tasks result in unsatisfactory efficiency of digital human generation tasks. Therefore, it is necessary to propose a solution for recognizing the actions of digital human action sequences.
[0003] How to improve the action recognition efficiency of digital human action sequences is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] This application provides a method, device, medium and product for digital human action recognition to at least solve the problem of the lack of an action recognition solution for digital human action sequences in related technologies.
[0005] This application provides a method for digital human action recognition, including:
[0006] Obtain digital human action samples;
[0007] Use a digital human action recognition model to encode according to the pose data of the action frames of the digital human action samples to obtain an action sequence representation, and perform multi-layer decoding calculations according to the action sequence representation and an action query vector;
[0008] Obtain the first predicted action output by the last decoding layer, and perform matching calculation on the first predicted action and the labeled action to obtain a first matching relationship;
[0009] According to the first matching relationship, calculate the loss values of the output results of multiple decoding layers compared with the labeled action, and calculate the model loss value according to multiple loss values to optimize the loss of the digital human action recognition model, and obtain the trained digital human action recognition model;
[0010] Use the trained digital human action recognition model with some decoding layers retained to perform the action recognition task of the digital human action sequence to be recognized.
[0011] This application also provides a method for training a digital human action recognition model, including:
[0012] Obtain digital human action samples;
[0013] The digital human motion recognition model encodes based on the pose data of the motion frames of the digital human motion samples to obtain a motion sequence representation, and performs multi-layer decoding calculations based on the motion sequence representation and a motion query vector;
[0014] Obtain the first predicted motion output by the last decoding layer, perform matching calculation on the first predicted motion and the labeled motion to obtain a first matching relationship;
[0015] According to the first matching relationship, calculate the loss values of the output results of multiple decoding layers compared with the labeled motion, and calculate the model loss value based on multiple loss values to optimize the loss of the digital human motion recognition model, and obtain the trained digital human motion recognition model.
[0016] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the above digital human motion recognition method or the steps of the above digital human motion recognition model training method when executing the computer program.
[0017] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above digital human motion recognition method or the steps of the above digital human motion recognition model training method.
[0018] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above digital human motion recognition method or the steps of the above digital human motion recognition model training method.
[0019] Through this application, action recognition of a digital human action sequence is achieved by training a digital human action model. During training, the digital human action recognition model encodes according to the pose data of the action frames of the digital human action samples to obtain an action sequence representation. Based on the action sequence representation and the action query vector, multi-layer decoding calculations are performed to obtain the first predicted action output by the last decoding layer. The first predicted action and the labeled action are matched and calculated to obtain the first matching relationship. According to the first matching relationship, the loss values of the output results of multiple decoding layers compared with the labeled action are calculated, and the model loss value is calculated based on the multiple loss values to optimize the loss of the digital human action recognition model. This not only makes up for the lack of a digital human action recognition solution in the related technology and improves the efficiency of digital human sequence action recognition, but also enables the outputs of multiple decoding layers of the model to directly participate in the loss calculation of digital human action recognition, ensuring that the model can contribute to digital human action recognition at each level. Therefore, the trained digital human action model with some output layers retained can be used to perform action recognition tasks, improving the recognition efficiency on the basis of ensuring recognition accuracy and being able to adapt to different application scenarios, dataset characteristics, and real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 Schematic diagram of a digital human action sequence with labels provided by an embodiment of the present invention;
[0022] Figure 2 Flowchart of a digital human action recognition method provided by an embodiment of the present invention;
[0023] Figure 3 Schematic diagram of the structure of a digital human action recognition system provided by an embodiment of the present invention;
[0024] Figure 4 Training flowchart of a digital human action recognition model provided by an embodiment of the present invention;
[0025] Figure 5 Inference flowchart of a digital human action recognition model provided by an embodiment of the present invention;
[0026] Figure 6 Schematic diagram of the structure of a digital human action recognition model provided by an embodiment of the present invention;
[0027] Figure 7Schematic diagram of another digital human action recognition model provided by an embodiment of the present invention;
[0028] Figure 8 Schematic diagram of a digital human action recognition device provided by an embodiment of the present invention. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0030] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0031] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0032] Here, some key terms used in the embodiments of the present invention will be explained first.
[0033] A digital human refers to a virtual human image created through digital technology and similar to the human image. It has features such as human appearance, behavior patterns, voice, emotional reactions, etc., and can operate and exist independently in the digital space. Digital humans are usually generated by technologies such as computer graphics, artificial intelligence, and natural language processing and have interaction capabilities.
[0034] A digital human action sequence refers to a series of data or instructions generated through computer technology for driving a digital human to perform coherent actions. These action sequences can include facial expressions, limb movements, gestures, etc., and can be created through preset animation data, motion capture technology, or generation models based on deep learning.
[0035] A digital human action sequence consists of multiple action frames. Figure 1 Schematic diagram of a digital human action sequence with labels provided by an embodiment of the present invention. As Figure 1As shown in the figure, a three-dimensional digital human model in the related art is composed of a hinged skeleton tree composed of multiple joints. The hierarchical structure of this skeleton tree defines the parent-child relationship of the nodes. Different postures of human body movements can be obtained by adjusting the rotation angle of the child node relative to the parent node. Each posture and action of the digital human based on the three-dimensional digital human model is determined by the relative rotation angle of these joints. Therefore, the action sequence of the digital human can be regarded as an action frame sequence composed of the rotation angles of multiple joints.
[0036] In order to realize the training of digital human generation models and various downstream applications, a large number of data sets have emerged in the field of digital humans. In order to be put into actual training, these data sets need to be labeled. The semantic labels corresponding to the action sequences of digital humans can include two levels: sequence level and action frame level. Figure 1 As shown, the digital human action sequence has 15 action frames, and the corresponding sequence-level label is "playing basketball". The action frame-level labels corresponding to some of the action frames include "catching the ball with both hands", "transition action", "transferring the basketball to the left hand", "sprinting", and "dribbling with the left hand".
[0037] The numerous digital human data sets bring a lot of labeling work. In order to realize the automatic labeling of digital human data sets, it is necessary to solve the problem of action recognition of digital human action sequences.
[0038] Currently, artificial intelligence solutions driven by deep learning technology have demonstrated excellent performance in many fields, especially in video character motion analysis. However, since the representation of digital human motion sequences is a new modality, the deep learning model used in video motion recognition tasks cannot be applied to the motion recognition of digital human motion sequences.
[0039] To solve the problem of action recognition for the action sequences of digital humans, an embodiment of the present invention provides a digital human action recognition solution. By training a digital human action recognition model to perform digital human action recognition tasks, during the training, the digital human action recognition model encodes based on the pose data of the action frames of the digital human action samples to obtain an action sequence representation. Multilayer decoding calculations are performed based on the action sequence representation and an action query vector to obtain a first predicted action output by the last decoding layer. A matching calculation is performed between the first predicted action and the labeled action to obtain a first matching relationship. The loss value of the output results of multiple decoding layers compared with the labeled action is calculated based on the first matching relationship, and the model loss value is calculated based on multiple loss values to optimize the loss of the digital human action recognition model. This not only makes up for the lack of a digital human action recognition solution in the related art and improves the efficiency of digital human sequence action recognition, but also enables the outputs of multiple decoding layers of the model to directly participate in the loss calculation of digital human action recognition, ensuring that the model can contribute to digital human action recognition at each level. Thus, the trained digital human action model with some output layers retained can be used to perform action recognition tasks, improving the recognition efficiency on the basis of ensuring recognition accuracy and being able to adapt to different application scenarios, dataset characteristics, and real-time requirements.
[0040] An embodiment of the present application provides a digital human action recognition method. Combining the execution process of the digital human action recognition method, the method will be described in detail below.
[0041] Figure 2 It is a flowchart of a digital human action recognition method provided by an embodiment of the present invention; Figure 3 It is a structural schematic diagram of a digital human action recognition system provided by an embodiment of the present invention; Figure 4 It is a training flowchart of a digital human action recognition model provided by an embodiment of the present invention; Figure 5 It is an inference flowchart of a digital human action recognition model provided by an embodiment of the present invention.
[0042] As Figure 2 shown, the digital human action recognition method provided by an embodiment of the present invention includes:
[0043] S201: Obtain digital human action samples;
[0044] S202: Use the digital human action recognition model to encode based on the pose data of the action frames of the digital human action samples to obtain an action sequence representation, and perform multilayer decoding calculations based on the action sequence representation and an action query vector;
[0045] S203: Obtain a first predicted action output by the last decoding layer, perform a matching calculation between the first predicted action and the labeled action to obtain a first matching relationship;
[0046] S204: According to the first matching relationship, calculate the loss value of the output results of multiple decoding layers compared with the labeled actions, and calculate the model loss value based on the multiple loss values to optimize the loss of the digital human action recognition model, so as to obtain the trained digital human action recognition model;
[0047] S205: Use the trained digital human action recognition model with some decoding layers reserved to perform the action recognition task of the digital human action sequence to be recognized.
[0048] In the embodiment of the present invention, the pose data of the action frame refers to the action parameters used to describe the actions of the digital human on the action frame, for example, it can be the joint rotation angle of the digital human.
[0049] To achieve digital human action recognition, the digital human action recognition method provided in the embodiment of the present invention mainly includes two links: training the digital human action recognition model and using the digital human action recognition model to perform the digital human action recognition task. In this regard, the digital human action recognition method provided in the embodiment of the present invention can be applied to, for example Figure 3 the digital human action recognition system shown in the figure. The digital human action recognition system includes a model training system and a digital human action recognition device.
[0050] As Figure 3 shown, the model training system can be supported by multiple artificial intelligence servers (artificial intelligence servers 1 to n) in terms of computing power and storage, and the digital human action recognition model is trained based on the model training system using digital human action samples. In some optional implementation manners of the embodiment of the present invention, a computing resource pool can be constructed based on the model training system. The computing resource pool includes multiple computing threads, so that the training task of the digital human action recognition model can be split into multiple subtasks for parallel execution. The model parallel training method or the data parallel training method can be adopted. The model training system can be deployed in a cloud computing center. The device structure of the cloud computing center can be similar to Figure 3 the digital human action recognition device shown in the figure, including basic settings such as artificial intelligence processors, storage, input / output devices, communication buses, and communication interfaces, and is composed of software environments such as operating systems, databases, and middleware, and application software. The model training system receives the digitally human action samples manually labeled, and can use the pre-trained model stored for fine-tuning to obtain the digital human action recognition model, or can also be trained using a randomly initialized model framework to obtain the digital human action recognition model.
[0051] The digital human action recognition device can be a computing device or a terminal device deployed in a cloud computing center, and its structure is as Figure 3As shown in the digital human action recognition device. If the computing device of the cloud computing center is adopted, after deploying the trained digital human action recognition model, the terminal device uploads the digital human action sequence to be recognized to the cloud computing center, and uses the trained digital human action recognition model to perform action recognition on the digital human action sequence to be recognized, and outputs the predicted action. If the digital human action recognition device adopts the terminal device, after the model training system trains the digital human action recognition model, it is sent to the terminal device through the communication unit and deployed in the storage of the terminal device. The terminal device receives the digital human action sequence to be recognized, performs prediction calculation through the local digital human action recognition model, and outputs the predicted action.
[0052] As Figure 4 shown, when training the digital human action recognition model, the pose data of the action frames of the digital human action samples is input into the digital human action recognition model to output the first predicted action. After matching the first predicted action with the labeled action, it is substituted into the loss function to calculate the model loss value, and the digital human action recognition model is optimized for loss using the model loss value.
[0053] In the specific implementation of the digital human action recognition method provided in the embodiments of the present invention, for S201, the digital human action samples can be sourced from the local storage of the artificial intelligence device where it is located, the shared memory of the model training system where it is located, or digital human action samples input by an external device.
[0054] For S202, the pose data of the action frames of the digital human action samples (for example, one frame of action includes the joint rotation angles of 24 joints) is input into the digital human action recognition model to output the first action category of the first predicted action and the first action interval of the first predicted action.
[0055] Denote the pose data of the action frames of the digital human action samples as (including action frames, represents the pose data of the th frame), denote the set of labeled actions of the digital human action samples as (including labeled actions), where represents the labeled action category of the th labeled action, which is a one-hot encoded vector of length ( is the number of all action categories in the action dataset), represents the coordinates of the labeled action interval of this labeled action, that is, is the serial number of the starting frame of the action, is the serial number of the ending frame of the action, that is, the action annotation category of the sub-action sequence is .
[0056] Input the pose data of the action frames of the digital human action samples into the digital human action recognition model, and the digital human action recognition model predicts a first predicted action ( is a hyperparameter, representing the maximum number of actions that the digital human action recognition model can recognize from a digital human action sequence. Usually, it is required that ), and the first predicted actions predicted by the digital human action recognition model form a first predicted action set , the th first predicted action includes the first action category of this first predicted action (a category confidence vector with a length of , where each element represents the confidence probability of a category, all of which are positive numbers and the sum is 1) and the first action interval (where ).
[0057] That is to say, for a digital human action sample containing frames, it has a total of annotated actions. Configure the digital human action recognition model to recognize first predicted actions according to the pose data of the action frames of the digital human action sample.
[0058] To meet the requirements of different scenarios, especially the model lightweighting requirements for downstream application scenarios that need to be deployed on terminal devices, the digital human action recognition model provided in the embodiments of the present invention includes multiple action decoding layers (i.e., decoding layers) during the training phase, and the outputs of multiple action decoding layers are directly connected to the output heads in the action output layer, and the output results of these action output layers are involved in the calculation of the model loss value, so that multiple action decoding layers are involved in the model loss optimization, ensuring that the model can contribute to digital human action recognition at each level. Based on this, during the model inference phase, some action decoding layers can be retained according to the lightweighting requirements to reduce the model parameters to be deployed and simplify the model calculation amount on the premise of ensuring high model accuracy.
[0059] For S203, first obtain the first action category and the first action interval of the first predicted action of the last decoding layer, obtain the annotated action category and the annotated action interval of the annotated action, and then perform matching calculations according to the parameters of the first predicted action and the parameters of the annotated action to determine the matching relationship between the first predicted action and the annotated action, which is denoted as the first matching relationship in the embodiments of the present invention. For example, screen out from the elements of the first predicted action set the ones that match the annotated actions in the annotated action set one by one A first prediction action. Without loss of generality, assume that the previous actions are the selected actions. Thus, the first prediction action and the annotation action are calculated for loss one by one for gradient calculation and model training.
[0060] During model training, the best model prediction confidence for each category can be set through the test set or the validation set , where . There are many setting methods. For example, if there are 50 validation data for the th category in the test set or the validation set, then it can be set to the lowest probability of the model's prediction for these 50 data, or the probability value of the action ranked 40th, etc.
[0061] For S204, according to the first matching relationship, the serial number of the first prediction action matched with the annotation action in the first prediction actions output by the current decoding layer can be determined, and this is used as the matching relationship of the output results of each decoding layer participating in the loss calculation. That is to say, for the first action category and the first action interval of the first prediction actions output by each decoding layer participating in the loss calculation, the corresponding annotation actions are all matched according to the first matching relationship, and the loss value corresponding to this decoding layer is calculated according to each first prediction action and the corresponding annotation action. Then, according to the loss values corresponding to each decoding layer, the model loss value of the current iterative training is calculated, and the digital human action recognition model is optimized for loss using the model loss value.
[0062] The training end condition can be that the number of iterative training times of the digital human action recognition model reaches the preset number of iterative times, or the model loss value of the digital human action recognition model is less than the preset loss value.
[0063] For S205, the trained digital human action recognition model is obtained and some action decoding layers are retained as needed. Additionally, only the action output layer connected to the last action decoding layer can be retained to achieve lightweight processing of the digital human action recognition model. After the deployment of the digital human action recognition model is completed, the pose data of the digital human action sequence to be recognized can be input into the digital human action recognition model, and the second action category and the second action interval of the second prediction action are output. The steps of using the digital human action recognition model to output the second prediction action are the same as the steps of calculating the first prediction action.
[0064] As Figure 5 shown, the pose data of the action frames of the digital human action sequence to be recognized is input into the digital human action recognition model, and the model predicts A second prediction action, including an action second action category and a second action interval (i.e., the confidence of the action category of the th action is , and the coordinate is . The largest element in is , that is, the most likely category of the model prediction action sequence is . If , then retain the prediction result; if , then discard the prediction result. In this way, the second prediction actions predicted by the model are processed one by one, and the
[0065] remaining
[0066] results are used as the output second prediction actions.
[0067] Figure 6 FIG.
[0068] In the embodiment of the present invention, the input of the digital human action recognition model is an action sequence and an action query vector . The former represents a complete action sequence, represents the action of the th frame of the sequence; the latter represents A query vector to be trained. Assume that at most actions need to be detected in any action sequence. The output of the digital human action recognition model is first predicted actions , indicating that the confidence probability vector of the action subsequence is .
[0069] As Figure 6 shown, the digital human action recognition model provided by the embodiment of the present invention may include an action mapping layer, a multi-scale feature layer, an action encoding layer (abbreviation: encoding layer), a query embedding layer, an action decoding layer (abbreviation: decoding layer), and an action output layer (abbreviation: output layer). Among them, the action mapping layer is used to map action frames into action representation vectors, laying a foundation for subsequent action feature extraction and action recognition. The multi-scale feature layer and the action encoding layer are used to extract multi-scale action features from the action representation vectors, perform encoding fusion, extract deep features of the action sequence, etc., providing rich semantic information for action recognition. The query embedding layer contains a series of action query vectors, which are used for the action decoding layer to query the action features output by the action encoding layer. The action decoding layer is used to decode the action representation by combining the action query vectors output by the query embedding layer to obtain the action decoding representation corresponding to each action query vector. The action output layer is used to perform action category prediction and coordinate prediction based on the output of the action decoding layer to achieve accurate recognition and positioning of digital human actions.
[0070] In the embodiment of the present invention, in S202, encoding is performed on the pose data of the action frames of the digital human action sample by using the digital human action recognition model to obtain an action sequence representation, which may include: inputting the pose data into the digital human action recognition model to convert the pose data into a dense vector representation, and performing self-attention calculation based on the position vector corresponding to the action frame and the dense vector representation to obtain the action sequence representation.
[0071] By converting the pose data into a dense vector representation, it helps the model to more accurately capture the potential connection between the action query vector and the action feature and the deep semantic information of the action feature contained in the pose data, and is conducive to easier similarity calculation and clustering analysis in the high-dimensional space. There is an order relationship between the action frames of the digital human action sequence. To enable the model to understand this order relationship, self-attention calculation is performed based on the position vector corresponding to the action frame and the dense vector representation to obtain the action sequence representation, so that the model can determine the position of each action frame in the digital human action sequence, thereby better expressing the front-back relationship and distance between actions.
[0072] Then the action mapping layer is applied. The input of the action mapping layer is the action sequence , and first each action therein is converted into a preset dimension The dense vector representation of, that is, the pose data of the th action frame, the corresponding dense vector representation is , then there is the following formula:
[0073] ;
[0074] Among them, , if has a dimension of , then has a dimension of , has a dimension of , and are both learnable parameters in the training process of the digital human action recognition model; has a dimension of .
[0075] The output of the action mapping layer is sequence, denoted as , that is .
[0076] In the embodiments of the present invention, the multi-scale feature layer and the action encoding layer are used to extract multi-scale action features from the action representation vector, perform encoding fusion, and extract the deep features of the action sequence. Then, the encoding according to the pose data of the action frames of the digital human action samples in S202 to obtain the action sequence representation may include: performing multi-round feature extraction of different scales on the original action features of the action frames corresponding to the pose data of the action frames, fusing the extracted local sequence features to obtain the first action feature fusion result; performing multi-layer self-attention calculation according to the first action feature fusion result to obtain the action sequence representation. As Figure 6 shown, in the digital human action recognition model provided by the embodiments of the present invention, multi-round feature extraction of different scales is first performed through the multi-scale feature layer, and then the multi-scale feature fusion layer ( Figure 6 not shown) is used to fuse the obtained action features of different scales, and then the action encoding layer including the self-attention module performs self-attention calculation according to the obtained first action feature fusion result, so that the digital human action recognition model can learn multi-scale features to adapt to long and short actions while better understanding the deep semantics of the action sequence.
[0077] In some alternative embodiments of the embodiments of the present invention, multi-round feature extraction of different scales on the original action features of the pose data corresponding to the action frames may include: in a single-round feature extraction, a one-dimensional convolutional kernel is used to perform convolutional calculation on the input action features, and local sequence features corresponding to adjacent multiple action features are output. Regarding the coordinates of multiple action frames of the digital human action samples as one-dimensional data, a one-dimensional convolution is used to extract the features of multiple consecutive action frames in sequence each time to obtain an action representation. By using multiple layers of such one-dimensional convolutions, the receptive field corresponding to the action representation extracted by the later layer is larger, so that action features corresponding to different receptive fields can be extracted for the digital human action recognition model to recognize long actions (corresponding to more action frames) and short actions (corresponding to fewer action frames).
[0078] In the embodiments of the present invention, the input of the multi-scale feature layer is the output of the action mapping layer . The multi-scale feature layer has layers of feature extraction layers, and each layer uses a feature extraction method based on one-dimensional convolution and pooling, aiming to capture local features between adjacent vector representations through convolution operations to generate feature vectors.
[0079] Specifically, specifically, the input of the th layer of the multi-scale feature layer is , and the output is , that is:
[0080] ;
[0081] Among them, and are the lengths of the input sequence and the output sequence , and are the convolutional kernel weights and bias terms, is the window size of the one-dimensional convolutional kernel, is the window size of the pooling layer, represents the max pooling calculation. In particular, is determined jointly by the convolutional kernel window size , window stride, padding size, dilation size, and the pooling layer window size . Then the length of the feature scale satisfies .
[0082] Thus, the low-scale feature is smaller than the high-scale feature The receptive field of low-scale features is small, and the range of feature scales is also small, which is beneficial for short action detection and recognition. In contrast, the receptive field of high-scale features is large, and the range of feature scales is also large, which is beneficial for long action detection and recognition.
[0083] The output of the multi-scale feature layer is , that is .
[0084] Through one-dimensional convolution, multiple consecutive action frames in the action sequence of the digital human action sample are used to extract information to obtain an action representation. For example, a one-dimensional convolutional kernel with a size corresponding to 3 action frames is used. Through the first layer, 3 consecutive action frames are extracted into an action representation, so the receptive field of the one-dimensional convolution in the first layer is 3 action frames; then, based on the first layer, the receptive field of the second layer of one-dimensional convolution is 3 action representations output by the first layer, corresponding to 5 action frames of the original action; the receptive field of the third layer corresponds to 7 action frames of the original action, and so on. Those that can sense 3 action frames can be used for short sequence recognition, and those that can sense 7 or more action frames can be used for long sequence recognition. Figure 6 In the multi-scale feature layer shown, the outputs of each feature extraction layer are concatenated together.
[0085] In some other optional embodiments of the present invention, for the original action features of the action frames corresponding to the pose data of the action frames, multiple rounds of feature extraction with different scales may further include: in a single round of feature extraction, among the input action features, adjacent multiple action features are sequentially concatenated in feature vectors and then subjected to multi-layer perceptron calculation to obtain local sequence features corresponding to the adjacent multiple action features. That is to say, a feature extraction layer can also be constructed using a multi-layer perceptron to improve the processing rate of the digital human action sequence.
[0086] Using a multi-layer perceptron to build a multi-scale feature layer, the input of the multi-scale feature layer is the output of the action mapping layer . The multi-scale feature layer can have layers of feature extraction layers, and each layer uses a feature extraction method based on a multi-layer perceptron and pooling, aiming to capture local features between adjacent action vector representations through multi-layer perceptron operations and generate feature vectors. Flatten the feature vectors within the local window, that is, concatenate all the feature vectors within the window to form a vector with a dimension of The long vectors are input into a multi-layer perceptron, which will non-linearly transform and fuse these features through its hidden layers to capture the complex relationships between features, thereby receiving the global information of all features within the window and performing more effective feature fusion. A pooling layer (max pooling or average pooling) is used to further reduce the number of feature vectors. This step will help reduce the dimensionality of the features, reduce the risk of overfitting, and improve the generalization ability of the model.
[0087] Specifically, the input of the -th layer of the multi-scale feature layer is and the output is That is:
[0088] ;
[0089] where and are the lengths of the input sequence and the output sequence , and are the convolutional kernel weights and bias terms, is the window size, is the scale after flattening the vectors within the window, that is , is the window size of the pooling layer, represents the calculation of the multi-layer perceptron, represents the pooling calculation. Specifically, is determined jointly by the window size , the window stride, the padding size, and the window size of the pooling layer. Then the length of the feature scale satisfies .
[0090] Therefore, the receptive field of the low-scale feature is smaller than that of the high-scale feature , and the feature scale range is also smaller, which is beneficial for short action detection and recognition, while the receptive field of the high-scale feature is larger and the feature scale range is also larger, which is beneficial for long action detection and recognition.
[0091] The output of the multi-scale feature layer is , that is .
[0092] To achieve multi-scale feature fusion, in some alternative embodiments of the present invention, fusing the extracted local sequence features may include: fusing the local sequence features corresponding to multiple rounds of feature extraction at different scales, the class vector corresponding to the digital human action sample, and the position vector corresponding to the action frame to obtain the action feature fusion result. That is to say, the action features output by the multi-scale feature layer are concatenated with the class (CLS) vector corresponding to the entire digital human action sample, which serves as the key technical feature for the model to obtain the overall representation information of the action sequence. By integrating the CLS vector in the multi-scale feature fusion layer, the model can not only capture the features at each scale in the action sequence but also further integrate these features to form a global understanding of the entire action sequence. The introduction of the CLS vector significantly improves the model's ability to grasp the overall structure and context information of the action sequence, thereby achieving a higher level of feature abstraction and more accurate action classification during the action detection and recognition process. Therefore, the model framework provided by the embodiments of the present invention can maintain sensitivity to local action details while also grasping the overall features of the action sequence. The above process can be represented by the following formula:
[0093] ;
[0094] ;
[0095] where, , the superscript is the th layer of the multi-scale feature layer, and the subscript is the th rd feature vector of the layer, is the CLS vector with a dimension of is the position vector of the CLS vector, is 's position vector, and their dimensions are all . The position vector can be set to a fixed value or a learnable model variable, mainly to distinguish and emphasize the feature vectors at different scales. In this way, by introducing a position vector for the action representation at each scale, the precise definition of the action scale and its front-back relationship is achieved. In this multi-scale action encoding and fusion layer, the introduction of the position vector is to retain the time scale information and sequential relationship of the action during the feature fusion process. This design enables the model to more accurately understand and process the long-term and short-term dependencies in the action sequence, thereby improving the accuracy and robustness of action recognition.
[0096] Thus, the input of the action encoding layer is obtained, and its length is Add a superscript to the elements to represent the layer number of the action encoding layer, that is . For the introduction of the action encoding layer, reference can be made to the above embodiments. After passing through the action encoding layer, the obtained action sequence representation is .
[0097] The above is the detailed description of the multi-scale feature layer in the embodiments of the present invention.
[0098] As Figure 6 shown, the action encoding layer can use the output of the multi-scale feature layer as input. In the embodiments of the present invention, the action encoding layer can be composed of layers of Transformer encoders connected in series. One layer of Transformer encoder includes two sub-structures: a self-attention module (Self-Attention) and a feed-forward neural network module (FFN). The former is the input of the latter. In addition, each sub-structure also includes a residual connection (Residual Connection) module and a layer normalization (LayerNorm) module. Since the multi-scale feature layer in the model framework provided by the embodiments of the present invention can extract and fuse multi-scale action features, the action encoding layer can be set with only one layer of Transformer encoder to improve the model training efficiency.
[0099] In the action encoding layer of the digital human action recognition model, the self-attention module can use a single-head self-attention module or a multi-head self-attention module (Multi-Head Self-Attention, MHSA).
[0100] If a single-head self-attention module is used, denote the input of the th layer of Transformer encoder as . The processing method of each layer of Transformer encoder is the same. Therefore, for the convenience of description, in the description of the Transformer encoder, is temporarily ignored. At this time, let , . In the single-head self-attention module, the vector passes through the query value transformation matrix , the key value transformation matrix and the embedding value transformation matrix respectively to generate the query value , the key value and the feature embedding value .
[0101] In the same Transformer encoder, all share a set of , and , and 、 and The element values of
[0102] ;
[0103] Among them, , and respectively represent vector spaces of dimensions and ; The dimensions of the query value and the key value are the same (both ), and their dimensions can be the same as the dimension of the feature embedding value ( ), or different ( ). In the embodiments of the present invention, .
[0104] Secondly, perform a matching calculation on the query value and the key value , that is, calculate the dot product value of any query value and the key value to obtain dot product values , where represents the transposed vector of ( is a column vector, is a row vector). Scale the above dot product values by to obtain ), and at this time let .
[0105] At the same time, the embodiments of the present invention define the mask matrix used in the action representation encoding layer as to represent whether and will perform attention calculation, that is:
[0106] . (1)
[0107] In the action encoding layer, perform attention calculation on both and , that is ( ).
[0108] Then, use the softmax operation to convert the above dot product values and mask values into probability values , that is:
[0109] ;
[0110] Among them , represents and 's logit prediction calculation result, , represents the mask matrix, represents and 's logit prediction calculation result.
[0111] It can be seen from formula (1) and the softmax of multi-head self-attention that when the Transformer encoder encodes the digital human action sequence, any two actions in the digital human action sequence will perform attention calculations on each other.
[0112] Finally, calculate corresponding feature embedding , as follows:
[0113] , where , represents the probability value obtained by converting the above dot product value and mask value using the softmax operation, represents the feature embedding value.
[0114] Therefore, is the output of the single-head self-attention module, where each element corresponds to the input of this module. If the self-attention module of the action encoding layer uses a single-head self-attention module, then the output of the single-head self-attention module can be used as the input of the feed-forward neural network module.
[0115] If a multi-head self-attention module with H heads is used, that is, the process of the above single-head self-attention module is performed in parallel H times, and each result is spliced together and then linearly mapped and output. Specifically, if the output of the th head is , then the rd element of the multi-head attention output is:
[0116] ;
[0117] Among them, 's dimension is , which is consistent with the dimension of , that is 's size is .
[0118] Therefore, is the output of the multi-head self-attention module, where each element corresponds to the input of the module . Then, residual connection and layer normalization operations are performed in sequence, as follows:
[0119] The residual connection is: , where , and here is the output of the residual connection.
[0120] The layer normalization layer is for any , and is calculated using the following formula:
[0121] ; (3)
[0122] Among them, is the mean of all elements in, that is, , where represents the -th element in. is the mean squared error, that is, , and are learnable parameters.
[0123] Therefore, is the output of the multi-head self-attention module, and then it is input into the feed-forward neural network module, where each element corresponds to the input of the module (that is, ).
[0124] The input of the feed-forward neural network module is the output of the self-attention module. Taking the multi-head self-attention module as an example, the input of the feed-forward neural network module is the output of the multi-head self-attention module. After passing through two layers of multi-layer perceptrons, residual connection and layer normalization are performed and then output, specifically:
[0125] ;
[0126] ;
[0127] Among them, is the output after residual connection in the feed-forward neural network module, is the output of layer normalization, is the activation function of the perceptron, which can be selected from sigmoid, tanh, relu, and gelu, etc., and The sizes are respectively and , and LayerNorm is a layer normalization operation, similar to the mechanism of formula (3), which will not be elaborated here.
[0128] Therefore, is the output of the feed-forward neural network module and also the action representation output by this action encoding layer, where each element corresponds to the input of this module (that is, ), that is, is , and they are also the inputs of the -th layer of the Transformer encoding layer.
[0129] The above is the detailed description of the action encoding layer in the embodiment of the present invention.
[0130] The query embedding layer contains a series of action query vectors, that is, the action query vectors , which can be denoted as , and are used for the action decoding layer to query the output of the action encoding layer, where is a preset hyperparameter. Assume that any digital human action sequence contains no more than sub-action sequences, and at the same time, the length of each is . The output of the query mapping layer is .
[0131] The input of the action decoding layer is the query vector and the output of the action encoding layer . Assume a group of 0 vectors with a length of , that is, , that is, each is a -dimensional 0 vector.
[0132] The action decoding layer can have layers, and each layer uses a Transformer decoder for modeling. In the embodiment of the present invention, the position vector of the encoder can also be added to each layer of the Transformer decoder to distinguish and strengthen different query positions. The input of the -th layer of the action decoding layer is the output of the -th layer , the query vector and the position vector , and the output of the -th layer is . First, and added to such that:
[0133] .
[0134] Then, and are input into the Transformer decoder to obtain , that is:
[0135] ;
[0136] That is, in the following form:
[0137] ;
[0138] wherein, , and are respectively the , and th elements of .
[0139] The query vector is added to the input of each layer of the Transformer decoder, so that the input carries the information of the query vector, strengthening the query of the output of the action encoding layer to be more targeted.
[0140] ;
[0141] wherein, , and are respectively the and th elements of . The output of the action decoding layer is .
[0142] Then, as shown in Figure 6 , in the embodiment of the present invention, the action decoding layer may be composed of layers of Transformer decoders connected in series and stacked. One layer of Transformer decoder includes a self-attention module (Self-Attention), a cross-attention module (Cross Attention), and a feed-forward neural network module (FFN). These three modules are connected in series and stacked, and the output of each module is the input of the next module, and the output of the feed-forward neural network module is the input of the self-attention module in the next Transformer decoder.
[0143] In the action encoding layer of the digital human action recognition model, the self-attention module can adopt a single-head self-attention module or a multi-head self-attention module (Multi-Head Self-Attention, MHSA), and its specific structure can refer to the above description. The cross-attention module can adopt a single-head cross-attention module or a multi-head cross-attention module (Multi-Head Cross Attention, MHCA).
[0144] If a single-head cross-attention module is adopted, assume that the input of the single-head cross-attention in the th layer is and .
[0145] At this time, let , , and let , .
[0146] In the single-head cross-attention module, the vector generates a query value after passing through the query value transformation matrix . The vector generates a key value and a feature embedding value
[0147]
[0148]
[0149] ;
[0148] Among them, , , and represent vector spaces with dimensions of and respectively.
[0149] The rest of the single-head cross-attention module is similar to the single-head self-attention module, and the relationship between the multi-head cross-attention module and the single-head cross-attention module is also similar to the relationship between the single-head self-attention module and the single-head self-attention module, which will not be elaborated here.
[0150] The feed-forward neural network module in the Transformer decoder has the same structure as the feed-forward neural network module in the Transformer encoder, and will not be restated here.
[0151] Therefore, in the cross-attention module of the Transformer decoder in the action decoding layer, the action query vector queries the action representation to obtain the first action category (category confidence) of the first predicted action and the first action interval (coordinates) of the first predicted action.
[0152] The above is a detailed introduction to the action decoding layer in the embodiments of the present invention.
[0153] The input of the action output layer is the output of the action decoding layer , that is .
[0154] First After passing through a fully connected layer, we get , as follows:
[0155] ;
[0156] Among them , " " is a matrix-vector multiplication operation, represents the logit prediction calculation result of the t-th output result, has a size of , has a size of dimensions, has a size of dimensions. is the category confidence vector of the -th action predicted by the model.
[0157] Secondly After passing through several layers of perceptrons, we get a two-dimensional vector , then after sigmoid activation, we get a value in the interval , multiply it by the length of the action sequence of the digital human action sample, and then take the integer to obtain the coordinates of the first action interval. The following is an example of a two-layer perceptron:
[0158] ;
[0159] Among them, " " is a matrix-vector multiplication operation, is the activation function of the perceptron, has a size of , The size is , The size is , The size is 2.
[0160] ;
[0161] Among them, is the multiplication operation of the vector and the scalar , is the starting coordinate value and the ending coordinate value of the rd first predicted action predicted by the digital human action recognition model.
[0162] The output of the action output layer , where , represents the class confidence vector and the action starting and ending coordinate values (the first action interval) of the nd first predicted action predicted by the digital human action recognition model. Each action query vector input to the action output layer corresponds to a decoded representation output by the action decoding layer.
[0163] Figure 7 is a schematic structural diagram of another digital human action recognition model provided by an embodiment of the present invention.
[0164] The above embodiment provides a method of first performing multi-scale feature extraction and then performing multi-scale feature fusion. In addition, in some alternative embodiments of the embodiments of the present invention, encoding the pose data of the action frames of the digital human action samples in S202 to obtain an action sequence representation may further include: performing multi-round feature extraction of different scales on the original action features of the action frames corresponding to the pose data of the action frames, fusing the extracted local sequence features to obtain a first action feature fusion result; performing self-attention calculation according to the first action feature fusion result to obtain a second action feature fusion result; after performing spatial size transformation on the second action feature fusion result, fusing it with the local sequence features of the same spatial size to obtain an action feature fusion result. As Figure 7 shown, in the digital human action recognition model provided by the embodiments of the present invention, first perform multi-round feature extraction of different scales through a multi-scale feature layer, then use the action encoding layer to fuse and encode the action features output by the multi-scale feature layer, and then use the spatial size transformation layer to perform spatial size transformation and fusion on the action features, so as to realize the multi-scale integration of action features. Similar to the Figure 6 introduction, adopting the Figure 7 shown model framework, the action encoding layer can also be set with only one layer of Transformer encoder to improve the model training efficiency.
[0165] In some alternative embodiments of the embodiments of the present invention, after performing a spatial size transformation on the second action feature fusion result, fusing it with the local sequence feature having the same spatial size to obtain an action feature fusion result may include: performing at least one upsampling and feature fusion on the second action feature fusion result to obtain an action feature fusion result; in a single upsampling, fusing the input action feature with the local sequence feature having the same spatial size to output a fused upsampled action feature; fusing the second action feature fusion result with the fused upsampled action feature output by the last upsampling to obtain an action feature fusion result. In the embodiments of the present invention, the upsampling layer interpolates the feature embedding to expand the number of features, thereby effectively transmitting the high-level semantic features to the low-level features.
[0166] In some other alternative embodiments of the embodiments of the present invention, after performing a spatial size transformation on the second action feature fusion result, fusing it with the local sequence feature having the same spatial size to obtain an action feature fusion result may further include: performing at least one upsampling and feature fusion on the second action feature fusion result to obtain an action feature fusion result; in a single upsampling, fusing the input action feature with the local sequence feature having the same spatial size to output a fused upsampled action feature; performing at least one downsampling and feature fusion on the fused upsampled action feature output by the last upsampling; in a single downsampling, fusing the fused upsampled action feature with the action feature in the second action feature fusion result having the same spatial size as the action feature input to the downsampling to output a fused downsampled action feature; fusing the fused upsampled action feature output by the last upsampling with the fused downsampled action feature to obtain an action feature fusion result.
[0167] As Figure 7 shown, the input of the spatial size transformation layer is the output of the action encoding layer and the output of the multi-scale feature layer , that is and .
[0168] The upsampling fusion layer can have layers. The input of the first upsampling fusion layer is and , and the output is ; the input of the second upsampling fusion layer is and , and the output is ; and so on. The input of the th upsampling fusion layer is and , and the output is ; until the The input of the layer is and , and the output is . That is:
[0169] ;
[0170] Among them, The operation is to splice two vectors, represents upsampling calculation, There are multiple methods, such as interpolating features and expanding the number of features.
[0171] Similarly, the downsampling fusion layer has a total of layers. The input of the first downsampling fusion layer is and , and the output ; the input of the second downsampling fusion layer is and , and the output ; and so on. The input of the th upsampling fusion layer is and , and the output , until the input of the th downsampling fusion layer is and , and the output . That is:
[0172] ;
[0173] Among them, The operation is to splice two vectors, represents downsampling calculation, There are multiple methods, such as interpolating features or using max pooling and average pooling, etc., to reduce the number of features.
[0174] In the embodiments of the present invention, the design of the parameter allows the model to be flexibly adjusted according to different application scenarios and performance requirements. By increasing the value, the model can integrate more feature layers, thereby improving the detection accuracy, but this will also increase the number of model parameters, increase the computational complexity and processing time. To balance the accuracy and efficiency of the model, the value can be appropriately reduced, which will reduce the model parameters and speed up the training and inference speed, but may sacrifice a certain amount of detection accuracy.
[0175] In addition, the embodiments of the present invention also provide a flexible feature fusion strategy, allowing to estimate features from Specific layers are selectively fused to adapt to different performance requirements and resource constraints. This design enables the model to be flexibly adjusted according to the actual application scenario while maintaining high efficiency, so as to achieve an optimal performance balance.
[0176] Meanwhile, the present invention also provides a general fusion strategy. The multi-scale feature layers selected from top to bottom and the multi-scale feature layers selected from bottom to top do not require one-to-one correspondence. Different numbers and levels of multi-scale features can be selected respectively, which can further enhance the robustness of the model to multi-scale features.
[0177] All stacked together, denoted as , which can be used as the output of the multi-scale feature fusion layer.
[0178] In some other alternative embodiments of the embodiments of the present invention, in addition to Figure 7 the structure shown, the spatial size transformation layer can also be arranged before the action encoding layer. Then, in S202, encoding is performed according to the pose data of the action frames of the digital human action samples to obtain an action sequence representation, which may further include: performing multi-round feature extraction of different scales on the original action features of the action frames corresponding to the pose data of the action frames, fusing the extracted local sequence features to obtain a first action feature fusion result; after performing spatial size transformation on the first action feature fusion result, fusing it with the local sequence features of the same spatial size to obtain a third action feature fusion result; performing self-attention calculation according to the third action feature fusion result to obtain an action feature fusion result. By performing spatial scale transformation and fusion on the basis of multi-scale feature sampling and then performing self-attention calculation, deep semantic information of the action sequence is extracted on the basis of realizing multi-scale integration of action features.
[0179] In some alternative embodiments of the embodiments of the present invention, after performing spatial size transformation on the first action feature fusion result and fusing it with the local sequence features of the same spatial size to obtain a third action feature fusion result, it may include: performing at least one upsampling and feature fusion according to the first action feature fusion result to obtain a third action feature fusion result; in a single upsampling, fusing the input action features with the local sequence features of the same spatial size to output the fused upsampled action features; fusing the first action feature fusion result with the fused upsampled action features output by the last upsampling to obtain an action feature fusion result.
[0180] In some other alternative embodiments of the embodiments of the present invention, after performing a spatial dimension transformation on the first action feature fusion result, feature fusion is performed with the local sequence feature having the same spatial dimension to obtain a third action feature fusion result, which may further include: performing at least one upsampling and feature fusion according to the second action feature fusion result to obtain an action feature fusion result; in a single upsampling, performing feature fusion on the input action feature and the local sequence feature having the same spatial dimension, and outputting the upsampled action feature after fusion; performing at least one downsampling and feature fusion according to the upsampled action feature after fusion output by the last upsampling; in a single downsampling, performing feature fusion on the upsampled action feature after fusion and the action feature having the same spatial dimension as the action feature input for downsampling in the second action feature fusion result, and outputting the downsampled action feature after fusion; performing feature fusion on the upsampled action feature after fusion output by the last upsampling and the downsampled action feature after fusion to obtain an action feature fusion result.
[0181] For the specific implementation of the above spatial dimension transformation (upsampling and downsampling), reference can be made to the introduction in the above embodiments. Similarly, only one layer of Transformer encoder can be set in the action encoding layer to improve the model training efficiency.
[0182] Figure 7 The action mapping layer, multi-scale feature layer and multi-scale feature fusion layer, action encoding layer, query embedding layer, action decoding layer and action output layer shown can be referred to Figure 6 for the introduction.
[0183] In the above embodiments, it is introduced that the digital human action recognition model outputs first predicted actions according to the pose data of the action frames of the digital human action samples. Assuming that there are labeled actions in the digital human action samples, it is necessary to match the first predicted actions with the labeled actions to calculate the model loss value.
[0184] In some alternative embodiments of the embodiments of the present invention, in S203, performing a matching calculation on the first predicted action and the labeled action to obtain a first matching relationship may include: calculating the category confidence loss value between the first action category of the first predicted action and the labeled action category in the labeled action; calculating the action interval loss value between the first action interval of the first predicted action and the labeled action interval of the labeled action; according to the category confidence loss value and the action interval loss value, calculating the first matching relationship between the first predicted action and the labeled action that minimizes the total loss value. That is to say, the loss values are calculated respectively from the two perspectives of the action category and the action interval, so as to determine the best matching relationship between the first predicted action and the labeled action.
[0185] The embodiments of the present invention provide a prediction action screening module to achieve the matching of the first prediction action and the labeled action. The input of the prediction action screening module is the first prediction action , and the labeled action . . The goal of the prediction action screening module is to screen out actions that match the true labeled action one by one from actions, and finally output candidate actions, which correspond one by one to actions of the true label.
[0186] In some alternative embodiments of the embodiments of the present invention, calculating the category confidence loss value between the first action category and the labeled action category in the labeled action may include: calculating the negative value of the product of the logarithm of the confidence value of the first action category and the labeled action to obtain the category confidence loss value.
[0187] Specifically, define the matching weight of the th first prediction action and the th labeled action as , which combines the category confidence loss and the coordinate intersection over union loss. The category confidence loss can be calculated by the following formula:
[0188] ;
[0189] wherein, represents the labeled action category of the th labeled action, which is a one-hot encoded vector with a length of ( is the number of all action categories in the action dataset). The goal of the category confidence loss is to make the predicted category consistent with the true category.
[0190] For the action interval loss value, in some alternative embodiments of the embodiments of the present invention, calculating the action interval loss value between the first action interval and the labeled action interval of the labeled action may include: calculating the first coordinate loss value between the starting action frame of the first action interval and the starting action frame of the labeled action interval; calculating the second coordinate loss value between the ending action frame of the first action interval and the ending action frame of the labeled action interval; calculating the action interval loss value based on the first coordinate loss value and the second coordinate loss value. Specifically, the action interval loss value can be calculated by the following formula:
[0191] ;
[0192] wherein, Indicates the calculation of the second norm. The goal of the action interval loss is to make the predicted coordinates of the action consistent with the true coordinates.
[0193] In some other alternative embodiments of the embodiments of the present invention, calculating the action interval loss value between the first action interval and the labeled action interval of the labeled action may further include: calculating the first coordinate length of the coordinate intersection part of the first action interval and the second action interval; calculating the second coordinate length of the coordinate union part of the first action interval and the second action interval; calculating the negative value of the natural logarithm of the ratio of the first coordinate length to the second coordinate length to obtain the action interval loss value. That is to say, the action interval loss value Can also be calculated by the following formula:
[0194] ;
[0195] Wherein, Represents the natural logarithm, the length of the intersection part between the first action interval output by the model and the labeled action interval is And the length of the union part of the two is . By defining the action coordinate matching weight, instead of using the Euclidean distance between the predicted coordinates and the true coordinates, the intersection over union of the predicted action interval and the true action interval is used. This intersection over union loss effectively solves the problem of unstable gradients in traditional detection methods based on Euclidean distance loss, and at the same time can avoid the inconsistency between model learning and model evaluation, thereby improving the accuracy and robustness of action recognition.
[0196] In still some other alternative embodiments of the embodiments of the present invention, calculating the action interval loss value between the first action interval and the labeled action interval of the labeled action may further include: calculating the minimum circumscribed length of the first action interval and the labeled action interval; calculating the first coordinate length of the coordinate intersection part of the first action interval and the second action interval; calculating the second coordinate length of the coordinate union part of the first action interval and the second action interval; calculating the action interval loss value based on the minimum circumscribed length, the first coordinate length and the second coordinate length. That is to say, the action interval loss value Can also be calculated by the following formula:
[0197]
[0198]
[0199] ;
[0200] Wherein, the minimum circumscribed length between the first action interval output by the model and the labeled action interval is , the length of the intersection part of the two is , and the length of the union part of the two is The defect of the Intersection over Union (IoU) loss is that when the predicted coordinate interval and the true coordinate interval do not intersect, it cannot reflect the distance between the two coordinate intervals, and the loss function is not differentiable at this time. That is, the IoU loss cannot optimize the situation where the two coordinate intervals do not intersect. Therefore, the circumscribed interval loss encloses the first action interval and the labeled action interval with a minimum circumscribed interval, which can measure the distance between the two, and the loss function is differentiable at this time, enabling better model training.
[0201] Using the class confidence loss and the action interval loss introduced above, according to the class confidence loss value and the action interval loss value, calculating the first matching relationship between the first predicted action and the labeled action that minimizes the total loss value can include: screening out the second number of first predicted actions from the first number of first predicted actions output by the digital human action recognition model, calculating the matching weight value between the first predicted action and the labeled action according to the class confidence loss value and the action interval loss value, and calculating the first matching relationship that minimizes the second number of matching weight values; where the second number is the total number of labeled actions corresponding to the digital human action samples.
[0202] Then the th first predicted action and the th labeled action's matching weight can be:
[0203] ;
[0204] where and are hyperparameters, which can be set as constants according to the dataset situation, or can be set as model-learnable parameters.
[0205] Thus, the weight matrix shown in Table 1 can be obtained. Table 1 is the weight matrix of the first predicted action and the labeled action, where each element represents the matching degree between a first predicted action and a labeled action.
[0206] Table 1
[0207]
[0208] The goal of establishing the first matching relationship is to find an optimal matching that minimizes the total loss of matching the first predicted actions selected from actions, one by one, with the labeled actions. Assuming that all selection methods of choosing numbers from form a set , then the matching goal is to select the best matching method , as follows:
[0209] ;
[0210] Among them, , represents the calculation of the minimum value, represents the th matching weight corresponding to the first matching relationship.
[0211] Solving the above formula, the best match is obtained. For example, the Hungarian algorithm or the like can be used.
[0212] The output of the predicted action screening module is the first matching relationship , that is, . Without loss of generality (because the first predicted actions output by the model are independent of each other and the order is irrelevant), it can be assumed that the first are the first matching relationships, that is, the output .
[0213] In the embodiments of the present invention, since multiple action decoding layers are all connected to corresponding action output layers, it is necessary to calculate the loss values corresponding to these action decoding layers and then calculate the total model loss value. That is to say, for the action decoding layers and action output layers participating in loss optimization, first calculate their corresponding first matching relationships, calculate the corresponding loss values according to the first matching relationships, and then calculate the model loss value according to each loss value.
[0214] In some alternative embodiments of the embodiments of the present invention, calculating the loss value of the corresponding decoding layer according to the first matching relationship in S204 may include: according to the first matching relationship, performing weighted calculation on the class confidence loss value and the action interval loss value between the matched first predicted action and the labeled action to obtain a prediction loss value; calculating the loss value according to the prediction loss values corresponding to each labeled action; wherein, the weights of the class confidence loss value and the action interval loss value are coefficients adaptively updated during the iterative training of the digital human action recognition model.
[0215] In the embodiments of the present invention, the model loss value can be calculated through the loss function module. The loss function module receives the output of the predicted action screening module and the corresponding labeled action.
[0216] Then, for the decoding layer participating in loss calculation, obtain the first action category (confidence value) and the first action interval (coordinate value) of the first predicted action output by it, and calculate the loss value with reference to the first matching relationship. For the th decoding layer participating in loss calculation, its corresponding loss value can be calculated by the following formula:
[0217] ;
[0218] Among them, and are model hyperparameters, both of which are positive numbers. They can be set as constants according to the dataset situation, or can be set as model learnable parameters and automatically learned during model training; represents the first action category of the th first predicted action determined according to the first matching relationship in the output result of the th decoding layer participating in loss calculation and the category loss value between it and the corresponding th labeled action, represents the action interval loss value between the first action category of the th first predicted action determined according to the first matching relationship in the output result of the th decoding layer participating in loss calculation and the corresponding th labeled action.
[0219] The calculation method of can adopt the calculation method of the above The calculation method of can adopt one of the above three The model loss value
[0220] .
[0221] It should be noted that here only represents the number of decoding layers participating in loss calculation, and is not necessarily the total number of decoding layers introduced in the above embodiments.
[0222] In some other optional implementation manners of the embodiments of the present invention, calculating the model loss value according to multiple loss values in S204 may also include: calculating the model loss value by performing weighted calculation on multiple loss values. In practical applications, the weights corresponding to each decoding layer can be set to be the same, or different weights can be set. In some optional implementation manners of the embodiments of the present invention, the weights of the action decoding layers closer to the action output layer can be set to be smaller, so that the action decoding layers in the front can also obtain better optimization, so that when performing lightweight calculation in the model inference stage, selecting to retain any action output layer can obtain better model accuracy.
[0223] In some other alternative embodiments of the embodiments of the present invention, in S203, when performing matching calculation on the first predicted action and the labeled action to obtain the first matching relationship, it may further include: determining the matching relationship between the first action category and the labeled action category according to the magnitude of the confidence value of the first action category output by the digital human action recognition model, and further determining the first matching relationship.
[0224] That is to say, the labeled action category can be matched with the first action category first. For example, the first predicted action with the largest category confidence in the first action category that is the same as the labeled action category of the labeled action can be used as the first predicted action matched with the labeled action, so as to quickly determine the best matching relationship between the first predicted action and the labeled action.
[0225] Determining the matching relationship between the first action category and the labeled action category according to the magnitude of the confidence value of the first action category output by the digital human action recognition model, and further determining the first matching relationship may include: screening out candidate actions with the same labeled action category as the first predicted action according to the magnitude of the confidence value of the first action category, and determining the first matching relationship.
[0226] Based on this first matching relationship, in S204, calculating the loss value of the corresponding decoding layer according to the first matching relationship may include: if the number of candidate actions is less than the number of labeled actions, calculating the loss value according to the action interval loss value corresponding to each first matching relationship and the unmatched labeled actions; if the number of candidate actions is equal to the number of labeled actions, calculating the loss value according to the action interval loss value corresponding to each first matching relationship. In practical applications, if the number of candidate actions is less than the number of labeled actions, a penalty value can be set according to the number of unmatched labeled actions, and the model loss value can be jointly calculated in combination with the action interval loss value corresponding to the first matching relationship. The calculation method of the action interval loss value can refer to the three methods introduced in the above embodiments.
[0227] Through the method for determining the first matching relationship and the method for calculating the model loss value provided by the embodiments of the present invention, the first matching relationship can be quickly determined and the model loss can be calculated.
[0228] Based on the digital human action recognition model provided by the embodiments of the present invention, the number of action decoding layers can be flexibly configured to adapt to different application scenarios, dataset characteristics, and real-time requirements. In the model training stage, all action decoding layers can participate in the calculation of the loss function to ensure that the model can comprehensively learn and optimize. In the model inference stage, the user can choose to use the results of some action output layers instead of all according to their own trade-off between accuracy and inference speed.
[0229] During the model training stage, by calculating the loss value for multiple action decoding layers and setting the weight coefficients with adaptive updates, the training process of the digital human action recognition model adaptively optimizes each action decoding layer, so as to ensure the model accuracy to the greatest extent when the digital human action recognition model is lightweight processed during the application stage. When performing the digital human action recognition task, the digital human action recognition model can be lightweight processed as needed, only retaining some action decoding layers, reducing the model parameters and the model calculation amount, while ensuring the model accuracy to the greatest extent.
[0230] It can be understood that there is a trade-off relationship between the selection of the action output layer and the inference speed and prediction accuracy of the model. If more action output layers are selected, the accuracy of the prediction result can be improved because the model can make decisions by synthesizing information from multiple levels. However, this also means an increase in the computational burden, thus reducing the inference speed. If fewer output layers are selected, the computational amount can be reduced and the inference speed can be improved, but it may sacrifice a certain prediction accuracy because the model loses the opportunity to synthesize information from more levels. Through this flexible design, users can make the best choice between the inference speed and the prediction accuracy according to specific application requirements, and achieve the optimization of the model performance.
[0231] In the embodiment of the present invention, when using the trained digital human action recognition model with some decoding layers retained to perform the action recognition task of the digital human action sequence to be recognized, it includes: receiving the model accuracy requirement parameter; determining the number of decoding layers to be retained according to the model accuracy requirement parameter, so as to use the trained digital human action recognition model with some decoding layers retained to perform the action recognition task.
[0232] Among them, determining the number of decoding layers to be retained according to the model accuracy requirement parameter may include: using the digital human action test samples to test the model accuracy test results of the trained digital human action recognition model with different layers of decoding layers retained; matching the model accuracy test results according to the model accuracy requirement parameter to determine the number of decoding layers to be retained.
[0233] In addition, as introduced in the above embodiments, one action encoding layer can be set during the model training stage, or multiple action encoding layers can be set. If multiple action encoding layers are set during the model training stage, only some action encoding layers can also be retained during the model inference stage. That is, in S205, using the trained digital human action recognition model with some decoding layers retained to perform the action recognition task of the digital human action sequence to be recognized may include: using the digital human action recognition model with some encoding layers and some decoding layers retained to perform the action recognition task; among them, the encoding layer is used to perform self-attention calculation.
[0234] Through the digital human action recognition model provided by the embodiments of the present invention, one or more action encoding layers can be set during the model training stage according to requirements, and the action encoding layer and the action decoding layer can be lightweight processed according to requirements during the model inference stage, so as to adapt to different scenario requirements.
[0235] The above embodiments introduce the digital human action recognition method. In addition, the embodiments of the present invention also provide a method for training a digital human action recognition model, which may include: obtaining digital human action samples; using the digital human action recognition model to encode according to the pose data of the action frames of the digital human action samples to obtain action sequence representations, and performing multi-layer decoding calculations according to the action sequence representations and action query vectors; obtaining a first predicted action output by the last decoding layer, and performing a matching calculation on the first predicted action and the labeled action to obtain a first matching relationship; according to the first matching relationship, calculating the loss values of the output results of multiple decoding layers compared with the labeled action, and calculating a model loss value according to the multiple loss values to optimize the loss of the digital human action recognition model, so as to obtain a trained digital human action recognition model.
[0236] The method for training a digital human action recognition model provided by the embodiments of the present invention can refer to the introduction in the above digital human action recognition method. By using the method for training a digital human action recognition model provided by the embodiments of the present invention, the output of each action decoding layer can be directly involved in the loss calculation of action detection during the training process of the digital human action recognition model, so as to ensure that the model can contribute to the action detection task at each level, so as to perform lightweight processing when needed.
[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation.
[0238] Figure 8 It is a schematic structural diagram of a digital human action recognition device provided by an embodiment of the present invention.
[0239] As Figure 8 shown, the embodiments of the present application also provide a digital human action recognition device, including:
[0240] An acquisition module 801, configured to acquire digital human action samples;
[0241] A model training module 802, configured to use the digital human action recognition model to encode according to the pose data of the action frames of the digital human action samples to obtain action sequence representations, and perform multi-layer decoding calculations according to the action sequence representations and action query vectors;
[0242] A prediction action screening module 803, configured to obtain a first predicted action output by a last decoding layer, perform matching calculation on the first predicted action and an annotated action, and obtain a first matching relationship;
[0243] A loss function module 804, configured to calculate a loss value of the output results of multiple decoding layers relative to the annotated action according to the first matching relationship, calculate a model loss value according to multiple loss values to optimize the loss of the digital human action recognition model, and obtain a trained digital human action recognition model;
[0244] An identification module 805, configured to use the trained digital human action recognition model with some decoding layers retained to perform an action recognition task on a digital human action sequence to be recognized.
[0245] In an embodiment of the present invention, the prediction action screening module 803 performs matching calculation on the first predicted action and the annotated action to obtain a first matching relationship, which may include: calculating a category confidence loss value between a first action category and an annotated action category in the annotated action; calculating an action interval loss value between a first action interval and an annotated action interval of the annotated action; and calculating a first matching relationship between the first predicted action and the annotated action that minimizes the total loss value according to the category confidence loss value and the action interval loss value.
[0246] In an embodiment of the present invention, the loss function module 804 calculates a model loss value according to multiple loss values, which may include: calculating a model loss value through weighted calculation according to multiple loss values.
[0247] In an embodiment of the present invention, the identification module 805 uses the trained digital human action recognition model with some decoding layers retained to perform an action recognition task on a digital human action sequence to be recognized, which may include: receiving a model accuracy requirement parameter; determining the number of decoding layers to be retained according to the model accuracy requirement parameter, so as to perform an action recognition task using the trained digital human action recognition model with some decoding layers retained.
[0248] In an embodiment of the present invention, the identification module 805 determines the number of decoding layers to be retained according to the model accuracy requirement parameter, which may include: using digital human action test samples to test the model accuracy test results of the trained digital human action recognition models with different numbers of decoding layers retained; and matching the model accuracy test results according to the model accuracy requirement parameter to determine the number of decoding layers to be retained.
[0249] In an embodiment of the present invention, the model training module 802 encodes the pose data of the action frames of the digital human action samples to obtain an action sequence representation, which may include: performing multi-round feature extraction of different scales on the original action features of the action frames corresponding to the pose data of the action frames, fusing the extracted local sequence features to obtain a first action feature fusion result; performing multi-layer self-attention calculation based on the first action feature fusion result to obtain an action sequence representation.
[0250] In an embodiment of the present invention, the recognition module 805 uses the trained digital human action recognition model with some decoding layers retained to perform the action recognition task of the digital human action sequence to be recognized, including: using the digital human action recognition model with some encoding layers and some decoding layers retained to perform the action recognition task; wherein, the encoding layer is used to perform self-attention calculation.
[0251] In an embodiment of the present invention, the model training module 802 encodes the pose data of the action frames of the digital human action samples to obtain an action sequence representation, which may include: performing multi-round feature extraction of different scales on the original action features of the action frames corresponding to the pose data of the action frames, fusing the extracted local sequence features to obtain a first action feature fusion result; performing self-attention calculation based on the first action feature fusion result to obtain a second action feature fusion result; after performing spatial size transformation on the second action feature fusion result, fusing it with the local sequence features of the same spatial size to obtain an action feature fusion result. Extracting action features of multiple scales from the pose data of the action frames, performing feature fusion and self-attention calculation to obtain an action sequence representation.
[0252] In an embodiment of the present invention, when the model training module 802 performs spatial size transformation on the second action feature fusion result and then fuses it with the local sequence features of the same spatial size to obtain an action feature fusion result, it may include: performing at least one upsampling and feature fusion based on the second action feature fusion result to obtain an action feature fusion result; in a single upsampling, fusing the input action features with the local sequence features of the same spatial size and outputting the fused upsampled action features; performing feature fusion based on the second action feature fusion result and the fused upsampled action features output by the last upsampling to obtain an action feature fusion result.
[0253] In an embodiment of the present invention, after performing a spatial size transformation on the second action feature fusion result, the model training module 802 performs feature fusion with local sequence features having the same spatial size to obtain an action feature fusion result, which may include: performing at least one upsampling and feature fusion on the second action feature fusion result to obtain an action feature fusion result; in a single upsampling, performing feature fusion on the input action features and local sequence features having the same spatial size, and outputting the upsampled action features after fusion; performing at least one downsampling and feature fusion on the upsampled action features after fusion output by the last upsampling; in a single downsampling, performing feature fusion on the upsampled action features after fusion and action features in the second action feature fusion result that have the same spatial size as the action features input for downsampling, and outputting the downsampled action features after fusion; and performing feature fusion on the upsampled action features after fusion output by the last upsampling and the downsampled action features after fusion to obtain an action feature fusion result.
[0254] In an embodiment of the present invention, the model training module 802 encodes the pose data of the action frames of the digital human action samples to obtain an action sequence representation, which may include: performing multi-round feature extraction of different scales on the original action features of the action frames corresponding to the pose data of the action frames, performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; performing a spatial size transformation on the first action feature fusion result, and performing feature fusion with local sequence features having the same spatial size to obtain a third action feature fusion result; and performing self-attention calculation based on the third action feature fusion result to obtain an action feature fusion result.
[0255] For the description of the features in the corresponding embodiments of the digital human action recognition device, reference may be made to the relevant descriptions in the corresponding embodiments of the digital human action recognition method, which will not be elaborated here one by one.
[0256] An embodiment of the present application further provides a digital human action recognition model training device, which may include:
[0257] An acquisition module 801, configured to acquire digital human action samples;
[0258] A model training module 802, configured to use a digital human action recognition model to encode the pose data of the action frames of the digital human action samples to obtain an action sequence representation, and perform multi-layer decoding calculations based on the action sequence representation and an action query vector;
[0259] A predicted action screening module 803, configured to obtain a first predicted action output by the last decoding layer, perform matching calculation on the first predicted action and a labeled action to obtain a first matching relationship;
[0260] A loss function module 804 is configured to calculate, according to a first matching relationship, a loss value of the output results of multiple decoding layers compared with the labeled actions, and calculate a model loss value based on the multiple loss values to optimize the loss of the digital human action recognition model, so as to obtain a trained digital human action recognition model.
[0261] For the description of the features in the corresponding embodiments of the digital human action recognition model training device, reference can be made to the relevant descriptions in the corresponding embodiments of the digital human action recognition method, which will not be elaborated here one by one.
[0262] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps of any one of the above digital human action recognition methods or the steps of the above digital human action recognition model training method.
[0263] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above digital human action recognition methods or the steps of the above digital human action recognition model training method when running.
[0264] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical disks and other media that can store computer programs.
[0265] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the above digital human action recognition methods or the steps of the above digital human action recognition model training method.
[0266] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above digital human action recognition method embodiments.
[0267] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0268] The above has introduced in detail a digital human action recognition method, device, medium, and product provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for digital human action recognition, characterized in that: include: Obtain digital human action samples; Using a digital human action recognition model to encode the posture data of the action frame of the digital human action sample to obtain an action sequence representation, and performing multi-layer decoding calculations based on the action sequence representation and the action query vector; Obtaining a first predicted action output by a last decoding layer, performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship; According to the first matching relationship, the loss values of the output results of the plurality of decoding layers compared to the labeled action are calculated, and the model loss value is calculated according to the plurality of loss values to perform loss optimization on the digital human action recognition model, so as to obtain the trained digital human action recognition model; Using the trained digital human action recognition model retaining part of the decoding layer to perform the action recognition task of the digital human action sequence to be recognized; According to the first matching relationship, calculating the loss values of the output results of the plurality of decoding layers compared to the labeled action includes: for each decoding layer participating in the loss calculation, determining the labeled action corresponding to the first predicted action output by the decoding layer according to the first matching relationship, so as to determine the loss value of the decoding layer according to the first predicted action and the corresponding labeled action; The model loss value is calculated according to the multiple loss values, including: the model loss value is obtained by weighted calculation according to the multiple loss values, and the weight of the decoding layer closer to the action output layer of the digital human action recognition model is smaller.
2. The method for digital human motion recognition according to claim 1, characterized in that: Performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship includes: Calculating a category confidence loss value of a first action category of the first predicted action and a labeled action category in the labeled action; Calculating an action interval loss value of a first action interval of the first predicted action and a marked action interval of the marked action; The first matching relationship between the first predicted action and the labeled action that minimizes the total loss value is calculated according to the category confidence loss value and the action interval loss value.
3. The method for digital human motion recognition according to claim 1, characterized in that: The action recognition task of the digital human action sequence to be recognized is performed by using the trained digital human action recognition model with a portion of the decoding layer retained, including: Receive model accuracy requirement parameters; The number of the retained decoding layers is determined according to the model accuracy requirement parameter, so as to perform the action recognition task by using the trained digital human action recognition model that retains part of the decoding layers.
4. The method for digital human motion recognition according to claim 3, characterized in that: Determining the number of the decoding layers to be retained according to the model accuracy requirement parameter includes: Using digital human action test samples to test and retain the model accuracy test results of the trained digital human action recognition model of different decoding layers; The number of the decoding layers to be retained is determined by matching the model accuracy test result with the model accuracy requirement parameter.
5. The method for digital human motion recognition according to claim 1, characterized in that: Encoding the posture data of the action frame of the digital human action sample to obtain an action sequence representation includes: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; A multi-layer self-attention calculation is performed according to the first action feature fusion result to obtain the action sequence representation.
6. The method for digital human motion recognition according to claim 5, characterized in that: The action recognition task of the digital human action sequence to be recognized is performed by using the trained digital human action recognition model with a portion of the decoding layer retained, including: Utilizing the digital human action recognition model that retains part of the encoding layer and retains part of the decoding layer to perform the action recognition task; The encoding layer is used to perform self-attention calculation.
7. The method for digital human motion recognition according to claim 1, characterized in that: Encoding the posture data of the action frame of the digital human action sample to obtain an action sequence representation includes: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; Performing self-attention calculation according to the first action feature fusion result to obtain a second action feature fusion result; After the second action feature fusion result is transformed in spatial size, feature fusion is performed with the local sequence features with the same spatial size to obtain the action feature fusion result.
8. The method for digital human motion recognition according to claim 7, characterized in that: After performing spatial size transformation on the second action feature fusion result, performing feature fusion with the local sequence feature having the same spatial size to obtain the action feature fusion result includes: Performing upsampling and feature fusion at least once according to the second action feature fusion result to obtain the action feature fusion result; In a single upsampling, the input action feature and the local sequence feature with the same spatial size are subjected to feature fusion, and the fused upsampled action feature is output; Feature fusion is performed based on the second action feature fusion result and the fused up-sampled action feature outputted from the last up-sampling to obtain the action feature fusion result.
9. The method for digital human motion recognition according to claim 7, characterized in that: After performing spatial size transformation on the second action feature fusion result, performing feature fusion with the local sequence feature having the same spatial size to obtain the action feature fusion result includes: Performing upsampling and feature fusion at least once according to the second action feature fusion result to obtain the action feature fusion result; In a single upsampling, the input action feature and the local sequence feature with the same spatial size are subjected to feature fusion, and the fused upsampled action feature is output; Perform at least one downsampling and feature fusion according to the fused upsampled action features outputted from the last upsampling; In a single downsampling, the fused upsampled motion features and the motion features in the second motion feature fusion result having the same spatial size as the downsampled input motion features are subjected to feature fusion, and the fused downsampled motion features are output; The up-sampled motion features after fusion and the down-sampled motion features after fusion outputted from the last up-sampling are subjected to feature fusion to obtain the motion feature fusion result.
10. The method for digital human motion recognition according to claim 1, characterized in that: Encoding the posture data of the action frame of the digital human action sample to obtain an action sequence representation includes: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; After performing spatial size transformation on the first action feature fusion result, feature fusion is performed with the local sequence features having the same spatial size to obtain a third action feature fusion result; A self-attention calculation is performed according to the third action feature fusion result to obtain the action feature fusion result.
11. A digital human action recognition model training method, characterized in that: include: Obtain digital human action samples; Using a digital human action recognition model to encode the posture data of the action frame of the digital human action sample to obtain an action sequence representation, and performing multi-layer decoding calculations based on the action sequence representation and the action query vector; Obtaining a first predicted action output by a last decoding layer, performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship; According to the first matching relationship, the loss values of the output results of the plurality of decoding layers compared to the labeled action are calculated, and the model loss value is calculated according to the plurality of loss values to perform loss optimization on the digital human action recognition model, so as to obtain the trained digital human action recognition model; Wherein, according to the first matching relationship, calculating the loss values of the output results of the plurality of the decoding layers compared to the labeled action comprises: for each of the decoding layers participating in the loss calculation, determining the labeled action corresponding to the first predicted action output by the decoding layer according to the first matching relationship, so as to determine the loss value of the decoding layer according to the first predicted action and the corresponding labeled action; The model loss value is calculated according to the multiple loss values, including: the model loss value is obtained by weighted calculation according to the multiple loss values, and the weight of the decoding layer closer to the action output layer of the digital human action recognition model is smaller.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the digital human action recognition method according to any one of claims 1 to 10 or the steps of the digital human action recognition model training method according to claim 11 when executing the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the digital human action recognition method according to any one of claims 1 to 10 or the steps of the digital human action recognition model training method according to claim 11.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the digital human action recognition method according to any one of claims 1 to 10 or the steps of the digital human action recognition model training method according to claim 11 are implemented.