A digital human motion recognition method, device, medium and product
By training the digital human action recognition model, extracting and fusing action features of multiple scales, the problem of low action recognition efficiency of digital human action sequences is solved, and accurate recognition and efficient annotation of action sequences of different lengths is achieved.
Patent Information
- Application Number
- CN202510233293.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The lack of action recognition schemes for digital human action sequences in the prior art, resulting in low efficiency in digital human generation tasks.
By training the digital human action recognition model, using the posture data of the action frame to extract action features from multiple scales for feature fusion, combining the action query vector to calculate and predict action categories and intervals, and updating model parameters through matching calculation and loss optimization.
The action recognition efficiency of digital human action sequences is improved, adapted to action sequences of different lengths, maintained the accuracy of long action recognition, and improved the performance of short action recognition, and improved the accuracy and labeling efficiency of action recognition of digital human action sequences.
Smart Images

Figure CN119741759B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital human technology, and in particular to a digital human action recognition method, device, medium and product. Background Art
[0002] Virtual digital human (abbreviated as "digital human") is one of the applications in the field of 3D vision and generative artificial intelligence, and is widely used in multiple fields such as augmented reality, virtual reality, digital twins, film and television, and game production. In order to achieve the accuracy of digital human generation in specific scenes, a large number of labeled data sets are required. However, the numerous labeling tasks lead to unsatisfactory efficiency of digital human generation tasks. Therefore, it is necessary to propose a solution for identifying the actions of digital human action sequences.
[0003] How to improve the efficiency of motion recognition of digital human motion sequences is a technical problem that those skilled in the art need to solve. Summary of the invention
[0004] The present application provides a method, device, medium and product for digital human motion recognition, so as to at least solve the problem of lack of motion recognition scheme for digital human motion sequence in the related art.
[0005] The present application provides a method for digital human action recognition, comprising:
[0006] Extracting motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample using the digital human motion recognition model and performing feature fusion, and calculating a first motion category of a first predicted motion and a first motion interval of the first predicted motion according to the motion feature fusion result and the motion query vector;
[0007] Performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship;
[0008] Calculate the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model, and obtain the trained digital human action recognition model;
[0009] Using the trained digital human action recognition model, according to the posture data of the digital human action sequence to be recognized, a second action category of a second predicted action and a second action interval of the second predicted action are calculated;
[0010] Among them, action features of different scales correspond to different numbers of the action frames.
[0011] The present application also provides a method for training a digital human action recognition model, comprising:
[0012] Extracting motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample using the digital human motion recognition model and performing feature fusion, and calculating a first motion category of a first predicted motion and a first motion interval of the first predicted motion according to the motion feature fusion result and the motion query vector;
[0013] Performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship;
[0014] Calculate the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model, and obtain the trained digital human action recognition model;
[0015] Among them, action features of different scales correspond to different numbers of the action frames.
[0016] The present application also provides an electronic device, comprising: a memory for storing a computer program; a processor for implementing the steps of the above-mentioned digital human action recognition method or the above-mentioned digital human action recognition model training method when executing the computer program.
[0017] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the above-mentioned digital human action recognition method or the steps of the above-mentioned digital human action recognition model training method are implemented.
[0018] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned digital human action recognition method or the above-mentioned digital human action recognition model training method.
[0019] Through the present application, the action recognition of the digital human action sequence is realized by training the digital human action model. During the training, the digital human action recognition model is used to extract action features of multiple scales according to the posture data of the action frame of the digital human action sample and perform feature fusion. The first action category and the first action interval of the first predicted action are calculated according to the action feature fusion result and the action query vector. The first predicted action and the marked action are matched and calculated to obtain a first matching relationship. Then, the model loss value is calculated according to the first matching relationship to update the parameters of the digital human action recognition model until the training is completed. Then, the trained digital human action recognition model is used to identify the action from the digital human action sequence to be identified. This not only makes up for the problem of the lack of digital human action recognition solutions in the related art and improves the efficiency of action recognition of the digital human sequence, but also can adapt to action sequences of different lengths, while maintaining the accuracy of long action recognition, the performance of short action recognition is significantly improved, thereby improving the accuracy of action recognition of the digital human action sequence, helping to improve the annotation efficiency of the digital human action sequence, thereby improving the efficiency of training the digital human generation model to solve downstream scene problems, and can also be used to test the digital human generation effect in order to optimize the digital human generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of a digital human action sequence with labels provided by an embodiment of the present invention;
[0022] Figure 2 A flowchart of a digital human action recognition method provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of the structure of a digital human motion recognition system provided by an embodiment of the present invention;
[0024] Figure 4 A training flow chart of a digital human action recognition model provided by an embodiment of the present invention;
[0025] Figure 5 A reasoning flow chart of a digital human action recognition model provided by an embodiment of the present invention;
[0026] Figure 6 A schematic diagram of the structure of a digital human action recognition model provided by an embodiment of the present invention;
[0027] Figure 7A schematic diagram of the structure of another digital human action recognition model provided by an embodiment of the present invention;
[0028] Figure 8 A schematic diagram of the structure of a digital human motion recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0032] Here, some key terms used in the embodiments of the present invention are explained.
[0033] Digital human refers to a virtual character image that is close to human image and created through digital technology. It has human appearance, behavior patterns, voice, emotional response and other characteristics, and can operate and exist independently in digital space. Digital human is usually generated by computer graphics, artificial intelligence, natural language processing and other technologies, and has interactive capabilities.
[0034] Digital human action sequences refer to a series of data or instructions generated by computer technology to drive digital humans to perform coherent actions. These action sequences can include facial expressions, body movements, gestures, etc., and can be created through preset animation data, motion capture technology, or generative models based on deep learning.
[0035] A digital human action sequence consists of multiple action frames. Figure 1 A schematic diagram of a digital human action sequence with labels provided by an embodiment of the present invention. Figure 1As shown in the figure, a three-dimensional digital human model in the related art is composed of a hinged skeleton tree composed of multiple joints. The hierarchical structure of this skeleton tree defines the parent-child relationship of the nodes. Different postures of human body movements can be obtained by adjusting the rotation angle of the child node relative to the parent node. Each posture and action of the digital human based on the three-dimensional digital human model is determined by the relative rotation angle of these joints. Therefore, the action sequence of the digital human can be regarded as an action frame sequence composed of the rotation angles of multiple joints.
[0036] In order to realize the training of digital human generation models and various downstream applications, a large number of data sets have emerged in the field of digital humans. In order to be put into actual training, these data sets need to be labeled. The semantic labels corresponding to the action sequences of digital humans can include two levels: sequence level and action frame level. Figure 1 As shown, the digital human action sequence has 15 action frames, and the corresponding sequence-level label is "playing basketball". The action frame-level labels corresponding to some of the action frames include "catching the ball with both hands", "transition action", "transferring the basketball to the left hand", "sprinting", and "dribbling with the left hand".
[0037] The numerous digital human data sets bring a lot of labeling work. In order to realize the automatic labeling of digital human data sets, it is necessary to solve the problem of action recognition of digital human action sequences.
[0038] Currently, artificial intelligence solutions driven by deep learning technology have demonstrated excellent performance in many fields, especially in video character motion analysis. However, since the representation of digital human motion sequences is a new modality, the deep learning model used in video motion recognition tasks cannot be applied to the motion recognition of digital human motion sequences.
[0039] In order to solve the problem of action recognition of digital human action sequences, an embodiment of the present invention provides a digital human action recognition solution. By training a digital human action recognition model to perform a digital human action recognition task, during the training, the digital human action recognition model is used to extract action features of multiple scales according to the posture data of the action frame of the digital human action sample and perform feature fusion. According to the action feature fusion result and the action query vector, a first action category and a first action interval of a first predicted action are calculated, and a matching calculation is performed on the first predicted action and the marked action to obtain a first matching relationship. Then, according to the first matching relationship, a model loss value is calculated to update the parameters of the digital human action recognition model until the training is completed. Then, the trained digital human action recognition model is used to identify actions from the digital human action sequence to be identified. This not only makes up for the problem of the lack of digital human action recognition solutions in the related art, improves the efficiency of action recognition of digital human sequences, but also can adapt to action sequences of different lengths, significantly improves the performance of short action recognition while maintaining the accuracy of long action recognition, thereby improving the accuracy of action recognition of digital human action sequences, helping to improve the annotation efficiency of digital human action sequences, thereby improving the efficiency of training digital human generation models to solve downstream scene problems, and can also be used to test digital human generation effects in order to optimize digital human generation effects.
[0040] The embodiment of the present application provides a method for digital human motion recognition. Combined with the execution flow of the method for digital human motion recognition, the method is described in detail below.
[0041] Figure 2 A flowchart of a digital human action recognition method provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a digital human motion recognition system provided by an embodiment of the present invention; Figure 4 A training flow chart of a digital human action recognition model provided by an embodiment of the present invention; Figure 5 A reasoning flow chart of a digital human action recognition model provided by an embodiment of the present invention.
[0042] like Figure 2 As shown, the digital human action recognition method provided by the embodiment of the present invention includes:
[0043] S201: extracting action features of multiple scales from the posture data of the action frames of the digital human action samples using a digital human action recognition model and performing feature fusion, and calculating a first action category of a first predicted action and a first action interval of the first predicted action according to the action feature fusion result and the action query vector;
[0044] S202: performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship;
[0045] S203: Calculating the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model, and obtaining a trained digital human action recognition model;
[0046] S204: using the trained digital human action recognition model to calculate the second action category of the second predicted action and the second action interval of the second predicted action according to the posture data of the digital human action sequence to be recognized;
[0047] Among them, action features of different scales correspond to different numbers of action frames.
[0048] In the embodiment of the present invention, the posture data of the action frame refers to the action parameters used to describe the action of the digital human in the action frame, for example, the rotation angle of the joints of the digital human.
[0049] To realize digital human motion recognition, the digital human motion recognition method provided by the embodiment of the present invention mainly includes two steps: training a digital human motion recognition model and using the digital human motion recognition model to perform digital human motion recognition tasks. In this regard, the digital human motion recognition method provided by the embodiment of the present invention can be applied to Figure 3 The digital human motion recognition system shown includes a model training system and a digital human motion recognition device.
[0050] like Figure 3 As shown, the model training system can be provided with computing power and storage support by multiple artificial intelligence servers (artificial intelligence servers 1~n), and the digital human action recognition model can be trained based on the model training system using digital human action samples. In some optional implementations of the embodiments of the present invention, a computing resource pool can be constructed based on the model training system, and the computing resource pool includes multiple computing threads, so that the training task of the digital human action recognition model can be split into multiple subtasks for parallel execution. The model parallel training method can be adopted, and the data parallel training method can also be adopted. The model training system can be deployed in a cloud computing center. The equipment structure of the cloud computing center can be similar to Figure 3 The digital human motion recognition device shown includes basic settings such as artificial intelligence processor, storage, input and output devices, communication bus, communication interface, etc., and is composed of software environments such as operating system, database, middleware, and application software. The model training system receives manually labeled digital human motion samples, and can use the stored pre-trained model for fine-tuning to obtain a digital human motion recognition model, or can use a randomly initialized model framework for training to obtain a digital human motion recognition model.
[0051] The digital human motion recognition device can be a computing device or terminal device deployed in a cloud computing center, and its structure is as follows: Figure 3As shown in the digital human action recognition device in. If the computing device of the cloud computing center is used, after the trained digital human action recognition model is deployed, the terminal device uploads the digital human action sequence to be recognized to the cloud computing center, and uses the trained digital human action recognition model to perform action recognition on the digital human action sequence to be recognized, and outputs the predicted action. If the digital human action recognition device uses a terminal device, the model training system trains the visual digital human action recognition model and sends it to the terminal device through the communication unit, and deploys it in the storage of the terminal device. The terminal device receives the digital human action sequence to be recognized, performs prediction calculation through the local digital human action recognition model, and outputs the predicted action.
[0052] like Figure 4 As shown, when training the digital human action recognition model, the posture data of the action frame of the digital human action sample is input into the digital human action recognition model, and the first predicted action is output. After matching the first predicted action with the labeled action, the loss value of the model is substituted into the loss function to calculate the model loss value, and the model loss value is used to optimize the loss of the digital human action recognition model.
[0053] In the specific implementation of the digital human motion recognition method provided in the embodiment of the present invention, the digital human motion samples can come from the local storage of the artificial intelligence device, the shared memory of the model training system, or the digital human motion samples received from an external device.
[0054] For S201, the posture data of the action frame of the digital human action sample (for example, one frame of action includes the joint rotation angles of 24 joints) is input into the digital human action recognition model, and the first action category and the first action interval of the first predicted action are output.
[0055] The posture data of the action frame of the digital human action sample is recorded as (Include Action frames, Representative Frame pose data), the set of labeled actions of the digital human action samples is recorded as (Include annotation actions), where Indicates The labeling action category of labeling actions is a length of The one-hot encoded vector of is the number of all action categories in the action dataset), Represents the coordinates of the annotation action interval of the annotation action, that is, is the sequence number of the action start frame, It is the sequence number of the action end frame, that is, the sub-action sequence The action annotation category is .
[0056] The posture data of the action frame of the digital human action sample Input the digital human action recognition model, and the digital human action recognition model predicts The first predicted action ( is a hyperparameter, representing the maximum number of actions that the digital human action recognition model can recognize from a digital human action sequence. It is usually required ), predicted by the digital human action recognition model The first predicted actions form a first predicted action set , No. The first predicted action includes a first action category of the first predicted action (Length is The category confidence vector, where each element represents the confidence probability of a category, all are positive and sum to 1) and the first action interval (in ).
[0057] That is to say, for The digital human action samples of the frame have The digital human action recognition model is configured to recognize the posture data of the action frame of the digital human action sample. The first predicted action.
[0058] Since the digital human action sequence may include actions of different lengths (i.e. corresponding to different numbers of action frames), in an embodiment of the present invention, a digital human action recognition model is used to extract action features of multiple scales based on the posture data of the action frames of the digital human action samples and perform feature fusion. The first action category of the first predicted action and the first action interval of the first predicted action are calculated based on the action feature fusion result and the action query vector, so that the digital human action recognition model can adapt to action sequences of different lengths (i.e. corresponding to different numbers of action frames), while maintaining the accuracy of long action recognition, the performance of short action recognition is significantly improved, and the accuracy of action recognition of digital human action sequences is improved.
[0059] For S202, a matching calculation is performed based on the parameters of the first predicted action and the parameters of the marked action to determine a matching relationship between the first predicted action and the marked action, which is recorded as a first matching relationship in the embodiment of the present invention. of Filter and mark action sets from elements The marked actions are matched one by one The first predicted action, without loss of generality, assume Before The actions are the selected actions, so the first predicted action and annotation actions Loss calculations are performed one-to-one for gradient calculation and model training.
[0060] During model training, the best model prediction confidence for each category can be set through the test set or validation set ,in There are many ways to set this up. For example, if There are 50 validation data for each category. It can be set to the lowest probability predicted by the model for these 50 data, or the probability value of the action ranked 40th, etc.
[0061] For S203, according to the first matching relationship, the parameters of the first predicted action and the parameters of the labeled action are substituted into the loss function, the loss value of the first predicted action and the corresponding labeled action is calculated, and the model loss value is calculated according to each loss value.
[0062] The training end condition may be that the number of iterative training times of the digital human motion recognition model reaches a preset number of iterations, or the model loss value of the digital human motion recognition model is less than a preset loss value.
[0063] For S204, the trained digital human action recognition model is used to perform the digital human action recognition task, the posture data of the digital human action sequence to be recognized is input into the trained digital human action recognition model, and the second action category of the second predicted action and the second action interval of the second predicted action are output. The step of outputting the second predicted action using the trained digital human action recognition model is the same as the step of calculating the first predicted action.
[0064] like Figure 5 As shown, the posture data of the action frame of the action sequence of the digital human to be recognized Input into the digital human action recognition model, the model predicts The second predicted action includes the second action category and the second action interval (i.e., the The category confidence of an action is , the coordinates are ,in . The largest element in is , that is, the model predicts the action sequence The most likely category is ,if , then keep the prediction result; if , then discard the prediction result. The second prediction actions are processed one by one, and the remaining The result is used as the second predicted action output.
[0065] The digital human action recognition method provided by the embodiment of the present invention performs the digital human action recognition task by training the digital human action recognition model. During the training, the digital human action recognition model is used to extract action features of multiple scales according to the posture data of the action frame of the digital human action sample and perform feature fusion. The first action category and the first action interval of the first predicted action are calculated according to the action feature fusion result and the action query vector. The first predicted action and the marked action are matched and calculated to obtain a first matching relationship. Then, the model loss value is calculated according to the first matching relationship to update the parameters of the digital human action recognition model until the training is completed. Then, the trained digital human action recognition model is used to identify the action from the digital human action sequence to be identified. This not only makes up for the problem of the lack of digital human action recognition solutions in the related art and improves the efficiency of action recognition of digital human sequences, but also can adapt to action sequences of different lengths, and significantly improves the performance of short action recognition while maintaining the accuracy of long action recognition, thereby improving the accuracy of action recognition of digital human action sequences, helping to improve the annotation efficiency of digital human action sequences, thereby improving the efficiency of training digital human generation models to solve downstream scene problems, and can also be used to test the digital human generation effect in order to optimize the digital human generation effect.
[0066] Based on the above embodiments, the embodiment of the present invention further illustrates the model framework of the digital human action recognition model.
[0067] Figure 6 A schematic diagram of the structure of a digital human action recognition model provided by an embodiment of the present invention.
[0068] In the embodiment of the present invention, the input of the digital human action recognition model is the action sequence and the action query vector , the former represents a complete action sequence, Represents the sequence Frame action; the latter means query vectors that need to be trained, assuming that at most The output of the digital human action recognition model is First Prediction Action , representing an action subsequence The confidence probability vector is .
[0069] like Figure 6As shown, the digital human action recognition model provided by the embodiment of the present invention may include an action mapping layer, a multi-scale feature layer, an action encoding layer (referred to as the encoding layer), a query embedding layer, an action decoding layer (referred to as the decoding layer), and an action output layer (referred to as the output layer). Among them, the action mapping layer is used to map the action frame into an action representation vector, laying the foundation for subsequent action feature extraction and action recognition. The multi-scale feature layer and the action encoding layer are used to extract multi-scale action features from the action representation vector and perform encoding fusion, extract deep features of the action sequence, etc., to provide rich semantic information for action recognition. The query embedding layer contains a series of action query vectors, which are used for the action decoding layer to query the action features output by the action encoding layer. The action decoding layer is used to decode the action representation in combination with the action query vector output by the query embedding layer, and obtain the action decoding representation corresponding to each action query vector. The action output layer is used to predict the category and coordinates of the action based on the output of the action decoding layer, so as to achieve accurate recognition and positioning of the digital human action.
[0070] In an embodiment of the present invention, the step of extracting original action features in S201 includes: converting the posture data of the action frame into a dense vector representation; and calculating the original action features corresponding to the action frame according to the position vector and the dense vector representation corresponding to the action frame.
[0071] By converting posture data into dense vector representation, the model can more accurately capture the potential connection between the action query vector and the action features, as well as the deep semantic information of the action features contained in the posture data, and can more easily perform similarity calculation and cluster analysis in high-dimensional space. There is a sequential relationship between the action frames in the digital human action sequence. In order to make the model understand this sequential relationship, self-attention calculation is performed based on the position vector and dense vector representation corresponding to the action frame to obtain the action sequence representation, so that the model can determine the position of each action frame in the digital human action sequence, thereby better expressing the causal relationship and distance between actions.
[0072] Then the action mapping layer is applied, and the input of the action mapping layer is the action sequence , first convert each action into a preset dimension The dense vector representation of Pose data of action frames The corresponding dense vector representation is , then we have the following formula:
[0073] ;
[0074] in, ,like The dimension is ,but The dimension is , The dimension is , and These are all learnable parameters in the training process of the digital human action recognition model; The dimension is .
[0075] The output of the action mapping layer is Sequence, denoted as ,Right now .
[0076] In the embodiment of the present invention, the multi-scale feature layer and the action coding layer are used to extract multi-scale action features from the action representation vector and perform coding fusion to extract deep features of the action sequence. Then, in S201, the action features of multiple scales are extracted and feature fused according to the posture data of the action frame of the digital human action sample, which may include: performing multiple rounds of feature extraction of different scales on the original action features of the action frame corresponding to the posture data of the action frame, performing feature fusion on the extracted local sequence features, and obtaining a first action feature fusion result; performing self-attention calculation based on the first action feature fusion result to obtain an action feature fusion result. Figure 6 As shown, in the digital human action recognition model provided by the embodiment of the present invention, multiple rounds of feature extraction at different scales are performed through the multi-scale feature layer, and then the multi-scale feature fusion layer ( Figure 6 (not shown) The obtained action features of different scales are feature fused, and then the action encoding layer including the self-attention module is used to perform self-attention calculation according to the obtained first action feature fusion result, so that the digital human action recognition model can learn multi-scale features to adapt to long and short actions while better understanding the deep semantics of the action sequence.
[0077] In some optional implementations of the embodiments of the present invention, performing multiple rounds of feature extraction of different scales on the original action features of the action frame corresponding to the posture data of the action frame may include: in a single round of feature extraction, using a one-dimensional convolution kernel to perform convolution calculation on the input action features, and outputting local sequence features corresponding to multiple adjacent action features. The coordinates of multiple action frames of the digital human action sample are regarded as one-dimensional data, and the features of multiple consecutive action frames are extracted each time in the order of the action frames using one-dimensional convolution to obtain an action representation. By using multiple layers of such one-dimensional convolution, the action representation extracted in the later layers has a larger receptive field, so that action features corresponding to different receptive fields can be extracted, so that the digital human action recognition model can recognize long actions (corresponding to more action frames) and short actions (corresponding to fewer action frames).
[0078] In this embodiment of the present invention, the input of the multi-scale feature layer is the output of the action mapping layer The multi-scale feature layer has Each layer uses a feature extraction method based on one-dimensional convolution and pooling, aiming to capture adjacent features through convolution operations. The local features between the vector representations are used to generate feature vectors.
[0079] Specifically, the multi-scale feature layer The input of the layer is , the output is ,Right now:
[0080] ;
[0081] in, and is the input sequence and the output sequence Length, and are the convolution kernel weights and bias terms, is the window size of the one-dimensional convolution kernel, is the window size of the pooling layer, represents the maximum pooling calculation. In particular, The size is determined by the convolution kernel window size , window step size, padding size, hole size and pooling layer window size Then the length of the characteristic scale satisfies .
[0082] Therefore, low-scale features High-scale features The receptive field of small-scale features is small, and the feature scale range is also small, which is conducive to short action detection and recognition, while the receptive field of high-scale features is large, and the feature scale range is also large, which is conducive to long action detection and recognition.
[0083] The output of the multi-scale feature layer is ,Right now .
[0084] Through one-dimensional convolution, multiple continuous action frames in the action sequence of the digital human action sample are extracted to obtain an action representation. For example, a one-dimensional convolution kernel with a size corresponding to three action frames is used to extract three consecutive action frames into an action representation through the first layer. The receptive field of the one-dimensional convolution of the first layer is three action frames; the receptive field of the one-dimensional convolution of the second layer is based on the first layer, and the receptive field is the three action representations output by the first layer, corresponding to the five action frames of the original action; the receptive field of the third layer corresponds to the seven action frames of the original action, and so on. The receptive field of three action frames can be used for short sequence recognition, and the receptive field of seven or more action frames can be used for long sequence recognition. Figure 6In the multi-scale feature layer shown, the outputs of each feature extraction layer are concat- ed together.
[0085] In some other optional implementations of the embodiments of the present invention, multiple rounds of feature extraction of different scales are performed on the original action features of the action frame corresponding to the posture data of the action frame, and the following may also be included: in a single round of feature extraction, in the input action features, multiple adjacent action features are sequentially connected in series by feature vectors and then multi-layer perceptron calculation is performed to obtain local sequence features corresponding to multiple adjacent action features. That is, the multi-layer perceptron may also be used to construct a feature extraction layer to improve the processing rate of the digital human action sequence.
[0086] Use a multi-layer perceptron to build a multi-scale feature layer. The input of the multi-scale feature layer is the output of the action mapping layer. The multi-scale feature layer can have Each layer uses a feature extraction method based on multi-layer perceptron and pooling, aiming to capture adjacent The local features between the action vector representations are generated into a feature vector. The feature vectors are flattened, that is, all the feature vectors in the window are connected in series to form a dimension of The long vector of is input into the multi-layer perceptron, so that the multi-layer perceptron will perform nonlinear transformation and fusion on these features through its hidden layer to capture the complex relationship between features, thereby receiving the global information of all features in the window, and performing more effective feature fusion. Use the pooling layer (maximum pooling or average pooling) to further reduce the number of feature vectors. This step will help reduce the dimension of the features, reduce the risk of overfitting, and improve the generalization ability of the model.
[0087] Specifically, the multi-scale feature layer The input of the layer is , the output is ,Right now:
[0088] ;
[0089] in, and is the input sequence and the output sequence Length, and are the convolution kernel weights and bias terms, is the window size, is the scale of the vector in the window after flattening, that is , is the window size of the pooling layer, represents the multi-layer perceptron computation, represents the pooling computation. In particular, The size is determined by the window size , window step size, padding size and pooling layer window size Then the length of the characteristic scale satisfies .
[0090] Therefore, low-scale features High-scale features The receptive field of small-scale features is small, and the feature scale range is also small, which is conducive to short action detection and recognition, while the receptive field of high-scale features is large, and the feature scale range is also large, which is conducive to long action detection and recognition.
[0091] The output of the multi-scale feature layer is ,Right now .
[0092] In order to realize multi-scale feature fusion, in some optional implementations of the embodiments of the present invention, the extracted local sequence features are subjected to feature fusion, which may include: performing feature fusion according to the local sequence features corresponding to the feature extraction of multiple rounds of different scales, the category vector corresponding to the digital human action sample, and the position vector corresponding to the action frame, to obtain the action feature fusion result. That is to say, the action features output by the multi-scale feature layer are spliced with the category (CLS) vector corresponding to the digital human action sample as a whole, and this is used as the key technical feature for the model to obtain the overall representation information of the action sequence. By integrating the CLS vector in the multi-scale feature fusion layer, the model can not only capture the features of each scale in the action sequence, but also further integrate these features to form a global understanding of the entire action sequence. The introduction of the CLS vector significantly improves the model's ability to grasp the overall structure and contextual information of the action sequence, thereby achieving a higher level of feature abstraction and more accurate action classification in the process of action detection and recognition. Therefore, the model framework provided by the embodiment of the present invention can grasp the overall characteristics of the action sequence while maintaining sensitivity to local action details. The above process can be expressed by the following formula:
[0093] ;
[0094] ;
[0095] in, , superscript is the first layer of the multi-scale feature layer. Layer, subscript It is Layer feature vectors, The dimension is The CLS vector, is the position vector of the CLS vector, yes The position vector of . The position vector can be set to a fixed value or a learnable model variable, mainly to distinguish and emphasize feature vectors at different scales. In this way, by introducing the position vector for the action representation of each scale, the action scale and its context are accurately defined. In the multi-scale action encoding fusion layer, the position vector is introduced to preserve the time scale information and sequential relationship of the action during the feature fusion process. This design enables the model to more accurately understand and process the long-term and short-term dependencies in the action sequence, thereby improving the accuracy and robustness of action recognition.
[0096] This gives the input of the action encoding layer , whose length is For ease of expression, The element of adds a superscript to indicate the number of action coding layers, that is, The introduction of the action coding layer can refer to the above embodiment. After the action coding layer, the action sequence representation obtained is for .
[0097] The above is a detailed description of the multi-scale feature layer in the embodiment of the present invention.
[0098] like Figure 6 As shown, the action coding layer can use the output of the multi-scale feature layer as input. In the embodiment of the present invention, the action coding layer can be composed of The layer Transformer encoder is composed of a series of stacked layers. A layer Transformer encoder includes two substructures: a self-attention module and a feed-forward neural network module (FFN). The former is the input of the latter. In addition, each substructure also includes a residual connection module and a layer normalization module.
[0099] In the action encoding layer of the digital human action recognition model, the self-attention module can adopt a single-head self-attention module or a multi-head self-attention module (Multi-Head Self-Attention, MHSA).
[0100] If a single-head self-attention module is used, The input of the layer Transformer encoder is , the processing method of each layer of Transformer encoder is the same, so for the sake of convenience, in the description of Transformer encoder, , at this time, let , In the single-head self-attention module, the vector After query value transformation matrix , key value transformation matrix and the embedding value transformation matrix , generate the query value , key value and feature embedding values .
[0101] In the same Transformer encoder, all Share a group , and ,and , and The element values of are learned during the model training process, namely:
[0102] ;
[0103] in, , and Representing dimensions and The query value and key value have the same dimension (both are ), whose dimensions can be the same as the feature embedding values ( ), or they can be different ( ). In the embodiment of the present invention, .
[0104] Secondly, for the query value With key value Perform matching calculation, that is, for any query value and key value The calculated dot product value is obtained Point product value ,in, represent The transposed vector of is a column vector, is a row vector). Perform the above dot product Zoom in, get ), this time .
[0105] At the same time, the embodiment of the present invention defines the mask matrix used in the action representation coding layer as To express and Will attention calculation be performed, that is:
[0106] . (1)
[0107] In the action encoding layer, and All of them perform attention calculation, that is, ( ).
[0108] Then, the softmax operation is used to convert the above dot product value and mask value into a probability value ,Right now:
[0109] ;
[0110] in , express and The logit prediction calculation results are: , represents the mask matrix, express and The logit prediction calculation results.
[0111] From formula (1) and the softmax of multi-head self-attention, we can see that when the Transformer encoder encodes the digital human action sequence, any two actions in the digital human action sequence will perform attention calculations on each other.
[0112] Finally, calculate The corresponding feature embedding ,as follows:
[0113] ,in , It represents the probability value obtained by converting the above point product value and mask value using the softmax operation. Represents the feature embedding value.
[0114] so, is the output of a single-head self-attention module, where each element Corresponding to the input of this module If the self-attention module of the action encoding layer adopts a single-head self-attention module, the output of the single-head self-attention module can be used As input to the feed-forward neural network module.
[0115] If a multi-head self-attention module with H heads is used, the process of the above single-head self-attention module is performed H times in parallel, and each Results After being concatenated together, they are linearly mapped and then output. Specifically, if The output of each head is , then the multi-head attention output Elements for:
[0116] ;
[0117] in, The dimension is ,and Dimensions Keep consistent, that is The size is .
[0118] so, is the output of the multi-head self-attention module, where each element Corresponding to the input of this module . Then the residual connection and layer normalization operations are performed in sequence, as follows:
[0119] The residual connection is: ,in , here is the output of the residual connection.
[0120] The layer normalization layer is a layer normalization layer for any , calculated using the following formula:
[0121] ; (3)
[0122] in, yes The mean of all elements in ,here express The elements. is the mean square error, that is , and is a learnable parameter.
[0123] so, is the output of the multi-head self-attention module, and then input into the feedforward neural network module, where each element Corresponding to the input of this module (Right now ).
[0124] The input of the feedforward neural network module is the output of the self-attention module. Taking the multi-head self-attention module as an example, the input of the feedforward neural network module is the output of the multi-head self-attention module. , after passing through two layers of multi-layer perceptrons, residual connection and layer normalization are performed before output. Specifically:
[0125] ;
[0126] ;
[0127] in, is the output of the residual connection in the feedforward neural network module, is the normalized output of the layer, It is the activation function of the perceptron, which can be selected from sigmoid, tanh, relu and gelu, etc. and The sizes are and , LayerNorm is a layer normalization operation, which is similar to the mechanism of formula (3) and will not be repeated here.
[0128] so, is the output of the forward feedback neural network module and the action representation output by the action encoding layer, where each element Corresponding to the input of this module (Right now ),Right now that is , they are also The input of the Transformer layer encoding layer.
[0129] The above is a detailed description of the action coding layer in the embodiment of the present invention.
[0130] The query embedding layer contains a series of action query vectors, namely, the action query vector , which can be written as , used for querying the action decoding layer on the output of the action encoding layer, where is a preset hyperparameter, assuming that any sequence of digital human actions The sub-action sequence contains no more than Each The length is The output of the query mapping layer is .
[0131] The input of the action decoding layer is the query vector and the output of the action encoding layer . Assume that a set of length , the dimension is The 0 vector ,Right now , that is, each Both Dimensional 0 vector.
[0132] The action decoding layer can have Layers, each layer is modeled using a Transformer decoder. In the embodiment of the present invention, each layer of the Transformer decoder can also be added to the encoder position vector , used to distinguish and strengthen different query positions. The input of the layer is Output of the layer , query vector and the position vector , No. The layer output is First, and Add to On, that is:
[0133] .
[0134] Then and Input into the Transformer decoder and get ,Right now:
[0135] ;
[0136] That is, in the following form:
[0137] ;
[0138] in, , and They are , and No. elements.
[0139] The query vector is added to the input of the Transformer decoder of each layer, so that the input contains the information of the query vector, which makes the query on the output of the action encoding layer more targeted.
[0140] ;
[0141] in, , and They are and No. elements. The output of the action decoding layer is .
[0142] If Figure 6 As shown, in the embodiment of the present invention, the action decoding layer can be composed of A layer of Transformer decoders is composed of a series of stacked modules. A layer of Transformer decoders includes a self-attention module, a cross attention module, and a feed-forward neural network module (FFN). These three modules are stacked in series. The output of each module is the input of the next module, and the output of the feed-forward neural network module is the input of the self-attention module in the next Transformer decoder.
[0143] In the action encoding layer of the digital human action recognition model, the self-attention module can use a single-head self-attention module or a multi-head self-attention module (Multi-Head Self-Attention, MHSA), and its specific structure can refer to the description above. The cross attention module can use a single-head cross attention module or a multi-head cross attention module (Multi-Head Cross Attention, MHCA).
[0144] If a single-head cross attention module is used, assuming that The input of the single-head cross attention layer is and .
[0145] At this time, , ,make , .
[0146] In the single-head criss-cross attention module, the vector After query value transformation matrix Generate query value ,vector After the key value transformation matrix and the embedding value transformation matrix , generate key value and feature embedding values In the same Transformer decoder single-head cross attention module, all Share a group ,all Share a group and ,and , and The element values of are learned during the model training process, namely:
[0147] ;
[0148] in, , , and Representing dimensions and The vector space of .
[0149] The rest of the single-head cross-attention module is similar to the single-head self-attention module. The relationship between the multi-head cross-attention module and the single-head cross-attention module is also similar to the relationship between the single-head self-attention module and the single-head self-attention module, which will not be repeated here.
[0150] The feedforward neural network module in the Transformer decoder has the same structure as the feedforward neural network module in the Transformer encoder and will not be repeated here.
[0151] Therefore, the cross-attention module in the Transformer decoder of the action decoding layer performs action query on the action query vector in the action representation to obtain the first action category (category confidence) of the first predicted action and the first action interval (coordinate) of the first predicted action.
[0152] The above is a detailed introduction to the action decoding layer in the embodiment of the present invention.
[0153] The input of the action output layer is the output of the action decoding layer ,Right now .
[0154] first, After a fully connected layer, we get ,as follows:
[0155] ;
[0156] in, , " " is the matrix-vector multiplication operation, represents the logit prediction calculation result of the t-th output result, The size is , The size is dimension, The size is dimension. The model predicts The class confidence vector of the action.
[0157] Secondly, After several layers of perceptrons, we get a two-dimensional vector , and then after sigmoid activation, we get value, multiplied by the length of the action sequence of the digital human action sample Then round off to get the coordinates of the first action interval , here is an example of a two-layer perceptron:
[0158] ;
[0159] in," " is the matrix-vector multiplication operation, is the activation function of the perceptron, The size is , The size is , The size is , The size is 2.
[0160] ;
[0161] in, Is a vector and a scalar The multiplication operation, It is the first The starting coordinate value and the ending coordinate value of the first predicted action.
[0162] Output of the action output layer ,in , represents the first The first predicted action category confidence vector and the action start and end coordinate values (first action interval). Each action query vector input to the action output layer corresponds to a decoded representation output by the action decoding layer.
[0163] Figure 7 A schematic diagram of the structure of another digital human action recognition model provided by an embodiment of the present invention.
[0164] The above embodiment provides a method of first performing multi-scale feature extraction and then performing multi-scale feature fusion. In addition, in other optional implementations of the embodiments of the present invention, extracting motion features of multiple scales and performing feature fusion based on the posture data of the motion frame of the digital human motion sample may also include: performing multiple rounds of feature extraction of different scales on the original motion features of the action frame corresponding to the posture data of the action frame, performing feature fusion on the extracted local sequence features to obtain a first motion feature fusion result; performing self-attention calculation based on the first motion feature fusion result to obtain a second motion feature fusion result; performing spatial dimension transformation on the second motion feature fusion result, and performing feature fusion with the local sequence features of the same spatial dimension to obtain a motion feature fusion result. Figure 7 As shown, in the digital human action recognition model provided by the embodiment of the present invention, multiple rounds of feature extraction at different scales are first performed through the multi-scale feature layer, and then the action features output by the multi-scale feature layer are fused and encoded by the action coding layer, and then the action features are spatially transformed and fused by the spatial size transformation layer, thereby realizing multi-scale integration of action features.
[0165] In some optional implementations of the embodiments of the present invention, after the second action feature fusion result is spatially transformed, feature fusion is performed with the local sequence features of the same spatial size to obtain the action feature fusion result, which may include: performing at least one upsampling and feature fusion according to the second action feature fusion result to obtain the action feature fusion result; in a single upsampling, the input action feature and the local sequence features of the same spatial size are feature fused to output the fused upsampled action feature; performing feature fusion according to the second action feature fusion result and the fused upsampled action feature output from the last upsampling to obtain the action feature fusion result. In the embodiment of the present invention, the upsampling layer interpolates the feature embedding to expand the number of features, thereby effectively transferring high-level semantic features to low-level features.
[0166] In some other optional implementations of the embodiments of the present invention, after the second action feature fusion result is transformed in spatial size, the local sequence features with the same spatial size are fused to obtain the action feature fusion result, which may also include: performing upsampling and feature fusion at least once according to the second action feature fusion result to obtain the action feature fusion result; in a single upsampling, the input action feature and the local sequence features with the same spatial size are fused to output the fused upsampled action feature; performing downsampling and feature fusion at least once according to the fused upsampled action feature output from the last upsampling; in a single downsampling, the fused upsampled action feature and the action feature with the same spatial size as the downsampled input action feature in the second action feature fusion result are fused to output the fused downsampled action feature; the fused upsampled action feature output from the last upsampling is fused with the fused downsampled action feature to obtain the action feature fusion result. That is to say, not only the semantic features of the high level are effectively transferred to the low level features through the upsampling layer, but also the number of features is reduced through the downsampling layer through downsampling processing, such as using pooling and other methods, so as to enhance the information of the low level features to the high level semantic features. In the bottom-up downsampling process, the feature output of each layer will be retained and finally used as the input of the action decoding layer. Not only the feature information from different levels is retained, but also the model's ability to fully capture action features is enhanced. Through this feature fusion strategy, the digital human action recognition model provided by the embodiment of the present invention can more accurately capture and decode action features, thereby improving the accuracy and robustness of action recognition.
[0167] like Figure 7 As shown, the input of the spatial size transformation layer is the output of the action coding layer And the output of the multi-scale feature layer ,Right now and .
[0168] The upsampling fusion layer can have layer, the input of the first upsampling fusion layer is and , output ; The input of the second upsampling fusion layer is and , output ; and so on, The input of the upsampling fusion layer is and , output ; until The input of the layer is and , the output is .Right now:
[0169] ;
[0170] in, The operation is to concatenate two vectors. represents the upsampling calculation, There are many methods, such as interpolating features and expanding the number of features.
[0171] Similarly, the downsampling fusion layer has layer, the input of the first downsampling fusion layer is and , output ; The input of the second downsampling fusion layer is and , output ; and so on, The input of the upsampling fusion layer is and , output , until The input of the downsampling fusion layer is and , output .Right now:
[0172] ;
[0173] in, The operation is to concatenate two vectors. represents the downsampling calculation, There are many ways to reduce the number of features, such as interpolating features or performing maximum pooling and average pooling.
[0174] In the embodiment of the present invention, the parameter The design allows the model to be flexibly adjusted according to different application scenarios and performance requirements. The model can integrate more feature layers to improve detection accuracy, but this will also increase the number of model parameters, increase computational complexity and processing time. In order to balance the accuracy and efficiency of the model, The value of can be appropriately reduced, which will reduce model parameters and speed up training and reasoning, but may sacrifice a certain degree of detection accuracy.
[0175] In addition, the embodiment of the present invention also provides a flexible feature fusion strategy, allowing the estimated features to be Specific layers are selectively integrated in the model to adapt to different performance requirements and resource constraints. This design allows the model to be flexibly adjusted according to actual application scenarios while maintaining high efficiency to achieve the optimal performance balance.
[0176] At the same time, the present invention also provides a general fusion strategy. The multi-scale feature layer selected from the top to the bottom is not required to correspond one to one with the multi-scale feature layer selected from the bottom to the top. Different numbers and levels of multi-scale features can be selected respectively, which can increase the robustness of the model to multi-scale features.
[0177] All Stacked together, denoted as , which can be used as the output of the multi-scale feature fusion layer.
[0178] In some other optional implementations of the embodiments of the present invention, in addition to Figure 7 In addition to the structure shown, the spatial scale transformation layer can also be set before the action coding layer, then in S201, the action features of multiple scales are extracted and feature fused according to the posture data of the action frame of the digital human action sample, and the feature fusion can also be included: performing multiple rounds of feature extraction of different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; performing spatial scale transformation on the first action feature fusion result, and then performing feature fusion with the local sequence features of the same spatial size to obtain a third action feature fusion result; performing self-attention calculation based on the third action feature fusion result to obtain the action feature fusion result. By performing spatial scale transformation and fusion on the basis of multi-scale feature sampling, and then performing self-attention calculation, the deep semantic information of the action sequence is extracted on the basis of realizing multi-scale integration of action features.
[0179] In some optional implementations of the embodiments of the present invention, after the first action feature fusion result is transformed in spatial size, feature fusion is performed with the local sequence features with the same spatial size to obtain a third action feature fusion result, which may include: performing at least one upsampling and feature fusion according to the first action feature fusion result to obtain the third action feature fusion result; in a single upsampling, feature fusion is performed on the input action feature and the local sequence features with the same spatial size to output the fused upsampled action feature; feature fusion is performed according to the first action feature fusion result and the fused upsampled action feature output from the last upsampling to obtain the action feature fusion result.
[0180] In some other optional implementations of the embodiments of the present invention, after the first action feature fusion result is transformed in spatial size, feature fusion is performed on the local sequence features with the same spatial size to obtain a third action feature fusion result, which may also include: performing at least one upsampling and feature fusion according to the second action feature fusion result to obtain the action feature fusion result; in a single upsampling, the input action feature and the local sequence features with the same spatial size are feature fused to output the fused upsampled action feature; according to the fused upsampled action feature output from the last upsampling, at least one downsampling and feature fusion are performed; in a single downsampling, the fused upsampled action feature and the action feature with the same spatial size as the downsampled input action feature in the second action feature fusion result are feature fused to output the fused downsampled action feature; the fused upsampled action feature output from the last upsampling and the fused downsampled action feature are feature fused to obtain the action feature fusion result.
[0181] The specific implementation of the above-mentioned spatial size conversion (upsampling and downsampling) can refer to the introduction of the above-mentioned embodiment.
[0182] Figure 7 The action mapping layer, multi-scale feature layer and multi-scale feature fusion layer, action encoding layer, query embedding layer, action decoding layer and action output layer shown in FIG. Figure 6 Introduction.
[0183] In the above embodiment, the digital human action recognition model is used to output the posture data of the action frame of the digital human action sample. The first predicted action, assuming that the digital human action samples have If there are a number of labeled actions, the first predicted action needs to be matched with the labeled action in order to calculate the model loss value.
[0184] In some optional implementations of the embodiments of the present invention, performing matching calculation on the first predicted action and the labeled action in S202 to obtain a first matching relationship may include: calculating the category confidence loss value of the first action category and the labeled action category in the labeled action; calculating the action interval loss value of the first action interval and the labeled action interval of the labeled action; and calculating the first matching relationship between the first predicted action and the labeled action that minimizes the total loss value based on the category confidence loss value and the action interval loss value. That is, the loss values are calculated from two perspectives, namely, the action category and the action interval are combined to determine the best matching relationship between the first predicted action and the labeled action.
[0185] The embodiment of the present invention provides a prediction action screening module to achieve matching between the first prediction action and the labeled action. The input of the prediction action screening module is the first prediction action. , and annotation actions , The goal of the predictive action filtering module is to The first predicted actions are selected to match the real labeled actions one by one. Actions, final output Candidate actions and the real labeled Each action corresponds to another one.
[0186] In some optional implementations of the embodiments of the present invention, calculating the category confidence loss value of the first action category and the labeled action category in the labeled action may include: calculating the negative value of the product of the logarithm of the confidence value of the first action category and the labeled action to obtain the category confidence loss value.
[0187] Specifically, define The first predicted action and the The matching weight of the labeled action is , which combines the category confidence loss and the coordinate intersection loss. Category confidence loss It can be calculated by the following formula:
[0188] ;
[0189] The goal of category confidence loss is to make the predicted category consistent with the true category.
[0190] For the action interval loss value, in some optional implementations of the embodiments of the present invention, calculating the action interval loss value of the first action interval and the marked action interval of the marked action may include: calculating the first coordinate loss value of the starting action frame of the first action interval and the starting action frame of the marked action interval; calculating the second coordinate loss value of the ending action frame of the first action interval and the ending action frame of the marked action interval; and calculating the action interval loss value according to the first coordinate loss value and the second coordinate loss value. Specifically, the action interval loss value It can be calculated by the following formula:
[0191] ;
[0192] in, The goal of the action interval loss is to make the predicted coordinates of the action consistent with the actual coordinates.
[0193] In some other optional implementations of the embodiments of the present invention, calculating the action interval loss value of the first action interval and the marked action interval of the marked action may also include: calculating the first coordinate length of the intersection of the coordinates of the first action interval and the second action interval; calculating the second coordinate length of the coordinate union of the first action interval and the second action interval; calculating the negative value of the natural logarithm of the ratio of the first coordinate length to the second coordinate length to obtain the action interval loss value. That is, the action interval loss value It can also be calculated by the following formula:
[0194] ;
[0195] in, represents the natural logarithm, and the length of the intersection of the first action interval output by the model and the marked action interval is , the length of the union of the two is By defining the action coordinate matching weight, instead of using the Euclidean distance between the predicted coordinates and the true coordinates, the intersection-over-union loss between the predicted action interval and the true action interval is used. This intersection-over-union loss effectively solves the gradient instability problem existing in the traditional detection method based on Euclidean distance loss, and at the same time avoids the inconsistency between model learning and model evaluation, thereby improving the accuracy and robustness of action recognition.
[0196] In some further optional implementations of the embodiments of the present invention, calculating the action interval loss value of the first action interval and the marked action interval of the marked action may also include: calculating the minimum circumscribed length of the first action interval and the marked action interval; calculating the first coordinate length of the intersection of the coordinates of the first action interval and the second action interval; calculating the second coordinate length of the coordinate union of the first action interval and the second action interval; and calculating the action interval loss value based on the minimum circumscribed length, the first coordinate length, and the second coordinate length. Action interval loss value It can also be calculated by the following formula:
[0197]
[0198]
[0199] ;
[0200] Among them, the minimum circumscribed length between the first action interval output by the model and the marked action interval is The length of the intersection of the two is , the length of the union of the two is The defect of the intersection-over-union loss is that when the predicted coordinate interval and the true coordinate interval do not intersect, it cannot reflect the distance between the two coordinate intervals, and the loss function is not differentiable at this time. That is, the intersection-over-union loss cannot optimize the situation where the two coordinate intervals do not intersect, so the circumscribed interval loss uses a minimum circumscribed interval to frame the first action interval and the labeled action interval, which can measure the distance between the two. At this time, the loss function is differentiable, which can better perform model training.
[0201] Using the category confidence loss and action interval loss introduced above, and calculating the first matching relationship between the first predicted action and the labeled action that minimizes the total loss value according to the category confidence loss value and the action interval loss value, it can include: screening out a second number of first predicted actions from the first number of first predicted actions output by the digital human action recognition model, calculating the matching weight value between the first predicted action and the labeled action according to the category confidence loss value and the action interval loss value, and calculating the first matching relationship that minimizes the second number of matching weight values; wherein the second number is the total number of labeled actions corresponding to the digital human action samples.
[0202] The first The first predicted action and the The matching weight of the labeled action Can be:
[0203] ;
[0204] in, and It is a hyperparameter, which can be set as a constant or a model learnable parameter according to the data set.
[0205] Thus, the weight matrix shown in Table 1 can be obtained. Table 1 is a weight matrix of the first predicted action and the labeled action, in which each element represents the matching degree between the first predicted action and the labeled action.
[0206] Table 1
[0207]
[0208] The goal of establishing the first matching relationship is to find an optimal match so that Select the first predicted action actions, and The total loss of matching the labeled actions one by one is the smallest. Assume that Select The total number of selections is a set , then the matching target selects the best matching method , as shown below:
[0209] ;
[0210] in, , Indicates the minimum value calculation. Indicates The matching weight corresponding to the first matching relationship.
[0211] Solve the above formula to get the best match , for example, the Hungarian algorithm can be used.
[0212] The output of the predicted action filtering module is the first matching relationship ,Right now , without loss of generality (because the model outputs The first predicted actions are independent of each other and have no order), we can assume that The first matching relationship is output .
[0213] Calculating the model loss value according to the first matching relationship may include: according to the first matching relationship, performing weighted calculation on the category confidence loss value and the action interval loss value between the matched first predicted action and the labeled action to obtain the predicted loss value; calculating the model loss value according to the predicted loss value corresponding to each labeled action; wherein the weight of the category confidence loss value and the weight of the action interval loss value are coefficients adaptively updated in the iterative training of the digital human action recognition model.
[0214] In an embodiment of the present invention, the model loss value can be calculated by a loss function module. The loss function module receives the output of the prediction action screening module. and the corresponding annotation actions .
[0215] If the first action interval loss calculation method provided above is used, the model loss value is calculated using the category confidence loss and action interval loss. , can be expressed by the following formula:
[0216] ;
[0217] in, and are model hyperparameters, all of which are positive numbers. They can be set as constants according to the data set, or they can be set as model learnable parameters and automatically learned during model training. Indicates that the first matching relationship The annotation action corresponds to The confidence value of the first action category of the first predicted action, Indicates that the first matching relationship The annotation action corresponds to a starting action frame of a first action interval of a first predicted action, Indicates that the first matching relationship The annotation action corresponds to The terminal action frame of the first action interval of the first predicted action.
[0218] If the second action interval loss calculation method provided above is used, the model loss value is calculated using the category confidence loss and action interval loss. , can be expressed by the following formula:
[0219] ;
[0220] in, and are model hyperparameters, all of which are positive numbers. They can be set as constants according to the data set, or they can be set as model learnable parameters and automatically learned during model training. Indicates that the first matching relationship The annotation action corresponds to The confidence value of the first action category of the first predicted action, Indicates that the first matching relationship The annotation action corresponds to a starting action frame of a first action interval of a first predicted action, Indicates that the first matching relationship The annotation action corresponds to The terminal action frame of the first action interval of the first predicted action.
[0221] If the third action interval loss calculation method provided above is used, the model loss value is calculated using the category confidence loss and action interval loss. , can be expressed by the following formula:
[0222] ;
[0223]
[0224]
[0225] ;
[0226] in, and are model hyperparameters, all of which are positive numbers. They can be set as constants according to the data set, or they can be set as model learnable parameters and automatically learned during model training. Indicates that the first matching relationship The annotation action corresponds to The confidence value of the first action category of the first predicted action, Indicates that the first matching relationship The annotation action corresponds to a starting action frame of a first action interval of a first predicted action, Indicates that the first matching relationship The annotation action corresponds to The terminal action frame of the first action interval of the first predicted action.
[0227] Therefore, combined with the network structure of the digital human action recognition model provided by the embodiment of the present invention, as well as the predicted action screening module and the loss function module, the model can be trained on digital human action samples, and the training can be stopped when the training end condition is reached, and then the model effect can be verified on the test set or the verification set.
[0228] In some other optional implementations of the embodiments of the present invention, S203 performs a matching calculation based on the first predicted action and the labeled action of the digital human action sample to obtain a first matching relationship between the first predicted action and the labeled action, and may also include: determining the matching relationship between the first action category and the labeled action category based on the size of the confidence value of the first action category output by the digital human action recognition model, and then determining the first matching relationship.
[0229] That is to say, the labeled action category can be matched first according to the first action category. For example, the first predicted action with the largest category confidence in the first action category that is the same as the labeled action category of the labeled action can be used as the first predicted action that matches the labeled action, thereby quickly determining the best matching relationship between the first predicted action and the labeled action.
[0230] Determining the matching relationship between the first action category and the labeled action category according to the size of the confidence value of the first action category output by the digital human action recognition model, and then determining the first matching relationship may include: selecting candidate actions with the same labeled action category from the first predicted action according to the size of the confidence value of the first action category, and determining the first matching relationship.
[0231] On the basis of this first matching relationship, calculating the model loss value according to the first matching relationship in S204 may include: if the number of candidate actions is less than the number of labeled actions, then calculating the model loss value according to the action interval loss values corresponding to each first matching relationship and the unmatched labeled actions; if the number of candidate actions is equal to the number of labeled actions, then calculating the model loss value according to the action interval loss values corresponding to each first matching relationship. In actual applications, if the number of candidate actions is less than the number of labeled actions, a penalty value may be set according to the number of unmatched labeled actions, and the model loss value may be calculated in combination with the action interval loss values corresponding to the first matching relationship. The method for calculating the action interval loss value may refer to the three methods described in the above embodiments.
[0232] Through the method for determining the first matching relationship and the method for calculating the model loss value provided in the embodiment of the present invention, the first matching relationship can be quickly determined and the model loss can be calculated.
[0233] Based on the model framework of the digital human action recognition model provided in the above embodiment, the embodiment of the present invention also provides another action output layer.
[0234] In an embodiment of the present invention, the first action category of the first predicted action and the first action interval of the first predicted action are calculated according to the action feature fusion result and the action query vector in S201, which may include: using the digital human action recognition model to encode the posture data of the action frame of the digital human action sample to obtain the action sequence representation, and performing multi-layer decoding calculation according to the action sequence representation and the action query vector. In S202, the first predicted action and the labeled action are matched and calculated to obtain the first matching relationship, which may include: obtaining the first predicted action output by the last decoding layer, matching the first predicted action with the labeled action, and obtaining the first matching relationship. In S203, the model loss value is calculated according to the first matching relationship to optimize the loss of the digital human action recognition model to obtain the trained digital human action recognition model, which may include: according to the first matching relationship, the output results of multiple decoding layers are calculated compared with the loss value of the labeled action, the model loss value is calculated according to the loss values corresponding to the multiple decoding layers, and the digital human action recognition model is optimized by using the model loss value to obtain the trained digital human action recognition model.
[0235] That is to say, in addition to using multiple layers of action decoding layers to output data layer by layer, the embodiment of the present invention also uses a multi-head action output layer, and sets the outputs of multiple action decoding layers to be directly connected to an output head in the output layer. This means that the outputs of multiple layers of action decoding layers can be involved in the loss calculation of action recognition, thereby ensuring that the model can contribute to the digital human action recognition task at multiple levels.
[0236] The digital human action recognition model provided by the embodiment of the present invention can allow the number of action decoding layers to be flexibly configured to adapt to different application scenarios, data set characteristics and real-time requirements. In the model training stage, all action decoding layers can participate in the calculation of the loss function to ensure that the model can be fully learned and optimized. In the prediction stage, users can choose to use the results of some action output layers instead of all of them based on their own trade-offs between accuracy and reasoning speed.
[0237] The model loss value is calculated based on multiple loss values, which may include: performing weighted calculation based on multiple loss values to obtain the model loss value. In practical applications, the weights corresponding to each decoding layer may be set to be the same, or different weights may be set. In some optional implementations of the embodiments of the present invention, the weight of the action decoding layer closer to the action output layer may be set to be smaller, so that the front action decoding layer can also be better optimized, so that when performing lightweight calculations in the model inference stage, any action output layer can be retained to obtain better model accuracy.
[0238] Thus, in S204, using the trained digital human action recognition model to calculate the second action category and the second action interval of the second predicted action according to the posture data of the digital human action sequence to be recognized can include: using the trained digital human action recognition model with a partial decoding layer retained to perform the action recognition task of the digital human action sequence to be recognized. Based on the model framework of the multi-action output layer provided by the embodiment of the present invention, when performing the digital human action recognition task, the digital human action recognition model can be lightweight processed as needed, and only a partial action decoding layer is retained to reduce model parameters and model calculation amount, while ensuring the model accuracy to the greatest extent.
[0239] It is understandable that there is a trade-off between the choice of action output layers and the model's reasoning speed and prediction accuracy. If more action output layers are selected, the accuracy of the prediction results can be improved because the model can synthesize information from multiple levels to make decisions. However, this also means an increased computational burden, which reduces the reasoning speed. If fewer output layers are selected, the amount of computation can be reduced and the reasoning speed can be increased, but a certain degree of prediction accuracy may be sacrificed because the model loses the opportunity to synthesize information from more levels. Through this flexible design, users can make the best choice between reasoning speed and prediction accuracy according to specific application requirements to optimize model performance.
[0240] In an embodiment of the present invention, performing an action recognition task of a digital human action sequence to be recognized using a trained digital human action recognition model that retains some decoding layers can include: receiving model accuracy requirement parameters; determining the number of retained decoding layers based on the model accuracy requirement parameters, so as to perform the action recognition task using the trained digital human action recognition model that retains some decoding layers.
[0241] Among them, determining the number of retained decoding layers according to the model accuracy requirement parameters can include: using digital human action test samples to test the model accuracy test results of the trained digital human action recognition model that retains different layers of decoding layers; matching the model accuracy test results according to the model accuracy requirement parameters to determine the number of retained decoding layers.
[0242] In addition, according to the digital human action recognition model provided by the above embodiment of the present invention, multiple action coding layers can be set in the model training stage, and the digital human action recognition model with some coding layers and some decoding layers retained can be used to perform action recognition tasks in the model reasoning stage. For example, if a multi-scale feature layer and / or a spatial dimension transformation layer is set, according to actual measurements, a layer of action coding layer can be set in the training stage, which can also ensure good model accuracy.
[0243] The above embodiment introduces a method for digital human motion recognition. In addition, an embodiment of the present invention also provides a method for training a digital human motion recognition model, which may include: using a digital human motion recognition model to extract motion features of multiple scales based on the posture data of the motion frame of the digital human motion sample and perform feature fusion, and calculate the first action category of the first predicted action and the first action interval of the first predicted action based on the action feature fusion result and the action query vector; perform matching calculation on the first predicted action and the labeled action to obtain a first matching relationship; calculate the model loss value based on the first matching relationship to perform loss optimization on the digital human motion recognition model to obtain a trained digital human motion recognition model; wherein, motion features of different scales correspond to different numbers of action frames.
[0244] The digital human action recognition model training method provided in the embodiment of the present invention can refer to the introduction of the digital human action recognition method mentioned above, and will not be repeated here.
[0245] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0246] Figure 8 A schematic diagram of the structure of a digital human motion recognition device provided by an embodiment of the present invention.
[0247] like Figure 8 As shown, the embodiment of the present application also provides a digital human action recognition device, including:
[0248] The model training module 801 is used to extract motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample using the digital human motion recognition model and perform feature fusion, and calculate the first motion category of the first predicted motion and the first motion interval of the first predicted motion according to the motion feature fusion result and the motion query vector;
[0249] The predicted action screening module 802 is used to perform matching calculation on the first predicted action and the marked action to obtain a first matching relationship;
[0250] A loss function module 803 is used to calculate the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model to obtain a trained digital human action recognition model;
[0251] The recognition module 804 is used to calculate the second action category of the second predicted action and the second action interval of the second predicted action according to the posture data of the action sequence of the digital human to be recognized by using the trained digital human action recognition model;
[0252] Among them, action features of different scales correspond to different numbers of action frames.
[0253] In an embodiment of the present invention, the model training module 801 extracts motion features of multiple scales based on the posture data of the motion frame of the digital human motion sample and performs feature fusion, which may include: performing multiple rounds of feature extraction of different scales on the original motion features of the action frame corresponding to the posture data of the action frame, performing feature fusion on the extracted local sequence features to obtain a first motion feature fusion result; performing self-attention calculation based on the first motion feature fusion result to obtain a motion feature fusion result.
[0254] In an embodiment of the present invention, the model training module 801 extracts motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample and performs feature fusion, and may also include: performing multiple rounds of feature extraction of different scales on the original motion features of the motion frame corresponding to the posture data of the motion frame, and performing feature fusion on the extracted local sequence features to obtain a first motion feature fusion result; performing self-attention calculation according to the first motion feature fusion result to obtain a second motion feature fusion result; performing spatial dimension transformation on the second motion feature fusion result, and then performing feature fusion with the local sequence features with the same spatial dimension to obtain a motion feature fusion result.
[0255] In an embodiment of the present invention, the model training module 801 performs spatial dimension transformation on the second motion feature fusion result, and then performs feature fusion with the local sequence features of the same spatial dimension to obtain the motion feature fusion result, which may include: performing at least one upsampling and feature fusion according to the second motion feature fusion result to obtain the motion feature fusion result; in a single upsampling, performing feature fusion on the input motion feature and the local sequence features of the same spatial dimension, and outputting the fused upsampled motion feature; performing feature fusion on the second motion feature fusion result and the fused upsampled motion feature outputted from the last upsampling to obtain the motion feature fusion result.
[0256] In an embodiment of the present invention, the model training module 801 performs spatial dimension transformation on the second motion feature fusion result, and then performs feature fusion with the local sequence features with the same spatial dimension to obtain the motion feature fusion result. It may also include: performing upsampling and feature fusion at least once according to the second motion feature fusion result to obtain the motion feature fusion result; in a single upsampling, the input motion feature and the local sequence features with the same spatial dimension are subjected to feature fusion, and the fused upsampled motion feature is output; according to the fused upsampled motion feature outputted from the last upsampling, at least one downsampling and feature fusion is performed; in a single downsampling, the fused upsampled motion feature and the motion feature with the same spatial dimension as the downsampled input motion feature in the second motion feature fusion result are subjected to feature fusion, and the fused downsampled motion feature is output; the fused upsampled motion feature outputted from the last upsampling and the fused downsampled motion feature are subjected to feature fusion to obtain the motion feature fusion result.
[0257] In an embodiment of the present invention, the model training module 801 extracts motion features of multiple scales and performs feature fusion based on the posture data of the motion frame of the digital human motion sample, and may also include: performing multiple rounds of feature extraction of different scales on the original motion features of the motion frame corresponding to the posture data of the motion frame, and performing feature fusion on the extracted local sequence features to obtain a first motion feature fusion result; performing spatial dimension transformation on the first motion feature fusion result, and then performing feature fusion with the local sequence features of the same spatial dimension to obtain a third motion feature fusion result; performing self-attention calculation based on the third motion feature fusion result to obtain a motion feature fusion result.
[0258] In an embodiment of the present invention, the model training module 801 performs multiple rounds of feature extraction of different scales on the original action features of the action frame corresponding to the posture data of the action frame, which may include: in a single round of feature extraction, using a one-dimensional convolution kernel to perform convolution calculation on the input action features, and outputting local sequence features corresponding to multiple adjacent action features.
[0259] In an embodiment of the present invention, the model training module 801 performs multiple rounds of feature extraction of different scales on the original action features of the action frame corresponding to the posture data of the action frame, and may also include: in a single round of feature extraction, in the input action features, the feature vectors of multiple adjacent action features are sequentially connected in series and then multi-layer perceptron calculation is performed to obtain local sequence features corresponding to the multiple adjacent action features.
[0260] In an embodiment of the present invention, the model training module 801 performs feature fusion on the extracted local sequence features, which may include: performing feature fusion on the local sequence features corresponding to multiple rounds of feature extraction at different scales, the category vectors corresponding to the digital human action samples, and the position vectors corresponding to the action frames to obtain action feature fusion results.
[0261] In an embodiment of the present invention, the step of extracting original action features in the model training module 801 may include: converting the posture data of the action frame into a dense vector representation; and calculating the original action features corresponding to the action frame based on the position vector and the dense vector representation corresponding to the action frame.
[0262] In an embodiment of the present invention, the predicted action screening module 802 performs matching calculations on the first predicted action and the labeled action to obtain a first matching relationship, which may include: calculating the category confidence loss value of the first action category and the labeled action category in the labeled action; calculating the action interval loss value of the first action interval and the labeled action interval of the labeled action; and calculating the first matching relationship between the first predicted action and the labeled action that minimizes the total loss value based on the category confidence loss value and the action interval loss value.
[0263] For the description of the features in the embodiment corresponding to the digital human motion recognition device, reference can be made to the relevant description of the embodiment corresponding to the digital human motion recognition method, which will not be repeated here.
[0264] The embodiment of the present application also provides a digital human action recognition model training device, which may include:
[0265] The model training module 801 is used to extract motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample using the digital human motion recognition model and perform feature fusion, and calculate the first motion category of the first predicted motion and the first motion interval of the first predicted motion according to the motion feature fusion result and the motion query vector;
[0266] The predicted action screening module 802 is used to perform matching calculation on the first predicted action and the marked action to obtain a first matching relationship;
[0267] A loss function module 803 is used to calculate the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model to obtain a trained digital human action recognition model;
[0268] Among them, action features of different scales correspond to different numbers of action frames.
[0269] For the description of the features in the embodiment corresponding to the digital human action recognition model training device, reference can be made to the relevant description of the embodiment corresponding to the digital human action recognition method, which will not be repeated here.
[0270] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any one of the above-mentioned digital human motion recognition methods or the above-mentioned digital human motion recognition model training methods.
[0271] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any one of the above-mentioned digital human action recognition methods or the above-mentioned digital human action recognition model training methods when running.
[0272] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0273] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned digital human action recognition methods or the above-mentioned digital human action recognition model training methods.
[0274] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned digital human action recognition method embodiments are implemented.
[0275] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0276] The above is a detailed introduction to a digital human motion recognition method, device, medium and product provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A method for digital human action recognition, characterized in that: include: Extracting motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample using the digital human motion recognition model and performing feature fusion, and calculating a first motion category of a first predicted motion and a first motion interval of the first predicted motion according to the motion feature fusion result and the motion query vector; Performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship; Calculate the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model, and obtain the trained digital human action recognition model; Using the trained digital human action recognition model, according to the posture data of the digital human action sequence to be recognized, a second action category of a second predicted action and a second action interval of the second predicted action are calculated; Wherein, action features of different scales correspond to different numbers of said action frames; The first action interval is an interval between an action start frame and an action end frame of the first predicted action; the second action interval is an interval between an action start frame and an action end frame of the second predicted action; A digital human action recognition model is used to extract action features of multiple scales according to the posture data of the action frame of the digital human action sample and perform feature fusion, including: using a multi-scale feature layer to perform multi-layer feature extraction starting from the original action feature corresponding to the action frame, and the input of the current feature extraction layer from the second layer is the action feature output by the previous layer, the feature extraction layer is used to extract action features corresponding to multiple consecutive action frames in the order of the action frames, and the more action features extracted by the later feature extraction layers correspond to the more action frames, and the action features output by each feature extraction layer are feature fused.
2. The method for digital human motion recognition according to claim 1, characterized in that: According to the posture data of the action frame of the digital human action sample, action features of multiple scales are extracted and feature fusion is performed, including: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; A self-attention calculation is performed according to the first action feature fusion result to obtain the action feature fusion result.
3. The method for digital human motion recognition according to claim 1, characterized in that: According to the posture data of the action frame of the digital human action sample, action features of multiple scales are extracted and feature fusion is performed, including: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; Performing self-attention calculation according to the first action feature fusion result to obtain a second action feature fusion result; After the second action feature fusion result is transformed in spatial size, feature fusion is performed with the local sequence features with the same spatial size to obtain the action feature fusion result.
4. The method for digital human motion recognition according to claim 3, characterized in that: After performing spatial size transformation on the second action feature fusion result, performing feature fusion with the local sequence feature having the same spatial size to obtain the action feature fusion result includes: Performing upsampling and feature fusion at least once according to the second action feature fusion result to obtain the action feature fusion result; In a single upsampling, the input action feature and the local sequence feature with the same spatial size are subjected to feature fusion, and the fused upsampled action feature is output; Feature fusion is performed based on the second action feature fusion result and the fused up-sampled action feature outputted from the last up-sampling to obtain the action feature fusion result.
5. The method for digital human motion recognition according to claim 3, characterized in that: After performing spatial size transformation on the second action feature fusion result, performing feature fusion with the local sequence feature having the same spatial size to obtain the action feature fusion result includes: Performing upsampling and feature fusion at least once according to the second action feature fusion result to obtain the action feature fusion result; In a single upsampling, the input action feature and the local sequence feature with the same spatial size are subjected to feature fusion, and the fused upsampled action feature is output; Perform at least one downsampling and feature fusion according to the fused upsampled action features outputted from the last upsampling; In a single downsampling, the fused upsampled motion features and the motion features in the second motion feature fusion result having the same spatial size as the downsampled input motion features are subjected to feature fusion, and the fused downsampled motion features are output; The up-sampled motion features after fusion and the down-sampled motion features after fusion outputted from the last up-sampling are subjected to feature fusion to obtain the motion feature fusion result.
6. The method for digital human motion recognition according to claim 1, characterized in that: According to the posture data of the action frame of the digital human action sample, action features of multiple scales are extracted and feature fusion is performed, including: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, and performing feature fusion on the extracted local sequence features to obtain a first action feature fusion result; After performing spatial size transformation on the first action feature fusion result, feature fusion is performed with the local sequence features having the same spatial size to obtain a third action feature fusion result; A self-attention calculation is performed according to the third action feature fusion result to obtain the action feature fusion result.
7. The method for digital human motion recognition according to any one of claims 2 to 6, characterized in that: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, including: In a single round of feature extraction, a one-dimensional convolution kernel is used to perform convolution calculation on the input action features, and the local sequence features corresponding to multiple adjacent action features are output.
8. The method for digital human motion recognition according to any one of claims 2 to 6, characterized in that: Performing multiple rounds of feature extraction at different scales on the original action features of the action frame corresponding to the posture data of the action frame, including: In a single round of feature extraction, in the input action features, multiple adjacent action features are sequentially connected in series with their feature vectors and then calculated using a multi-layer perceptron to obtain local sequence features corresponding to the multiple adjacent action features.
9. The method for digital human motion recognition according to any one of claims 2 to 6, characterized in that: The extracted local sequence features are subjected to feature fusion, including: Feature fusion is performed based on the local sequence features corresponding to multiple rounds of feature extraction at different scales, the category vectors corresponding to the digital human action samples, and the position vectors corresponding to the action frames to obtain the action feature fusion result.
10. The method for digital human motion recognition according to any one of claims 2 to 6, characterized in that: The step of extracting the original action features comprises: Converting the posture data of the action frame into a dense vector representation; The original action feature corresponding to the action frame is obtained by calculation based on the position vector corresponding to the action frame and the dense vector representation.
11. The method for digital human motion recognition according to claim 1, characterized in that: Performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship includes: Calculating category confidence loss values of the first action category and the labeled action category in the labeled actions; Calculating the action interval loss value of the first action interval and the marked action interval of the marked action; The first matching relationship between the first predicted action and the labeled action that minimizes the total loss value is calculated according to the category confidence loss value and the action interval loss value.
12. A digital human action recognition model training method, characterized in that: include: Extracting motion features of multiple scales according to the posture data of the motion frame of the digital human motion sample using the digital human motion recognition model and performing feature fusion, and calculating a first motion category of a first predicted motion and a first motion interval of the first predicted motion according to the motion feature fusion result and the motion query vector; Performing a matching calculation on the first predicted action and the marked action to obtain a first matching relationship; Calculate the model loss value according to the first matching relationship to perform loss optimization on the digital human action recognition model, and obtain the trained digital human action recognition model; Wherein, action features of different scales correspond to different numbers of said action frames; The first action interval is an interval between an action start frame and an action end frame of the first predicted action; A digital human action recognition model is used to extract action features of multiple scales according to the posture data of the action frame of the digital human action sample and perform feature fusion, including: using a multi-scale feature layer to perform multi-layer feature extraction starting from the original action feature corresponding to the action frame, and the input of the current feature extraction layer from the second layer is the action feature output by the previous layer, the feature extraction layer is used to extract action features corresponding to multiple consecutive action frames in the order of the action frames, and the more action features extracted by the later feature extraction layers correspond to the more action frames, and the action features output by each feature extraction layer are feature fused.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the digital human action recognition method according to any one of claims 1 to 11 or the steps of the digital human action recognition model training method according to claim 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the digital human action recognition method according to any one of claims 1 to 11 or the steps of the digital human action recognition model training method according to claim 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the digital human action recognition method according to any one of claims 1 to 11 or the steps of the digital human action recognition model training method according to claim 12 are implemented.
Citation Information
Patent Citations
Human body posture image intelligent identification method and system
CN115527269A
Fine-grained action recognition method based on cross-modal knowledge alignment
CN118196888A