A multi-condition human motion generation method and system based on scenarios and text

By using a bounding box to represent scene data, a scene encoder and text encoder are constructed, and combined with a denoising diffusion model, the problem of lack of physical rationality and environmental adaptability of action generation in the prior art is solved, and the stability and accuracy of action generation are achieved, which is especially suitable for real-time interaction in areas such as virtual reality and game development.

CN119741408BActive Publication Date: 2025-07-08SUN YAT SEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510246686.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-08
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

When generating actions that interact with complex scenes, the prior art lacks physical rationality and environmental adaptability, and it is difficult to effectively integrate scenes and text information, resulting in the generated actions lacking rationality and accuracy.

Method used

By using a bounding box to represent scene data, a scene encoder and text encoder are constructed, combined with the denoising diffusion model, the preprocessed ternary data set is used for feature extraction and noise addition processing, and the scene encoder and denoising device are trained jointly to generate actions that conform to the text description and interact reasonably with the scene.

Benefits of technology

It realizes the stability and accuracy of action generation, can quickly and efficiently capture scene information, and the generated actions meet the text semantic requirements and interact reasonably with specific scenes, improving the generality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741408B_ABST
    Figure CN119741408B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-condition human motion generation method and system based on scenarios and text. The steps of this method include: First, obtain a ternary dataset and perform preprocessing. Then, extract features from the scenario data and text data to obtain scenario features and text features. After that, perform noise addition processing on the motion data in the preprocessed ternary dataset, and at the same time, input the combined scenario features and text features into a denoiser to obtain denoised motion data. Use the motion data and the denoised motion data to jointly train a scenario encoder and a denoiser. Finally, the user uses the trained scenario encoder and denoiser to generate new motion data. The present invention uses a bounding box to represent the scenario, which can quickly and efficiently capture scenario information. At the same time, by fusing scenario data and text descriptions, the generated motions not only meet the semantic requirements of the text descriptions but also can reasonably interact with specific scenarios, effectively improving the stability and accuracy of the generated motions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer graphics, 3D digital humans, and motion generation technology. More specifically, it relates to a multi-condition human motion generation method and system based on scenarios and text. Background Art

[0002] Human motion generation models are widely used in fields such as virtual reality, animation production, game development, and human-computer interaction. Such models can automatically generate high-quality motions that meet specific requirements according to various input conditions such as speech, text, and scenarios, significantly improving production efficiency, enhancing the user experience, and promoting the research progress of technologies such as action recognition and pose estimation. Especially in the field of automated motion generation, how to accurately combine semantic information and environmental constraints has always been the focus of research.

[0003] Existing generation models usually require input of motion data, text descriptions, and scene information. Motion data is mostly represented by sequences of human key points or morphological parameters (such as SMPL / SMPL-X models); text input is usually a sequence of words or texts, describing the type, details, and context of the motion; while scene data is mostly in the form of point cloud data or bounding boxes, providing physical and spatial constraints. One of the key challenges of existing technologies is how to process and fuse these heterogeneous data to generate natural motions that conform to semantics and scenarios.

[0004] Current research mostly relies on large-scale motion-text-scene datasets, and the scale and quality of these datasets directly affect the generalization ability of the generation model. Generalization ability refers to the ability of the model to still generate reasonable motions when facing texts and scenarios that are not seen or rarely seen in the dataset. Due to the limitations of the dataset, many generation models are difficult to effectively process complex and dynamically changing scene information. Especially in the actual application of motion generation, the diversity and complexity of scene conditions are often not fully considered.

[0005] In addition, although some methods attempt to enhance the robustness of the generation model by introducing technologies such as self-supervised learning and generative adversarial networks, due to the high correlation between scenes and motions, these methods usually face difficulties in fusing scene and text information. Therefore, how to accurately integrate multi-modal information and improve the quality and adaptability of motion generation remains an important research direction in this field.

[0006] Currently, a patent application with the publication number "CN117953113A" discloses a human motion generation method that can interact with 3D scene targets and users. This technology extracts and encodes features of the scene, semantics, and human motions by inputting the scene, text description, and training motion sequences, and fuses these features and inputs them into an iterative optimization trajectory decoder. Using a dual Transformer structure, this method iteratively learns the path planning in the scene and the human motion sequence controlled by the text, generates a 3D motion sequence corresponding to the user's needs and the scene, and realizes controllable text scene action generation for human-machine scene interaction. The drawback of this technology is that the fusion process of scene and text features may be insufficient to capture complex scene constraints and semantic relationships, which may lead to the generated motions lacking physical rationality and semantic accuracy, and the generalization ability for new scenes and diverse descriptions may be limited.

[0007] A patent application with the publication number "CN118644895A" also discloses a text-conditioned human motion generation method based on a discrete diffusion model. This technology includes four steps: First, obtain a human 3D key point data set. Then, use this data set to pre-train an action quantization variational autoencoder to obtain action latent variables. Next, use the human 3D key point data set and action latent variables to train a lightweight discrete diffusion model. Finally, input the given text condition into a text encoder to generate a text feature vector, and decode it through the trained diffusion model and variational autoencoder to obtain a 3D key point motion sequence corresponding to the text condition. The drawback of this technology is that although it can generate high-precision motions that conform to the text description, it cannot consider the scene as a condition, only relies on text features to generate motions, lacks an understanding of the 3D scene environment, and therefore lacks physical rationality and environmental adaptability when generating motions for interacting with complex scenes.

[0008] In addition, the publicly available paper "Scaling up dynamic human-scene interaction modeling, CVPR2024" proposes a human action generation method based on an autoregressive action diffusion model. This technology first generates a motion sequence step by step through a conditional diffusion model according to the 3D scene and the target position. Then, the overall target is divided into smaller sub-goals to ensure precise control of the action implementation in each action segment. At the same time, frame-by-frame action labels are introduced to make the action evolve naturally in different segments, enhancing the coherence and realism of the sequence. Finally, by optimizing the interaction with dynamic objects, it ensures that the generated human action sequence matches the environmental dynamics. The disadvantages of this technology are that it cannot interact using text, which limits the flexibility of users. Secondly, dividing the sub-goals and frame-by-frame action labels make the operation more complex, increasing the learning cost for users. In addition, although it supports real-time generation, when dealing with complex scenes, the computational requirements may lead to a slower generation speed. Summary of the Invention

[0009] To overcome the deficiencies of the above-mentioned prior art in lacking physical rationality and environmental adaptability when generating actions for interacting with complex scenes, the present invention provides a multi-conditional human action generation method and system based on scenes and text. By using bounding boxes to represent the scene, it can quickly and efficiently capture scene information. At the same time, by fusing scene data and text descriptions, the generated actions not only meet the semantic requirements of the text description but also can interact reasonably with the specific scene, effectively improving the accuracy of the generated actions.

[0010] To solve the above technical problems, the technical solution of the present invention is as follows:

[0011] A multi-conditional human action generation method based on scenes and text, comprising the following steps:

[0012] S1: Obtain a ternary data set containing initial action data, scene data, and text data. Extract the bounding box for the scene data, and perform dimensional filling on the initial action data and the scene data after bounding box extraction to obtain a preprocessed ternary data set;

[0013] S2: Construct a scene encoder and a text encoder. For the scene data and text data in the preprocessed ternary data set, use the scene encoder and the text encoder to encode them respectively to obtain scene features and text features ;

[0014] S3: Construct a denoiser, randomly sample time steps , and perform noise addition processing on the action data in the preprocessed ternary data set to obtain the noise-added action data ; The noisy action data , scene features and text features are jointly input into the denoiser to obtain denoised action data ;

[0015] S4: Using the action data in the preprocessed triple dataset and the denoised action data , jointly train the scene encoder and the denoiser using a preset loss function to obtain a trained scene encoder and a trained denoiser;

[0016] S5: Obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features and new text features respectively through the trained scene encoder and the text encoder, generate initial noise by random sampling , and use the trained denoiser to perform step iterative denoising process to finally generate new action data , where is a positive integer.

[0017] Preferably, in step S1, the triple dataset includes a number of triple data samples, and each triple data sample includes an initial action data, a scene data, and a text data respectively;

[0018] The initial action data is specifically a frame sequence of human key points; the scene data is specifically a point cloud of each object in the scene; the text data is specifically a text description of the action data and the scene data.

[0019] Preferably, in step S1, obtain initial action data including a number of actions, the data dimension of a single action is (F, 24, 3), where F is the number of action frames, each frame contains a static pose of the human body, and the static pose is represented by the three-dimensional coordinates of 24 human key points, and 3 indicates that the position of each human key point is represented by three-dimensional space coordinates; use the zero-padding method to uniformly pad the data dimensions of all actions to a fixed number of frames , and at the same time generate a mask to complete the dimension padding of the initial action data to obtain action data ;

[0020] For the scene data in the triple dataset, extract the bounding boxes of each object in the scene; each bounding box represents an object, and the parameters of the i-th bounding box include shape information and category information , where represents the central position coordinates of the i-th bounding box; represents the size of the i-th bounding box; represents the rotation angle of the i-th bounding box around the vertical axis;

[0021] The scene data after bounding box extraction contains bounding boxes of N objects. The scene data after bounding box extraction is uniformly filled to a fixed number using zero-padding method to complete the dimension filling of the scene data after bounding box extraction.

[0022] Preferably, in the step S2, the constructed scene encoder is based on the Transformer encoder structure, and the initial parameters are random parameters;

[0023] For the bounding box of the i-th object , use two juxtaposed multi-layer perceptrons to process the shape information and class information of the i-th bounding box respectively, to obtain shape features and class features; concatenate the shape features and class features to obtain the feature vector of the i-th bounding box ;

[0024] For multiple bounding box features in the scene data , use the Transformer encoder for scene-level global feature aggregation; the feature vectors of each bounding box are aggregated through the self-attention mechanism in the Transformer encoder to capture the relationships between objects and global context information, and finally obtain the scene features containing global information ;

[0025] Use the text encoder in the pre-trained CLIP model to encode the text data to obtain the text features .

[0026] Preferably, the step S3 includes the following steps:

[0027] Randomly sample time steps from the uniform distribution on the interval [0, T] , and perform noise addition processing on the action data according to the time steps to generate the noisy action data , where represents the i-th frame in the action data , represents the i-th frame in the noisy action data ;

[0028] The formula for the noise addition processing is as follows:

[0029]

[0030] in, is standard Gaussian noise; is a predefined time step Decreasing noise weight function;

[0031] The scene features and text features The dimensions are aligned through a linear layer respectively, and the two features after dimension alignment are added and fused to obtain the joint feature. , expressed as:

[0032]

[0033] For the noise-added motion data Each frame in , using a multi-layer perceptron to process each frame As an action code , a total of action code units, represented by ;

[0034] The time step is calculated using the following formula: Perform cosine position encoding:

[0035]

[0036]

[0037] in, and Represents the time step Encoded values ​​on even and odd dimensions, is the dimension index, is the preset embedding dimension;

[0038] The cosine position code is combined with the joint feature Add together to get the conditional code ;

[0039] The denoiser is based on the Transformer encoder structure, and the initial parameters are random parameters; As the input of the denoiser, the output of the denoiser is the denoised codeword sequence , the action codewords in the denoised codeword sequence Through a linear layer, we get the denoised action data .

[0040] Preferably, in the step S4, the preset loss function is specifically the mean square error loss function, and the calculation formula is:

[0041]

[0042] Where is the mean square error loss function value; is the total number of triplet data samples in the triplet dataset; is the i-th action data ; is the i-th denoised action data .

[0043] Preferably, the step S5 includes the following steps:

[0044] Obtain new scene data and perform bounding box extraction, and new text data, and generate new scene features and new text features respectively through the trained scene encoder and the text encoder;

[0045] Input the initial noise generated by the random sampling, as well as the new scene features and the new text features into the trained denoiser together;

[0046] The trained denoiser obtains the action data after the first round of denoising; then perform a noise addition operation on the time step to obtain the noise at the second round of iteration; then use the trained denoiser to perform the second round of denoising to obtain the action data after the second round of iteration;

[0047] Repeat the above process, and use the trained denoiser to perform rounds of iteration process, and finally obtain the action data after the -th round of iteration, and use it as the new action data , which is expressed as: .

[0048] The present invention also provides a multi-condition human action generation system based on scenes and texts, applying the above multi-condition human action generation method based on scenes and texts, including:

[0049] Data preprocessing unit: It is used to obtain a ternary dataset containing initial action data, scene data, and text data, extract bounding boxes from the scene data, and perform dimensional filling on the initial action data and the scene data after bounding box extraction to obtain a preprocessed ternary dataset;

[0050] Feature extraction unit: It is used to construct a scene encoder and a text encoder. For the scene data and text data in the preprocessed ternary dataset, the scene encoder and the text encoder are respectively used for encoding to obtain scene features and text features ;

[0051] Prediction result generation unit: It is used to construct a denoiser, randomly sample time steps , perform noise addition processing on the action data in the preprocessed ternary dataset to obtain the noise-added action data ; The noise-added action data , scene features and text features are jointly input into the denoiser to obtain the denoised action data ;

[0052] Joint training unit: It is used to utilize the action data in the preprocessed ternary dataset and the denoised action data , and use a preset loss function to jointly train the scene encoder and the denoiser to obtain a trained scene encoder and a trained denoiser;

[0053] Action generation unit: It is used to obtain new scene data and perform bounding box extraction, as well as new text data, and respectively generate new scene features and new text features through the trained scene encoder and the text encoder, generate initial noise by random sampling , use the trained denoiser to go through steps of iterative denoising process, and finally generate new action data , where is a positive integer.

[0054] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0055] The present invention also provides an electronic device, including a processor and a memory. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the above method are run.

[0056] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:

[0057] The present invention provides a multi-condition human motion generation method and system based on scenarios and text, and the main problem to be solved is: how to combine scenario data with text descriptions and use a denoising diffusion model for motion generation to ensure that the generated motions not only meet the text semantic requirements but also closely integrate with scenario information; the steps of this method are as follows: First, obtain a ternary dataset and perform preprocessing, then extract features from the scenario data and text data to obtain scenario features and text features; after that, add noise to the motion data in the preprocessed ternary dataset, and at the same time input the scenario features and text features into the denoiser to obtain denoised motion data; use the motion data and the denoised motion data to jointly train the scenario encoder and the denoiser; finally, the user uses the trained scenario encoder and denoiser to generate new motion data.

[0058] The present invention has the following advantages:

[0059] 1) By using bounding boxes to represent scenarios, the present invention can quickly and efficiently capture scenario information, especially suitable for the real-time interaction requirements of dynamic scenarios such as games; as a simple scenario representation method, bounding boxes have lower computational overhead and higher processing efficiency compared to complex 3D models or point cloud data.

[0060] 2) The present invention integrates scenario and text conditions, and the generated motions not only meet the semantic requirements of the text description but also can interact reasonably with specific scenarios; by extracting scenario features and text features through the scenario encoder and text encoder respectively, the model can accurately understand environmental constraints and motion details, thereby generating more physically reasonable and scenario-adaptive motion sequences.

[0061] 3) The present invention uses a denoising diffusion model for step-by-step denoising, ensuring the stability and accuracy of motion generation; compared with traditional methods, the present invention can effectively combine scenario and text information, avoid the limitation of ignoring environmental constraints, improve the generality of the model, and is particularly suitable for real-time interaction in fields such as virtual reality and game development. Brief Description of the Drawings

[0062] Figure 1 It is the overall flowchart of a multi-condition human motion generation method based on scenarios and text provided in Embodiment 1.

[0063] Figure 2 It is the specific implementation flowchart of a multi-condition human motion generation method based on scenarios and text provided in Embodiment 2.

[0064] Figure 3 It is the schematic diagram of the denoising process provided in Embodiment 2.

[0065] Figure 4 It is a structural diagram of a multi - condition human motion generation system based on scenarios and text provided in Embodiment 3. Specific implementation manners

[0066] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on the present invention;

[0067] To better illustrate this embodiment, some components in the accompanying drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;

[0068] For those skilled in the art, it is understandable that some well - known structures and their descriptions in the accompanying drawings may be omitted.

[0069] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0070] Embodiment 1

[0071] As Figure 1 shown, this embodiment provides a multi - condition human motion generation method based on scenarios and text, including the following steps:

[0072] S1: Obtain a ternary data set containing initial motion data, scenario data, and text data, extract the bounding box of the scenario data, and perform dimension filling on the initial motion data and the scenario data after bounding box extraction to obtain a pre - processed ternary data set;

[0073] S2: Construct a scenario encoder and a text encoder. For the scenario data and text data in the pre - processed ternary data set, use the scenario encoder and the text encoder to perform encoding respectively to obtain a scenario feature and a text feature ;

[0074] S3: Construct a denoiser, randomly sample time steps , perform noise addition processing on the motion data in the pre - processed ternary data set to obtain the noise - added motion data ; Input the noise - added motion data , the scenario feature and the text feature into the denoiser together to obtain the denoised motion data ;

[0075] S4: Utilize the motion data in the pre - processed ternary data set and the denoised motion data , jointly train the scene encoder and the denoiser using a preset loss function to obtain a trained scene encoder and a trained denoiser;

[0076] S5: Obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features through the trained scene encoder and the text encoder respectively and new text features , generate initial noise through random sampling , use the trained denoiser to go through steps of iterative denoising process to finally generate new action data , where is a positive integer.

[0077] In the specific implementation process, first obtain a ternary dataset containing initial action data, scene data, and text data, perform bounding box extraction on the scene data, and perform dimension filling on the initial action data and the scene data after bounding box extraction to obtain a preprocessed ternary dataset;

[0078] Then construct a scene encoder and a text encoder, and respectively use the scene encoder and the text encoder to encode the scene data and text data in the preprocessed ternary dataset to obtain scene features and text features ;

[0079] Subsequently, construct a denoiser, randomly sample the time step , add noise to the action data in the preprocessed ternary dataset to obtain the noisy action data ; Input the noisy action data , scene features and text features into the denoiser together to obtain the denoised action data ;

[0080] After that, use the action data in the preprocessed ternary dataset and the denoised action data , jointly train the scene encoder and the denoiser using a preset loss function to obtain a trained scene encoder and a trained denoiser;

[0081] Finally, obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features through the trained scene encoder and the text encoder respectively and new text features , generate initial noise through random sampling , use the trained denoiser to go through In the step - iterative denoising process, new action data is finally generated. ;

[0082] By using the bounding box to represent the scene, this method can quickly and efficiently capture scene information. At the same time, by fusing scene data and text descriptions, the generated actions not only meet the semantic requirements of the text descriptions but also can interact reasonably with the specific scene, effectively improving the stability and accuracy of the generated actions.

[0083] Embodiment 2

[0084] As Figure 2 shown, this embodiment provides a multi - condition human action generation method based on scene and text, including the following steps:

[0085] S1: Obtain a ternary data set containing initial action data, scene data, and text data. Extract the bounding box from the scene data, and perform dimension filling on the initial action data and the scene data after bounding box extraction to obtain a pre - processed ternary data set.

[0086] S2: Construct a scene encoder and a text encoder. For the scene data and text data in the pre - processed ternary data set, use the scene encoder and the text encoder to encode them respectively to obtain scene features and text features ;

[0087] S3: Construct a denoiser, randomly sample time steps , and perform noise addition processing on the action data in the pre - processed ternary data set to obtain the noise - added action data ; Input the noise - added action data , scene features and text features into the denoiser together to obtain the denoised action data ;

[0088] S4: Use the action data in the pre - processed ternary data set and the denoised action data , and use a preset loss function to jointly train the scene encoder and the denoiser to obtain a trained scene encoder and a trained denoiser.

[0089] S5: Obtain new scene data and perform bounding box extraction, as well as new text data. Generate new scene features and new text features respectively through the trained scene encoder and the text encoder, and generate initial noise by random sampling , through the -step iterative denoising process using the trained denoiser, new action data is finally generated , where is a positive integer;

[0090] In the step S1, the triple dataset includes a number of triple data samples, and each of the triple data samples includes an initial action data, a scene data, and a text data respectively;

[0091] The initial action data is specifically a frame sequence of human body key points; the scene data is specifically the point cloud of each object in the scene; the text data is specifically the text description of the action data and the scene data;

[0092] In the step S1, the initial action data including a number of actions is obtained, and the data dimension of a single action is (F, 24, 3), where F is the number of action frames, each frame includes a static pose of the human body, and the static pose is represented by the three-dimensional coordinates of 24 human body key points, and 3 indicates that the position of each human body key point is represented by three-dimensional space coordinates; the zero-padding method is used to uniformly pad the data dimensions of all actions to a fixed number of frames , and at the same time a mask is generated to complete the dimension padding of the initial action data and obtain the action data ;

[0093] For the scene data in the triple dataset, the bounding boxes of each object in the scene are extracted; each of the bounding boxes represents an object, and the parameters of the i-th bounding box include shape information and category information , where represents the central position coordinates of the i-th bounding box; represents the size of the i-th bounding box; represents the rotation angle of the i-th bounding box around the vertical axis;

[0094] After the bounding boxes are extracted, the scene data contains the bounding boxes of N objects, and the zero-padding method is used to uniformly pad the scene data after the bounding boxes are extracted to a fixed number to complete the dimension padding of the scene data after the bounding boxes are extracted;

[0095] In the step S2, the constructed scene encoder is based on the Transformer encoder structure, and the initial parameters are random parameters;

[0096] For the bounding box of the i-th object , two juxtaposed multi-layer perceptrons are used to process the shape information of the i-th bounding box and the category information , shape features and class features are obtained; the shape features and class features are concatenated to obtain the i-th bounding box 's feature vector ;

[0097] For multiple bounding box features in the scene data , use a Transformer encoder for scene-level global feature aggregation; the feature vectors of each bounding box are aggregated through the self-attention mechanism in the Transformer encoder to capture the relationships between objects and global context information, and finally obtain the scene features containing global information ;

[0098] Use the text encoder in the pre-trained CLIP model to encode the text data to obtain the text features ;

[0099] The step S3 includes the following steps:

[0100] Randomly sample time steps from a uniform distribution on the interval [0, T] , and according to the time step Perform noise addition processing on the action data to generate the noisy action data , where represents the i-th frame in the action data , represents the i-th frame in the noisy action data ;

[0101] The formula for the noise addition processing is as follows:

[0102]

[0103] where is standard Gaussian noise; is a predefined noise weight function that decreases with the time step ;

[0104] Align the dimensions of the scene features and the text features respectively through a linear layer, add and fuse the two features after dimension alignment to obtain the joint feature , expressed as:

[0105]

[0106] For each frame in the noisy action data , use a multi-layer perceptron for processing, and process each frame after processing As an action token , a total of action tokens are obtained, denoted as ;

[0107] Use the following formula to perform cosine position encoding on the time step :

[0108]

[0109]

[0110] where and respectively represent the encoding values of the time step on even and odd dimensions, is the dimension index, is the preset embedding dimension;

[0111] Add the cosine position encoding to the joint feature to obtain the conditional token ;

[0112] The denoiser is based on the Transformer encoder structure, and the initial parameters are random parameters; the token sequence is used as the input of the denoiser, and the output of the denoiser is the denoised token sequence . Pass the action tokens in the denoised token sequence through a linear layer to obtain the denoised action data ;

[0113] In step S4, the preset loss function is specifically the mean squared error loss function, and the calculation formula is:

[0114]

[0115] where is the mean squared error loss function value; is the total number of triple data samples in the triple dataset; is the i-th action data ; is the i-th denoised action data ;

[0116] Step S5 includes the following steps:

[0117] Obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features and new text features ;

[0118] Input the initial noise generated by the random sampling , as well as the new scene features and the new text features jointly into the trained denoiser;

[0119] After the first round of denoising by the trained denoiser, the action data after the first round of iteration is obtained ; Subsequently, add noise to the time step to obtain the noise at the second round of iteration ; Then use the trained denoiser for the second round of denoising to obtain the action data after the second round of iteration ;

[0120] Repeat the above process, and use the trained denoiser for rounds of iteration process, and finally obtain the action data after the round of iteration , and use it as the new action data , which is expressed as: .

[0121] In the specific implementation process, first obtain the "action-text-scene" three-way dataset containing the initial action data, scene data, and text data. In this embodiment, the data samples of the three-way dataset are from the public datasets HUMANISE (from the paper "HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes") and Trumans (from the paper "Scaling Up Dynamic Human-Scene Interaction Modeling"); among them, the action data is the frame sequence of human key points, such as an action sequence of a person sitting on the bed; the scene data is the point cloud, such as a bed and other common objects that may exist; the text data is a simple description of the action and the scene, such as "walk to the bedside", "sit on the toilet";

[0122] After obtaining the three-way dataset, preprocessing is still required. Extract the bounding box for the scene data, and perform dimension filling on the initial action data and the scene data after the bounding box extraction to obtain the preprocessed three-way dataset;

[0123] For the action data in the ternary dataset, the dimension of a single action data is (F, 24, 3), where F represents the number of frames, and each frame contains a static pose of the human body, represented as the three-dimensional coordinates of 24 human key points; in human skeleton action capture data, "3" usually indicates that the position of each human key point is represented by three-dimensional space coordinates, which is used to accurately locate the position of human skeleton joints in three-dimensional space. This is also a conventional representation method in the field of action capture;

[0124] The zero-padding method is adopted to uniformly pad each action data to a fixed number of frames ; the purpose of zero-padding is to unify all action data to the same time step, ensuring the consistency of data dimensions in subsequent model training and calculations; the padded data will be accompanied by a mask, which is used to identify which frames are generated by the padding operation. The mask will be used to exclude the invalid padded part in the calculation of the subsequent loss function to avoid interference with the model learning process;

[0125] For the scene data in the ternary dataset, the existing open-source object detection toolbox MMdetection3D is used to extract the bounding boxes of the objects in the scene; each bounding box represents an object in the scene and is defined by the following four pieces of information: the center position , the size , the rotation angle around the vertical axis ( axis) and the class label ; after the extraction operation, a single scene data in the ternary dataset is represented as N an unordered set of bounding boxes; for the scene data, zero-padding is also used to uniformly pad each scene data to a fixed quantity

[0126] Then, a scene encoder and a text encoder are constructed. For the scene data and text data in the preprocessed ternary dataset, the scene encoder and the text encoder are used for encoding respectively to obtain the scene feature and the text feature ;

[0127] In this embodiment, the constructed scene encoder is based on the Transformer encoder structure, and the initial parameters are random parameters;

[0128] For the bounding box of the i-th object , two juxtaposed multi-layer perceptrons are used to process the shape information and the class information , shape features and category features are obtained; the shape features and category features are concatenated to obtain the i-th bounding box of the feature vector ;

[0129] For multiple bounding box features in the scene data , a Transformer encoder is used for global feature aggregation at the scene level; the feature vectors of each bounding box are aggregated through the self-attention mechanism in the Transformer encoder to capture the relationships between objects and global context information, and finally scene features containing global information are obtained ;

[0130] For the text encoder, in this embodiment, the text encoder in the publicly available pre-trained CLIP model (whose parameters are frozen and do not participate in subsequent training) is used to encode the text data to obtain text features ;

[0131] As Figure 3 shown, a denoiser is then constructed, and time steps are randomly sampled , and the action data in the preprocessed triple dataset is subjected to noise addition processing to obtain the noisy action data ; the noisy action data , scene features and text features are jointly input into the denoiser to obtain the denoised action data ;

[0132] Specifically, time steps are randomly sampled from a uniform distribution on the interval [0, T] , and the action data is subjected to noise addition processing according to the time step to generate the noisy action data , where represents the i-th frame in the action data , represents the i-th frame in the noisy action data ;

[0133] The formula for noise addition processing is as follows:

[0134]

[0135] where is standard Gaussian noise, randomly sampled from a multivariate normal distribution with a mean of 0 and a covariance matrix of the identity matrix I; is a predefined value that varies with the time step A decreasing noise weight function, whose specific form can be selected according to actual needs. Common noise weight functions are usually functions that decrease with the time step t, such as exponential decay functions or linear decay functions. Selecting different decay functions can affect the noise control strategy during the generation process, thus having different effects on the generation results. In specific applications, the most suitable noise decay function can be selected according to experimental results or task requirements;

[0136] Align the scene features and text features respectively through a linear layer to ensure that their dimensions are consistent; Add and fuse the two features after dimension alignment to form a 512-dimensional joint feature , denoted as:

[0137]

[0138] For each frame in the noisy action data , use a multi-layer perceptron for processing, and use each processed frame as an action token . Each action token has a feature dimension of 512 dimensions, and a total of action tokens are obtained, denoted as ;

[0139] Given the time step and the embedding dimension = 512, use the following formula to perform cosine position encoding on the time step :

[0140]

[0141]

[0142] where and respectively represent the encoding values of the time step on the even and odd dimensions, is the dimension index;

[0143] Add the cosine position encoding to the joint feature to obtain the conditional token ;

[0144] The denoiser is based on the Transformer encoder structure, and the initial parameters are random parameters; Use the token sequence as the input of the denoiser, and the output of the denoiser is the denoised token sequence . Take the action tokens in the denoised token sequence Through a linear layer, the denoised action data is obtained. ;

[0145] After that, using the action data in the preprocessed triple dataset and the denoised action data , the scene encoder and the denoiser are jointly trained using the mean squared error loss function to obtain the trained scene encoder and the trained denoiser;

[0146] In the training process of this embodiment, the mean squared error is used as the loss function to measure the difference between the action data output by the denoiser and the real unnoisy action data. The calculation formula of the loss function is as follows:

[0147]

[0148] Where is the value of the mean squared error loss function; is the total number of triple data samples in the triple dataset; is the i-th action data ; is the i-th denoised action data ;

[0149] Finally, new scene data is obtained and bounding box extraction is performed, as well as new text data. New scene features and new text features are generated respectively through the trained scene encoder and text encoder; Set the time step , randomly sample to generate the initial noise , the time step is a fixed value, and common values are such as 50, 500, 1000;

[0150] The initial noise generated by random sampling, as well as the new scene features and the new text features are jointly input into the trained denoiser;

[0151] The trained denoiser obtains the action data after the first round of denoising after the first round of denoising ; Subsequently, the time step is noise-added to obtain the noise at the second round of iteration; After that, the trained denoiser is used for the second round of denoising to obtain the action data after the second round of denoising ;

[0152] Repeat the above process, and use the trained denoiser for rounds of iteration process, and finally obtain the Action data after round iteration and use it as the new action data , expressed as: ;

[0153] By using bounding boxes to represent the scene, this method can quickly and efficiently capture scene information. At the same time, by fusing scene data and text descriptions, the generated actions not only meet the semantic requirements of the text descriptions but also can interact reasonably with the specific scene, effectively improving the stability and accuracy of the generated actions.

[0154] Embodiment 3

[0155] As Figure 4 shown, this embodiment provides a multi-condition human action generation system based on scenes and text, applying the above-mentioned multi-condition human action generation method based on scenes and text, including:

[0156] Data preprocessing unit 301: used to obtain a ternary data set containing initial action data, scene data, and text data, extract bounding boxes from the scene data, and perform dimension filling on the initial action data and the scene data after bounding box extraction to obtain a preprocessed ternary data set;

[0157] Feature extraction unit 302: used to construct a scene encoder and a text encoder, and respectively encode the scene data and text data in the preprocessed ternary data set using the scene encoder and the text encoder to obtain scene features and text features ;

[0158] Prediction result generation unit 303: used to construct a denoiser, randomly sample time steps , perform noise addition processing on the action data in the preprocessed ternary data set to obtain the noise-added action data ; input the noise-added action data , scene features and text features into the denoiser together to obtain the denoised action data ;

[0159] Joint training unit 304: used to jointly train the scene encoder and the denoiser using the action data in the preprocessed ternary data set and the denoised action data using a preset loss function to obtain a trained scene encoder and a trained denoiser;

[0160] Action generation unit 305: It is used to obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features through the trained scene encoder and the text encoder respectively and new text features , generate initial noise through random sampling , and use the trained denoiser through step iterative denoising process to finally generate new action data , where is a positive integer.

[0161] In the specific implementation process, first, the data preprocessing unit 301 obtains a ternary dataset containing initial action data, scene data, and text data, performs bounding box extraction on the scene data, and performs dimension padding on the initial action data and the scene data after bounding box extraction to obtain the preprocessed ternary dataset;

[0162] Next, the feature extraction unit 302 constructs a scene encoder and a text encoder. For the scene data and text data in the preprocessed ternary dataset, they are encoded using the scene encoder and the text encoder respectively to obtain scene features and text features ;

[0163] Subsequently, the prediction result generation unit 303 constructs a denoiser, randomly samples time steps , and performs noise addition processing on the action data in the preprocessed ternary dataset to obtain the noise-added action data ; The noise-added action data , scene features and text features are jointly input into the denoiser to obtain the denoised action data ;

[0164] After that, the joint training unit 304 uses the action data in the preprocessed ternary dataset and the denoised action data , and jointly trains the scene encoder and the denoiser using a preset loss function to obtain the trained scene encoder and the trained denoiser;

[0165] Finally, the action generation unit 305 obtains new scene data and performs bounding box extraction, as well as new text data, and generates new scene features and new text features , generate initial noise through random sampling , and use the trained denoiser through The step iterative denoising process finally generates new action data ;

[0166] By using bounding boxes to represent the scene, this system can quickly and efficiently capture scene information. At the same time, by fusing scene data and text descriptions, the generated actions not only meet the semantic requirements of the text descriptions but also can interact reasonably with the specific scene, effectively improving the stability and accuracy of the generated actions.

[0167] Like or similar reference numerals correspond to like or similar components;

[0168] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as limiting the present invention;

[0169] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A multi-condition human motion generation method based on scenarios and text, characterized in that It includes the following steps: S1: Obtain a ternary dataset containing initial action data, scene data, and text data, perform bounding box extraction on the scene data, and perform dimensionality padding on the initial action data and the scene data after bounding box extraction to obtain a preprocessed ternary dataset; Among them, initial action data containing several actions is obtained. The data dimension of a single action is (F, 24, 3), where F is the number of action frames, and each frame contains a static pose of the human body. The static pose is represented by the three-dimensional coordinates of 24 human key points, and 3 indicates that the positions of each human key point are represented by three-dimensional space coordinates. The zero-padding method is used to uniformly pad the data dimensions of all actions to a fixed number of frames , and at the same time, a mask is generated to complete the dimension padding of the initial action data, obtaining action data ; For the scene data in the ternary dataset, extract the bounding boxes of each object in the scene; each said bounding box represents an object, and the parameters of the i-th bounding box include shape information and category information , where represents the central position coordinates of the i-th bounding box; represents the size of the i-th bounding box; represents the rotation angle of the i-th bounding box around the vertical axis; The scene data after bounding box extraction contains the bounding boxes of N objects, and the scene data after bounding box extraction is uniformly filled to a fixed number using the zero-padding method , completing the dimension filling of the scene data after bounding box extraction; S2: Construct a scene encoder and a text encoder, and use the scene encoder and the text encoder to encode the scene data and text data in the preprocessed triple dataset respectively to obtain scene features and text features ; The constructed scene encoder is based on the Transformer encoder structure, and the initial parameters are random parameters; Bounding box for the i-th object For the shape information of the i-th bounding box, two juxtaposed multi-layer perceptrons are used for processing and the category information to obtain the shape feature and the category feature; the shape feature and the category feature are concatenated to obtain the feature vector of the i-th bounding box ; For multiple bounding box features in the scene data , use a Transformer encoder for global feature aggregation at the scene level; the feature vectors of each bounding box are aggregated through the self-attention mechanism in the Transformer encoder to capture the relationships between objects and global context information, and finally, scene features containing global information are obtained ; Encode the text data using the text encoder in the pre-trained CLIP model to obtain the text features ; S3: Build a denoiser and randomly sample time steps , and add noise to the action data in the preprocessed three - tuple dataset to obtain the noisy action data ; Input the noisy action data , scene features and text features into the denoiser together to obtain the denoised action data ; S4: Use the action data in the preprocessed triple dataset and the denoised action data , and use a preset loss function to jointly train the scene encoder and the denoiser to obtain a trained scene encoder and a trained denoiser; S5: Obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features through the trained scene encoder and the text encoder respectively and new text features , generate initial noise through random sampling , and use the trained denoiser through step iterative denoising process to finally generate new action data , where is a positive integer.

2. The multi-condition human motion generation method based on scenario and text according to claim 1, characterized in that, In the step S1, the ternary dataset includes a number of ternary data samples, and each ternary data sample includes an initial action data, a scene data, and a text data respectively; The initial action data is specifically a frame sequence of human key points; the scene data is specifically the point cloud of each object in the scene; the text data is specifically a text description of the action data and the scene data.

3. A method for generating multi-condition human actions based on scenarios and texts according to claim 1, characterized in that, The step S3 includes the following steps: Randomly sample time steps from a uniform distribution on the interval [0, T] , and based on the time steps add noise to the action data to generate the noisy action data , where represents the i-th frame in the action data , and represents the i-th frame in the noisy action data ; The formula for the noise addition process is as follows: Among them, is standard Gaussian noise; is a predefined noise weight function that decreases with time step ; Align the dimensions of the described scene features and text features respectively through a linear layer, add and fuse the two features after dimension alignment to obtain a joint feature , expressed as: For each frame in the noisy action data , it is processed using a multi-layer perceptron, and each processed frame is used as an action code element . A total of action code elements are obtained, denoted as ; Perform cosine position encoding on the said time step using the following formula : Among them, and respectively represent the encoded values of the time step on the even and odd dimensions, is the dimension index, is the preset embedding dimension; Add the cosine position encoding to the combined feature to obtain a conditional code element ; The denoiser is based on the Transformer encoder structure and its initial parameters are random parameters; the symbol sequence is used as the input of the denoiser, and the output of the denoiser is the denoised symbol sequence . The action symbols in the denoised symbol sequence pass through a linear layer to obtain the denoised action data .

4. A multi-condition human motion generation method based on scenarios and text according to claim 1, characterized in that, In the step S4, the preset loss function is specifically the mean square error loss function, and the calculation formula is: Among them, is the mean square error loss function value; is the total number of triple data samples in the triple dataset; is the i-th action data ; is the i-th denoised action data .

5. A multi-condition human motion generation method based on scenarios and text according to claim 1, characterized in that The step S5 includes the following steps: Obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features through the trained scene encoder and the text encoder respectively and new text features ; Input the initial noise generated by the random sampling , along with the new scene features and the new text features jointly into the trained denoiser; The trained denoiser obtains the action data after the first round of denoising after the first iteration ; Subsequently, noise addition operation is performed on the time step to obtain the noise at the second iteration ; Then, the trained denoiser is used for the second round of denoising to obtain the action data after the second iteration ; Repeat the above process and use the trained denoiser for rounds of iterative processes, and finally obtain the action data after the round of iteration, and use it as the new action data , which is expressed as: .

6. A multi-condition human motion generation system based on scenarios and text, applying the multi-condition human motion generation method based on scenarios and text described in any one of claims 1 to 5, characterized in that, It includes: Data preprocessing unit: used to obtain a ternary dataset containing initial action data, scene data, and text data, perform bounding box extraction on the scene data, and perform dimensionality padding on the initial action data and the scene data after bounding box extraction to obtain a preprocessed ternary dataset; Feature extraction unit: used to construct a scene encoder and a text encoder, and respectively encode the scene data and text data in the preprocessed triple dataset using the scene encoder and the text encoder to obtain scene features and text features ; Prediction result generation unit: used to construct a denoiser and randomly sample time steps , and perform noise addition processing on the action data in the preprocessed three-way dataset to obtain the noisy action data ; The noisy action data , scene features and text features are jointly input into the denoiser to obtain the denoised action data ; Joint training unit: used to utilize the action data in the preprocessed triple dataset and the denoised action data , and jointly train the scene encoder and the denoiser using a preset loss function to obtain a trained scene encoder and a trained denoiser; Action generation unit: used to obtain new scene data and perform bounding box extraction, as well as new text data, and generate new scene features through the trained scene encoder and the text encoder respectively and new text features , generate initial noise through random sampling , use the trained denoiser through step iterative denoising process to finally generate new action data , where is a positive integer.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the method described in any one of claims 1 to 5.

8. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the method described in any one of claims 1 to 5 are run.

Citation Information

Patent Citations

  • Human body action generation method capable of interacting with three-dimensional scene target and user

    CN117953113A

  • Generating action data for the animation of characters

    CA2314712A1

  • Action generation method, device and equipment based on action generation model

    CN116702707A