Generative model training method, generation method and system of human motion sequence
By training the sample set to generate stickman and combining the diffusion model, the multi-condition fusion of the human body motion sequence generation model is optimized, which solves the problem of inconsistent stickman and text representation, and achieves a more natural and accurate human body motion sequence generation.
Patent Information
- Application Number
- CN202510399014.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The prior art is difficult to quickly generate natural and accurate human motion sequences, especially when dealing with complex and detailed human motion descriptions, stickman inconsistent with text representations leads to a degradation in the generation performance and high computational volume.
Stickman is generated by training the sample set, and the diffusion forward noise addition module, motion encoder, stickman encoder, text encoder and multi-layer multi-condition fusion module are used, combined with the diffusion reverse noise denoising module, a human motion sequence generation model is built, and the stickman and text vectors are fused through a multi-layer multi-condition fusion module and attention mechanism to optimize the generation process.
It reduces the need for complex text description data, improves the degree of fit between generated motion and user imagination, improves recognition accuracy and multi-condition fusion performance, and achieves more natural and accurate human motion sequence generation.
Smart Images

Figure CN120298554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human motion sequence generation, and in particular, to a method, a generation method and a system for training a generation model of a human motion sequence. Background Art
[0002] The human motion sequence generation technology is widely used in fields such as film and television production, virtual reality simulation, and the game industry. Among them, the text-driven human motion sequence generation technology can generate natural human motion sequences according to language descriptions, reducing the burden on 3D animators to manually set key frames.
[0003] In the prior art, a description of human motion, such as "squat after running a few steps", is input into a deep learning-based model, and then a human motion sequence is output. This motion can be represented by 3D joint coordinates or joint rotation matrices. Simply relying on text description data often makes it difficult to accurately capture the complex and detailed human motions in the user's mind. For example, for the simple description of "kick the leg forward high", the specific posture of the arm cannot be determined. Previous studies have focused on generating fine-grained motions through complex text description data. Flame allows modifying the character motion sequence by attaching text description data based on the diffusion model; FineMoGen controls each body part of the 3D character through detailed descriptions. However, these methods require users to provide more detailed text description data to generate motion sequences that meet expectations. In terms of multi-condition fusion and data processing, existing methods also have deficiencies. When processing multiple condition inputs, previous work achieved condition combination by using masking operations in the self-attention module, but unnecessary computational complexity is introduced when calculating the attention of masked tokens, and the inconsistent representation between the stick figure and the text will lead to a decline in the performance of generating human motion sequences. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, a generation method and a system for training a generation model of a human motion sequence to eliminate or improve one or more defects existing in the prior art, and solve the problem that a natural and accurate human motion sequence cannot be quickly generated in the prior art.
[0005] One aspect of the present invention provides a method for training a human motion sequence generation model, the method comprising the following steps:
[0006] Obtain a training sample set, the training sample set including a plurality of samples, each sample containing a human motion sequence for the sample action and its text description data; generate a corresponding stick figure for the human motion sequence in each sample based on a preset stick figure generation algorithm; the human motion sequence is labeled with the true value of the human motion sequence; the stick figure is labeled with the true value of the position index.
[0007] The initial human motion sequence generation model is trained using the training sample set; the initial human motion sequence generation model includes a diffusion forward noise addition module, a motion encoder, a stick figure encoder, a text encoder, a multi-layer multi-condition fusion module, a backbone network, and a diffusion backward denoising module; the multi-layer multi-condition fusion module layer includes a first feature decoder, a second feature decoder, and a latent encoder; the diffusion forward noise addition module receives the human motion sequence and gradually adds noise to obtain a noisy motion sequence, encodes the noisy motion sequence into a motion vector through the motion encoder, encodes the stick figure into a stick figure vector through the stick figure encoder, encodes the text description data into a text vector through the text encoder, the multi-layer multi-condition fusion module at least batches and combines the stick figure vector and the text vector to obtain multiple combined input vectors and integrates them into the motion vector, the first combined input vector at least includes the text vector, the second combined input vector at least includes the text vector and the stick figure vector, the third combined input vector at least includes the stick figure vector, the fourth combined input vector does not include the text vector and the stick figure vector, the first feature decoder at least receives the first combined input vector and the second combined input vector, the second feature decoder at least receives the third combined input vector and the fourth combined input vector and obtains a text influence motion offset and a stick figure influence motion offset after integrating them into the motion vector through an attention mechanism, adds them to the motion vector to obtain a target motion vector, re-encodes it through the latent encoder and inputs it into the backbone network, and outputs a predicted motion noise and a predicted value of the position index of the stick figure; the human motion sequence prediction value is obtained by gradually removing the predicted motion noise through the diffusion backward denoising module;
[0008] A motion loss function is constructed through the human motion sequence prediction value and the true value of the human motion sequence, an index loss function is constructed through the predicted value of the position index and the true value of the position index, the motion loss function and the index loss function are fused to obtain a total loss function, and the parameters of the motion encoder, the multi-layer multi-condition fusion module, the backbone network, and the diffusion backward denoising module are updated with the goal of minimizing the total loss function, and the human motion sequence generation model is constructed and obtained.
[0009] In some embodiments, the index loss function The motion loss function And the total loss function The expressions are:
[0010]
[0011] where M represents a preset weight coefficient, L represents the sequence length, Denote the predicted position index of the stickman in the l-th frame of the predicted human motion sequence. Denote the human motion sequence generated under the second combined input vector and the third combined input vector. Denote the human motion sequences generated under all conditional combinations, x l Denote the action in the l-th frame of the true value of the human motion sequence, x i Denote the action of the stickman generated from the true motion sequence through the stickman generation algorithm.
[0012] In some embodiments, the method further includes supervising the motion loss using a classifier-free diffusion guidance algorithm, and the supervision expression is:
[0013]
[0014] Wherein, Denote the calculation of the expected value for the noise true value ∈ t , time step t, and initial motion x0; ∈ θ Denote the predicted motion noise; C(stick) denotes the input text vector; C(stick) denotes the input stickman vector.
[0015] In some embodiments, when obtaining the predicted value of the human motion sequence by gradually removing the predicted motion noise through the diffusion denoising module, the expression of the predicted motion noise is:
[0016]
[0017] When t ∈ [T, T / 10],
[0018] When t ∈ [T / 10, 0], (w1 = 1; w2, w3, w4 = 0);
[0019] Wherein, stick denotes the stickman, text denotes the text, w and are custom weight parameters.
[0020] In some embodiments, the process of generating the corresponding stickman for the human motion sequence in each sample based on a preset stickman generation algorithm includes:
[0021] Normalize the 3D joint point coordinates in the preset joint coordinate dataset using the lengths of the human arms and legs as the denominator, and obtain the 2D normalized coordinates through a front view pose with the line of sight perpendicular to the target human pelvis plane.
[0022] Generate a standard circle, truncate it at random positions on the circle, and then randomly extend it to obtain the head line. Connect each of the 2D normalized coordinates to obtain the limb lines and the spine line;
[0023] After adding a first preset number of interpolation coordinate points representing the line trajectory to the head line, the limb lines, and the spine line, accumulate and add a randomly generated noise array with the same shape as the line and evenly distributed to the line to obtain a noise-added line. After smoothing the noise-added line through Gaussian filtering, align it with the original head line, the limb lines, and the spine line through affine transformation, and add a generated small-amplitude evenly distributed noise array to obtain a hand-drawn style line;
[0024] After obtaining the head hand-drawn style line, the spine hand-drawn style line, and the limb hand-drawn style lines, reasonably combine the head hand-drawn style line and the spine hand-drawn style line in the neck direction, connect the arm hand-drawn style lines to the head of the spine hand-drawn style line, and connect the leg hand-drawn style lines to the tail of the spine hand-drawn style line to obtain an initial stick figure. Perturb the overall position of each line of the initial stick figure by generating two noise values and adding them to the horizontal and vertical coordinates of the line, and perform equidistant sampling on each line of the initial stick figure at a certain distance interval to unify the number of coordinate points of each line to obtain a stick figure.
[0025] In some embodiments, encode the noise motion sequence into a motion vector through a motion encoder, encode the stick figure into a stick figure vector through a stick figure encoder, and encode the text description data into a text vector through a text encoder, including:
[0026] The motion encoder uses a simple linear layer to encode the noise motion sequence to obtain a motion vector;
[0027] The stick figure encoder uses a transform encoder to encode each line of the stick figure to obtain a stick figure vector;
[0028] The text encoder uses CLIP ViT-B / 32 to encode the text description data to obtain a text vector.
[0029] On the other hand, the present invention also provides a method for generating a human motion sequence. The method is performed in a human motion sequence generation model obtained by the training method of the human motion sequence generation model described in any one of the above. The method includes the following steps:
[0030] Input a randomly generated noise motion sequence, a stick figure, and text description data about the target action into the human motion sequence generation model, and output the human motion sequence of the target action.
[0031] On the other hand, the present invention also provides a human motion sequence generation system, including a processor, a memory, and computer programs / instructions stored on the memory, characterized in that the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed, the system implements the steps of the method described in any one of the above.
[0032] On the other hand, the present invention also provides a computer-readable storage medium, on which computer programs / instructions are stored, characterized in that when the computer programs / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0033] On the other hand, the present invention also provides a computer program product, including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0034] In the training method of the human motion sequence generation model based on stick figures of the present invention, the human motion sequence is generated to correspond to a stick figure according to a preset stick figure generation algorithm, and the control of the stick figure is introduced to move the details, reducing the demand for complex text description data and improving the fit between the generated motion and the user's imagination; the noise motion sequence, the stick figure, and the text description data are input into the initial human motion sequence generation model. After each line of the stick figure is separately encoded by the stick figure encoder, through simple transformation encoder interaction, an embedded representation of the stick figure is obtained, reducing the computational demand and improving the recognition accuracy. The multi-layer multi-condition fusion module has a lower computational complexity and better performance, effectively solving the problem of inconsistent representation between the stick figure and the text, and better integrating multiple input conditions to generate a high-quality motion sequence; the first feature decoder receives at least the first combined input vector and the second combined input vector, and the second feature decoder receives at least the third combined input vector and the fourth combined input vector and incorporates the motion vector through an attention mechanism to obtain the text influence motion offset and the stick figure influence motion offset. After adding them to the motion vector, the target motion vector is obtained. The target motion vector is input into the latent encoder for re-encoding and then input into the backbone network to output the predicted motion noise and the position index prediction value of the stick figure. The noise in the predicted noise is gradually removed to obtain the human motion sequence prediction value, improving the multi-condition fusion performance, optimizing the generation process, realizing a more natural and accurate human motion sequence generation, and providing a simpler and faster method for obtaining an ideal human motion sequence.
[0035] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows, and will partly become obvious to those of ordinary skill in the art upon examination of the following, or may be learned by practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structure particularly pointed out in the specification and the drawings.
[0036] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other objects that the present invention can achieve will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings described herein are used to provide a further understanding of the present invention, and form a part of this application, and do not limit the present invention. In the drawings:
[0038] Figure 1 is a schematic flow chart of a method for training a human motion sequence generation model according to an embodiment of the present invention.
[0039] Figure 2 is a schematic structural diagram of a method for training a human motion sequence generation model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the objects, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0041] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0042] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0043] Here, it should also be noted that if not otherwise specified, the term "connection" in this document can not only refer to direct connection, but also represent indirect connection with an intermediate.
[0044] In the following, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0045] In the prior art, when a description of human motion, such as "squat down after running a few steps", is input into a deep learning-based model and a human motion sequence is output, it is often difficult to accurately capture the complex and detailed human motion in the user's mind solely relying on text description data. For example, for a simple description like "kick high forward", the specific posture of the arm cannot be determined. Previous research has been dedicated to generating fine-grained motion through complex text description data. Flame allows modifying the character motion sequence based on a diffusion model with additional text description data; FineMoGen controls each body part of a 3D character through detailed descriptions. However, these methods require users to provide more detailed text description data to generate a motion sequence that meets expectations. In terms of multi-condition fusion and data processing, existing methods also have deficiencies. When dealing with multiple condition inputs, previous work achieved condition combination by using a masking operation in the self-attention module. However, unnecessary computational effort is introduced when calculating the attention of masked tokens, and the inconsistent representation between the stick figure and the text leads to performance degradation. The present invention proposes a method, a generation method and a system for training a generation model of a human motion sequence. Each sample in the training sample set includes a human motion sequence for the sample action and its text description data. For the human motion sequence in each sample, a corresponding stick figure is generated based on a preset stick figure generation algorithm; the initial human motion sequence generation model is trained using the training sample set; noise is gradually added to the human motion sequence to obtain a noisy motion sequence. The noisy motion sequence is encoded into a motion vector by a motion encoder, the stick figure is encoded into a stick figure vector by a stick figure encoder, and the text description data is encoded into a text vector by a text encoder. The multi-layer multi-condition fusion module at least batches and combines the stick figure vector and the text vector to obtain multiple combined input vectors and integrates them into the motion vector. The first combined input vector includes at least the text vector, the second combined input vector includes at least the text vector and the stick figure vector, the third combined input vector includes at least the stick figure vector, and the fourth combined input vector does not include the text vector and the stick figure vector. The first feature decoder receives at least the first combined input vector and the second combined input vector, the second feature decoder receives at least the third combined input vector and the fourth combined input vector, and after integrating them into the motion vector through an attention mechanism, a text-influenced motion offset and a stick figure-influenced motion offset are obtained. After adding them to the motion vector, a target motion vector is obtained, which is re-encoded by a latent encoder and then input into the backbone network to output a predicted motion noise and a predicted value of the position index of the stick figure; the predicted motion noise is gradually removed through a diffusion denoising module to obtain a predicted value of the human motion sequence;Construct a motion loss function using the predicted value and the true value of the human motion sequence, construct an index loss function using the predicted value and the true value of the position index, fuse the motion loss function and the index loss function to obtain the total loss function, and update the parameters of the motion encoder, the multi-layer multi-condition fusion module, the backbone network, and the diffusion denoising module in the reverse direction with the goal of minimizing the total loss function, and construct a human motion sequence generation model.
[0046] Figure 1 FIG. 4 is a schematic flow chart of a training method for the human motion sequence generation model according to an embodiment of the present invention. Specifically, one aspect of the present invention provides a training method for a human motion sequence generation model, and the method includes the following steps S101 to S103:
[0047] Step S101: Obtain a training sample set. The training sample set includes a plurality of samples, and each sample includes a human motion sequence for the sample action and its text description data; generate a corresponding stick figure for the human motion sequence in each sample based on a preset stick figure generation algorithm; the human motion sequence is labeled with the true value of the human motion sequence; the stick figure is labeled with the true value of the position index.
[0048] Step S102: Train an initial human motion sequence generation model using the training sample set; the initial human motion sequence generation model includes a diffusion forward noise addition module, a motion encoder, a stick figure encoder, a text encoder, a multi-layer multi-condition fusion module, a backbone network, and a diffusion denoising module in the reverse direction; the multi-layer multi-condition fusion module includes a first feature decoder, a second feature decoder, and a latent encoder; the diffusion forward noise addition module receives the human motion sequence and gradually adds noise to obtain a noisy motion sequence, encodes the noisy motion sequence into a motion vector through the motion encoder, encodes the stick figure into a stick figure vector through the stick figure encoder, encodes the text description data into a text vector through the text encoder, the multi-layer multi-condition fusion module at least batches and combines the stick figure vector and the text vector to obtain a plurality of combined input vectors and integrates them into the motion vector, the first combined input vector includes at least the text vector, the second combined input vector includes at least the text vector and the stick figure vector, the third combined input vector includes at least the stick figure vector, the fourth combined input vector does not include the text vector and the stick figure vector, the first feature decoder receives at least the first combined input vector and the second combined input vector, the second feature decoder receives at least the third combined input vector and the fourth combined input vector and obtains a text influence motion offset and a stick figure influence motion offset after integrating the motion vector through an attention mechanism, adds them to the motion vector to obtain a target motion vector, re-encodes it through the latent encoder and inputs it into the backbone network, outputs a predicted motion noise and a predicted value of the position index of the stick figure; obtains a predicted value of the human motion sequence by gradually removing the predicted motion noise through the diffusion denoising module in the reverse direction.
[0049] Step S103: Construct a motion loss function through the predicted value and the true value of the human motion sequence, construct an index loss function through the predicted value and the true value of the position index, fuse the motion loss function and the index loss function to obtain the total loss function, and update the parameters of the motion encoder, the multi-layer multi-condition fusion module, the backbone network, and the diffusion denoising module with the goal of minimizing the total loss function, and construct a human motion sequence generation model.
[0050] In step S101, a plurality of human motion sequences are preset in the motion sequence dataset. The human motion sequence is a set of actions arranged in chronological order, and is often used to describe the motion changes of a person or an object over a period of time. The stickman generation algorithm automatically generates a stickman based on the 3D joint coordinates in the existing joint coordinate dataset. The stickman algorithm converts the 3D joint coordinates into 2D handwriting coordinate sequence data, and represents the stickman through a series of coordinate sequences. During the stickman generation process, stroke smoothness, position deviation, and scaling differences are considered. In some embodiments, the process of generating a corresponding stickman for the human motion sequence in each sample based on the preset stickman generation algorithm includes steps S1011 to S1014:
[0051] Step S1011: Normalize the 3D joint coordinates in the preset joint coordinate dataset using the lengths of the human arms and legs as the denominator, and obtain 2D normalized coordinates by observing the front view pose with the line of sight perpendicular to the target human pelvis plane.
[0052] Step S1012: Generate a standard circle, randomly truncate it at a random position on the circle, and then randomly extend it to obtain a head line. Connect each 2D normalized coordinate to obtain limb lines and a spine line.
[0053] Step S1013: After adding a first preset number of interpolation coordinate points representing the line trajectory to the head line, limb lines, and spine line, add a randomly generated noise array with the same shape as the line and uniformly distributed to the line to obtain a noisy line. Smooth the noisy line through Gaussian filtering, then align it with the original head line, limb lines, and spine line through affine transformation, and add a small uniformly distributed noise array generated to obtain a hand-drawn style line.
[0054] Step S1014: After obtaining the head hand-drawn style line, spine hand-drawn style line, and limb hand-drawn style lines, the head hand-drawn style line and spine hand-drawn style line are reasonably combined in the neck direction, and the arm hand-drawn style line is connected to the head of the spine hand-drawn style line, and the leg hand-drawn style line is connected to the tail of the spine hand-drawn style line to obtain the initial stick figure. The overall position of each line of the initial stick figure is perturbed by adding two noise values to the horizontal and vertical coordinates of the line, and equidistant sampling is performed on each line of the initial stick figure at a certain distance interval to unify the number of coordinate points of each line to obtain the stick figure.
[0055] In step S102, the diffusion forward noise addition module adds Gaussian noise to the human motion sequence during the diffusion forward process of the diffusion model, gradually adding Gaussian noise ∈ t ~N(0, I) from the initial time t = 0 to the final time t = T to the initial actual motion x0. The Gaussian noise follows a normal distribution with a mean of 0 and a covariance matrix of the identity matrix I, and the weight of adding noise is an increasing sequence β t , and the probability distribution formula of the forward process is where q(x 1:T |x0) represents the joint probability of this series of noisy motion sequences from x1 to x T given the initial motion x0; is the product symbol, indicating the product of the probabilities q(x t |x t-1 ) from t = 1 to t = T for each time step; indicates that given the motion x t-1 at the previous time step, the motion x t at the current time step follows a normal distribution with a mean of and a variance of (1 - α t )I. Here, α t is related to the weight β t , and α t = 1 - β t ; this formula is equivalent to where is the product of all α s from 1 to the t-th moment.
[0056] In some embodiments, a noisy motion sequence is encoded into motion vectors by a motion encoder, a stick figure is encoded into stick figure vectors by a stick figure encoder, and text description data is encoded into text vectors by a text encoder. Specifically, the motion encoder uses a simple linear layer to encode the noisy motion sequence to obtain motion vectors. The stick figure encoder uses a transform encoder to encode each line of the stick figure to obtain stick figure vectors. The text encoder uses CLIP ViT-B / 32 to encode the text description data to obtain text vectors. Specifically, the noisy motion sequence, the stick figure, and the text description data are respectively encoded into vectors of shapes [L m , E], [L s , E], and [L t , E], where L m , L s , L t represent the sequence lengths of the respective inputs, and E represents the dimension of the input encoding. The stick figure encoder uses a Transformer encoder to separately encode each of the six lines representing the head, torso, and limbs of the stick figure, and then obtains the embedded representation of the stick figure through encoder interaction, reducing the computational requirements and improving the accuracy.
[0057] Furthermore, each layer in the multi-layer multi-condition fusion module divides and combines the current batch of input stick figure vectors and text vectors according to the batch dimension to obtain multiple input condition combinations: the first combined input vector B1, the second combined input vector B2, the third combined input vector B3, and the fourth combined input vector B4. For the condition combinations in the four groups of input conditions that contain text vectors, only the first feature decoder (Feat Decoder) of the text vector is considered. For the condition combinations that contain stick figure vectors, only the second feature decoder of the stick figure vector is considered as the value vector and the key vector, and they are fused with the motion vector as the query vector through the self-attention mechanism to obtain the text influence motion offset representing the direction and degree of adjustment of the text vector to the motion vector and the stick figure influence motion offset representing the direction and degree of adjustment of the stick figure vector to the motion vector. The target motion vector is re-encoded by the latent encoder to further fuse information. The backbone network uses a Transformer. The motion noise of the noisy motion sequence is gradually removed to obtain the predicted value of the human motion sequence, and the human motion sequence is visually displayed according to the 3D coordinates. The position index score is used to represent the position of the action represented by the input stick figure in the noisy motion sequence, assisting the user in judging whether the output human motion sequence meets the requirements of the stick figure and the text description data.
[0058] In some embodiments, when obtaining the predicted value of the human motion sequence by gradually removing the predicted motion noise through the diffusion denoising module, the expression of the predicted motion noise is:
[0059]
[0060] When \(t\in[T, T / 10]\),
[0061] When \(t\in[T / 10, 0]\), \((w1 = 1; w2, w3, w4 = 0)\);
[0062] where, stick represents a stick figure, text represents text, and w and are user-defined weight parameters.
[0063] Specifically, during the training process of the human motion sequence generation model, by adjusting the conditional mixing weights w1, w2, w3, w4, the generated human motion sequence is biased towards different conditional combinations. Through the classifier-free diffusion guidance algorithm, the output human motion sequence reaches a balance between the stick figure vector and the text vector. The motion loss is supervised by calculating the expected value through predicting the motion noise and the true value of the noise. The probabilities of the appearance of the stick figure vector input condition and the text vector input condition during the generation process of predicting the motion noise are both 0.7. When denoising in the diffusion reverse denoising module, by adjusting the conditional mixing weights w1, w2, w3, w4 and \(w1 + w2 + w3 + w4 = 1\), the output predicted motion noise is biased towards different combined input vectors, and the weight distribution in the initial stage and the final stage is defined to optimize the predicted noise; (stick, text) represents the combined input vector containing the stick figure vector and the text vector, represents the combined input vector containing the stick figure vector, represents the combined input vector containing the text vector, represents the combined input vector that does not contain the stick figure vector and the text vector.
[0064] In step S103, a motion loss function, an index loss function, and a fusion loss function are constructed. In some embodiments, the index loss function Motion loss function And the total loss function The expressions are:
[0065]
[0066] where, M represents a preset weight coefficient, L represents the sequence length, represents the predicted value of the position index of the stick figure in the \(l\)-th frame in the human motion sequence prediction value, represents the human motion sequence generated under the second combined input vector and the third combined input vector, represents the human motion sequence generated under all conditional combinations, x lDenote the action of the l-th frame in the true value of the human motion sequence, x i Denote the action of generating a stick figure from the true motion sequence through the stick figure generation algorithm.
[0067] In some embodiments, the method further includes supervising the motion loss by using a classifier-free diffusion guidance algorithm, and the supervision expression is:
[0068]
[0069] Wherein, Denote calculating the expected value for the noise true value ∈ t , time step t, and initial motion x0; ∈ θ Denote the predicted motion noise; C(stick) denotes the input text vector; C(stick) denotes the input stick figure vector.
[0070] On the other hand, the present invention also provides a method for generating a human motion sequence, which is performed in the human motion sequence generation model obtained by the training method of the human motion sequence generation model described in any one of the above, and the method includes the following steps:
[0071] Input a randomly generated noise motion sequence and the stick figure and text description data about the target action into the human motion sequence generation model, and output the human motion sequence of the target action.
[0072] Specifically, the noise motion sequence is obtained by adding Gaussian noise through a diffusion forward process for a series of action sets arranged in chronological order. After selecting the target action therefrom, the stick figure of the target action and the text description data for verbally describing the action represented by the stick figure are input into the trained human motion sequence generation model. The target action is a 3D human pose randomly selected near the start, middle, and end positions of the noise motion sequence, and the human motion sequence is output.
[0073] On the other hand, the present invention also provides a human motion sequence generation system, including a processor, a memory, and a computer program / instructions stored on the memory. It is characterized in that the processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method described in any one of the above.
[0074] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored. It is characterized in that when the computer program / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0075] On the other hand, the present invention also provides a computer program product, including a computer program / instructions. It is characterized in that when the computer program / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0076] The present invention will be described below in conjunction with a specific embodiment:
[0077] Figure 2 It is a structural schematic diagram of a training method for a human motion sequence generation model according to an embodiment of the present invention. The present invention proposes a human motion sequence generation method based on stick figure conditions. By introducing stick figures as conditions for controlling human motion details and combining text description data, the dependence on complex texts is reduced. At the same time, an efficient multi-condition fusion module and a dynamic supervision strategy are designed to improve the multi-condition fusion performance, optimize the generation process, achieve more natural and accurate human motion generation, and provide a more convenient and fast method for creators to obtain ideal human motions. The human motion sequence generation model (StickMotion) takes stick figures and text description data as input methods, and users can provide any combination of stick figures and text description data at the starting, intermediate, and ending positions; the human motion sequence generation model utilizes the forward and backward processes of the diffusion model, combines the multi-condition fusion module to fuse different input conditions, and generates 3D human motion sequences that meet the conditions.
[0078] 1. Stick figure generation algorithm: Since there is a lack of hand-drawn stick figures in the existing training dataset, a stick figure generation algorithm can be used to automatically generate stick figures. An example of the stick figure generation algorithm is the simple genetic algorithm (SGA: Simple Genetic Algorithm). This algorithm automatically generates stick figures based on the 3D coordinates of human joints in the existing joint coordinate dataset. The generation process considers factors such as stroke smoothness, position deviation, and scaling differences. The specific generation process includes:
[0079] (1) 3D human pose coordinate normalization: Normalize the 3D joint point coordinates using the lengths of the arms and legs as the denominator, and stipulate that the 2D normalized coordinates are obtained by observing the human pose from the front, with the line of sight perpendicular to the pelvic plane of the pose; connect the 2D normalized coordinates to obtain the limb lines and the spine line.
[0080] (2) Head line: First, generate a standard circle, truncate it at a random position from the circle and extend it randomly as the intersection of the hand-drawn circle.
[0081] (3) After obtaining the head line, spine line, and limb lines, generate a unified hand-drawn style for the lines. The detailed steps are as follows: 1) Interpolate the lines composed of fewer coordinate points by N times. When the arm is composed of 3 coordinate points, increase the number of coordinate points representing the line trajectory, including but not limited to interpolating N coordinate points between the shoulder and the elbow, to enhance the line granularity for receiving the jitter noise of hand-drawing next. 2) Randomly generate a noise array with the same shape as the line, distributed uniformly, and accumulate this array from the beginning to simulate the large errors of hand-drawing. 3) Add the accumulated noise array to the interpolated line and smooth the noisy line using Gaussian filtering. 4) Align the processed line with the original line through affine transformation. 5) Finally, generate a small uniform distribution noise array again and add it to the line to simulate the small jitter of hand-drawing.
[0082] (4) Line combination and sampling: After obtaining the hand-drawn style lines of the head, spine, and limbs, it is necessary to combine the lines and introduce certain perturbations to the combined positions to represent hand-drawing mistakes, and finally perform equidistant sampling. The specific steps include: 1) Combine the head line and the spine line reasonably according to the direction of the neck vector, and connect the arm lines and leg lines to the head and tail of the spine line respectively. 2) Subsequently, perform an overall position offset perturbation on each combined line separately, that is, generate two noise values and add them to the horizontal and vertical coordinates of the line to randomly offset the line, representing the hand-drawing positioning error. 3) Perform equidistant sampling on each line, that is, sample the line at a certain distance interval to unify the number of coordinate points of each line.
[0083] (5) Information encoding: To balance user convenience and computational efficiency, the user only needs to draw six one-stroke lines representing the head, torso, and limbs. After each line is encoded separately, through simple interaction with the Transformer encoder, obtain the embedded representation of the stick figure, reducing computational requirements and improving recognition accuracy.
[0084] 2. Diffusion-based motion generation: Use the diffusion model to decompose the task into forward and backward processes.
[0085] 2.1 Forward process: Gradually add Gaussian noise ∈ t ~N(0, I) to the actual motion x0 from t = 0 to t = T. Here, t represents the time step, gradually advancing from the initial time 0 to the final time T; x0 is the initial actual motion sequence; ∈ t is the Gaussian noise added at time step t, following a normal distribution with a mean of 0 and a covariance matrix of the identity matrix I. The weight of adding noise is an increasing sequence β t . The probability distribution formula of the forward process is where q(x 1:T |x0) represents the probability from x1 to x given the initial motion x0T The joint probability of this series of noisy motion sequences; is the product symbol, indicating the multiplication of the probabilities q(x t |x t-1 ) from each time step t = 1 to t = T. And This indicates that given the previous motion x t-1 , the current motion x t follows a normal distribution with mean and variance (1 - α t )I. Here, α t is related to the weight β t , and α t = 1 - β t . This formula is equivalent to where is the product of all α s from 1 to time step t. This formula further explains the relationship between the noisy motion x, the initial motion x0, and the noise ∈ t , and is used for model training.
[0086] 2.2 Reverse process: Use the trained model to gradually remove the noise X T to obtain the true motion sequence. The formula is where x t-1 is the motion sequence of the previous time step to be restored; x t is the noisy motion sequence at the current time step; α t has the same meaning as in the forward process; ∈ θ is the noise predicted by the model and output by the trained model; σ t is the weighted Gaussian noise, which is used to appropriately introduce randomness during the denoising process to assist the model in better restoring the true motion sequence.
[0087] 2.3 Conditional mixing: To balance the influence of the stick figure and text conditions on the final generation result during the reverse process, classifier-free diffusion guidance is adopted during the forward process (training process). The supervision formula is where represents the expectation. Here, it is the expectation with respect to the noise ∈ t , the time step t, and the initial motion x0; ∥·∥ 2 is to find the squared norm; ∈ t is the true noise; ∈ θ (x t , t, L, C(stick), C(text)) is the model based on the current noisy motion sequence x t, the time step t, the length L of the generated action sequence, the stickman condition C(stick), and the text condition C(text) are used to predict the motion noise. During training Here and represent the probabilities of the stickman condition and the text condition occurring respectively, both set to 0.7.
[0088] During the reverse process (generation process), by adjusting the conditional mixing weights (w1, w2, w3, w4) such that w1 + w2 + w3 + w4 = 1, the generated results are biased towards different conditional combinations, i.e., Four conditions. In the initial stage (t ∈ [T, T / 10]), the conditional mixing formula is (w1 = w; w4 = 1 - 2·w), where w and are custom weight parameters. Through such weight allocation, the influence of different conditions is balanced in the initial stage; in the final stage (t ∈ [T / 10, 0]), (w1 = 1; w2, w3, w4 = 0) is set to further optimize the results, i.e., in the final stage, the results are mainly optimized depending on the conditions (sitck, text).
[0089] 3. Structure diagram of the 3D human motion sequence generation method.
[0090] 3.1 Input encoding: The input data includes the noise motion sequence (Motion), the stickman (Stickman), and the text description data (Text), which are encoded into vectors with shapes of [L m , E], [L s , E], and [L t , E] respectively. Among them, L m , L s , L t represent the sequence lengths of each input respectively, and E represents the dimension of the input. The noise motion sequence is encoded using a simple linear layer; the text description data is encoded using CLIP ViT - B / 32 (with 154M parameters); the stickman is encoded through a standard Transformer encoder and decoded back to 3D poses to achieve loss supervision. The above text encoder and stickman encoder will be frozen during training, i.e., they do not participate in the training process of the human motion sequence generation model.
[0091] 3.2 Multi-Condition Module (MCM): The multi-condition module is used to fuse conditions. Through the condition fusion module, the stickman vector and the text vector are integrated into the motion vector in the latent space, and then re-encoded by the latent encoder to further fuse information. The data is divided into four parts (B1, B2, B3, B4) along the batch dimension, representing four conditional combinations of the text vector and the stickman vector; two standard Transformer decoder layers: the Feature Decoder respectively considers the text input and the stickman input, and obtains a new motion vector by adding it to the corresponding motion vector, reducing the computational complexity and improving the performance. At the same time, an efficient attention mechanism is adopted in the feature decoder and the latent encoder to further reduce the computational amount.
[0092] 3.3 Output: The human motion sequence generation model consists of multiple multi-condition fusion modules and outputs the predicted noise ∈ of the noisy motion sequence θ and the position index score of the input stickman The index score is used for training supervision and user interaction. During training, it supervises the distance between the dynamically assigned pose and the stickman, and during generation, it indicates the position of the stickman in the motion sequence, helping the user decide whether to adjust the generation result. The predicted noise ∈ θ is used for the reverse process of the diffusion model, that is, removing the noise from the input noise to restore the final human motion sequence. The index score is used for training supervision and user interaction. During training, it supervises the distance between the dynamically assigned pose and the stickman, and during generation, it indicates the position of the stickman in the motion sequence, helping the user decide whether to adjust the generation result.
[0093] 3.4 Model Training: The overall training process of the human motion sequence generation model includes obtaining 3D human motion sequences from the training dataset, selecting 3 3D human postures near the start, middle, and end positions of the sequence and converting them into 2D hand-drawn style stickmen through the stickman generation algorithm as inputs. At this time, the stickman, the noisy motion sequence after adding noise, and the corresponding text description data are used as three inputs and encoded and input into the human motion sequence generation model, and finally the predicted noise ∈ θ and the position index score of the input stickman position
[0094] 4. Dynamic Supervision. To solve the problem of determining the stickman index in the generated motion sequence, the user only needs to specify the approximate position of the stickman at the start, middle, or end, and the network automatically adjusts the index dynamically near the specified position to optimize the naturalness of the generation and the fit with the text description data. By randomly sampling human postures at the start, middle, and end positions of the motion sequence and creating masks to simulate random user inputs, the loss function is divided into the index loss function Motion loss function and total loss function The formulas are respectively where is the predicted index score of the l-th frame stickman in the generated motion sequence, is the motion sequence generated under the (stickman, text) and (stickman, empty) condition combinations, is the predicted pose of the l-th frame under all condition combinations, x l is the pose of the l-th frame of the real motion sequence, x i is the pose of the stickman generated from the real motion sequence through the stickman generation algorithm.
[0095] In summary, the present invention provides a method, a generation method and a system for training a generation model of a human motion sequence. The training method includes: obtaining a training sample set of each sample including a human motion sequence for a sample action and its text description data; generating a stick figure corresponding to the human motion sequence based on a preset stick figure generation algorithm; training an initial human motion sequence generation model using the training sample set; the diffusion forward noise addition module receives the human motion sequence and gradually adds noise to obtain a noisy motion sequence, which is encoded into a motion vector by the motion encoder, the stick figure is encoded into a stick figure vector by the stick figure encoder, and the text description data is encoded into a text vector by the text encoder. The multi-layer multi-condition fusion module at least batches and combines the stick figure vector and the text vector to obtain multiple combined input vectors and integrates them into the motion vector. The first combined input vector at least includes the text vector, the second combined input vector at least includes the text vector and the stick figure vector, the third combined input vector at least includes the stick figure vector, and the fourth combined input vector does not include the text vector and the stick figure vector. The first feature decoder at least receives the first combined input vector and the second combined input vector, and the second feature decoder at least receives the third combined input vector and the fourth combined input vector, and after integrating them into the motion vector through the attention mechanism, obtains a text-influenced motion offset and a stick figure-influenced motion offset. After adding them to the motion vector, a target motion vector is obtained, which is re-encoded and input into the backbone network to output a predicted motion noise and a predicted value of the position index of the stick figure; the diffusion backward denoising module gradually removes the predicted motion noise to obtain a predicted value of the human motion sequence; a motion loss function is constructed through the predicted value of the human motion sequence and the true value of the human motion sequence, an index loss function is constructed through the predicted value of the position index and the true value of the position index, and the motion loss function and the index loss function are fused to obtain a total loss function. The motion encoder, the multi-layer multi-condition fusion module, the backbone network and the diffusion backward denoising module are updated with parameters with the goal of minimizing the total loss function, and the human motion sequence generation model is constructed and obtained.
[0096] Correspondingly, the present invention also provides a human motion sequence generation system, including a processor, a memory and a computer program / instructions stored on the memory. It is characterized in that the processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method described above.
[0097] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing edge computing server deployment method are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.
[0098] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0099] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0100] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0101] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A training method for a human motion sequence generation model, characterized in that The method includes the following steps: Obtain a training sample set, where the training sample set includes multiple samples, and each sample contains a human motion sequence for the sample action and its text description data; generate corresponding stick figures for the human motion sequences in each sample based on a preset stick figure generation algorithm; the human motion sequence is labeled with the true value of the human motion sequence; the stick figure is labeled with the true value of the position index; Train an initial human motion sequence generation model using the training sample set; the initial human motion sequence generation model includes a diffusion forward noise addition module, a motion encoder, a stick figure encoder, a text encoder, a multi-layer multi-condition fusion module, a backbone network, and a diffusion reverse denoising module; the multi-layer multi-condition fusion module layer includes a first feature decoder, a second feature decoder, and a latent encoder; the diffusion forward noise addition module receives the human motion sequence and gradually adds noise to obtain a noisy motion sequence, encodes the noisy motion sequence into a motion vector through the motion encoder, encodes the stick figure into a stick figure vector through the stick figure encoder, encodes the text description data into a text vector through the text encoder, the multi-layer multi-condition fusion module at least batches and combines the stick figure vector and the text vector to obtain multiple combined input vectors and integrates them into the motion vector, the first combined input vector includes at least the text vector, the second combined input vector includes at least the text vector and the stick figure vector, the third combined input vector includes at least the stick figure vector, the fourth combined input vector does not include the text vector and the stick figure vector, the first feature decoder receives at least the first combined input vector and the second combined input vector, the second feature decoder receives at least the third combined input vector and the fourth combined input vector and obtains a text influence motion offset and a stick figure influence motion offset after integrating them into the motion vector through an attention mechanism, adds them to the motion vector to obtain a target motion vector, re-encodes through the latent encoder and inputs it into the backbone network, and outputs a predicted motion noise and a predicted value of the position index of the stick figure; gradually remove the predicted motion noise through the diffusion reverse denoising module to obtain a predicted value of the human motion sequence; Construct a motion loss function through the predicted value of the human motion sequence and the true value of the human motion sequence, construct an index loss function through the predicted value of the position index and the true value of the position index, fuse the motion loss function and the index loss function to obtain a total loss function, and update the parameters of the motion encoder, the multi-layer multi-condition fusion module, the backbone network, and the diffusion reverse denoising module with the goal of minimizing the total loss function and construct the human motion sequence generation model.
2. The training method of the human motion sequence generation model according to claim 1, wherein The index loss function The motion loss function and the total loss function are expressed as follows: Where M represents a preset weight coefficient and L represents the sequence length, represents the predicted value of the position index of the stick figure in the l-th frame of the predicted value of the human motion sequence, represents the human motion sequence generated under the second combined input vector and the third combined input vector, represents the human motion sequences generated under all conditional combinations, x l represents the action of the l-th frame in the true value of the human motion sequence, x i represents the action of the stick figure generated from the true motion sequence through the stick figure generation algorithm.
3. The training method of the human motion sequence generation model according to claim 2, characterized in that The method further includes supervising the motion loss using a classifier-free diffusion guidance algorithm, and the supervision expression is: Among them, represents calculating the expected value for the noise true value ∈ t , time step t, and initial motion x0; ∈ θ represents the predicted motion noise; C(stick) represents the input text vector; C(stick) represents the input stick figure vector.
4. The training method of the human motion sequence generation model according to claim 1, characterized in that The expression of the predicted motion noise when gradually removing the predicted motion noise through the diffusion reverse denoising module to obtain a predicted value of the human motion sequence is: When \(t\in[T, T / 10]\), \(w1 = w\); w4 = 1 - 2\cdot w; When \(t\in[T / 10,0]\), \(w1 = 1\); \(w2, w3, w4 = 0\); Among them, stick represents a stick figure, text represents text, and w and are custom weight parameters.
5. The training method of the human motion sequence generation model according to claim 1, characterized in that, The process of generating the corresponding stick figure for the human motion sequence in each sample based on a preset stick figure generation algorithm includes: Normalize the 3D joint point coordinates in the preset joint coordinate dataset using the lengths of the human arms and legs as the denominator, and obtain 2D normalized coordinates through a front view pose with the line of sight perpendicular to the target human pelvis plane; Generate a standard circle, truncate it at a random position on the circle, and then randomly extend it to obtain the head line. Connect each of the 2D normalized coordinates to obtain the limb lines and the spine line; After adding a first preset number of interpolation coordinate points representing the line trajectory to the head line, the limb lines, and the spine line, accumulate a randomly generated noise array with the same shape as the line and evenly distributed and add it to the line to obtain a noisy line. Smooth the noisy line through Gaussian filtering and then align it with the original head line, the limb lines, and the spine line through affine transformation, and add a generated small amplitude evenly distributed noise array to obtain a hand-drawn style line; After obtaining the hand-drawn style head line, the hand-drawn style spine line, and the hand-drawn style limb lines, reasonably combine the hand-drawn style head line and the hand-drawn style spine line according to the neck direction, connect the hand-drawn style arm lines to the head of the hand-drawn style spine line, and connect the hand-drawn style leg lines to the tail of the hand-drawn style spine line to obtain the initial stick figure. Perturb the overall position of each line of the initial stick figure by generating two noise values and adding them to the horizontal and vertical coordinates of the line, and perform equidistant sampling on each line of the initial stick figure at a certain distance interval to unify the number of coordinate points of each line to obtain the stick figure.
6. The training method of the human motion sequence generation model according to claim 1, wherein Encode the noise motion sequence into a motion vector through the motion encoder, encode the stick figure into a stick figure vector through the stick figure encoder, and encode the text description data into a text vector through the text encoder, including: The motion encoder uses a simple linear layer to encode the noise motion sequence to obtain a motion vector; The stick figure encoder uses a transformation encoder to encode each line of the stick figure to obtain a stick figure vector; The text encoder uses CLIP ViT-B / 32 to encode the text description data to obtain a text vector.
7. A method for generating a human motion sequence, characterized in that The method is performed in the human motion sequence generation model obtained by the training method of the human motion sequence generation model according to any one of claims 1 to 6. This method includes the following steps: Input a randomly generated noise motion sequence, the stick figure of the target action, and the text description data into the human motion sequence generation model, and output the human motion sequence of the target action.
8. A human motion sequence generation system, comprising a processor, a memory, and computer programs / instructions stored on the memory, characterized in that, The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the system implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for generating action sequence for driving virtual character to move according to text
CN116883555A
Action generation method and device based on diffusion model, equipment and medium
CN118037907A
Posture-guided figure image synthesis method and system based on multi-condition diffusion model
CN118446887A
Potential space human motion generation method based on diffusion model
CN119169205A
Method and device for generating video clip from text description and sequence of key points synthesized by diffusion model
RU2823216C1
Cited By
Human body action editing method and device based on sketch guidance, terminal and medium
CN122244401A