A method and apparatus for generating human body movements based on text description
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请提供一种基于文本描述的肢体协调人体动作生成方法及装置,以解决相关技术中缺少对人体不同局部部分的理解,泛化能力较差,且难以应对复杂文本描述的生成等问题
[0047]本申请实施例可以基于描述文本将人体全身运动拆分为多部分运动,对每部分运动进行生成并协调得到协调后的每部分运动,进一步组合每部分运动得到全身运动,从而能够处理复杂或抽象的文本,生成符合语义描述且协调自然的人体运动,并且具有更强的泛化能力。由此,解决了相关技术中缺少对人体不同局部部分的理解,泛化能力较差,且难以应对复杂文本描述的生成等技术问题。
Smart Images

Figure CN118012984B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human motion synthesis technology, and in particular to a method for generating human motion coordination based on text description. Background Technology
[0002] The field of human motion synthesis encompasses various types of tasks, which can be categorized into unconditional motion generation and conditional motion generation based on the input signal. Among these, text-driven motion generation aims to generate human actions based on input text descriptions.
[0003] While related technologies can easily generate human movements based on text, there remains a problem of discrepancies between the generated human motion and the input text description when dealing with complex textual descriptions. Current mainstream methods treat human motion as a whole, lacking understanding of different local parts, poor generalization ability, and difficulty in generating complex textual descriptions. Although SCA divides the skeleton into upper and lower body, its completely independent structure prevents the generated upper and lower body movements from coordinating, and it still lacks a more granular understanding of different local parts. Summary of the Invention
[0004] This application provides a method and apparatus for generating human body movements based on text description, in order to solve the problems in related technologies such as lack of understanding of different parts of the human body, poor generalization ability, and difficulty in generating complex text descriptions.
[0005] The first aspect of this application provides a method for generating human body movements based on text description, including the following steps: obtaining a descriptive text of the whole body movement and discretizing the whole body movement into partial movements; generating and coordinating the encoding of each partial movement based on the descriptive text; decoding the encoding of each partial movement and combining the decoded partial movements to obtain the whole body movement of the human body that conforms to the descriptive text.
[0006] Optionally, generating and coordinating the encoding of each part of the motion based on the descriptive text includes: encoding each part of the motion using a pre-trained encoder, and coordinating the encoding of each part of the motion to obtain the coordinated encoding of each part of the motion.
[0007] Optionally, the encoder training process includes: independently encoding each part of the motion, wherein the independent encoding formula is:
[0008]
[0009] Among them, E i For encoding the i-th part of the motion, P i For the motion of the i-th part, For the motion of the i-th part in the l-th frame, L is the number of frames of the output motion, r is the encoder downsampling rate, and C is the number of frames of the input video.
[0010] Based on the learnable codebook, the code for each part of the motion is discretized into multiple discrete motion codes, where the discretization formula is:
[0011]
[0012] Among them, Q i This is the motion encoding for the i-th part after discretization. Let v represent the motion of the i-th part after discretization in the l-th frame, v be the learnable encoding dictionary, and k be the index in the encoding dictionary. Let L be the index of the i-th part in the dictionary v corresponding to the motion in the l-th frame, and L be the number of frames of the output motion.
[0013] Optionally, motion coordination is performed on the encoding of each part of the motion, including: obtaining the conditional distribution of the permissible limb parts coordinating with each other for each part of the motion; and coordinating according to the conditional distribution of each part of the motion to obtain the coordinated motion encoding of each part.
[0014] Alternatively, the formula for the conditional distribution is:
[0015]
[0016]
[0017] Among them, K i This is the encoding of the i-th part. Let L be the number of frames in the output motion, t be the input descriptive text, S be the number of parts, and p(K) be the motion encoding of the i-th part in the h-th frame. i |t) is K i The probability under condition t.
[0018] Alternatively, the decoding formula is:
[0019]
[0020] Among them, Q i For the discretized o-th part of the motion encoding, Decoder i To decode the i-th part of the motion, This represents the i-th part of the motion after decoding.
[0021] Optionally, after obtaining the full-body human motion that matches the descriptive text for each part of the motion after combination decoding, the method further includes: training and optimizing the full-body human motion, wherein the training and optimization objective is:
[0022]
[0023] in, To optimize the training objective, E M,t~p(M,t) For maximum likelihood estimation, p(M|t) is the probability of M under condition t, and p(K) is the probability of M under condition t. i |t) is K i The probability K under condition t i Let M be the encoding of the i-th part, M be the encoding of the whole, and t be the input text description.
[0024] A second aspect of this application provides a limb coordination human motion generation device based on text description, comprising: an acquisition module for acquiring descriptive text of whole-body motion and discretizing the whole-body motion into partial motions; a generation module for generating and coordinating the encoding of each partial motion based on the descriptive text; and a combination module for decoding the encoding of each partial motion and combining the decoded partial motions to obtain whole-body motion of the human body that conforms to the descriptive text.
[0025] Optionally, the generation module is further configured to: encode each part of the motion using a pre-trained encoder, and coordinate the encoding of each part of the motion to obtain a coordinated encoding of each part of the motion.
[0026] Optionally, the encoder training process includes:
[0027] Each motion component is encoded independently, and the independent encoding formula is as follows:
[0028]
[0029] Among them, W i For encoding the i-th part of the motion, P i For the motion of the i-th part, For the motion of the i-th part in the l-th frame, L is the number of frames of the output motion, r is the encoder downsampling rate, and C is the number of frames of the input video.
[0030] Based on the learnable codebook, the code for each part of the motion is discretized into multiple discrete motion codes, where the discretization formula is:
[0031]
[0032] Among them, Q i This is the motion encoding for the i-th part after discretization. Let v represent the motion of the i-th part after discretization in the l-th frame, v be the learnable encoding dictionary, and k be the index in the encoding dictionary. Let L be the index of the i-th part in the dictionary v corresponding to the motion in the l-th frame, and L be the number of frames of the output motion.
[0033] Optionally, the generation module is further configured to: obtain the conditional distribution of permissible limb parts coordinating with each other for each part of the movement; and coordinate each part of the movement according to the conditional distribution of each part of the movement to obtain the coordinated movement coding sequence of each part.
[0034] Alternatively, the formula for the conditional distribution is:
[0035]
[0036]
[0037] Among them, K i This is the encoding of the i-th part. Let L be the number of frames in the output motion, t be the input descriptive text, S be the number of parts, and p(K) be the motion encoding of the i-th part in the h-th frame. i |t) is K i The probability under condition t.
[0038] Alternatively, the decoding formula is:
[0039]
[0040] Among them, Q i For the discretized i-th part of the motion encoding, Decoder i To decode the i-th part of the motion, This represents the i-th part of the motion after decoding.
[0041] Optionally, it also includes: an optimization module, used to train and optimize the whole-body human motion after each part of the motion is combined and decoded to obtain the whole-body human motion that conforms to the description text, wherein the training and optimization objective is:
[0042]
[0043] in, To optimize the training objective, E M,t~p(m,t) For maximum likelihood estimation, p(M|t) is the probability of M under condition t, and p(K) is the probability of M under condition t. i |t) is K i The probability K under condition t i Let M be the encoding of the i-th part, M be the encoding of the whole, and t be the input text description.
[0044] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to perform the text-based limb coordination human motion generation method as described in the above embodiments.
[0045] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to perform the text-based limb coordination human motion generation method as described in the above embodiments.
[0046] Therefore, this application has at least the following beneficial effects:
[0047] This application's embodiments can break down whole-body human movement into multiple parts based on descriptive text, generate and coordinate each part to obtain coordinated movement, and further combine these parts to obtain whole-body movement. This allows for the processing of complex or abstract text, generating semantically accurate and naturally coordinated human movements, and exhibiting stronger generalization capabilities. Thus, it solves the technical problems in related technologies, such as a lack of understanding of different parts of the human body, poor generalization ability, and difficulty in generating complex text descriptions.
[0048] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0049] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0050] Figure 1 This is a schematic diagram illustrating the generation of human motion in related technologies;
[0051] Figure 2 This is a flowchart of a text-based method for generating human body movements based on descriptions, according to an embodiment of this application.
[0052] Figure 3 This is a flowchart illustrating a partial motion discretization process according to an embodiment of this application;
[0053] Figure 4 This is a flowchart illustrating the text-based method for generating human body movements based on text descriptions, according to an embodiment of this application.
[0054] Figure 5 This is a flowchart illustrating the method for generating human body movements in accordance with embodiments of this application.
[0055] Figure 6 This is a schematic diagram of motion coding for each part according to an embodiment of this application;
[0056] Figure 7 This is a schematic diagram illustrating some motion coordination and decoding according to embodiments of this application;
[0057] Figure 8This is a structural diagram of the Part Coordination module provided according to an embodiment of this application;
[0058] Figure 9 This is a schematic diagram illustrating the coordinated output action for each part of the motion according to an embodiment of this application;
[0059] Figure 10 This is a schematic diagram of a text-based limb coordination human motion generation device provided in an embodiment of this application;
[0060] Figure 11 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the application. Detailed Implementation
[0061] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0062] Early approaches employed a joint latent model approach, adding a text encoder to the unconditional motion generation motion encoder to assist in motion generation. Text2Action generates short text motion and uses a recursive model for training; TEMOS is similar, but both text and motion use self-encoder structures and KL divergence constraints; T2M-GPT replaces the encoder with a transformer / GRU structure, achieving good results; MotionCLIP directly introduces the CLIP text encoder with strong zero-shot capabilities, training it according to the Language2Pose text-pose alignment method, and also rendering images, introducing a CLIP image encoder for auxiliary supervision. Later, diffusion model-based schemes emerged. MDM introduces a probability mapping-based diffusion model to achieve the generation of human motion sequences from text descriptions. A schematic diagram illustrating the generation of human motion from input text is shown below. Figure 1 As shown.
[0063] The following description, with reference to the accompanying drawings, describes a method and apparatus for generating coordinated human movements based on text descriptions according to embodiments of this application. Addressing the problems mentioned in the background section regarding the lack of understanding of different local parts, poor generalization ability, and difficulty in handling complex text descriptions in current text-based human movement generation methods, this application provides a method for generating coordinated human movements based on text descriptions. In this method, the whole-body movement is divided into multiple partial movements, and partial movements are generated and coordinated according to the input text description to generate coordinated human movements. This solves the problems in related technologies, such as the lack of understanding of different local parts of the human body, poor generalization ability for whole-body movement, and difficulty in handling complex text descriptions.
[0064] Specifically, Figure 2 This is a flowchart illustrating a method for generating human body movements based on text description, provided in an embodiment of this application.
[0065] like Figure 2 As shown, the text-based method for generating coordinated human movements includes the following steps:
[0066] In step S101, the description text of the whole body movement is obtained, and the whole body movement is discretized into partial movements.
[0067] The text describes the body's full-body movements, such as moving to the right, turning the head to the right, turning the body, moving from the right to the middle, etc.
[0068] It is understood that the embodiments of this application can divide the whole body movement into multiple parts of movement, that is, divide the whole body movement into the movement of multiple limbs, specifically into 6 parts, including the right arm, left arm, right leg, left leg, upper body (representing the spine and skull), and pelvis.
[0069] In step S102, the code for coordinating the motion of each part is generated based on the description text.
[0070] It is understood that the embodiments of this application can generate and coordinate the encoding of each part of the movement based on the descriptive text, thereby realizing the coordination between the movements of multiple parts of the human body, enabling it to handle complex text input. The encoding and coordination methods will be described in the following embodiments.
[0071] In this embodiment of the application, the encoding of each part of the motion is generated and coordinated based on the descriptive text, including: encoding each part of the motion using a pre-trained encoder, and coordinating the encoding of each part of the motion to obtain the coordinated encoding of each part of the motion.
[0072] It is understood that the embodiments of this application can use a pre-trained encoder to encode each part of the motion and coordinate each part of the motion to obtain the encoding of each coordinated part of the motion. Specifically, VQ-VAE independent encoding can be used.
[0073] In this process, motion coordination can be achieved by inserting a component coordinator (except for the first part of the motion) before the encoder of each part of the motion. The component coordinator can coordinate with the encoders of other parts of the motion to obtain the coordinated motion code for each part. The specific coordination process is described in the following embodiments.
[0074] In this embodiment of the application, the training process for the toilet includes:
[0075] Each motion component is encoded independently, and the independent encoding formula is as follows:
[0076]
[0077] Among them, E i For encoding the i-th part of the motion, P i For the motion of the i-th part, For the motion of the i-th part in the l-th frame, L is the number of frames of the output motion, r is the encoder downsampling rate, and C is the number of frames of the input video.
[0078] Based on the learnable codebook, the code for each part of the motion is discretized into multiple discrete motion codes, where the discretization formula is:
[0079]
[0080] Among them, Q i This is the motion encoding for the i-th part after discretization. Let v represent the motion of the i-th part after discretization in the l-th frame, v be the learnable encoding dictionary, and k be the index in the encoding dictionary. Let L be the index of the i-th part in the dictionary v corresponding to the motion in the l-th frame, and L be the number of frames of the output motion.
[0081] It is understood that the embodiments of this application can encode each part of the motion independently, using a separate encoder, so that each part of the motion has an independent representation space, and the encoding of each part of the motion can be discretized into multiple discrete motion codes according to the learnable codebook.
[0082] This application embodiment can use an Encoder. i Obtain the encoding of the i-th part of the motion. Based on learnable codebook E i Discretized Where J is the code number in the codebook, and index k is obtained by finding the most similar code:
[0083]
[0084] In this way, each part of the motion is discretized into multiple discrete partial motion representations {Q}. i}, i=1,…,S, the process of partially discrete motion is as follows Figure 3 As shown.
[0085] In this embodiment of the application, coordinating the encoding of each part of the movement includes: obtaining the conditional distribution of each part of the movement that allows the limb parts to coordinate with each other; and coordinating according to the conditional distribution of each part of the movement to obtain the coordinated encoding of each part of the movement.
[0086] It is understood that, according to the embodiments of this application, the conditional distribution of each movement can be further coordinated to obtain the coordinated movement code of each movement based on the conditional distribution of the permissible limb parts to coordinate with each other.
[0087] In this embodiment of the application, the formula for the conditional distribution is:
[0088]
[0089]
[0090] Among them, K i This is the encoding of the i-th part. Let L be the number of frames in the output motion, t be the input descriptive text, s be the number of parts, and p(K) be the number of frames in the output motion. i |t) is K i The probability under condition t.
[0091] In step S103, the encoding of each part of the motion is decoded, and the decoded parts of the motion are combined to obtain the whole-body motion of the human body that matches the description text.
[0092] It is understood that the embodiments of this application can decode the encoding of each part of the motion to obtain each part of the motion, and combine each part of the motion to obtain the whole body motion of the human body that conforms to the description text, thereby realizing the combination of partial motion to obtain whole body motion.
[0093] In this embodiment of the application, the decoding formula is:
[0094]
[0095] Among them, Q i For the discretized i-th part of the motion encoding, Decoder iTo decode the i-th part of the motion, This represents the i-th part of the motion after decoding.
[0096] In addition, to train VQ-VAE, a decoder is used to reconstruct the i-th part of the motion. Reconstruction loss The optimization objective for the i-th part of VQ-VAE is:
[0097]
[0098] Here, sg represents the stopping gradient operation. The first term is the reconstruction loss function, ensuring that VQ-VAE can reconstruct the original part motion from the encoding. The second term is the codebook loss function. The third term is the commitment loss, which aims to make the representation output by the encoder as close as possible to the code u contained in the codebook. The weight of this loss is controlled by the hyperparameter β.
[0099] In this embodiment of the application, after obtaining the full-body human motion that conforms to the description text after combining and decoding each part of the motion, the method further includes: training and optimizing the full-body human motion, wherein the training and optimization objective is:
[0100]
[0101] in, To optimize the training objective, E M,t~p(M,t) For maximum likelihood estimation, p(M|t) is the probability of M under condition t, and the output is the probability of M, p(K) i |t) is K i The probability K under condition t i Let M be the encoding of the i-th part, M be the encoding of the whole, and t be the input text description.
[0102] This application embodiment can model the entire text-to-motion generation process as an estimation of the distribution p(M|t) of the motion M of a given text t. Since the whole-body motion is discretized into partial motions, the conditional distribution p(K) of each partial motion can be estimated. i |t) is used to model the entire motion distribution, where,
[0103] To model the motion distribution for each part, this embodiment of the application employs an autoregressive distribution that allows for coordination between limb parts:
[0104]
[0105]
[0106] When predicting the token k for the i-th part, the prediction depends not only on all the tokens it predicts during the period from time 1 to h-1, but also on the results of predictions from all other parts within the same time frame. Similarly, after the i-th part generator (using a transformer as the generator in this embodiment) predicts token K, the current part generator and all other part generators will use that prediction to predict the next token. Finally, the motion of the entire body is understood by estimating the distribution of motion across all parts.
[0107]
[0108] This application embodiment also includes a partial coordination generator (transformer), a transformer capable of coordinating with other transformers, approximating p(K). i |t}. A part coordinator (also called a part coordination layer, except for the first transformer layer) is inserted before each transformer layer. For each token x output by the previous transformer layer, it will pass through our coordination module. The coordination module coordinates with other part motion generators to merge the tokens with tokens from other part transformers:
[0109]
[0110] y = {x j},j≠i,j∈[1,…,S].
[0111] Where LN represents the LayerNorm operation, and y represents a token from another transformer layer. (Merge Token) Then it is input into the subsequent transformer layer.
[0112] Specifically, such as Figure 4 As shown, the human motion generation method based on text description of limb coordination in this application embodiment can be divided into 5 steps: First, input text description, wherein the text description describes the movement of the human body; Second, discretize the partial limb movements, and discretize and encode the partial limb movements of each limb; Third, generate partial movements, and generate partial movements of different limbs; Fourth, integrate the whole body movements, and integrate the partial movements to obtain the complete whole body movements.
[0113] In summary, the limb coordination human motion generation method of this application embodiment has the following motion generation process: Figure 5 As shown, it mainly includes the following two stages:
[0114] Phase 1: The entire body motion is divided into multiple partial motions, and each partial motion is independently encoded using VQ-VAE. This ensures that each part motion has an independent representation space, providing prior knowledge about the concept of part motion for the next phase. The entire body motion is divided into six parts: right arm, left arm, right leg, left leg, upper body (also called trunk), and pelvis. A schematic diagram of encoding the motion of each part is shown below. Figure 6 As shown, the trunk represents the spine and skull, and the pelvis represents the movement information of the pelvic joints.
[0115] First, use Encoder i Obtain the encoding of the i-th part of the motion. Then based on the learnable codebook E i Discretized Index k is obtained by finding the most similar code:
[0116]
[0117] In this way, we discretize motion M into S discrete partial motion representations, {Q i}, i = 1, ..., S, In order to train VQ-VAE, the decoder is used to reconstruct the i-th part of the motion. Reconstruction loss The optimization objective for the i-th part of VQ-VAE is:
[0118]
[0119] Phase 2: Using generators, the overall human body movement is generated based on the input text. Unlike current methods that use a single generator to generate the entire movement, this application uses multiple smaller generators (i.e., partial motion generators) to generate movements for different limbs. Based on the partial limb motion encoding generated in Phase 1, these partial motion generators are aware of the previously constructed concept of "partial motion." However, relying solely on this separate and independent design of the partial motion generators means that they are unaware of the motion of other parts, preventing them from cooperating with other parts to complete the overall movement. Therefore, this application proposes a partial coordination module to facilitate communication between all partial motion generators. This module allows all partial motion generators to communicate with each other and collaboratively complete the overall full-body movement, thereby enabling the processing of complex text input and the generation of human body movements that conform to the semantics of the text description. Figure 7 As shown ( Figure 7PartCoordination represents partial motion coordination (part motion coordination). It uses six generators to generate code index sequences of part motions based on the input text "a man is raising his left hand while kicking something with his right foot." These six generators collaborate and coordinate with each other through the PartCoordination module. These generated index sequences are decoded by the VQ-VAE decoder into the raw representation of the part motions. These local motions combine to form the overall motion. The structure of the PartCoordination module is as follows: Figure 8 As shown, this module is designed because the movements of different parts (left arm, right arm, etc.) are not independent but interconnected. Therefore, the Part Coordination module achieves coordinated movement between different parts (left arm, right arm, etc.) by integrating features from other parts at the same stage during the generation of each independent part (left arm, right arm, etc.). A schematic diagram of the coordination module coordinating and outputting actions for partial movements is shown below. Figure 9 As shown, where, Figure 9 In this context, ParCo represents the coordination module, Text represents the input text, and G1-G7 represent the generation of different motion components.
[0120] It should be noted that text-driven motion generation is a key computer vision task, aiming to generate human motion consistent with text descriptions and demonstrate coordinated movements. It helps obtain desired motion from text descriptions, benefiting many applications in industrial scenarios such as animation, AR / VR applications, video games, autonomous driving, and robotics, while reducing the cost of data acquisition. With the improvements in this application, the results of text-driven motion generation are more accurate, and the generated human motions are closer to the semantics of the text, regardless of whether the input is long, short, or abstract. Compared to the original method, the different limb movements generated in this application are more coordinated and have stronger generalization performance, showing better performance for detailed text descriptions in specific tasks. This application further improves the practicality and scalability of text-driven motion generation in industrial applications, producing smoother and more realistic human motions, providing users with a more authentic experience in animation, AR, VR, and video games.
[0121] The method for generating human body movements based on text description proposed in the embodiments of this application can break down the whole body movement into multiple parts based on the descriptive text, generate and coordinate each part of the movement to obtain the coordinated part of the movement, and further combine each part of the movement to obtain the whole body movement. This method can handle complex or abstract text, generate human body movements that conform to semantic description and are coordinated and natural, and has a stronger generalization ability.
[0122] Next, referring to the accompanying drawings, a text-based limb coordination human motion generation device according to an embodiment of this application is described.
[0123] Figure 10 This is a block diagram of a text-based limb coordination human motion generation device according to an embodiment of this application.
[0124] like Figure 10 As shown, the text-based limb coordination human motion generation device 10 includes: an acquisition module 100, a generation module 200, and a combination module 300.
[0125] The acquisition module 100 is used to acquire the descriptive text of the whole body movement and discretize the whole body movement into partial movements; the generation module 200 is used to generate and coordinate the encoding of each part of the movement based on the descriptive text; and the combination module 300 is used to decode the encoding of each part of the movement and combine the decoded part of the movement to obtain the whole body movement of the human body that conforms to the descriptive text.
[0126] In this embodiment of the application, the generation module 200 is further configured to: encode each part of the motion using a pre-trained encoder, and coordinate the encoding of each part of the motion to obtain the coordinated encoding of each part of the motion.
[0127] In this embodiment of the application, the encoder training process includes:
[0128] Each motion component is encoded independently, and the independent encoding formula is as follows:
[0129]
[0130] Among them, W i For encoding the i-th part of the motion, P i For the motion of the i-th part, For the motion of the i-th part in the l-th frame, L is the number of frames of the output motion, r is the encoder downsampling rate, and C is the number of frames of the input video.
[0131] Based on the learnable codebook, the code for each part of the motion is discretized into multiple discrete motion codes, where the discretization formula is:
[0132]
[0133] Among them, Q i This is the motion encoding for the i-th part after discretization. Let v represent the motion of the i-th part after discretization in the l-th frame, v be the learnable encoding dictionary, and k be the index in the encoding dictionary. Let L be the index of the i-th part in the dictionary v corresponding to the motion in the l-th frame, and L be the number of frames of the output motion.
[0134] In this embodiment of the application, the generation module 200 is further configured to: obtain the conditional distribution of the permissible coordination of limb parts for each part of the movement; and coordinate the movements according to the conditional distribution of each part of the movement to obtain the coordinated movement code for each part.
[0135] In this embodiment of the application, the formula for the conditional distribution is:
[0136]
[0137]
[0138] Among them, K i This is the encoding of the i-th part. Let L be the number of frames in the output motion, t be the input descriptive text, S be the number of parts, and p(K) be the motion encoding of the i-th part in the h-th frame. i |t) is K i The probability under condition t.
[0139] In this embodiment of the application, the decoding formula is:
[0140]
[0141] Among them, Q i For the discretized i-th part of the motion encoding, Decoder i To decode the i-th part of the motion, This represents the i-th part of the motion after decoding.
[0142] In this embodiment of the application, the apparatus 10 further includes an optimization module.
[0143] The optimization module is used to train and optimize the whole-body human motion after each part of the motion is combined and decoded to obtain a whole-body human motion that matches the description text. The training and optimization objective is:
[0144]
[0145] in, To optimize the training objective, E M,t~p(M,t)For maximum likelihood estimation, p(M|t) is the probability of M under condition t, and p(K) is the probability of M under condition t. i |t) is K i The probability K under condition t i Let M be the encoding of the i-th part, M be the encoding of the whole, and t be the input text description.
[0146] It should be noted that the foregoing explanation of the embodiment of the text-based limb coordination human motion generation method also applies to the text-based limb coordination human motion generation device of this embodiment, and will not be repeated here.
[0147] According to the embodiments of this application, the limb coordination human motion generation device based on text description can break down the whole body movement into multiple parts of movement based on the description text, generate and coordinate each part of movement to obtain the coordinated part of movement, and further combine each part of movement to obtain the whole body movement. Thus, it can process complex or abstract text, generate human motion that conforms to semantic description and is coordinated and natural, and has a stronger generalization ability.
[0148] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0149] The memory 1101, the processor 1102, and the computer program stored on the memory 1101 and executable on the processor 1102.
[0150] When the processor 1102 executes the program, it implements the text-based limb coordination human motion generation method provided in the above embodiments.
[0151] Furthermore, electronic devices also include:
[0152] Communication interface 1103 is used for communication between memory 1101 and processor 1102.
[0153] The memory 1101 is used to store computer programs that can run on the processor 1102.
[0154] The memory 1101 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0155] If the memory 1101, processor 1102, and communication interface 1103 are implemented independently, then the communication interface 1103, memory 1101, and processor 1102 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0156] Optionally, in a specific implementation, if the memory 1101, processor 1102, and communication interface 1103 are integrated on a single chip, then the memory 1101, processor 1102, and communication interface 1103 can communicate with each other through an internal interface.
[0157] The processor 1102 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0158] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating limb coordination human movements based on text description.
[0159] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0160] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0161] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0162] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0163] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
Claims
1. A method for generating human body movements based on text description, characterized in that, Includes the following steps: Obtain the descriptive text of the whole-body movement and discretize the whole-body movement into partial movements; The code for generating and coordinating the motion of each part is based on the described text; The step of generating and coordinating the encoding of each part of the movement based on the description text includes: encoding each part of the movement using a pre-trained encoder, and coordinating the encoding of each part of the movement to obtain the coordinated encoding of each part of the movement; the step of coordinating the movement encoding of each part of the movement includes: obtaining the conditional distribution of the permissible limb parts coordinating with each other for each part of the movement; and coordinating according to the conditional distribution of each part of the movement to obtain the coordinated encoding of each part of the movement. The encoding of each part of the motion is decoded, and the decoded parts of the motion are combined to obtain the whole-body motion of the human body that conforms to the description text.
2. The method for generating limb coordination human movements based on text description according to claim 1, characterized in that, The training process of the encoder includes: Each motion component is independently encoded, wherein the independent encoding formula is: ; in, For the first Encoding of some motion, For the first Partial movement, For the first Part in Frame motion, , The output frame rate is r, where r is the encoder downsampling rate. The number of frames in the input video; Based on the learnable codebook, the code for each part of the motion is discretized into multiple discrete motion codes, wherein the discretization formula is: ]; in, For the discretized first... Partial motion coding, For the discretized first... Part in Frame motion, For a learnable encoding dictionary, This is the sequence number in the encoding dictionary. For the first Part in Dictionary of frame motion The serial number in This is the number of frames output for the motion.
3. The method for generating limb coordination human movements based on text description according to claim 1, characterized in that, The formula for the conditional distribution is: ; in, For the first Partial encoding, For the first Part in Motion coding of frames, To output the number of frames in the motion, For the input description text, The number of parts to be divided. for The probability under condition t.
4. The method for generating limb coordination human movements based on text description according to claim 1, characterized in that, The decoding formula is: ; in, For the discretized first... Partial motion coding, To decode the first Partial movement, For the decoded first Partial movement.
5. The method for generating limb coordination human movements based on text description according to claim 1, characterized in that, After each part of the motion, after being combined and decoded, yields a full-body human motion that matches the described text, the process further includes: The training and optimization of the human body's whole-body movements are performed, wherein the training and optimization objective is: ; in, To optimize the training objective, For maximum likelihood estimation, for Under conditions The probability of the following for Under conditions The probability of the following For the first Partial encoding, For the overall encoding, This is the text description you input.
6. A device for generating human body movements based on text description, characterized in that, include: The acquisition module is used to acquire descriptive text of the whole-body movement and discretize the whole-body movement into partial movements; A generation module is used to generate and coordinate the codes for each part of the motion based on the description text; The step of generating and coordinating the encoding of each part of the movement based on the description text includes: encoding each part of the movement using a pre-trained encoder, and coordinating the encoding of each part of the movement to obtain the coordinated encoding of each part of the movement; the step of coordinating the movement encoding of each part of the movement includes: obtaining the conditional distribution of the permissible limb parts coordinating with each other for each part of the movement; and coordinating according to the conditional distribution of each part of the movement to obtain the coordinated encoding of each part of the movement. The combination module is used to decode the encoded sequence of each part of the motion, and combine the decoded parts of the motion to obtain the whole-body motion of the human body that conforms to the description text.
7. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the text-based method for generating limb-coordinated human movements as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the text-based method for generating limb-coordinated human movements as described in any one of claims 1-5.
Citation Information
Patent Citations
ZigBee coordinator
CN109257394A
Image figure behavior description generation method based on multi-stage image context coding and decoding
CN113449801A