Text-based virtual human model driving method and device and computer device
By using a text-driven virtual human model approach, pre-trained models are used to generate and filter pose frames, solving the hardware device dependency problem in existing technologies and realizing a simplified virtual human model driving process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2023-03-23
- Publication Date
- 2026-04-17
AI Technical Summary
Existing methods for driving virtual human models require the pre-deployment of hardware devices with high accuracy in acquiring motion information or the pre-preparation of corresponding motion video materials, making it difficult to drive virtual human models conveniently and quickly.
By acquiring the text to be processed, a first action sequence is generated based on the text. Target pose frames are then selected and integrated to drive a preset virtual human model. The similarity between the pose frames and the text is calculated using a pre-trained action sequence generation model and a CLIP model to generate a target action sequence and drive the virtual human model.
It enables simple and quick driving of virtual human models without the need for high-precision hardware equipment and video materials, improving the convenience and efficiency of the driving process.
Smart Images

Figure CN116361512B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology in artificial intelligence, and more specifically, to a text-based virtual human model driving method, apparatus, and computer device. Background Technology
[0002] Currently, virtual human technology is applied in many fields (such as investment education in finance and online video teaching of pathology in medicine). The generation and driving of the virtual human model are the two most crucial aspects for the realization of a virtual human product. Modeling generates the virtual human's physical body, while the driving method gives the virtual human its soul. Current mainstream driving methods include vision-based methods and motion capture device-based methods. For example, vision-based methods rely on cameras or video to acquire real-world human motion data and map this data frame by frame onto the model, thus causing the model to move. Motion capture device-based methods use motion capture equipment (such as motion capture suits) to directly bind the wearer's motion data to the model's skeleton, thereby driving the model. As can be seen, mainstream driving methods require pre-deployed hardware with high precision in acquiring motion information or pre-prepared motion video materials, making it difficult to conveniently and quickly drive virtual human models. Summary of the Invention
[0003] The main objective of this application is to provide a text-based virtual human model driving method, apparatus, and computer device, aiming to solve the technical problem that the process of driving virtual human models using existing virtual human driving methods is relatively complex.
[0004] To achieve the above-mentioned objectives, this application provides a text-based virtual human model driving method, comprising:
[0005] Get the text to be processed;
[0006] A first action sequence is obtained based on the text, wherein the first action sequence includes multiple first pose frames;
[0007] Based on each first pose frame, a first preset format file corresponding to each first pose frame is obtained, wherein the first preset format file includes the joint coordinates of the first pose frame.
[0008] Determine whether the alignment of the joint coordinates in each of the first preset format files with the skeletal joints of the preset virtual human model satisfies the first preset condition.
[0009] If the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model satisfies the first preset condition, then the first pose frame corresponding to the first preset file is set as the target pose frame.
[0010] Integrate all the target pose frames to obtain the target action sequence;
[0011] The preset virtual human model is driven based on the target action sequence.
[0012] In one embodiment, the step of obtaining the first action sequence based on the text includes:
[0013] The text is input into a pre-trained action sequence generation model to generate an initial action sequence, wherein the initial action sequence includes multiple initial pose frames;
[0014] The text and each initial pose frame are respectively input into the pre-trained CLIP model to obtain the similarity between each initial pose frame and the text;
[0015] Determine whether the similarity between each initial pose frame and the text satisfies the second preset condition;
[0016] If the similarity between the initial pose frame and the text meets the second preset condition, then the initial pose frame is set as the first pose frame;
[0017] Integrate all the first pose frames to obtain the first action sequence.
[0018] In one embodiment, the step of obtaining a first preset format file corresponding to each first attitude frame based on each first attitude frame includes:
[0019] Each first pose frame is input into a pre-trained pose calculation model to obtain a first preset format file corresponding to each first pose frame.
[0020] In one embodiment, the step of integrating all the target pose frames to obtain the target action sequence includes:
[0021] Determine whether there are at least two target pose frames among all the target pose frames whose similarity satisfies the third preset condition;
[0022] If so, deduplication is performed to obtain an optimized target pose frame set;
[0023] If not, then the set of all the target pose frames is set as the optimized target pose frame set;
[0024] The target attitude frames in the target attitude frame set are sorted according to a preset rule to obtain the target action sequence.
[0025] In one embodiment, prior to the step of inputting the text into the pre-trained action sequence generation model, the method further includes:
[0026] The text is preprocessed by keyword extraction.
[0027] In one embodiment, the pre-trained CLIP model includes a text encoder and an image encoder, and the similarity between the initial pose frame and the text is obtained by the following formula:
[0028] S = 1 - norm(f) p )*norm(f T ),
[0029] Where S is the similarity, f p f is the encoded value output by the image encoder for the pose frame. T The encoded value of the text output by the text encoder.
[0030] In one embodiment, the training loss of the pre-trained CLIP model is obtained by the following formula:
[0031]
[0032] Where Loss is the training loss, L is the total number of pose frames in the training sample sequence, i is the i-th pose frame in the training sample sequence, γ(i) is the regularization parameter for the i-th pose frame, and S i Let be the similarity between the i-th pose frame and the text.
[0033] This application also provides a text-based virtual human model driving device, comprising:
[0034] The text acquisition module is used to acquire the text to be processed.
[0035] The first action sequence acquisition module is used to obtain a first action sequence based on the text, wherein the first action sequence includes multiple first pose frames;
[0036] The first preset format file acquisition module is used to obtain a first preset format file corresponding to each first attitude frame based on each first attitude frame, wherein the first preset format file includes the joint coordinates of the first attitude frame.
[0037] The judgment module is used to determine whether the alignment of the joint coordinates in each of the first preset format files with the skeletal joints of the preset virtual human model meets the first preset condition.
[0038] The first execution module is used to set the first pose frame corresponding to the first preset file as the target pose frame when the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model meets the first preset condition.
[0039] The target action sequence acquisition module is used to integrate all the target pose frames to obtain the target action sequence;
[0040] The virtual human model driving module is used to drive the preset virtual human model based on the target action sequence.
[0041] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the text-based virtual human model driving method provided in any of the above embodiments.
[0042] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the text-based virtual human model driving method provided in any of the above embodiments.
[0043] This application provides a text-based virtual human model driving method, apparatus, and computer device, comprising: acquiring text to be processed; obtaining a first action sequence based on the text, wherein the first action sequence includes multiple first pose frames; obtaining a first preset format file corresponding to each first pose frame, wherein the first preset format file includes the joint coordinates of the first pose frame; determining whether the alignment of the joint coordinates in each first preset format file with the skeletal joints of a preset virtual human model satisfies a first preset condition; if the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model satisfies the first preset condition, then setting the first pose frame corresponding to the first preset file as a target pose frame; integrating all the target pose frames to obtain a target action sequence; and driving the preset virtual human model based on the target action sequence. This application drives the virtual human model by inputting text. Compared with the traditional method of driving the virtual human model by using video or real people, the virtual human model can be driven by inputting a piece of text. This eliminates the need to set up high-precision or bulky motion capture equipment (such as cameras, motion capture suits, etc.) or prepare video materials in advance, thus making the driving process of the virtual human model simpler and more convenient. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating a text-based virtual human model driving method according to an embodiment of this application;
[0045] Figure 2 This is a flowchart illustrating step S20 in a text-based virtual human model driving method according to an embodiment of this application.
[0046] Figure 3 This is a flowchart illustrating step S60 in a text-based virtual human model driving method according to an embodiment of this application.
[0047] Figure 4 This is a flowchart illustrating step S20 in a text-based virtual human model driving method according to another embodiment of this application.
[0048] Figure 5 This is a schematic diagram of the structure of a text-based virtual human model driving device according to an embodiment of this application;
[0049] Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of this application. Detailed Implementation
[0050] Virtual humans refer to beings existing in the non-physical world, created and used through computer technology such as computer graphics, image rendering, motion capture, deep learning, and speech synthesis. They are comprehensive products possessing multiple human characteristics (physical features, human performance abilities, human interaction abilities, etc.) and are also known as virtual avatars, virtual digital humans, or digital humans. The three main characteristics of virtual digital humans are the combined maturity of multiple technologies such as virtualization, NLP (Natural Language Processing), CV (Computer Vision), and speech, resulting in a high degree of human-likeness. This high degree of human-likeness brings users a sense of intimacy, care, and immersion, which is the core driving force for most users. Currently, virtual human technology is applied in many fields (such as investment education in the financial field and pathology teaching in the medical field using virtual human online video teaching). However, existing virtual human model driving methods require pre-deployed hardware devices with high accuracy in acquiring motion information or pre-prepared corresponding motion video materials, making it difficult to conveniently and quickly drive virtual human models. Therefore, it is necessary to explore new virtual human driving methods to simplify the driving process and thus conveniently and quickly drive virtual human models.
[0051] Please refer to Figure 1 This application provides a text-based virtual human model driving method, which includes steps S10-S70. The detailed description of each step of the method is as follows.
[0052] In one embodiment, the text-based virtual human model driving method includes:
[0053] S10. Obtain the text to be processed;
[0054] S20. Obtain a first action sequence based on the text, wherein the first action sequence includes multiple first pose frames;
[0055] S30. Based on each first attitude frame, obtain a first preset format file corresponding to each first attitude frame, wherein the first preset format file includes the joint coordinates of the first attitude frame.
[0056] S40. Determine whether the alignment of the joint coordinates in each of the first preset format files with the skeletal joints of the preset virtual human model satisfies the first preset condition.
[0057] S50. If the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model satisfies the first preset condition, then the first pose frame corresponding to the first preset file is set as the target pose frame.
[0058] S60. Integrate all the target pose frames to obtain the target action sequence;
[0059] S70. Drive the preset virtual human model based on the target action sequence.
[0060] As described in step S10 above, the text to be processed is first obtained. This text can be content related to investment education in the financial field, medical teaching in the medical field, etc. For example, the input text could be something like "An investment training expert is lecturing on the podium" or "A doctor is demonstrating hand rehabilitation exercises to a patient."
[0061] As described in step S20 above, a first action sequence is obtained based on the text to be processed, wherein the first action sequence includes multiple first pose frames. The action sequence can be viewed as an object encapsulating multiple actions (each action corresponding to a pose frame), and when this object is executed, the encapsulated actions are executed sequentially. For example, a hand rehabilitation training action sequence may include multiple pose frames such as hanging down, clenching a fist, straightening the arm to the chest, raising the hand to the head, and repeated up-and-down stretching. Specifically, the text to be processed can be input into a pre-trained first action sequence generation model, which generates a first action sequence corresponding to the text to be processed. It should be noted that due to differences in the composition of the action sequence generation model and the choice of training method, the same text input into different action sequence generation models may yield action sequences of different forms (character models, action expressions, etc.).
[0062] As described in step S30 above, based on each first pose frame in the first action sequence, a first preset format file (generally a BVH or FBX format file, for easy input into model animation generation software) is obtained for each first pose frame. The first preset format file includes the joint coordinates of the corresponding first pose frame. The pose frames in the first action sequence obtained in step S20 above may not necessarily be effectively combined with the preset virtual human model. To avoid the virtual human model failing to effectively match the pose frames, resulting in motion deformation or errors, the first pose frames need to be pre-screened to ensure a good match with the virtual human model. Preferably, in some embodiments, the coordinates of joints at important positions in the pose frame can be mapped to the coordinates of corresponding skeletal joints in the preset virtual human model to determine whether it can be combined with the preset virtual human model. In other embodiments, parameters such as the motion matrix in the pose frame can also be used as the basis for judgment.
[0063] As described in steps S40-S50 above, it is determined whether the alignment of the joint coordinates in each first preset format file with the skeletal joints of the preset virtual human model meets the first preset condition. If the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model meets the first preset condition, then the first pose frame corresponding to the preset file is set as the target pose frame. In this embodiment, the first preset condition can be set as the alignment of the joint coordinates of the limbs, head, and chest in the first preset format file with the skeletal joints of the limbs, head, and chest of the preset virtual human model. In other embodiments, the number of joint alignments can be appropriately increased or decreased (for example, if higher accuracy is desired, the number of joint alignments can be increased, while if lower accuracy is required, the number of joint alignments can be decreased). If the alignment of the joint coordinates in the first preset format file with the skeletal joints of the preset virtual human model meets the first preset condition, it indicates that the corresponding first pose frame has a high degree of matching with the preset virtual human model and can achieve a good matching effect. Then, the corresponding first pose frame is set as the target pose frame for subsequent virtual human model driving.
[0064] As described in step S60 above, after selecting the target posture frames (i.e., posture frames with a high degree of matching with the preset virtual human model) that meet the first preset conditions in the first action sequence, all target posture frames are integrated (e.g., the order of each posture frame is determined) to obtain the target action sequence to drive the virtual human model.
[0065] As described in step S70 above, the preset virtual human model is driven based on the obtained target action sequence, so that it completes the corresponding action output effect according to the target action sequence. Specifically, files such as bvh and fbx generated from the target action sequence that conform to the format of the pre-selected model generation software are input into the pre-selected model generation software to obtain the corresponding animated GIF or video of the virtual human model's action output effect for subsequent use.
[0066] This application provides a text-based virtual human model driving method, comprising: acquiring text to be processed; obtaining a first action sequence based on the text, wherein the first action sequence includes multiple first pose frames; obtaining a first preset format file corresponding to each first pose frame, wherein the first preset format file includes the joint coordinates of the first pose frame; determining whether the alignment of the joint coordinates in each first preset format file with the skeletal joints of a preset virtual human model satisfies a first preset condition; if the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model satisfies the first preset condition, then setting the first pose frame corresponding to the first preset file as a target pose frame; integrating all the target pose frames to obtain a target action sequence; and driving the preset virtual human model based on the target action sequence. This application drives the virtual human model by inputting text. Compared to traditional methods of driving virtual human models using video or live actors, this method only requires inputting text to complete the driving of the virtual human model, eliminating the need for high-precision or bulky motion capture equipment (such as cameras, motion capture suits, etc.) or pre-prepared video materials, thus making the driving process of the virtual human model simpler and more convenient.
[0067] In some embodiments, please refer to Figure 2 The step of obtaining the first action sequence based on the text includes:
[0068] S201. Input the text into a pre-trained action sequence generation model to generate an initial action sequence, wherein the initial action sequence includes multiple initial pose frames.
[0069] S202. Input the text and each initial pose frame into the pre-trained CLIP model to obtain the similarity between each initial pose frame and the text.
[0070] S203. Determine whether the similarity between each initial pose frame and the text satisfies the second preset condition.
[0071] S204. If the similarity between the initial pose frame and the text meets the second preset condition, then the initial pose frame is set as the first pose frame.
[0072] S205. Integrate all the first pose frames to obtain the first action sequence.
[0073] As described in step S201 above, in this embodiment, a pre-trained action sequence generation model, such as a VAE (Variational Autoencoder) model, can first generate pose frames related to the text to be processed. For example, when the text to be processed is "running man", multiple running human figures can be generated through the pre-trained action sequence generation model (the multiple running human figures do not necessarily have the same character model or action pose). The aforementioned action sequence generation model can be composed of an action encoder, a reparameter module, and an action decoder. The loss function of this model is generally set as the MSE (Mean Square Error) between the input of the action encoder and the output of the action decoder. It should be noted that the model structure and training method of the pre-trained action sequence generation model are existing technologies and will not be elaborated here. Please refer to the relevant existing technologies for details.
[0074] As described in step S202 above, the text to be processed and each initial pose frame are input into the pre-trained CLIP (Contrastive Language-Image Pre-training) model to obtain the similarity between each initial pose frame and the text to be processed. The CLIP model is an open-source pre-trained neural network model for matching images and text, which has shown outstanding performance in many task processing tasks. The CLIP model includes a text encoder and an image encoder. The CLIP model can calculate the similarity between a single person pose frame and the text description to be processed. The resulting series of pose frames can be processed appropriately to construct the desired target action sequence.
[0075] As described in steps S203-S205 above, it is determined whether the similarity between each initial pose frame in the initial action sequence and the text to be processed meets the second preset condition. If the similarity between an initial pose frame and the text to be processed meets the second preset condition, then the initial pose frame is set as the first pose frame. After obtaining all the first pose frames through the above determination process, all the first pose frames are integrated (e.g., the order of the pose frames is determined) to obtain the first action sequence. In this embodiment, the second preset condition can be set to the similarity between the initial pose frame and the text to be processed being greater than 90%. Therefore, when the similarity between an initial pose frame and the text to be processed is greater than 90%, it means that the initial pose frame basically conforms to the action pose described by the text to be processed, and then the initial pose frame is set as the first pose frame. It should be noted that in other embodiments, the similarity threshold (i.e., the requirement to meet the second preset condition) can also be reasonably set according to the actual design requirements (the level of accuracy), and is not limited here.
[0076] In some embodiments, the step of obtaining a first preset format file corresponding to each first attitude frame based on each first attitude frame includes:
[0077] S301. Input each first pose frame into the pre-trained pose calculation model to obtain the first preset format file corresponding to each first pose frame.
[0078] As described in step S301 above, each first pose frame is input into a pre-trained pose calculation model to obtain a first preset format file corresponding to each first pose frame. The first preset format file includes the coordinates of the human joints in that first pose frame. Specifically, human pose calculation is an important task in computer vision and an essential step for computers to understand human actions and behaviors. Using machine learning (deep learning) for human pose calculation has become mainstream. In actual human pose calculation, the calculation is generally transformed into the problem of determining the human joints. That is, from the obtained human skeleton (usually in image form) and based on prior knowledge, the spatial relationship between joints is determined to obtain the position coordinates of each joint. Currently, mainstream model generation software generally primarily supports BVH and FBX file formats. Therefore, a first preset format (such as BVH or FBX format) file is obtained through a pre-trained pose calculation model to adapt to mainstream model generation software. It should be noted that the model structure and training method of the pre-trained pose calculation model are also existing technologies, which will not be elaborated here. Please refer to the relevant existing technologies for details.
[0079] In some embodiments, please refer to Figure 3The step of integrating all the target pose frames to obtain the target action sequence includes:
[0080] S601. Determine whether there are at least two target pose frames among all the target pose frames whose similarity satisfies the third preset condition;
[0081] S602. If so, perform deduplication to obtain the optimized target pose frame set;
[0082] S603. If not, then set the set of all the target attitude frames as the optimized target attitude frame set;
[0083] S604. Sort the target attitude frames in the target attitude frame set according to a preset rule to obtain the target action sequence.
[0084] As described in steps S601-S604 above, determine whether there are at least two target pose frames among all target pose frames whose similarity satisfies the third preset condition; if so, perform deduplication processing to obtain an optimized target pose frame set; if not, set the set of all target pose frames as the optimized target pose frame set. In this embodiment, to avoid having multiple highly similar pose frames in the target pose frames, which would prevent the subsequently generated target action sequence from exhibiting a large range of action (i.e., the action sequence does not clearly represent the action process and has no good application value), deduplication is required. Specifically, at least two highly similar pose frames that meet the third preset condition (e.g., image similarity reaches 95%) are selected and retained, and the remaining pose frames are discarded to obtain an optimized target pose frame set. If none of the target pose frames have at least two target pose frames whose similarity meets the third preset condition, it means that all target pose frames have different action poses, and the set including all target pose frames can be directly set as the optimized target pose frame set. After determining the optimized target pose frame set, the target pose frames in the target pose frame set are sorted according to a preset rule (i.e., the order of human action poses inferred from prior knowledge, such as the human body pose at the start, small step, large step, small step, and end of running in running) to obtain the target action sequence.
[0085] In some embodiments, please refer to Figure 4 Before the step of inputting the text into the pre-trained action sequence generation model, the method further includes:
[0086] S200, Perform keyword extraction preprocessing on the text.
[0087] As described in step S200 above, in order to improve the processing efficiency of the action sequence generation model, the text to be processed is preprocessed by keyword extraction. For example, if the text to be processed is "the doctor demonstrates hand rehabilitation training movements to the patient", then after the keyword extraction preprocessing, the keyword "hand rehabilitation training" is obtained and input into the action sequence generation model, thereby improving the processing efficiency of the action sequence generation model.
[0088] In some embodiments, the pre-trained CLIP model includes a text encoder and an image encoder. The similarity between the initial pose frame and the text to be processed can be obtained by the following formula (where norm is a vector processing function that assigns length and size to vectors in vector space):
[0089] S = 1 - norm(f) p )*norm(f T ),
[0090] Where S is the similarity, f p f is the encoded value output by the image encoder for the pose frame. T The encoded value of the text output by the text encoder.
[0091] In some embodiments, the training loss of the pre-trained CLIP model is obtained by the following formula:
[0092]
[0093] Where Loss is the training loss, L is the total number of pose frames in the training sample sequence, i is the i-th pose frame in the training sample sequence, γ(i) is the regularization parameter for the i-th pose frame, and S i Let be the similarity between the i-th pose frame and the text.
[0094] The goal of model training is to minimize its training loss. The training process involves adjusting the model's parameters based on the training loss value obtained in each training iteration to ensure that it meets preset training conditions (e.g., training loss less than 2%). The regularization parameters mentioned above are set to prevent overfitting, i.e., to improve the model's generalization ability and thus enhance its versatility. The regularization parameters need to be set according to actual design requirements and in conjunction with prior knowledge.
[0095] Please refer to Figure 5 This application also provides a text-based virtual human model driving device, comprising:
[0096] Text acquisition module 501 is used to acquire the text to be processed;
[0097] The first action sequence acquisition module 502 is used to obtain a first action sequence based on the text, wherein the first action sequence includes a plurality of first pose frames;
[0098] The first preset format file acquisition module 503 is used to obtain a first preset format file corresponding to each first posture frame based on each first posture frame, wherein the first preset format file includes the joint coordinates of the first posture frame.
[0099] The judgment module 504 is used to determine whether the alignment of the joint coordinates in each of the first preset format files with the skeletal joints of the preset virtual human model meets the first preset condition.
[0100] The first execution module 505 is used to set the first pose frame corresponding to the first preset file as the target pose frame when the joint coordinates in the first preset file are aligned with the skeletal joints of the preset virtual human model, which satisfies the first preset condition.
[0101] The target action sequence acquisition module 506 is used to integrate all the target pose frames to obtain the target action sequence;
[0102] The virtual human model driving module 507 is used to drive the preset virtual human model based on the target action sequence.
[0103] In this embodiment, the text to be processed is first acquired by the text acquisition module 501. The text to be processed can be content such as investment education in the financial field or medical teaching in the medical field. For example, the input text could be something like "An investment training expert is lecturing on the podium" or "A doctor is demonstrating hand rehabilitation exercises to a patient."
[0104] In this embodiment, the first action sequence acquisition module 502 also obtains a first action sequence based on the text to be processed. The first action sequence includes multiple first pose frames. An action sequence can be viewed as an object encapsulating multiple actions (each action corresponding to a pose frame), and when this object is executed, the encapsulated actions are executed sequentially. For example, a hand rehabilitation training action sequence may include multiple pose frames such as hanging down, clenching a fist, straightening the arm to the chest, raising the hand to the head, and repeated up-and-down stretching. Specifically, the text to be processed can be input into a pre-trained first action sequence generation model, which generates a first action sequence corresponding to the text. It should be noted that due to differences in the composition of the action sequence generation model and the choice of training method, the same text input into different action sequence generation models may yield action sequences of different forms (character models, action expressions, etc.).
[0105] In this embodiment, the first preset format file acquisition module 503 also obtains a first preset format file (generally a BVH or FBX format file, for easy input into the model animation generation software) for each first pose frame in the first action sequence. The first preset format file includes the joint coordinates of the corresponding first pose frame. The pose frames in the first action sequence obtained by the first action sequence acquisition module 502 may not necessarily be effectively combined with the preset virtual human model. To avoid the virtual human model failing to effectively match the pose frames, resulting in motion deformation or errors, the first pose frames need to be pre-screened to ensure a good match with the virtual human model. Preferably, in some embodiments, the coordinates of joints at important positions in the pose frame can be mapped to the coordinates of corresponding skeletal joints in the preset virtual human model to determine whether it can be combined with the preset virtual human model. In other embodiments, parameters such as the motion matrix in the pose frame can also be used as the basis for judgment.
[0106] In this embodiment, the judgment module 504 further determines whether the alignment of the joint coordinates in each first preset format file with the skeletal joints of the preset virtual human model meets the first preset condition; and when the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model meets the first preset condition, the first execution module 505 sets the first pose frame corresponding to a preset file as the target pose frame. In this embodiment, the first preset condition can be set as the alignment of the joint coordinates of the limbs, head, and chest in the first preset format file with the skeletal joints of the limbs, head, and chest of the preset virtual human model. In other embodiments, the number of joint alignments can be appropriately increased or decreased (for example, if higher accuracy is desired, the number of joint alignments can be increased, while if lower accuracy is required, the number of joint alignments can be decreased). If the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model meets the first preset condition, it indicates that the corresponding first pose frame has a high degree of matching with the preset virtual human model and can achieve a good matching effect. Then, the corresponding first pose frame is set as the target pose frame for subsequent virtual human model driving.
[0107] In this embodiment, after selecting the target posture frames (i.e., posture frames with a high degree of matching with the preset virtual human model) that meet the first preset conditions in the first action sequence, the target action sequence acquisition module 506 integrates all the target posture frames (e.g., determines the order of each posture frame) to obtain the target action sequence to drive the virtual human model.
[0108] In this embodiment, the virtual human model driving module 507 also drives a preset virtual human model based on the target action sequence obtained above, so that it completes the corresponding action output effect according to the target action sequence. Specifically, the virtual human model driving module 507 inputs files such as bvh and fbx, which are generated from the target action sequence and conform to the format of the pre-selected model generation software, into the pre-selected model generation software to obtain the corresponding animated GIF or video of the virtual human model's action output effect for subsequent use.
[0109] In some embodiments, the first action sequence acquisition module 502 includes an initial action sequence generation unit, a similarity acquisition unit, a first judgment unit, a first execution unit, and a first action sequence acquisition unit. The initial action sequence generation unit inputs the text into a pre-trained action sequence generation model to generate an initial action sequence, wherein the initial action sequence includes multiple initial pose frames. The similarity acquisition unit inputs the text and each initial pose frame into a pre-trained CLIP model to obtain the similarity between each initial pose frame and the text. The first judgment unit determines whether the similarity between each initial pose frame and the text satisfies a second preset condition. The first execution unit sets the initial pose frame as a first pose frame when the similarity between the initial pose frame and the text satisfies the second preset condition. The first action sequence acquisition unit integrates all the first pose frames to obtain a first action sequence.
[0110] In this embodiment, the text is input into a pre-trained action sequence generation model, such as a VAE (Variational Autoencoder) model, through an initial action sequence generation unit to generate pose frames related to the text to be processed. For example, when the text to be processed is "running man", multiple running human figures can be generated by the pre-trained action sequence generation model (the multiple running human figures do not necessarily have the same character model or action pose). The aforementioned action sequence generation model can be composed of an action encoder, a reparameter module, and an action decoder. The loss function of this model is generally set as the MSE (Mean Square Error) between the input of the action encoder and the output of the action decoder. It should be noted that the model structure and training method of the pre-trained action sequence generation model are existing technologies and will not be elaborated here. Please refer to the relevant existing technologies for details.
[0111] In this embodiment, a similarity acquisition unit further inputs the text to be processed and each initial pose frame into a pre-trained CLIP (Contrastive Language-Image Pre-training) model to obtain the similarity between each initial pose frame and the text to be processed. The CLIP model is an open-source pre-trained neural network model for matching images and text, and it has demonstrated impressive performance in many task processing tasks. The CLIP model includes a text encoder and an image encoder. The CLIP model can calculate the similarity between a single person pose frame and the text description to be processed. A series of pose frames obtained from this model, after appropriate processing, can construct the desired target action sequence.
[0112] In this embodiment, the first judgment unit further judges whether the similarity between each initial pose frame in the initial action sequence and the text to be processed meets the second preset condition; and when the similarity between one of the initial pose frames and the text to be processed meets the second preset condition, the first execution unit sets the initial pose frame as the first pose frame; and after all the first pose frames are obtained through the above judgment process, the first action sequence acquisition unit integrates all the first pose frames (such as determining the order of each pose frame) to obtain the first action sequence.
[0113] In some embodiments, the first preset format file acquisition module 503 includes a first preset format file acquisition unit, which is used to input each first pose frame into a pre-trained pose calculation model to obtain a first preset format file corresponding to each first pose frame.
[0114] In this embodiment, each first pose frame is input into a pre-trained pose calculation model through a first preset format file acquisition unit, thereby obtaining a first preset format file corresponding to each first pose frame. The first preset format file includes the coordinates of the human joints in the first pose frame. Specifically, human pose calculation is an important task in computer vision and an essential step for computers to understand human actions and behaviors. Using machine learning (deep learning) for human pose calculation has become mainstream. In actual human pose calculation, the calculation is generally transformed into the problem of determining human joints. That is, from the obtained human skeleton (usually in image form) and based on prior knowledge, the spatial relationship between joints is determined to obtain the position coordinates of each joint. Currently, mainstream model generation software generally primarily supports BVH and FBX file formats. Therefore, a first preset format (such as BVH or FBX format) file is obtained through a pre-trained pose calculation model to adapt to mainstream model generation software. It should be noted that the model structure and training method of the pre-trained pose calculation model are also existing technologies, which will not be elaborated here. Please refer to the relevant existing technologies for details.
[0115] In some embodiments, the target action sequence acquisition module 506 includes a second judgment unit, a second execution unit, a third execution unit, and a target action sequence acquisition unit. The second judgment unit is used to determine whether at least two target pose frames among all the target pose frames have a similarity that satisfies a third preset condition. The second execution unit is used to perform deduplication processing to obtain an optimized target pose frame set when at least two target pose frames among all the target pose frames have a similarity that satisfies the third preset condition. The third execution unit is used to set the set of all target pose frames as the optimized target pose frame set when no at least two target pose frames among all the target pose frames have a similarity that satisfies the third preset condition. The target action sequence acquisition unit is used to sort the target pose frames in the target pose frame set according to a preset rule to obtain a target action sequence.
[0116] In this embodiment, the second judgment unit determines whether there are at least two target pose frames among all target pose frames whose similarity satisfies the third preset condition; when there are at least two target pose frames among all target pose frames whose similarity satisfies the third preset condition, the second execution unit performs deduplication processing to obtain an optimized target pose frame set; and when there are no at least two target pose frames among all target pose frames whose similarity satisfies the third preset condition, the third execution unit sets the set of all target pose frames as the optimized target pose frame set. In this embodiment, to avoid having multiple highly similar pose frames in the target pose frames, which would prevent the subsequently generated target action sequence from exhibiting a large range of action (i.e., the action sequence does not clearly represent the action process and has little application value), deduplication is required. Specifically, the second execution unit selects one of at least two highly similar pose frames that meet the third preset condition (e.g., image similarity reaches 95%) and discards the remaining pose frames to obtain an optimized target pose frame set. If none of the target pose frames have at least two target pose frames whose similarity meets the third preset condition, it means that all target pose frames have different action poses, and the third execution unit can directly set the set including all target pose frames as the optimized target pose frame set. After determining the optimized target pose frame set, the target action sequence acquisition unit sorts the target pose frames in the target pose frame set according to a preset rule (i.e., the order of human action poses inferred from prior knowledge, such as the starting, small step, large step, small step, and the human pose at the end of the running action in running) to obtain the target action sequence.
[0117] In some embodiments, the first action sequence acquisition module 502 further includes a keyword extraction unit, which is used to preprocess the text by extracting keywords. In this embodiment, in order to improve the processing efficiency of the action sequence generation model, the text to be processed is preprocessed by the keyword extraction unit to extract keywords. For example, if the text to be processed is "a doctor demonstrates hand rehabilitation training movements to a patient", then after the keyword extraction model preprocesses the text to be processed by extracting keywords, the keyword "hand rehabilitation training" is obtained and input into the action sequence generation model, thereby improving the processing efficiency of the action sequence generation model.
[0118] In some embodiments, the pre-trained CLIP model described above includes a text encoder and an image encoder. The similarity between the initial pose frame and the text to be processed can be obtained by the following formula (where norm is a vector processing function that assigns length and size to a vector in the vector space):
[0119] S = 1 - norm(f) p)*norm(f T ),
[0120] Where S is the similarity, f p f is the encoded value output by the image encoder for the pose frame. T The encoded value of the text output by the text encoder.
[0121] In some embodiments, the training loss of the pre-trained CLIP model described above is obtained by the following formula:
[0122]
[0123] Where Loss is the training loss, L is the total number of pose frames in the training sample sequence, i is the i-th pose frame in the training sample sequence, γ(i) is the regularization parameter for the i-th pose frame, and S i Let be the similarity between the i-th pose frame and the text.
[0124] The goal of model training is to minimize its training loss. The training process involves adjusting the model's parameters based on the training loss value obtained in each training iteration to ensure that it meets preset training conditions (e.g., training loss less than 2%). The regularization parameters mentioned above are set to prevent overfitting, i.e., to improve the model's generalization ability and thus enhance its versatility. The regularization parameters need to be set according to actual design requirements and in conjunction with prior knowledge.
[0125] It is understood that the components of the text-based virtual human model driving device proposed in this application can realize the function of any of the text-based virtual human model driving methods provided in any of the above embodiments, and the specific structure will not be described in detail.
[0126] Please refer to Figure 6 This application also provides a computer device whose internal structure can be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor is designed to provide computing and control capabilities. The memory includes a storage medium and internal memory. The storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the storage medium. The database stores data related to a text-based virtual human model driving method. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the text-based virtual human model driving method provided in any of the above embodiments.
[0127] This application also provides a computer-readable storage medium, which may be non-volatile or volatile, storing a computer program thereon. When the computer program is executed by a processor, it implements the text-based virtual human model driving method provided in any of the above embodiments.
[0128] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), expanded SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0129] This application provides a text-based virtual human model driving method, apparatus, and computer device, comprising: acquiring text to be processed; obtaining a first action sequence based on the text, wherein the first action sequence includes multiple first pose frames; obtaining a first preset format file corresponding to each first pose frame, wherein the first preset format file includes the joint coordinates of the first pose frame; determining whether the alignment of the joint coordinates in each first preset format file with the skeletal joints of a preset virtual human model satisfies a first preset condition; if the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model satisfies the first preset condition, then setting the first pose frame corresponding to the first preset file as a target pose frame; integrating all the target pose frames to obtain a target action sequence; and driving the preset virtual human model based on the target action sequence. This application drives the virtual human model by inputting text. Compared with the traditional method of driving the virtual human model by using video or real people, the virtual human model can be driven by inputting a piece of text. This eliminates the need to set up high-precision or bulky motion capture equipment (such as cameras, motion capture suits, etc.) or prepare video materials in advance, thus making the driving process of the virtual human model simpler and more convenient.
[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0131] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A text-based virtual human model driving method, characterized in that, include: Get the text to be processed; A first action sequence is obtained based on the text, wherein the first action sequence includes multiple first pose frames; Based on each first pose frame, a first preset format file corresponding to each first pose frame is obtained, wherein the first preset format file includes the joint coordinates of the first pose frame. Determine whether the alignment of the joint coordinates in each of the first preset format files with the skeletal joints of the preset virtual human model satisfies the first preset condition. If the alignment of the joint coordinates in the first preset format file with the skeletal joints of the preset virtual human model satisfies the first preset condition, then the first pose frame corresponding to the first preset format file is set as the target pose frame. Integrate all the target pose frames to obtain the target action sequence; The preset virtual human model is driven based on the target action sequence; The first preset condition is set to the joint coordinates of the limbs, head, and chest in the first preset format file and the skeletal joints of the limbs, head, and chest of the preset virtual human model. The step of obtaining the first action sequence based on the text includes: The text is input into a pre-trained action sequence generation model to generate an initial action sequence, wherein the initial action sequence includes multiple initial pose frames; The text and each initial pose frame are respectively input into the pre-trained CLIP model to obtain the similarity between each initial pose frame and the text; Determine whether the similarity between each initial pose frame and the text satisfies the second preset condition; If the similarity between the initial pose frame and the text meets the second preset condition, then the initial pose frame is set as the first pose frame; Integrate all the first pose frames to obtain the first action sequence; The step of obtaining a first preset format file corresponding to each first attitude frame based on each first attitude frame includes: Each first pose frame is input into a pre-trained pose calculation model to obtain a first preset format file corresponding to each first pose frame. 2.The text-based virtual human model driving method of claim 1, wherein, The step of integrating all the target pose frames to obtain the target action sequence includes: Determine whether there are at least two target pose frames among all the target pose frames whose similarity satisfies the third preset condition; If so, deduplication is performed to obtain an optimized target pose frame set; If not, then the set of all the target pose frames is set as the optimized target pose frame set; The target attitude frames in the target attitude frame set are sorted according to a preset rule to obtain the target action sequence. 3.The text-based virtual human model driving method of claim 1, wherein, Prior to the step of inputting the text into the pre-trained action sequence generation model, the method further includes: The text is preprocessed by keyword extraction. 4.The text-based virtual human model driving method of claim 1, wherein, The pre-trained CLIP model includes a text encoder and an image encoder, and the similarity between the initial pose frame and the text is obtained by the following formula: , in, For similarity, The encoded value output by the image encoder for the pose frame. The encoded value of the text output by the text encoder.
5. The text-based virtual human model driving method according to claim 4, characterized in that, The training loss of the pre-trained CLIP model is obtained by the following formula: , in, For training loss, The total number of pose frames in the training sample sequence. For the training sample sequence of the th One pose frame. For the first Regularization parameters for each pose frame. For the first The similarity between each pose frame and the text.
6. A character-based virtual human model driving apparatus, which refers to the character-based virtual human model driving method according to claim 1, characterized by, include: The text acquisition module is used to acquire the text to be processed. The first action sequence acquisition module is used to obtain a first action sequence based on the text, wherein the first action sequence includes multiple first pose frames; The first preset format file acquisition module is used to obtain a first preset format file corresponding to each first attitude frame based on each first attitude frame, wherein the first preset format file includes the joint coordinates of the first attitude frame. The judgment module is used to determine whether the alignment of the joint coordinates in each of the first preset format files with the skeletal joints of the preset virtual human model meets the first preset condition. The first execution module is used to set the first pose frame corresponding to the first preset file as the target pose frame when the alignment of the joint coordinates in the first preset file with the skeletal joints of the preset virtual human model meets the first preset condition. The target action sequence acquisition module is used to integrate all the target pose frames to obtain the target action sequence; The virtual human model driving module is used to drive the preset virtual human model based on the target action sequence.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the text-based virtual human model driving method according to any one of claims 1-5.
8. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the text-based virtual human model driving method according to any one of claims 1-5.
Citation Information
Patent Citations
Method and device for driving virtual human in real time, electronic equipment and medium
CN113689879A
Virtual human driving method and system
CN115546365A