Digital human video generation method and device

By generating behavioral description text and action sequences through text expansion models and deep learning models, and combining them with human images to generate digital human videos, this technology solves the problem of existing technologies being unable to handle diverse user commands, achieving greater flexibility and richness in action generation, and reducing the difficulty of system maintenance.

CN120807719APending Publication Date: 2025-10-17BEIJING XIAOBING YUEDONG TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510854297.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing digital human video generation methods cannot handle diverse user commands and complex user intentions. They require predefined action sequence templates and mechanical matching, which makes system maintenance difficult.

Method used

The text expansion model generates behavioral description text based on user-input behavioral keywords and prompt word templates. Based on the behavioral description text, it generates action sequences and combines them with human images to generate digital human videos. The diversity and accuracy of actions are achieved through large models and deep learning models such as MotionDiffuse and AnimateAnyone.

Benefits of technology

It enables flexible responses to diverse user commands and complex intents, reduces reliance on predefined action templates, improves the richness and fluency of action generation, and reduces the difficulty of system maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807719A_ABST
    Figure CN120807719A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a digital human video generation method and device, and the method comprises the steps: generating a behavior description text related to a behavior keyword according to the behavior keyword inputted by a user and a cue word template through a text expansion model; generating an action sequence based on actions in the behavior description text; generating a digital human video based on a human image and the action sequence; wherein the text expansion model is obtained by training on the basis of a sample behavior keyword and a label behavior description text corresponding to the sample behavior keyword. According to the method, a richer and more flexible text basis can be provided for subsequent action generation by utilizing semantic comprehension and generation capability of the text expanding and writing model, the expanded and written text not only can reserve a core intention input by a user, but also can be more diversified in expression mode, and the user experience is improved. Therefore, diversified user instructions and complex user intentions can be coped with.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a digital human video generation method and device. BACKGROUND

[0002] In the prior art, the action exhibited by the digital human is mainly controlled by a predefined template. Specifically, the prior method usually needs to provide a template of an action sequence in advance to realize the action of the finally generated digital human. Therefore, there is usually an action template library prepared in advance. When generating a digital human video, a simple keyword matching is performed on the action template library according to a text or voice instruction input by a user, and a corresponding action is selected from the action template library according to the matched keyword to be played, thereby generating a digital human video.

[0003] In the above digital human video generation method, the template of the action sequence needs to be predefined and the action matching is performed according to the keyword, and the preset action is played mechanically. The method cannot cope with diversified user instructions and complex user intentions, and the action template library needs to be maintained, which leads to difficulty in system maintenance. SUMMARY

[0004] The present application provides a digital human video generation method and device to solve the problem that the prior art digital human video generation method cannot cope with diversified user instructions and complex user intentions.

[0005] The present application provides a digital human video generation method, comprising: generating an action description text related to a behavior keyword input by a user according to the behavior keyword and a prompt word template by using a text expansion model; generating an action sequence based on an action in the action description text; generating a digital human video based on a character image and the action sequence; wherein the text expansion model is trained based on a sample behavior keyword and a label action description text corresponding to the sample behavior keyword.

[0006] According to the digital human video generation method provided by the present application, the text expansion model is a large model. The action description text related to the behavior keyword input by the user is generated according to the behavior keyword and the prompt word template by using the text expansion model, which comprises: inputting the behavior keyword into the large model, and the large model generating a target prompt word according to the prompt word template and the behavior keyword; the large model generating the action description text according to the target prompt word; The large model is obtained by fine-tuning training based on a sample behavior keyword and a label behavior description text corresponding to the sample behavior keyword.

[0007] According to the digital human video generation method provided by the application, the action sequence is generated based on the actions in the behavior description text, and the action sequence generation method comprises the following steps: Extracting a sentence including a body part keyword in the behavior description text, each of the sentences including a body part; Generating a sub-action sequence corresponding to each body part based on the action keywords in each of the sentences, the sub-action sequence including pixel coordinate values of the key points of the corresponding body part at different time points; Merging each of the sub-action sequences to generate the action sequence.

[0008] According to the digital human video generation method provided by the application, the sentence including the body part keyword in the behavior description text is extracted, each of the sentences including a body part, and the method comprises the following steps: Obtaining multiple text segments by punctuation in the behavior description text; Checking the body part keyword and the corresponding action keyword in any text segment, if the body part keyword and the corresponding action keyword are checked at the same time, the any text segment is determined as a sentence; if only the action keyword is checked in the subsequent text segment, the nearest preceding text segment in which the body part keyword is checked is determined, and the subsequent text segment and the determined sentence of the preceding text segment are merged into one sentence; If multiple sentences include the same body part keyword, the multiple sentences are merged into one sentence.

[0009] According to the digital human video generation method provided by the application, the action sequence is generated by merging each of the sub-action sequences, and the method comprises the following steps: Extracting the pixel coordinate value sequence of the key points of each body part corresponding to each of the sub-action sequences, and merging the pixel coordinate value sequences of the key points of each body part to generate the action sequence.

[0010] According to the digital human video generation method provided by the application, the digital human video is generated based on the character image and the action sequence, and the method comprises the following steps: Preprocessing the character image, the preprocessing including edge enhancement and texture extraction; Recognizing the character in the preprocessed character image, determining a character region bounding box, dividing the surrounding of the character region bounding box into multiple peripheral regions by drawing lines on each side of the character region bounding box, and expanding the peripheral regions so that the area of each peripheral region is not less than the area of the character region bounding box; Generate a digital human video based on the character image after the extended peripheral region and the action sequence.

[0011] The digital human video generation method provided by the application further comprises the following steps before generating the digital human video based on the character image after the extended peripheral region and the action sequence: Converting the key point pixel coordinates of each body part in the action sequence to key point pixel coordinates under the resolution of the digital human video to be generated to obtain the converted action sequence. Generating the digital human video based on the character image after the extended peripheral region and the action sequence comprises the following steps: Generating the digital human video based on the character image after the extended peripheral region and the converted action sequence.

[0012] The application further provides a digital human video generation device comprising: The behavior description text generation module is configured to generate a behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user by using a text expansion model. The action sequence generation module is configured to generate an action sequence based on the action in the behavior description text. The digital human generation module is configured to generate a digital human video based on a character image and the action sequence. The text expansion model is trained based on sample behavior keywords and label behavior description texts corresponding to the sample behavior keywords.

[0013] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the digital human video generation method according to any one of the above embodiments when executing the program.

[0014] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the digital human video generation method according to any one of the above embodiments.

[0015] The digital human video generation method and device provided by the application generate a behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user by using a text expansion model, generate an action sequence based on the action in the behavior description text, and generate a digital human video based on a character image and the action sequence. The semantic understanding and generation capability of the text expansion model are utilized in this embodiment to provide more abundant and flexible text basis for subsequent action generation. The expanded text not only retains the core intention input by the user but also is more diversified in expression manner, thereby being able to cope with diversified user instructions and complex user intentions. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is a flow chart of the digital human video generation method provided by the present invention.

[0018] Figure 2 It is a structural diagram of the digital human video generation device provided by the present invention.

[0019] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0021] The digital human video generation method of the embodiment of the present invention is as follows: Figure 1 As shown, it includes steps S110 to S130.

[0022] Step S110: Generate a behavior description text related to the behavior keyword using the text expansion model based on the behavior keyword and prompt word template input by the user. In this step, the text expansion model is used to expand the behavior keyword based on the behavior keyword and prompt word template input by the user to generate a behavior description text that describes the behavior keyword.

[0023] For example, if the user input is "wave," and the prompt template is "Please express [user input] with different actions," where [user input] is "wave." Based on this user input and the prompt template, the text expansion model expands the action description to "Stretch out your arm, wave gently, and look forward with a smile."

[0024] In this embodiment, the text expansion model is trained based on sample behavior keywords and the label behavior description text corresponding to the sample behavior keywords. By training the text expansion model with sample behavior keywords and corresponding label behavior description text, the text expansion model has semantic understanding ability and text generation ability for behavior keywords. Through behavior keyword expansion, more rich behavior description text is generated, and diversified action description for behavior keywords is realized to meet the user's demand for digital human action diversity.

[0025] Step S120: generating an action sequence based on the action in the behavior description text. In this step, the text semantics in the behavior description text and the action features are combined through a deep learning model to generate a pose sequence that meets the user's intention. This process involves complex semantic understanding and action mapping, which not only ensures that the generated action is natural and smooth, but also accurately reflects the user's input intention.

[0026] For example, in this step, a diffusion model (Diffusion Model) is used to capture the relationship between complex text semantics and action features. For example, the MotionDiffuse model is used to generate an action sequence from a behavior description text to a human body. The MotionDiffuse model can generate a diversified and fine-grained action sequence based on the behavior description text. The action sequence includes the pixel coordinate values of the key points of each body part at different time points, and the different pixel coordinate values of each body part at different time points reflect the action of the human body. The pixel coordinate values (row and column of each pixel) of the key points of each body part at different time points are determined based on the default resolution of the video frame. For example, the MotionDiffuse model is based on a video frame with a default resolution of 512x512, and generates an action sequence with pixel coordinate values in the range of 0-512.

[0027] Step S130: generating a digital human video based on the character image and the action sequence. In the generated digital human video, the character in the character image performs each action in the action sequence. For example, the character image and the action sequence can be input into the Animate Anyone model, and the Animate Anyone model can convert the character image into a digital human video controlled by the action sequence.

[0028] In the digital human video generation method of the embodiment, a text expansion model is used to generate behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user; an action sequence is generated based on the action in the behavior description text; and a digital human video is generated based on the character image and the action sequence. The semantic understanding and generation capability of the text expansion model can provide more abundant and flexible text basis for subsequent action generation. The expanded text can not only retain the core intention of the user input, but also be more diversified in expression, so as to cope with diversified user instructions and complex user intentions. Moreover, compared with the traditional method, the dependence on a large number of predefined action templates is reduced. With the change of business scenarios and the update of user requirements, the text expansion model can automatically adapt to new user input and action requirements through training, better meet the user's demand for digital human action diversity, make the behavior of the digital human more lively and interesting, and save the development and maintenance of manpower and material resources without frequent code modification, database maintenance and complex action scripts.

[0029] In some embodiments, the text expansion model can be a large model. Based on this, the text expansion model is used to generate behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user, which includes the following specific steps.

[0030] The behavior keyword is input into the large model, and the large model generates a target prompt word according to the prompt word template and the behavior keyword. The prompt word template is pre-stored in the large model. After receiving the behavior keyword input by the user, the large model generates a target prompt word according to the prompt word template and the behavior keyword, and expands the behavior keyword input by the user based on the target prompt word.

[0031] For example, the prompt word template is "please express [user input] with different actions" or "generate several different [user input] actions", and the like. After receiving the behavior keyword input by the user, the large model replaces "[user input]" with "behavior keyword" to generate a target prompt word.

[0032] For example, the user input is "wave to say hello", the prompt word template is "please express [user input] with different actions", and the generated target prompt word is "please express wave to say hello with different actions".

[0033] The large model generates the behavior description text according to the target prompt word. The large model restates the action keywords according to the target prompt word to increase diversity while ensuring that the basic content remains unchanged. The large model is obtained by fine-tuning training based on sample action keywords and label behavior description texts corresponding to the sample action keywords. Before fine-tuning training of the large model, a data set needs to be prepared. Specifically, texts related to actions generated by users on the current network can be collected, such as commentary texts in dance teaching videos and sports event commentary texts. These data have detailed action descriptions and certain action keywords. The action keywords can be used as sample action keywords, and the action descriptions corresponding to the action keywords can be used as labels to form a data set for fine-tuning training of the large model.

[0034] During fine-tuning training of the large model, the data set is divided into a training set and a validation set according to a certain proportion. The training set is used to fine-tune the model, and the validation set is used to monitor the performance of the model to prevent overfitting.

[0035] In each iteration, the large model adjusts its parameters according to the feedback of the training data to minimize the difference between the predicted results and the true labels. This process continues for multiple epochs until the performance of the large model on the validation set tends to be stable or reaches a preset stopping condition. The evaluation of the fine-tuning training effect of the large model can be determined by the proportion of action-related words in the generated results, the input-to-output length ratio, and whether the large model meets the requirements.

[0036] In this embodiment, the large model can be an open-source model such as Llama3.1 or Qwen2.5. The large model is fine-tuned to better meet the semantic and style requirements of a specific field. Moreover, only the large model is fine-tuned, and the training cost is lower than that of a dedicated text expansion model.

[0037] When the behavior description text contains changes in the actions of multiple body parts, the deep learning model is too complex, resulting in low accuracy of the generated action sequence, or even missing some actions. To improve the accuracy of the action sequence, in some embodiments, the action sequence is generated based on the actions in the behavior description text, including: Extracting sentences in the behavior description text that include body part keywords, each of which includes a body part.

[0038] Generating a sub-action sequence corresponding to each body part based on the action keywords in each of the sentences. The sub-action sequence includes pixel coordinate values of the key points of the corresponding body part at different time points. For example, for the head, the sub-action sequence of the head includes pixel coordinate values of the key points of the head at different time points.

[0039] merge each of the sub-action sequences to generate the action sequence.

[0040] Specifically, the behavior description text is split into multiple sentences, each of which includes only one body part. The MotionDiffuse model can be used to generate a sub-action sequence, i.e., each sentence is input into the MotionDiffuse model, and the MotionDiffuse model outputs a sub-action sequence of the corresponding body part. Since there are no more than two body parts in each sentence, the MotionDiffuse model is more efficient and can generate more accurate sub-action sequences.

[0041] In this embodiment, the behavior description text is split by body part, and each sentence includes only one body part, so that the corresponding sub-action sequence can be generated more accurately and some actions can not be missed.

[0042] In some embodiments, sentences including body part keywords in the behavior description text are extracted, each of which includes one body part, including: The text is segmented by punctuation marks in the behavior description text, and the punctuation marks are commas, semicolons, or periods. The punctuation marks are used as segmentation points to quickly segment the text and ensure that each sentence is semantically coherent.

[0043] The body part keyword and the corresponding action keyword in any text segment are checked. If the body part keyword and the corresponding action keyword are checked at the same time, the any text segment is determined to be a sentence. If only the action keyword is checked in the subsequent text segment, the nearest preceding text segment in which the body part keyword is checked is determined, and the subsequent text segment and the determined sentence of the preceding text segment are merged into one sentence. If multiple sentences include the same body part keyword, the multiple sentences are merged into one sentence.

[0044] The body part keyword and the corresponding action keyword can be checked by pre-establishing a body part keyword table, in which multiple action keywords corresponding to each body part are set. The characters in the text segment are queried in the table to check the body part keyword and the corresponding action keyword.

[0045] For example, the behavior description text is "walked up with light steps, with a smile on his face, eyes curved into crescent moons, and a light in his eyes. The body slightly leaned forward, the right hand raised high, the arm naturally curved, and the palm gently swayed."

[0046] The text can be divided into multiple text segments according to the commas and periods in the text. The body part keywords and corresponding action keywords in any text segment are checked. For the first text segment "walks near with light steps", the body part is "foot", the corresponding action keywords are "walks" and "walks", and the body part keywords and corresponding action keywords are checked. Therefore, "walks near with light steps" is a clause.

[0047] For the text segment "eyelids curve into crescent", the body part is "eyes", and the corresponding action keyword is "curve into". Therefore, "eyelids curve into crescent" is a clause. For the text segment "shines with light", only the action keyword "shines" is checked, and the previous text segment of the latter text segment is determined by searching forward, that is, the text segment "eyelids curve into crescent", and the body part keyword is checked. Therefore, "eyelids curve into crescent" and "shines with light" are combined into a clause "eyelids curve into crescent, shines with light".

[0048] For the text segments "right hand high up", "arm naturally curved" and "palm gently sways", they are three clauses, but the same body part keyword "hand" is checked, so the three clauses are combined into a clause "right hand high up, arm naturally curved, palm gently sways".

[0049] Finally, the above behavior description text is divided into five different body part corresponding clauses: ①Walks near with light steps.

[0050] ②Face full of smiles.

[0051] ③Eyes curve into crescent, shines with light.

[0052] ④Body slightly forward.

[0053] ⑤Right hand high up, arm naturally curved, palm gently sways.

[0054] In this embodiment, the behavior description text is divided into text segments according to the punctuation marks, and the clauses are determined by checking the body part keywords and corresponding action keywords, so that only one body part and corresponding action are included in the determined clause, avoiding the generation of inaccurate sub-action sequences caused by too many body parts.

[0055] In some embodiments, the merging of each of the sub-action sequences to generate the action sequence comprises: extracting the pixel coordinate value sequence of the key points of each body part corresponding to each of the sub-action sequences, and merging the pixel coordinate value sequence of the key points of each body part to generate the action sequence. For example, the sub-action sequence of the head is {[100, 50]}, {[110, 60]}, and {[105, 55]}, the sub-action sequence of the shoulder is {[80, 120]}, {[85, 125]}, and {[95, 135]}, and the merged action sequence is {[100, 50], [80, 120]}, {[110, 60], [85, 125]}, and {[105, 55], [95, 135]}. Of course, the pixel coordinate values of the key points of all body parts at the corresponding time point are included in the merged {} in the above example, and the {} only shows two body parts.

[0056] It should be noted that the MotionDiffuse model generates an action sequence containing the pixel coordinate values of the entire human body at multiple time points, and therefore, when generating a sub-action sequence, the pixel coordinate values of the key points of the corresponding body part are valid values, and the pixel coordinate values of the key points of the remaining body parts are default values.

[0057] In some embodiments, generating a digital human video based on the character image and the action sequence comprises: The preprocessing of the character image comprises edge enhancement and texture extraction. Since the character image contains not only the character but also the background and other objects, in order to highlight the appearance details of the character in the character image, such as the texture of the hair and the wrinkles of the clothes, etc., the character image is preprocessed by edge enhancement and texture extraction to highlight the character region in the character image, so that only the character has corresponding actions in the subsequently generated digital human video, and the background or objects at the edge of the character do not follow the character actions to cause shaking.

[0058] The character in the preprocessed character image is recognized, a character region bounding box is determined, each side of the character region bounding box is drawn, the surrounding of the character region bounding box is divided into multiple peripheral regions, and the peripheral regions are expanded so that the area of each peripheral region is not less than the area of the character region bounding box, thereby avoiding the character actions exceeding the image range.

[0059] The digital human video is generated based on the character image with the extended peripheral region and the action sequence. Exemplarily, the character image with the extended peripheral region and the action sequence are input into an AnimateAnyone model, and the AnimateAnyone model converts the static image of the character into a digital human video controlled by the pose sequence. The deficiencies of the prior art in maintaining appearance detail consistency, action control capability and video inter-frame continuity are solved. By introducing a diffusion model and proposing an innovative network architecture, AnimateAnyone can generate high-quality videos while maintaining appearance detail consistency with the reference image and video fluency.

[0060] The core method includes three key components: ReferenceNet, Pose Guider and temporal layer. ReferenceNet is used to extract the appearance features of the reference image and integrate them into the denoising process through a spatial attention mechanism, ensuring that the generated video is highly consistent with the appearance details of the reference image. Pose Guider efficiently integrates the pose control signal into the denoising process through a lightweight convolutional layer, achieving precise control of the character's actions. The temporal layer models the temporal relationship between video frames to ensure the continuity of the actions and the fluency of the video.

[0061] In some embodiments, the image resolution based on which the key point pixel coordinates of each body part in the action sequence may not be consistent with the resolution of the digital human video to be generated, i.e., the image coordinate systems are not consistent, resulting in a deviation between the actions of the generated digital human video and the body parts performing the actions. Therefore, in order to ensure the accurate alignment between the actions of the digital human video and the body parts performing the actions, before generating the digital human video based on the character image with the extended peripheral region and the action sequence, the following steps are further included: The key point pixel coordinates of each body part in the action sequence are converted to key point pixel coordinates at the resolution of the digital human video to be generated, obtaining a converted action sequence. For example, the MotionDiffuse model is used to generate a sub-action sequence, and the AnaymateAnyone model is used to generate a digital human video. The key point pixel coordinates of each body part in the action sequence are determined based on the default character image resolution of the MotionDiffuse model, and the resolution of the digital human video to be generated is determined based on the AnaymateAnyone model. The character image resolution (e.g., 512x512) and the resolution of the digital human video to be generated (e.g., 1024x1024) are inconsistent. The key point pixel coordinates of each body part in the action sequence (1 x , y ) are converted to (2 x , 2 y ).

[0062] Based on this, the digital human video is generated based on the character image after the extended peripheral region and the converted action sequence, so that the action of the digital human video is accurately aligned with the body part performing the action.

[0063] The digital human video generation device provided by the present application is described below. The digital human video generation device described below can be referred to in correspondence with the digital human video generation method described above.

[0064] The digital human video generation device of the embodiment of the present application, as shown in Figure 2 includes: The behavior description text generation module 210 is configured to generate a behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user by using a text expansion model. The action sequence generation module 220 is configured to generate an action sequence based on the action in the behavior description text. The digital human generation module 230 is configured to generate a digital human video based on the character image and the action sequence. The text expansion model is trained based on a sample behavior keyword and a label behavior description text corresponding to the sample behavior keyword.

[0065] The digital human video generation device provided by the present application generates a behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user by using a text expansion model, generates an action sequence based on the action in the behavior description text, and generates a digital human video based on the character image and the action sequence. This embodiment utilizes the semantic understanding and generation capability of the text expansion model to provide a richer and more flexible text basis for subsequent action generation. The expanded text not only retains the core intention input by the user, but also can be more diversified in expression, so as to cope with diversified user instructions and complex user intentions.

[0066] In some embodiments, the text expansion model is a large model, and the behavior description text generation module 210 is specifically configured to: input the behavior keyword into the large model, and the large model generates a target prompt word according to the prompt word template and the behavior keyword.

[0067] The large model generates the behavior description text according to the target prompt word.

[0068] The large model is fine-tuned and trained based on a sample behavior keyword and a label behavior description text corresponding to the sample behavior keyword.

[0069] In some embodiments, the action sequence generation module 220 comprises: a sentence extraction module configured to extract sentences including body part keywords from the action description text, each of the sentences including a body part.

[0070] a sub-action sequence generation module configured to generate a sub-action sequence corresponding to each body part based on the action keywords in each of the sentences, the sub-action sequence including pixel coordinate values of key points of the corresponding body part at different time points; an action sequence synthesis module configured to merge the sub-action sequences to generate the action sequence.

[0071] In some embodiments, the sentence extraction module comprises: a text segment division module configured to divide the action description text into text segments based on punctuation marks in the action description text.

[0072] a sentence determination module configured to check for body part keywords and corresponding action keywords in any of the text segments, and determine the any of the text segments as a sentence if both the body part keywords and the corresponding action keywords are checked; if only the action keywords are checked in a subsequent text segment, determine a preceding text segment in which the body part keywords are last checked, and merge the subsequent text segment and the determined sentence of the preceding text segment into one sentence; and merge multiple sentences into one sentence if the multiple sentences include the same body part keywords.

[0073] In some embodiments, the action sequence synthesis module is specifically configured to extract pixel coordinate value sequences of key points of the body parts corresponding to each of the sub-action sequences, and merge the pixel coordinate value sequences of the key points of the body parts to generate the action sequence.

[0074] In some embodiments, the digital human generation module 230 specifically comprises: an image preprocessing module configured to perform preprocessing on the person image, the preprocessing including edge enhancement and texture extraction.

[0075] an image expansion module configured to identify a person region bounding box in the preprocessed person image, divide the person region bounding box into multiple peripheral regions based on each side of the person region bounding box, and expand the peripheral regions so that the area of each peripheral region is not less than the area of the person region bounding box.

[0076] a video generation module configured to generate a digital human video based on the person image after the peripheral regions are expanded and the action sequence.

[0077] In some embodiments, the digital human generation module 230 further includes an action sequence conversion module configured to convert the key point pixel coordinates of each body part in the action sequence to key point pixel coordinates at a resolution of the digital human video to be generated, to obtain a converted action sequence.

[0078] Based on this, the video generation module is specifically configured to generate a digital human video based on the character image after the extended peripheral region and the converted action sequence.

[0079] Figure 3 An example of a schematic diagram of the physical structure of an electronic device is shown in Figure 3 As shown, the electronic device can include a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can invoke the logical instructions in the memory 330 to execute a digital human video generation method, which includes: Generating a behavior description text related to the behavior keyword according to the behavior keyword and the prompt word template input by the user using a text expansion model.

[0080] Generating an action sequence based on the actions in the behavior description text.

[0081] Generating a digital human video based on the character image and the action sequence.

[0082] The text expansion model is trained based on sample behavior keywords and label behavior description texts corresponding to the sample behavior keywords.

[0083] In addition, the logical instructions in the memory 330 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0084] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored in a non-transitory computer-readable storage medium, and the computer program being capable of executing the digital human video generation method provided by the above-mentioned method when executed by a processor, the method comprising: generating a behavior description text related to the behavior keyword according to the user input behavior keyword and the prompt word template by using a text expansion model.

[0085] generating a motion sequence based on the actions in the behavior description text.

[0086] generating a digital human video based on the character image and the motion sequence.

[0087] wherein the text expansion model is trained based on sample behavior keywords and label behavior description texts corresponding to the sample behavior keywords.

[0088] In another aspect, the present application also provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is capable of executing the digital human video generation method provided by the above-mentioned method when executed by a processor, the method comprising: generating a behavior description text related to the behavior keyword according to the user input behavior keyword and the prompt word template by using a text expansion model.

[0089] generating a motion sequence based on the actions in the behavior description text.

[0090] generating a digital human video based on the character image and the motion sequence.

[0091] wherein the text expansion model is trained based on sample behavior keywords and label behavior description texts corresponding to the sample behavior keywords.

[0092] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0093] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0094] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating a digital human video, characterized in that: include: Generate a behavior description text related to the behavior keyword based on the behavior keyword and prompt word template input by the user using a text expansion model; generating an action sequence based on the actions in the behavior description text; generating a digital human video based on the character image and the action sequence; The text expansion model is obtained by training based on sample behavior keywords and label behavior description texts corresponding to the sample behavior keywords.

2. The method for generating a digital human video according to claim 1, wherein: The text expansion model is a large model; The text expansion model is used to generate behavior description text related to the behavior keywords input by the user and the prompt word template, including: The behavior keyword is input into the big model, and the big model generates a target prompt word according to the prompt word template and the behavior keyword; The large model generates the behavior description text according to the target prompt word; The large model is obtained by fine-tuning and training based on sample behavior keywords and the label behavior description text corresponding to the sample behavior keywords.

3. The method for generating a digital human video according to claim 1, wherein: Generating an action sequence based on the actions in the behavior description text includes: Extracting sentences containing body part keywords from the behavior description text, each of the sentences containing a body part; generating a sub-action sequence corresponding to each body part based on the action keywords in each of the sentences, wherein the sub-action sequence includes pixel coordinate values ​​of key points of the corresponding body part at different time points; The sub-action sequences are merged to generate the action sequence.

4. The method for generating a digital human video according to claim 3, wherein: Extracting sentences containing body part keywords from the behavior description text, each of which contains a body part, including: Using punctuation marks in the behavior description text as segments to obtain multiple text segments; Checking body part keywords and corresponding action keywords in any text segment; if both the body part keywords and the corresponding action keywords are detected, determining the text segment as a sentence; if only the action keywords are detected in a subsequent text segment, determining the previous text segment in which the body part keywords were most recently detected, and merging the sentences determined in the subsequent text segment and the previous text segment into one sentence; If multiple clauses contain the same body part keyword, the clauses are merged into one clause.

5. The method for generating a digital human video according to claim 3, wherein: Merging the sub-action sequences to generate the action sequence includes: A sequence of pixel coordinate values ​​of key points of the body parts corresponding to each of the sub-action sequences is extracted, and the sequence of pixel coordinate values ​​of the key points of each body part is merged to generate the action sequence.

6. The method for generating a digital human video according to any one of claims 1 to 5, characterized in that: Generating a digital human video based on the character image and the action sequence, including: Preprocessing the character image, wherein the preprocessing includes edge enhancement and texture extraction; Recognize a person in the pre-processed person image, determine a person region labeling frame, draw lines along the edges of the person region labeling frame, divide the area around the person region labeling frame into a plurality of peripheral regions, and expand the peripheral regions so that the area of ​​each peripheral region is not less than the area of ​​the person region labeling frame; A digital human video is generated based on the character image after the peripheral area is expanded and the action sequence.

7. The method for generating a digital human video according to claim 6, wherein: Before generating a digital human video based on the character image after the peripheral area is expanded and the action sequence, the method further includes: Converting the pixel coordinates of key points of each body part in the action sequence to the pixel coordinates of key points at the resolution of the digital human video to be generated, to obtain a converted action sequence; Generating a digital human video based on the character image after the peripheral area is expanded and the action sequence, including: Generate digital human video based on the character image with expanded surrounding area and the converted action sequence.

8. A digital human video generation device, characterized in that: include: A behavior description text generation module is used to generate behavior description text related to the behavior keywords according to the behavior keywords and prompt word templates input by the user using a text expansion model; An action sequence generation module, configured to generate an action sequence based on the actions in the behavior description text; A digital human generation module, configured to generate a digital human video based on a human image and the action sequence; The text expansion model is obtained by training based on sample behavior keywords and label behavior description texts corresponding to the sample behavior keywords.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the digital human video generation method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the digital human video generation method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Backboard video generation method for real-time interactive digital human and related device

    CN121619475A