Virtual human video generation method and device, computer program product and electronic equipment
Through the large language model and Flux model combined with the skeleton posture diagram, the storyboard planning information and action images of virtual human videos are automatically generated, which solves the problems of low efficiency and high cost in the existing technology, and achieves efficient and accurate virtual human video generation.
Patent Information
- Application Number
- CN202510315941.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the operation efficiency of generating virtual human videos is low, the production cycle is long and the cost is high.
The audio content is recognized through a large language model and the storyboard planning information is generated. Combined with the Flux model and the skeleton posture diagram, the scene and character action images are automatically output, and the virtual human video is combined to generate.
It improves the operation efficiency of generating virtual human videos, shortens production cycles, reduces costs, and increases the convenience and accuracy of generation.
Smart Images

Figure CN120182448A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to a method and apparatus for generating virtual human videos, a computer program product, and an electronic device. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The descriptions herein are not admitted to be prior art merely by virtue of their inclusion in this section.
[0003] In the traditional process of producing virtual human videos based on audio, it is generally necessary to first plan the characters and scenes involved in the audio according to the audio content, then manually conceive the storyboard, determine the storyboard content, then produce the character and scene images, and finally make the static images into corresponding action videos. Summary of the Invention
[0004] However, in related technologies, the operation efficiency of generating virtual human videos is low, the production cycle is long, and the cost is relatively high.
[0005] Therefore, there is a great need for an improved method for generating virtual human videos to automatically generate virtual human videos according to large language models.
[0006] In this context, embodiments of the present disclosure are expected to provide a method for generating virtual human videos, an apparatus for generating virtual human videos, a computer program product, and an electronic device.
[0007] According to one aspect of the present disclosure, there is provided a method for generating a virtual human video, including: recognizing corresponding text content according to the original audio, and determining skeleton pose diagrams of multiple groups of human body actions;
[0008] Processing the text content and the skeleton pose diagrams through a large language model to output storyboard planning information corresponding to the text content; the storyboard planning information includes the overall style, scene description, character description, and storyboard attribute information of each storyboard; generating a scene animation based on the overall style and the scene description, and generating a character static image based on the overall style and the character description; generating a character action animation for each storyboard according to the character static image bound to each storyboard and the storyboard attribute information, and combining the scene animation and the character action animations of multiple storyboards to generate a virtual human video in the same scene.
[0009] In an exemplary embodiment of the present disclosure, the generating a scene animation based on the overall style and the scene description includes: using the overall style and the scene description as prompt words to generate a scene static image, and converting the scene static image into the scene animation.
[0010] In an exemplary embodiment of the present disclosure, the conversion of the static scene graph to the dynamic scene graph includes: predicting a plurality of sequential frame images located after the static scene graph based on the static scene graph; combining the plurality of sequential frame images to obtain the dynamic scene graph.
[0011] In an exemplary embodiment of the present disclosure, the storyboard planning information includes storyboard attribute information; the storyboard attribute information includes the storyboard duration, the character to which the storyboard belongs, and the preset actions of the character to which the storyboard belongs; the generation of the dynamic character action graph for each storyboard according to the static character graph bound to each storyboard and the storyboard attribute information includes: obtaining the character name corresponding to the character to which each storyboard belongs, and obtaining the static character graph bound to the character name according to the character name; obtaining the skeleton pose graph corresponding to the preset action according to the preset action of the character to which the storyboard belongs; generating the dynamic character action graph for each storyboard with the number of video frames corresponding to the storyboard duration based on the bound static character graph and the skeleton pose graph corresponding to the preset action.
[0012] In an exemplary embodiment of the present disclosure, the generation of the dynamic character action graph for each storyboard with the number of video frames corresponding to the storyboard duration based on the bound static character graph and the skeleton pose graph corresponding to the preset action includes: generating an image according to the static character graph and the skeleton pose graph corresponding to the preset action to obtain the character action pose graph for each storyboard; the character action pose graph includes the first-frame character action pose graph and the last-frame character action pose graph; combining the first-frame character action pose graph and the last-frame character action pose graph to generate the dynamic character action graph for each storyboard.
[0013] In an exemplary embodiment of the present disclosure, the generation of the character action pose graph for each storyboard by generating an image according to the static character graph and the skeleton pose graph corresponding to the preset action includes: extracting the identity feature information of the static character graph; parsing the skeleton pose graph to obtain pose control information; fusing the identity feature information and the pose control information to determine the fusion feature, and generating the character action pose graph based on the fusion feature; wherein, the character action pose graph conforms to the pose control information and retains the identity feature information.
[0014] In an exemplary embodiment of the present disclosure, the combination of the first-frame character action pose graph and the last-frame character action pose graph to generate the dynamic character action graph for each storyboard includes:
[0015] Perform intermediate frame prediction based on the first-frame human action pose map and the last-frame human action pose map to generate an intermediate-frame human action pose map; determine the number of reference video frames according to the storyboard duration and frame rate; based on the first-frame human action pose map, the intermediate-frame human action pose map, and the last-frame human action pose map, compose a human action animation with the number of video frames being the number of reference video frames.
[0016] In an exemplary embodiment of the present disclosure, the storyboard attribute information further includes the scene to which the storyboard belongs; the combining the scene animation and the human action animations of multiple storyboards to generate a virtual human video in the same scene includes: determining, according to the scene to which the storyboard belongs, the target scene animation corresponding to each storyboard from the scene animations; performing matte extraction on the human action animations to obtain human action animations with transparent backgrounds; merging the target scene animations with the human action animations with transparent backgrounds of each storyboard to generate target scene human action animations for each storyboard; splicing the target scene human action animations of the multiple storyboards to generate a virtual human video in the same scene.
[0017] In an exemplary embodiment of the present disclosure, the merging the target scene animation with the human action animations with transparent backgrounds of each storyboard to generate target scene human action animations for each storyboard includes: using the target scene animation as the lower layer and the human action animations with transparent backgrounds of each storyboard as the upper layer; performing layer merging on the upper layer and the lower layer to generate target scene human action animations for each storyboard.
[0018] In an exemplary embodiment of the present disclosure, the storyboard planning information includes storyboard attribute information; the storyboard attribute information includes the storyboard duration and the scene to which the storyboard belongs; the method further includes: determining, according to the scene to which the storyboard belongs, the target scene animation corresponding to each storyboard from the scene animations; determining the number of video frames of the target scene animation according to the storyboard duration; determining a virtual human video in the same scene according to the target scene animation.
[0019] According to one aspect of the present disclosure, there is provided a virtual human video generation device, including: a text content recognition module, configured to recognize corresponding text content according to the original audio and determine skeleton pose diagrams of multiple groups of human body actions; a storyboard planning information determination module, configured to process the text content and the skeleton pose diagrams through a large language model and output storyboard planning information corresponding to the text content; the storyboard planning information includes the overall style, scene description, character description, and storyboard attribute information of each storyboard; an image generation module, configured to generate a scene moving picture based on the overall style and the scene description, and generate a character static picture based on the overall style and the character description; a video generation module, configured to generate a character action moving picture of each storyboard according to the character static picture bound to the storyboard and the storyboard attribute information, and combine the scene moving picture and the character action moving pictures of multiple storyboards to generate a virtual human video in the same scene.
[0020] According to one aspect of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor, implements the virtual human video generation method as described in any one of the above.
[0021] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions; wherein the processor is configured to execute the virtual human video generation method as described in any one of the above by executing the executable instructions.
[0022] In the virtual human video generation method, device, computer program product, and electronic device according to the embodiments of the present disclosure, on the one hand, by processing the text content corresponding to the original audio and the skeleton pose diagrams through a large language model, it is possible to automatically output the storyboard planning information corresponding to the text content, avoiding the problem of low operation efficiency in generating virtual human videos in the related art and improving the operation efficiency of generating storyboard planning information. On the other hand, a scene moving picture can be generated according to the overall style and the scene description in the storyboard planning information, and a character static picture can be generated based on the overall style and the character description;
[0023] A character action moving picture of each storyboard is generated according to the character static picture bound to the storyboard and the storyboard attribute information, and the scene moving picture and the character action moving pictures of multiple storyboards are combined to generate a virtual human video in the same scene, which can accurately control the character action moving pictures of each storyboard while keeping the characters consistent, improving the accuracy of the virtual human video, increasing the convenience of generating the virtual human video, shortening the production cycle, reducing the cost of generating the virtual human video, and also increasing the versatility and application scope. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understandable. In the drawings, several embodiments of the present disclosure are shown by way of illustration and not limitation, wherein:
[0025] Figure 1 A schematic diagram of the system architecture of the application scenario of the embodiment of the present disclosure is shown schematically.
[0026] Figure 2 A schematic diagram of the flow of the virtual human video generation method in the embodiment of the present disclosure is shown schematically.
[0027] Figure 3 A schematic diagram of the flow of generating the character action animation for each storyboard based on the static character image and the storyboard attribute information bound to each storyboard in the embodiment of the present disclosure is shown schematically.
[0028] Figure 4 A flowchart of generating the character action animation for each storyboard based on the static character image and the skeleton pose image in the embodiment of the present disclosure is shown schematically.
[0029] Figure 5 A flowchart of obtaining the character action animation for each storyboard according to the first-frame character action pose image and the last-frame character action pose image in the embodiment of the present disclosure is shown schematically.
[0030] Figure 6 A schematic diagram of the overall flow of generating a virtual human video in the embodiment of the present disclosure is shown schematically.
[0031] Figure 7 A schematic block diagram of the virtual human video generation device in the embodiment of the present disclosure is shown schematically.
[0032] Figure 8 A block diagram of an electronic device in the embodiment of the present disclosure is shown schematically.
[0033] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed Embodiments
[0034] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present disclosure, and not to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0035] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0036] According to the embodiments of the present disclosure, a virtual human video generation method, a virtual human video generation device, a computer program product, and an electronic device are provided.
[0037] In addition, the number of any elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0038] Next, with reference to several representative embodiments of the present disclosure, the principles and spirit of the present disclosure will be elaborated in detail.
[0039] In the related art, first, the characters and scenes involved in the audio are planned according to the audio content, then the storyboard is conceived, after determining the storyboard content, the character and scene images are produced, and finally the static images are made into corresponding action videos. Therefore, the efficiency of generating virtual human videos is low and the production cycle is long.
[0040] Based on the above, the embodiments of the present disclosure generate the storyboard planning information of the original audio through a large language model, and combine the Flux model to process the overall style, character description, and storyboard attribute information of each storyboard in the storyboard planning information to generate the character action animated images of each storyboard. Further, the scene animated images obtained from the scene description and the overall style and the character action animated images of multiple storyboards are merged into layers to generate a virtual human video corresponding to the original audio.
[0041] After introducing the basic principles of the present disclosure, the various non-limiting embodiments of the present disclosure will be specifically introduced below.
[0042] It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0043] First, refer to Figure 1 , Figure 1 FIG. shows a schematic diagram of the system architecture of an exemplary application scenario of the embodiments of the present disclosure. This virtual human video generation method can be used in any scenario where a virtual human video is generated according to audio. For example, it can be applied to scenarios such as audiobooks where there is only audio material and video material needs to be expanded, or it can also be applied to scenarios where virtual human videos need to be generated in games. Specific limitations are not made here. As Figure 1As shown, the system architecture 100 includes a terminal 101, a network 102, and a server 103. The server 103 identifies the original audio extracted from the terminal 101 to obtain the text content, and then processes the text content and the skeleton postures of multiple groups of human body movements through a large language model to generate the overall style, scene description, character description, and the storyboard attribute information of each storyboard corresponding to the text content. Further, a scene animation and static character images are generated, and the character action animations of each storyboard are generated according to the static character images and the storyboard attribute information bound to each storyboard. Then, the scene animation and the character action animations of multiple storyboards are merged into layers to generate a virtual human video in the same scene. Those skilled in the art should understand, Figure 1 The schematic framework shown is only an example in which the embodiments of the present disclosure can be implemented. The scope of application of the embodiments of the present disclosure is not limited by any aspect of this framework.
[0044] It should be noted that the server 103 can be a local server or a remote server. In addition, the server 103 can also be other products that can provide storage functions or processing functions, such as cloud servers. The embodiments of the present disclosure are not specifically limited herein. The server can also be composed of a terminal device with fast computing capabilities or a vehicle-mounted device, etc., which is not limited here. The terminal 101 can be any device that can generate user behavior data, such as a smart phone, a tablet computer, and a computer.
[0045] It should be understood that in the application scenarios of the present disclosure, the actions of the embodiments of the present disclosure can be executed by the server 103 or by a terminal with computing capabilities. The present disclosure is not limited in terms of the execution subject as long as the actions disclosed in the embodiments of the present disclosure are executed.
[0046] Next, in combination with Figure 1 the application scenario of, refer to Figure 2 to describe the virtual human video generation method according to an exemplary embodiment of the present disclosure. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0047] Figure 2 shows a flowchart of the virtual human video generation method according to an embodiment of the present disclosure. Referring to Figure 2 as shown, the virtual human video generation method may include the following steps:
[0048] In step S210, the text content corresponding to the original audio is identified, and the skeleton posture diagrams of multiple groups of human body movements are determined;
[0049] In step S220, the large language model processes the text content and the skeleton pose diagram, and outputs the storyboard planning information corresponding to the text content; the storyboard planning information includes the overall style, scene description, character description, and the storyboard attribute information of each storyboard.
[0050] In step S230, a scene animation is generated based on the overall style and the scene description, and a character static image is generated based on the overall style and the character description.
[0051] In step S240, a character action animation for each storyboard is generated according to the character static image bound to each storyboard and the storyboard attribute information, and the scene animation and the character action animations of multiple storyboards are combined to generate a virtual human video in the same scene.
[0052] In the embodiments of the present disclosure, on the one hand, by processing the text content corresponding to the original audio and the skeleton pose diagram through the large language model, the storyboard planning information corresponding to the text content can be automatically output, avoiding the problem of low operation efficiency in generating virtual human videos in the related art, and improving the operation efficiency of generating the storyboard planning information. On the other hand, a scene animation can be generated according to the overall style and the scene description in the storyboard planning information, and a character static image can be generated based on the overall style and the character description; a character action animation for each storyboard is generated according to the character static image bound to each storyboard and the storyboard attribute information, and the scene animation and the character action animations of multiple storyboards are combined to generate a virtual human video in the same scene, which can accurately control the character action animations of each storyboard while keeping the characters consistent, improving the accuracy of the virtual human video, increasing the convenience of generating the virtual human video, shortening the production cycle, reducing the cost of generating the virtual human video, and also increasing the versatility and application scope.
[0053] Next, the virtual human video generation method in the embodiments of the present disclosure will be explained in detail with reference to the accompanying drawings.
[0054] In step S210, the text content corresponding to the original audio is recognized, and the skeleton pose diagrams of multiple groups of human body actions are determined.
[0055] In the embodiments of the present disclosure, the original audio can be an audio of any length and any type, and the original audio can be a pre-determined audio. For example, the original audio can be an audiobook or other types of audio, and the language type of the original audio can be of any type. Audio recognition can be performed on the original audio to identify the text content corresponding to the original audio. Exemplarily, audio features can be extracted from the original audio, and the audio features can be, for example, MFCC features. In some embodiments, an acoustic model can be used to perform acoustic modeling on the audio features, mapping the audio features to phonemes or sub-word units. The acoustic model can be a deep learning model, such as an LSTM or Transformer model. Combining with a language model to decode the obtained phonemes or sub-word units to predict the probability of the word sequence. The decoder combines the acoustic model and the language model to determine the most likely text and output it to obtain the text content corresponding to the original audio.
[0056] In some embodiments, multiple sets of skeleton pose diagrams of human body movements can also be pre-configured. The human body movements can be one or more of standing, raising hands, shaking the head, and giving a thumbs up. The multiple sets of pre-configured skeleton pose diagrams of human body movements can be configured according to the action type and actual requirements. The skeleton pose diagrams are used to represent various body movements and postures of a person through a skeleton model. The skeleton model therein can be composed of a combination of lines and joint points to clearly display the dynamics and structural relationships of a person.
[0057] In step S220, the large language model processes the text content and the skeleton pose diagrams to output the storyboard planning information corresponding to the text content; the storyboard planning information includes the overall style, scene description, character description, and the storyboard attribute information of each storyboard.
[0058] In the embodiments of the present disclosure, the text content and the skeleton pose diagrams can be input into the large language model, and the large language model is controlled by a prompt to output the storyboard planning information corresponding to the text content. Among them, the prompt can be used to specify the input and output content of the storyboard planning information and can also specify the output format. The prompt can be, for example, "Output the storyboard planning information according to the text content and the skeleton pose diagrams and output the storyboard planning information in format 1". The storyboard planning information is used to plan and describe the detailed content of each storyboard.
[0059] In some embodiments, the large language model can generally be a language model based on the Transformer architecture. The text content and the action types corresponding to the skeleton pose diagrams are processed into text prompts, and the action types can be represented by action names. After inputting the text prompts, tokenization and encoding are first performed in the large language model to convert the text into digital encodings, and special tokens such as the beginning token (bos_token) and the ending token (eos_token) are added to assist the model in identifying the start and end of the text and facilitate sequence padding or truncation operations. Subsequently, the digital encodings are converted into high-dimensional vectors, i.e., embedding vectors, through the embedding layer. In the encoder part, three matrices, namely query (Q), key (K), and value (V), are calculated for the embedding vectors. With the help of the multi-head attention mechanism, the attention weight matrix is obtained through matrix dot product, scaling, and softmax operations. Through weighted summation and multi-head parallel calculation, the multi-head attention output is generated. After that, the multi-head attention output passes through a feed-forward neural network, combined with residual connection and layer normalization, to achieve feature extraction. Finally, the features extracted are decoded cyclically through the decoder to generate the output text word by word to determine the storyboard planning information corresponding to the text content.
[0060] Specifically, the storyboard planning information may include the overall style, scene description, character description, and the storyboard attribute information of each storyboard. The overall style can be, for example, ancient style, modern style, and so on. The scene description is used to represent one or more of the objects included in the scene, the characteristics of each object, and the relationships between different objects. The character description is used to represent the characteristics of each character, such as facial expressions, physical features, age, occupation, personality, and so on. The storyboard attribute information of each storyboard refers to the content parameters of each storyboard, and the storyboard attribute information is used to generate the character action animation of each storyboard. Among them, when a storyboard contains characters, the storyboard attribute information may include the storyboard duration, the character to which the storyboard belongs, and the preset action of the character to which the storyboard belongs. The preset action can match any one of the skeleton pose diagrams of multiple groups of human body actions configured in advance. When a storyboard does not contain characters, the storyboard attribute information may include the storyboard duration, without including the character to which the storyboard belongs and the preset action of the character to which the storyboard belongs. The storyboard duration can be represented by the storyboard time range, the character to which the storyboard belongs can be determined by the character name, and the preset action of the character to which the storyboard belongs can be represented by the action name.
[0061] For example, when the text content and the skeleton pose diagrams are input into the large language model, the generated storyboard planning information can be represented in the following form:
[0062] Style: "Ancient style". Scene description: "Palace: Interior view of a magnificent ancient palace, filled with rows of ministers, solemn, ancient style". Character description: "A: Ancient XX, middle-aged man, black hair, serious expression, wearing a crown, Han-style clothing, gold, ancient, ancient style"; "B: Minister, middle-aged man, black hair, gloomy, frowning, dark complexion, wearing black armor, Han-style clothing, ancient, ancient style". Storyboard timeline: "00:03: Scene, palace; Character, B; Action: raise hand"; "00:19: Scene, palace; Character, A; Action: nod".
[0063] In the embodiments of the present disclosure, by outputting the storyboard planning information corresponding to the text content through the large language model, it avoids the problems of long production cycle and slow production progress in the related art where first the characters and scenes are planned according to the audio content, and then the storyboard is manually conceived, and the character and scene images are produced after determining the storyboard content. By predicting the storyboard planning information through the large language model, the operation efficiency is improved, and the rationality and accuracy of the storyboard planning information can be improved.
[0064] In step S230, generate a scene animated image based on the overall style and the scene description, and generate a character static image based on the overall style and the character description.
[0065] In the embodiments of the present disclosure, the scene animated image refers to the content of a certain scene presented in the form of a dynamic image (GIF or video clip). Exemplarily, based on the Flux model, the overall style and the scene description can be used as prompt words together to generate a scene static image, and the scene static image can be converted into a scene animated image. The Flux model can input prompt words to generate an image corresponding to the content. Specifically, the overall style and the scene description can be combined to generate prompt words. The combination here can be direct splicing or other combination methods, which are not specifically limited here. Further, the prompt words can be input into the Flux model, so that the model converts the prompt words into a vector representation understandable by the Flux model, and performs diffusion on the vector representation to generate an image to obtain the scene static image.
[0066] On this basis, multiple sequential frame images after the scene static image can be predicted based on the scene static image, and the multiple sequential frame images can be combined to obtain the scene animated image. Exemplarily, based on SVD (Stable VideoDiffusion, a video generation model based on the diffusion model), the corresponding video can be generated according to the input scene static image. Specifically, the scene static image can be input into the UNET network, and multiple sequential frame images after the scene static image can be predicted frame by frame, and the multiple sequential frame images can be synthesized into the scene animated image in the generation order. Figure 1 Frame by frame to predict multiple sequential frame images after the scene static image, and the multiple sequential frame images can be synthesized into the scene animated image in the generation order.
[0067] In addition, static character images can also be generated based on the overall style and character description. Exemplarily, based on the Flux model, the character description and the overall style can be used as prompts for text-to-image generation to generate corresponding static character images. Specifically, a text encoder can be used to convert the prompts into text feature vectors, and the text feature vectors are input into the generator of the Flux model, so that the generator generates static character images based on the text feature vectors. Further, the static character images can be bound to the character names, so that there is a correspondence between the static character images and the character names.
[0068] In the embodiments of the present disclosure, scene animated images are generated through scene descriptions and the overall style, and static character images are generated through character descriptions and the overall style. Since the overall style is integrated in the process of generating images, the style of the generated images can be accurately controlled, and the accuracy of the generated scenes and characters can be improved.
[0069] Next, continue to refer to Figure 2 As shown in, in step S240, animated character action images for each shot are generated according to the static character images bound to each shot and the shot attribute information, and the scene animated image and the animated character action images of multiple shots are combined to generate a virtual human video in the same scene.
[0070] In the embodiments of the present disclosure, for each shot, animated character action images for each shot can be generated according to the static character images bound to the shot and the shot attribute information output by the large language model. Among them, the static character image of the shot can be determined according to the character name set for each shot and the correspondence between the character name and the static character image. When the shot includes a character, the shot attribute information may include the shot duration, the character to which the shot belongs, and the preset action of the character to which the shot belongs. Or when the shot does not include a character, the shot attribute information may only include the shot duration. Here, the process of generating a virtual human video in the case where the shot includes a character is specifically described first.
[0071] Figure 3 The flowchart of generating animated character action images for each shot according to the static character images bound to each shot and the shot attribute information is schematically shown in, refer to Figure 3 As shown in, it mainly includes the following steps:
[0072] In step S310, obtain the character name corresponding to the character to which each shot belongs, and obtain the static character image bound to the character name according to the character name;
[0073] In step S320, according to the preset action of the character to which the shot belongs, obtain the skeleton pose image corresponding to the preset action;
[0074] In step S330, based on the bound static figure of the person and the skeleton pose figure corresponding to the preset action, a moving figure of the person's action for each storyboard corresponding to the number of video frames and the storyboard duration is generated.
[0075] Among them, since the storyboard attribute information includes the storyboard duration, the person to whom the storyboard belongs, and the preset action of the person to whom the storyboard belongs. Therefore, first, the person name corresponding to the person to whom the storyboard belongs can be determined. After the person name is determined, the static figure of the person bound to the person name can be determined according to the correspondence between the person name and the static figure of the person.
[0076] Furthermore, the skeleton pose figure corresponding to the preset action can be obtained according to the preset action of the person to whom the storyboard belongs included in the storyboard attribute information. Since multiple groups of skeleton pose figures of the person's body movements are configured in advance, based on this, the preset action can be matched with multiple groups of the person's body movements, and then the skeleton pose figure corresponding to the person's body movement with a successful match can be used as the skeleton pose figure corresponding to the preset action of the person to whom the storyboard belongs. For example, if the preset action of the person to whom storyboard 1 belongs is "raising the hand", the skeleton pose figure corresponding to "raising the hand" can be used as the skeleton pose figure of storyboard 1. If the preset action of the person to whom storyboard 2 belongs is "nodding", the skeleton pose figure corresponding to "nodding" can be used as the skeleton pose figure of storyboard 2.
[0077] On this basis, the bound static figure of the person and the skeleton pose figure corresponding to the preset action can be input into the Flux model, and the moving figure of the person's action for each storyboard is generated by combining the PuLID (Pure and Lightning ID) technology and the Controlnet-Openpose technology of the Flux model. Exemplarily, the static figure of the person can be used as the input reference figure of the PuLID technology of the Flux model, and the skeleton pose figure corresponding to the preset action of the person to whom the storyboard belongs can be used as the input control figure of the Controlnet-Openpose technology of the Flux model. By combining the two technologies, image generation can be performed based on the Flux model according to the static figure of the person and the skeleton pose figure corresponding to the preset action to obtain the moving figure of the person's action for each storyboard. The PuLID technology of the Flux model is an image generation technology for ensuring person consistency. It can generate different scene figures with the same person characteristics as the reference figure according to the input person image, ensuring person consistency during the image generation process. The Controlnet-Openpose technology of the Flux model is used to precisely control the human body posture when generating images or videos. It can stably generate images of the person corresponding to the posture according to the input skeleton pose figure. Specifically, the human key point information extracted by OpenPose can be used as a conditional input to guide the generation model to generate images that conform to a specific posture. Among them, the number of video frames of the moving figure of the person's action generated for each storyboard corresponds to the storyboard duration in the storyboard attribute information.
[0078] Figure 4 schematically shows a static figure of a person based on binding and a skeleton pose figure corresponding to a preset action, and a flowchart for generating a moving figure of the person's action for each storyboard, which belongs to the specific implementation process of step S330. Refer to Figure 4 As shown in , it mainly includes the following steps:
[0079] In step S410, image generation is performed according to the static figure of the person and the skeleton pose figure corresponding to the preset action, and a person's action pose figure for each storyboard is obtained; the person's action pose figure includes the first-frame person's action pose figure and the last-frame person's action pose figure;
[0080] In step S420, the first-frame person's action pose figure and the last-frame person's action pose figure are combined to generate a moving figure of the person's action for each storyboard.
[0081] In the embodiment of the present disclosure, when generating the person's action pose figure for each storyboard, the static figure of the person can be used as the input reference figure of the PuLID technology of the Flux model, and the skeleton pose figure corresponding to the preset action of the person to which the storyboard belongs can be used as the input control figure of the Controlnet-Openpose technology of the Flux model. Based on this, the identity feature information of the static figure of the person can be extracted according to the PuLID technology of the Flux model; the identity feature information can include one or more of facial features, hairstyle features, and clothing features. The PuLID technology can generate different scene figures with the same person features as the reference image according to the input person image, ensuring the consistency of the person in the image generation process, but it cannot ensure the action.
[0082] Analyze the skeleton pose figure corresponding to the preset action of the person to which the storyboard belongs according to the OpenPose technology of the Flux model, and extract the key point information in the skeleton pose figure. The key information can be joint positions, limb directions, etc. Further, these key point information can be encoded into pose control signals and input into the ControlNet network. The ControlNet network combines the pose control signals with the intermediate features of the generation model to output pose control information, and the pose control information can be a skeleton key point figure or a pose feature figure.
[0083] Based on this, the personal identity characteristics and pose control information can be input into the Flux model. Through multimodal feature fusion, the Flux model fuses the personal identity characteristics and the pose control information to obtain fused features, and inputs the fused features into the generator network. The generator network is used to generate a personal action pose graph that conforms to the pose control information and retains the personal identity characteristics. The generator network can be a GAN (Generative Adversarial Network) or a diffusion model. The personal action pose graph generated by the Flux model not only conforms to the skeleton pose graph corresponding to the preset action of the character in the storyboard, but also retains the personal identity characteristics of the input static personal graph, so as to ensure the consistency of the characters generated in multiple storyboards.
[0084] It should be noted that the personal action pose graph here includes the first-frame personal action pose graph and the last-frame personal action pose graph. For example, the first-frame personal action pose graph can be a person standing, and the last-frame personal action pose graph can be a person raising their hand. In the process of generating the personal action animation for each storyboard, the first-frame personal action pose graph and the last-frame personal action pose graph can be combined to obtain the personal action animation for each storyboard. For each storyboard, the first-frame personal action pose graph is fixed, and what changes for each storyboard is the last-frame personal action pose graph.
[0085] Figure 5 schematically shows the flow chart of combining the first-frame personal action pose graph and the last-frame personal action pose graph to obtain the personal action animation for each storyboard. Figure 5 The steps in are the specific implementation manners of step S420. Refer to Figure 5 As shown in, it mainly includes the following steps:
[0086] In step S510, intermediate-frame prediction is performed based on the first-frame personal action pose graph and the last-frame personal action pose graph to generate an intermediate-frame personal action pose graph;
[0087] In step S520, the reference video frame number is determined according to the storyboard duration and the frame rate;
[0088] In step S530, based on the first-frame personal action pose graph, the intermediate-frame personal action pose graph, and the last-frame personal action pose graph, a personal action animation with the video frame number being the reference video frame number is composed.
[0089] In the embodiments of the present disclosure, the first frame of the human action pose diagram can be used as the first frame, the last frame of the human action pose diagram can be used as the last frame, and the first frame of the human action pose diagram and the last frame of the human action pose diagram can be used as the input of the ToonCrafter technology to perform intermediate frame prediction and generate intermediate frame human action pose diagrams. The ToonCrafter technology can complete the continuous intermediate frames based on the input first frame image and the last frame image, thereby obtaining the corresponding video.
[0090] Exemplarily, the first frame of the human action pose diagram and the last frame of the human action pose diagram can be aligned to ensure that the human position and proportion are consistent. Next, a pose estimation model is used to extract the key point information of the human, and the key point information is converted into a vector representation. The key points can be, for example, joint positions. The pose estimation model can be, for example, the OpenPose model, the MediaPipe model, etc. Further, an interpolation algorithm is used to generate the key point positions of the intermediate frames, and an intermediate frame human action pose diagram is generated based on the interpolated key points using a generation model. Among them, the interpolation algorithm can be linear interpolation or spline interpolation, and the generation model can be a GAN model or a diffusion model, etc. The overall style of the generated intermediate frame human action pose diagram is consistent with that of the first frame of the human action pose diagram and the last frame of the human action pose diagram. Alternatively, the motion displacement between the first frame of the human action pose diagram and the last frame of the human action pose diagram can also be analyzed to infer the motion estimation result; further, an intermediate frame human action pose diagram is generated according to the motion estimation result, and the intermediate frame human action pose diagram can be generated according to a video frame interpolation model.
[0091] For example, the first frame of the human action pose diagram can be, for example, a human standing diagram, and the last frame of the human action pose diagram can be, for example, a human raising hand diagram. The human standing diagram and the human raising hand diagram can be used as the input of the ToonCrafter technology to generate a corresponding human action animation, and the human action animation can be, for example, a human raising hand action diagram.
[0092] For the storyboard attribute information, it may further include the storyboard duration. The storyboard duration here can be used to control the duration of the animated GIF of the character actions corresponding to each storyboard or the number of video frames. Specifically, the reference video frame number can be determined according to the storyboard duration and the frame rate. The reference video frame number is used to represent the number of video frames of the animated GIF of the character actions, and the number of video frames refers to the total length of the video. Specifically, the reference video frame number can be determined according to the storyboard duration and the frame rate. For example, the storyboard duration and the frame rate can be multiplied, and the product of the two can be used as the reference video frame number. The frame rate refers to the number of frames displayed per second, and the frame rate is a key parameter determining the smoothness of the animation or video. The frame rate can be comprehensively determined according to the specific application scenario, user type, device parameters, and content type. The application scenario can be, for example, film and television animation, games, or television and live broadcasts, etc. The user type can be ordinary users, professional users, or game players. The device parameters can be hardware performance, file size, or bandwidth. The content type can be action scenes, static or slow-paced scenes. For example, the frame rate can be 16 or other appropriate values. For example, if the storyboard duration in the storyboard attribute information is 3 seconds and the frame rate is 16, the reference video frame number is 3×16 = 48 frames.
[0093] After obtaining the reference video frame number, the first-frame character action pose diagram, the middle-frame character action pose diagram, and the last-frame character action pose diagram can be combined into a sequence in the order of the video frames to obtain an animated GIF of the character actions with the number of video frames being the reference video frame number.
[0094] In the embodiments of the present disclosure, by using the static character diagrams of each storyboard and the storyboard attribute information, the animated GIFs of the character actions of each storyboard are generated based on the first-frame character action pose diagram and the last-frame character action pose diagram, which can accurately control the character actions and postures while keeping the characters consistent, and accurately generate the animated GIFs of the character actions of each storyboard.
[0095] Next, the scene animated GIF and the animated GIFs of the character actions of multiple storyboards can be combined to generate a virtual human video in the same scene. Since the storyboard attribute information may further include the scene to which the storyboard belongs, and different scenes to which the storyboards belong may have different corresponding scene animated GIFs, when combining the scene animated GIF with the animated GIFs of the character actions of multiple storyboards, the target scene animated GIF of each storyboard can be determined from the scene animated GIFs determined according to the scene description and the overall style according to the scene to which the storyboard belongs.
[0096] In addition, since the background of the obtained animated GIF of the human action is not set by the storyboard, in order to improve the accuracy, the animated GIF of the human action can be cropped frame by frame to obtain an animated GIF of the human with a transparent background. Based on this, the animated GIF of the target scene of the storyboard can be merged with the animated GIF of the human with a transparent background obtained by cropping each storyboard to generate the animated GIF of the human in the target scene for each storyboard. Specifically, the animated GIF of the target scene can be used as the lower layer, and the animated GIF of the human with a transparent background for each storyboard can be used as the upper layer; the upper layer and the lower layer are merged to merge the animated GIF of the human with a transparent background for each storyboard onto the layer above the animated GIF of the target scene, generating the animated GIF of the human in the target scene for each storyboard.
[0097] For each storyboard, the first-frame human action pose diagram and the last-frame human action pose diagram of each storyboard can be generated through the static human diagram and the skeleton pose diagram corresponding to the preset action, and then the intermediate-frame human action pose diagram can be predicted based on the first-frame human action pose diagram and the last-frame human action pose diagram. The first-frame human action pose diagram, the intermediate-frame human action pose diagram, and the last-frame human action pose diagram are spliced in chronological order to form an animated GIF of the human action with the number of video frames being the reference number of video frames. And the animated GIF of the human action for each storyboard can be cropped to obtain an animated GIF of the human with a transparent background for each storyboard.
[0098] Furthermore, for each storyboard, the animated GIF of the target scene corresponding to the scene can be determined according to the scene to which each storyboard belongs, and the animated GIF of the target scene is merged with the animated GIF of the human with a transparent background of the storyboard to obtain the animated GIF of the human in the target scene for each storyboard. On this basis, the animated GIFs of the human in the target scene of multiple storyboards can be spliced and merged in sequence according to the arrangement order of the storyboards to obtain a virtual human video with good continuity in the same scene. The arrangement order of the storyboards can be determined according to the logical order of the text content corresponding to the original audio.
[0099] In the embodiments of the present disclosure, the animated GIF of the target scene is determined through the scene to which each storyboard belongs, and the animated GIF of the target scene and the animated GIF of the human with a transparent background for each storyboard are merged in layers to generate the animated GIF of the human in the target scene for each storyboard, which can improve the accuracy of the animated GIF of the human action for each storyboard, and further improve the accuracy of the entire virtual human video.
[0100] In some embodiments, if the shot planning information output by the large language model includes shot attribute information, but the shot attribute information of a certain shot only includes the shot duration and the scene to which the shot belongs, and does not include the character to which the shot belongs and the preset actions of the character to which the shot belongs. At this time, it can be considered that this shot is a pure scene shot without characters. Therefore, this shot can only display the scene animation without generating and displaying characters. Specifically, for a shot without characters, the target scene animation corresponding to each shot can be determined from the scene animations according to the scene to which the shot belongs; among them, the number of video frames of the target scene animation is determined according to the reference number of video frames represented by the product of the shot duration and the frame rate. After obtaining the target scene animation, the target scene animation can be directly determined as the virtual human video in the same scene. It should be noted that there are no characters in this virtual human video.
[0101] Figure 6 schematically shows a flowchart for generating a virtual human video, refer to Figure 6 as shown in, mainly includes the following processes:
[0102] First, the original audio can be subjected to audio recognition to obtain the text content corresponding to the original audio, and multiple sets of skeleton pose diagrams of human body movements can be pre-configured. The text content and multiple sets of skeleton pose diagrams of human body movements are input into the large language model LLM for processing, and the shot planning information corresponding to the text content is output. The shot planning information can include the overall style, scene description, character description, and shot attribute information of each shot; the shot attribute information is, for example, the shot planning timeline, etc.
[0103] Further, according to the overall style and scene description, a static scene graph can be generated based on the Flux model, and sequential frame image prediction can be performed on the static scene graph based on the SVD model to convert the static scene graph into an animated scene graph. Meanwhile, image generation can be performed on the overall style and character description based on the Flux model to generate a static character graph, and the character name can be bound to the static character graph. For each storyboard, the first-frame and last-frame character action pose graphs of each storyboard can be generated based on the PuLID technology and Controlnet-Openpose technology of the Flux model, that is, the first and last frame action pose graphs of the character are generated. After obtaining the first-frame and last-frame character action pose graphs, the first-frame and last-frame character action pose graphs can be used as the input of the ToonCrafter technology for intermediate frame prediction to generate intermediate-frame character action pose graphs; furthermore, by combining the first-frame character action pose graph, the intermediate-frame character action pose graph, and the last-frame character action pose graph, layer merging is performed to generate the animated character action graph of each storyboard. Frame-by-frame matting is performed on the animated character action graph of each storyboard to obtain an animated character graph with a transparent background. The target animated scene graph of each storyboard is determined according to the scene to which each storyboard belongs, and the target animated scene graph of each storyboard is merged with the animated character graph with a transparent background of this storyboard to generate the target scene character animated graph of each storyboard. Moreover, the target scene character animated graphs of multiple storyboards can be spliced according to the arrangement order of the storyboards to generate a virtual human video.
[0104] In the embodiments of the present disclosure, when the user has the original audio, by inputting the original audio into the large language model, the storyboard planning information is output by the large language model. Then, based on the overall style, character description, and storyboard attribute information of each storyboard in the storyboard planning information, the animated character action graph of each storyboard is generated, and the animated scene graph is generated according to the overall style and scene description. By combining the animated scene graph and the animated character action graph, the virtual human video corresponding to the original audio can be produced with one key, which improves the operation efficiency of generating the virtual human video, shortens the production cycle, reduces the production cost, improves the convenience and versatility of generating the virtual human video, and also improves the accuracy of video generation.
[0105] Next, refer to Figure 7 to describe the virtual human video generation device of the exemplary embodiment of the present disclosure. As Figure 7 shown, the virtual human video generation device 700 may include:
[0106] A text content recognition module 710, configured to recognize the corresponding text content according to the original audio and determine the skeleton pose graphs of multiple groups of human body actions;
[0107] The storyboard planning information determination module 720 is configured to process the text content and the skeleton pose diagram through a large language model, and output the storyboard planning information corresponding to the text content; the storyboard planning information includes the overall style, scene description, character description, and the storyboard attribute information of each storyboard;
[0108] The image generation module 730 is configured to generate a scene animation based on the overall style and the scene description, and generate a character static image based on the overall style and the character description;
[0109] The video generation module 740 is configured to generate a character action animation for each storyboard according to the character static image bound to the storyboard and the storyboard attribute information, and combine the scene animation and the character action animations of multiple storyboards to generate a virtual human video in the same scene.
[0110] In an exemplary embodiment of the present disclosure, the generating a scene animation based on the overall style and the scene description includes:
[0111] Using the overall style and the scene description as prompt words to generate a scene static image, and converting the scene static image into the scene animation.
[0112] In an exemplary embodiment of the present disclosure, the converting the scene static image into the scene animation includes:
[0113] Predicting a plurality of sequential frame images located after the scene static image based on the scene static image;
[0114] Combining the plurality of sequential frame images to obtain the scene animation.
[0115] In an exemplary embodiment of the present disclosure, the storyboard planning information includes storyboard attribute information; the storyboard attribute information includes the storyboard duration, the character to which the storyboard belongs, and the preset action of the character to which the storyboard belongs;
[0116] The generating a character action animation for each storyboard according to the character static image bound to the storyboard and the storyboard attribute information includes:
[0117] Obtaining the character name corresponding to the character to which each storyboard belongs, and obtaining the character static image bound to the character name according to the character name;
[0118] According to the preset action of the character to which the storyboard belongs, obtaining the skeleton pose diagram corresponding to the preset action;
[0119] Generating a character action animation for each storyboard with the number of video frames corresponding to the storyboard duration based on the bound character static image and the skeleton pose diagram corresponding to the preset action.
[0120] In an exemplary embodiment of the present disclosure, generating a character action animation for each scene corresponding to the video frame number and the scene duration based on the bound static character image and the skeleton pose image corresponding to the preset action includes:
[0121] Performing image generation based on the static character image and the skeleton pose image corresponding to the preset action to obtain the character action pose image for each scene; the character action pose image includes the first-frame character action pose image and the last-frame character action pose image;
[0122] Combining the first-frame character action pose image and the last-frame character action pose image to generate the character action animation for each scene.
[0123] In an exemplary embodiment of the present disclosure, the performing image generation based on the static character image and the skeleton pose image corresponding to the preset action to obtain the character action pose image for each scene includes:
[0124] Extracting the identity feature information of the static character image;
[0125] Parsing the skeleton pose image to obtain pose control information;
[0126] Fusing the identity feature information and the pose control information to determine the fusion feature, and generating the character action pose image based on the fusion feature;
[0127] Wherein, the character action pose image conforms to the pose control information and retains the identity feature information.
[0128] In an exemplary embodiment of the present disclosure, the combining the first-frame character action pose image and the last-frame character action pose image to generate the character action animation for each scene includes:
[0129] Performing intermediate frame prediction based on the first-frame character action pose image and the last-frame character action pose image to generate intermediate-frame character action pose images;
[0130] Determining the reference video frame number according to the scene duration and the frame rate;
[0131] Based on the first-frame character action pose image, the intermediate-frame character action pose images, and the last-frame character action pose image, forming a character action animation with the reference video frame number.
[0132] In an exemplary embodiment of the present disclosure, the storyboard attribute information further includes the scene to which the storyboard belongs; the combining the scene animation and the character action animations of multiple storyboards to generate a virtual human video in the same scene includes:
[0133] Determining, according to the scene to which the storyboard belongs, a target scene animation corresponding to each storyboard from the scene animations;
[0134] Performing matte extraction on the character action animations to obtain character action animations with a transparent background;
[0135] Merging the target scene animations with the character action animations with a transparent background of each storyboard to generate a target scene character animation for each storyboard;
[0136] Stitching the target scene character animations of the multiple storyboards to generate a virtual human video in the same scene.
[0137] In an exemplary embodiment of the present disclosure, the merging the target scene animations with the character action animations with a transparent background of each storyboard to generate a target scene character animation for each storyboard includes:
[0138] Taking the target scene animations as the lower layer and the character action animations with a transparent background of each storyboard as the upper layer;
[0139] Performing layer merging on the upper layer and the lower layer to generate a target scene character animation for each storyboard.
[0140] In an exemplary embodiment of the present disclosure, the storyboard planning information includes storyboard attribute information; the storyboard attribute information includes the storyboard duration and the scene to which the storyboard belongs; the method further includes:
[0141] Determining, according to the scene to which the storyboard belongs, a target scene animation corresponding to each storyboard from the scene animations; the number of video frames of the target scene animation is determined according to the storyboard duration;
[0142] Determining a virtual human video in the same scene according to the target scene animation.
[0143] It should be noted that the specific details of each module of the virtual human video generation device have been described in detail in the steps of the corresponding virtual human video generation method, so they will not be repeated here.
[0144] Next, refer to Figure 8 to describe the electronic device 800 according to this embodiment of the present disclosure. Figure 8 The shown electronic device 800 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0145] As shown Figure 8 in the figure, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840. The bus 830 may include a data bus, an address bus, and a control bus.
[0146] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of the present specification above. For example, the processing unit 810 may execute the steps as shown Figure 2 in the figure: recognize the corresponding text content according to the original audio, and determine the skeleton pose diagrams of multiple groups of human body movements; process the text content and the skeleton pose diagrams through a large language model, and output the storyboard planning information corresponding to the text content; the storyboard planning information includes the overall style, scene description, character description, and the storyboard attribute information of each storyboard; generate a scene animation based on the overall style and the scene description, and generate a character static image based on the overall style and the character description; generate the character action animation of each storyboard according to the character static image and the storyboard attribute information bound to each storyboard, and combine the scene animation and the character action animations of multiple storyboards to generate a virtual human video in the same scene.
[0147] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.
[0148] The storage unit 820 may further include a program / utility 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0149] The bus 830 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration interface, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0150] The electronic device 800 can also communicate with one or more external devices 900 (such as a keyboard, a pointing device, a Bluetooth device, etc.). Such communication can be carried out through the input / output (I / O) interface 850. Moreover, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0151] It should be noted that in some embodiments of the present disclosure, a computer program product is also provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0152] In one implementation, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing a computer program. The computer-readable storage medium can be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), hard disk drive (HDD), solid state drive (SSD), etc. Exemplarily, the computer program product can be implemented as a non-volatile storage medium storing a computer program, such as read-only memory, Nand Flash, etc.
[0153] In one implementation, the computer program product can be an intangible product containing a computer program. Exemplarily, the computer program product can be implemented as a virtual digital product, such as an executable file storing a computer program, an installation package, and other digital files.
[0154] The code of the computer program can be written in one or more programming languages. Programming languages such as C language, Java, C++, etc. The program code can be executed completely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or completely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (for example, through an Internet connection provided by an operator).
[0155] A computer program can be carried or transmitted by signals such as electrical, magnetic, optical, electromagnetic, infrared, etc. An electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, to cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure.
[0156] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0157] In addition, the above drawings are only schematic illustrations of the processes included in the methods according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0158] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0159] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the content disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0160] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for generating a virtual human video, characterized in that: include: Identify the corresponding text content based on the original audio and determine the skeleton posture graphs of multiple groups of character body movements; Processing the text content and the skeleton posture graph through a large language model, and outputting storyboard planning information corresponding to the text content; the storyboard planning information includes overall style, scene description, character description, and storyboard attribute information of each storyboard; Generate a scene animated image based on the overall style and the scene description, and generate a character static image based on the overall style and the character description; A character action animated image of each storyboard is generated according to the character static image bound to each storyboard and the storyboard attribute information, and the scene animated image and the character action animated images of multiple storyboards are combined to generate a virtual human video in the same scene.
2. The method for generating a virtual human video according to claim 1, characterized in that: The generating a scene animated picture based on the overall style and the scene description includes: The overall style and the scene description are used as prompt words to generate a scene static image, and the scene static image is converted into the scene dynamic image.
3. The method for generating a virtual human video according to claim 2, characterized in that: The converting the scene static image into the scene dynamic image comprises: Predicting a plurality of sequence frame images located after the scene static image based on the scene static image; The multiple sequence frame images are combined to obtain the scene dynamic image.
4. The method for generating a virtual human video according to claim 1, characterized in that: The storyboard planning information includes storyboard attribute information; the storyboard attribute information includes the storyboard duration, the character to which the storyboard belongs, and the preset action of the character to which the storyboard belongs; The step of generating a character action animated image of each storyboard according to the character static image bound to each storyboard and the storyboard attribute information comprises: Obtain the character name corresponding to the character to which each storyboard belongs, and obtain the character static image bound to the character name according to the character name; According to the preset action of the character to which the storyboard belongs, obtaining a skeleton posture diagram corresponding to the preset action; Based on the bound static image of the character and the skeleton posture image corresponding to the preset action, a character action dynamic image of each storyboard corresponding to the number of video frames and the length of the storyboard is generated.
5. The method for generating a virtual human video according to claim 4, characterized in that: The method of generating a character action animated picture of each storyboard corresponding to the number of video frames and the storyboard duration based on the bound character static picture and the skeleton posture picture corresponding to the preset action comprises: An image is generated according to the character static image and the skeleton posture image corresponding to the preset action, to obtain the character action posture image of each storyboard; the character action posture image includes a first frame character action posture image and a last frame character action posture image; The first frame of the character action posture diagram and the last frame of the character action posture diagram are combined to generate a character action animation for each storyboard.
6. The method for generating a virtual human video according to claim 5, characterized in that: The step of generating an image based on the static image of the character and the skeleton posture image corresponding to the preset action to obtain the character action posture image of each storyboard includes: Extracting identity feature information of the static image of the person; Parsing the skeleton posture graph to obtain posture control information; Fusing the identity feature information and the posture control information to determine a fusion feature, and generating the character action posture graph based on the fusion feature; The character action posture diagram conforms to the posture control information and retains the identity feature information.
7. The method for generating a virtual human video according to claim 5, characterized in that: The step of combining the first frame character action posture graph and the last frame character action posture graph to generate a character action dynamic graph of each storyboard includes: Performing intermediate frame prediction based on the first frame character action posture graph and the last frame character action posture graph to generate an intermediate frame character action posture graph; Determine the number of reference video frames based on the length of the storyboard and the frame rate; Based on the first frame character action posture diagram, the middle frame character action posture diagram and the last frame character action posture diagram, a character action animation having a video frame number equal to the reference video frame number is formed.
8. A virtual human video generation device, characterized in that: include: A text content recognition module, used to recognize the corresponding text content according to the original audio, and determine the skeleton posture graphs of multiple groups of character body movements; A storyboard planning information determination module is used to process the text content and the skeleton posture graph through a large language model, and output the storyboard planning information corresponding to the text content; the storyboard planning information includes overall style, scene description, character description and storyboard attribute information of each storyboard; An image generation module, configured to generate a scene dynamic image based on the overall style and the scene description, and to generate a character static image based on the overall style and the character description; The video generation module is used to generate a character action dynamic image for each storyboard based on the character static image bound to the storyboard and the storyboard attribute information, and to combine the scene dynamic image and the character action dynamic images of multiple storyboards to generate a virtual human video in the same scene.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a virtual human video according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: processor; as well as A memory for storing executable instructions; Wherein, the processor is configured to execute the virtual human video generation method described in any one of claims 1 to 7 by executing the executable instructions.
Citation Information
Cited By
Generation method of digital human broadcast video, electronic equipment, medium and program product
CN120935429A
Video generation method and device, electronic equipment, storage medium and program product
CN121000952A
AI video content integrated generation method and system based on script structuring
CN122053940A