Method and device for generating digital human
The local image of digital people is generated and integrated through computing devices, which solves the problem of inconsistent quality of the face image uploaded by users and achieves the high-quality fusion effect of digital people's images.
Patent Information
- Application Number
- CN202410101094.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-07
- Filing Date
- 2024-01-24
- Publication Date
- 2025-05-09
AI Technical Summary
In the existing digital human image customization solution, the quality of the face images uploaded by users is inconsistent, resulting in poor integration with the digital human template.
The local image of a digital person that meets the fusion requirements is generated through computing devices, and the local image generation model and fusion model are used for backpropagation, and the model parameters are updated to improve image quality.
The integration effect of local image and digital human templates has been improved, and the image quality and fusion matching of digital human image have been improved.
Smart Images

Figure CN119967225A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 7, 2023, with application number 202311473683.8 and invention name “Image generation method, device, computing device cluster and storage medium”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The embodiments of the present application relate to the field of computers, and in particular to a method and device for a digital human. Background Art
[0003] Digital human image customization refers to the use of computer technology and design tools to allow users to customize a digital human image according to their preferences and needs. With the application of digital human image customization in scenarios such as social platforms, games, and e-commerce, users can customize a rich digital human image in a virtual environment. For example, users can select and adjust the appearance, style, and characteristics of the digital human image according to their own aesthetics and personal preferences to express a unique personality and identity.
[0004] In the current digital human image customization solution, users first need to upload the digital human face image, select the digital human template in the digital human template library, and then fuse the face image with the digital human template to obtain the full-body digital human image corresponding to the face image. Due to the uneven quality of the face images uploaded by users in the current solution, the fusion effect of the face image and the digital human template is poor. Summary of the invention
[0005] The embodiment of the present application provides a method for generating a digital human, wherein a computing device can add restrictions to the generation process of a partial image, generate a partial image of a digital human that meets the fusion requirements, thereby improving the fusion effect of the partial image and the digital human template. The embodiment of the present application also provides a digital human generation device, a computing device, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the method for generating a digital human.
[0006] In a first aspect, an embodiment of the present application provides a method for generating a digital human, which can be executed by a computing device, or by a component of the computing device, such as a processor, a chip or a chip system of the computing device, or by a logic module or software that can realize all or part of the functions of the computing device. The method provided in the first aspect includes: the computing device obtains a first partial image generation condition, the first partial image generation condition is used to indicate the characteristics of the partial image of the first digital human to be generated, the type of the partial image of the first digital human includes one or more of a face, clothes, and a hairstyle, and the type of the first partial image generation condition includes one or more of text, picture, and voice. The computing device inputs the first partial image generation condition into a partial image generation model, and performs forward calculation through the partial image generation model to generate at least one first partial image of the digital human. The computing device inputs the at least one first partial image and the first digital human template into a fusion model to generate at least one image of the digital human, and the at least one image of the digital human is the result of the fusion of the at least one first partial image and the first digital human template. The computing device performs back propagation on the local image generation model based on the at least one first digital human image and the at least one first target digital human image to update the parameters of the local image generation model, wherein the at least one first target digital human image is a digital human image corresponding to the first local image generation condition.
[0007] In the embodiment of the present application, the computing device can generate a partial image of a digital human based on the partial image generation condition and the partial image generation model. Compared with the partial image uploaded by the user, the computing device generates the partial image based on the partial image generation condition, thereby improving the image quality of the partial image to be fused. At the same time, the computing device can train the partial image generation model based on the fused digital human image and the target digital human image, thereby further improving the image quality of the partial image generated by the partial image generation model, so that the generation result of the partial image generation model meets the generation requirements of the digital human.
[0008] In one possible implementation, the computing device inputs the second partial image generation condition into the trained partial image generation model, the second partial image generation condition is used to indicate the characteristics of the partial image of the second digital person to be generated, the type of the partial image of the second digital person includes one or more of face, clothing, and hairstyle, and the type of the second partial image generation condition includes one or more of text, picture, and voice, to generate at least one second partial image. The computing device inputs the at least one second partial image and the second digital human template into the fusion model to generate at least one second digital human image, and the at least one second digital human image is the result of the fusion of the at least one second partial image and the second digital human template.
[0009] In the embodiment of the present application, the computing device can input the second partial image generation conditions into the trained partial image generation model, and generate the model to generate the second partial image, thereby improving the image quality of the second partial image and further improving the fusion effect of the digital human.
[0010] In one possible implementation, the restriction condition is an explicit restriction condition. In a process in which the computing device generates at least one local image using a local image generation model based on the local image generation condition and the restriction condition, the computing device generates at least one local image using a local image generation model based on the local image generation condition and the explicit restriction condition. The local image satisfies the explicit restriction condition. The explicit restriction condition is used to indicate the explicit requirements of the local image fusion model for the local image.
[0011] In the embodiment of the present application, the computing device generates a local image generation model based on explicit restriction constraints, so that the local image generated by the local image generation model meets the explicit restriction conditions, thereby improving the fusion matching of the local image and the digital human template.
[0012] In a possible implementation, when the local image generated by the local image generation model does not meet the restriction condition, that is, the local image generation condition conflicts with the restriction condition, and the local image generation model cannot generate a local image that meets the restriction condition, the computing device sends a prompt message to the user device, and the prompt message is used to prompt the user device to modify the local image generation condition. The prompt message is also used to prompt the explicit restriction condition content that the current local image generation condition does not meet.
[0013] In the embodiment of the present application, when there is a conflict between the local image generation condition and the restriction condition, the computing device can send a prompt message to the user, prompting the user device to modify the local image generation condition, thereby improving the efficiency of the computing device in generating the local image.
[0014] In a possible implementation, the computing device performs one or more of the following operations on the second partial image based on the restriction condition: adding, modifying and deleting features of the second partial image, so that the second partial image satisfies the restriction condition.
[0015] In the embodiment of the present application, the computing device can also add, modify and delete the features of the second partial image based on the restriction conditions, thereby improving the image quality of the partial image and further improving the fusion effect of the digital human.
[0016] In one possible implementation, the restriction condition is an implicit condition. In the process of generating at least one local image using a local image generation model based on the local image generation condition and the restriction condition, the computing device generates at least one local image using a local image generation model based on the local image generation condition and the implicit restriction condition, and the vector features of the local image generation model satisfy the implicit restriction condition.
[0017] In the embodiment of the present application, the computing device constrains the vector features of the local image generation model in the process of generating the local image based on implicit constraints, so that the local image generation model generates a local image that meets the requirements, thereby improving the fusion matching of the local image and the digital human template.
[0018] In a possible implementation, before the computing device generates at least one partial image using the partial image generation model, the computing device may train the partial image generation model. Specifically, the computing device updates the model parameters of the partial image generation model based on the calculation result of the loss function, and the calculation result of the loss function is determined based on the digital human generated by the partial image fusion model and the target digital human.
[0019] In the embodiment of the present application, the computing device can update the model parameters of the local image generation model based on the calculation result of the loss function, thereby improving the accuracy of the local image generation model in generating the local image based on the local image generation conditions.
[0020] In a possible implementation, before the computing device generates the partial image based on the partial image generation model, the computing device may train the partial image fusion model. Specifically, the computing device can update the model parameters of the partial image fusion model based on the calculation result of the loss function, and the calculation result of the loss function is determined based on the digital human generated by the partial image fusion model and the target digital human.
[0021] In the embodiment of the present application, the computing device can update the model parameters of the local image fusion model based on the calculation result of the loss function, thereby improving the fusion effect of the local image fusion model on the local image and the digital human template.
[0022] In one possible implementation, after the computing device generates at least one local image using the local image generation model, the computing device filters multiple local images based on screening indicators, where the screening indicators include quality screening indicators and algorithm restriction screening indicators corresponding to the fusion model, wherein the quality screening indicators are used to indicate the quality requirements of the screening module for the local image, and the algorithm restriction screening indicators are used to indicate the requirements of the local image fusion model for the local image.
[0023] In the embodiment of the present application, the computing device can filter the local image generated by the local image generation model based on a variety of filtering indicators, thereby improving the quality of the local image.
[0024] In a possible implementation, the computing device displays the filtered partial images on a display interface, and the user can manually select a partial image as the partial image to be fused based on the filtering results displayed on the display interface.
[0025] In the embodiment of the present application, the user can manually screen the partial images screened by the screening module to determine the partial images to be fused, thereby improving the feasibility of the solution.
[0026] In one possible implementation, the computing device generates quality scores of multiple partial images based on a quality assessment algorithm, and the quality scores include one or more of the following: visible light index score, posture index score, posture skin texture resolution, and facial feature clarity. The computing device selects at least one partial image from the multiple partial images based on a screening index, and the screening index includes that the quality score of the partial image is greater than or equal to a threshold.
[0027] In the embodiment of the present application, the computing device generates quality scores of multiple local images based on a quality assessment algorithm, and filters the local images based on the quality scores corresponding to different filtering indicators, thereby improving the quality of the local images.
[0028] In a possible implementation, before the computing device uses the partial image fusion model to fuse the partial image with the digital human template in the digital human template library, it selects a digital human template from the digital human template library, and the digital human template is used to fuse the partial image. For example, when the partial image is a face image, the user needs to select a digital human template from the digital human template library, and the digital human template includes the full body image of the digital human.
[0029] In the embodiment of the present application, the user can select a digital human template from the model image library, and generate a digital human based on the fusion of the selected digital human template, thereby improving the richness of the digital human.
[0030] In a possible implementation, the computing device composes a sequence of actions for the digital human based on the action template of the digital human and generates a video corresponding to the digital human, wherein the action template of the digital human includes different actions of the digital human. The computing device can play the video of the digital human on the display interface.
[0031] The computing device in the embodiment of the present application can choreograph the actions of the fused digital human and generate a video of the digital human, thereby enhancing the richness of the application scenarios of the digital human.
[0032] In a possible implementation, for professional users, the local image generation condition may also be a vectorized feature, which may directly indicate the feature of the local image of the digital human to be generated. The local image generated by the local image generation model of the computing device includes the local features corresponding to the vectorized feature generation.
[0033] The local image generation conditions in the embodiments of the present application can also be vectorized features, and the local image generation model can directly generate the local image based on the vectorized features, thereby improving the efficiency of generating the local image from the local image.
[0034] In one possible implementation, the computing device provides multiple input templates of partial image generation conditions, and the user selects the partial image generation condition from the input templates. For example, the input templates include face, hairstyle, and clothing, etc., wherein the face input template includes gender, age, and skin color, etc., the hairstyle input template includes long hair, short hair, and hair color, etc., and the clothing input template includes shirt, pants, skirt, and clothing color, etc.
[0035] In the embodiment of the present application, the computing device may provide an input template of the local image generation conditions, and the user may select the local image generation conditions, thereby improving the efficiency of the local image generation model in generating the local image according to the local image generation conditions.
[0036] In a second aspect, an embodiment of the present application provides a digital human generation device, the device comprising a transceiver unit and a processing unit. The transceiver unit is used to obtain a first partial image generation condition, the first partial image generation condition is used to indicate the characteristics of the partial image of the first digital human to be generated, the type of the partial image of the first digital human includes one or more of a face, clothes, and a hairstyle, and the type of the first partial image generation condition includes one or more of text, picture, and voice. The processing unit is used to input the first partial image generation condition into a partial image generation model, and perform forward calculation through the partial image generation model to generate at least one first partial image of a digital human. The processing unit is also used to input at least one first partial image and a first digital human template into a fusion model to generate at least one image of a digital human, and the at least one image of a digital human is the result of the fusion of at least one first partial image and the first digital human template. The processing unit is also used to perform back propagation on the partial image generation model based on at least one first digital human image and at least one first target digital human image to update the parameters of the partial image generation model, and at least one first target digital human image is a digital human image corresponding to the first partial image generation condition.
[0037] In a possible implementation, the processing unit is further used to input the second partial image generation condition into the trained partial image generation model, the second partial image generation condition is used to indicate the characteristics of the partial image of the second digital human to be generated, the type of the partial image of the second digital human includes one or more of face, clothing, and hairstyle, and the type of the second partial image generation condition includes one or more of text, picture, and voice, so as to generate at least one second partial image. The at least one second partial image and the second digital human template are input into the fusion model to generate at least one second digital human image, and the at least one second digital human image is the result of the fusion of the at least one second partial image and the second digital human template.
[0038] In one possible implementation, the processing unit is specifically used to input the second local image generation conditions and restriction conditions into a trained local image generation model to generate at least one second local image. The restriction conditions are used to indicate the requirements for the local image when the fusion model fuses the digital human template and the local image. The types of restriction conditions include text.
[0039] In a possible implementation manner, the transceiver unit is further configured to send a prompt message to the user device when the local image generated by the local image generation model does not meet the restriction condition, and the prompt message is configured to prompt the user device to modify the local image generation condition.
[0040] In a possible implementation manner, the processing unit is further configured to perform one or more of the following operations on the second partial image based on the restriction condition: adding, modifying and deleting features of the second partial image.
[0041] In a possible implementation, the processing unit is further used to filter at least one second partial image based on a filtering index, where the filtering index includes a quality filtering index and an algorithm restriction filtering index corresponding to the fusion model, and the algorithm restriction filtering index is used to indicate the requirements of the fusion model for the second partial image.
[0042] In a possible implementation, the processing unit is further configured to generate a quality score of at least one second partial image based on a quality assessment algorithm, the quality score comprising one or more of the following: a visible light index score, a posture index score, a posture skin texture resolution, and a facial feature clarity. The processing unit is specifically configured to filter out a plurality of partial images from at least one second partial image based on a screening index, the screening index comprising that the quality score of the partial image is greater than or equal to a threshold.
[0043] In a possible implementation manner, the processing unit is further used to select a digital human template from a digital human template library, and the digital human template is used to fuse the second partial image.
[0044] In a possible implementation, the processing unit is further used to arrange an action sequence for the digital human based on a digital human action template to generate a video corresponding to the digital human, wherein the digital human action template includes different digital human actions.
[0045] In a third aspect, an embodiment of the present application provides a computing device, comprising a processor, the processor being coupled to a memory, the processor being used to store instructions, and when the instructions are executed by the processor, the computing device executes the method described in the first aspect or any possible implementation manner of the first aspect.
[0046] In a fourth aspect, an embodiment of the present application provides a computing device cluster, the computing device cluster includes one or more computing devices, the computing device includes a processor, the processor is coupled to a memory, the processor is used to store instructions, when the instructions are executed by the processor, the computing device cluster executes the method described in the first aspect or any possible implementation method of the first aspect.
[0047] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed, the computer executes the method described in the first aspect or any possible implementation manner of the first aspect.
[0048] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed, the computer implements the method described in the first aspect or any possible implementation method of the first aspect.
[0049] It can be understood that the beneficial effects that can be achieved by any of the digital human generation devices, computing devices, computing device clusters, computer-readable media or computer program products provided above can refer to the beneficial effects in the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of the system architecture of a digital human generation system provided in an embodiment of the present application;
[0051] Figure 2 A schematic diagram of a flow chart of a method for generating a digital human provided in an embodiment of the present application;
[0052] Figure 3 A schematic diagram of a flow chart of another method for generating a digital human provided in an embodiment of the present application;
[0053] Figure 4 A schematic diagram of a flow chart of another method for generating a digital human provided in an embodiment of the present application;
[0054] Figure 5A schematic diagram of a flow chart of a computing device screening layout image provided in an embodiment of the present application;
[0055] Figure 6 A schematic diagram of training a local image generation model provided in an embodiment of the present application;
[0056] Figure 7 A schematic diagram of a digital human generation device provided in an embodiment of the present application;
[0057] Figure 8 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0058] Fig. 9 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0059] Fig.10 A schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The embodiments of the present application provide a method and device for generating a digital image, which are used to enhance the fusion effect of a partial image of a digital human and a digital human template.
[0061] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0062] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0063] First, some terms involved in the embodiments of the present application are introduced to facilitate technical personnel in this field to understand the technical solution.
[0064] Generative artificial intelligence (AIGC) is an artificial intelligence technology based on deep learning neural network models. It can automatically generate language, images, audio and other content by learning large amounts of data and simulating human thinking.
[0065] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.
[0066] See also Figure 1 , Figure 1 A schematic diagram of the system architecture of a digital human generation system provided as an example of this application. Figure 1 In the example shown, the digital human generation system 10 includes an input module 101, a generation module 102, a screening module 103, a digital human template library 104, a fusion module 105, a restriction module 106 and a result display module 107. The specific functions of each module in the digital human generation system 10 are described below.
[0067] The input module 101 is used to receive partial image generation conditions, which are used to indicate the characteristics of the partial image of the digital person to be generated. The types of partial image generation conditions include one or more of text, picture and voice. For example, the user can input text as the partial image generation condition through the input module 101, and the text input condition is such as "youth, male, sunny". For another example, the user can also input text and picture as the partial image generation condition through the input module 101, and the text and picture input conditions are such as a male picture and the text "wide face".
[0068] The generation module 102 is used to generate a partial image of a digital human based on a partial image generation model, and the partial image generation model is a generative artificial intelligence AIGC model. Specifically, the generation module 102 generates a partial image of a digital human using the partial image generation model according to the partial image generation conditions and restriction conditions, wherein the type of the partial image includes one or more of a face, clothes, and a hairstyle. Among them, the restriction condition is used to indicate the requirements of the fusion module 105 for the partial image generated by the generation module 102, and the restriction condition is a partial image generation condition that the user does not perceive.
[0069] The screening module 103 is used to screen the partial images of the digital human generated by the generation module 102 based on the screening index. The screening process is that the screening module 103 automatically screens based on the screening index. The screening index includes a quality screening index and an algorithm restriction screening index. The quality screening index is used to indicate the quality requirements of the screening module 103 for the partial images generated by the generation module 102, and the algorithm restriction screening index is used to indicate the quality requirements of the fusion module 105 for the partial images generated by the generation module 102.
[0070] It is understandable that, since the user may not perceive the screening process of the screening module 103 , the screening module 103 may also be integrated into the generating module 102 as a submodule of the generating module 102 , without any specific limitation.
[0071] The digital human template library 104 is used to provide digital human templates that are integrated with the partial image of the digital human. The digital human template library 104 has multiple digital human templates pre-set, and the user can select a digital human template to be integrated with the partial image through the digital human template library 104. The digital human template is a pre-set digital human image. For example, the digital human template includes a face image, a body image, and a clothing image.
[0072] The fusion module 105 is used to fuse the partial image with the digital human template to obtain a fused digital human. Specifically, based on the partial image fusion model, the fusion module 105 replaces the partial image generated by the generation module 102 with the digital human template selected in the digital human template library 104, thereby obtaining a fused digital human. For example, the fusion module 104 replaces the facial image generated by the generation module 102 with the facial image of the digital human template, thereby achieving a face-changing operation on the digital human template.
[0073] The restriction module 106 is used to generate restriction conditions for the local image generation model. Specifically, the restriction module 106 obtains the requirements of the local image fusion model of the fusion module 105 for the local image to be fused, and generates restriction conditions for the local image generation model based on the requirements. Among them, the restriction conditions include explicit restriction conditions and implicit restriction conditions. Explicit restriction conditions refer to restriction conditions that can be reflected on the local image. For example, explicit restriction conditions restrict the local image generation model from generating the image of a face wearing glasses. Implicit restriction conditions refer to restrictions on the vector features of the local image generation model.
[0074] The result display module 107 is used to display the digital human image obtained by the fusion module 105. The user can also modify the partial image of the digital human through the result display module 107, and the user can also choreograph actions for the digital human image through the result display module 107 to obtain a video related to the digital human image.
[0075] It is understandable that the digital human generation system 10 provided in the embodiment of the present application can be applied to a variety of digital human application scenarios, without specific limitation, such as digital human game scenarios, digital human anchor live broadcast scenarios, digital human teacher and lecture scenarios, and digital human customer service scenarios.
[0076] based on Figure 1The digital human generation system 10 shown in the figure, the present application also provides a digital human generation method. The digital human generation method provided in the present application embodiment is introduced below in conjunction with the embodiment.
[0077] See also Figure 2 , Figure 2 A schematic diagram of a method for generating a digital human provided in an embodiment of the present application. Figure 2 In the example shown, the method includes the following steps:
[0078] 201. The computing device obtains a first partial image generation condition, where the first partial image generation condition is used to indicate features of a partial image of a digital person to be generated, and the type of the first partial image generation condition includes one or more of text, picture, and voice.
[0079] The computing device receives the first partial image generation condition, which is the partial image generation requirement provided by the user. The first partial image generation condition is used to indicate the characteristics of the partial image of the digital human to be generated, and the type of the partial image includes one or more of face, clothes, and hairstyle. Specifically, the user inputs the first partial image generation condition on the display interface of the digital human generation system 10, and the type of the first partial image generation condition includes one or more of text, picture, and voice.
[0080] When the type of the first partial image generation condition is text, the user can enter the text requirements of the partial image on the display interface. When the type of the first partial image generation condition is text and picture, the user can upload the picture and enter the text requirements of the partial image on the display interface. When the type of the first partial image generation condition is voice, the computing device can convert the voice into text as the first partial image generation condition.
[0081] For example, the user may input the first partial image generation condition of text type, such as "old, female, kind". For another example, the user may also input the first partial image generation condition of text and picture type, such as a picture of a young woman and the text "old, kind". For another example, the user may input a voice, such as "please generate a kind, elderly female face image".
[0082] It should be noted that the first local image generation condition in the embodiment of the present application refers to the local image generation condition input into the local image generation model during the training phase of the local image generation model, and correspondingly, the second local image generation condition refers to the local image generation condition input into the local image generation model during the reasoning or application phase of the local image generation model.
[0083] In some descriptions of the embodiments of the present application, the first local image generation condition and the second local image generation condition are not distinguished and are collectively referred to as local image generation conditions. Correspondingly, the first local image and the second local image are collectively referred to as local images, and the first digital human module and the second digital human template are collectively referred to as digital human templates.
[0084] See also Figure 3 , Figure 3 A schematic diagram of a method for generating a digital human provided in an embodiment of the present application. Figure 3 In the example shown, the digital human generation system 10 receives partial image generation conditions. Specifically, the user inputs the partial image generation conditions through the input module 101. The partial image generation conditions can indicate the user's requirements for the partial image.
[0085] exist Figure 3 In the example shown, the partial image is, for example, a human face image. If the user inputs the text-type partial image generation condition as "youth, male, sunny", the user wishes to generate a sunny young male face image. For another example, if the user uploads a picture-type partial image generation condition as a male face photo, and inputs the text-type partial image generation condition as "wide face", the user wishes to generate a wide-faced male face photo based on the male face photo.
[0086] In a possible implementation, for professional users, the local image generation condition may also be a vectorized feature, which may directly indicate the features of the local image of the digital human to be generated to the generation module 102. The local image generated by the local image generation model of the generation module 102 includes the local features corresponding to the vectorized feature generation.
[0087] In a possible implementation, the computing device may also provide the user with multiple input templates of partial image generation conditions, and the user may select a partial image generation condition from the input templates. For example, the input templates include face, hairstyle, and clothing, where the face input template includes gender, age, and skin color, the hairstyle input template includes long hair, short hair, and hair color, and the clothing input template includes shirt, pants, skirt, and clothing color.
[0088] See also Figure 4 , Figure 4 A flowchart of another method for generating a digital human provided in an embodiment of the present application. Figure 4 In the example step b shown, the input module 101 receives local image generation conditions and sends the local image generation conditions to the generation module 102. The local image generation conditions include text, pictures, voice and vectorized features.
[0089] exist Figure 4In the example shown, the partial image generation condition received by the input module 101 may also be a partial image generation condition selected by the user based on an input template. For example, the input module 101 provides a face image template, a hairstyle image template, and a clothing image template, etc., from which the user can select a face image template as the type of partial image to be generated. Further, after the user selects a face image template, the user can further select partial image generation conditions such as gender, age, and skin color.
[0090] 202. The computing device inputs the first partial image generation condition into the partial image generation model, performs forward calculation through the partial image generation model, and generates a first partial image of at least one digital human.
[0091] After the computing device obtains the first partial image generation condition, the computing device inputs the first partial image generation condition into a partial image generation model, and uses the partial image generation model to generate a first partial image of at least one digital human.
[0092] In one possible implementation, during the inference phase of the local image generation model, after the computing device obtains the second local image generation condition, the second local image generation condition is input into the trained local image generation model, where the second local image generation condition is used to indicate features of the local image of the second digital person to be generated, where the type of the local image of the second digital person includes one or more of a face, clothes, and a hairstyle, and where the type of the second local image generation condition includes one or more of text, a picture, and a voice, so as to generate at least one second local image.
[0093] In a possible implementation, when the local image generation model of the computing device generates at least one local image of a digital human, the computing device generates the local image based on the local image generation model and the restriction condition. The restriction condition is used to indicate the requirements for the local image when the local image fusion model fuses the local image with the digital human template in the digital human template library. The requirements include requirements based on image quality and requirements based on application scenarios. The requirements based on image quality are such as no occlusion, clarity, etc., and the requirements based on application scenarios are such as front view, portrait without bangs, etc.
[0094] Please continue reading Figure 3 ,exist Figure 3 In the example shown, the partial image to be generated is a human face image. When the user inputs the text type of partial image generation condition "youth, male, sunny", the generation module 102 of the computing device generates multiple facial images of sunny young males that meet the partial image generation conditions based on the partial image generation model.
[0095] exist Figure 3In the example shown, when the partial image generation condition is a male face photo of picture type and a text description "wide face", the generation module 102 of the computing device generates one or more wide-faced young male face images based on the male face photo based on the partial image generation model.
[0096] In a possible implementation, before the computing device generates at least one partial image based on the restriction condition, the computing device needs to obtain the restriction condition. Specifically, the generation module 102 of the computing device receives the restriction condition sent by the restriction module 106, which is the restriction condition determined by the restriction module 106 based on the constraint of the fusion module 105 on the partial image.
[0097] It is understandable that the process of the generation module 102 of the computing device receiving the restriction conditions is not perceptible to the user. The user can only perceive the restriction conditions when there is a conflict between the local image generation conditions and the restriction conditions, and the generation module 103 of the computing device cannot generate a local image that meets the local image generation conditions and prompts the user to modify the local image generation conditions.
[0098] Please continue reading Figure 4 ,exist Figure 4 In step a and step c of the example shown, the fusion module 105 sends the constraints of the local image fusion model on the local image to the restriction module 106. The restriction module 106 generates restriction conditions based on the constraints of the local image fusion model on the local image, and sends the restriction conditions to the generation module 102. The generation module 102 needs to meet the restriction conditions when generating the local image.
[0099] exist Figure 4 In the example shown, when the restriction module 105 sends the restriction conditions to the generation module 102, the restriction conditions can be added to the local image generation model through the training stage, or can be feature fused with the local image generation conditions in the reasoning stage and then input into the local image generation model to ensure that the generation model is controlled by both the local image generation conditions and the restriction conditions.
[0100] Among them, in the feature fusion method with the local image generation conditions at the inference stage, the computing device directly converts the restriction conditions into feature codes, and then splices or fuses the restriction conditions with the local image generation conditions into a text description through semantic fusion, and then inputs it into the local image generation model. For example, the local image generation condition is a short-haired square-faced male, and the restriction condition is a frontal image with an unobstructed forehead. After semantic fusion, it is a frontal photo of a short-haired square-faced male with an unobstructed forehead. The semantic feature is then encoded and input into the local image generation model.
[0101] The constraints in the embodiments of the present application include explicit constraints and implicit constraints, wherein the explicit constraints refer to constraints that can be directly reflected in the features of the local image, for example, the explicit constraints restrict the local image generation model to generate a side face image. The implicit constraints refer to the restrictions on the vector features of the local image generation model, and the implicit constraints are not directly reflected in the visual features of the local image.
[0102] In one possible implementation, when the constraint is an explicit constraint, the computing device generates at least one local image using a local image generation model based on the local image generation condition and the explicit constraint, and the local image satisfies the explicit constraint. The explicit constraint is used to indicate the explicit requirements of the local image fusion model for the local image.
[0103] For example, if the explicit restriction condition received by the generation module 102 of the computing device is to restrict the generation of a face image of a person wearing glasses, then no matter what the partial image generation condition of the user is, the generation module 102 of the computing device will not generate a face image of a person wearing glasses. For example, if the explicit restriction condition received by the generation module 102 of the computing device is to restrict the generation of a face image of a side face, then no matter what the partial image generation condition of the user is, the generation module 102 of the computing device will not generate a face image of a side face.
[0104] In one possible implementation, when the constraint is an implicit constraint, the computing device generates at least one local image using a local image generation model based on the local image generation condition and the implicit constraint, and the vector features of the local image generation model satisfy the implicit constraint.
[0105] In a possible implementation, when the local image generated by the local image generation model of the generation module 102 does not meet the restriction condition, the computing device sends a prompt message to the user device, and the prompt message is used to prompt the user device to modify the local image generation condition. The prompt message is also used to prompt the explicit restriction condition that the current local image generation condition does not meet, and to prompt suggestions for modifying the local image generation condition.
[0106] For example, the partial image generation condition received by the input module 101 of the computing device is "wearing glasses, young, male", and the explicit restriction condition obtained by the generation module 102 is to restrict the generation of a face image wearing glasses. If the partial images generated by the partial image generation model of the generation module 102 based on the partial image generation condition do not meet the restriction condition, the computing device sends a prompt message to the user device, such as "wearing glasses does not meet the restriction condition, please modify". The user modifies the partial image generation condition based on the prompt message.
[0107] In a possible implementation, the computing device performs one or more of the following operations on the partial image based on the restriction condition: adding, modifying and deleting the features of the second partial image. That is, the computing device can operate the partial image based on the partial image generation condition so that the generated partial image meets the restriction condition.
[0108] In one possible implementation, after the computing device generates a local image based on the local image generation conditions and restriction conditions, the filtering module 103 of the computing device filters multiple local images based on filtering indicators, where the filtering indicators include quality filtering indicators and algorithm restriction filtering indicators corresponding to the fusion model, wherein the quality filtering indicators are used to indicate the quality requirements of the filtering module for the local image, and the algorithm restriction filtering indicators are used to indicate the requirements of the local image fusion model for the local image.
[0109] Please continue reading Figure 4 ,exist Figure 4 In step d to step e of the example shown, the generating module 102 generates one or more partial images based on the partial image generating condition, and obtains at least one partial image based on the screening index. The generating module 102 displays the screened partial images to the user, wherein each partial image has a unique identity. The user can further manually screen the displayed partial images based on the identity to select the partial images to be fused.
[0110] In a possible implementation, when the computing device screens multiple partial images based on the screening index, the computing device generates quality scores of the multiple partial images based on the quality assessment algorithm, and the quality scores include visible light index score, posture index score, posture skin texture resolution, and facial features clarity, etc. The computing device screens out at least one partial image from the multiple partial images based on the screening index, and the screening index includes that the quality score of the partial image is greater than or equal to a threshold.
[0111] In the process of a computing device generating quality scores of a plurality of partial images based on a quality assessment algorithm, when the partial image is a face image, the quality assessment algorithm is, for example, a face quality assessment (FQA) algorithm.
[0112] It is understandable that the user is unaware of the process in which the filtering module 103 of the computing device filters the partial images generated by the generating module 102 , and the partial images displayed to the user by the computing device are the partial images that have been automatically filtered and meet the filtering criteria.
[0113] See also Figure 5 , Figure 5 A schematic diagram of a flow chart of a computing device screening layout image provided in an embodiment of the present application. Figure 5In the example shown, after the generating module 102 generates the partial images, the screening module 103 calculates the quality score of each partial image based on the quality evaluation algorithm, and quantitatively screens the partial images based on the quality screening index and the algorithm restriction screening index, that is, the screening module 103 screens out the partial images whose quality scores satisfy both the quality screening index and the algorithm restriction screening index. The user can further manually screen the partial images screened out by the screening module 103 to manually screen out the partial images to be fused.
[0114] exist Figure 5 In the example shown, the quality screening index depends on the settings in the screening module 103, and the user can modify the quality screening index, while the algorithm restriction screening index depends on the local image fusion model. The quality screening index and the algorithm restriction screening index may have some of the same indicators, which are not specifically limited.
[0115] In one possible implementation, after the filtering module 103 of the computing device obtains the partial image based on the filtering index, the filtered partial image is displayed on the display interface. The user can manually select a partial image as the partial image to be fused based on the filtering results displayed on the display interface.
[0116] Please continue reading Figure 3 ,exist Figure 3 In the example shown, the computing device generates multiple facial images of young men based on local image generation conditions, and the user needs to select one of the multiple facial images of young men as the facial image to be fused.
[0117] 203. The computing device inputs at least one first partial image and a digital human template into a fusion model to generate at least one image of the digital human.
[0118] In the training phase of the partial image generation model, the computing device inputs at least one first partial image and the digital human template into the fusion model to generate at least one image of the digital human. The at least one image of the digital human is the result of the fusion of at least one first partial image and the first digital human template. The fusion result can be an image or a video, which is not specifically limited.
[0119] In one possible implementation, during the inference phase of the local image generation model, the computing device inputs at least one second local image and a second digital human template into a fusion model to generate at least one second digital human image, wherein the at least one second digital human image is the result of the fusion of the at least one second local image and the second digital human template.
[0120] Specifically, after the computing device determines the partial image to be fused, the fusion module 105 of the computing device fuses the partial image with the digital human template in the digital human template library based on the partial image fusion model to obtain a digital human corresponding to the partial image. For example, when the partial image to be fused is a face image, the computing device fuses the face image with the digital human template selected from the digital human template library to obtain a digital human including the face image.
[0121] In a possible implementation, before the computing device fuses the partial images, the user needs to select a digital human template from a digital human template library, and the digital human template is used to fuse the partial images. For example, when the partial image is a face image, the user needs to select a digital human template from the digital human template library, and the digital human template includes the full body image of the digital human.
[0122] Please continue reading Figure 3 ,exist Figure 3 In the example shown, after the computing device determines the facial image to be fused, the user needs to select the digital human template to be fused in the digital human template library. For example, the user can select the digital human template wearing a suit. The fusion module 105 of the computing device fuses the partial image to be fused with the digital human template based on the partial image fusion model. For example, the computing device fuses the facial image of a young male to be fused with the digital human template wearing a suit, and thus obtains a young male digital human wearing a suit.
[0123] Please continue reading Figure 4 ,exist Figure 4 In step f to step g of the example shown, after the user manually screens and determines the partial image to be fused, a digital human template is selected from the digital human template library, and the fusion module 105 fuses the partial image to be fused with the digital human template to obtain a digital human containing the partial image. After the fusion module 105 generates the fused digital human, it sends the digital human to the result display module 107.
[0124] 204. The computing device performs back-propagation on the local image generation model based on at least one first digital human image and at least one first target digital human image to update the parameters of the local image generation model.
[0125] In the training phase of the local image generation model, the computing device performs back propagation on the local image generation model based on at least one first digital human image and at least one first target digital human image to update the parameters of the local image generation model, wherein the at least one first target digital human image is a digital human image corresponding to the first local image generation condition. Specifically, in the process of training the local image generation model, the computing device can update the model parameters of the local image generation model based on the calculation result of the loss function, and the calculation result of the loss function is determined based on the first digital human and the first target digital human generated by the local image fusion model.
[0126] It is understandable that the computing device can also train the local image fusion model. For example, the computing device can also update the model parameters of the local image generation model based on the calculation results of the loss function.
[0127] See also Figure 6 , Figure 6 A schematic diagram of training a local image generation model provided in this application. Figure 6 In the example shown, the computing device generates a partial image based on the partial image generation model, and fuses the partial image with a digital human template based on the partial image fusion model to obtain a digital human.
[0128] exist Figure 6 In the example shown, the computing device further calculates a loss function for the digital human obtained by the local image fusion model and the target digital human, and updates the model parameters of the local image generation model based on the calculation results of the loss function, and updates the model parameters of the local image fusion model based on the calculation results of the loss function.
[0129] In a possible implementation, after the computing device obtains the digital human corresponding to the partial image through fusion, the digital human corresponding to the partial image is displayed on the display interface based on the result display module 106. The computing device can also compose an action sequence for the digital human based on the action template of the digital human, generate a video corresponding to the digital human, and play the video of the digital human through the display interface. The action template includes one or more actions of the digital human, such as standing, walking, waving, and expression.
[0130] Please continue reading Figure 3 ,exist Figure 3 In the example shown, the computing device fuses the facial image of the male youth to be fused with the template of the digital human wearing a suit to obtain a young male digital human wearing a suit. The computing device then choreographs an action sequence for the digital human based on the action template of the digital human, and generates a video corresponding to the young male digital human wearing a suit based on the choreographed action sequence.
[0131] It can be seen from the above embodiments that the computing device in the embodiments of the present application can generate a partial image of a digital human based on partial image generation conditions and restriction conditions, thereby improving the quality of the partial image. At the same time, the partial image generated by the computing device meets the requirements of the partial image fusion model for the partial image, thereby improving the fusion matching of the partial image and the digital human template.
[0132] Based on the above method embodiment, the embodiment of the present application also provides a digital human generation device. The digital human generation device provided by the embodiment of the present application is described in detail below.
[0133] See also Figure 7 , Figure 7 A schematic diagram of the structure of a digital human generation device provided in an embodiment of the present application. Figure 7 In the example shown, the digital human generation device 700 is used to implement the various steps performed by the computing device in the above embodiments. The digital human generation device 700 includes a transceiver unit 701 and a processing unit 702 .
[0134] The transceiver unit 701 is used to obtain the first partial image generation condition, the first partial image generation condition is used to indicate the characteristics of the partial image of the first digital human to be generated, the type of the partial image of the first digital human includes one or more of face, clothes, and hairstyle, and the type of the first partial image generation condition includes one or more of text, picture, and voice. The processing unit 702 is used to input the first partial image generation condition into the partial image generation model, and perform forward calculation through the partial image generation model to generate at least one first partial image of the digital human. The processing unit 702 is also used to input at least one first partial image and the first digital human template into the fusion model to generate at least one image of the digital human, and the at least one image of the digital human is the result of the fusion of at least one first partial image and the first digital human template. The processing unit 702 is also used to perform back propagation on the partial image generation model based on the at least one first digital human image and the at least one first target digital human image to update the parameters of the partial image generation model, and the at least one first target digital human image is a digital human image corresponding to the first partial image generation condition.
[0135] In a possible implementation, the processing unit 702 is further used to input the second partial image generation condition into the trained partial image generation model, the second partial image generation condition is used to indicate the characteristics of the partial image of the second digital human to be generated, the type of the partial image of the second digital human includes one or more of face, clothing, and hairstyle, and the type of the second partial image generation condition includes one or more of text, picture, and voice, so as to generate at least one second partial image. The at least one second partial image and the second digital human template are input into the fusion model to generate at least one second digital human image, and the at least one second digital human image is the result of the fusion of the at least one second partial image and the second digital human template.
[0136] In a possible implementation, the processing unit 702 is specifically used to input the second local image generation conditions and restriction conditions into the trained local image generation model to generate at least one second local image. The restriction conditions are used to indicate the requirements for the local image when the fusion model fuses the digital human template and the local image. The types of restriction conditions include text.
[0137] In a possible implementation manner, the transceiver unit 701 is further configured to send a prompt message to the user device when the local image generated by the local image generation model does not meet the restriction condition, and the prompt message is used to prompt the user device to modify the local image generation condition.
[0138] In a possible implementation manner, the processing unit 702 is further configured to perform one or more of the following operations on the second partial image based on the restriction condition: adding, modifying and deleting features of the second partial image.
[0139] In a possible implementation, the processing unit 702 is further used to filter at least one second partial image based on a filtering index, where the filtering index includes a quality filtering index and an algorithm restriction filtering index corresponding to the fusion model, and the algorithm restriction filtering index is used to indicate the requirements of the fusion model for the second partial image.
[0140] In a possible implementation, the processing unit 702 is further configured to generate a quality score of at least one second partial image based on a quality assessment algorithm, the quality score comprising one or more of the following: a visible light index score, a posture index score, a posture skin texture resolution, and a facial feature clarity. The processing unit 702 is specifically configured to filter out a plurality of partial images from the at least one second partial image based on a screening index, the screening index comprising that the quality score of the partial image is greater than or equal to a threshold.
[0141] In a possible implementation, the processing unit 702 is further configured to select a digital human template from a digital human template library, and the digital human template is used to fuse the second partial image.
[0142] In a possible implementation, the processing unit 702 is further configured to arrange an action sequence for the digital human based on a digital human action template to generate a video corresponding to the digital human, wherein the digital human action template includes different digital human actions.
[0143] It is understandable that the transceiver unit 701 and the processing unit 702 in the digital human generation device 700 can be used as functional modules and Figure 1 There is a mapping relationship between the modules in the digital human generation system 10, so as to realize the functions of the modules in the digital human generation system 10. For example, the transceiver unit 701 may correspond to the input module 101, and the processing unit 702 may correspond to the generation module 102, the screening module 103, the digital human template library 104, the fusion module 105, the restriction module 106 and the result display module 107.
[0144] It should be understood that the division of the units in the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And the units in the device can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some units can also be implemented in the form of software calling through processing elements, and some units can be implemented in the form of hardware. For example, each unit can be a separately established processing element, or it can be integrated in a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, and called and executed by a certain processing element of the device. The function of the unit. In addition, all or part of these units can be integrated together, or they can be implemented independently. The processing element described here can also be a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each unit above can be implemented by an integrated logic circuit of hardware in the processor element or in the form of software calling through a processing element.
[0145] It is worth noting that, for the above method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the described order of actions. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required for the present application.
[0146] Other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with the fact that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0147] See also Figure 8 , Figure 8A schematic diagram of the structure of a computing device provided in an embodiment of the present application. Figure 8 As shown, the computing device 800 includes: a processor 801, a memory 802, a communication interface 803 and a bus 804. The processor 801, the memory 802 and the communication interface 803 are coupled via a bus (not marked in the figure). The memory 802 stores instructions. When the execution instructions in the memory 802 are executed, the computing device 800 executes the method executed by the computing device in the above method embodiment.
[0148] The computing device 800 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), or a combination of at least two of these integrated circuit forms. For another example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For another example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0149] The processor 801 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.
[0150] The memory 802 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. Among them, the nonvolatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0151] The memory 802 stores executable program codes, and the processor 801 executes the executable program codes to respectively implement the functions of the aforementioned units or modules, thereby implementing the aforementioned digital human generation method. That is, the memory 802 stores instructions for executing the aforementioned digital human generation method.
[0152] The communication interface 803 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or a communication network.
[0153] In addition to the data bus, the bus 804 may also include a power bus, a control bus, a status signal bus, etc. The bus may be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0154] See also Fig. 9 , Fig. 9 A schematic diagram of a computing device cluster provided in an embodiment of the present application. Fig. 9 As shown, the computing device cluster 900 includes at least one computing device 800 .
[0155] like Fig. 9 As shown, the computing device cluster 900 includes at least one computing device 800. The memory 802 in one or more computing devices 800 in the computing device cluster 900 may store the same instructions for executing the above-mentioned digital human generation method.
[0156] In some possible implementations, the memory 802 of one or more computing devices 800 in the computing device cluster 900 may also respectively store some instructions for executing the above-mentioned method for generating a digital human. In other words, the combination of one or more computing devices 800 may jointly execute the instructions for executing the above-mentioned method for generating a digital human.
[0157] It should be noted that the memory 802 in different computing devices 800 in the computing device cluster 900 can store different instructions, which are respectively used to execute part of the functions of the above-mentioned task processing device. That is, the instructions stored in the memory 802 in different computing devices 800 can realize the functions of one or more modules in the transceiver unit and the processing unit.
[0158] In some possible implementations, one or more computing devices 800 in the computing device cluster 900 may be connected via a network, which may be a wide area network or a local area network.
[0159] See also Fig.10 , Fig.10A schematic diagram of computer devices in a computer cluster connected via a network provided in an embodiment of the present application. Fig.10 As shown, two computing devices 800A and 800B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device.
[0160] In a possible implementation, the memory in the computing device 800A stores instructions for executing the functions of the transceiver unit, and the memory in the computing device 800B stores instructions for executing the functions of the processing unit.
[0161] It should be understood that Fig.10 The functions of the computing device 800A shown in FIG. 8 may also be completed by multiple computing devices. Similarly, the functions of the computing device 800B may also be completed by multiple computing devices.
[0162] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the above method embodiment.
[0163] In another embodiment of the present application, a computer program product is provided, the computer program product includes computer execution instructions, the computer execution instructions are stored in a computer readable storage medium. When the processor of the device executes the computer execution instructions, the device executes the method executed by the computing device in the above method embodiment.
[0164] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0165] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0166] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0168] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A method for generating a digital human, characterized in that: The method comprises: Acquire a first partial image generation condition, the first partial image generation condition is used to indicate the characteristics of the partial image of the first digital person to be generated, the type of the partial image of the first digital person includes one or more of face, clothes, and hairstyle, and the type of the first partial image generation condition includes one or more of text, picture, and voice; Inputting the first partial image generation condition into a partial image generation model, and performing forward calculation through the partial image generation model to generate at least one first partial image of the digital human; Inputting the at least one first partial image and the first digital human template into a fusion model to generate at least one image of the digital human, wherein the at least one image of the digital human is a result of fusion of the at least one first partial image and the first digital human template; Based on the at least one first digital human image and at least one first target digital human image, the local image generation model is back-propagated to update the parameters of the local image generation model, and the at least one first target digital human image is a digital human image corresponding to the first local image generation condition.
2. The method according to claim 1, characterized in that The method further comprises: Inputting a second partial image generation condition into the trained partial image generation model, the second partial image generation condition is used to indicate the features of the partial image of the second digital person to be generated, the type of the partial image of the second digital person includes one or more of face, clothes, and hairstyle, and the type of the second partial image generation condition includes one or more of text, picture, and voice, so as to generate at least one second partial image; The at least one second partial image and the second digital human template are input into the fusion model to generate at least one second digital human image, where the at least one second digital human image is the fusion result of the at least one second partial image and the second digital human template.
3. The method according to claim 2, characterized in that The step of inputting the second partial image generation condition into the trained partial image generation model to generate at least one second partial image comprises: The second partial image generation condition and restriction condition are input into the trained partial image generation model to generate at least one second partial image. The restriction condition is used to indicate the requirements for the partial image when the fusion model fuses the digital human template and the partial image. The type of the restriction condition includes text.
4. The method according to claim 3, characterized in that The method further comprises: When the partial image generated by the partial image generation model does not satisfy the restriction condition, a prompt message is sent to the user equipment, wherein the prompt message is used to prompt the user to modify the second partial image generation condition.
5. The method according to claim 3 or 4, characterized in that: The method further comprises: Based on the restriction condition, one or more of the following operations are performed on the second partial image: adding, modifying and deleting features of the second partial image.
6. The method according to any one of claims 2 to 5, characterized in that The method further comprises: The at least one second partial image is screened based on a screening index, wherein the screening index includes a quality screening index and an algorithm restriction screening index corresponding to the fusion model, and the algorithm restriction screening index is used to indicate the requirements of the fusion model for the second partial image.
7. The method according to claim 6, characterized in that The method further comprises: Generating a quality score of the at least one second partial image based on a quality assessment algorithm, the quality score comprising one or more of the following: a visible light index score, a posture index score, a posture skin texture resolution, and a facial feature clarity; The filtering of the at least one second partial image based on the filtering index includes: At least one candidate partial image is screened out from the at least one second partial image based on a screening index, wherein the screening index includes that a quality score of the partial image is greater than or equal to a threshold.
8. The method according to any one of claims 2 to 6, characterized in that The method further comprises: Receive a digital human template selection instruction from a user; Based on the digital human template selection instruction, the second digital human template is selected from a digital human template library.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: A video corresponding to the digital human is generated based on a digital human choreographed action sequence corresponding to a digital human action template, wherein the digital human action template includes different digital human actions.
10. A digital human generation device, characterized in that: The device comprises: a transceiver unit, configured to obtain a first partial image generation condition, wherein the first partial image generation condition is used to indicate features of a partial image of a first digital person to be generated, wherein the type of the partial image of the first digital person includes one or more of a face, clothes, and a hairstyle, and the type of the first partial image generation condition includes one or more of text, a picture, and a voice; A processing unit, configured to input the first partial image generation condition into a partial image generation model, and perform forward calculation through the partial image generation model to generate at least one first partial image of the digital human; The processing unit is further used to input the at least one first partial image and the first digital human template into the fusion model to generate at least one image of the digital human, wherein the at least one image of the digital human is the result of the fusion of the at least one first partial image and the first digital human template; The processing unit is also used to perform back propagation on the local image generation model based on the at least one first digital human image and at least one first target digital human image to update the parameters of the local image generation model, and the at least one first target digital human image is a digital human image corresponding to the first local image generation condition.
11. The device according to claim 10, characterized in that The processing unit is also used for: Inputting a second partial image generation condition into the trained partial image generation model, the second partial image generation condition is used to indicate the features of the partial image of the second digital person to be generated, the type of the partial image of the second digital person includes one or more of face, clothes, and hairstyle, and the type of the second partial image generation condition includes one or more of text, picture, and voice, so as to generate at least one second partial image; The at least one second partial image and the second digital human template are input into the fusion model to generate at least one second digital human image, where the at least one second digital human image is the fusion result of the at least one second partial image and the second digital human template.
12. The device according to claim 11, characterized in that The processing unit is specifically used for: The second partial image generation condition and restriction condition are input into the trained partial image generation model to generate at least one second partial image. The restriction condition is used to indicate the requirements for the partial image when the fusion model fuses the digital human template and the partial image. The type of the restriction condition includes text.
13. The device according to claim 12, characterized in that The transceiver unit is also used for: When the partial image generated by the partial image generation model does not satisfy the restriction condition, a prompt message is sent to the user equipment, wherein the prompt message is used to prompt the user to modify the second partial image generation condition.
14. The device according to claim 12 or 13, characterized in that The processing unit is also used for: Based on the restriction condition, one or more of the following operations are performed on the second partial image: adding, modifying and deleting features of the second partial image.
15. The device according to any one of claims 11 to 14, characterized in that The processing unit is also used for: The at least one second partial image is screened based on a screening index, wherein the screening index includes a quality screening index and an algorithm restriction screening index corresponding to the fusion model, and the algorithm restriction screening index is used to indicate the requirements of the fusion model for the second partial image.
16. The device according to claim 15, characterized in that The processing unit is also used for: Generating a quality score of at least one second partial image of the at least one second partial image based on a quality assessment algorithm, the quality score comprising one or more of the following: a visible light index score, a posture index score, a posture skin texture resolution, and a facial feature clarity; A plurality of partial images are screened out from the at least one second partial image based on a screening index, wherein the screening index includes a quality score of the partial image being greater than or equal to a threshold.
17. The device according to any one of claims 11 to 16, characterized in that The processing unit is also used for: Receive a digital human template selection instruction from a user; Based on the digital human template selection instruction, the second digital human template is selected from a digital human template library.
18. The device according to any one of claims 10 to 17, characterized in that The processing unit is also used for: A video corresponding to the digital human is generated based on a digital human choreographed action sequence corresponding to a digital human action template, wherein the digital human action template includes different digital human actions.
19. A computing device, characterized in that The electronic device comprises a processor coupled to a memory, wherein the processor is used to store instructions. When the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 9.
20. A computing device cluster, characterized in that: The method comprises at least one computing device, wherein the computing device comprises a processor, wherein the processor is coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method according to any one of claims 1 to 9.
21. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed, the computer is caused to perform the method according to any one of claims 1 to 9.
22. A computer program product, comprising instructions, characterized in that: When the instructions are executed, the computer implements the method according to any one of claims 1 to 9.