Digital human model processing method, virtual digital human mouth shape driving method and computer equipment
By updating the model parameters associated with regions of good image quality in the digital human model, the problem of poor driving effect caused by poor image quality is solved, and high-quality virtual face generation and lip-syncing are achieved under low-quality image conditions.
Patent Information
- Application Number
- CN202510892069.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-21
AI Technical Summary
In the prior art, the poor quality of image data results in a poor driving effect of the digital human, which cannot meet user requirements.
By acquiring sample videos of the target image, extracting sample audio and real face sample images, and using a pre-trained general digital human model to output virtual face sample images, the model parameters are updated based on the differences between the virtual face and the real face to generate the target digital human model. The key is to update the model parameters associated with regions with good image quality, rather than updating regions with poor image quality, thereby achieving image restoration and improving the driving effect.
When the image quality does not meet the preset conditions, higher quality virtual face images are generated, which improves the driving effect and display quality of the digital human.
Smart Images

Figure CN120823296A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for processing a digital human model, a method and apparatus for driving the lip movements of a virtual digital human, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Artificial Intelligence (AI) is a comprehensive field of computer science. By studying the design principles and implementation methods of various intelligent machines, AI aims to empower them with the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, including natural language processing and machine learning and deep learning. With technological advancements, AI will be applied in even more areas and play an increasingly important role.
[0003] Voice-driven mouth and face generation technology allows developers to quickly build digital human-based applications, such as virtual hosts, virtual customer service representatives, and virtual teachers. However, when faced with poor image data, data limitations may not meet the requirements for digital human training or inference, resulting in the digital human driving effect failing to meet user expectations.
[0004] Therefore, there is a problem in the related technology that the driving effect of the digital human is poor. Summary of the Invention
[0005] Based on this, it is necessary to provide a digital human model processing method, a virtual digital human lip-shaping driving method, device, computer equipment, computer-readable storage medium and computer program product that can improve the driving effect of the digital human in response to the above technical problems.
[0006] In a first aspect, the present application provides a method for processing a digital human model, comprising:
[0007] Obtaining a sample video of a target image, and extracting sample audio and a real face sample image from the sample video; the real face sample image includes a first area whose image quality does not meet a preset condition;
[0008] Inputting the real human face sample image and the sample audio into a pre-trained universal digital human model, and outputting a virtual human face sample image; the similarity between the preset image corresponding to the universal digital human model and the target image meets a preset similarity condition;
[0009] According to the difference between the virtual face sample image and the real face sample image, some model parameters in the universal digital human model are updated to obtain a target digital human model;
[0010] Among them, the partial model parameters include model parameters associated with the second area in the general digital human model, and the second area is the area in the real face sample image other than the first area; the target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
[0011] In one embodiment, updating some model parameters of the universal digital human model according to the difference between the virtual human face sample image and the real human face sample image to obtain the target digital human model includes:
[0012] Determining a target backpropagation link in the universal digital human model based on regional position information of the first region in the real face sample image; the target backpropagation link is at least a portion of all backpropagation links in the universal digital human model that needs to participate in gradient calculation;
[0013] Performing a backpropagation operation on the model loss value determined based on the difference according to the target backpropagation link to calculate a gradient for the portion of the model parameters;
[0014] According to the gradient, the part of the model parameters is updated until the target digital human model is obtained.
[0015] In one embodiment, determining the target backpropagation link in the universal digital human model based on the regional position information of the first region in the real face sample image includes:
[0016] performing a fixing operation on the model parameters to be fixed in the universal digital human model according to the regional position information of the first region in the real human face sample image;
[0017] The target back-propagation link is determined according to model parameters in the universal digital human model that are not fixed by the fixing operation.
[0018] In one embodiment, the method further comprises:
[0019] Inputting the real face sample image into a pre-trained semantic segmentation model to obtain a semantic segmentation result of the pre-trained semantic segmentation model on the first region;
[0020] A region mask of the first region is generated according to a semantic segmentation result of the first region; the region mask is used to represent region position information of the first region in the real face sample image.
[0021] In one embodiment, the method further comprises:
[0022] Obtaining pre-trained candidate digital human models; each candidate digital human model is trained using a real human face sample image of the corresponding candidate image; the real human face sample image of the candidate image does not include the first area;
[0023] The image similarity between the target image and each of the candidate images is obtained, and at least one candidate digital human model ranked higher in image similarity is selected from the candidate digital human models as the universal digital human model.
[0024] In one embodiment, obtaining the image similarity between the target image and each of the candidate images includes:
[0025] For any candidate image among the candidate images, respectively obtain image attribute information of at least one common image attribute of the candidate image and the target image;
[0026] The image similarity between the target image and any of the candidate images is determined based on the image attribute information of any of the candidate images and the image attribute information of the target image.
[0027] In a second aspect, the present application provides a lip-sync driving method for a virtual digital human, comprising:
[0028] Obtain a video of the target image to be processed, extract audio and multiple frames of real facial images from the video to be processed; input each frame of the real facial image and the audio into a target digital human model, and output multiple frames of virtual facial images of a virtual digital human corresponding to the target image lip-synced according to the audio; the target digital human model is obtained by processing the digital human model according to the above-mentioned processing method.
[0029] In a third aspect, the present application further provides a digital human model processing device, comprising:
[0030] A sample video acquisition module is used to acquire a sample video of a target image and extract sample audio and a real face sample image from the sample video; the real face sample image includes a first area whose image quality does not meet a preset condition;
[0031] A sample image generation module is configured to input the real human face sample image and the sample audio into a pre-trained universal digital human model and output a virtual human face sample image; the similarity between the preset image corresponding to the universal digital human model and the target image satisfies a preset similarity condition;
[0032] An updating module, configured to update some model parameters of the universal digital human model according to the difference between the virtual human face sample image and the real human face sample image, to obtain a target digital human model;
[0033] Among them, the partial model parameters include model parameters associated with the second area in the general digital human model, and the second area is the area in the real face sample image other than the first area; the target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
[0034] In a fourth aspect, the present application further provides a lip-activated device for a virtual digital human, comprising:
[0035] A video acquisition module is used to acquire a to-be-processed video of a target image and extract audio and multiple frames of real face images from the to-be-processed video;
[0036] The image generation module is used to input each frame of the real human face image and the audio into the target digital human model, and output multiple frames of virtual human face images corresponding to the target image, in which the virtual digital human is lip-driven according to the audio; the target digital human model is obtained by processing the digital human model according to the processing method described above.
[0037] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0038] In a sixth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.
[0039] In a seventh aspect, the present application further provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0040] The above-mentioned processing method, device, computer equipment, computer-readable storage medium and computer program product of the digital human model obtain a sample video of the target image, extract sample audio and real face sample image from the sample video; the real face sample image includes a first area whose image quality does not meet the preset conditions; the real face sample image and sample audio are input into a pre-trained universal digital human model, and a virtual face sample image is output; the similarity between the preset image corresponding to the universal digital human model and the target image meets the preset similarity conditions; based on the difference between the virtual face sample image and the real face sample image, some model parameters in the universal digital human model are updated to obtain a target digital human model; wherein the some model parameters include model parameters associated with the second area in the universal digital human model, and the second area is the area in the real face sample image other than the first area; the target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
[0041] In this way, by extracting sample audio and a real face sample image containing a first area whose image quality does not meet the preset conditions from the sample video of the target image, a pre-trained universal digital human model is used to output a virtual face sample image based on the real face sample image and the sample audio, and the similarity between the preset image corresponding to the universal digital human model and the target image is selected to meet the preset similarity condition, thereby ensuring the similarity between the output virtual face sample image and the target image, and updating some model parameters in the universal digital human model through the difference between the virtual face sample image and the real face sample image to obtain the target digital human model, wherein some model parameters include model parameters associated with the second area in the universal digital human model, and the second area is the area in the real face sample image other than the first area; in this way, only the virtual face sample image other than the real face sample image is updated. In this image, some model parameters associated with the second area other than the first area with poor image quality are updated, and the model parameters associated with the first area are not updated, so as to fix the model parameters associated with the first area and reduce the interference of the first area whose image quality does not meet the preset conditions on the model. In this way, during the fine-tuning training of the universal digital human model, the texture information corresponding to the universal digital human model can be migrated to the first area to perform image repair on the first area. This allows the trained target digital human model to generate a higher quality virtual face image even when the image quality of the input face image does not meet the preset conditions, thereby improving the display effect of the virtual digital human. The virtual digital human in the virtual face image can also be lip-driven according to the audio signal, effectively improving the driving effect of the digital human. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0043] Figure 1 1 is a flow chart of a method for processing a digital human model in one embodiment;
[0044] Figure 2 A schematic flow chart of the steps of updating some model parameters in a general digital human model to obtain a target digital human model in one embodiment;
[0045] Figure 3 1 is a flow chart of a method for driving the lip movements of a virtual digital human in one embodiment;
[0046] Figure 4 is a flow chart of a method for driving the lip movements of a virtual digital human in another embodiment;
[0047] Figure 5 is a structural block diagram of a processing device for a digital human model in one embodiment;
[0048] Figure 6 is a structural block diagram of a lip-activated device for a virtual digital human in one embodiment;
[0049] Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0051] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0052] In one embodiment, Figure 1 As shown, a method for processing a digital human model is provided. This embodiment uses the method applied to a computer device as an example. It is understood that the computer device can be a terminal, a server, or a system including a terminal and a server. In this embodiment, the method includes the following steps:
[0053] Step S110 , obtaining a sample video of the target image, and extracting sample audio and a real face sample image from the sample video.
[0054] The target image may be a real person image for which a corresponding virtual digital human needs to be generated.
[0055] The sample video may refer to a video used as a training sample. In this embodiment, the sample video may include a video of the target image speaking.
[0056] The sample audio refers to the audio extracted from the sample video.
[0057] The real face sample image refers to a video frame image containing the face information of the target image extracted from the sample video.
[0058] The real face sample image includes a first area whose image quality does not meet a preset condition.
[0059] In practical applications, areas in real face sample images where the face is unclear or the face texture information is destroyed can be used as the first area.
[0060] Among them, the first area can be a face area that can be clearly divided in the real face sample image. For example, if the chin area in the real face sample image is unclear, the chin area can be used as the first area; for example, if the texture information of the tooth area in the real face sample image is destroyed, the tooth area can be used as the first area.
[0061] In practical applications, the first area may also be named as a defect area.
[0062] In a specific implementation, a computer device can obtain a sample video of the target image, such as obtaining a video clip of a target image with flaws on its face through the Internet as a sample video, and extract audio from the sample video as sample audio, and extract a video frame image containing facial information of the target image from the sample video as a real face sample image.
[0063] Step S120: input the real face sample image and sample audio into the pre-trained universal digital human model, and output the virtual face sample image.
[0064] The similarity between the preset image corresponding to the universal digital human model and the target image satisfies a preset similarity condition.
[0065] Among them, the general digital human model is obtained by training real face sample images with preset images.
[0066] The image quality of the real face sample image of the preset image meets the preset conditions.
[0067] In practical applications, a face image with a preset image and clear face and continuous texture can be obtained as a real face sample image.
[0068] The virtual face sample image contains the face information of the virtual digital person corresponding to the target image.
[0069] In a specific implementation, a computer device can input real face sample images and sample audio into a pre-trained general digital human model. The pre-trained general digital human model can generate an image containing facial information of a virtual digital human matching the target image based on the sample image features corresponding to the real face sample images and the sample audio features corresponding to the sample audio, as a virtual face sample image corresponding to the target image.
[0070] The sample audio feature may be a Mel-frequency cepstrum.
[0071] Step S130 , updating some model parameters in the universal digital human model according to the difference between the virtual face sample image and the real face sample image, to obtain a target digital human model.
[0072] In a specific implementation, the computer device can update some model parameters in the universal digital human model according to the difference between the virtual face sample image and the real face sample image to obtain the target digital human model. In this way, the pre-trained universal digital human model can be fine-tuned according to the real face sample image and sample audio, instead of randomly initializing the weights to train a digital human model for the target image. This can not only improve the efficiency of the target digital human model, but also make the digital human image generated by the target digital human model more similar to the target image.
[0073] Part of the model parameters includes model parameters associated with the second region in the universal digital human model, and the second region is the region other than the first region in the real face sample image.
[0074] The image quality of the second area meets the preset condition. In practical applications, the second area can be named as a normal area.
[0075] The target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
[0076] In some embodiments, the computer device may obtain regional position information of the first region in the real human face sample image, determine the region of the real human face sample image other than the first region as the second region, and thereby determine model parameters associated with the second region in the universal digital human model as partial model parameters. In this way, based on the difference between the virtual human face sample image and the real human face sample image, some model parameters in the universal digital human model can be updated without updating the model parameters associated with the first region, thereby achieving the purpose of fixing the model parameters associated with the first region.
[0077] In the above-mentioned processing method of the digital human model, a sample video of the target image is obtained, and sample audio and a real face sample image are extracted from the sample video; the real face sample image includes a first area whose image quality does not meet the preset conditions; the real face sample image and the sample audio are input into a pre-trained universal digital human model, and a virtual face sample image is output; the similarity between the preset image corresponding to the universal digital human model and the target image meets the preset similarity conditions; based on the difference between the virtual face sample image and the real face sample image, some model parameters in the universal digital human model are updated to obtain a target digital human model; wherein, some model parameters include model parameters associated with the second area in the universal digital human model, and the second area is the area in the real face sample image other than the first area; the target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
[0078] In this way, by extracting sample audio and a real face sample image containing a first area whose image quality does not meet the preset conditions from the sample video of the target image, a pre-trained universal digital human model is used to output a virtual face sample image based on the real face sample image and the sample audio, and the similarity between the preset image corresponding to the universal digital human model and the target image is selected to meet the preset similarity condition, thereby ensuring the similarity between the output virtual face sample image and the target image, and updating some model parameters in the universal digital human model through the difference between the virtual face sample image and the real face sample image to obtain the target digital human model, wherein some model parameters include model parameters associated with the second area in the universal digital human model, and the second area is the area in the real face sample image other than the first area; in this way, only the virtual face sample image other than the real face sample image is updated. Some model parameters associated with the second area outside the first area with poor image quality in the image are updated, and the model parameters associated with the first area are not updated, so as to fix the model parameters associated with the first area and reduce the interference of the first area whose image quality does not meet the preset conditions on the model. In this way, during the fine-tuning training of the universal digital human model, the texture information corresponding to the universal digital human model can be migrated to the first area to perform image repair on the first area. This allows the trained target digital human model to generate a higher quality virtual face image even when the image quality of the input face image does not meet the preset conditions, thereby improving the display effect of the virtual digital human. The virtual digital human in the virtual face image can also be lip-driven according to the audio signal, effectively improving the driving effect of the digital human.
[0079] In some embodiments, when a computer device updates some model parameters in a universal digital human model according to the difference between a virtual face sample image and a real face sample image to obtain a target digital human model, the computer device may obtain a model loss value of the universal digital human model according to the difference between the virtual face sample image and the real face sample image; and update some model parameters of the universal digital human model according to the target back propagation link based on the model loss value through a back propagation algorithm to obtain a target digital human model.
[0080] The target back-propagation link is at least a part of all back-propagation links in the universal digital human model that needs to participate in gradient calculation.
[0081] Furthermore, if Figure 2 As shown, step S130, based on the difference between the virtual face sample image and the real face sample image, updates some model parameters in the universal digital human model to obtain the target digital human model, including the following steps:
[0082] Step S1302: determining a target back-propagation link in the universal digital human model according to the regional position information of the first region in the real face sample image.
[0083] In a specific implementation, the computer device can determine the target back-propagation link in the universal digital human model based on the regional position information of the first region in the real face sample image.
[0084] The target backpropagation link does not include the backpropagation link associated with the first region in all backpropagation links in the universal digital human model.
[0085] Step S1304 : Perform a backpropagation operation on the model loss value determined based on the difference according to the target backpropagation link to calculate the gradient of some model parameters.
[0086] In a specific implementation, the computer device can perform a backpropagation operation on the model loss value determined based on the difference according to the target backpropagation link to calculate the gradient for some model parameters.
[0087] Step S1306: Update some model parameters according to the gradient until the target digital human model is obtained.
[0088] In a specific implementation, the computer device can update some model parameters according to the gradient of some model parameters until the general digital human model converges to obtain the target digital human model.
[0089] The technical solution of this embodiment determines a target backpropagation link in a universal digital human model based on the regional position information of a first region in a real face sample image. The target backpropagation link is at least a portion of all backpropagation links in the universal digital human model that need to participate in gradient calculation. Backpropagation is performed on the model loss value determined based on the difference according to the target backpropagation link to calculate the gradients for some model parameters. Based on the gradients, some model parameters are updated until a target digital human model is obtained. In this way, the target backpropagation link in the universal digital human model is determined based on the regional position information of the first region in the real face sample image. By determining the target backpropagation link, only some backpropagation links are involved in gradient calculation, thereby reducing the computational complexity. Furthermore, based on the regional position information of the first region whose image quality does not meet preset conditions, the target backpropagation link that needs to participate in gradient calculation can be accurately determined, avoiding the backpropagation links associated with the first region from participating in gradient calculation. This can prevent the model from being interfered with by low-quality image regions, making parameter updates more stable and reliable, and helping the model converge to a more optimal solution.
[0090] In one embodiment, a target back propagation link in a universal digital human model is determined based on regional position information of a first region in a real human face sample image, including: performing a fixing operation on model parameters to be fixed in the universal digital human model based on the regional position information of the first region in the real human face sample image; and determining the target back propagation link based on model parameters in the universal digital human model that are not fixed by the fixing operation.
[0091] The model parameters to be fixed may refer to the model parameters associated with the first region in the universal digital human model. Therefore, the model parameters associated with the first region in the universal digital human model, i.e., the model parameters to be fixed, can be determined based on the regional position information of the first region in the real human face sample image. A fixing operation can then be performed on the model parameters to be fixed to sever the backpropagation link associated with the first region. Consequently, the target backpropagation link can be determined based on the model parameters in the universal digital human model that are not fixed by the fixing operation. Specifically, the backpropagation links corresponding to the model parameters in the universal digital human model that are not fixed by the fixing operation constitute the target backpropagation link.
[0092] The model parameters in the universal digital human model that are not fixed by the fixed operation are part of the model parameters in the above embodiment.
[0093] The technical solution of this embodiment performs a fixing operation on the model parameters to be fixed in the universal digital human model based on the regional position information of the first region in the real human face sample image; and determines the target backpropagation link based on the model parameters in the universal digital human model that are not fixed by the fixing operation. In this way, by fixing the model parameters to be fixed associated with the first region, there is no need to perform gradient calculation and update on these parameters, which can reduce the amount of calculation and speed up the training of the universal digital human model. In addition, since the first region is a region whose image quality does not meet the preset conditions, the model parameters associated with this region are fixed to achieve the purpose of severing the backpropagation link associated with the first region, which can prevent the region whose image quality does not meet the preset conditions from negatively affecting the model parameter update of the universal digital human model and ensure the stability of model training. Furthermore, based on the model parameters in the universal digital human model that are not fixed by the fixing operation, the target backpropagation link in the universal digital human model that needs to participate in the gradient calculation can be accurately determined.
[0094] In one embodiment, the method further includes: inputting a real face sample image into a pre-trained semantic segmentation model to obtain a semantic segmentation result of the pre-trained semantic segmentation model for the first region; and generating a region mask of the first region based on the semantic segmentation result of the first region.
[0095] The region mask is used to represent the region position information of the first region in the real face sample image.
[0096] The pre-trained semantic segmentation model is used to divide the face image into different semantic regions and label and classify each region. For example, the pre-trained semantic segmentation model can divide the input face image into the teeth region, mouth region, eye region, nose region, etc.
[0097] In a specific implementation, when a computer device obtains the regional position information of a first region in a real face sample image, the computer device can input the real face sample image into a pre-trained semantic segmentation model. The pre-trained semantic segmentation model can divide the real face sample image into different semantic regions, and label and classify each region, thereby obtaining the semantic segmentation result of the pre-trained semantic segmentation model for the first region. Thus, a regional mask of the first region can be generated based on the semantic segmentation result of the first region; the regional mask is used to characterize the regional position information of the first region in the real face sample image.
[0098] In practical applications, the pre-trained semantic segmentation model can be used with a face parsing model. By inputting a real face sample image into the face parsing model, a regional mask for the first region can be generated. For example, if the image quality of the teeth region in the real face sample image does not meet preset requirements, a regional mask for the teeth region can be generated based on the semantic segmentation results of the teeth region obtained by the face parsing model.
[0099] The technical solution of this embodiment is to input a real face sample image into a pre-trained semantic segmentation model to obtain the semantic segmentation result of the pre-trained semantic segmentation model for the first region; based on the semantic segmentation result of the first region, a region mask of the first region is generated; the region mask is used to represent the regional position information of the first region in the real face sample image. In this way, the pre-trained semantic segmentation model can accurately divide the real face sample image into different regions to obtain the semantic segmentation result of the first region, and by generating the region mask of the first region, the specific position of the first region in the real face sample image can be clearly determined.
[0100] In one embodiment, the method further includes: obtaining each pre-trained candidate digital human model; each candidate digital human model is obtained by training using a real face sample image of the corresponding candidate image; the real face sample image of the candidate image does not include the first area; obtaining the image similarity between the target image and each candidate image, and selecting at least one candidate digital human model with a higher image similarity ranking from each candidate digital human model as the general digital human model.
[0101] The real face sample image of the candidate image does not include the first region, wherein the first region is a region whose image quality does not meet the preset condition. That is, the image quality of the real face sample image of the candidate image meets the preset condition.
[0102] In practical applications, a face image of the candidate image with a clear face and continuous texture can be obtained as a real face sample image.
[0103] Among them, different candidate images have different corresponding image attribute information under at least one same image attribute.
[0104] Among them, image attributes refer to various attributes or characteristics used to describe and portray image characteristics.
[0105] Image attribute information refers to the specific descriptions and data related to image attributes. It is a detailed description of image attributes and may exist in various forms, such as text descriptions, numerical values, labels, vectors, etc. Image attribute information can provide more precise information to facilitate image recognition, classification, comparison, and processing.
[0106] In a specific implementation, a computer device can obtain various pre-trained candidate digital human models; wherein, each candidate digital human model is obtained by training using a real face sample image of the corresponding candidate image; the real face sample image of the candidate image does not include the first area; by obtaining the image similarity between the target image and each candidate image, at least one candidate digital human model with a higher image similarity ranking is selected from each candidate digital human model as a universal digital human model.
[0107] In some embodiments, the candidate digital human model with the highest image similarity ranking may be selected from the candidate digital human models as the universal digital human model.
[0108] The technical solution of this embodiment obtains pre-trained candidate digital human models; each candidate digital human model is trained using a real human face sample image of the corresponding candidate image; the real human face sample image of the candidate image does not include the first region; obtains the image similarity between the target image and each candidate image, and selects at least one candidate digital human model with the highest image similarity ranking from each candidate digital human model as the universal digital human model. In this way, by calculating the image similarity between the target image and the candidate images corresponding to each pre-trained candidate digital human model, and selecting the candidate digital human model with the highest ranking as the universal digital human model, the digital human model that best matches the target image can be selected from multiple pre-trained models. The digital human image generated by the selected universal digital human model is closer to the target image, and can provide users with a more realistic and more user-friendly digital human image.
[0109] In one embodiment, obtaining the image similarity between the target image and each candidate image includes: for any candidate image among the candidate images, respectively obtaining image attribute information of any candidate image and the target image on at least one common image attribute; and determining the image similarity between the target image and any candidate image based on the image attribute information of any candidate image and the image attribute information of the target image.
[0110] In a specific implementation, when a computer device obtains the image similarity between a target image and each candidate image, for any candidate image among the candidate images, the computer device can respectively obtain the image attribute information of the any candidate image and the target image on at least one identical image attribute, and determine the image similarity between the target image and the any candidate image based on the image attribute information of the any candidate image and the image attribute information of the target image.
[0111] In this way, based on the same method, the computer device can obtain the image similarity between the target image and each candidate image.
[0112] The technical solution of this embodiment obtains image attribute information for at least one common image attribute between each candidate image and the target image; then, based on the image attribute information of each candidate image and the target image, determines the image similarity between the target image and each candidate image. By selecting at least one common image attribute and obtaining and comparing the corresponding image attribute information for the candidate and target images under the same image attribute, the image similarity calculation is more accurate.
[0113] In some embodiments, the pre-trained candidate digital human model can be improved and trained using a U-Net model. Specifically, the U-Net model's input can be modified to multiple frames of real facial images and audio features, and an attention structure can be added. The U-Net output can be modified to multiple frames of virtual facial images, enabling it to generate corresponding virtual faces based on audio features.
[0114] Furthermore, during the process of training the U-Net model to obtain pre-trained candidate digital human models, the computer device may obtain sample videos of the candidate images. The sample videos of the candidate images may include videos of the candidate images speaking. For example, sample videos of candidate images with clear faces and continuous textures may be obtained via the internet, or captured by an image acquisition device.
[0115] In this way, multiple frames of real face sample images of the candidate image and sample audio of the candidate image can be extracted from the sample video of the candidate image, and the sample audio features corresponding to the sample audio of the candidate image can be extracted. The real face sample images of each frame of the candidate image and the sample audio features corresponding to the sample audio of the candidate image are input into the U-Net model to be trained. The U-Net model to be trained can generate virtual face sample images of each frame of the candidate image according to the sample audio features corresponding to the sample audio of the candidate image and the sample image features corresponding to the real face sample images of each frame of the candidate image. According to the difference between the virtual face sample images of each frame of the candidate image and their corresponding real face sample images, the model loss value of the U-Net model to be trained is calculated. The gradient value of the U-Net model to be trained is obtained from the model loss value for neural network back propagation, and the model parameters of the U-Net model to be trained are updated. The above operation is continuously trained until the trained U-Net model is fitted to obtain the pre-trained candidate digital human model corresponding to the candidate image.
[0116] Among them, the U-Net model to be trained, in the process of generating virtual face sample images of each frame of the candidate image based on the sample audio features corresponding to the sample audio of the candidate image and the sample image features corresponding to the real face sample images of each frame of the candidate image, can adjust the weights of the sample image features corresponding to the real face sample images of the candidate image and the sample audio features corresponding to the sample audio of the candidate image according to the attention mechanism, and generate the virtual face sample images of the candidate image according to the weight adjustment results.
[0117] In yet another embodiment, Figure 3 As shown, a flow chart of a method for driving the lip movements of a virtual digital human is provided. The method is illustrated by applying the method to a computer device. The method includes the following steps:
[0118] Step S310: Obtain a video to be processed of the target image, and extract audio and multiple frames of real face images from the video to be processed.
[0119] The video to be processed may be a video of the target image speaking.
[0120] The real face image may include a first region whose image quality does not meet a preset condition.
[0121] In a specific implementation, a computer device can obtain a video to be processed of a target image, and extract audio and multiple frames of real face images from the video to be processed.
[0122] Step S320: input each frame of real human face image and audio into the target digital human model, and output multiple frames of virtual human face images corresponding to the target image, which are driven by the lip movements of the virtual digital human according to the audio.
[0123] In a specific implementation, a computer device can input each frame of real facial images and audio into a target digital human model corresponding to the target image, wherein the target digital human model is obtained by processing according to the digital human model processing method as described above. In this way, the target digital human model can generate multiple frames of virtual facial images driven by lip movements according to the audio features corresponding to the audio and the image features corresponding to each frame of real facial images.
[0124] In the above-mentioned method for driving the lip movement of a virtual digital human, a video to be processed of a target image is obtained, and audio and multiple frames of real facial images are extracted from the video to be processed; each frame of real facial image and audio is input into a target digital human model, and multiple frames of virtual facial images of the virtual digital human corresponding to the target image driven by lip movement according to the audio are output; wherein the target digital human model is obtained by processing according to the above-mentioned method for processing the digital human model.
[0125] In this way, by extracting sample audio and a real face sample image containing a first area whose image quality does not meet the preset conditions from the sample video of the target image, a pre-trained universal digital human model is used to output a virtual face sample image based on the real face sample image and the sample audio, and the similarity between the preset image corresponding to the universal digital human model and the target image is selected to meet the preset similarity condition, thereby ensuring the similarity between the output virtual face sample image and the target image, and updating some model parameters in the universal digital human model through the difference between the virtual face sample image and the real face sample image to obtain the target digital human model, wherein some model parameters include model parameters associated with the second area in the universal digital human model, and the second area is the area in the real face sample image other than the first area; in this way, only the virtual face sample image other than the real face sample image is updated. In this image, some model parameters associated with the second area other than the first area with poor image quality are updated, and the model parameters associated with the first area are not updated, so as to fix the model parameters associated with the first area and reduce the interference of the first area whose image quality does not meet the preset conditions on the model. In this way, during the fine-tuning training of the universal digital human model, the texture information corresponding to the universal digital human model can be migrated to the first area to perform image repair on the first area. This allows the trained target digital human model to generate a higher quality virtual face image even when the image quality of the input face image does not meet the preset conditions, thereby improving the display effect of the virtual digital human. The virtual digital human in the virtual face image can also be lip-driven according to the audio signal, effectively improving the driving effect of the digital human.
[0126] In another embodiment, Figure 4 As shown, a flowchart of a method for driving the lip movements of a virtual digital human is provided, comprising the following steps:
[0127] Step S402: Obtain a sample video of the target image, and extract sample audio and a real face sample image from the sample video.
[0128] Step S404: Obtain each pre-trained candidate digital human model, and for any candidate image among the candidate images, obtain image attribute information of at least one common image attribute between any candidate image and the target image.
[0129] Step S406 : determining the image similarity between the target image and any candidate image based on the image attribute information of any candidate image and the image attribute information of the target image.
[0130] Step S408 : selecting at least one candidate digital human model ranked higher in image similarity from among the candidate digital human models as a universal digital human model.
[0131] Step S410: inputting real face sample images and sample audio into a pre-trained universal digital human model, and outputting virtual face sample images.
[0132] Step S412: performing a fixing operation on the model parameters to be fixed in the universal digital human model according to the region position information of the first region in the real face sample image.
[0133] Step S414: determining a target back-propagation link according to the model parameters in the universal digital human model that are not fixed by the fixed operation.
[0134] Step S416: Perform a backpropagation operation on the model loss value determined based on the difference according to the target backpropagation link to calculate the gradient of some model parameters.
[0135] Step S418: Update some model parameters according to the gradient until the target digital human model is obtained.
[0136] Step S420: Obtain a video to be processed of the target image, and extract audio and multiple frames of real face images from the video to be processed.
[0137] Step S422: input each frame of real human face image and audio into the target digital human model, and output multiple frames of virtual human face images corresponding to the target image, which are driven by the lip movements of the virtual digital human according to the audio.
[0138] It should be noted that the specific definitions of the above steps can refer to the specific definitions of a digital human model processing method and a virtual digital human lip-sync driving method described above.
[0139] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0140] Based on the same inventive concept, embodiments of the present application also provide a digital human model processing device for implementing the aforementioned digital human model processing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of the one or more digital human model processing device embodiments provided below can be found in the aforementioned limitations of the digital human model processing method and will not be further elaborated here.
[0141] In an exemplary embodiment, Figure 5 As shown, a digital human model processing device is provided, comprising: a sample video acquisition module 510, a sample image generation module 520 and an update module 530, wherein:
[0142] The sample video acquisition module 510 is used to acquire a sample video of the target image, and extract sample audio and a real face sample image from the sample video; the real face sample image includes a first area whose image quality does not meet a preset condition.
[0143] The sample image generation module 520 is used to input the real face sample image and the sample audio into a pre-trained universal digital human model and output a virtual face sample image; the similarity between the preset image corresponding to the universal digital human model and the target image meets the preset similarity condition.
[0144] An updating module 530 is configured to update some model parameters of the general digital human model according to the difference between the virtual human face sample image and the real human face sample image to obtain a target digital human model;
[0145] Among them, the partial model parameters include model parameters associated with the second area in the general digital human model, and the second area is the area in the real face sample image other than the first area; the target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
[0146] In one embodiment, the update module 530 is specifically used to determine the target back propagation link in the universal digital human model based on the regional position information of the first region in the real face sample image; the target back propagation link is at least a part of all the back propagation links in the universal digital human model that needs to participate in the gradient calculation; according to the target back propagation link, the model loss value determined based on the difference is back propagated to calculate the gradient for the partial model parameters; according to the gradient, the partial model parameters are updated until the target digital human model is obtained.
[0147] In one embodiment, the update module 530 is specifically used to perform a fixing operation on the model parameters to be fixed in the universal digital human model based on the regional position information of the first region in the real face sample image; and determine the target backpropagation link based on the model parameters in the universal digital human model that are not fixed by the fixing operation.
[0148] In one embodiment, the device also includes: a position determination module, used to input the real face sample image into a pre-trained semantic segmentation model to obtain the semantic segmentation result of the pre-trained semantic segmentation model for the first region; based on the semantic segmentation result of the first region, generate a region mask of the first region; the region mask is used to represent the region position information of the first region in the real face sample image.
[0149] In one embodiment, the device also includes: a screening module for obtaining each pre-trained candidate digital human model; each candidate digital human model is trained using a real face sample image of the corresponding candidate image; the real face sample image of the candidate image does not include the first area; obtaining the image similarity between the target image and each candidate image, and selecting at least one candidate digital human model with a higher image similarity ranking from each candidate digital human model as the universal digital human model.
[0150] In one embodiment, the screening module is specifically used to obtain, for any candidate image among the candidate images, image attribute information of the candidate image and the target image on at least one common image attribute; and determine the image similarity between the target image and the candidate image based on the image attribute information of the candidate image and the image attribute information of the target image.
[0151] Based on the same inventive concept, embodiments of the present application also provide a virtual digital human lip-actuating device for implementing the aforementioned virtual digital human lip-actuating method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of the one or more virtual digital human lip-actuating device embodiments provided below can be found in the aforementioned limitations of the virtual digital human lip-actuating method and will not be further elaborated here.
[0152] In an exemplary embodiment, Figure 6 As shown, a lip-activated device for a virtual digital human is provided, comprising: a video acquisition module 610 and an image generation module 620, wherein:
[0153] The video acquisition module 610 is used to acquire a to-be-processed video of a target image, and extract audio and multiple frames of real face images from the to-be-processed video.
[0154] The image generation module 620 is used to input each frame of the real human face image and the audio into the target digital human model, and output multiple frames of virtual human face images corresponding to the target image, driven by the lip movements of the virtual digital human according to the audio; the target digital human model is obtained according to the digital human model processing method described above.
[0155] The various modules in the aforementioned digital human model processing device and virtual digital human lip-activated device can be implemented in whole or in part through software, hardware, or a combination thereof. These modules can be embedded in or independent of a processor within a computer device in the form of hardware, or stored in a computer device's memory in the form of software, allowing the processor to call and execute the corresponding operations of each module.
[0156] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication, and the wireless communication can be achieved via Wi-Fi, a mobile cellular network, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a lip-sync driving method for a virtual digital human and / or a method for processing a digital human model. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0157] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0158] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0159] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0160] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0162] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0163] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0164] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for processing a digital human model, characterized in that: The method comprises: Obtaining a sample video of a target image, and extracting sample audio and a real face sample image from the sample video; the real face sample image includes a first area whose image quality does not meet a preset condition; Inputting the real human face sample image and the sample audio into a pre-trained universal digital human model, and outputting a virtual human face sample image; the similarity between the preset image corresponding to the universal digital human model and the target image meets a preset similarity condition; According to the difference between the virtual face sample image and the real face sample image, some model parameters in the universal digital human model are updated to obtain a target digital human model; Among them, the partial model parameters include model parameters associated with the second area in the general digital human model, and the second area is the area in the real face sample image other than the first area; the target digital human model is used to lip-drive the virtual digital human corresponding to the target image.
2. The method according to claim 1, characterized in that The updating of some model parameters in the universal digital human model according to the difference between the virtual human face sample image and the real human face sample image to obtain the target digital human model includes: Determining a target backpropagation link in the universal digital human model based on regional position information of the first region in the real face sample image; the target backpropagation link is at least a portion of all backpropagation links in the universal digital human model that needs to participate in gradient calculation; Performing a backpropagation operation on the model loss value determined based on the difference according to the target backpropagation link to calculate a gradient for the portion of the model parameters; According to the gradient, the part of the model parameters is updated until the target digital human model is obtained.
3. The method according to claim 2, characterized in that The step of determining a target backpropagation link in the universal digital human model based on the regional position information of the first region in the real face sample image includes: performing a fixing operation on the model parameters to be fixed in the universal digital human model according to the regional position information of the first region in the real human face sample image; The target back-propagation link is determined according to model parameters in the universal digital human model that are not fixed by the fixing operation.
4. The method according to claim 2, characterized in that The method further comprises: Inputting the real face sample image into a pre-trained semantic segmentation model to obtain a semantic segmentation result of the pre-trained semantic segmentation model on the first region; A region mask of the first region is generated according to a semantic segmentation result of the first region; the region mask is used to represent region position information of the first region in the real face sample image.
5. The method according to claim 1, wherein The method further comprises: Obtaining pre-trained candidate digital human models; each candidate digital human model is trained using a real human face sample image of the corresponding candidate image; the real human face sample image of the candidate image does not include the first area; The image similarity between the target image and each of the candidate images is obtained, and at least one candidate digital human model ranked higher in image similarity is selected from the candidate digital human models as the universal digital human model.
6. The method according to claim 5, characterized in that The obtaining of the image similarity between the target image and each of the candidate images includes: For any candidate image among the candidate images, respectively obtain image attribute information of at least one common image attribute of the candidate image and the target image; The image similarity between the target image and any of the candidate images is determined based on the image attribute information of any of the candidate images and the image attribute information of the target image.
7. A lip-sync driving method for a virtual digital human, characterized in that: The method comprises: Obtaining a to-be-processed video of a target image, and extracting audio and multiple frames of real face images from the to-be-processed video; Input each frame of the real human face image and the audio into a target digital human model, and output multiple frames of virtual human face images of a virtual digital human corresponding to the target image driven by lip movements according to the audio; the target digital human model is obtained by processing according to the digital human model processing method described in any one of claims 1 to 6.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.