Face texture image generation method, 3D digital human generation method and 3D digital human video generation method

Through the target diffusion model, high-quality face texture images are generated using the face images collected by ordinary shooting equipment, solving the problems of high cost and complex operation in the prior art, and achieving low-cost and high-quality 3D digital human generation.

CN120355832APending Publication Date: 2025-07-22MOFA (SHANGHAI) INFORMATION TECH CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510194844.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, generating high-quality face texture images requires expensive professional equipment and cumbersome operations. Ordinary shooting equipment cannot obtain high-quality face texture images, resulting in poor quality 3D digital people.

Method used

The target diffusion model is used, and the face image map image is used as a constraint condition to denoise the noise image to generate high-quality face texture images, which can be achieved through the face image collected by ordinary shooting equipment.

Benefits of technology

It reduces the cost and equipment operation threshold for generating high-quality face texture images, and improves the quality and efficiency of 3D digital human generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355832A_ABST
    Figure CN120355832A_ABST
Patent Text Reader

Abstract

The invention provides a human face texture image generation method, a 3D digital human generation method and a 3D digital human video generation method, and relates to the technical field of digital humans, the human face texture image generation method can obtain a high-quality human face texture image to meet user requirements by taking a human face image mapping image as a constraint condition of a target diffusion model, and the user experience is improved. The method can be applied to the field of 3D digital human generation. According to the method, the problems of high implementation cost and high professional equipment operation threshold for generating a high-quality human face texture image in the prior art can be avoided, and the human face texture image is obtained only by adopting a human face image acquired by common shooting equipment such as a mobile phone or a digital camera; therefore, the high-quality face texture image can be obtained through conversion of the target diffusion model, the implementation cost is low, and the equipment operation threshold is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital humans, and in particular, to a method for generating a face texture image, a method for generating a 3D digital human, and a method for generating a 3D digital human video. Background Art

[0002] With the increasing development of virtual reality (VR) and augmented reality (AR) technologies, digital human technology has also become a research hotspot in the field of computer vision. In digital human technology, generating high-quality face textures is of utmost importance.

[0003] In the prior art, traditional face texture image generation techniques are usually used to generate face texture images, that is, given a face image as input, a face texture image is generated to support the generation and rendering of 3D faces. The traditional face texture image generation technique mainly uses a high-precision 3D face scanning device to perform millimeter-level facial scans on models, and based on the scanned face model and face texture map, the face of a hyper-realistic 3D digital human can be generated. Then, by splicing and fusing with the template torso of the 3D digital human, a hyper-realistic 3D digital human identical to the appearance of the model can be created.

[0004] However, in the above technical solution, to generate a high-quality face texture image, it is necessary to use expensive professional equipment to capture face images, resulting in a high implementation cost, and the operation of professional equipment is relatively cumbersome with too high a threshold. If ordinary shooting equipment such as mobile phones or digital cameras is used to collect face images, high-quality face texture images cannot be obtained, thereby resulting in poor quality of the generated 3D digital humans. Summary of the Invention

[0005] The present invention provides a method for generating a face texture image, a method for generating a 3D digital human, and a method for generating a 3D digital human video to solve the defects existing in the related technologies.

[0006] The present invention provides a method for generating a face texture image, including: Receiving a face image mapping image of a target object; Based on a target diffusion model, using the face image mapping image as a constraint condition, denoising a noise image to obtain the face texture image of the target object; wherein the target diffusion model is obtained by performing forward diffusion and reverse diffusion training based on face texture image samples using face image mapping image samples as constraint conditions.

[0007] The present invention further provides a method for generating a 3D digital human, including: Receiving at least one face image of a target object; Based on the at least one face image, determine the face image mapping image and facial geometric information of the target object; Based on the face image mapping image, apply the above-mentioned face texture image generation method to generate a face texture image; Based on the facial geometric information and the face texture image, generate a 3D digital face corresponding to the target object, and based on the 3D digital face, generate a target 3D digital human.

[0008] The present invention also provides a method for generating a 3D digital human video, including: Obtain a target 3D digital human, where the target 3D digital human is generated based on the above-mentioned 3D digital human generation method; Based on the target 3D digital human, generate a 3D digital human video.

[0009] The present invention also provides a device for generating a face texture image, including: A first image receiving module, configured to receive a face image mapping image of a target object; An image denoising module, configured to denoise a noise image based on a target diffusion model using the face image mapping image as a constraint condition to obtain the face texture image of the target object; Wherein, the target diffusion model is trained by forward diffusion and reverse diffusion based on face texture image samples using face image mapping image samples as constraint conditions.

[0010] The present invention also provides a 3D digital human generation system, including: A second image receiving module, configured to receive at least one face image of a target object; A texture geometry image determination module, configured to determine the face image mapping image and facial geometric information of the target object based on the at least one face image; A face texture image generation device, configured to generate a face texture image based on the face image mapping image by applying the above-mentioned face texture image generation method; A digital human generation module, configured to generate a 3D digital face corresponding to the target object based on the facial geometric information and the face texture image, and generate a target 3D digital human based on the 3D digital face.

[0011] The present invention also provides a 3D digital human video generation system, including: A 3D digital human generation system, configured to obtain a target 3D digital human, where the target 3D digital human is generated based on the above-mentioned 3D digital human generation method; A video generation module, configured to generate a 3D digital human video based on the target 3D digital human.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the face texture image generation method, or the 3D digital human generation method, or the 3D digital human video generation method as described in any one of the above.

[0013] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the face texture image generation method, or the 3D digital human generation method, or the 3D digital human video generation method as described in any one of the above.

[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the face texture image generation method, or the 3D digital human generation method, or the 3D digital human video generation method as described in any one of the above.

[0015] For the face texture image generation method, 3D digital human generation method, and 3D digital human video generation method provided by the present invention, the face texture image generation method first receives a face image mapping image of a target object; then, based on a target diffusion model, uses the face image mapping image as a constraint condition to denoise a noise image, and obtains a face texture image of the target object. By using the face image mapping image as a constraint condition for the target diffusion model, a high-quality face texture image can be obtained to meet user requirements, and then it can be applied to the field of 3D digital human generation. This method can avoid the problems of high implementation cost and high operation threshold of professional equipment faced in the prior art for generating high-quality face texture images. Only by using a face texture image obtained from a face image collected by an ordinary shooting device such as a mobile phone or a digital camera, a high-quality face texture image can be obtained through conversion by the target diffusion model, with low implementation cost and low equipment operation threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 is one of the flow diagrams of the face texture image generation method provided by the present invention.

[0018] Figure 2 is the flow diagram of the forward diffusion of the initial diffusion model in the face texture image generation method provided by the present invention.

[0019] Figure 3 It is a schematic flow diagram of the reverse diffusion of the initial diffusion model in the face texture image generation method provided by the present invention.

[0020] Figure 4 It is an overall schematic diagram of the training process of a conventional diffusion model.

[0021] Figure 5 It is a schematic flow diagram of training the initial diffusion model in the face texture image generation method provided by the present invention.

[0022] Figure 6 It is the second schematic flow diagram of the face texture image generation method provided by the present invention.

[0023] Figure 7 It is a schematic flow diagram of the first denoising step of the latent features by the UNet neural network in the face texture image generation method provided by the present invention.

[0024] Figure 8 It is the first schematic flow diagram of the 3D digital human generation method provided by the present invention.

[0025] Figure 9 It is the second schematic flow diagram of the 3D digital human generation method provided by the present invention.

[0026] Figure 10 It is a schematic flow diagram of the 3D digital human video generation method provided by the present invention.

[0027] Figure 11 It is a schematic structural diagram of the face texture image generation device provided by the present invention.

[0028] Figure 12 It is a schematic structural diagram of the 3D digital human generation system provided by the present invention.

[0029] Figure 13 It is a schematic structural diagram of the 3D digital human video generation system provided by the present invention.

[0030] Figure 14 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0032] In the existing technology, when facing the technical need to generate high-quality face texture images, it is impossible to use face images captured by ordinary shooting devices such as mobile phones or digital cameras in a low-cost and high-quality manner. Based on this, an embodiment of the present invention provides a method for generating a face texture image.

[0033] Figure 1 It is a schematic flowchart of a method for generating a face texture image provided in an embodiment of the present invention, as Figure 1 shown, the method includes: S11, receiving a face image mapping image of a target object; S12, based on a target diffusion model, using the face image mapping image as a constraint condition, denoising a noise image to obtain a face texture image of the target object; Among them, the target diffusion model is obtained by performing forward diffusion and reverse diffusion training based on face texture image samples using face image mapping image samples as constraint conditions.

[0034] Specifically, for the method for generating a face texture image provided in an embodiment of the present invention, the execution subject is a face texture image generation device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0035] First, step S11 is executed to receive a face image mapping image of a target object. The target object refers to a person whose face image mapping image is required to generate a high-quality face texture image.

[0036] The face image mapping image in the embodiment of the present invention can be a two-dimensional face texture image obtained by parsing at least one face image of a target object, and can be generated through the following steps: by parsing at least one face image of a target object, using 3D face reconstruction technology to generate a 3D face corresponding to the face image, and then generating an image of the 3D face surface in the two-dimensional texture space (UV space) based on the two-dimensional texture coordinates (UV coordinates) of each point on the 3D face surface, thereby establishing a mapping relationship between the face image and the two-dimensional texture image. The generated two-dimensional texture image is the face image mapping image. Specifically, at least one face image of the target object can be input into a three-dimensional deformable face model (3DMorphable Model, 3DMM) to obtain the facial geometric information of the face and the face image mapping image. Among them, the facial geometric information of the face can be used as position reference information when generating the face image mapping image, that is, used to determine the two-dimensional texture coordinates of each point on the 3D face surface. The specific generation process of the face image mapping image is not elaborated here.

[0037] It can be understood that in the embodiments of the present invention, at least one face image of a target object is parsed to generate a face image mapping image. However, the defects existing in the face image will result in a low quality of the generated face image mapping image, and further processing of the face image mapping image through image processing techniques is required to obtain a high-quality face texture image.

[0038] Then, step S12 is executed. By using a target Diffusion model, which can use the face image mapping image as a constraint condition, denoising of the noise image can be performed to obtain the face texture image of the target object.

[0039] During the training process of the target Diffusion model, the face texture image samples in the training data can be randomly selected from a high-quality texture image library. Then, face image mapping image samples can be synthesized through the face texture image samples. Pairing the face texture image samples with the synthesized face image mapping image samples can obtain training sample pairs. It should be understood that by performing image processing operations such as illumination processing, shadow processing, defect processing, resolution processing, and background processing on the low-quality face image mapping image, a high-quality face texture image can be obtained. Therefore, by performing the reverse operations of the above image processing operations on the face texture image samples, the face image mapping image samples can be obtained. Each face image mapping image sample and the corresponding face texture image sample can form a training sample pair for training to obtain the target Diffusion model.

[0040] The target Diffusion model can take both the face image mapping image and the noise image as inputs. By using the face image mapping image to assist in the denoising operation of the noise image, a face texture image with the same resolution as the noise image and highly correlated with the content of the face image mapping image can be obtained as the output. In the above example, the low-quality face image mapping image is inferior to the high-quality face texture image in terms of resolution, noise, artifacts, contrast, etc.

[0041] Here, the noise image used can be a Gaussian noise map or a random noise map, and no specific limitation is made here.

[0042] The target diffusion model may include a UNet neural network. Through the UNet neural network, with the face image mapping image as a constraint (condition), the noise image can be gradually denoised. Here, the denoising steps of the target diffusion model for the noise image can be set as needed. For example, it can be set to T denoising steps. While taking both the face image mapping image and the noise image as inputs, the noise addition step identifier corresponding to the current denoising step can also be taken as an input together to guide the target diffusion model to denoise the noise image. The target diffusion model will perform different processes on the input noise image according to the noise addition step identifier to achieve denoising. Among them, if the target diffusion model involves a total of T noise addition steps during training, the noise addition step identifiers can be respectively represented as 1, 2,..., T - 1, T. Correspondingly, the noise addition step identifiers corresponding to each denoising step can be respectively represented as T, T - 1,..., 2, 1.

[0043] The target diffusion model can be obtained by training an initial diffusion model. When training the initial diffusion model, the initial diffusion model can be used to perform progressive forward diffusion and reverse diffusion on texture image samples, and use the face image mapping image sample as a constraint condition during each step of reverse diffusion. Then, calculate the loss through the texture image sample and the reverse diffusion result of each step, and use this loss to iteratively train the initial diffusion model to finally obtain the target diffusion model.

[0044] As Figure 2 shown, during the training process of a conventional diffusion model, forward diffusion is a process of gradually increasing the complexity of data, that is, the noise addition process (Forward Diffusion Process). Through a series of reversible and progressive modifications, a certain amount of noise is introduced into the original image sample at each step. As the noise is continuously introduced, the complexity of the original image sample increases continuously. Finally, a noisy image (including the original image sample and the superimposed T noises) very similar to the desired complex data distribution (such as a Gaussian distribution) is obtained, thus transforming a structured data distribution into a noise distribution. Figure 2 Among them, x0 is the original image sample, x t-1 is the noise addition result of the (t - 1)-th step, x t is the noise addition result of the t-th step, x T is the noise addition result of the T-th step, that is, the noisy image. Among them, T is the total number of noise addition steps in the noise addition process.

[0045] As Figure 3 shown, during the training process of a conventional diffusion model, reverse diffusion, that is, the denoising process, can, by gradually performing denoising steps opposite to those of forward diffusion, start from the noise addition result x TA restored image is obtained. It can be understood that the total number of noise addition steps in the noise addition process should be the same as the total number of noise removal steps in the noise removal process. In some improved solutions, the noise removal process can be accelerated by means of "skipping steps" to reduce the consumption of computing power resources and time, but the basic principle of the noise removal process is not changed, and it will not be elaborated here.

[0046] As Figure 4 shown, it is an overall schematic diagram of the training process of a conventional diffusion model. Figure 4 In it, the original image sample x0 becomes a noisy image x through a noise addition process including T noise addition steps T , and the noisy image x T obtains a restored image through a noise removal process including T noise removal steps (i.e., Repeat T times) . Among them, the conventional diffusion model is composed of a UNet neural network.

[0047] Considering that each noise addition step is repeatedly generating random noise (specifically, random noise can be sampled from a standard Gaussian distribution) and adding noise, and each noise removal step is repeatedly predicting the random noise introduced during noise addition and removing the noise. For the convenience of explanation, only a single noise addition step and the corresponding noise removal step will be described.

[0048] Since the image restoration scenarios in the noise removal process usually include two types, one is generating random images, and the other is generating images that conform to the text description based on the input text. It can be understood that if it is to generate random images, only images need to be used as training samples to train the diffusion model, and if it is to generate images that conform to the text description, both the text description and the corresponding images need to be used as training samples to train the diffusion model.

[0049] A. Scenario of generating random images: In the t-th noise addition step, random noise is introduced into the input image to obtain a noise-added image, and the random noise in the t-th noise addition step is recorded as the training data for the corresponding noise removal step.

[0050] In the (T - t)-th noise removal step, based on the noise-added image and the noise addition step identifier corresponding to the (T - t)-th noise removal step (i.e., the value of t), the random noise introduced in the t-th noise addition step is predicted, and the loss between the predicted noise and the actual random noise introduced in the t-th noise addition step is calculated.

[0051] For example, in the first noise removal step, the corresponding noise addition step identifier is T, and the embedding vector of T (i.e., Time step embedding T) and the noisy image x TInput into the UNet neural network, and the UNet neural network predicts the random noise introduced in the T-th noise addition step. In the T-th denoising step, the corresponding noise addition step is labeled as 1. The embedding vector of 1 (i.e., Time step embedding 1) and the noisy image x1 are input into the UNet neural network, and the UNet neural network predicts the random noise introduced in the first noise addition step.

[0052] Based on the loss calculated for each denoising step, adjust the parameters of the UNet neural network, and perform denoising and loss calculation again. Repeat the above steps multiple times until the UNet neural network converges.

[0053] B. Scenario of generating an image that conforms to the text description based on the input text: In the t-th noise addition step, introduce random noise to the input image to obtain a noisy image. Record the random noise in the t-th noise addition step as the training data for the corresponding denoising step.

[0054] In the (T - t)-th denoising step, based on the text description corresponding to the image, the noisy image, and the noise addition step label corresponding to the (T - t)-th denoising step (i.e., the value of t), predict the random noise introduced in the t-th noise addition step, and calculate the loss between the predicted noise and the actual random noise introduced in the t-th noise addition step.

[0055] Based on the loss calculated for each denoising step, adjust the parameters of the UNet neural network, and perform denoising and loss calculation again. Repeat the above steps multiple times until the UNet neural network converges.

[0056] (2)Usage of the conventional diffusion model Based on the foregoing description of the basic principle of the conventional diffusion model, the usage process of the conventional diffusion model is similar to the denoising process in the training process of the conventional diffusion model, and only includes T denoising steps of the reverse diffusion process (i.e., denoising the noisy image into an image). In each denoising step, a certain amount of noise is removed from the input noisy image until an image is generated. Each denoising step is processed by the same UNet neural network, and the UNet neural network will perform different processes on the input noisy image according to the denoising step label.

[0057] For the sake of illustration, only a single denoising step is still described here.

[0058] A. Scenario of generating a random image: In the (T - t)-th denoising step, based on the noisy image and the value of t, predict the random noise introduced in the t-th noise addition step, and remove the predicted noise from the noisy image to complete the denoising operation of the (T - t)-th denoising step. Here, the value of t needs to be encoded to obtain an embedding vector.

[0059] B. Scenario of generating an image that conforms to the text description based on the input text: In the (T - t)-th denoising step, based on the text description corresponding to the image, the noisy image, and the value of t, predict the random noise introduced in the t-th noise addition step, and remove the predicted noise from the noisy image to complete the denoising operation of the (T - t)-th denoising step.

[0060] After completing all denoising steps, the generation of the image is completed.

[0061] In the embodiment of the present invention, compared with the training process of the conventional diffusion model, the training process of the initial diffusion model is different in that a face image mapping image sample is introduced as a constraint condition. When training the initial diffusion model, calculate the loss through the face texture image sample and the restored image, and iteratively train the initial diffusion model according to the loss to obtain the target diffusion model.

[0062] The face texture image generation method provided in the embodiment of the present invention first receives the face image mapping image of the target object; then, based on the target diffusion model, uses the face image mapping image as a constraint condition to denoise the noise image to obtain the texture image of the target object. By using the face image mapping image as a constraint condition for the target diffusion model, a high-quality face texture image can be obtained to meet the user's needs, and thus it can be applied to the field of 3D digital human generation. This method can avoid the problems of high implementation cost and high professional equipment operation threshold faced in the prior art for generating high-quality face texture images. Only by using the face texture image obtained from the face image collected by ordinary shooting devices such as mobile phones or digital cameras, a high-quality face texture image can be obtained through conversion by the target diffusion model, with low implementation cost and low equipment operation threshold.

[0063] Since the image space is a high-dimensional space, directly adding noise to and denoising the image itself requires a large amount of data to be processed, resulting in a large consumption of computing power. Based on this, based on the target diffusion model, using the face image mapping image as a constraint condition to denoise the noise image to obtain the face texture image of the target object, including: Based on the target diffusion model, respectively extract the target image features of the face image mapping image and the latent features of the noise image; Use the target image features as a constraint condition to denoise the latent features to obtain the target denoised features, and generate the face texture image based on the target denoised features.

[0064] Specifically, the target diffusion model adopted in the embodiments of the present invention can introduce an encoding and decoding function. When generating a face texture image using the target diffusion model, the target diffusion model can be first used to encode the face image mapping image and the noise image respectively to obtain the target image features of the face image mapping image and the latent features of the noise image, thereby converting the data processing from the high-dimensional image space to the low-dimensional feature space. This low-dimensional feature space has the same data processing effect as the high-dimensional image space and is therefore also called the latent space (Latent Space). Here, both the target image features and the latent features are one-dimensional feature vectors.

[0065] After that, using the target image features as a constraint condition, the latent features are denoised to obtain the target denoised features. Among them, the process of denoising the latent features is the same as the process of denoising the noise image in the above embodiments, and the only difference is the object of denoising and the dimension of data processing.

[0066] In the process of denoising the noise image by the target diffusion model, the embedding features (time step embedding) of the noise addition step identifier corresponding to each denoising step can be extracted, and in each denoising step of using the target image features as a constraint condition to denoise the latent features, the input target image features and the embedding vector of the noise addition step identifier corresponding to this denoising step are combined to jointly perform noise prediction, and the difference between the input previous denoising result and the predicted noise is used as the denoising result of this denoising step.

[0067] After determining the target denoised features, the target diffusion model can be used to decode the target denoised features, and then the face texture image can be generated.

[0068] In the embodiments of the present invention, by the target diffusion model, the target image features of the face image mapping image and the latent features of the noise image are respectively extracted, and the data processing of the image dimension is converted into the feature dimension of the latent space, which can reduce the amount of data required to be processed and reduce the computing power consumption while achieving the same data processing effect.

[0069] Based on the above embodiments, using the target image features as a constraint condition to denoise the latent features to obtain the target denoised features includes: Using the target image features as a constraint condition and adopting a cross-attention mechanism to denoise the latent features to obtain the target denoised features.

[0070] Specifically, in the embodiments of the present invention, since the data sources of the target image features and the latent features are different, a cross-attention component can be connected to the target diffusion model to introduce the cross-attention mechanism (Cross-Attention Mechanism, CAM). The target image features and the latent features are combined through CAM to realize noise prediction and denoising of the noisy image, and the target denoised features can be obtained. The UNet neural network includes several encoding blocks and decoding blocks, and each encoding block and decoding block respectively includes several neural network layers. Therefore, the cross-attention component can be selectively connected to the specified layers in the encoding blocks and decoding blocks to introduce the cross-attention mechanism. The specified layer can include one or more, and no specific limitation is made here. Furthermore, the UNet neural network can include one or more cross-attention components. The target image features and the latent features can be input into the input layer of the UNet neural network, and the target image features are input into each cross-attention component, so that each cross-attention component combines the input features.

[0071] In the embodiments of the present invention, introducing the cross-attention mechanism to denoise the latent features of the noisy image can enable the target diffusion model to effectively align and focus on the features from different data sources, thereby helping the target diffusion model better capture the correlation between the features from different data sources, and enabling the target image features as the constraint conditions to directly affect the specified layers of the UNet neural network, avoiding the deep network of the UNet neural network forgetting the constraint conditions.

[0072] Based on the above embodiments, the target diffusion model includes a perceptual image compression module, a latent diffusion model, and a conditional encoding module; the perceptual image compression module includes an encoder and a decoder; The face texture image is specifically generated through the following steps: Extract the target image features from the face image mapping image based on the conditional encoding module; Extract the latent features from the noisy image based on the encoder; Based on the latent diffusion model, use the target image features as the constraint conditions to denoise the latent features to obtain the target denoised features; Based on the decoder, apply the target denoised features to generate the face texture image.

[0073] Specifically, the target diffusion model can include a perceptual image compression module (Perceptual Image Compression), a latent diffusion model (Latent Diffusion Model, LDM), and a conditional encoding module. The perceptual image compression module includes an encoder and a decoder. At this time, the target diffusion model can be a stable diffusion model (Stable Diffusion).

[0074] The perceptual image compression module can be a trained Variational Auto Encoder (VAE). The encoder (E) of the VAE is the encoder of the perceptual image compression module, and the decoder (D) of the VAE is the decoder of the perceptual image compression module.

[0075] The encoder of the perceptual image compression module can map an image into a distribution in the Latent Space, specifically in the form of discrete latent features in the latent space. The decoder has the opposite function, that is, it maps from the latent space to the image space according to the distribution. The input of the encoder of the perceptual image compression module is a noisy image. This encoder is used to encode the noisy image, extract latent features from the noisy image and use them as the output, so as to convert high-dimensional image data into low-dimensional feature vectors for processing.

[0076] The conditional encoding module can specifically be a trained Contrastive Language-Image Pre-training (CLIP) model. The conditional encoding module can include multiple types of encoders, which can encode inputs of different modalities respectively. For example, the conditional encoding module can at least include an image encoder for encoding the input image and converting the input image into image features. This image encoder can be a convolutional neural network (such as ResNet, etc.) or a Transformer model (such as ViT, etc.). On this basis, the conditional encoding module can also include a text encoder for encoding the input text and converting the input text into text features. This text encoder can be a Transformer model or other structures.

[0077] In the embodiments of the present invention, the encoders included in the conditional encoding module can share a vector space to achieve cross-modal information interaction and fusion.

[0078] The input of the conditional encoding module includes a face image mapping image. The face image mapping image is encoded by the image encoder in the conditional encoding module, and target image features are extracted from the face image mapping image and used as an output of the conditional encoding module.

[0079] The latent diffusion model can be a UNet neural network, which includes several encoding blocks and decoding blocks. Each encoding block and decoding block respectively contains several neural network layers. Specified layers can be selected in the encoding block and decoding block respectively to access the cross-attention component, so as to introduce the cross-attention mechanism. The specified layer can include one or more, and no specific limitation is made here. Furthermore, the latent diffusion model can include one or more cross-attention components. The latent feature and the target image feature can be input into the input layer of the UNet neural network, and the target image feature can be input into each cross-attention component, so that each cross-attention component combines the input features.

[0080] The outputs of the encoder of the perceptual image compression module and the conditional encoding module are both used as the input of the latent diffusion model. The latent diffusion model uses the target image feature as a constraint condition to denoise the latent feature, and obtains the target denoised feature as the output. Different from the implementation manner of the target diffusion model in the above embodiment, the denoising object of this latent diffusion model is not the noise image in the image space, but the discrete latent feature of the noise image in the latent space.

[0081] Through the UNet neural network, with the target image feature as the constraint condition (condition), the latent feature can be gradually denoised. Here, the denoising steps of the latent feature by the target diffusion model can be set as needed. For example, it can be set to T denoising steps.

[0082] The output of the latent diffusion model is used as the input of the decoder, and the decoder is used to decode the target denoised feature to obtain the face texture image as the output, so as to restore the processed low-dimensional feature vector to high-dimensional image data.

[0083] In the embodiment of the present invention, the specific structure of the target diffusion model is given. Through the perceptual image compression module and the conditional encoding module, the input of the latent diffusion model is a feature, so as to reduce the amount of data to be processed and reduce the computing power consumption.

[0084] On the basis of the above embodiment, the latent diffusion model is obtained by training the initial diffusion model based on the following steps: Based on the encoder, extract the sample latent feature of the face texture image sample, and apply target noise to the sample latent feature to obtain the noise feature; Based on the conditional encoding module, extract the sample image feature of the face image mapping image sample; Based on the initial diffusion model, use the sample image feature as the constraint condition to perform noise prediction on the noise feature to obtain the predicted noise; Based on the target noise and the predicted noise, calculate the noise loss, and based on the noise loss, iteratively train the initial diffusion model to obtain the latent diffusion model.

[0085] Specifically, the latent diffusion model can be obtained by training the initial diffusion model, which has the same structure as the latent diffusion model, with the only difference being the neural network parameters.

[0086] It can be seen from this that in the target diffusion model, both the perceptual image compression module and the conditional encoding module are trained and directly applicable functional modules, while the latent diffusion model is a functional module that needs to be trained before it can be applied. Furthermore, the process of training the target diffusion model is the same as the process of training the latent diffusion model. After obtaining the latent diffusion model, by splicing it with the perceptual image compression module and the conditional encoding module, the target diffusion model can be obtained.

[0087] In the embodiments of the present invention, as Figure 5 shown, during the process of training the initial diffusion model, the face texture image sample y0 can be first input into the encoder (E) in the perceptual image compression module, and the sample latent feature z0 of the face texture image sample is extracted through this encoder (E).

[0088] After that, the sample latent feature z0 can be gradually noise-added through a noise-adding process, that is, the target noise is gradually applied to the sample latent feature z0, and after T noise-adding steps, the noise map z T is obtained. The target noise applied in each step can be sampled from Gaussian distribution or approximately Gaussian distribution noise, and the target noise applied in all noise-adding steps can form a Gaussian distribution or approximately Gaussian distribution through superposition.

[0089] Synchronously, the face image mapped image sample can be input into the conditional encoding module τ θ , and the sample image feature of the face image mapped image sample is extracted through the conditional encoding module τ θ , and the noise feature of the noise map z T is extracted through the encoder (E). .

[0090] After that, through a denoising process, both the sample image feature and the noise feature are used as the input of the initial diffusion model. Using the initial diffusion model and using the sample image feature as a constraint condition, the noise feature is predicted for noise prediction to obtain the predicted noise. The difference between the noise feature and the predicted noise is the first denoising result during the training process . It can be understood that when the sample image feature and the noise feature While all are used as inputs to the initial diffusion model, the embedding vector of the noise addition step identifier T corresponding to the first denoising step can also be input into the initial diffusion model to guide the orderly progress of the denoising steps.

[0091] After that, using the target noise and the predicted noise, a suitable loss function can be selected to calculate the noise loss. Using the noise loss, the initial diffusion model can be iteratively trained until the noise loss converges or reaches a preset number of iterations to obtain the latent diffusion model. Figure 5 The preset number of iterations in it is T. In the T-th denoising step, the (T - 1)-th denoising result , the sample image features, and the embedding vector of the noise addition step identifier 1 corresponding to the T-th denoising step are input into the initial diffusion model together. After the T-th denoising step, denoised features can be obtained. These denoised features pass through the decoder (D) in the perceptual image compression module to obtain the noise features corresponding restored image.

[0092] In the embodiments of the present invention, only the latent diffusion model needs to be trained, and other modules in the target diffusion model are all pre-trained, so that the model training cost can be greatly reduced and the generality of the model can be improved.

[0093] Based on the above embodiments, based on the target diffusion model, using the face image mapping image as a constraint condition, the noise image is denoised to obtain the face texture image of the target object, and it further includes: Receiving the image description information of the face texture image; Based on the target diffusion model, using the face image mapping image and the image description information as constraint conditions, the noise image is denoised to obtain the face texture image; Among them, the target diffusion model is specifically trained by forward diffusion and backward diffusion based on the sample description information of the face image mapping image sample and the face texture image sample as constraint conditions.

[0094] Specifically, in the embodiments of the present invention, during the process of the target diffusion model obtaining the face texture image, the execution entity can also receive the image description information of the face texture image, and the image description information can include at least one of information such as description text (text) and semantic map.

[0095] After that, the image description information, the face image mapping image, and the noise image can be jointly input into the target diffusion model. The target diffusion model uses the face image mapping image and the image description information as joint constraint conditions to denoise the noise image and obtain the face texture image.

[0096] Here, as Figure 6 shown, when the target diffusion model includes a perceptual image compression module, a latent diffusion model, and a conditional encoding module, the image description information can be encoded by the conditional encoding module to extract description information features from the image description information.

[0097] For example, the description text in the image description information is input into the text encoder in the conditional encoding module, and the text features in the input description text are extracted by the text encoder. The semantic map in the image description information is input into the map encoder in the conditional encoding module, and the map features in the input semantic map are extracted by the map encoder.

[0098] Similar to the operations in the above embodiments, the face image mapping image needs to be input into the image encoder in the conditional encoding module, and the face image mapping image is encoded by the image encoder to extract target image features from the face image mapping image. The noise image is input into the encoder of the perceptual image compression module, and latent features are extracted from the noise image by this encoder.

[0099] After that, the target image features, the description information features, and the latent features are input into the latent diffusion model, which adopts a UNet neural network. Using the latent diffusion model, the target image features and the description information features are used as joint constraint conditions to denoise the latent features and obtain target denoised features. Here, the latent features of the noise image, the embedding features of the noise addition step identifier corresponding to the current denoising step, the target image features, and the description information features can be jointly input into the input layer of the latent diffusion model, and the target image features and the description information features are input into each cross-attention component, so that each cross-attention component combines the input features.

[0100] As Figure 7 shown, taking the first denoising step of the latent features by the UNet neural network as an example, the noise addition step identifier is T, and its embedding features can be represented as p T , and the input of the UNet neural network includes the embedding features p T , the latent features z T , the target image features I, and the description information features τ. Among them, the embedding features p T and the latent features z TIt only needs to be input into the input layer of the UNet neural network once. The target image feature I and the description information feature τ need to be input not only into the input layer of the UNet neural network but also into each cross-attention component. Finally, through the first denoising step of the UNet neural network, the first denoising result z is obtained. T-1 。

[0101] After that, the first denoising result is subjected to a second denoising step, that is, the first denoising result z T-1 , the embedding feature of the noise addition step identifier corresponding to the second denoising step, the target image feature I, and the description information feature τ are re-input into the UNet neural network. At this time, the noise addition step identifier becomes T-1, and the embedding feature becomes p T-1 . Iterate in this way until the Tth denoising step is completed, that is, the target denoising feature is obtained.

[0102] Finally, using the decoder, the target denoising feature is decoded to generate a face texture image.

[0103] In the embodiments of the present invention, by introducing the image description information of the face texture image, the target diffusion model can denoise the noise image based on more information, so that the obtained face texture image can better meet the requirements.

[0104] Based on the above embodiments, as Figure 8 shown, in the embodiments of the present invention, a 3D digital human generation method is provided, and the method includes: S21, receiving at least one face image of a target object; S22, based on at least one face image, determining the face image mapping image and facial geometry information of the target object; S23, based on the face image mapping image, applying the face texture image generation method provided in the above embodiments to generate a face texture image; S24, based on the facial geometry information and the face texture image, generating a 3D digital face corresponding to the target object, and based on the 3D digital face, generating a target 3D digital human.

[0105] Specifically, for the 3D digital human generation method provided in the embodiments of the present invention, the execution subject is a 3D digital human generation system, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0106] First, step S21 is executed to receive at least one face image of the target object uploaded by the user. The target object can be a person for whom a 3D digital human with the same face is required. Among the at least one face image, it can include at least one of the frontal face image, profile face image, or face images at other angles of the target object. The profile face image can include the left profile face image and the right profile face image. Other angles can include 45 degrees obliquely, 60 degrees obliquely, etc. Face images at other angles can include the left / right front 45-degree profile face image, the left / right front 60-degree profile face image, etc.

[0107] Here, each face image can be a two-dimensional image collected by an ordinary photographing device such as a mobile phone or a digital camera. Although the data accuracy of each face image taken by an ordinary photographing device is not high, it still records the face features and facial skin texture of the target object. The face image can still be analyzed and processed by means of an image processing algorithm to create a hyper-realistic 3D digital human that is relatively similar to the target object. Compared with the high-precision 3D face reconstruction scheme, although the similarity is a bit lower, the cost has been greatly reduced and can be borne by ordinary users. In usage scenarios such as short video production and online social networking, it is sufficient to meet the usage needs of ordinary users.

[0108] After receiving the face image, preprocessing can also be performed on the face image. For example, image processing algorithms such as EyeGlasses GAN can be used to eliminate accessories such as glasses and earrings in the face image to obtain a pure face image.

[0109] Then, step S22 is executed. Using the received face image, the face image mapping image and facial geometric information of the target object are determined. Here, facial feature points can be directly extracted from the face image using a 3D face reconstruction model, and then the face image mapping image and facial geometric information are obtained.

[0110] The facial geometric information can be facial geometry features (identity), can be a two-dimensional feature image, or can be a one-dimensional feature vector, which is not specifically limited here.

[0111] The face image mapping image and facial geometric information can be obtained by processing the face image through a neural network model trained with multiple tasks.

[0112] It can be understood that when the face image uploaded by the user is one, the frontal face image can be preferably selected. When the face images uploaded by the user are multiple, one of them that contains a frontal face image can be preferably selected.

[0113] When extracting facial geometric information from multiple face images, the frontal face image can be used as the main one, and the remaining face images can be used to supplement the three-dimensional side features of the face. The facial geometric information of each face image at different angles can be extracted separately first. Based on the frontal face image, the distribution features of the facial features and the frontal face features can be accurately extracted, and based on the side face image, the three-dimensional side features can be accurately extracted. Then, the facial geometric information extracted from the face images at different angles can be combined to obtain accurate and complete facial geometric information.

[0114] When extracting the mapped images of face images from multiple face images, the frontal face image can be directly determined from the multiple face images uploaded by the user, and the mapped images of face images are extracted only based on the frontal face image. This implementation method has a relatively small amount of image processing and a high extraction efficiency of the mapped images of face images. In addition, when extracting the mapped images of face images from multiple face images separately, the mapped images of face images extracted through the frontal face image are used as the basis and main part, and the mapped images of face images extracted from the remaining face images are used to supplement and improve it. This implementation method has a high extraction accuracy of the mapped images of face images.

[0115] Subsequently, step S23 is executed. Using the mapped images of face images, the face texture image generation method provided in the above embodiments is applied to generate a face texture image. Since the face texture image generation method can be used to improve the quality of the image, in the case where the quality of the mapped images of face images is low, its quality can be improved through this method to make it a face texture image.

[0116] It should be noted that in addition to the face texture image generation method provided in the above embodiments, there are other image processing methods that can generate high-quality face texture images based on low-quality mapped images of face images, such as the Pixel2Pixel algorithm model, which will not be elaborated here.

[0117] Finally, step S24 is executed. Using the facial geometric information and the face texture image, a 3D digital face corresponding to the target object is generated. Here, the facial geometric information can be used first, and a three-dimensional geometric model of the face can be constructed using three-dimensional modeling software (such as Blender, Maya, etc.). Then, the face texture image is mapped onto the three-dimensional geometric model to obtain the 3D digital face.

[0118] The 3D digital face is spliced with a preset 3D digital human template torso to obtain a complete 3D digital human. Among them, the 3D digital human template torso can be randomly selected from the torso library, or can be selected by the user from the torso library, which is not specifically limited here. The torso library can include male template torsos and female template torsos of different torso types.

[0119] In the 3D digital human generation method provided in the embodiments of the present invention, since the face texture image generation method provided in the above embodiments is introduced, the quality of the image is improved, which can make the details of the finally obtained 3D digital human clearer, highly similar to the target object in appearance, and of higher quality. Moreover, for this method, the user only needs to upload at least one face image collected by an ordinary shooting device, with simple operation and low cost.

[0120] Based on the above embodiments, a 3D digital face corresponding to the target object is generated based on the facial geometric information and the face texture image, including: Classify the eyebrow shape and / or hairstyle in at least one face image to determine the type of the eyebrow shape and / or hairstyle; Determine a template eyebrow shape of the same type based on the type of the eyebrow shape, and / or determine a template hairstyle of the same type based on the type of the hairstyle; Generate a 3D digital face based on the template eyebrow shape and / or template hairstyle, as well as the facial geometric information and the face texture image.

[0121] Specifically, during the process of generating the 3D digital face, in order to make the generated 3D digital face have a similar eyebrow shape and / or hairstyle to the face in the received face image, the category of the eyebrow shape and / or hairstyle in the received face image can also be recognized, and the same eyebrow shape and / or hairstyle can be bound to the 3D digital human.

[0122] Here, to classify the eyebrow shape and / or hairstyle in the face image, a classification and recognition model can be introduced, and this classification and recognition model can be the SegFormer network. The SegFormer network can include an encoder and a decoder. The encoder can include multiple layers of Transformer structures, and the decoder can include a lightweight Multilayer Perceptron (MLP).

[0123] Input each face image into the classification and recognition model, and the type of the eyebrow shape and / or hairstyle in the face image can be output through the classification and recognition model.

[0124] During the process of training the classification and recognition model, the initial SegFormer network can be trained first using a large amount of training data, and then, with the decoder remaining unchanged, the encoder can be trained in a second stage using refined training data with type labels of the eyebrow shape and / or hairstyle manually labeled, and finally the classification and recognition model is obtained.

[0125] After that, the type of the eyebrow shape can be used to determine the pre-stored template eyebrow shape of the same type in the eyebrow shape library, or the type of the hairstyle can be used to determine the pre-stored template hairstyle of the same type in the hairstyle library.

[0126] Finally, a 3D digital human face is generated by using the template eyebrow shape and / or template hairstyle, as well as the facial geometric information and the human face texture image. That is, the template eyebrow shape and / or template hairstyle are superimposed on the initial digital human face obtained from the facial geometric information and the human face texture image to generate the final 3D digital human face.

[0127] In the embodiment of the present invention, by adding a template eyebrow shape and / or template hairstyle that conforms to the human face image to the 3D digital human face, the generated 3D digital human face can be made more similar to the human face image input by the user.

[0128] On the basis of the above embodiment, a 3D digital human face corresponding to the target object is generated based on the facial geometric information and the human face texture image, and then it further includes: Receiving an adjustment instruction from the user for a specified area of the 3D digital human face; Adjusting the specified area based on the adjustment instruction.

[0129] Specifically, after generating the 3D digital human face, the 3D digital human face can be displayed to the user. The user can operate on a specified area of the 3D digital human face through a mouse, keyboard or touch screen, so that the executing entity responds to the operation and receives an adjustment instruction from the user for the specified area. The specified area may include one or more of areas such as the whole human face, the nose area, and the eye area. The adjustment instruction may be to widen / narrow the face shape, enlarge / shrink the nose, move / rotate the eyes, etc.

[0130] Thereafter, the executing entity can use the adjustment instruction to adjust the specified area, so that the generated 3D digital human face can better meet the user's needs.

[0131] On the basis of the above embodiment, before determining the face image mapping image and the facial geometric information of the target object based on at least one human face image, it includes: Detecting at least one of facial expressions, abnormal content, human face proportion, human face position, and human face integrity in at least one human face image, and detecting whether there are human face images at multiple angles in the case where the at least one human face image includes multiple human face images; If any of the human face images fails to pass the detection or there are no human face images at multiple angles, a prompt message is sent to the user.

[0132] Specifically, in the embodiments of the present invention, to avoid the influence of facial expressions in the face images on the generated 3D digital face, the face images uploaded by the user need to have no facial expressions. To ensure the legality of the content of the face images, the face images uploaded by the user cannot contain abnormal content such as pornographic, violent, or political content. Moreover, to ensure the quality of the generated 3D digital face, the face images uploaded by the user need to meet the requirements that the face is included and the proportion of the face is greater than or equal to a preset threshold. The position of the face can be centered, and the face cannot be blocked, etc. Here, the preset threshold can be set as needed. For example, it can be set to be greater than or equal to 50%, or it can be set to other values, which are not specifically limited here.

[0133] Therefore, after receiving at least one face image and before applying each face image, at least one of the facial expressions, abnormal content, face proportion, face position, and face integrity in each face image can be detected. For any face image, when detecting the facial expression, if there is a facial expression in the face image, the facial expression detection of the face image fails. When detecting the abnormal content, if there is abnormal content in the face image, the abnormal content detection of the face image fails. When detecting the face proportion, if the proportion of the face in the face image is less than the preset threshold, the face proportion detection of the face image fails. When detecting the face position, if the deviation of the face position in the face image from the center position of the image is greater than the preset value, the face position detection of the face image fails. When detecting the face integrity, if the face in the face image is blocked, the face integrity detection of the face image fails.

[0134] Meanwhile, in the case of multiple face images, it is also necessary to detect whether there are face images at multiple angles in the multiple face images, that is, whether there are at least two images among the frontal face image and the side face images at various angles at the same time. If so, the detection passes; if not, the detection fails.

[0135] After that, if there is a face image with a failed detection, or there are no face images at multiple angles, it is considered that a 3D digital face cannot be generated based on the face images uploaded by the user. Furthermore, a prompt message needs to be sent to the user to prompt the reason for the failed detection and prompt the user to re-upload one or more face images.

[0136] To ensure that the face images uploaded by the user pass the detection, the preferred shooting method can also be displayed to the user on the interface for uploading face images. For example, it can be to use a lens with a 2x focal length to shoot at a distance of 1.5 meters from the face.

[0137] In the embodiments of the present invention, by detecting the face images uploaded by the user, the quality of the generated 3D digital face can be ensured.

[0138] Based on the above embodiments, based on at least one face image, determining a face image mapping image and facial geometry information of a target object, including: Based on at least one face image, determining a face image mapping image, facial geometry information, and facial auxiliary features, where the facial auxiliary features include at least one of expression features, pose features, and lighting features.

[0139] Correspondingly, based on the facial geometry information and the face texture image, generating a 3D digital face corresponding to the target object, including: Based on the facial geometry information, determining a target face model, and based on the face texture image, performing texture mapping on the target face model to obtain an initial 3D face; Based on the facial auxiliary features, optimizing the initial 3D face to obtain a 3D digital face.

[0140] Specifically, when determining the face image mapping image and facial geometry information of the target object, each face image can be used, and a neural network model for multi-task learning can be adopted to synchronously determine the face image mapping image, facial geometry information, and facial auxiliary features. The facial auxiliary features can include at least one of expression features, pose features, and lighting features.

[0141] Furthermore, when generating a 3D digital face corresponding to the target object, the target face model can be determined first using the facial geometry information. Here, the facial geometry information can be used to process a template face model to obtain the target face model. The template face model can be a pre-determined face model, and both the template face model and the target face model are three-dimensional face models.

[0142] Then, the target face model can be subjected to texture mapping processing using the face texture image to obtain a 3D face, which can be referred to as an initial 3D face. The facial auxiliary features also need to be used to optimize the initial 3D face, that is, to adjust the initial 3D face through the facial auxiliary features to obtain a 3D digital face, which can further improve the matching degree between the 3D digital face and the face image input by the user.

[0143] Based on the above embodiments, based on at least one face image, determining a face image mapping image, facial geometry information, and facial auxiliary features, including: Inputting at least one face image into a three-dimensional deformable face model to obtain a face image mapping image, facial geometry information, and facial auxiliary features determined by the three-dimensional deformable face model; Among them, the three-dimensional deformable face model is trained based on face image samples at different angles.

[0144] Specifically, the neural network model for multi-task learning used in the embodiments of the present invention can be a three-dimensional deformable face model (3D Morphable Model, 3DMM). The 3DMM can be a backbone network. For example, ConvNeXt can be used as the backbone network. The 3DMM can include a face geometry basis recognition module and a texture basis recognition module. The face geometry basis recognition module is used to determine the facial geometry information of the target object according to at least one input face image, and the texture basis recognition module is used to determine the face image mapping image of the target object according to at least one input face image. For facial auxiliary features, they can be directly extracted by the trained 3DMM when determining the facial geometry information and the face image mapping image of the target object.

[0145] Here, the face geometry basis recognition module and the texture basis recognition module can be trained separately to complete the training of the 3DMM. For the training of the face geometry basis recognition module, it can be achieved through face image samples at different angles and pre-annotated principal component analysis coefficients. For the training of the texture basis recognition module, it can be achieved through face image samples at different angles and corresponding high-quality texture images.

[0146] As Figure 9 shown, it is a schematic diagram of the complete process of the 3D digital face generation method. Inputting at least one face image input by the user into the 3DMM, a face image mapping image, facial geometry information, and facial auxiliary features can be obtained. Inputting the face image mapping image into the target diffusion model, a face texture image can be obtained. Combining the facial geometry information, the face texture image, and the facial auxiliary features, a 3D digital face can be generated.

[0147] Based on the above embodiments, generating a target 3D digital human based on the 3D digital face includes: Receiving the target torso type selected by the user; Generating a target 3D digital human based on the template torso corresponding to the target torso type and the 3D digital face; Or, Extracting the style features in at least one face image; Generating a target 3D digital human based on the template torso matching the style features and the 3D digital face.

[0148] Specifically, when generating the target 3D digital human, the 3D digital human template torso used can be selected by the user. Therefore, the execution entity can first receive the target torso type selected by the user, and then search for the template torso corresponding to the target torso type in the torso library according to the target torso type. This template torso is used as the 3D digital human template torso and is spliced with the 3D digital face to generate the target 3D digital human.

[0149] In addition, the 3D digital human template torso adopted can also be determined by at least one face image input by the user. Therefore, the execution entity can first extract the style features in each face image. For example, each face image can be input into the style feature extraction model, and the style features in each face image can be extracted through the style feature extraction model.

[0150] After that, a template torso that matches the style features can be searched for in the torso library as the 3D digital human template torso, and it can be spliced with the 3D digital human face to obtain the generated target 3D digital human.

[0151] In the embodiments of the present invention, two methods for determining the 3D digital human template torso are given. Determining through the face image input by the user can make the obtained 3D digital human template torso more in line with the user's needs. Automatically matching by the execution entity can make the obtained 3D digital human template torso more matching the input face image.

[0152] Based on the above embodiments, after generating the target 3D digital human based on the 3D digital human face, it includes: Receiving the 3D digital human display space selected by the user; Displaying the target 3D digital human in the 3D digital human display space; Or, Displaying the target 3D digital human in the 3D digital human display space that matches the style features.

[0153] Specifically, after generating the target 3D digital human, the target 3D digital human can also be displayed to the user. At this time, a 3D digital human display space needs to be determined for the target 3D digital human, and this 3D digital human display space can be selected by the user from the display space set. The display space set can include various types of display spaces determined in advance. Furthermore, the execution entity can receive the 3D digital human display space selected by the user and display the target 3D digital human in the 3D digital human display space, so as to achieve a display effect that meets the user's expectations.

[0154] In addition, the 3D digital human display space can also be obtained by automatic matching of the execution entity, that is, the execution entity can determine the display space that matches the style features in the display space set according to the style features extracted from at least one face image as the 3D digital human display space for displaying the target 3D digital human, so that the display effect can be more in line with the style type of the received face image.

[0155] Based on the above embodiments, after generating the target 3D digital human based on the 3D digital human face, it includes: Receiving the modification instruction of the user for the target 3D digital human; Update the target 3D digital human based on the modification instruction.

[0156] Specifically, after the target 3D digital human is generated, professionals as users can also perform manual refinement on the target 3D digital human. At this time, the execution entity can also receive the modification instruction of the user for the target 3D digital human. The modification instruction can be obtained by operating on the area to be modified of the target 3D digital human through a mouse, keyboard or touch screen.

[0157] The execution entity can use this modification instruction to update the target 3D digital human to improve the production quality of the target 3D digital human.

[0158] As Figure 10 shown, based on the above embodiments, the embodiments of the present invention also provide a method for generating a 3D digital human video, which includes: S31, obtain a target 3D digital human, which is generated based on the 3D digital human generation method provided in the above embodiments; S32, generate a 3D digital human video based on the target 3D digital human.

[0159] Specifically, for the method for generating a 3D digital human video provided in the embodiments of the present invention, the execution entity is a 3D digital human video generation system, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0160] For the target 3D digital human generated by the 3D digital human generation method provided in the above embodiments, a 3D digital human video is generated by using video production technology. For example, the characters in an existing video can be replaced with the target 3D digital human to obtain a 3D digital human video, or a 3D digital human video can be generated based on the existing text content with the target 3D digital human as the character image in the video, and no specific limitation is made here.

[0161] In addition, the 3D digital human video can also be used to drive the 3D digital human for live broadcast, and the 3D digital human can be used as a game character, etc.

[0162] For the method for generating a 3D digital human video provided in the embodiments of the present invention, by generating a 3D digital human video through the target 3D digital human generated by the foregoing 3D digital human generation method, the generated 3D digital human video can be made clearer, the texture of the target 3D digital human in the 3D digital human video is more delicate, the sensory effect of the user can be improved, and the user experience can be enhanced.

[0163] As Figure 11 shown, based on the above embodiments, the embodiments of the present invention provide a device for generating a face texture image, including: The first image receiving module 91 is configured to receive the face image mapping image of the target object; The image denoising module 92 is configured to denoise the noise image based on the target diffusion model by using the face image mapping image as a constraint condition to obtain the face texture image of the target object; Wherein, the target diffusion model is obtained by performing forward diffusion and reverse diffusion training based on the face texture image samples by using the face image mapping image samples as constraint conditions.

[0164] Specifically, the functions of the modules in the face texture image generation device provided in the embodiments of the present invention correspond one-to-one to the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not be elaborated herein.

[0165] As Figure 12 shown, on the basis of the above embodiments, the embodiments of the present invention provide a 3D digital human generation system, including: The second image receiving module 11 is configured to receive at least one face image of the target object; The texture geometry image determining module 12 is configured to determine the face image mapping image and the facial geometry information of the target object based on at least one face image; The face texture image generation device 13 is configured to generate a face texture image based on the face image mapping image by applying the face texture image generation method provided in the above embodiments; The digital human generation module 14 is configured to generate a 3D digital face corresponding to the target object based on the facial geometry information and the face texture image, and generate a target 3D digital human based on the 3D digital face.

[0166] Specifically, the functions of the modules in the 3D digital human generation system provided in the embodiments of the present invention correspond one-to-one to the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not be elaborated herein.

[0167] As Figure 13 shown, on the basis of the above embodiments, the embodiments of the present invention provide a 3D digital human video generation system, including: The 3D digital human generation system 101 is configured to obtain a target 3D digital human, and the target 3D digital human is generated based on the 3D digital human generation method provided in the above embodiments; The video generation module 102 is configured to generate a 3D digital human video based on the target 3D digital human.

[0168] Specifically, in the 3D digital human video generation system provided in the embodiments of the present invention, the functions of each module correspond one-to-one to the operation processes of each step in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not elaborate herein.

[0169] Figure 14 FIG. illustrates a schematic physical structure diagram of an electronic device, as Figure 14 shown. The electronic device may include: a processor 110, a communications interface 120, a memory 130, and a communication bus 140. Among them, the processor 110, the communications interface 120, and the memory 130 complete mutual communication through the communication bus 140. The processor 110 can call the logical instructions in the memory 130 to execute the face texture image generation method, or the 3D digital human generation method, or the 3D digital human video generation method provided in the above embodiments.

[0170] In addition, when the logical instructions in the above-mentioned memory 130 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0171] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the face texture image generation method, or the 3D digital human generation method, or the 3D digital human video generation method provided in the above embodiments.

[0172] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the face texture image generation method, or the 3D digital human generation method, or the 3D digital human video generation method provided in the above embodiments. The computer-readable storage medium can be either a non-transitory computer-readable storage medium or a transitory computer-readable storage medium, and no specific limitation is imposed here.

[0173] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0174] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a face texture image, characterized in that, Including: Receiving a face image mapping image of a target object; Based on a target diffusion model, using the face image mapping image as a constraint condition, denoising a noise image to obtain a face texture image of the target object; Wherein, the target diffusion model is trained by forward diffusion and reverse diffusion based on face texture image samples using face image mapping image samples as constraint conditions.

2. The method for generating a face texture image according to claim 1, wherein The step of, based on the target diffusion model, using the face image mapping image as a constraint condition, denoising the noise image to obtain the face texture image of the target object includes: Based on the target diffusion model, respectively extracting target image features of the face image mapping image and latent features of the noise image; Using the target image features as a constraint condition, denoising the latent features to obtain target denoised features, and generating the face texture image based on the target denoised features.

3. The method for generating a facial texture image according to claim 2, wherein The step of using the target image features as a constraint condition, denoising the latent features to obtain target denoised features includes: Using the target image features as a constraint condition, and adopting a cross-attention mechanism to denoise the latent features to obtain the target denoised features.

4. The method for generating a facial texture image according to claim 2, wherein The target diffusion model includes a perceptual image compression module, a latent diffusion model, and a conditional encoding module; the perceptual image compression module includes an encoder and a decoder; The face texture image is specifically generated through the following steps: Based on the conditional encoding module, extracting the target image features from the face image mapping image; Based on the encoder, extracting the latent features from the noise image; Based on the latent diffusion model, using the target image features as a constraint condition, denoising the latent features to obtain the target denoised features; Based on the decoder, applying the target denoised features to generate the face texture image.

5. The method for generating a facial texture image according to claim 4, wherein The latent diffusion model is obtained by training an initial diffusion model based on the following steps: Based on the encoder, extracting sample latent features of the face texture image sample, and applying target noise to the sample latent features to obtain noise features; Based on the conditional encoding module, extracting sample image features of the face image mapping image sample; Based on the initial diffusion model, using the sample image features as a constraint condition, performing noise prediction on the noise features to obtain predicted noise; Based on the target noise and the predicted noise, calculating a noise loss, and based on the noise loss, performing iterative training on the initial diffusion model to obtain the latent diffusion model.

6. The method for generating a face texture image according to any one of claims 1-5, characterized in that, The step of, based on the target diffusion model, using the face image mapping image as a constraint condition, denoising the noise image to obtain the face texture image of the target object further includes: Receiving image description information of the face texture image; Based on the target diffusion model, using the face image mapping image and the image description information as constraint conditions, denoising the noise image to obtain the face texture image; Among them, the target diffusion model is specifically trained by forward diffusion and reverse diffusion based on the face texture image sample, using the sample description information of the face image mapping image and the face texture image sample as constraint conditions.

7. A 3D digital human generation method, characterized in that, It includes: Receiving at least one face image of a target object; Based on the at least one face image, determining the face image mapping image and facial geometric information of the target object; Based on the face image mapping image, applying the face texture image generation method according to any one of claims 1-6 to generate a face texture image; Based on the facial geometric information and the face texture image, generating a 3D digital face corresponding to the target object, and based on the 3D digital face, generating a target 3D digital human.

8. The 3D digital life generation method according to claim 7, characterized in that, The generating a 3D digital face corresponding to the target object based on the facial geometric information and the face texture image includes: Classifying the eyebrow shape and / or hairstyle in the at least one face image to determine the type of the eyebrow shape and / or the hairstyle; Determining a template eyebrow shape of the same type based on the type of the eyebrow shape, and / or determining a template hairstyle of the same type based on the type of the hairstyle; Based on the template eyebrow shape and / or the template hairstyle, as well as the facial geometric information and the face texture image, generating the 3D digital face.

9. The 3D digital life generation method according to claim 7, wherein After generating the 3D digital face corresponding to the target object based on the facial geometric information and the face texture image, it further includes: Receiving an adjustment instruction from the user for a specified area of the 3D digital face; Based on the adjustment instruction, adjusting the specified area.

10. The 3D digital human generation method according to claim 7, wherein Before determining the face image mapping image and facial geometric information of the target object based on the at least one face image, it includes: Detecting at least one of facial expressions, abnormal content, face proportion, face position, and face integrity in the at least one face image, and detecting whether there are face images at multiple angles in the case where the at least one face image includes multiple face images; If any face image fails the detection or there are no face images at multiple angles, a prompt message is sent to the user.

11. The 3D digital human generation method according to any one of claims 7-10, characterized in that, Determining the face image mapping image and facial geometric information of the target object based on the at least one face image includes: Based on the at least one face image, determining the face image mapping image, the facial geometric information, and facial auxiliary features, where the facial auxiliary features include at least one of expression features, pose features, and lighting features; Correspondingly, generating a 3D digital face corresponding to the target object based on the facial geometric information and the face texture image includes: Based on the facial geometric information, determining a target face model, and based on the face texture image, performing texture mapping on the target face model to obtain an initial 3D face; Based on the facial auxiliary features, optimizing the initial 3D face to obtain the 3D digital face.

12. The 3D digital life generation method according to claim 11, wherein Determining the face image mapped image, the facial geometry information, and the facial auxiliary features based on the at least one face image includes: Inputting the at least one face image into a three-dimensional deformable face model to obtain the face image mapped image, the facial geometry information, and the facial auxiliary features determined by the three-dimensional deformable face model; Wherein, the three-dimensional deformable face model is trained based on face image samples at different angles.

13. The 3D digital human generation method according to any one of claims 7-10, characterized in that Generating a target 3D digital human based on the 3D digital face includes: Receiving a target torso type selected by a user; Generating the target 3D digital human based on the template torso corresponding to the target torso type and the 3D digital face; Or, Extracting the style features from the at least one face image; Generating the target 3D digital human based on the template torso matching the style features and the 3D digital face.

14. The 3D digital life generation method according to claim 13, characterized in that After generating the target 3D digital human based on the 3D digital face, it includes: Receiving a 3D digital human display space selected by a user; Displaying the target 3D digital human in the 3D digital human display space; Or, Displaying the target 3D digital human in a 3D digital human display space matching the style features.

15. The 3D digital life generation method according to any one of claims 7-10, characterized in that, After generating the target 3D digital human based on the 3D digital face, it includes: Receiving a modification instruction for the target 3D digital human from a user; Updating the target 3D digital human based on the modification instruction.

16. A method for generating a 3D digital human video, characterized in that, It includes: Obtaining a target 3D digital human, where the target 3D digital human is generated based on the 3D digital human generation method according to any one of claims 7-15; Generating a 3D digital human video based on the target 3D digital human.

17. A device for generating a face texture image, characterized in that, It includes: A first image receiving module for receiving a face image mapped image of a target object; An image denoising module for denoising a noise image based on a target diffusion model using the face image mapped image as a constraint condition to obtain a face texture image of the target object; Wherein, the target diffusion model is trained by forward diffusion and reverse diffusion based on face texture image samples using face image mapped image samples as constraint conditions.

18. A 3D digital human generation system, characterized in that, It includes: A second image receiving module for receiving at least one face image of a target object; A texture geometry image determining module for determining a face image mapped image and facial geometry information of the target object based on the at least one face image; A face texture image generating device for generating a face texture image based on the face image mapped image by applying the face texture image generating method according to any one of claims 1-6; A digital human generating module for generating a 3D digital face corresponding to the target object based on the facial geometry information and the face texture image, and generating a target 3D digital human based on the 3D digital face.

19. A 3D digital human video generation system, characterized in that, It includes: A 3D digital human generating system for obtaining a target 3D digital human, where the target 3D digital human is generated based on the 3D digital human generation method according to any one of claims 7-15; A video generation module, configured to generate a 3D digital human video based on the target 3D digital human.

20. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the face texture image generation method according to any one of claims 1-6, or the 3D digital human generation method according to any one of claims 7-15, or the 3D digital human video generation method according to claim 16.

21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the face texture image generation method according to any one of claims 1-6, or the 3D digital human generation method according to any one of claims 7-15, or the 3D digital human video generation method according to claim 16.

22. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the face texture image generation method according to any one of claims 1-6, or the 3D digital human generation method according to any one of claims 7-15, or the 3D digital human video generation method according to claim 16.

Citation Information

Cited By

  • Normal map generation method and digital human video generation method

    CN121544760A

  • Normal map generation method and digital human video generation method

    CN121544760B