Image processing method and device, storage medium and equipment

Through adversarial model training and potential space vector optimization, the problem of cumbersome and high cost of texture reconstruction process is solved, efficient and low-cost texture reconstruction is achieved, and more realistic target texture images are generated.

CN120047593APending Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311599069.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the texture reconstruction process is cumbersome and costly, and the expression ability of 3DMM is weak, making it difficult to reconstruct complex facial features.

Method used

By obtaining face image samples, constructing texture image samples, and training the preset adversarial model, a texture image similar to the texture image samples are generated. Then, the image is reconstructed using the latent space vector and face rendering parameters, and the latent space vector is adjusted by the optimization function to generate the target texture image.

Benefits of technology

It reduces the processing cost of texture reconstruction, greatly improves the efficiency of image processing, and generates more realistic target texture images, which can better reconstruct complex face features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047593A_ABST
    Figure CN120047593A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method and device, a storage medium and equipment, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic and aided driving, texture image samples are constructed through face image samples, a preset confrontation model is trained through the texture image samples, and the trained preset confrontation model is obtained; extracting face rendering parameters of the to-be-processed face image; inputting the potential spatial vector into a trained preset confrontation model, outputting a corresponding predicted texture image, and reconstructing the image according to the predicted texture image and the face rendering parameter to obtain a to-be-judged face image; constructing an optimization function corresponding to the potential space vector according to the difference between the to-be-judged face image and the to-be-processed face image, and calculating a target potential space vector enabling the difference to be minimum according to the optimization function; and obtaining a target texture image output by the target potential space vector in the trained adversarial model. And the image processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and particularly to an image processing method, apparatus, storage medium and device. Background Art

[0002] In computer vision technology, texture reconstruction (also known as texture remodeling) is often required. Specifically, texture technology is used to map texture images, etc. onto the target three-dimensional geometric surface to form the color information of the object.

[0003] In the related art, a 3D light field acquisition device can be used to achieve texture reconstruction. For example, the structured light of the 3D light field acquisition device is used to scan the human face, so that the geometric shape and texture image of the human face can be obtained, and texture reconstruction can be achieved. However, although the effect of this method is relatively good, it is limited by the requirements of the device and the site, resulting in a very cumbersome processing process and a high processing cost. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, apparatus, storage medium and device, which can improve the efficiency of image processing and save the processing cost.

[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0006] An image processing method, comprising:

[0007] Obtain a human face image sample, construct a texture image sample based on the human face image sample, and train a preset adversarial model through the texture image sample to obtain a trained preset adversarial model. The trained preset adversarial model learns the distribution of the texture image sample and is used to generate a texture image similar to the texture image sample;

[0008] Obtain a human face image to be processed, and extract the human face rendering parameters of the human face image to be processed;

[0009] Input a latent space vector into the trained preset adversarial model. The latent space vector is a randomly generated texture basis parameter, which is used to be transformed into a texture image similar to the texture image sample, output a corresponding predicted texture image, and reconstruct an image according to the predicted texture image and the human face rendering parameters to obtain a human face image to be judged;

[0010] Construct an optimization function corresponding to the latent space vector according to the difference between the human face image to be judged and the human face image to be processed, and calculate a target latent space vector that minimizes the difference according to the optimization function;

[0011] Obtain the target texture image output by the target latent space vector in the trained adversarial model.

[0012] An image processing apparatus, comprising:

[0013] A first acquisition unit, configured to acquire a face image sample, construct a texture image sample based on the face image sample, and train a preset adversarial model through the texture image sample to obtain a trained preset adversarial model. The trained preset adversarial model learns the distribution of the texture image sample and is used to generate a texture image similar to the texture image sample;

[0014] A second acquisition unit, configured to acquire a face image to be processed and extract face rendering parameters of the face image to be processed;

[0015] A reconstruction unit, configured to input a latent space vector into the trained preset adversarial model. The latent space vector is a randomly generated texture basis parameter and is used to be transformed into a texture image similar to the texture image sample, output a corresponding predicted texture image, and reconstruct an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged;

[0016] A calculation unit, configured to construct an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed, and calculate a target latent space vector that minimizes the difference according to the optimization function;

[0017] A third acquisition unit, configured to obtain the target texture image output by the target latent space vector in the trained adversarial model.

[0018] In some embodiments, the first acquisition unit includes:

[0019] An acquisition subunit, configured to acquire a face image sample;

[0020] A perspective enhancement subunit, configured to perform perspective enhancement processing on the face image sample to obtain face image samples under different perspectives;

[0021] A generation subunit, configured to extract face geometry data of the face image sample under each perspective and generate a three-dimensional face mesh under each perspective according to the face geometry data;

[0022] An inverse projection subunit, configured to inverse project the three-dimensional face mesh into the camera space for color sampling, and obtain an initial texture image corresponding to the three-dimensional face mesh under each perspective through linear interpolation;

[0023] A fusion subunit, configured to perform multi-perspective fusion on the initial texture images under different perspectives to obtain a texture image sample;

[0024] A training subunit, configured to train a preset adversarial model with the texture image samples to obtain a trained preset adversarial model.

[0025] In some embodiments, the perspective enhancement subunit is configured to:

[0026] Preprocess the face image samples to obtain intermediate face image samples, where the preprocessing includes at least one of expression normalization, illumination normalization, or hair removal processing;

[0027] Perform perspective enhancement processing on the intermediate face image samples to obtain face image samples from different perspectives.

[0028] In some embodiments, the fusion subunit includes:

[0029] An acquisition sub-module, configured to acquire predefined masks corresponding to the initial texture images from each perspective;

[0030] A blending sub-module, configured to perform alpha blending on the initial texture images from different perspectives based on the predefined masks to obtain texture image samples.

[0031] In some embodiments, the blending sub-module is configured to:

[0032] Perform alpha blending on the initial texture images from different perspectives based on the predefined masks to obtain a first intermediate texture image;

[0033] Determine the distorted areas of the first intermediate texture image, and complement the distorted areas with a preset human template to obtain a second intermediate texture image;

[0034] Determine the missing areas of the second intermediate texture image, and complement the missing areas with a preset human template to obtain texture image samples.

[0035] In some embodiments, the preset adversarial model at least includes a generator and a discriminator, and the training subunit is configured to:

[0036] Input the texture image samples into the discriminator in the preset adversarial model for training to obtain a trained discriminator;

[0037] Input a training latent space vector into the generator of the preset adversarial model to output a corresponding training texture image;

[0038] Input the training texture image into the trained discriminator to output a discrimination result;

[0039] Training the generator according to the discrimination difference between the discrimination result and the preset label until the discrimination difference converges, to obtain a trained preset adversarial model, where the trained preset adversarial model at least includes a trained generator.

[0040] In some embodiments, the reconstruction unit includes:

[0041] An input subunit, configured to input a latent space vector into the trained generator of the trained preset adversarial model, and output a corresponding predicted texture image;

[0042] A reconstruction subunit, configured to reconstruct an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged.

[0043] In some embodiments, the calculation unit includes:

[0044] A first construction subunit, configured to construct a first constraint condition according to the pixel difference between the face image to be judged and the face image to be processed;

[0045] A second construction subunit, configured to construct a second constraint condition according to the identity difference between the face image to be judged and the face image to be processed;

[0046] A generation subunit, configured to generate a third constraint condition for controlling the change amplitude of the latent space vector;

[0047] A third construction subunit, configured to construct an optimization function corresponding to the latent space vector based on at least one of the first constraint condition, the second constraint condition, and the third constraint condition;

[0048] A calculation subunit, configured to calculate a target latent space vector that minimizes the difference according to the optimization function.

[0049] In some embodiments, the calculation subunit includes:

[0050] An update sub-module, configured to update the latent space vector according to the constraints of the optimization function;

[0051] An input sub-module, configured to input the updated latent space vector into the trained preset adversarial model, output an updated predicted texture image, and reconstruct an image according to the updated predicted texture image and the face rendering parameters to obtain an updated face image to be judged;

[0052] A result sub-module, configured to obtain a target latent space vector when it is detected that the updated face image to be judged satisfies the constraints of the optimization function.

[0053] In some embodiments, the result sub-module is configured to:

[0054] When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed satisfies the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed satisfies the second constraint condition, determine the current updated potential space vector as the target potential space vector.

[0055] In some embodiments, the calculation sub-unit further includes:

[0056] The re-execution sub-module is configured to re-execute the update of the potential space vector according to the constraints of the optimization function when it is detected that the updated face image to be judged does not satisfy the constraints of the optimization function.

[0057] In some embodiments, the face rendering parameters at least include identity parameters, expression parameters, linear texture images, lighting parameters, and camera parameters. The reconstruction sub-unit is configured to:

[0058] Replace the linear texture image with the predicted texture image;

[0059] Through differentiable rendering, reconstruct an image based on the identity parameters, expression parameters, predicted texture image, lighting parameters, and camera parameters to obtain the image to be judged.

[0060] A computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the above-mentioned image processing method.

[0061] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned image processing method is implemented.

[0062] A computer program product or computer program includes computer instructions, and the computer instructions are stored in a storage medium. The processor of the computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions to implement the above-mentioned image processing method.

[0063] In an embodiment of the present application, a face image sample is obtained, a texture image sample is constructed based on the face image sample, and a preset adversarial model is trained with the texture image sample to obtain a trained preset adversarial model; a face image to be processed is obtained, and face rendering parameters of the face image to be processed are extracted; a latent space vector is input into the trained preset adversarial model to output a corresponding predicted texture image, and an image is reconstructed according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; an optimization function corresponding to the latent space vector is constructed based on the difference between the face image to be judged and the face image to be processed, and a target latent space vector that minimizes the difference is calculated according to the optimization function; a target texture image output by the trained adversarial model for the target latent space vector is obtained. In this way, a corresponding texture image sample is constructed through the face image sample, and the distribution of the texture image sample is learned through the preset adversarial model to obtain a trained preset adversarial model. Furthermore, an image is reconstructed based on the non-linear basis learned by the trained preset adversarial model to obtain a face image to be judged. An optimization function is constructed based on the difference between the face image to be judged and the face image to be processed for optimization, the optimal target latent space vector is solved, and a target texture image close to the original image is generated. Compared with the related technology that uses a 3D light field acquisition device to implement texture reconstruction, the embodiment of the present application reduces the processing cost and greatly improves the efficiency of image processing.

[0064] Other features and advantages of the present disclosure will be described in subsequent specifications, and, in part, will be obvious from the specifications, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures specifically pointed out in the specifications, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0066] Figure 1 It is a schematic diagram of the scenario of the image processing system provided by the embodiment of the present application.

[0067] Figure 2 It is an application schematic diagram of the image processing method provided by the embodiment of the present application.

[0068] Figure 3 It is a flowchart of the image processing method provided by the embodiment of the present application.

[0069] Figure 4 Scene schematic diagram of the image processing method provided by the embodiments of the present application.

[0070] Figure 5 Structural schematic diagram of the preset adversarial model provided by the embodiments of the present application.

[0071] Figure 6 Generator structure of StyleGAN provided by the embodiments of the present application.

[0072] Figure 7 Another process schematic diagram of the image processing method provided by the embodiments of the present application.

[0073] Figure 8 Another process schematic diagram of the image processing method provided by the embodiments of the present application.

[0074] Figure 9 Structural schematic diagram of the image processing device provided by the embodiments of the present application.

[0075] Figure 10 Structural schematic diagram of the terminal provided by the embodiments of the present application.

[0076] Figure 11 Structural schematic diagram of the server provided by the embodiments of the present application. Detailed implementation manners

[0077] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope protected by the present application.

[0078] It can be understood that in the specific implementation manners of the present application, data related to face images and the like are involved. When the above embodiments of the present application are applied to specific products or technologies, permission or consent of the object needs to be obtained, and the collection, use, and processing of the relevant data need to comply with relevant laws, regulations, and standards.

[0079] In addition, when the embodiments of the present application need to obtain data related to face images and the like, a separate permission or separate consent for the data related to face images and the like will be obtained by means of a pop-up window or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent for the data related to face images and the like, the necessary data related to face images and the like for enabling the embodiments of the present application to operate normally will be obtained.

[0080] It should be noted that in some processes described in the specification, claims, and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second", or "target" in this document are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0081] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0082] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0083] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0084] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition, monitoring, and measurement, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0085] Image Texture: Image texture refers to the visual texture effect presented in an image, which is composed of regular changes in pixel color and brightness.

[0086] 3DMM (Deformable 3D Model), whose model is represented as a 3D Mesh of the human face, has a mean shape and a mean texture, and its vertex number is fixed. Its core idea is that the 3D structure of the human face can be obtained by weighted addition / multiplication of many independently controlled facial features. In 3DMM, the human face is a linear space spanned by a shape vector and a texture vector, and each point in the space is a face. By randomly selecting a set of parameters, a new face can be generated.

[0087] The linear texture basis refers to the linearly compressed texture library obtained by performing PCA (Principal Component Analysis, used for data dimensionality reduction, i.e., the technology adopted by 3DMM) on the texture database (the regression is not necessarily accurate and the expression ability is weak). The non-linear texture basis is mainly a neural generation network. The non-linear texture basis has a richer expression ability.

[0088] StyleFlow, an open-source tool for 2D image editing, can continuously adjust multiple attribute dimensions of a human face while keeping the human face identity unchanged.

[0089] Laplacian Pyramid Blending is a traditional image template fusion technology that can make the overlapping area of ​​two images transition naturally and smoothly.

[0090] DECA (Detailed Expression Capture and Animation), a deep learning-based face reconstruction technology, can robustly generate facial detail parameters and general expression parameters from a low-dimensional latent representation. The core idea is to obtain detailed facial expression information through deep learning algorithms and apply it to the expression and actions of virtual characters.

[0091] At present, texture reconstruction technology plays an important role in the application of Metaverse digital human construction. The quality of texture reconstruction directly affects the production effect of 3D digital human and the later display effect. Therefore, the texture reconstruction technology of human face is an indispensable part of the 3D digital human production process.

[0092] In related technologies, 3D light field acquisition equipment can be used to achieve texture reconstruction. For example, the face can be scanned by the structured light of the 3D light field acquisition equipment to obtain the geometric shape and texture image of the face and achieve texture reconstruction. It is also possible to optimize the linear texture base by constructing 3DMM and solve the corresponding texture coefficients. The texture coefficients act on the linear texture base to obtain the corresponding texture image and achieve texture reconstruction.

[0093] However, although the effect of texture reconstruction using 3D light field acquisition equipment is relatively good, the processing process is very cumbersome and the processing cost is high due to the requirements of equipment and venue. Since 3DMM is a linear basis, the final face reconstruction result is obtained by linear combination of predicted parameters, which is limited by the expression ability of the linear basis. For example, for some facial features that are difficult to reconstruct or have strong nonlinearity, such as deep nasolabial folds and wrinkles, it is difficult for 3DMM to reconstruct them. Therefore, the expression ability of 3DMM is weak, and the reconstruction result produced has a strong CG (Computer Graphics, the English abbreviation of computer graphics, also called digital graphics) feeling.

[0094] In order to solve the above problems, the embodiments of the present application propose a texture image processing method that can achieve both high-quality texture reconstruction and low cost, which can improve the efficiency of image processing and reduce costs. Please refer to the following specific embodiments for details.

[0095] See also Figure 1 , Figure 1 1 is a schematic diagram of a scene of an image processing system provided in an embodiment of the present application, which includes a terminal 140, the Internet 130, a gateway 120, a server 110, and the like.

[0096] The terminal 140 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Additionally, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0097] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with ordinary terminals 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) partitioned from a high-performance computer, a combination of parts (such as virtual machines) partitioned from multiple high-performance computers, etc.

[0098] The gateway 120 is also called an inter-network connector and protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion function. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 needs to be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 also needs to be sent to the corresponding terminal 140 through the gateway 120.

[0099] The image processing method of the embodiments of the present disclosure can be fully implemented on the terminal 140, fully implemented on the server 110, or partially implemented on the terminal 140 and partially implemented on the server 110.

[0100] When the image processing method is fully implemented on the terminal 140, a face image sample is obtained on the terminal 140, a texture image sample is constructed based on the face image sample, and a preset adversarial model is trained through the texture image sample to obtain a trained preset adversarial model; a face image to be processed is obtained, and the face rendering parameters of the face image to be processed are extracted; a latent space vector is input into the trained preset adversarial model, and a corresponding predicted texture image is output, and an image is reconstructed according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; an optimization function corresponding to the latent space vector is constructed according to the difference between the face image to be judged and the face image to be processed, and a target latent space vector that minimizes the difference is calculated according to the optimization function; the target texture image output by the target latent space vector in the trained adversarial model is obtained.

[0101] When the image processing method is fully implemented on the server 110, a face image sample is obtained on the server 110, a texture image sample is constructed based on the face image sample, and the preset adversarial model is trained through the texture image sample to obtain the trained preset adversarial model; a face image to be processed is obtained, and the face rendering parameters of the face image to be processed are extracted; the latent space vector is input into the trained preset adversarial model, a corresponding predicted texture image is output, and an image is reconstructed according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; an optimization function corresponding to the latent space vector is constructed according to the difference between the face image to be judged and the face image to be processed, and the target latent space vector that minimizes the difference is calculated according to the optimization function; the target texture image output by the target latent space vector in the trained adversarial model is obtained.

[0102] When a part of the image processing method is implemented on the terminal 140 and the other part is implemented on the server 110, generally, model training is implemented on the server 110, and each terminal 140 provides face image samples for training. Each terminal 140 sends the collected face image samples to the server 110, and the server 110 trains the preset adversarial model on this basis.

[0103] Embodiments of the present disclosure can be applied in various scenarios, such as Figure 2 the scenario of the image processing system shown.

[0104] The scenario of the image processing system:

[0105] An image processing system refers to a system that can automatically obtain a corresponding target texture image according to a face image to be processed provided by an object, and it can identify an accurate target texture image from the face image to be processed through computer vision technology.

[0106] It is possible to obtain the face image 11 to be processed input by the object, extract the face rendering parameters of the face image 11 to be processed, input the latent space vector into the trained preset adversarial model, output a corresponding predicted texture image, and reconstruct an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged. An optimization function corresponding to the latent space vector is constructed according to the difference between the face image to be judged and the face image to be processed, and the target latent space vector that minimizes the difference is calculated according to the optimization function; the target texture image 12 output by the target latent space vector in the trained adversarial model is obtained, thereby realizing the rapid reconstruction of the texture image.

[0107] It should be noted that Figure 1The schematic diagram of the scenario of the image processing system shown is only an example. The image processing system and the scenario described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of image processing and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0108] In this embodiment, the description will be made from the perspective of the image processing device. The image processing device may be specifically integrated in a computer device with a storage unit and a microprocessor installed and having computing capabilities. The computer device may be a terminal or a server. In this embodiment, the computer device is described as a server for illustration.

[0109] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the image processing method provided by the embodiments of the present application. The image processing method includes:

[0110] In step 201, a face image sample is obtained, a texture image sample is constructed based on the face image sample, and the preset adversarial model is trained through the texture image sample to obtain the trained preset adversarial model.

[0111] Among them, the face image sample refers to the basic image sample for subsequent model training. The face image sample is a two-dimensional image, and each face image sample contains the corresponding face. Here, the face image sample can be single or multiple. For example, each face image in a face image set can serve as a face image sample. For the case of multiple face image samples, each face image sample can be sequentially trained. The more face image samples there are, the more times the training is performed. During each round of execution, the preset adversarial model is slightly adjusted. Therefore, the more face image samples there are, the finer the adjustment of the preset adversarial model and the better the training effect.

[0112] Correspondingly, since the texture image needs to be generated subsequently, a corresponding texture image sample also needs to be constructed based on the face image sample, that is, the texture corresponding to the face in the face image sample is extracted as the texture image sample.

[0113] In some embodiments, the extraction of the texture image sample can be an unsupervised extraction method. That is, the construction of the texture image sample based on the face image sample may include:

[0114] (1) Perform perspective enhancement processing on the face image sample to obtain face image samples from different perspectives;

[0115] (2) Extract the facial geometry data of the facial image samples from each perspective, and generate a 3D facial mesh for each perspective based on the facial geometry data;

[0116] (3) Back-project the 3D facial mesh into the camera space for color sampling, and obtain the initial texture image corresponding to the 3D facial mesh for each perspective through linear interpolation;

[0117] (4) Perform multi-view fusion on the initial texture images from different perspectives to obtain the texture image samples.

[0118] Among them, in step (1), the StyleFlow can be used to perform perspective enhancement processing on the facial image samples. The concept of perspective enhancement processing is to generate the invisible parts of the face in the facial image samples, which is convenient for subsequent full-face texture reconstruction. That is, by performing perspective enhancement processing on the facial image samples, facial image samples from the left perspective, the middle perspective, and the right perspective can be obtained.

[0119] In some embodiments, in order to improve the accuracy of subsequent back-projection, the facial image samples can be pre-optimized before performing perspective enhancement processing. That is, by performing perspective enhancement processing on the facial image samples to obtain facial image samples from different perspectives, it may include:

[0120] (1.1) Perform preprocessing on the facial image samples to obtain intermediate facial image samples, where the preprocessing includes at least one of expression normalization, illumination normalization, or hair removal;

[0121] (1.2) Perform perspective enhancement processing on the intermediate facial image samples to obtain facial image samples from different perspectives.

[0122] Among them, the StyleFlow can be used to perform expression normalization on the facial image samples, and edit the facial expression of the facial image samples into a neutral expression. For example, the expression parameter of the facial image samples can be set to 0 through the StyleFlow. Correspondingly, the StyleFlow can also be used to perform illumination normalization on the facial image samples, and set the illumination of the facial image samples to be irradiated by a unified light source. For example, the vector of the 9D illumination parameters of the facial image samples can be set to (1, 0, 0,..., 0) through the StyleFlow. Finally, the StyleFlow can also be used to remove the hair on the face of the facial image samples. For example, the hair attribute can be set to 0 through the StyleFlow. In this way, the intermediate facial image samples can be obtained through at least one of the above processing methods.

[0123] Furthermore, the intermediate face image samples can be processed by StyleFlow for perspective enhancement to obtain face image samples from the left perspective, middle perspective, and right perspective, which can be determined by the yaw attribute. For example, the yaw is set to -90 for the left perspective, 0 for the middle perspective, and 90 for the right perspective. Due to the previous optimizations of expression normalization, illumination normalization, and hair removal, the face image samples under different perspectives are standardized, which can improve the accuracy of subsequent back-projection. For example, please refer to Figure 4 As shown, through StyleFlow, the face image samples are subjected to expression normalization, illumination normalization, and hair removal to obtain intermediate face image samples, and the intermediate face image samples are processed for perspective enhancement to obtain face image samples from the left perspective, middle perspective, and right perspective. Moreover, the expressions and illuminations of the face images under each perspective are unified and there is no hair.

[0124] In step (2), the face geometric data includes the shape parameters and expression parameters of the face. The face geometric data of the face image samples from each perspective can be extracted by DECA. Based on this, the corresponding 3D face meshes for each perspective can be generated by substituting the face geometric data into the formula of 3DMM. For example, the formula of 3DMM is as follows:

[0125]

[0126] Among them, the is the average value of the texture, the α id is the covariance matrix of the shape parameters, the α exp is the covariance matrix of the expression parameters, the A id is the shape parameter, the A exp is the expression parameter. Since the α id and α exp are known parameters, therefore, only the shape parameter A id and the expression parameter α exp need to be substituted to form the 3D face mesh formed by mapping the face image samples onto the 3D face model. It should be noted that the 3D face mesh is composed of many triangular mesh patches. Therefore, for any pixel of the texture map corresponding to the face image sample, it corresponds to the corresponding triangular mesh patch.

[0127] In step (3), in order to obtain the texture image, the three-dimensional face meshes of each view can be respectively back-projected into the camera space (x, y) for color sampling. That is to say, it can be understood that the three-dimensional face meshes are spread out into a two-dimensional space, and the three-dimensional coordinates of the three-dimensional face meshes are mapped in the two-dimensional space. Based on this, for each pixel of the texture map formed on the three-dimensional face model corresponding to the face image sample, the triangle mesh patch index of the back-projection mapping can be found. Since the RGB (red, green, and blue) values of the three vertices of the triangle mesh patch are known, therefore, the RGB value of each pixel of the texture map formed by the face image sample can be obtained through linear interpolation. Furthermore, the RGB values of each pixel are combined to obtain the initial texture image of the texture map formed by the face image sample. By analogy, the initial texture images corresponding to the three-dimensional face meshes of each view are obtained. For example, please continue to refer to Figure 4 As shown in Figure 4 , the facial geometry data of the face image samples from different views are extracted, and the three-dimensional face meshes of different views are generated according to the facial geometry data. The three-dimensional face meshes are back-projected into the camera space for color sampling, and the initial texture images corresponding to the three-dimensional face meshes of each view are obtained through linear interpolation.

[0128] In step (4), since there are problems such as projection deviation and incomplete projection area in the results of a single view, therefore, the initial texture images of different views can be fused from multiple views to obtain an accurate texture image sample. This multi-view fusion can be performed by using the Alpha-Blending method. The principle is to blend two overlapping graphics to achieve different degrees of transparency visually.

[0129] In some embodiments, for each facial part with high reconstruction accuracy in each view, generally there is an empirical value, which can be understood as a mask with high credibility, or it can be understood as a mask (alphamask) of a smoothing function, used to implement alpha blending. Thus, this step (4) can include:

[0130] (4.1) Obtain the predefined mask corresponding to the initial texture image of each view;

[0131] (4.2) Perform alpha blending on the initial texture images of different views based on the predefined mask to obtain the texture image sample.

[0132] Among them, the predefined masks corresponding to the initial texture images in the left view, middle view, and right view can be obtained respectively. The predefined mask is a mask with high credibility and is used to implement subsequent alpha blending. Thus, based on the predefined mask, the initial texture images in each view can be masked respectively, and the initial texture images in each view after masking are fused to obtain a texture image sample. Specifically, the following formula can be referred to for understanding:

[0133] T = M l T l +M f T f +M r T r

[0134] Among them, the M l 、M f and M r are the predefined masks of the initial texture image in the left view, the predefined mask of the initial texture image in the middle view, and the predefined mask of the initial texture image in the right view respectively. The T l 、T f and T r are the initial texture image in the left view, the initial texture image in the middle view, and the initial texture image in the right view respectively. Based on the above formula, the initial texture images in different views are alpha-blended through the predefined mask to obtain a texture image sample T. The texture image sample T is fused through the initial texture images of multiple views, which can solve the problem of poor texture image quality caused by projection deviation and incomplete projection area.

[0135] In some embodiments, since the initial texture images of different views are fused, artifacts and missing regions will appear. Therefore, it can also be supplemented by a general human template (Human). The general human template is a general three-dimensional human template and can be used to supplement some missing parts. That is, the above step (4.2) can include:

[0136] (4.2.1) Based on the predefined mask, the initial texture images in different views are alpha-blended to obtain a first intermediate texture image;

[0137] (4.2.2) Determine the distorted region of the first intermediate texture image, and complete the distorted region through a preset human template to obtain a second intermediate texture image;

[0138] (4.2.3) Determine the missing region of the second intermediate texture image, and complete the missing region through a preset human template to obtain a texture image sample.

[0139] Among them, based on the predefined masks corresponding to the initial texture images in the left view, middle view, and right view, alpha blending can be performed on the initial texture images in the left view, middle view, and right view to obtain a first intermediate texture image. For example, please continue to refer to Figure 4 As shown, predefined masks for each view can be obtained, and based on the predefined masks for each view, alpha blending is performed on the initial texture images from different views to obtain a first intermediate texture image, which has more comprehensive and rich details.

[0140] Since distortion and missing regions will occur when fusing the initial texture images from different views, it is also necessary to complete the distorted regions of the first intermediate texture image. Specifically, through distortion detection, the distorted regions are determined, and the preset human template is used to complete the distorted regions to form a smooth second intermediate texture image, making the fusion of the second intermediate texture image smoother and more complete. For example, please continue to refer to Figure 4 As shown, the distorted regions of the first intermediate texture image can be detected by means of detecting the distorted regions, and the preset human template is used to complete the distorted regions to obtain the second intermediate texture image.

[0141] Furthermore, even when fusing the initial texture images from different views, missing regions that cannot be covered by multiple views will appear, such as missing regions of the head, chin, and neck. Therefore, the missing regions can also be completed according to the preset human template to finally obtain a complete texture image sample. For example, please continue to refer to Figure 4 As shown, the missing regions of the second intermediate texture image can be determined. For example, the missing regions are the head region and the neck region, and the preset human template is used to complete the missing regions to obtain a complete texture image sample.

[0142] In the embodiments of the present application, the texture image sample undergoes alpha blending, artifact correction (i.e., completion of distorted regions), and completion of missing regions of the initial texture images from different views, and both the integrity and details of the texture are greatly improved, which can improve the accuracy of subsequent preset adversarial training.

[0143] It should be noted that the preset adversarial model can be a GAN (Generative Adversarial Networks) model or a StyleGAN (Style Adversarial Generative Network). The principle is to use a set of real images as input and attempt to generate similar images that can pass for real. In the embodiments of the present application, texture image samples are used as input to attempt to generate similar texture images that can pass for real, so as to achieve texture reconstruction. Thus, the preset adversarial model can be trained with the texture image samples to obtain a trained preset adversarial model. This trained preset adversarial model has learned the distribution of the texture image samples and has the ability to generate texture images similar to the texture image samples. Therefore, this trained preset adversarial model can be used to generate texture images similar to the texture image samples. Moreover, since the texture images generated by the trained preset adversarial model are non-linear textures, that is, a non-linear texture base is built, compared with the texture images generated by the linear texture base, the expressive ability is richer and more vivid.

[0144] To better illustrate the embodiments of the present application, the following describes the specific training process of the preset adversarial model, that is, training the preset adversarial model with the texture image samples to obtain a trained preset adversarial model, including:

[0145] (A) Input the texture image samples into the discriminator in the preset adversarial model for training to obtain a trained discriminator;

[0146] (B) Input the training latent space vectors into the generator of the preset adversarial model to output corresponding training texture images;

[0147] (C) Input the training texture images into the trained discriminator to output discrimination results;

[0148] (D) Train the generator according to the discrimination difference between the discrimination results and the preset labels until the discrimination difference converges to obtain a trained preset adversarial model.

[0149] Among them, for a better understanding of the training process, please refer to Figure 5 for understanding. The preset adversarial model at least includes a generator (abbreviated as G, also known as the generative network) and a discriminator (abbreviated as D, also known as the discriminative network). The generator is used to generate new data, and the basis for generating data is often a set of noises or random numbers, such as a noise vector with a high-dimensional normal distribution. The discriminator is used to judge which of the generated data and the real data is real. The generator has no label and is an unsupervised network, while the discriminator has a label and is a supervised network, and its label is "false or true" (0 and 1).

[0150] In the embodiments of the present application, the texture image sample can be input into the discriminator in a pre-trained adversarial model for training. That is, the discriminator first learns how to identify texture images. The label of the texture image sample is 1, which is a "true" label. Based on this, the discriminator will output a corresponding prediction result according to the texture image sample, and compare the prediction result with the label 1 to obtain the training discrimination difference, and adjust the model parameters of the generator according to the training discrimination difference. Iteratively repeat the above process until the training discrimination difference converges to obtain the trained discriminator. The trained discriminator has the ability to identify texture images as "false or true".

[0151] The training latent space vector is the basis for the generator to generate data. It can be a noise vector with a high-dimensional normal distribution. Input the training latent space vector into the generator of the preset adversarial model, perform an affine transformation on the training latent space vector, map it to a higher dimension, and then rearrange it into a rectangle. The arrangement of the rectangle is closer to the texture image. Then, the generator processes the rectangle through convolution, pooling, and activation functions to obtain a noise matrix with the same size as the texture image sample, that is, the corresponding training texture image is obtained.

[0152] Further, in order to generate a texture image similar to the texture image sample, the generator needs to be trained. Therefore, the training texture image can be input into the trained discriminator to output a corresponding discrimination result. The training purpose is to expect the trained discriminator to determine that the discrimination result of this training texture image is true. Therefore, the discrimination result can be compared with a preset label (which is 1, representing "true") to obtain the discrimination difference, and the model parameters of the generator can be adjusted according to the discrimination difference. Iteratively repeat the above process until the discrimination difference converges to obtain the trained generator, that is, the trained preset adversarial model. The trained preset adversarial model includes the trained generator, and the trained generator has learned the distribution of the texture image sample and has the ability to generate a texture image similar to the texture image sample. Moreover, since the texture image generated by the trained generator is a non-linear texture, that is, a non-linear texture basis is built, compared with the texture image generated by the linear texture basis, the expression ability is richer and more vivid.

[0153] In some embodiments, when the preset adversarial model is StyleGAN, its training principle is the same, but its generator is different from that of GAN. Please refer to Figure 6 as shown Figure 6It is the structural diagram of the generator of StyleGAN. The generator of this StyleGAN consists of a mapping network and a generation network. The mapping network always consists of 8 fully connected layers. The input is a 512-dimensional noise vector. After passing through 8 fully connected layers, the 512-dimensional training latent space vector in the embodiments of the present application is obtained. That is, different from GAN, this training latent space vector is obtained through the processing of 8 fully connected layers. The advantage of this is to get rid of the influence of the input vector on the input data distribution, so that there is a better linear relationship between the training latent space vector and the subsequently generated training texture image, which is conducive to the attribute control of the training texture image. The generation network is a structure with gradually increasing resolution, consisting of 9 stylized modules. By inputting the training latent space vector into this generation network, the training texture image can be obtained. The specific training process is the same as that of GAN. Please refer to the training method of the above GAN, and no specific description will be given here.

[0154] In step 202, the face image to be processed is obtained, and the face rendering parameters of the face image to be processed are extracted.

[0155] Among them, the face image to be processed is the given two-dimensional image of the object. The object needs to perform texture reconstruction based on this face image to be processed to realize the production of a 3D digital human. Therefore, after obtaining the face image to be processed, the face rendering parameters of the face image to be processed can be quickly extracted through DECA. The face rendering parameters at least include identity parameters, expression parameters, linear texture image, lighting parameters, and camera parameters. The identity parameters are used to control the identity of the face and can be composed of a 100-dimensional vector. The expression parameters are used to control the expression of the face and can be composed of a 40-dimensional vector. The linear texture image can be composed of linear texture parameters and is used to control the linear texture of the face and can be composed of a 100-dimensional vector. The lighting parameters are used to control the lighting of the face and can be composed of a 9-dimensional vector. The camera parameters are used to control the imaging angle of the face and can be composed of camera parameters.

[0156] In step 203, the latent space vector is input into the trained preset adversarial model, and the corresponding predicted texture image is output. Then, an image is reconstructed based on the predicted texture image and the face rendering parameters to obtain the face image to be judged.

[0157] Among them, since the trained preset adversarial model has learned the distribution of texture image samples and has the ability to generate texture images similar to the texture image samples, a latent space vector can be randomly generated. This latent space vector is a randomly generated texture basis parameter and can be a noise vector of a random high-dimensional normal distribution, which is used to transform into a texture image similar to the texture image samples. The trained preset adversarial model is the non-linear texture basis for generating texture images with non-linear expressions. The key point for adjusting the style of the texture image with non-linear expressions is the above-mentioned texture basis parameter. The latent space vector is input into the trained preset adversarial model, and a corresponding predicted texture image is randomly output. The predicted texture image is similar to the texture image samples used for training before. However, since the latent space vector is random, the predicted texture image is not the corresponding texture image of the face image to be processed.

[0158] Therefore, it is necessary to adjust the latent space vector to make it approximate the texture image of the face image to be processed. In the embodiment of the present application, a reconstructed image can be obtained by combining the predicted texture image and the face rendering parameters. The specific reconstruction method can be a differentiable rendering method. The differentiable rendering method is used to solve the inverse rendering problem. Inverse rendering is a process opposite to rendering. This process obtains the geometric representation, material, scene, lighting, and camera parameters of an object from an image to express the image in reverse and seek optimization. That is, the face rendering parameters extracted before and the predicted texture image can be used to express the image in reverse to generate the face image to be judged. Since the face rendering parameters are fixed, the predicted texture image determines the representation of the generated face image to be judged.

[0159] In some embodiments, inputting the latent space vector into the trained preset adversarial model to output a corresponding predicted texture image includes: inputting the latent space vector into the trained generator of the trained preset adversarial model to output a corresponding predicted texture image.

[0160] Among them, since the trained generator in the trained preset adversarial model has learned the distribution of texture image samples and has the ability to generate texture images similar to the texture image samples, a latent space vector can be randomly generated, and the latent space vector is input into the trained generator of the trained preset adversarial model, and a corresponding predicted texture image is randomly output. The predicted texture image is similar to the texture image samples used for training before.

[0161] In some embodiments, reconstructing an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged includes:

[0162] Replacing the linear texture image with the predicted texture image;

[0163] Through differentiable rendering, an image is reconstructed based on identity parameters, expression parameters, predicted texture images, lighting parameters, and camera parameters to obtain an image to be judged.

[0164] Among them, the face rendering parameters at least include identity parameters, expression parameters, linear texture images, lighting parameters, and camera parameters. Since the linear texture image is linear and limited by its expressive ability, it cannot represent strong face features, such as deep nasolabial folds, wrinkles, etc. Therefore, in the embodiments of the present application, the predicted texture image can be used to replace the linear texture image, that is, the linear texture image is deleted, and an image is reconstructed through differentiable rendering using identity parameters, expression parameters, predicted texture images, lighting parameters, and camera parameters to obtain an image to be judged. The following formula can be referred to for understanding:

[0165] I = render(p id , p exp , G tex (z), p cam , p light )

[0166] Among them, the render() represents a differentiable rendering function, the p id is the identity parameter, the, p exp is the expression parameter, G tex (z) is the predicted texture image (i.e., the non-linear texture basis), p cam is the lighting parameter, p light is the camera parameter. Thus, through the above formula, the image to be judged I is obtained.

[0167] In step 204, an optimization function corresponding to the latent space vector is constructed based on the difference between the face image to be judged and the face image to be processed, and the target latent space vector that minimizes the difference is calculated according to the optimization function.

[0168] Among them, the difference can be the pixel distribution difference. Since the face rendering parameters are fixed, the predicted texture image determines the representation of the generated face image to be judged. That is, the greater the difference between the face image to be judged and the face image to be processed, the greater the difference between the predicted texture image and the texture image of the face image to be processed, and the less accurate the predicted texture image. Correspondingly, the smaller the difference between the face image to be judged and the face image to be processed, until convergence, the closer the predicted texture image is to the texture image of the face image to be processed, and the more accurate the predicted texture image.

[0169] It should be noted that the predicted texture image is generated by a preset adversarial model after training using a latent space vector. Therefore, the predicted texture image can be modified by adjusting the latent space vector, and further, the representation of the face image to be judged can be adjusted. Thus, in the embodiments of the present application, an optimization function for the latent space vector can be constructed according to the difference between the face image to be judged and the face image to be processed. It is expected that by continuously adjusting the latent space vector, the difference between the face image to be judged reconstructed from the predicted texture image output by the latent space vector and the face image to be processed becomes smaller and smaller, achieving difference convergence and realizing the output of an accurate predicted texture image.

[0170] Therefore, the latent space vector can be continuously optimized and updated according to the above optimization function, and finally, a target latent space vector that makes the above difference converge (i.e., the difference is minimized) is obtained. The predicted texture image output by inputting the target latent space vector into the preset adversarial model after training is the most accurate.

[0171] In some embodiments, constructing an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed may include:

[0172] Constructing a first constraint condition according to the pixel difference between the face image to be judged and the face image to be processed;

[0173] Constructing a second constraint condition according to the identity difference between the face image to be judged and the face image to be processed;

[0174] Generating a third constraint condition for controlling the change amplitude of the latent space vector;

[0175] Based on at least one of the first constraint condition, the second constraint condition, and the third constraint condition, constructing an optimization function corresponding to the latent space vector.

[0176] Wherein, the pixel difference refers to the sum of the absolute values of the differences of each pixel between the face image to be judged and the face image to be processed. Thus, a first constraint condition can be constructed according to the pixel difference between the face image to be judged and the face image to be processed, and the constraint of the first constraint condition is to make the pixel difference converge as much as possible.

[0177] It should be noted that each face image carries an identity parameter, which is used to identify the identity of the face and can be composed of a 100-dimensional vector. The identity difference is the absolute value of the difference between the identity parameter extracted from the face image to be judged and the identity parameter extracted from the face image to be processed. Thus, a second constraint condition can be constructed according to the identity difference between the face image to be judged and the face image to be processed, and the constraint of the second constraint condition is to make the identity difference converge as much as possible.

[0178] Correspondingly, since the optimization process is a multi-round iterative process, in order to avoid the change amplitude of the potential space vector being too large, resulting in too large a change in the corresponding predicted texture image and affecting the convergence of the optimization, it is necessary to control the amplitude of the change of the potential space vector before and after, that is, to generate a third constraint condition for controlling the change amplitude of the potential space vector. The constraint of the third constraint condition is to make the change amplitude of the potential space vector as small as possible.

[0179] Furthermore, based on at least one of the first constraint condition, the second constraint condition, and the third constraint condition, an optimization function corresponding to the potential space vector can be constructed, and based on this optimization function, the optimal solution corresponding to the potential space vector can be solved. For a better understanding of the embodiments of the present application, please refer to the following formula together:

[0180] L = λ pix L pix + λ id L id + λ reg L reg

[0181] Among them, this L pix is the first constraint condition, L id is the second constraint condition, L reg is the third constraint condition, this γ pix is the first weight of the first constraint condition, γ id is the second weight of the second constraint condition, λ reg is the third weight of the third constraint condition. The magnitudes of the first weight, the second weight, and the third weight are all preset settings. For example, the first weight is 0.3, the second weight is 0.4, and the third weight is 0.3, which are used to adjust the importance of each constraint in the optimization process. Based on this, the optimization function L corresponding to the potential space vector is constructed.

[0182] In some embodiments, calculating the target potential space vector that minimizes the difference according to the optimization function may include:

[0183] Updating the potential space vector according to the constraints of the optimization function;

[0184] Inputting the updated potential space vector into the trained preset adversarial model, outputting the updated predicted texture image, and reconstructing an image according to the updated predicted texture image and the face rendering parameters to obtain the updated face image to be judged;

[0185] When it is detected that the updated face image to be judged does not satisfy the constraints of the optimization function, re-execute updating the potential space vector according to the constraints of the optimization function;

[0186] When it is detected that the updated face image to be judged meets the constraints of the optimization function, a target latent space vector is obtained.

[0187] Among them, according to the constraints of the optimization function (that is, it is expected to continuously adjust the latent space vector so that the difference between the face image to be judged reconstructed from the predicted texture image output by the latent space vector and the face image to be processed becomes smaller and smaller until it converges), the latent space vector can be updated.

[0188] Correspondingly, in order to verify the progress of the update, it is necessary to input the updated latent space vector into the trained generator of the trained preset adversarial model again to generate an updated predicted texture image, and reconstruct the image again through differentiable rendering according to the updated predicted texture image and face rendering parameters to obtain an updated face image to be judged.

[0189] Furthermore, it is possible to detect whether the difference between the updated face image to be judged and the face image to be processed converges to determine whether the updated face image to be judged meets the constraints of the optimization function. When it is detected that the difference between the updated face image to be judged and the face image to be processed converges, it means that the predicted texture image generated by the updated latent space vector is already very close to the texture image of the face image to be processed, and it is determined that the updated face image to be judged meets the constraints of the optimization function, and the current updated latent space vector is determined as the target latent space vector.

[0190] When it is detected that the difference between the updated face image to be judged and the face image to be processed does not converge, it means that there is still a large difference between the predicted texture image generated by the updated latent space vector and the texture image of the face image to be processed. It is determined that the updated face image to be judged does not meet the constraints of the optimization function, and return to execute the update of the latent space vector according to the constraints of the optimization function until the updated face image to be judged meets the constraints of the optimization function.

[0191] In some embodiments, when it is detected that the updated face image to be judged meets the constraints of the optimization function, obtaining the target latent space vector may include:

[0192] When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed meets the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition, the current updated latent space vector is determined as the target latent space vector.

[0193] Among them, under the constraints of the optimization function including the first constraint condition, the second constraint condition, and the third constraint condition, the above update of the latent space vector according to the constraints of the optimization function can be to update the latent space vector according to the third constraint condition of the optimization function to control the amplitude of the front and back changes of the latent space vector.

[0194] Correspondingly, in the embodiment of the present application, after obtaining the updated face image to be judged, it is necessary to synchronously detect whether the pixel difference between the updated face image to be judged and the face image to be processed meets the first constraint condition, and whether the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition.

[0195] When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed meets the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition, it indicates that the updated predicted texture image generated according to the current updated latent space vector, and the updated face image to be judged obtained by reconstructing the image are close in pixels to the face image to be processed, and the identity is also the same object, indicating that the current updated latent space vector is accurate. Therefore, the current updated latent space vector is determined as the target latent space vector.

[0196] When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed does not meet the first constraint condition, or the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition, it indicates that there is still a large difference between the predicted texture image generated by the updated latent space vector and the texture image of the face image to be processed. Return to execute the update of the latent space vector according to the constraints of the optimization function until the pixel difference between the updated face image to be judged and the face image to be processed meets the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition.

[0197] Therefore, on the basis of restricting the change amplitude of the latent space vector by the third constraint condition in the embodiment of the present application, the latent space vector is continuously optimized based on the first constraint condition of pixel difference and the second constraint condition of identity difference, so as to continuously optimize the predicted texture image, make the similarity between the predicted texture image of texture reconstruction and the image to be processed closer, and improve the accuracy of texture reconstruction.

[0198] In step 205, obtain the target texture image output by the target latent space vector in the trained adversarial model.

[0199] Among them, since the target potential space vector is input into the preset adversarial model after training, the output target texture image and the face rendering parameters are used to reconstruct the image through differentiable rendering, and the difference between the face image to be judged and the face image to be processed converges, indicating that the two face images are very close, and also indicating that the target texture image is very close to the texture image of the face image to be processed. Therefore, the target texture image output by the target potential space vector in the trained adversarial model can be directly obtained as the texture image reconstructed according to the corresponding texture of the image to be processed. Since the target texture image is a non-linear texture image, the expression ability of the target texture image is richer and the effect is more realistic.

[0200] Based on this, a corresponding realistic 3D digital human can be automatically produced according to the target texture image, in combination with 3D reconstruction technology and hair generation technology.

[0201] As can be seen from the above, in the embodiment of the present application, a face image sample is obtained, a texture image sample is constructed based on the face image sample, and the preset adversarial model is trained through the texture image sample to obtain the trained preset adversarial model; the face image to be processed is obtained, and the face rendering parameters of the face image to be processed are extracted; the potential space vector is input into the trained preset adversarial model, the corresponding predicted texture image is output, and the image is reconstructed according to the predicted texture image and the face rendering parameters to obtain the face image to be judged; an optimization function corresponding to the potential space vector is constructed according to the difference between the face image to be judged and the face image to be processed, and the target potential space vector that minimizes the difference is calculated according to the optimization function; the target texture image output by the target potential space vector in the trained adversarial model is obtained. In this way, a corresponding texture image sample is constructed through the face image sample, and the distribution of the texture image sample is learned through the preset adversarial model to obtain the trained preset adversarial model. Furthermore, the image is reconstructed based on the non-linear basis learned by the trained preset adversarial model to obtain the face image to be judged. An optimization function is constructed according to the difference between the image to be judged and the image to be processed for optimization, and the optimal target potential space vector is solved to generate a target texture image close to the original image. Compared with the related technology that uses a 3D light field acquisition device to achieve texture reconstruction, the embodiment of the present application reduces the processing cost and greatly improves the efficiency of image processing.

[0202] Combined with the method described in the above embodiments, the following will give further detailed examples.

[0203] In this embodiment, it will be described by taking the specific integration of the image processing device in the server as an example.

[0204] To better illustrate the embodiments of the present application, please refer to Figure 7 , Figure 7Another flowchart of the image processing method provided by the embodiments of this application. It includes:

[0205] In step 301, the server obtains a face image sample, performs expression normalization, illumination normalization, and hair removal on the face image sample to obtain an intermediate face image sample, and performs perspective enhancement processing on the intermediate face image sample to obtain face image samples from different perspectives.

[0206] Among them, the face image sample refers to the basic image sample for subsequent model training. The face image sample is a two-dimensional image, and each face image sample contains a corresponding face. Here, the face image sample can be single or multiple. For the case of multiple face image samples, training can be performed on each face image sample in turn. For each face image sample, the training is performed once. In the process of each round of execution, the preset adversarial model is slightly adjusted. Therefore, the more face image samples there are, the more refined the adjustment of the preset adversarial model is, and the better the training effect is.

[0207] The face image sample can be expression-normalized through StyleFlow, and the facial expression of the face image sample can be edited into a non-expression. For example, the expression parameter of the face image sample is set to 0 through StyleFlow. Correspondingly, the face image sample can also be illumination-normalized through StyleFlow, and the illumination of the face image sample is set to be irradiated by a unified light source. For example, the vector of the 9-dimensional illumination parameter of the face image sample is set to (1, 0, 0,..., 0) through StyleFlow. Finally, the hair on the face of the face image sample can also be removed through StyleFlow. For example, the hair attribute is set to 0 through StyleFlow. In this way, through the above processing methods, an intermediate face image sample is obtained.

[0208] Furthermore, the intermediate face image sample can be perspective-enhanced through StyleFlow to obtain face image samples from the left perspective, the middle perspective, and the right perspective, which can be determined by the yaw attribute. For example, the left perspective is to set the yaw of the intermediate face image sample to -90, the middle perspective is to set the yaw of the intermediate image sample to 0, and the right perspective is to set the yaw of the intermediate face image sample to 90. Due to the optimization of the previous expression normalization, illumination normalization, and hair removal processing, the face image samples from different perspectives obtained are standardized, which can improve the accuracy of subsequent back-projection. For example, please continue to refer to Figure 4As shown, through StyleFlow, the expressions of the face image samples are normalized, the lighting is normalized, and the hair is removed to obtain the intermediate face image samples, and the intermediate face image samples are subjected to perspective enhancement processing to obtain the face image samples from the left perspective, the middle perspective, and the right perspective. Moreover, the expressions and lighting of the face images under each perspective are unified and there is no hair.

[0209] In step 302, the server extracts the face geometry data of the face image samples from each perspective and generates a three-dimensional face mesh for each perspective based on the face geometry data.

[0210] Among them, the face geometry data includes the shape parameters and expression parameters of the face. The face geometry data of the face image samples from each perspective can be extracted through DECA. Thus, according to the face geometry data, substituting it into the formula of 3DMM, the corresponding three-dimensional face mesh for each perspective can be generated. For example, the formula of 3DMM is as follows:

[0211]

[0212] Among them, the is the average value of the texture, the α id is the covariance matrix of the shape parameters, the α exp is the covariance matrix of the expression parameters, the Λ id is the shape parameter, the A exp is the expression parameter. Since the α id and α exp are known parameters, therefore, only the shape parameter A id and the expression parameter α exp need to be substituted, and then a three-dimensional face mesh formed by mapping the face image samples onto the three-dimensional face model can be constructed. It should be noted that the three-dimensional face mesh is composed of many triangular mesh patches. Therefore, for any pixel of the texture map corresponding to the face image sample, it corresponds to the corresponding triangular mesh patch.

[0213] In step 303, the server back-projects the three-dimensional face mesh into the camera space to pick colors and obtains the initial texture image corresponding to the three-dimensional face mesh from each perspective through linear interpolation.

[0214] Among them, in order to obtain the texture image, the 3D face meshes of each view can be respectively back-projected into the camera space (x, y) for color sampling. That is to say, it can be understood as spreading the 3D face mesh to the 2D space and mapping the 3D coordinates of the 3D face mesh in the 2D space. Based on this, for each pixel of the texture map formed on the 3D face model corresponding to the face image sample, the triangle mesh patch index of the back-projection mapping can be found. Since the RGB values of the three vertices of the triangle mesh patch are known, the RGB value of each pixel of the texture map formed by the face image sample can be obtained through linear interpolation. Furthermore, the RGB values of each pixel are combined to obtain the initial texture image of the texture map formed by the face image sample. By analogy, the initial texture images corresponding to the 3D face meshes of each view are obtained. For example, please continue to refer to Figure 4 As shown, extract the face geometry data of the face image samples from different views, generate 3D face meshes from different views, back-project the 3D face meshes into the camera space for color sampling, and obtain the initial texture images corresponding to the 3D face meshes of each view through linear interpolation.

[0215] In step 304, the server obtains the predefined masks corresponding to the initial texture images of each view, and performs alpha blending on the initial texture images of different views based on the predefined masks to obtain the first intermediate texture image.

[0216] Among them, the predefined masks corresponding to the initial texture images of the left view, the middle view, and the right view can be respectively obtained. The predefined mask is a mask with high credibility and is used to implement subsequent alpha blending. In this way, the initial texture images of each view can be respectively masked based on the predefined mask, and the masked initial texture images of each view are fused to obtain the texture image sample. Specifically, the following formula can be referred to for understanding:

[0217] T = M l T l + M f T f + M r T r

[0218] Among them, the M l 、M f and M r are respectively the predefined masks of the initial texture image of the left view, the predefined mask of the initial texture image of the middle view, and the predefined mask of the initial texture image of the right view. The T l 、T f and T rThey are the initial texture images from the left view, the initial texture images from the middle view, and the initial texture images from the right view respectively. Based on the above formula, the initial texture images from different views are subjected to alpha blending through a predefined mask to obtain a texture image sample T. The texture image sample T is fused by the initial texture images from multiple views, which can solve the problem of poor texture image quality caused by projection deviation and incomplete projection area. For example, please continue to refer to Figure 4 As shown, the predefined masks for each view can be obtained, and the initial texture images from different views are subjected to alpha blending based on the predefined masks for each view to obtain a first intermediate texture image, and the details of the first intermediate texture image are more comprehensive and rich.

[0219] In step 305, the server determines the distorted area of the first intermediate texture image, and fills in the distorted area through a preset human body template to obtain a second intermediate texture image, determines the missing area of the second intermediate texture image, and fills in the missing area through a preset human body template to obtain a texture image sample.

[0220] Since the initial texture images from different views are fused, there will be distorted and missing areas. Therefore, it is also necessary to fill in the distorted area of the first intermediate texture image. Specifically, through distortion detection, the distorted area is determined, and the distorted area is filled in through a preset human body template to form a smooth second intermediate texture image, making the second intermediate texture image smoother and more complete in fusion. For example, please continue to refer to Figure 4 As shown, the distorted area of the first intermediate texture image can be detected by detecting the distorted area means, and the distorted area is filled in through a preset human body template to obtain a second intermediate texture image.

[0221] Furthermore, even if the initial texture images from different views are fused, there will be missing areas that cannot be covered by multiple views, such as missing areas in the head, chin, and neck. Therefore, the missing areas can also be filled in according to the preset human body template to finally obtain a complete texture image sample. For example, please continue to refer to Figure 4 As shown, the missing area of the second intermediate texture image can be determined. For example, the missing areas are the head area and the neck area, and the missing areas are filled in through a preset human body template to obtain a complete texture image sample.

[0222] In the embodiment of the present application, the texture image sample undergoes alpha blending, artifact correction (i.e., filling in the distorted area), and filling in the missing area of the initial texture images from different views, and both the integrity and details of the texture are greatly improved, which can improve the accuracy of subsequent preset adversarial training.

[0223] In step 306, the server inputs the texture image samples into the discriminator in a preset adversarial model for training to obtain a trained discriminator, inputs the training latent space vectors into the generator of the preset adversarial model, and outputs corresponding training texture images.

[0224] It should be noted that the preset adversarial model can be StyleGAN, and the StyleGAN at least includes a generator and a discriminator. Please continue to refer to Figure 6 as shown in Figure 6 FIG. 7 is a structural diagram of the generator of StyleGAN. The generator of the StyleGAN is used to generate training texture images, and the discriminator is used to determine whether the generated training texture images are real.

[0225] Therefore, the texture image samples can be input into the discriminator in StyleGAN for training. That is, the discriminator first learns how to recognize texture images. The label of the texture image samples is 1, which is a "real" label. Based on this, the discriminator will output corresponding prediction results according to the texture image samples, compare the prediction results with the label 1 to obtain the training discrimination difference, and adjust the model parameters of the generator according to the training discrimination difference. Iteratively repeat the above process until the training discrimination difference converges to obtain a trained discriminator. The trained discriminator has the ability to recognize texture images as "false or real".

[0226] The generator consists of a mapping network and a generation network. The mapping network always consists of 8 fully connected layers. The input is a 512-dimensional noise vector. After passing through 8 fully connected layers, the 512-dimensional training latent space vector in the embodiment of the present application is obtained. That is, the training latent space vector is obtained by processing the noise vector through 8 fully connected layers. The advantage of this is to get rid of the influence of the input vector on the input data distribution, so that the training latent space vector has a better linear relationship with the subsequent generated training texture images, which is beneficial to the attribute control of the training texture images. The generation network is a structure with gradually increasing resolution, consisting of 9 stylized modules. By inputting the training latent space vector into the generation network, the training texture images can be obtained.

[0227] In step 307, the server inputs the training texture images into the trained discriminator, outputs the discrimination results, and trains the generator according to the discrimination difference between the discrimination results and the preset labels until the discrimination difference converges to obtain a trained preset adversarial model.

[0228] Among them, in order to generate a texture image similar to the texture image sample, the generator needs to be trained. Therefore, the trained discriminator can be input with the training texture image, and the corresponding discrimination result is output. The training objective is that it is expected that the discrimination result of the trained discriminator for this training texture image is true. Therefore, the discrimination result can be compared with the preset label (i.e., 1, representing "true") to obtain the discrimination difference, and the model parameters of the generator can be adjusted according to the discrimination difference. The above process is iteratively repeated until the discrimination difference converges, and the trained generator is obtained, that is, the trained preset adversarial model is obtained. The trained preset adversarial model includes the trained generator, and the trained generator has learned the distribution of the texture image sample and has the ability to generate a texture image similar to the texture image sample. Moreover, since the texture image generated by the trained generator is a non-linear texture, that is, a non-linear texture base is built, compared with the texture image generated by the linear texture base, the expression ability is richer and more vivid.

[0229] In step 308, the server obtains the face image to be processed, extracts the face rendering parameters of the face image to be processed, and inputs the latent space vector into the trained generator of the trained preset adversarial model to output the corresponding predicted texture image.

[0230] Among them, the face image to be processed is the given two-dimensional image of the object. The object needs to perform texture reconstruction based on the face image to be processed to realize the production of a 3D digital human. Therefore, after obtaining the face image to be processed, the face rendering parameters of the face image to be processed can be quickly extracted through DECA. The face rendering parameters at least include identity parameters, expression parameters, linear texture image, lighting parameters, and camera parameters. The identity parameters are used to control the identity of the face and can be composed of a 100-dimensional vector. The expression parameters are used to control the expression of the face and can be composed of a 40-dimensional vector. The linear texture image can be composed of linear texture parameters and is used to control the linear texture of the face and can be composed of a 100-dimensional vector. The lighting parameters are used to control the lighting of the face and can be composed of a 9-dimensional vector. The camera parameters are used to control the imaging angle of the face and can be composed of camera parameters.

[0231] Since the trained generator in the trained preset adversarial model has learned the distribution of the texture image sample and has the ability to generate a texture image similar to the texture image sample, a latent space vector can be randomly generated and input into the trained generator of the trained preset adversarial model to randomly output the corresponding predicted texture image, and the predicted texture image is similar to the texture image sample used for training before.

[0232] In step 309, the server replaces the linear texture image with the predicted texture image, and through differentiable rendering, reconstructs an image based on the identity parameter, expression parameter, predicted texture image, lighting parameter, and camera parameter to obtain the image to be judged.

[0233] Among them, the face rendering parameters at least include the identity parameter, expression parameter, linear texture image, lighting parameter, and camera parameter. Since the linear texture image is linear and limited by its expressive ability, it cannot represent strong face features, such as deep nasolabial folds, wrinkles, etc. Therefore, in the embodiment of the present application, the predicted texture image can be used to replace the linear texture image, that is, the linear texture image is deleted, and an image is reconstructed through differentiable rendering using the identity parameter, expression parameter, predicted texture image, lighting parameter, and camera parameter to obtain the image to be judged. The following formula can be referred to for understanding:

[0234] I = render(p id , p exp , G tex (z), p cam , p light )

[0235] Among them, the render() represents the differentiable rendering function, p id is the identity parameter, p exp is the expression parameter, G tex (z) is the predicted texture image (i.e., the non-linear texture basis), p cam is the lighting parameter, p light is the camera parameter. Thus, through the above formula, the image to be judged I is obtained.

[0236] In step 310, the server constructs a first constraint condition based on the pixel difference between the image of the face to be judged and the image of the face to be processed, constructs a second constraint condition based on the identity difference between the image of the face to be judged and the image of the face to be processed, generates a third constraint condition for controlling the change amplitude of the latent space vector, and constructs an optimization function corresponding to the latent space vector based on the first constraint condition, the second constraint condition, and the third constraint condition.

[0237] Among them, the pixel difference refers to the sum of the absolute values of the differences of each pixel between the image of the face to be judged and the image of the face to be processed. Thus, a first constraint condition can be constructed based on the pixel difference between the image of the face to be judged and the image of the face to be processed, and the constraint of the first constraint condition is to make the pixel difference converge as much as possible.

[0238] It should be noted that all face images carry identity parameters. These identity parameters are used to identify the identity of the face and can be composed of 100-dimensional vectors. The identity difference is the absolute value of the difference between the identity parameters extracted from the face image to be judged and the identity parameters extracted from the face image to be processed. Based on this, a second constraint condition can be constructed according to the identity difference between the face image to be judged and the face image to be processed. The constraint of this second constraint condition is to make the identity difference converge as much as possible.

[0239] Correspondingly, since the optimization process is a multi-round iterative process, in order to avoid the change range of the latent space vector being too large, resulting in too large a change in the corresponding predicted texture image and affecting the convergence of the optimization, it is necessary to control the change range of the latent space vector before and after, that is, to generate a third constraint condition for controlling the change range of the latent space vector. The constraint of this third constraint condition is to make the change range of the latent space vector as small as possible.

[0240] Furthermore, an optimization function corresponding to the latent space vector can be constructed based on the first constraint condition, the second constraint condition, and the third constraint condition. Based on this optimization function, the optimal solution corresponding to the latent space vector can be solved. For a better understanding of the embodiments of the present application, please refer to the following formula together:

[0241] L = λ pix L pix + λ id L id + λ reg L reg

[0242] Among them, this L pix is the first constraint condition, L id is the second constraint condition, L reg is the third constraint condition, this λ pix is the first weight of the first constraint condition, λ id is the second weight of the second constraint condition, λ reg is the third weight of the third constraint condition. The magnitudes of the first weight, the second weight, and the third weight are all preset settings. For example, the first weight is 0.3, the second weight is 0.4, and the third weight is 0.3, which are used to adjust the importance of each constraint in the optimization process. Based on this, the optimization function L corresponding to the latent space vector is constructed.

[0243] In step 311, the server updates the latent space vector according to the constraints of the optimization function, inputs the updated latent space vector into the trained preset adversarial model, outputs the updated predicted texture image, and reconstructs an image based on the updated predicted texture image and the face rendering parameters to obtain the updated face image to be judged.

[0244] Update the latent space vector according to the third constraint condition of the optimization function to control the amplitude of the change of the latent space vector before and after, and obtain the updated latent space vector.

[0245] Correspondingly, in order to verify the progress of the update, it is necessary to input the updated latent space vector into the trained generator of the trained StyleGAN again to generate an updated predicted texture image, and reconstruct the image again through differentiable rendering according to the updated predicted texture image and the face rendering parameters to obtain an updated face image to be judged.

[0246] In step 312, when the server detects that the pixel difference between the updated face image to be judged and the face image to be processed satisfies the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed satisfies the second constraint condition, the current updated latent space vector is determined as the target latent space vector.

[0247] Among them, it is necessary to synchronously detect whether the pixel difference between the updated face image to be judged and the face image to be processed satisfies the first constraint condition, and whether the identity difference between the updated face image to be judged and the face image to be processed satisfies the second constraint condition.

[0248] When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed satisfies the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed satisfies the second constraint condition, it indicates that the updated predicted texture image generated according to the current updated latent space vector and the updated face image to be judged obtained by reconstructing the image are close in pixels and have the same identity as the face image to be processed, indicating that the current updated latent space vector is accurate. Therefore, the current updated latent space vector is determined as the target latent space vector.

[0249] In step 313, when the server detects that the updated face image to be judged does not satisfy the constraints of the optimization function, re - execute the update of the latent space vector according to the constraints of the optimization function.

[0250] Among them, when it is detected that the pixel difference between the updated face image to be judged and the face image to be processed does not meet the first constraint condition, or the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition, it indicates that there is still a large difference between the predicted texture image generated by the updated latent space vector and the texture image of the face image to be processed. It is determined that the updated face image to be judged does not meet the constraints of the optimization function, and step 311 is returned to execute. The latent space vector is updated according to the constraints of the optimization function until the pixel difference between the updated face image to be judged and the face image to be processed meets the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed meets the second constraint condition.

[0251] In step 314, the server obtains the target texture image output by the target latent space vector in the trained adversarial model.

[0252] Among them, since the target texture image output by inputting the target latent space vector into the trained preset adversarial model and the difference between the face rendering parameters and the face image to be judged obtained by differentiable rendering and reconstructing the image converges with the face image to be processed, it indicates that the two face images are very close, and it also indicates that the target texture image is very close to the texture image of the face image to be processed. Therefore, the target texture image output by the target latent space vector in the trained adversarial model can be directly obtained as the texture image reconstructed according to the corresponding texture of the image to be processed. Since the target texture image is a non-linear texture image, the expression ability of the target texture image is richer and the effect is more realistic.

[0253] Based on this, corresponding realistic 3D digital humans can be automatically produced according to the target texture image, in cooperation with 3D reconstruction technology and hair generation technology.

[0254] As can be seen from the above, in the embodiment of the present application, a face image sample is obtained, a texture image sample is constructed based on the face image sample, and the preset adversarial model is trained through the texture image sample to obtain a trained preset adversarial model; a face image to be processed is obtained, and the face rendering parameters of the face image to be processed are extracted; the latent space vector is input into the trained preset adversarial model, and the corresponding predicted texture image is output, and the image is reconstructed according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; an optimization function corresponding to the latent space vector is constructed according to the difference between the face image to be judged and the face image to be processed, and the target latent space vector that minimizes the difference is calculated according to the optimization function; the target texture image output by the target latent space vector in the trained adversarial model is obtained. In this way, a corresponding texture image sample is constructed through the face image sample, and the distribution of the texture image sample is learned through the preset adversarial model to obtain a trained preset adversarial model. Furthermore, based on the non-linear basis learned by the trained preset adversarial model, an image is reconstructed to obtain a face image to be judged, and an optimization function is constructed according to the difference between the face image to be judged and the face image to be processed for optimization, and the optimal target latent space vector is solved to generate a target texture image close to the original image. Compared with the related technology that uses a 3D light field acquisition device to implement texture reconstruction, the embodiment of the present application reduces the processing cost and greatly improves the efficiency of image processing.

[0255] Furthermore, in the embodiment of the present application, the accuracy of subsequent backprojection can be improved by performing preprocessing on the face image sample, such as expression normalization, illumination normalization, and hair removal, so as to better improve the accuracy of the texture. In addition, the optimization function can be specifically optimized according to the first constraint condition of pixel difference and the second constraint condition of identity difference, so that the similarity between the target texture image after texture reconstruction and the original image is better.

[0256] In some embodiments, to better illustrate the embodiments of the present application, please refer to Figure 8 As shown, the server can select a face image sample from the face photo set, and construct a texture image sample through an offline texture production algorithm. By training the preset adversarial model according to the texture image sample, non-linear texture basis learning is performed to obtain a trained preset adversarial model, that is, a non-linear texture basis is built. Furthermore, the face image to be processed input by the receiving object is received, and in combination with the principle of the 3DMM face geometry basis, an optimization function is constructed to solve the corresponding texture basis parameters (i.e., the target latent space vector), and finally the high-definition target texture image corresponding to the target latent space vector is obtained. That is, in the embodiment of the present application, an unsupervised deep learning-based texture reconstruction method can be used to perform texture reconstruction on portrait photos in any scenario at low cost. On the basis of cost savings, the efficiency and accuracy of texture reconstruction can also be improved.

[0257] For the specific implementation of each of the above steps, reference may be made to the previous embodiments, which will not be elaborated here.

[0258] To facilitate the better implementation of the image processing method provided in the embodiments of the present application, the embodiments of the present application also provide an apparatus based on the above image processing method. The meanings of the terms are the same as those in the above image processing method, and the specific implementation details can be referred to the descriptions in the method embodiments.

[0259] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the image processing apparatus provided in the embodiments of the present application. The image processing apparatus is applied to a computer device, and the image processing apparatus may include a first acquisition unit 401, a second acquisition unit 402, a reconstruction unit 403, a calculation unit 404, and a third acquisition unit 405, etc.

[0260] The first acquisition unit 401 is configured to acquire a face image sample, construct a texture image sample based on the face image sample, and train a preset adversarial model through the texture image sample to obtain a trained preset adversarial model. The trained preset adversarial model learns the distribution of the texture image sample and is used to generate a texture image similar to the texture image sample.

[0261] In some embodiments, the first acquisition unit 401 includes:

[0262] An acquisition subunit (not labeled) for acquiring a face image sample;

[0263] A perspective enhancement subunit (not labeled) for performing perspective enhancement processing on the face image sample to obtain face image samples under different perspectives;

[0264] A generation subunit (not labeled) for extracting the face geometric data of the face image sample under each perspective and generating a three-dimensional face mesh under each perspective according to the face geometric data;

[0265] An inverse projection subunit (not labeled) for inverse projecting the three-dimensional face mesh into the camera space for color sampling and obtaining an initial texture image corresponding to the three-dimensional face mesh under each perspective through linear interpolation;

[0266] A fusion subunit (not labeled) for performing multi-perspective fusion on the initial texture images under different perspectives to obtain a texture image sample;

[0267] A training subunit (not labeled) for training a preset adversarial model through the texture image sample to obtain a trained preset adversarial model.

[0268] In some embodiments, the perspective enhancer unit (not labeled) is configured to:

[0269] Preprocess the face image sample to obtain an intermediate face image sample, where the preprocessing includes at least one of expression normalization, illumination normalization, or hair removal;

[0270] Perform perspective enhancement processing on the intermediate face image sample to obtain face image samples from different perspectives.

[0271] In some embodiments, the fusion unit (not labeled) includes:

[0272] An acquisition sub-module for acquiring a predefined mask corresponding to the initial texture image from each perspective;

[0273] A mixing sub-module for performing alpha blending on the initial texture images from different perspectives based on the predefined mask to obtain a texture image sample.

[0274] In some embodiments, the mixing sub-module (not labeled) is configured to:

[0275] Perform alpha blending on the initial texture images from different perspectives based on the predefined mask to obtain a first intermediate texture image;

[0276] Determine the distorted area of the first intermediate texture image, and complete the distorted area through a preset human body template to obtain a second intermediate texture image;

[0277] Determine the missing area of the second intermediate texture image, and complete the missing area through a preset human body template to obtain a texture image sample.

[0278] In some embodiments, the preset adversarial model includes at least a generator and a discriminator. The training sub-unit (not labeled) is configured to:

[0279] Input the texture image sample into the discriminator in the preset adversarial model for training to obtain a trained discriminator;

[0280] Input the training latent space vector into the generator of the preset adversarial model to output a corresponding training texture image;

[0281] Input the training texture image into the trained discriminator to output a discrimination result;

[0282] Train the generator according to the discrimination difference between the discrimination result and the preset label until the discrimination difference converges to obtain a trained preset adversarial model. The trained preset adversarial model includes at least a trained generator.

[0283] A second acquisition unit 402, configured to acquire a face image to be processed and extract face rendering parameters of the face image to be processed.

[0284] A reconstruction unit 403, configured to input a latent space vector into the trained preset adversarial model, where the latent space vector is randomly generated texture base parameters for converting into a texture image similar to a texture image sample, output a corresponding predicted texture image, and reconstruct an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged.

[0285] In some embodiments, the reconstruction unit 403 includes:

[0286] An input subunit (not labeled), configured to input the latent space vector into the trained generator of the trained preset adversarial model and output a corresponding predicted texture image;

[0287] A reconstruction subunit (not labeled), configured to reconstruct an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged.

[0288] In some embodiments, the face rendering parameters at least include identity parameters, expression parameters, a linear texture image, lighting parameters, and camera parameters. The reconstruction subunit (not labeled) is configured to:

[0289] Replace the linear texture image with the predicted texture image;

[0290] Reconstruct an image based on the identity parameters, expression parameters, predicted texture image, lighting parameters, and camera parameters through differentiable rendering to obtain an image to be judged.

[0291] A calculation unit 404, configured to construct an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed, and calculate a target latent space vector that minimizes the difference according to the optimization function.

[0292] In some embodiments, the calculation unit 404 includes:

[0293] A first construction subunit (not labeled), configured to construct a first constraint condition according to the pixel difference between the face image to be judged and the face image to be processed;

[0294] A second construction subunit (not labeled), configured to construct a second constraint condition according to the identity difference between the face image to be judged and the face image to be processed;

[0295] A generation subunit (not labeled), configured to generate a third constraint condition for controlling the change amplitude of the latent space vector;

[0296] A third construction subunit (not labeled) for constructing an optimization function corresponding to the latent space vector based on at least one of the first constraint condition, the second constraint condition, and the third constraint condition;

[0297] A calculation subunit (not labeled) for calculating a target latent space vector that minimizes the difference according to the optimization function.

[0298] In some embodiments, the calculation subunit (not labeled) includes:

[0299] An update sub-module (not labeled) for updating the latent space vector according to the constraints of the optimization function;

[0300] An input sub-module (not labeled) for inputting the updated latent space vector into the trained preset adversarial model, outputting an updated predicted texture image, and reconstructing an image according to the updated predicted texture image and the face rendering parameters to obtain an updated face image to be judged;

[0301] A result sub-module (not labeled) for obtaining a target latent space vector when it is detected that the updated face image to be judged satisfies the constraints of the optimization function.

[0302] In some embodiments, the result sub-module (not labeled) is used for:

[0303] When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed satisfies the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed satisfies the second constraint condition, determining the current updated latent space vector as the target latent space vector.

[0304] In some embodiments, the calculation subunit (not labeled) further includes:

[0305] A re-execution sub-module (not labeled) for re-executing the update of the latent space vector according to the constraints of the optimization function when it is detected that the updated face image to be judged does not satisfy the constraints of the optimization function.

[0306] A third acquisition unit 405 for acquiring a target texture image output by the target latent space vector in the trained adversarial model.

[0307] For the specific implementation of each of the above units, reference may be made to the previous embodiments and will not be elaborated here.

[0308] As described above, in the embodiment of the present application, the first acquisition unit 401 acquires a face image sample, constructs a texture image sample based on the face image sample, and trains a preset adversarial model with the texture image sample to obtain a trained preset adversarial model; the second acquisition unit 402 acquires a face image to be processed and extracts the face rendering parameters of the face image to be processed; the reconstruction unit 403 inputs a latent space vector into the trained preset adversarial model, outputs a corresponding predicted texture image, and reconstructs an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; the calculation unit 404 constructs an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed, and calculates a target latent space vector that minimizes the difference according to the optimization function; the third acquisition unit 405 acquires a target texture image output by the target latent space vector in the trained adversarial model. In this way, a corresponding texture image sample is constructed through the face image sample, and the distribution of the texture image sample is learned through a preset adversarial model to obtain a trained preset adversarial model. Furthermore, an image is reconstructed based on the non-linear basis learned by the trained preset adversarial model to obtain a face image to be judged. An optimization function is constructed according to the difference between the face image to be judged and the face image to be processed for optimization, and the optimal target latent space vector is solved to generate a target texture image close to the original image. Compared with the related technology that uses a 3D light field acquisition device to implement texture reconstruction, the embodiment of the present application reduces the processing cost and greatly improves the efficiency of image processing.

[0309] For the specific implementation of each of the above units, reference may be made to the previous embodiments and will not be elaborated herein.

[0310] Refer to Figure 10 , Figure 10 FIG. is a block diagram of a part of the structure of the terminal 140 for implementing the embodiment of the present disclosure. The terminal 140 includes: a radio frequency (RF) circuit 510, a memory 515, an input unit 530, a display unit 540, a sensor 550, an audio circuit 560, a wireless fidelity (WiFi) module 570, a processor 580, and a power supply 590, etc. Those skilled in the art can understand that Figure 10 The structure of the terminal 140 shown does not constitute a limitation on a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0311] The RF circuit 510 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information of the base station, it is given to the processor 580 for processing; in addition, the uplink data designed is sent to the base station.

[0312] The memory 515 can be used to store software programs and modules. The processor 580 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 515.

[0313] The input unit 530 can be used to receive input digital or character information and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 530 can include a touch panel 531 and other input devices 532.

[0314] The display unit 540 can be used to display the input information or provided information and various menus of the terminal. The display unit 540 can include a display panel 541.

[0315] The audio circuit 560, speaker 561, and microphone 562 can provide an audio interface.

[0316] In this embodiment, the processor 580 included in the terminal 140 can execute the image processing method of the previous embodiment.

[0317] The terminal 140 of the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0318] Figure 11 It is a structural block diagram of a part of the server 110 for implementing the embodiments of the present disclosure. The server 110 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 can be transient storage or persistent storage. The programs stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 600. Further, the central processing unit 622 can be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the server 600.

[0319] The server 600 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0320] The central processing unit 622 in the server 600 can be used to execute the image processing method of the embodiments of the present disclosure.

[0321] The embodiments of the present disclosure also provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the image processing methods of the foregoing various embodiments.

[0322] The embodiments of the present disclosure also provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned image processing method.

[0323] In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0324] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0325] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0326] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0327] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0328] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0329] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0330] It should also be understood that the various implementation manners provided by the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.

[0331] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0332] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. An image processing method, characterized in that, comprising: obtaining a face image sample, constructing a texture image sample based on the face image sample, and training a preset adversarial model through the texture image sample to obtain a trained preset adversarial model. The trained preset adversarial model learns the distribution of the texture image sample and is used to generate a texture image similar to the texture image sample; obtaining a face image to be processed and extracting the face rendering parameters of the face image to be processed; inputting a latent space vector into the trained preset adversarial model. The latent space vector is a randomly generated texture basis parameter used to be transformed into a texture image similar to the texture image sample, outputting a corresponding predicted texture image, and reconstructing an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; constructing an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed, and calculating a target latent space vector that minimizes the difference according to the optimization function; obtaining the target texture image output by the target latent space vector in the trained adversarial model.

2. The image processing method according to claim 1, characterized in that, the constructing a texture image sample based on the face image sample includes: performing perspective enhancement processing on the face image sample to obtain face image samples under different perspectives; extracting the face geometric data of the face image sample under each perspective, and generating a three-dimensional face mesh under each perspective according to the face geometric data; back-projecting the three-dimensional face mesh into the camera space for color sampling, and obtaining an initial texture image corresponding to the three-dimensional face mesh under each perspective through linear interpolation; performing multi-perspective fusion on the initial texture images under different perspectives to obtain a texture image sample.

3. The image processing method according to claim 2, characterized in that, the performing perspective enhancement processing on the face image sample to obtain face image samples under different perspectives includes: performing preprocessing on the face image sample to obtain an intermediate face image sample, where the preprocessing includes at least one of expression normalization, illumination normalization, or hair removal processing; performing perspective enhancement processing on the intermediate face image sample to obtain face image samples under different perspectives.

4. The image processing method according to claim 2, characterized in that, the performing multi-perspective fusion on the initial texture images under different perspectives to obtain a texture image sample includes: obtaining a predefined mask corresponding to the initial texture image under each perspective; performing alpha blending on the initial texture images under different perspectives based on the predefined mask to obtain a texture image sample.

5. The image processing method according to claim 4, characterized in that, the performing alpha blending on the initial texture images under different perspectives based on the predefined mask to obtain a texture image sample includes: performing alpha blending on the initial texture images under different perspectives based on the predefined mask to obtain a first intermediate texture image; Determine the distorted area of the first intermediate texture image, and complete the distorted area through a preset human body template to obtain a second intermediate texture image; Determine the missing area of the second intermediate texture image, and complete the missing area through a preset human body template to obtain a texture image sample.

6. The image processing method according to any one of claims 1 to 5, characterized in that, the preset adversarial model at least includes a generator and a discriminator, and training the preset adversarial model with the texture image sample to obtain a trained preset adversarial model includes: Inputting the texture image sample into the discriminator in the preset adversarial model for training to obtain a trained discriminator; Inputting the training latent space vector into the generator of the preset adversarial model to output a corresponding training texture image; Inputting the training texture image into the trained discriminator to output a discrimination result; Training the generator according to the discrimination difference between the discrimination result and a preset label until the discrimination difference converges to obtain a trained preset adversarial model, and the trained preset adversarial model at least includes a trained generator.

7. The image processing method according to claim 6, characterized in that, the inputting the latent space vector into the trained preset adversarial model to output a corresponding predicted texture image includes: Inputting the latent space vector into the trained generator of the trained preset adversarial model to output a corresponding predicted texture image.

8. The image processing method according to claim 1, characterized in that, the constructing an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed includes: Constructing a first constraint condition according to the pixel difference between the face image to be judged and the face image to be processed; Constructing a second constraint condition according to the identity difference between the face image to be judged and the face image to be processed; Generating a third constraint condition for controlling the change amplitude of the latent space vector; Constructing an optimization function corresponding to the latent space vector based on at least one of the first constraint condition, the second constraint condition, and the third constraint condition.

9. The image processing method according to claim 8, characterized in that, the calculating a target latent space vector that minimizes the difference according to the optimization function includes: Updating the latent space vector according to the constraints of the optimization function; Inputting the updated latent space vector into the trained preset adversarial model to output an updated predicted texture image, and reconstructing an image according to the updated predicted texture image and the face rendering parameters to obtain an updated face image to be judged; When it is detected that the updated face image to be judged satisfies the constraints of the optimization function, a target latent space vector is obtained.

10. The image processing method according to claim 9, characterized in that, the when it is detected that the updated face image to be judged satisfies the constraints of the optimization function, obtaining a target latent space vector includes: When it is detected that the pixel difference between the updated face image to be judged and the face image to be processed satisfies the first constraint condition, and the identity difference between the updated face image to be judged and the face image to be processed satisfies the second constraint condition, the current updated latent space vector is determined as the target latent space vector.

11. The image processing method according to claim 9, wherein, the method further includes: when it is detected that the updated face image to be judged does not satisfy the constraint of the optimization function, re - execute the update of the latent space vector according to the constraint of the optimization function.

12. The image processing method according to claim 1, wherein, the face rendering parameters at least include identity parameters, expression parameters, linear texture images, lighting parameters, and camera parameters; the reconstructing an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged includes: replacing the linear texture image with the predicted texture image; reconstructing an image based on the identity parameters, expression parameters, predicted texture image, lighting parameters, and camera parameters through differentiable rendering to obtain an image to be judged.

13. An image processing apparatus, wherein, it includes: a first acquisition unit, configured to acquire a face image sample, construct a texture image sample based on the face image sample, and train a preset adversarial model through the texture image sample to obtain a trained preset adversarial model, where the trained preset adversarial model learns the distribution of the texture image sample and is used to generate a texture image similar to the texture image sample; a second acquisition unit, configured to acquire a face image to be processed and extract the face rendering parameters of the face image to be processed; a reconstruction unit, configured to input a latent space vector into the trained preset adversarial model, where the latent space vector is randomly generated texture basis parameters and is used to be transformed into a texture image similar to the texture image sample, output a corresponding predicted texture image, and reconstruct an image according to the predicted texture image and the face rendering parameters to obtain a face image to be judged; a calculation unit, configured to construct an optimization function corresponding to the latent space vector according to the difference between the face image to be judged and the face image to be processed, and calculate a target latent space vector that minimizes the difference according to the optimization function; a third acquisition unit, configured to acquire the target texture image output by the target latent space vector in the trained adversarial model.

14. A computer - readable storage medium, wherein, the computer - readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the image processing method according to any one of claims 1 to 12.

15. A computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the image processing method according to any one of claims 1 to 12 is implemented.

16. A computer program product, including a computer program or instruction, wherein, When the computer program or instruction is executed by a processor, it implements the image processing method according to any one of claims 1 to 12.