Image processing method, device, computer equipment and storage medium

Through the combination of three-dimensional facial reconstruction and two-dimensional projection with facial identity feature fusion processing, the low efficiency problem of existing image processing methods is solved, and the automatic generation and efficient processing of target images are achieved.

CN113570684BActive Publication Date: 2025-09-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110088576.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-22
Publication Date
2025-09-26
Estimated Expiration
2041-01-27

AI Technical Summary

Technical Problem

Existing image processing methods require manual operation by users, resulting in low efficiency and high requirements on users' hands-on ability.

Method used

By obtaining the facial features of the initial template image and the input image, three-dimensional facial reconstruction and two-dimensional projection are performed to generate a target template image, which is then fused according to facial identity features to achieve automatic generation of the target image.

Benefits of technology

The automatic generation of target images is achieved, the efficiency of image processing is improved, the consistency of facial identity features and shape features is ensured, and the tedious operation of manual processing is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113570684B_ABST
    Figure CN113570684B_ABST
Patent Text Reader

Abstract

The present application relates to an image processing method, apparatus, computer device, and storage medium. The method includes: obtaining an initial template image and an initial input image containing a facial region, obtaining facial state features of the initial template image, and obtaining initial facial shape features of the initial input image; performing three-dimensional facial reconstruction on the initial template image and the initial input image based on the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image; performing two-dimensional projection on the three-dimensional reconstructed facial image to obtain reconstructed facial shape features corresponding to the obtained two-dimensional reconstructed facial image; adjusting the facial region of the initial template image according to the reconstructed facial shape features to obtain a target template image; obtaining facial identity features of the initial input image, and performing fusion processing based on the facial identity features and the target template image to obtain a target image. This method can improve image processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the continuous development of artificial intelligence in image processing technology, it is becoming increasingly common to personalize images and generate new images on computer devices. For example, after a user takes a photo through a terminal, they can perform personalized processing such as beautification on the photo to generate a new image.

[0003] However, this current image processing method requires manual operation by the user, such as manually selecting the image area to be processed or manually selecting the materials for beautifying the image. This image processing method is cumbersome and requires high manual skills from the user, resulting in low image processing efficiency. Summary of the Invention

[0004] Based on this, it is necessary to provide an image processing method, apparatus, computer equipment and storage medium that can improve image processing efficiency in response to the above technical problems.

[0005] An image processing method, comprising:

[0006] Obtain an initial template image and an initial input image containing a facial region;

[0007] Acquiring facial state features of the initial template image and acquiring initial facial shape features of the initial input image;

[0008] Performing three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image;

[0009] Performing two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image;

[0010] Acquire a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image;

[0011] The facial identity features of the initial input image are acquired, and a fusion process is performed based on the facial identity features and the target template image to obtain a target image.

[0012] An image processing device, comprising:

[0013] An image acquisition module is used to acquire an initial template image containing a facial region and an initial input image;

[0014] A feature acquisition module, configured to acquire facial state features of the initial template image and initial facial shape features of the initial input image;

[0015] a three-dimensional reconstruction module, configured to perform three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features, to obtain a three-dimensional reconstructed facial image;

[0016] a two-dimensional projection module, configured to perform two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image;

[0017] an adjustment module, configured to obtain a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image;

[0018] The fusion module is used to obtain the facial identity features of the initial input image, and perform fusion processing based on the facial identity features and the target template image to obtain a target image.

[0019] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0020] Obtain an initial template image and an initial input image containing a facial region;

[0021] Acquiring facial state features of the initial template image and acquiring initial facial shape features of the initial input image;

[0022] Performing three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image;

[0023] Performing two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image;

[0024] Acquire a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image;

[0025] The facial identity features of the initial input image are acquired, and a fusion process is performed based on the facial identity features and the target template image to obtain a target image.

[0026] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0027] Obtain an initial template image and an initial input image containing a facial region;

[0028] Acquiring facial state features of the initial template image and acquiring initial facial shape features of the initial input image;

[0029] Performing three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image;

[0030] Performing two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image;

[0031] Acquire a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image;

[0032] The facial identity features of the initial input image are acquired, and a fusion process is performed based on the facial identity features and the target template image to obtain a target image.

[0033] The above-mentioned image processing method, device, computer equipment and storage medium, after obtaining an initial template image and an initial input image containing a facial area, further obtain facial state features of the initial template image and obtain initial facial shape features of the initial input image, perform three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image, then perform two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image, obtain reconstructed facial shape features corresponding to the two-dimensional reconstructed facial image, adjust the facial area of ​​the initial template image according to the reconstructed facial shape features to obtain a target template image, and finally obtain facial identity features of the initial input image, perform fusion processing based on the facial identity features and the initial input image to obtain a target image, thereby realizing automatic generation of the target image, avoiding the tedious operations of manual processing, and greatly improving the efficiency of image processing.

[0034] Furthermore, since the fusion is performed based on the facial identity features of the initial input image and the target template image, and the facial shape features of the target template image match the facial shape features of the initial input image, the target image finally obtained is similar to the facial identity features of the initial input image and similar to the facial shape features of the initial input image, thereby ensuring the consistency of the facial identity features of the target image and the initial input image, and at the same time ensuring the subjective similarity between the target image and the initial input image. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A diagram showing an application environment of an image processing method in one embodiment;

[0036] Figure 2 1 is a flow chart of an image processing method according to an embodiment;

[0037] Figure 3 is a flowchart of an image processing method in another embodiment;

[0038] Figure 3A Schematic diagram of a 3DMM library in one embodiment;

[0039] Figure 3B A schematic diagram of a portion of expressions in a row of a 3DMM library in another embodiment;

[0040] Figure 4 Schematic diagram of a process for optimizing an initial template image in one embodiment;

[0041] Figure 5 Schematic diagram of the structure of a generator of a Pix2PixHD model in one embodiment;

[0042] Figure 6 A schematic diagram of a fusion process in one embodiment;

[0043] Figure 7A This is an example of an ID card template in one embodiment;

[0044] Figure 7B This is an example of a visa template for each country in one embodiment;

[0045] Figure 7C This is an example of a resume photo template in one embodiment;

[0046] Figure 8A A schematic diagram of a process for optimizing an ID photo template in one embodiment;

[0047] Figure 8B A schematic diagram of an image fusion process in one embodiment;

[0048] Figure 8C is a schematic diagram of a portrait enhancement process in one embodiment;

[0049] Figure 9 is a structural block diagram of an image processing device in one embodiment;

[0050] Figure 10 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0052] The image processing method provided in this application can be applied to Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. Both terminal 102 and server 104 can be used independently to execute the video data processing method provided in the embodiments of the present application. Terminal 102 and server 104 can also be used in conjunction to execute the video data processing and generation method provided in the embodiments of the present application.

[0053] For example, the server 104 may store a set of template images containing facial areas. When performing image processing, the server 104 returns a set of model images to the terminal 102 according to the request of the terminal 102. The user of the terminal selects a template image from the template image set as the initial template image. At the same time, the user takes a selfie through the terminal to obtain an initial input image containing the facial area and sends it to the server. The server can thereby obtain the initial template image and the initial input image. The server 104 further obtains the facial state features of the initial template image and the initial facial shape features of the initial input image. Then, based on the facial state features and the initial facial shape features, the initial template image and the initial input image are subjected to three-dimensional facial reconstruction to obtain a three-dimensional reconstructed facial image. The three-dimensional reconstructed facial image is subjected to two-dimensional projection to obtain a two-dimensional reconstructed facial image. The server further obtains the reconstructed facial shape features corresponding to the two-dimensional reconstructed facial image, adjusts the facial area of ​​the initial template image according to the reconstructed facial shape features to obtain a target template image. Finally, the server obtains the facial identity features of the initial input image, performs fusion processing based on the facial identity features and the target template image to obtain a target image, and returns the target image to the terminal 102.

[0054] The server 104 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 102 may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these. The terminal 102 and the server 104 may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0055] It should be noted that the image processing method provided in the embodiment of the present application is intended to generate a corresponding target image based on an initial template image and an initial input image. The facial identity features and facial shape features of the generated target image are similar to those of the initial input image, and other attribute features in the target image other than the facial identity features and facial shape features (including hairstyle, clothing, background, light, posture, expression, etc.) are consistent with the initial template image.

[0056] The image processing method provided in the embodiments of the present application relates to the field of artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0057] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0058] The solutions provided in the embodiments of this application mainly involve artificial intelligence computer vision technology and machine learning technology. Among them:

[0059] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, where cameras and computers replace the human eye in identifying and measuring objects, performing further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0060] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0061] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0062] This application illustrates the computer vision technology and machine learning technology involved through the following embodiments.

[0063] In one embodiment, Figure 2 As shown, an image processing method is provided, which is applied to Figure 1 The following steps are used as an example to illustrate the terminal in the figure:

[0064] Step 202: Acquire an initial template image containing a facial region and an initial input image.

[0065] The initial template image and the initial input image both include facial regions. The facial regions may be image regions corresponding to the face of a target object, such as a natural person, an animal, or a virtual character.

[0066] It should be noted that, based on the purpose of the image processing method provided in this application, the image used to provide facial identity features is the initial input image, and the image used to be fused with the facial identity features of the initial input image is the template image.

[0067] Specifically, the terminal obtains an initial template image and an initial input image, both of which include facial regions. Typically, the facial regions included in the initial template image and the initial input image are facial regions corresponding to different target objects.

[0068] In one embodiment, the initial input image can be an image provided by the user, such as a photo of a person taken by the user via a terminal. The initial template image can be an image provided by the terminal for the user to select as a template, such as an image of a game character or a public figure. In another embodiment, both the initial template image and the initial input image can be user-provided images. In this case, the user needs to specify which of the provided images should be used as the initial input image and which should be used as the initial input image.

[0069] In a specific embodiment, an image processing application may be run on the terminal, and the terminal may start the image processing application according to user operations. The image processing application may obtain a photo taken and selected by the user as an initial input image, and obtain an image selected by the user from a template image set as an initial template image.

[0070] Step 204: Acquire facial state features of the initial template image and acquire initial facial shape features of the initial input image.

[0071] Facial shape features refer to features related to the shape and outline of the face, such as data related to facial features, face shape, etc. Facial state features refer to features related to the state of the face, and facial state features may include facial posture, facial expression, and facial position angle.

[0072] Specifically, after obtaining the initial template image and the initial input image containing the facial area, the terminal can further obtain facial state features of the initial template image and further obtain facial shape features of the initial input image to obtain initial facial shape features.

[0073] It is understandable that, based on the purpose of the image processing method provided in this application, the initial facial shape features and facial state features can be different types of data depending on the different three-dimensional reconstruction technologies used in the specific implementation. Three-dimensional reconstruction here refers to the process of reconstructing a three-dimensional model based on one or more images (usually referred to as three-dimensional modeling). Three-dimensional modeling technology includes but is not limited to modeling methods based on image technology, modeling methods based on deep learning, and modeling methods based on a three-dimensional face model database. The embodiments of this application are mainly described with reference to the modeling method based on a three-dimensional face model database, and specific reference is made to the description of the subsequent embodiments.

[0074] In one embodiment, the terminal may select a machine learning-based neural network to extract facial state features from the initial template image and facial shape features from the initial input image. It is understood that different neural networks are used to extract facial state features from the initial template image and to extract facial shape features from the initial input image. The neural network for extracting facial state features from the initial template image is trained using images containing faces and expression annotations; the neural network for extracting facial shape features from the initial input image is trained using images containing faces and facial shape annotations.

[0075] In another embodiment, the terminal can obtain three-dimensional face model data from a pre-established three-dimensional face model database, and further obtain the shape weight coefficients of the initial input image corresponding to each three-dimensional shape base in the database as the initial facial shape features of the initial input image, and at the same time obtain the expression weight coefficients of the initial template image corresponding to each three-dimensional expression base in the database as the facial state features of the initial input image.

[0076] Step 206 , performing three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the facial shape features to obtain a three-dimensional reconstructed facial image.

[0077] 3D facial reconstruction refers to the process of reconstructing a 3D facial model (i.e., the 3D modeling mentioned above). A 3D reconstructed facial image refers to a 3D model containing a face. In one embodiment, when the facial region contained in the initial template image and the initial input image is a human face region, the 3D reconstructed facial image is a 3D face model.

[0078] Specifically, the terminal performs three-dimensional facial reconstruction on the initial template image and the initial input image based on the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image, the facial state features of the three-dimensional reconstructed facial image matching the facial state features of the initial template image, and the facial shape features matching the initial facial shape features of the initial input image.

[0079] In one embodiment, the terminal can input facial state features and initial facial shape features into a trained machine learning-based neural network. The neural network first fuses the facial state features and initial facial shape features, and performs three-dimensional facial reconstruction based on the fused features to obtain a three-dimensional reconstructed facial image.

[0080] In another embodiment, the terminal can determine the three-dimensional reconstructed facial shape based on each three-dimensional shape base and the determined shape weight coefficient of each three-dimensional shape base, and determine the three-dimensional reconstructed facial expression based on each three-dimensional expression base and the determined expression weight coefficient of each three-dimensional expression base, and finally generate a three-dimensional reconstructed facial image based on the determined three-dimensional reconstructed facial shape and the three-dimensional reconstructed facial expression.

[0081] Step 208 : Perform two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image.

[0082] Among them, two-dimensional projection refers to projection onto a two-dimensional plane, and the result of two-dimensional projection is a two-dimensional plane image.

[0083] Specifically, the terminal first establishes a spatial coordinate system for the three-dimensional reconstructed facial image, selects the observation point P, and then transforms the original spatial coordinate system into a spatial coordinate system with the observation point as the origin and PO as the z-axis through spatial coordinate transformation. The terminal further determines the projection plane based on the facial state characteristics of the initial template image, and finally maps the three-dimensional spatial coordinates to a predetermined projection plane to obtain a two-dimensional reconstructed facial image.

[0084] In one embodiment, the terminal may determine the projection plane based on the facial state features of the initial template image. For example, the terminal may determine the projection plane based on the position angle of the initial template image.

[0085] It can be understood that during the two-dimensional projection process, the facial shape features will not change significantly. The facial shape features of the obtained two-dimensional reconstructed facial image match the facial shape features of the three-dimensional reconstructed facial image, and the facial state features of the three-dimensional reconstructed facial image match the facial state features of the initial template image. Therefore, the facial state features of the obtained two-dimensional reconstructed facial image match the facial state features of the initial template image.

[0086] Step 210 , obtaining a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjusting the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image.

[0087] Specifically, the terminal obtains facial shape features of the two-dimensional reconstructed facial image to obtain the reconstructed facial shape features, and deforms and adjusts the facial area of ​​the initial template image according to the reconstructed facial shape features to obtain the target template image.

[0088] It can be understood that since the target template image in the embodiment of the present application is obtained by adjusting the facial area of ​​the initial template image according to the reconstructed facial shape features, and the reconstructed facial shape features are matched with the initial facial shape features, the facial shape features of the obtained target template image are similar to the initial facial shape features.

[0089] It can also be understood that the target template image in the embodiment of the present application is obtained by adjusting the facial area of ​​the initial template image according to the reconstructed facial shape features of the two-dimensional reconstructed facial image, and the facial state features of the two-dimensional reconstructed facial image match the initial template image. Then the deformation adjustment of the present application has nothing to do with the posture of the initial input image. Therefore, the present application has no restrictions on the posture of the initial input image, and the initial input image can be an image of any posture.

[0090] In one embodiment, the terminal may perform facial registration on the 2D reconstructed facial image to obtain facial feature points of the 2D reconstructed facial image as reconstructed facial shape features of the 2D reconstructed facial image. The facial feature points may include, but are not limited to, key points such as the eyes, nose, mouth, eyebrows, and facial contours.

[0091] In one embodiment, the terminal may deform the facial area of ​​the initial template image by reconstructing facial shape features through triangular patch stretching or pixel resampling, so that the facial shape features of the initial template image match the initial input image to obtain a target template image.

[0092] Step 212: Acquire facial identity features of the initial input image, perform fusion processing based on the facial identity features and the initial template image to obtain a target image.

[0093] Among them, facial identity features refer to features that can be used for identity recognition. Fusion refers to representing more than one data through one data and including the information expressed by these more than one data.

[0094] Specifically, the terminal can perform face recognition on the initial input image through mathematical calculations or a neural network based on computational learning to obtain facial identity features, and further obtain the attribute characteristics of the target template image, and perform fusion processing based on the facial identity features and attribute characteristics to obtain a target image, which can simultaneously express the facial identity features of the initial input image and the attribute characteristics of the target template image.

[0095] It should be noted that attribute features refer to other features in the facial area in addition to facial identity features. Attribute characteristics include facial shape features, as well as hairstyle, clothing, background, light, posture, etc.

[0096] In the above image processing method, after obtaining the initial template image and the initial input image containing the facial area, the facial state features of the initial template image are further obtained, and the initial facial shape features of the initial input image are obtained. The initial template image and the initial input image are subjected to three-dimensional facial reconstruction according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image. The three-dimensional reconstructed facial image is then two-dimensionally projected to obtain a two-dimensional reconstructed facial image. The reconstructed facial shape features corresponding to the two-dimensional reconstructed facial image are obtained. The facial area of ​​the initial template image is adjusted according to the reconstructed facial shape features to obtain a target template image. Finally, the facial identity features of the initial input image are obtained. The facial identity features and the initial input image are fused to obtain a target image. This realizes the automatic generation of the target image, avoids the tedious manual processing operations, and greatly improves the efficiency of image processing.

[0097] Furthermore, since the fusion is performed based on the facial identity features of the initial input image and the target template image, and the facial shape features of the target template image match the facial shape features of the initial input image, the target image finally obtained is similar to the facial identity features of the initial input image and similar to the facial shape features of the initial input image, thereby ensuring the consistency of the facial identity features of the target image and the initial input image, and at the same time ensuring the subjective similarity between the target image and the initial input image.

[0098] In one embodiment, Figure 3 As shown, an image processing method is provided, including the steps of optimizing a template image and fusion processing, wherein the step of optimizing the template image includes:

[0099] Step 302: Acquire an initial template image containing a facial region and an initial input image.

[0100] Step 304: Acquire 3D face model data from a pre-established 3D face model database; the 3D face model data includes a 3D shape basis set and a 3D expression basis set.

[0101] The 3D face model database refers to a database used to store 3D face model data. The 3D face model here is a 3D Morphable Face Model (3D Morphable Face Model), so the 3D face model database can be referred to as a 3DMM library. The 3DMM library includes a preset number of 3D shape bases and 3D expression bases. The 3D shape base is a 3D shape base model, and the 3D expression base is a 3D expression base model. The 3DMM library can scan multiple sets of 3D facial data together with high precision and align them. Principal Component Analysis (PCA) is then used to derive a lower-dimensional subspace from these 3D shape and color data. Variability is reflected in the ability to combine and deform these PCA subspaces, transferring the characteristics of one face to another or generating a new face.

[0102] like Figure 3A Figure 1 shows a schematic diagram of a 3DMM library in one embodiment. Each row represents the same person, and there are m people in total, so there are m rows. Each person corresponds to a shape, so there are m different shapes. Each column in a row corresponds to a different expression, and there are n expressions in total, so there are n columns.

[0103] It is understandable that in other embodiments, the expressions in each column may also be expressions at various positions and angles. Figure 3B , is a schematic diagram of some expressions of a row in the 3DMM library in one embodiment, where 301, 302, 303, and 304 correspond to different positions, angles, and expressions, respectively.

[0104] Step 306: Obtain shape weight coefficients corresponding to the three-dimensional shape bases of the initial input image, and determine the shape weight coefficients as initial facial shape features of the initial input image.

[0105] Specifically, the terminal can first obtain the feature points of each feature part in the initial input image, and then perform 3D fitting processing on each feature point to obtain the three-dimensional feature point corresponding to each feature point. The 3D fitting process is the process of adding a depth value to a two-dimensional image. The obtained three-dimensional feature point can be expressed as: (x, y, z), where x represents the horizontal coordinate value of the pixel point corresponding to the three-dimensional feature point; y represents the horizontal coordinate value of the pixel point corresponding to the three-dimensional feature point; and z represents the depth value of the three-dimensional feature point. Among them, x and y are the same as the x and y values ​​of the feature points of the initial input image.

[0106] In one embodiment, after obtaining the three-dimensional feature points of each feature point, the terminal can match the shape formed by the three-dimensional feature points with each three-dimensional shape basis to determine the weight coefficients of each three-dimensional shape basis that can form the shape of the initial input image. This can be understood as the process of solving the coefficients in the linear equation f(x) = a1x1 + a2x2 + a3x3 + ... + aixi + ... + anxn, where f(x) represents the shape formed by the three-dimensional feature points of the initial input image; xi represents the i-th three-dimensional shape basis; and ai represents the weight coefficient of the i-th three-dimensional shape basis. Based on the above process, the weight coefficients of each three-dimensional shape basis that can form the shape of the human face in the initial input image can be determined. The terminal uses the determined weight coefficients of each three-dimensional shape basis as the initial facial shape features of the initial input image.

[0107] In another embodiment, the terminal may average the pixel values ​​of the three-dimensional feature points obtained from the initial input image, and average the pixel values ​​of the three-dimensional feature points of each three-dimensional shape base, and determine the ratio of the average value corresponding to the three-dimensional feature points obtained based on the initial input image to the average value corresponding to the three-dimensional feature points of each three-dimensional shape base as the weight coefficient corresponding to each three-dimensional shape base.

[0108] Step 308: Obtain expression weight coefficients corresponding to the three-dimensional expression bases of the initial template image, and determine the expression weight coefficients as facial state features of the initial template image.

[0109] In one embodiment, the terminal can first obtain the feature points of each feature part in the initial template image, and then perform 3D fitting processing on each feature point respectively to obtain the three-dimensional feature point corresponding to each feature point. The terminal further constructs the three-dimensional feature points obtained from the initial template image into a matrix C, and constructs the three-dimensional feature points of each three-dimensional expression base into a matrix Mi respectively. Then, the difference matrix between the matrix C and Mi can be determined, and then the absolute value of the difference matrix is ​​taken, recorded as Di, and then the sum of the elements in each Di can be determined. Since the smaller the sum of the elements, the closer the expression displayed in the three-dimensional expression base is to the expression displayed in the initial template image, and the larger the sum of the elements of Di, the greater the difference between the expression displayed in the three-dimensional expression base and the expression displayed in the initial template image. Based on this, the weight coefficient of the basic expression base model with the smallest sum of elements can be set to a larger value, and the weight coefficient of the basic expression base model with the largest sum of elements can be set to a smaller value, thereby obtaining the weight coefficient of each three-dimensional expression base corresponding to the expression displayed in the initial template image.

[0110] Step 310: Determine a three-dimensional reconstructed face shape based on each three-dimensional shape basis and the determined shape weight coefficients of each three-dimensional shape basis.

[0111] Step 312 : Determine the three-dimensional reconstructed facial expression based on each three-dimensional expression base and the determined expression weight coefficient of each three-dimensional expression base.

[0112] It can be understood that since the three-dimensional shape basis and the three-dimensional expression basis are essentially composed of matrices, the terminal can perform weighted summation processing on the matrices of each three-dimensional shape basis and the weight coefficients corresponding to each three-dimensional shape basis, and the weighted summation result obtained is the three-dimensional reconstructed facial shape, and perform weighted summation processing on the matrices of each three-dimensional expression basis and the weight coefficients corresponding to each three-dimensional expression basis, and the weighted summation result obtained is the three-dimensional reconstructed facial expression.

[0113] Step 314 : generating a 3D reconstructed facial image based on the 3D reconstructed facial shape and the 3D reconstructed facial expression.

[0114] Specifically, the terminal performs summation processing on the matrix corresponding to the three-dimensional reconstructed face shape and the matrix corresponding to the three-dimensional reconstructed face expression, and the summation result is the three-dimensional reconstructed facial image.

[0115] Step 316 , performing two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image.

[0116] Step 318: Obtain the reconstructed facial shape features corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape features to obtain a target template image.

[0117] like Figure 4 FIG. 1 is a flow chart of optimizing the initial template image in a specific embodiment, referring to FIG. Figure 4 After obtaining the initial input image and the initial template image, the terminal first performs face detection and registration on the initial input image and the initial template image respectively, and obtains the facial feature points corresponding to the initial input image and the initial template image respectively. Then, based on the obtained facial feature points, the 3DMM parameters of the facial feature points corresponding to the initial input image and the initial template image are obtained (i.e., the shape weight coefficient and the expression weight coefficient mentioned above). The facial features and shape parameters (i.e., the shape weight coefficient) in the 3DMM parameters corresponding to the initial input image and the position, posture, and expression parameters (i.e., the expression weight coefficient) in the 3DMM parameters corresponding to the initial template image are taken, and a three-dimensional facial image is reconstructed based on the 3DMM library. The reconstructed three-dimensional facial image is projected two-dimensionally to obtain a two-dimensional reconstructed facial image of the three-dimensional facial image under the position, posture, and expression corresponding to the initial template image. The facial features and shape registration points (i.e., facial feature points) of the two-dimensional reconstructed facial image are obtained, and the facial features and shape registration points of the two-dimensional reconstructed facial image are used to deform and adjust the facial features and shape of the facial area of ​​the initial template image to obtain a target template image.

[0118] Further, continue to refer to Figure 3 The fusion processing steps are as follows: Step 320, obtaining the facial identity features of the initial input image, performing fusion processing based on the facial identity features and the target template image to obtain the target image.

[0119] In this embodiment, a three-dimensional reconstructed facial image is generated by obtaining three-dimensional facial model data from a pre-established three-dimensional facial model database. Since the three-dimensional facial model data includes a three-dimensional shape basis set and a three-dimensional expression basis set, the shape weight coefficient corresponding to each three-dimensional shape basis of the initial input image is obtained, and the shape weight coefficient is determined as the initial facial shape feature of the initial input image. The expression weight coefficient corresponding to each three-dimensional expression basis of the initial template image is obtained, and the expression weight coefficient is determined as the facial state feature of the initial template image. The facial shape feature of the initial input image and the facial state feature of the initial template image can be accurately expressed, thereby generating a three-dimensional reconstructed facial image whose facial shape feature has a high degree of match with the initial input image and whose facial state feature has a high degree of match with the initial template image.

[0120] In one embodiment, in the above step 210, the facial area of ​​the initial template image is adjusted according to the reconstructed facial shape features to obtain the target template image, which includes: triangulating the facial area of ​​the initial template image to obtain multiple triangular facets corresponding to the initial template image; and deforming each triangular facet according to the reconstructed facial shape to obtain the target template image.

[0121] Triangulation refers to splitting a plane into fragments, where each fragment must meet the following conditions: (1) Each fragment is a triangle; and (2) Any two triangles must either not intersect or intersect at a common edge (they cannot intersect at two or more edges simultaneously). In the embodiment of the present application, the facial region of the initial template image is triangulated, and the resulting fragments are triangular facets.

[0122] Specifically, in this embodiment, when the terminal performs two-dimensional projection, it determines the projection plane according to the position angle of the initial template image. Then, the position angle of the obtained two-dimensional reconstructed facial image is the same as the initial template image. Then, the terminal can triangulate the facial area of ​​the initial template image to obtain multiple triangular facets corresponding to the initial template image, and perform deformation and stretching processing on the two-dimensional reconstructed facial image.

[0123] In this embodiment, the reconstructed facial shape features are facial feature points. The terminal can determine the facial feature proportions and the specific shape of the face based on the facial feature points. Then, the terminal can use the determined facial feature proportions and the specific shape of the face as a standard to deform each triangular facet so that the deformed facial feature proportions and the specific shape of the face match the initial template image. When performing triangulation, the terminal can use a common triangulation algorithm, for example, an algorithm based on Delaunay triangles, including a flanging algorithm, a point-by-point insertion algorithm, a segmentation and merging algorithm, and the like.

[0124] In the above embodiment, the facial shape features of the initial template image are adjusted by stretching the triangular facets obtained by triangulating the facial area of ​​the initial template image, which is equivalent to dividing the initial template image into multiple sub-areas for fine adjustment. The facial area of ​​the initial template image can be adjusted comprehensively and accurately, and the facial features proportions and specific shape of the target template image obtained have a high degree of similarity with the initial template image.

[0125] In another embodiment, in the above step 210, adjusting the facial area of ​​the initial template image according to the reconstructed facial shape features to obtain the target template image includes: performing pixel resampling on the facial area of ​​the initial template image according to the reconstructed facial shape features; and determining the target template image based on the pixel resampling result.

[0126] Image resampling refers to resampling a digital image to the desired pixel positions or inter-pixel spacing to construct a new image after geometric transformation. The resampling process is essentially an image restoration process. It reconstructs a two-dimensional continuous function representing the original image from the input discrete digital image and then resamples it to the new pixel spacing and positions. The mathematical process involves estimating or interpolating the values ​​of the new sampling points using the values ​​of several surrounding pixels based on the reconstructed continuous function (surface).

[0127] Specifically, after the terminal obtains the facial feature points of the two-dimensional reconstructed facial image, it uses them as the reconstructed facial features, resamples the pixels of the facial area of ​​the initial template image according to the facial feature points, reconstructs the facial area of ​​the initial template image based on the pixel resampling results, and replaces the facial area of ​​the initial template image with the reconstructed image area to obtain the target template image.

[0128] In the above embodiment, by performing pixel resampling on the facial region of the initial template image to obtain the template image, the target template image can be quickly determined, thereby further improving the image processing efficiency.

[0129] In one embodiment, the above-mentioned image processing method further includes: inputting the target image into a trained beautification model, performing beautification processing on the target image through the beautification model to obtain a beautification image; and obtaining the beautification image output by the beautification model.

[0130] The beauty processing includes at least one of enhancing skin color brightness, improving skin quality, and enhancing the naturalness of makeup. The beauty model is a deep learning-based network model, obtained through supervised training using beauty training samples; the beauty training samples include the original image and the beauty-enhanced image corresponding to the original image. In one embodiment, the original image can be a regular selfie without beauty processing, and the beauty-enhanced image corresponding to the original image can be the original image after parameters are adjusted using the beauty application.

[0131] In one embodiment, the structure of the beauty model can adopt the Pix2PixHD model, which is a network composed of a generator and a discriminator, wherein the generator adopts a coarse-to-fine generator and the discriminator adopts a multi-scale discriminator. Figure 5 The following is a schematic diagram of the structure of the generator of the Pix2PixHD model. Figure 5 The generator consists of two parts, G1 and G2. The image is first downsampled by a factor of 2 through the convolutional layer of one generator, G1. Then, another generator, G2, is used to generate a low-resolution image. The result is multiplied element-by-element with the downsampled image, and then added. The result is output to the subsequent network of G1 to generate a high-resolution image. In one embodiment, the above-mentioned image processing method further includes: inputting the target image into a trained clarity enhancement model, performing clarity enhancement processing on the target image using the clarity enhancement model to obtain a clear image; and obtaining the clear image output by the clarity enhancement model.

[0132] Clarity enhancement refers to improving image resolution. The clarity enhancement model is a deep learning-based network model, trained through supervised training using clear training samples. These samples include the original image and a blurred image obtained by degrading the original image's clarity. Degradation processing refers to reducing image clarity. Degradation methods include, but are not limited to, image scaling, image downsampling, and blur kernel filtering.

[0133] In one embodiment, the structure of the beauty model can also adopt the Pix2PixHD model.

[0134] In one embodiment, the above-mentioned image processing method also includes: inputting the target image into a deep beautification model to obtain a deep beautification result image with improved skin color, skin quality and natural makeup, inputting the beautification result image into a portrait clarity enhancement model to improve the face resolution, and obtaining the final result image.

[0135] In one embodiment, obtaining facial identity features of an initial input image and fusing the facial identity features with a target template image to obtain a target image includes the following steps:

[0136] First, the initial input image and the target template image are encoded respectively to obtain the facial identity features of the initial input image and the target attribute features of the target template image.

[0137] Encoding is the process of converting information from one form or format to another. Encoding the initial input image is the process of expressing one type of characteristic information included in the initial input image. This characteristic information can specifically be a facial identity feature. Encoding the target template image is the process of expressing another type of characteristic information included in the target template image. This characteristic information can specifically be an attribute feature.

[0138] Specifically, the terminal may select a traditional encoding function to encode the initial input image and the target template image separately. Traditional encoding functions, such as encoding functions based on the SIFT (Scale Invariant Feature Transform) algorithm or the HOG (Histogram of Oriented Gradient) algorithm, etc. In another embodiment, the terminal may also select a neural network based on machine learning to encode the initial input image and the target template image. The neural network used for encoding may specifically be an encoding model based on convolution operations, etc. The present application mainly implements encoding through a neural network based on machine learning. The specific implementation process can refer to the description of subsequent embodiments.

[0139] Then, the facial identity features and target attribute features are fused to obtain the target features.

[0140] Fusion refers to representing more than one data item through one data item and including the information expressed by the more than one data item. In this embodiment, fusing more than one feature into one feature can remove the discreteness of the data and facilitate the subsequent decoding process.

[0141] Specifically, the terminal can combine, splice or weighted sum the combined identity features and attribute features, or further operate on the results of the combination, splicing or weighted summing operations through a neural network to obtain a target feature that integrates the two feature information.

[0142] In one embodiment, the terminal may perform multi-level fusion processing on the facial identity features and the target attribute features to obtain the target features. When performing the fusion, the terminal may perform feature fusion by using a step-by-step channel superposition method.

[0143] Then, the target features are decoded to obtain a target image; the target image matches the facial identity features of the initial input image and matches the target attribute features of the target template image.

[0144] Decoding is the inverse process of encoding. Decoding restores data expressed in another form to its original form or format, reconstructing a new image with the same form or format as the original image.

[0145] Specifically, after obtaining the target features, the terminal decodes them to restore the target image. Because the target features incorporate the facial identity features of the initial input image and the attribute features of the target template image, the target image maintains consistency with the initial input image in terms of facial identity features and with the target template image in terms of attribute features. The terminal can use either a traditional decoding function or a neural network to decode the target features.

[0146] In one embodiment, encoding a target template image to obtain attribute features of the target template image includes: encoding the target template image through an attribute feature encoding model to obtain the attribute features of the target template image; fusing facial identity features and target attribute features to obtain target features: fusing facial identity features and attribute features through a feature fusion model to obtain target features; decoding the target features to obtain a target image includes: decoding the target features through a decoding model to obtain the target image; wherein the feature fusion model, the decoding model and the attribute feature encoding model are obtained by jointly training by alternating unsupervised image samples and self-supervised image samples.

[0147] The attribute feature encoding model, feature fusion model, and decoding model are all machine learning models. These three models are trained by alternating between unsupervised and self-supervised image samples.

[0148] Unsupervised image samples are image samples without training labels and are used for unsupervised training. They include multiple pairs of samples, each of which consists of an initial facial image sample and a template facial image sample. Self-supervised image samples are image samples for which training labels can be automatically generated and are used for self-supervised training. They include multiple pairs of samples, each of which consists of an initial facial image sample and a template facial image sample, with the initial and template facial image samples in each pair being identical.

[0149] In one embodiment, the terminal may also use a machine learning model to encode the initial input image. Specifically, the terminal may encode the initial input image separately using a recognition feature encoding model to obtain facial identity features corresponding to the initial input image.

[0150] The identification feature encoding model is trained using universal image samples. These universal image samples serve as training samples for a machine learning model with universal facial identity feature encoding capabilities. This machine learning model is widely used in various face recognition scenarios. If the facial identity features encoded by a machine learning model with universal facial identity feature encoding capabilities meet the facial identity feature requirements of the image processing method provided herein, then the machine learning model with universal facial identity feature encoding capabilities can be used as the identification feature encoding model for the image processing method provided herein.

[0151] In other words, in this embodiment, the terminal processes the acquired target template image and the initial input image using four models (recognition feature encoding model, attribute feature encoding model, feature fusion model, and decoding model) to obtain the target image. The combined processing of these four models significantly improves the efficiency and accuracy of image processing.

[0152] In this embodiment, feature encoding and decoding are achieved through a deep learning neural network. Leveraging the powerful learning capabilities of neural networks, the target image is reconstructed, preserving the facial identity features of the initial input image and the attribute features of the target template image, based on the desired useful features encoded from the initial input image and the target template image. Furthermore, the feature fusion model, decoding model, and attribute feature encoding model are jointly trained using alternating unsupervised and self-supervised image samples. This allows unsupervised learning to be supplemented by self-supervised learning, resulting in a more effective model for image generation. Furthermore, the model training process does not require sample labeling, significantly reducing costs.

[0153] In one embodiment, the attribute feature encoding model, the feature fusion model, and the decoding model are included in the generation network; the training sample acquisition step, the unsupervised training step, the self-supervised training step, and the loop step of the generation network are as follows:

[0154] Training sample acquisition step: obtain unsupervised image samples and self-supervised image samples.

[0155] Among them, the unsupervised image sample includes a first initial facial image sample and a first template facial image sample; the first initial facial image sample and the first template facial image sample are different image samples; the self-supervised image sample includes a second initial facial image sample and a second template facial image sample; the second initial facial image sample and the second template facial image sample are the same image sample.

[0156] Unsupervised training step: Perform unsupervised training on the generative network based on unsupervised image samples, and adjust the model parameters of the attribute feature encoding model, feature fusion model, and decoding model.

[0157] The unsupervised image samples include several sets of unsupervised sample pairs, each of which includes a first initial facial image sample Source and a first template facial image sample Reference, i.e., Unsupervised(Source, Reference). Unsupervised image samples are used for unsupervised training, also known as unsupervised learning, a method in which a machine learning model learns based on unlabeled sample data.

[0158] It should be noted that the generative network is typically combined with the discriminative network to form a generative adversarial network (GAN). During training, the two networks learn through a game of mutual competition. The generative network randomly samples from the latent space as input, and its output is required to closely mimic the real samples in the training set. The discriminative network takes real samples or the output of the generative network as input, and its goal is to distinguish the output of the generative network from real samples as much as possible. The generative network, on the other hand, is required to deceive the discriminative network as much as possible. The two networks compete with each other, constantly adjusting their parameters, ultimately producing images that are indistinguishable from the real ones.

[0159] Therefore, in this embodiment, the terminal can construct an unsupervised training loss function for jointly training the discriminant network and the generative network based on the unsupervised image samples, perform training based on the unsupervised training loss function, and adjust the model parameters of the attribute feature encoding model, feature fusion model, and decoding model. The discriminant network can be a universal discriminant network, and the image samples in the supervised image samples can all be considered real samples and can be used as positive samples for the discriminant network. The target image samples generated by the generative network based on the initial facial image samples and the template facial image samples are generated images and can be used as negative samples for the discriminant network. The discriminant network learns to distinguish the output of the generative network from the real samples as much as possible.

[0160] In one embodiment, the generation network further includes an identification feature coding model, and the generation network is unsupervisedly trained according to the unsupervised image sample, and the model parameters of the attribute feature coding model, the feature fusion model and the decoding model are adjusted, including: encoding the first initial facial image sample through the identification feature coding model to obtain the facial identity feature of the first initial facial image sample; encoding the first template facial image sample through the attribute feature coding model to obtain the attribute feature of the first template facial image sample; inputting the facial identity feature of the first initial facial image sample and the attribute feature of the first template facial image sample into the feature fusion model and the decoding model in sequence to obtain the first target facial image sample; The identification feature encoding model and the attribute feature encoding model respectively encode the first target facial image sample to obtain the facial identity feature and attribute feature of the first target facial image sample; a discriminant network is obtained, and at least one of the first initial facial image sample and the first template facial image sample is used as a positive sample of the discriminant network, and the first target facial image sample is used as a negative sample of the discriminant network; based on the discriminant loss of the discriminant network, the difference in facial identity features between the first initial facial image sample and the first target facial image sample, and the difference in attribute features between the first template facial image sample and the first target facial image sample, the model parameters of the attribute feature encoding model, the feature fusion model and the decoding model are adjusted.

[0161] In this embodiment, the generation network includes an identification feature encoding model, an attribute feature encoding model, a feature fusion model, and a decoding model. The facial identity features encoded by a machine learning model with universal facial identity feature encoding capabilities meet the facial identity feature requirements of the image processing method provided in this application. Therefore, the machine learning model with universal facial identity feature encoding capabilities can be used as the identification feature encoding model for the image processing method provided in this application. The identification feature encoding model in the embodiment of this application can be pre-trained separately. During subsequent training, the model parameters of the identification feature encoding model are fixed, while the model parameters of the attribute feature encoding model, feature fusion model, and decoding model are adjusted.

[0162] It can be understood that in this embodiment, since unsupervised training is performed and there are no corresponding training labels, the terminal can respectively obtain the loss between the discrimination result of the discriminant network and the sample label, the loss between the facial identity feature Xid of the first target facial image sample Result and the facial identity feature Zid of the first initial facial image sample Source, and the loss between the attribute feature Xatt of the first target facial image sample Result and the attribute feature Zatt of the first template facial image sample Reference.

[0163] In this way, the terminal can use the weighted sum of the discriminator loss, the identity loss (the difference between Xid and Zid), and the attribute loss (the difference between Xatt and Yatt) as the unsupervised training loss function for the adversarial training generator and discriminator networks. Based on this unsupervised training loss function, the model parameters of the attribute feature encoding model, feature fusion model, and decoding model are adjusted. The weight distribution can be customized based on the importance of the loss to the generated results and the actual image processing requirements.

[0164] Self-supervised training steps: Perform self-supervised training on the generative network based on the self-supervised image samples, and adjust the model parameters of the attribute feature encoding model, feature fusion model, and decoding model.

[0165] In this embodiment, considering that pure unsupervised training is very difficult, self-supervised training can be supplemented by constructing self-supervised image samples for self-supervised training. The self-supervised image samples include several groups of self-supervised sample pairs, each of which includes a second initial facial image sample Source and a second template facial image sample Source, i.e., Self-supervised (Source, Source). Self-supervised image samples are used for self-supervised training, which can also be called self-supervised learning. It can be regarded as an "ideal state" of machine learning, where the machine learning model directly learns to generate labels from unlabeled data without the need for annotated data.

[0166] Specifically, the terminal can construct a self-supervised training loss function for jointly training the discriminant network and the generative network based on the self-supervised image samples, and train the same generative adversarial network (generative network + discriminant network) according to the self-supervised training loss function.

[0167] In one embodiment, the self-supervised training of the generative network is performed based on the self-supervised image sample, and the model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted, including: encoding the second initial facial image sample through the recognition feature encoding model to obtain the facial identity feature of the second initial facial image sample; encoding the second template facial image sample through the attribute feature encoding model to obtain the attribute feature of the second template facial image sample; inputting the facial identity feature of the second initial facial image sample and the attribute feature of the second template facial image sample into the feature fusion model and the decoding model in sequence to obtain the second target facial image sample; respectively encoding the second template facial image sample through the recognition feature encoding model and the attribute feature encoding model to obtain the attribute feature of the second template ... The second target facial image sample is encoded to obtain facial identity features and attribute features of the second target facial image sample; at least one of the second initial facial image sample and the second template facial image sample is used as a positive sample of the discriminant network, and the second target facial image sample is used as a negative sample of the discriminant network; based on the discriminant loss of the discriminant network, the pixel difference between the second target facial image sample and the second initial facial image sample, the difference in facial identity features between the second initial facial image sample and the second target facial image sample, and the difference in attribute features between the second template facial image sample and the second target facial image sample, the model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted.

[0168] Specifically, the terminal can input the second initial facial image sample Source into the recognition feature coding model of the generation network to obtain the facial identity feature Zid of the second initial facial image sample Source; input the second template facial image sample Source into the attribute feature coding model of the generation network to obtain the attribute feature Zatt of the second template facial image sample Source; and after the facial identity feature Zid and the attribute feature Zatt pass through the feature fusion model and decoding model of the generation network in sequence, the second target facial image sample Result is obtained.

[0169] Furthermore, the terminal inputs the second target facial image sample Result into the identification feature encoding model and the attribute feature encoding model of the generation network respectively to obtain the facial identity feature Xid and the attribute feature Xatt of the second target facial image sample Result.

[0170] It can be understood that since the first initial facial image sample and the first template facial image sample in the self-supervised training sample use the same image sample Source, ideally, the generated image should be identical to Source. That is, during the self-supervised training process, the model will automatically generate a training label, which is the image Source corresponding to the first initial facial image sample. Therefore, when constructing the self-supervised training loss function, the terminal can obtain the pixel loss between the second target facial image sample Result and the training label Source as the pixel reconstruction loss (Reconstruction Loss), and construct the loss function of the generative network based on this pixel reconstruction loss. Furthermore, since the generative network also includes two encoding branches that respectively encode facial identity features and attribute features, the loss function of the generative network can also include the loss for the difference in facial identity features (the difference between Xid and Zid) between the second target facial image sample Result and the second initial facial image sample Source, as well as the loss for the difference in attribute features (the difference between Xatt and Zatt) between the second target facial image sample Result and the second template facial image sample Reference.

[0171] In this way, the terminal can use the weighted sum of the discriminator loss (Discriminator Loss), pixel reconstruction loss (Reconstruction Loss), facial identity feature difference loss (the difference between Xid and Zid) and attribute feature difference loss (the difference between Xatt and Zatt) as the self-supervised training loss function for the adversarial training generative network and discriminant network. Based on this self-supervised training loss function, the model parameters of the attribute feature encoding model, feature fusion model, and decoding model are adjusted. The weight distribution can be customized based on the importance of the loss to the generated results and the actual image processing requirements.

[0172] Loop step: Repeat the self-supervised training step and the unsupervised training step so that unsupervised training and self-supervised training are performed alternately until the training stop condition is met.

[0173] Specifically, the terminal alternates between unsupervised image samples and self-supervised image samples to train the same generative adversarial network, so that unsupervised training and self-supervised training are performed alternately until the generation effect is stable, and the facial identity features of the output target facial image sample Result are significantly close to the facial identity features of the initial facial image sample Source, and the attribute features of the target facial image sample Result are significantly close to the attribute features of the template facial image sample Reference. That is, from a perceptual perspective, the generative network can generate a target facial image whose identity (Identity) is consistent with the initial facial image sample Source, and whose other features (posture, expression, lighting, and background, etc.) are consistent with the facial image sample Reference.

[0174] In this embodiment, model training is performed alternately using unsupervised and self-supervised data. This significantly reduces the cost of model training because unsupervised data training eliminates the need for sample labeling. Furthermore, the introduction of self-supervised data to assist in training the generative network significantly improves its stability under various circumstances. Furthermore, since both self-supervised and unsupervised training require no training labels, samples of their respective poses can be introduced during training. This allows the trained generative network to handle any facial image without any pose restrictions. This significantly improves image processing efficiency when the trained generative network is used for image processing.

[0175] In another embodiment, when training the generative network, initial facial images with correct postures can be selected from unsupervised image samples or self-supervised image samples for early training, and initial facial images with other postures can be added for training in the later stage of training. This not only improves the convergence speed of the model training, but also makes the trained model more stable.

[0176] It is understandable that the terminal primarily processes the facial region during fusion processing. Typically, the facial region accounts for a relatively small proportion of an image (except for close-up facial images). Therefore, the terminal can pre-process the image, capturing the facial region from the initial input image and the target template image, and then perform subsequent image processing based on the captured facial image. This can reduce the computational effort during image processing and improve image processing efficiency.

[0177] Specifically, the terminal can align facial feature points of the initial input image based on a traditional feature point positioning algorithm or a machine learning model, determine the facial feature points in the initial image, locate the facial area determined in the initial image according to the facial feature points determined in the initial image, and capture the facial area image according to the facial area located in the initial image.

[0178] For the target template image, the terminal can capture the facial region image in the same manner as the initial facial image. However, the processing of the target template image can be performed in advance, which improves image processing efficiency, or in real time, which reduces the device's storage burden.

[0179] It is understood that when the terminal performs fusion, it fuses facial region images extracted from the initial input image and the target template image. Therefore, after obtaining the fused image, it is necessary to reverse-paste the fused image to restore the image size and image content. Therefore, in this embodiment, after obtaining the fused image, the terminal can reverse-paste the fused image to the facial region in the target template image to obtain the target image. The resulting target image retains the facial identity features of the facial region in the initial image and the attribute features of the facial region in the target template image, and the portion outside the facial region is consistent with the template image.

[0180] For example, refer to Figure 6 , which shows a schematic flow chart of the fusion process in one embodiment. After acquiring the target template image, the terminal can align the facial feature points of the initial input image (i.e., facial detection and registration), and then determine a facial region screenshot based on the determined facial feature points (i.e., cutout based on the registration points) to obtain the initial facial image (i.e., posture-aligned facial image). In addition, the terminal can also align the facial feature points of the target template image (i.e., facial detection and registration), and then determine a facial region screenshot based on the determined facial feature points (i.e., cutout based on the registration points) to obtain the template facial image (i.e., posture-aligned facial image).

[0181] After that, the terminal can input the initial facial image into the recognition feature coding model to encode the facial identity features corresponding to the initial facial image, and input the template facial image into the attribute feature coding model to encode the attribute features. Then, the facial identity features and attribute features are input into the multi-level feature fusion model for feature fusion, and then the target facial image is obtained through the decoding model.

[0182] Furthermore, after obtaining the target facial image, the terminal may reversely paste the target facial image back onto the template image to obtain the target image.

[0183] In the above embodiment, when performing image processing, only the facial area is cut out for image processing, which not only reduces the amount of image processing data and improves image processing efficiency; but also eliminates the need to wastefully process areas outside the facial area, thereby avoiding wasting computing resources.

[0184] This application also provides an application scenario, which applies the above-mentioned image processing method. Specifically, the application of the image processing method in this application scenario is as follows:

[0185] In this application scenario, the image processing application running on the terminal executes the image processing method of this application to generate an ID photo. The user can use the image processing application to call the terminal's camera to take any selfie as the initial input image (without restrictions on head posture, expression, and lighting), and then select the standard ID photo template of the corresponding scene as the initial template image according to actual needs (dress requirements, background color requirements, photo size ratio requirements, etc.). Figures 7A-7C , which shows an example of ID photo templates provided by the terminal. 7A is an ID photo template, 7B is a visa template for various countries, and 7C is a resume photo template. It can be seen that resume photo templates are not limited to full frontal photos and can be in any pose. In this application scenario, the pose of the resulting ID photo is consistent with the non-face recognition features of the ID photo template (such as pose, expression, lighting, and background). The ID photo maintains facial recognition consistency with the user's input selfie and has a high subjective similarity to the user's input selfie.

[0186] To ensure that the generated ID photo is more usable, the selected standard ID photo template needs to keep the gender consistent with the user input image. The terminal's image processing application can then generate the ID photo through the following three stages.

[0187] Phase 1: Optimize ID photo template.

[0188] refer to Figure 8AAfter obtaining the initial input image and the initial template image, the terminal matches the initial input image with 105 three-dimensional shape bases in the pre-established 3DMM library (the facial features of each three-dimensional shape base are different) to obtain the shape weight coefficients of each three-dimensional shape base corresponding to the initial input image, and matches the initial template image with 101 three-dimensional expression bases in the pre-established 3DMM library (where the three-dimensional expression base is the expression base at each position and posture, that is, for any two expression bases, at least one of their positions, angles, and expressions is different) to obtain the expression weight coefficients of each three-dimensional expression base corresponding to the initial input image. Based on each three-dimensional shape base and the determined shape weight coefficients of each three-dimensional shape base, the terminal obtains the expression weight coefficients of each three-dimensional expression base. Number, determine the three-dimensional reconstructed face shape, based on each three-dimensional expression base and the expression weight coefficient of each determined three-dimensional expression base, determine the three-dimensional reconstructed face expression, based on the three-dimensional reconstructed face shape and the three-dimensional reconstructed face expression, generate a three-dimensional reconstructed facial image, further perform two-dimensional projection on the three-dimensional reconstructed facial image, obtain the two-dimensional reconstructed facial image of the user's face under the template position angle expression, obtain the facial features of the two-dimensional reconstructed facial image. The facial area of ​​the initial template image is triangulated to obtain multiple triangular facets corresponding to the initial template image, and each triangular facet is deformed according to the facial features of the two-dimensional reconstructed facial image. Finally, an optimized ID photo template image is obtained, that is, an optimized template image. The facial features proportions and face shape of the optimized template image are consistent with the initial input Figure 1 To.

[0189] Stage 2: Image fusion.

[0190] refer to Figure 8B The terminal inputs the optimized template image into the attribute feature encoding model to encode the attribute features, inputs the initial input image into the recognition feature encoding model to encode the facial identity features, inputs the attribute features and facial identity features into the fusion module for multi-level feature fusion to obtain the target features, and inputs the target features into the decoding model for decoding to obtain the fusion result image. The face recognition features of the fusion result image are consistent with the initial input image, and other non-face recognition features are consistent with the optimized template. Figure 1 To.

[0191] The recognition feature encoding model is a public face recognition model, and the attribute encoding model, fusion model, and decoding model are models in the generative network sent by the server. The server trains the generative network through the following steps:

[0192] First, the server obtains a generation network, unsupervised image samples and self-supervised image samples; the unsupervised image samples include a first initial facial image sample and a first template facial image sample; the first initial facial image sample and the first template facial image sample are different image samples; the self-supervised image samples include a second initial facial image sample and a second template facial image sample; the second initial facial image sample and the second template facial image sample are the same image sample.

[0193] Furthermore, the server performs unsupervised training. Specifically, the first initial facial image sample is encoded using the recognition feature encoding model to obtain a facial identity feature of the first initial facial image sample, the first template facial image sample is encoded using the attribute feature encoding model to obtain an attribute feature of the first template facial image sample, the facial identity feature of the first initial facial image sample and the attribute feature of the first template facial image sample are sequentially input into a feature fusion model and a decoding model to obtain a first target facial image sample, the first target facial image sample is encoded using the recognition feature encoding model and the attribute feature encoding model respectively to obtain a facial identity feature and an attribute feature of the first target facial image sample, a discriminant network is obtained, at least one of the first initial facial image sample and the first template facial image sample is used as a positive sample of the discriminant network, and the first target facial image sample is used as a negative sample of the discriminant network, and model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted based on the discriminant loss of the discriminant network, the difference in facial identity features between the first initial facial image sample and the first target facial image sample, and the difference in attribute features between the first template facial image sample and the first target facial image sample.

[0194] Furthermore, the server performs self-supervised training. Specifically, the server encodes the second initial facial image sample using the recognition feature encoding model to obtain a facial identity feature of the second initial facial image sample, encodes the second template facial image sample using the attribute feature encoding model to obtain an attribute feature of the second template facial image sample, sequentially inputs the facial identity feature of the second initial facial image sample and the attribute feature of the second template facial image sample into a feature fusion model and a decoding model to obtain a second target facial image sample, encodes the second target facial image sample using the recognition feature encoding model and the attribute feature encoding model respectively to obtain a facial identity feature and an attribute feature of the second target facial image sample, uses at least one of the second initial facial image sample and the second template facial image sample as a positive sample of the discriminant network, uses the second target facial image sample as a negative sample of the discriminant network, and adjusts model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model based on the discriminant loss of the discriminant network, the pixel difference between the second target facial image sample and the second initial facial image sample, the difference in facial identity features between the second initial facial image sample and the second target facial image sample, and the difference in attribute features between the second template facial image sample and the second target facial image sample.

[0195] Furthermore, the server repeats the unsupervised training and self-supervised training steps so that unsupervised training and self-supervised training are performed alternately until the generation effect of the generation network is stable, and the facial identity features of the output target facial image sample are significantly close to the facial identity features of the initial facial image sample, and the attribute features of the target facial image sample are significantly close to the attribute features of the template facial image sample. In other words, from a perceptual perspective, the generation network can generate a target facial image whose identity is consistent with the initial facial image sample and whose other features (posture, expression, lighting, background, etc.) are consistent with the facial image sample.

[0196] Among them, before encoding the initial input image and the initial template image, the terminal can align the facial feature points of the initial input image and the initial template image respectively, and locate the facial areas in the initial input image and the initial template image; capture facial images according to the facial areas located in the initial input image and the initial template image respectively, encode the facial image corresponding to the initial input image and the facial image corresponding to the initial template image, and finally obtain a fused facial image, and post it back to the facial area corresponding to the initial template image to obtain the final fusion result image.

[0197] Stage three: portrait enhancement.

[0198] refer to Figure 8CThe terminal first inputs the fusion result image into the beautification model, and performs deep beautification processing through the beautification model, including improving skin color and skin quality, and improving the naturalness of makeup, to obtain a beautification result image. The beautification result image is input into the portrait clarity enhancement model to improve the face resolution and obtain the final ID photo.

[0199] This application also provides an application scenario that applies the above-mentioned image processing method. In this application scenario, the terminal obtains an arbitrary portrait image and a standard template image (i.e., an image with standard posture, expression, and lighting) required to establish a facial model. Based on the arbitrary portrait image and the standard template image, a standard facial image is generated, thereby achieving better facial modeling. The specific implementation process can be referred to the description of the above embodiment, and this application will not repeat it here.

[0200] It should be understood that although Figure 1 The steps in the flowchart of -8 are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 -At least part of the steps in 8 may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0201] In one embodiment, Figure 9 As shown, an image data processing device 900 is provided. The device can be a software module or a hardware module, or a combination of the two to form a part of a computer device. The device specifically includes:

[0202] An image acquisition module 902 is configured to acquire an initial template image and an initial input image containing a facial region;

[0203] A feature acquisition module 904 is used to acquire facial state features of the initial template image and to acquire initial facial shape features of the initial input image;

[0204] A 3D reconstruction module 906 is configured to perform 3D facial reconstruction on the initial template image and the initial input image based on facial state features and initial facial shape features to obtain a 3D reconstructed facial image;

[0205] A two-dimensional projection module 908 is used to perform two-dimensional projection on the three-dimensional reconstructed facial image to obtain a two-dimensional reconstructed facial image;

[0206] An adjustment module 910 is configured to obtain a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image;

[0207] The fusion module 912 is used to obtain the facial identity features of the initial input image, and perform fusion processing based on the facial identity features and the target template image to obtain the target image.

[0208] In the above-mentioned image processing device, after obtaining the initial template image and the initial input image containing the facial area, the facial state features of the initial template image are further obtained, and the initial facial shape features of the initial input image are obtained. The initial template image and the initial input image are subjected to three-dimensional facial reconstruction according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image. The three-dimensional reconstructed facial image is then two-dimensionally projected to obtain a two-dimensional reconstructed facial image. The reconstructed facial shape features corresponding to the two-dimensional reconstructed facial image are obtained, and the facial area of ​​the initial template image is adjusted according to the reconstructed facial shape features to obtain a target template image. Finally, the facial identity features of the initial input image are obtained, and fusion processing is performed on the facial identity features and the initial input image to obtain a target image. This realizes the automatic generation of the target image, avoids the tedious operations of manual processing, and greatly improves the efficiency of image processing.

[0209] Furthermore, since the fusion is performed based on the facial identity features of the initial input image and the target template image, and the facial shape features of the target template image match the facial shape features of the initial input image, the target image finally obtained is similar to the facial identity features of the initial input image and similar to the facial shape features of the initial input image, thereby ensuring the consistency of the facial identity features of the target image and the initial input image, and at the same time ensuring the subjective similarity between the target image and the initial input image.

[0210] In one embodiment, the feature acquisition module is used to: obtain three-dimensional face model data from a pre-established three-dimensional face model database; the three-dimensional face model data includes a three-dimensional shape basis set and a three-dimensional expression basis set; obtain the shape weight coefficient of the initial input image corresponding to each three-dimensional shape basis, and determine the shape weight coefficient as the initial facial shape feature of the initial input image; obtain the expression weight coefficient of the initial template image corresponding to each three-dimensional expression basis, and determine the expression weight coefficient as the facial state feature of the initial template image.

[0211] In one embodiment, the three-dimensional reconstruction module is used to: determine the three-dimensional reconstructed face shape based on each three-dimensional shape base and the determined shape weight coefficient of each three-dimensional shape base; determine the three-dimensional reconstructed face expression based on each three-dimensional expression base and the determined expression weight coefficient of each three-dimensional expression base; and generate a three-dimensional reconstructed facial image based on the three-dimensional reconstructed face shape and the three-dimensional reconstructed face expression.

[0212] In one embodiment, the adjustment module is used to: triangulate the facial area of ​​the initial template image to obtain multiple triangular facets corresponding to the initial template image; and deform each triangular facet according to the reconstructed facial shape features to obtain the target template image.

[0213] In another embodiment, the adjustment module is configured to: perform pixel resampling on the facial region of the initial template image according to the reconstructed facial shape features; and determine the target template image according to the pixel resampling result.

[0214] In one embodiment, the above-mentioned device also includes a beautification module, which is used to: input the target image into a trained beautification model, and perform beautification processing on the target image through the beautification model to obtain a beautified image; the beautification model is obtained through supervised training of beautification training samples; the beautification training samples include the original image and the beautification image corresponding to the original image; and obtain the beautification image output by the beautification model.

[0215] In one embodiment, the above-mentioned device also includes a clarity enhancement module, which is used to: input the target image into a trained clarity enhancement model, and perform clarity enhancement processing on the target image through the clarity enhancement model to obtain a clear image; the clarity enhancement model is obtained by supervised training through clear training samples; the clear training samples include the original image and the blurred image obtained by degrading the clarity of the original image; and obtain the clear image output by the clarity enhancement model.

[0216] In one embodiment, a fusion module is used to encode the initial input image and the target template image respectively to obtain the facial identity features of the initial input image and the target attribute features of the target template image; fuse the facial identity features and the target attribute features to obtain the target features; decode the target features to obtain the target image; the target image matches the facial identity features of the initial input image and the target attribute features of the target template image.

[0217] In one embodiment, the fusion module is also used to encode the target template image through the attribute feature encoding model to obtain the attribute features of the target template image; fuse the facial identity features and the target attribute features through the feature fusion model to obtain the target features; decode the target features through the decoding model to obtain the target image; wherein, the feature fusion model, the decoding model and the attribute feature encoding model are obtained by alternating the use of unsupervised image samples and self-supervised image samples for joint training.

[0218] In one embodiment, the attribute feature encoding model, the feature fusion model and the decoding model are included in the generation network; the above-mentioned device also includes a training module for obtaining unsupervised image samples and self-supervised image samples; the unsupervised image samples include a first initial facial image sample and a first template facial image sample; the first initial facial image sample and the first template facial image sample are different image samples; the self-supervised image samples include a second initial facial image sample and a second template facial image sample; the second initial facial image sample and the second template facial image sample are the same image samples; the generation network is unsupervisedly trained according to the unsupervised image samples, and the model parameters of the attribute feature encoding model, the feature fusion model and the decoding model are adjusted; the generation network is self-supervisedly trained according to the self-supervised image samples, and the model parameters of the attribute feature encoding model, the feature fusion model and the decoding model are adjusted; the step of performing unsupervised training on the generation network according to the unsupervised image samples is repeated so that the unsupervised training and the self-supervised training are performed alternately until the training is terminated when the training stop condition is met.

[0219] In one embodiment, the generation network also includes a recognition feature encoding model, and the training module is further used to: encode the first initial facial image sample using the recognition feature encoding model to obtain facial identity features of the first initial facial image sample; encode the first template facial image sample using the attribute feature encoding model to obtain attribute features of the first template facial image sample; input the facial identity features of the first initial facial image sample and the attribute features of the first template facial image sample into the feature fusion model and the decoding model in sequence to obtain a first target facial image sample; encode the first target facial image sample using the recognition feature encoding model and the attribute feature encoding model respectively to obtain facial identity features and attribute features of the first target facial image sample; obtain a discriminant network, use at least one of the first initial facial image sample and the first template facial image sample as a positive sample of the discriminant network, and use the first target facial image sample as a negative sample of the discriminant network; and adjust model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model based on the discriminant loss of the discriminant network, the difference in facial identity features between the first initial facial image sample and the first target facial image sample, and the difference in attribute features between the first template facial image sample and the first target facial image sample.

[0220] In one embodiment, the training module is further configured to: encode the second initial facial image sample using the recognition feature encoding model to obtain facial identity features of the second initial facial image sample; encode the second template facial image sample using the attribute feature encoding model to obtain attribute features of the second template facial image sample; input the facial identity features of the second initial facial image sample and the attribute features of the second template facial image sample into the feature fusion model and the decoding model in sequence to obtain a second target facial image sample; encode the second target facial image sample using the recognition feature encoding model and the attribute feature encoding model respectively to obtain facial identity features and attribute features of the second target facial image sample; use at least one of the second initial facial image sample and the second template facial image sample as a positive sample of the discriminant network, and use the second target facial image sample as a negative sample of the discriminant network; and adjust model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model based on the discriminant loss of the discriminant network, the pixel difference between the second target facial image sample and the second initial facial image sample, the difference in facial identity features between the second initial facial image sample and the second target facial image sample, and the difference in attribute features between the second template facial image sample and the second target facial image sample.

[0221] For the specific definition of the image processing device, please refer to the definition of the image processing method above and will not be repeated here. Each module in the above-mentioned image processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0222] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 10As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an image data processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0223] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0224] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0225] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0226] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of each of the above-described method embodiments.

[0227] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0228] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0229] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. An image processing method, characterized in that: The method comprises: Obtain an initial template image and an initial input image containing a facial region; Acquiring facial state features of the initial template image and acquiring initial facial shape features of the initial input image; Performing three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features to obtain a three-dimensional reconstructed facial image; Projecting the three-dimensional spatial coordinates of the three-dimensional reconstructed facial image onto a projection plane determined according to facial state features of the initial template image to obtain a two-dimensional reconstructed facial image; Acquire a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image; The facial identity features of the initial input image are acquired, and a fusion process is performed based on the facial identity features and the target template image to obtain a target image.

2. The method according to claim 1, characterized in that The step of obtaining facial state features of the initial template image and initial facial shape features of the initial input image includes: Acquire three-dimensional face model data from a pre-established three-dimensional face model database; the three-dimensional face model data includes a three-dimensional shape basis set and a three-dimensional expression basis set; Obtaining shape weight coefficients corresponding to the three-dimensional shape bases of the initial input image, and determining the shape weight coefficients as initial facial shape features of the initial input image; The expression weight coefficients corresponding to the three-dimensional expression bases of the initial template image are obtained, and the expression weight coefficients are determined as the facial state features of the initial template image.

3. The method according to claim 2, characterized in that The performing three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state feature and the initial facial shape feature to obtain a three-dimensional reconstructed facial image comprises: Determining a three-dimensional reconstructed face shape based on each three-dimensional shape basis and the determined shape weight coefficients of each three-dimensional shape basis; Determining a three-dimensional reconstructed facial expression based on each three-dimensional expression base and the determined expression weight coefficient of each three-dimensional expression base; A three-dimensional reconstructed facial image is generated based on the three-dimensional reconstructed facial shape and the three-dimensional reconstructed facial expression.

4. The method according to claim 1, wherein The step of adjusting the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image comprises: triangulate the facial region of the initial template image to obtain a plurality of triangular facets corresponding to the initial template image; Deformation processing is performed on each of the triangular facets according to the reconstructed facial shape features to obtain a target template image.

5. The method according to claim 1, wherein The step of adjusting the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image comprises: resampling pixels of the facial region of the initial template image according to the reconstructed facial shape features; The target template image is determined based on the pixel resampling result.

6. The method according to claim 1, wherein The method further comprises: Inputting the target image into a trained beautification model, and performing beautification processing on the target image using the beautification model to obtain a beautification image; the beautification model is obtained by supervised training using beautification training samples; the beautification training samples include the original image and the beautification image corresponding to the original image; Obtain a beautified image output by the beautification model.

7. The method according to claim 1, characterized in that The method further comprises: Inputting the target image into a trained clarity enhancement model, and performing clarity enhancement processing on the target image using the clarity enhancement model to obtain a clear image; the clarity enhancement model is obtained by supervised training using clear training samples; the clear training samples include an original image and a blurred image obtained by degrading the clarity of the original image; Obtain a clear image output by the clarity enhancement model.

8. The method according to claim 1, characterized in that The acquiring of facial identity features of the initial input image and performing fusion processing with the target template image according to the facial identity features to obtain a target image comprises: Encoding the initial input image and the target template image respectively to obtain facial identity features of the initial input image and target attribute features of the target template image; Fusing the facial identity features and the target attribute features to obtain target features; The target features are decoded to obtain a target image; the target image matches the facial identity features of the initial input image and matches the target attribute features of the target template image.

9. The method according to claim 8, characterized in that Said encoding of the initial input image and the target template image respectively to obtain the facial identity features of the initial input image and the target attribute features of the target template image comprises: Encoding the target template image through an attribute feature encoding model to obtain attribute features of the target template image; The target feature obtained by fusing the facial identity feature and the target attribute feature includes: The facial identity feature and the target attribute feature are fused by a feature fusion model to obtain a target feature; The decoding of the target feature to obtain a target image includes: Decoding the target features through a decoding model to obtain a target image; The feature fusion model, the decoding model and the attribute feature encoding model are obtained by jointly training by alternately using unsupervised image samples and self-supervised image samples.

10. The method according to claim 9, characterized in that The attribute feature encoding model, the feature fusion model, and the decoding model are included in a generation network; the training steps of the generation network include: Obtaining an unsupervised image sample and a self-supervised image sample; the unsupervised image sample includes a first initial facial image sample and a first template facial image sample; the first initial facial image sample and the first template facial image sample are different image samples; the self-supervised image sample includes a second initial facial image sample and a second template facial image sample; the second initial facial image sample and the second template facial image sample are the same image sample; Performing unsupervised training on the generative network according to the unsupervised image samples, and adjusting model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model; Performing self-supervisory training on the generative network according to the self-supervised image samples, and adjusting model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model; The step of performing unsupervised training on the generative network according to the unsupervised image samples is repeated so that the unsupervised training and the self-supervised training are performed alternately until the training is terminated when a training stop condition is met.

11. The method according to claim 10, characterized in that The generation network further includes a recognition feature encoding model, and the performing of unsupervised training on the generation network according to the unsupervised image sample and adjusting model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model include: encoding the first initial facial image sample using the identification feature encoding model to obtain facial identity features of the first initial facial image sample; Encoding the first template facial image sample using the attribute feature encoding model to obtain attribute features of the first template facial image sample; Inputting the facial identity feature of the first initial facial image sample and the attribute feature of the first template facial image sample into the feature fusion model and the decoding model in sequence to obtain a first target facial image sample; Encoding the first target facial image sample using the identification feature encoding model and the attribute feature encoding model respectively to obtain facial identity features and attribute features of the first target facial image sample; Obtaining a discriminant network, using at least one of the first initial facial image sample and the first template facial image sample as a positive sample of the discriminant network, and using the first target facial image sample as a negative sample of the discriminant network; Based on the discriminant loss of the discriminant network, the difference in facial identity features between the first initial facial image sample and the first target facial image sample, and the difference in attribute features between the first template facial image sample and the first target facial image sample, model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted.

12. The method according to claim 10, characterized in that The performing self-supervisory training on the generative network according to the self-supervised image samples and adjusting the model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model includes: encoding the second initial facial image sample using a recognition feature encoding model to obtain a facial identity feature of the second initial facial image sample; Encoding the second template facial image sample using the attribute feature encoding model to obtain attribute features of the second template facial image sample; inputting the facial identity features of the second initial facial image sample and the attribute features of the second template facial image sample into the feature fusion model and the decoding model in sequence to obtain a second target facial image sample; Encoding the second target facial image sample using the identification feature encoding model and the attribute feature encoding model respectively to obtain facial identity features and attribute features of the second target facial image sample; using at least one of the second initial facial image sample and the second template facial image sample as a positive sample of a discriminant network, and using the second target facial image sample as a negative sample of the discriminant network; Based on the discriminative loss of the discriminative network, the pixel difference between the second target facial image sample and the second initial facial image sample, the difference in facial identity features between the second initial facial image sample and the second target facial image sample, and the difference in attribute features between the second template facial image sample and the second target facial image sample, the model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted.

13. An image processing device, characterized in that: The device comprises: An image acquisition module is used to acquire an initial template image containing a facial region and an initial input image; A feature acquisition module, configured to acquire facial state features of the initial template image and initial facial shape features of the initial input image; a three-dimensional reconstruction module, configured to perform three-dimensional facial reconstruction on the initial template image and the initial input image according to the facial state features and the initial facial shape features, to obtain a three-dimensional reconstructed facial image; a two-dimensional projection module, configured to project the three-dimensional spatial coordinates of the three-dimensional reconstructed facial image onto a projection plane determined according to facial state features of the initial template image, to obtain a two-dimensional reconstructed facial image; an adjustment module, configured to obtain a reconstructed facial shape feature corresponding to the two-dimensional reconstructed facial image, and adjust the facial region of the initial template image according to the reconstructed facial shape feature to obtain a target template image; The fusion module is used to obtain the facial identity features of the initial input image, and perform fusion processing based on the facial identity features and the target template image to obtain a target image.

14. The device according to claim 13, characterized in that The feature acquisition module is further used for: Acquiring three-dimensional face model data from a pre-established three-dimensional face model database; the three-dimensional face model data including a three-dimensional shape basis set and a three-dimensional expression basis set; obtaining shape weight coefficients corresponding to the three-dimensional shape basis of the initial input image, and determining the shape weight coefficients as initial facial shape features of the initial input image; The expression weight coefficients corresponding to the three-dimensional expression bases of the initial template image are obtained, and the expression weight coefficients are determined as the facial state features of the initial template image.

15. The device according to claim 14, characterized in that The three-dimensional reconstruction module is also used for: Based on each three-dimensional shape base and the determined shape weight coefficients of each three-dimensional shape base, the three-dimensional reconstructed facial shape is determined; based on each three-dimensional expression base and the determined expression weight coefficients of each three-dimensional expression base, the three-dimensional reconstructed facial expression is determined; based on the three-dimensional reconstructed facial shape and the three-dimensional reconstructed facial expression, a three-dimensional reconstructed facial image is generated.

16. The device according to claim 13, characterized in that The adjustment module is further configured to: The facial region of the initial template image is triangulated to obtain a plurality of triangular facets corresponding to the initial template image; and each triangular facet is deformed according to the reconstructed facial shape feature to obtain a target template image.

17. The device according to claim 13, characterized in that The adjustment module is further configured to: Pixel resampling is performed on the facial region of the initial template image according to the reconstructed facial shape feature; and a target template image is determined according to the pixel resampling result.

18. The device according to claim 13, characterized in that The device also includes a beauty module for: The target image is input into a trained beautification model, and the target image is beautified by the beautification model to obtain a beautification image; the beautification model is obtained by supervised training with beautification training samples; the beautification training samples include the original image and the beautification image corresponding to the original image; and the beautification image output by the beautification model is obtained.

19. The device according to claim 13, characterized in that The device further comprises a definition enhancement module, configured to: The target image is input into a trained clarity enhancement model, and the target image is subjected to clarity enhancement processing by the clarity enhancement model to obtain a clear image; the clarity enhancement model is obtained by supervised training using clear training samples; the clear training samples include the original image and a blurred image obtained by degrading the clarity of the original image; and the clear image output by the clarity enhancement model is obtained.

20. The device according to claim 13, wherein The fusion module is also used to Encoding the initial input image and the target template image respectively to obtain facial identity features of the initial input image and target attribute features of the target template image; Fusing the facial identity features and the target attribute features to obtain target features; The target features are decoded to obtain a target image; the target image matches the facial identity features of the initial input image and matches the target attribute features of the target template image.

21. The device according to claim 20, characterized in that The fusion module is further configured to: Encoding the target template image through an attribute feature encoding model to obtain the attribute features of the target template image; fusing the facial identity features and the target attribute features through a feature fusion model to obtain target features; The target features are decoded by a decoding model to obtain a target image; wherein the feature fusion model, the decoding model and the attribute feature encoding model are obtained by jointly training by alternately using unsupervised image samples and self-supervised image samples.

22. The device according to claim 21, characterized in that The attribute feature encoding model, the feature fusion model and the decoding model are included in the generation network; the device also includes a training module, which is used to: obtain unsupervised image samples and self-supervised image samples; the unsupervised image samples include a first initial facial image sample and a first template facial image sample; the first initial facial image sample and the first template facial image sample are different image samples; the self-supervised image samples include a second initial facial image sample and a second template facial image sample; the second initial facial image sample and the second template facial image sample are the same image samples; the generation network is unsupervisedly trained according to the unsupervised image samples, and the model parameters of the attribute feature encoding model, the feature fusion model and the decoding model are adjusted; the generation network is self-supervisedly trained according to the self-supervised image samples, and the model parameters of the attribute feature encoding model, the feature fusion model and the decoding model are adjusted; and the step of performing unsupervised training on the generation network according to the unsupervised image samples is repeated so that the unsupervised training and the self-supervised training are performed alternately until the training is terminated when the training stop condition is met.

23. The device according to claim 22, characterized in that The generation network also includes a recognition feature encoding model, and the training module is further used to: Encode the first initial facial image sample using the identification feature encoding model to obtain facial identity features of the first initial facial image sample; encode the first template facial image sample using the attribute feature encoding model to obtain attribute features of the first template facial image sample; input the facial identity features of the first initial facial image sample and the attribute features of the first template facial image sample into the feature fusion model and the decoding model in sequence to obtain a first target facial image sample; encode the first target facial image sample using the identification feature encoding model and the attribute feature encoding model respectively to obtain facial identity features and attribute features of the first target facial image sample; obtain a discriminant network, use at least one of the first initial facial image sample and the first template facial image sample as a positive sample of the discriminant network, and use the first target facial image sample as a negative sample of the discriminant network; Based on the discriminant loss of the discriminant network, the difference in facial identity features between the first initial facial image sample and the first target facial image sample, and the difference in attribute features between the first template facial image sample and the first target facial image sample, model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted.

24. The device according to claim 22, characterized in that The training module is also used to: Encoding the second initial facial image sample using the recognition feature coding model to obtain facial identity features of the second initial facial image sample; encoding the second template facial image sample using the attribute feature coding model to obtain attribute features of the second template facial image sample; inputting the facial identity features of the second initial facial image sample and the attribute features of the second template facial image sample into the feature fusion model and the decoding model in sequence to obtain a second target facial image sample; encoding the second target facial image sample using the recognition feature coding model and the attribute feature coding model respectively to obtain facial identity features and attribute features of the second target facial image sample; using at least one of the second initial facial image sample and the second template facial image sample as a positive sample of a discriminant network, and using the second target facial image sample as a negative sample of the discriminant network; Based on the discriminative loss of the discriminative network, the pixel difference between the second target facial image sample and the second initial facial image sample, the difference in facial identity features between the second initial facial image sample and the second target facial image sample, and the difference in attribute features between the second template facial image sample and the second target facial image sample, the model parameters of the attribute feature encoding model, the feature fusion model, and the decoding model are adjusted.

25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

26. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

27. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Face three-dimensional image generation method, device and readable medium

    CN109377544A

  • Image processing method and device, image beautifying method and device and storage medium

    CN110070484A

  • Image processing method and device, model training method and device, computer equipment and storage medium

    CN111401216A