Image Processing Method, Apparatus, Device, and Storage Medium

Through the eye position editing solution based on the generative adversarial network, the editing scene problem in the prior art that cannot adapt to single image input is solved, flexible editing of eye position in face images is achieved, and the authenticity of editing processing is improved.

CN114998483BActive Publication Date: 2025-06-20XIAMEN MEITUZHIJIA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210682932.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-06-20
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

The prior art cannot adapt to the eye position editing scenarios of single image input, and it is necessary to pre-record the eye state at different positions.

Method used

Using an eye position editing scheme based on a generative adversarial network, by acquiring the single face image to be processed and the target position, the original eye area image and the mask image of the eyeball at the target position are determined, and input it to the pre-trained target image generator model to generate the target eye area image, and finally fuse it into the face image.

Benefits of technology

It realizes editing of the eyeball position in a reasonable range in a face image, solves the problem of editing scenes that cannot adapt to the input of a single image, and maintains the consistency between other areas except the eyes and the original image, improving the authenticity of editing processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998483B_ABST
    Figure CN114998483B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, apparatus, device and storage medium, relating to the technical field of image processing. The method includes: obtaining a face image to be processed and a target position input by a user; determining a first image and a second image according to the face image and the target position; inputting the first image and the second image into a pre-trained target image generator model to obtain a third image; and performing fusion processing on the eye region image in the face image according to the face image and the third image. Compared with the prior art, the solution proposed in the present application can realize the editing of the eyeballs at any position within a reasonable range in the face image only relying on a single face image, effectively solving the technical problem that the prior art cannot adapt to the editing scenario of single-picture input; at the same time, the solution proposed in the present application focuses on the generation of the eye region image, so as to maintain the consistency of other regions except the eyes with the original image, making the editing effect more realistic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art

[0002] Nowadays, with the development of the Internet and the popularization of digital products such as mobile phones, most users will share their life photos through social network platforms. Usually, in order to pursue a higher-quality image beautification effect, people will use some beauty retouching software to beautify the face images taken, so as to improve the overall beauty of the face images. In particular, the eye features in the face image are one of the key beautification areas.

[0003] Currently, in the existing eye position editing solutions for face images, mainly the eye tracking technology is used to pre-record the movement trajectory of the eyes and the eye states at different positions. During actual editing, the eye states at the corresponding positions are taken out and pasted to the current eye position to obtain the target image after eye position editing.

[0004] However, in order to ensure that the eye position editing effect is close to the real state, the existing such eye position editing solutions need to obtain the eye states at different positions in advance, which leads to the inability to adapt to the editing scenario of single-image input. Summary of the Invention

[0005] The purpose of the present application is to provide an image processing method, apparatus, device, and storage medium for the above-mentioned deficiencies in the existing technology, so as to solve the problem that the existing technology cannot adapt to the editing scenario of single-image input when performing eye position editing processing.

[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:

[0007] In a first aspect, an embodiment of the present application provides an image processing method, including:

[0008] Obtain a face image to be processed and a target position input by a user, where the target position is the position where the eyes in the face image are to be moved;

[0009] Determine a first image and a second image according to the face image and the target position, where the first image is the original eye region image in the face image, and the second image is the mask image in which the eyes in the original eye region image are at the target position;

[0010] Input the first image and the second image into a pre-trained target image generator model to obtain a third image, where the third image is the target eye region image, and the eyes in the third image are at the target position;

[0011] Based on the face image and the third image, perform fusion processing on the eye region image in the face image to obtain a processed target image.

[0012] Optionally, the target image generator model is trained in the following manner:

[0013] Obtain the original input images and original mask images of multiple face samples, where the original input images are the original eye region images in the face samples, and the original mask images are the mask images of the actual positions of the eyeballs in the face images;

[0014] Based on the original input images and original mask images of the multiple face images, perform iterative training on the initial image generator model to obtain the target image generator model.

[0015] Optionally, the performing iterative training on the initial image generator model based on the original input images and original mask images of the multiple face samples to obtain the target image generator model includes:

[0016] Input the original input image of the first face sample and the original mask image of the second face sample into the initial image generator model to obtain the generated image of the second face sample output by the initial image generator model;

[0017] Input the generated image of the second face sample and the original mask image of the first face sample into the initial image generator model to obtain the restored image of the first face sample output by the initial image generator model;

[0018] Calculate the loss value of the initial image generator model according to the original input image of the first face sample and the restored image of the first face sample, and correct the parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a new initial image generator model;

[0019] Input the generated image of the second face sample and the restored image of the first face sample into the initial mask generator model respectively to obtain the restored mask image of the second face sample and the generated mask image of the first face sample;

[0020] Calculate the loss value of the initial mask generator model according to the restored mask image of the second face sample and the original mask image of the second face sample, and correct the parameters of the initial mask generator model according to the loss value of the initial mask generator model to obtain a new initial mask generator model;

[0021] Execute the above steps in a loop until the loss value of the initial image generator model meets a preset first loss threshold and the loss value of the initial mask generator model meets a preset second loss threshold. Then, use the initial image generator model that meets the first loss threshold as the target image generator model and the initial mask generator model that meets the second loss threshold as the target mask generator model.

[0022] Optionally, when the loss value of the initial image generator model meets a preset first loss threshold and the loss value of the initial mask generator model meets a preset second loss threshold, using the initial image generator model that meets the first loss threshold as the target image generator model and the initial mask generator model that meets the second loss threshold as the target mask generator model includes:

[0023] If the loss value of the initial image generator model meets the first loss threshold and the loss value of the initial mask generator model meets the second loss threshold, then input the original input image of the first face sample and the generated image of the second face sample into an image discriminator to determine whether the output result of the image discriminator is false;

[0024] If so, use the initial image generator model as the target image generator model;

[0025] Input the generated image of the second face sample and the original mask image of the second face sample into a mask discriminator to determine whether the output result of the mask discriminator is false;

[0026] If so, use the initial mask generator model as the target mask generator model.

[0027] Optionally, the determining the first image and the second image according to the face image includes:

[0028] Extract the eye feature information in the face image, where the eye feature information includes: the center point of the eyeball, the width and height of the eye region, the upper and lower eyelid points, the upper and lower vertex points of the eyeball, and the radius of the eyeball;

[0029] Calculate the vertex coordinates of the eye bounding box according to the center point of the eyeball and the width and height of the eye region;

[0030] Calculate a first affine transformation matrix according to the vertex coordinates of the eye bounding box and a preset post - cropped eye image, where the first affine transformation matrix is the affine transformation matrix from the original image to the local image, and the image size of the post - cropped eye image is a preset value;

[0031] Perform an affine transformation on the face image according to the first affine transformation matrix to obtain an initial image including the eye region;

[0032] Obtain the first image according to the upper and lower eyelid points and the initial image including the eye region;

[0033] Obtain the second image according to the target position and the position mapping relationship between the target position and the upper and lower vertices of the eyeball in the eye feature information.

[0034] Optionally, the performing a fusion process on the eye region image in the face image according to the face image and the third image to obtain a processed target image includes:

[0035] Use an eye segmentation model to perform eye segmentation on the first image and the third image respectively to obtain the original eyeball region in the first image and the eyeball region in the third image;

[0036] Perform Poisson fusion on the original eyeball region in the first image and the eyeball region in the third image to obtain a processed third image;

[0037] Perform a fusion process on the eye region image in the face image according to the face image and the processed third image to obtain a processed target image.

[0038] In a second aspect, an embodiment of the present application further provides an image processing device, where the device includes:

[0039] An acquisition module, configured to acquire a face image to be processed and a target position input by a user, where the target position is a position where the eyeball in the face image is to move to;

[0040] A processing module, configured to determine a first image and a second image according to the face image and the target position, where the first image is an original eye region image in the face image, and the second image is a mask image in which the eyeball in the original eye region image is at the target position; input the first image and the second image into a pre-trained target image generator model to obtain a third image, where the third image is a target eye region image, and the eyeball in the third image is located at the target position;

[0041] A fusion module, configured to perform a fusion process on the eye region image in the face image according to the face image and the third image to obtain a processed target image.

[0042] Optionally, the device further includes:

[0043] An acquisition module for acquiring the original input images and original mask images of multiple face samples, where the original input images are the original eye region images in the face images, and the original mask images are the mask images of the actual positions of the eyeballs in the face images;

[0044] A training module for iteratively training an initial image generator model based on the original input images and original mask images of the multiple face samples to obtain the target image generator model.

[0045] Optionally, the training module is further configured to:

[0046] Input the original input image of the first face sample and the original mask image of the second face sample into the initial image generator model to obtain the generated image of the second face sample output by the initial image generator model;

[0047] Input the generated image of the second face sample and the original mask image of the first face sample into the initial image generator model to obtain the restored image of the first face sample output by the initial image generator model;

[0048] Calculate the loss value of the initial image generator model according to the original input image of the first face sample and the restored image of the first face sample, and correct the parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a new initial image generator model;

[0049] Input the generated image of the second face sample and the restored image of the first face sample into the initial mask generator model respectively to obtain the restored mask image of the second face sample and the generated mask image of the first face sample;

[0050] Calculate the loss value of the initial mask generator model according to the restored mask image of the second face sample and the original mask image of the second face sample, and correct the parameters of the initial mask generator model according to the loss value of the initial mask generator model to obtain a new initial mask generator model;

[0051] Loop through the above steps until the loss value of the initial image generator model meets a preset first loss threshold and the loss value of the initial mask generator model meets a preset second loss threshold. Take the initial image generator model that meets the first loss threshold as the target image generator model and the initial mask generator model that meets the second loss threshold as the target mask generator model.

[0052] Optionally, the training module is further configured to:

[0053] If the loss value of the initial image generator model satisfies the first loss threshold and the loss value of the initial mask generator model satisfies the second loss threshold, the original input image of the first face sample and the generated image of the second face sample are input into an image discriminator to determine whether the output result of the image discriminator is false;

[0054] If so, the initial image generator model is used as the target image generator model;

[0055] The generated image of the second face sample and the original mask image of the second face sample are input into a mask discriminator to determine whether the output result of the mask discriminator is false;

[0056] If so, the initial mask generator model is used as the target mask generator model.

[0057] Optionally, the processing module is further configured to:

[0058] Extract eye feature information from the face image, where the eye feature information includes: the center point of the eyeball, the width and height of the eye region, the upper and lower eyelid points, the upper and lower vertex points of the eyeball, and the radius of the eyeball;

[0059] Calculate the vertex coordinates of the eye bounding box according to the center point of the eyeball and the width and height of the eye region;

[0060] Calculate a first affine transformation matrix according to the vertex coordinates of the eye bounding box and a preset cropped eye image, where the first affine transformation matrix is an affine transformation matrix from the original image to the local image, and the image size of the cropped eye image is a preset value;

[0061] Perform an affine transformation on the face image according to the first affine transformation matrix to obtain an initial image including the eye region;

[0062] Obtain the first image according to the upper and lower eyelid points and the initial image including the eye region;

[0063] Obtain the second image according to the target position and the position mapping relationship between the target position and the upper and lower vertex points of the eyeball in the eye feature information.

[0064] Optionally, the fusion module is further configured to:

[0065] Use an eye segmentation model to perform eye segmentation on the first image and the third image respectively to obtain the original eyeball region in the first image and the eyeball region in the third image;

[0066] Perform Poisson fusion on the original eyeball region in the first image and the eyeball region in the third image to obtain the processed third image;

[0067] According to the face image and the processed third image, perform fusion processing on the eye region image in the face image to obtain the processed target image.

[0068] In a third aspect, an embodiment of the present application further provides a processing unit device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the processing unit device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the method provided in the first aspect.

[0069] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is run by a processor, it executes the steps of the method provided in the first aspect.

[0070] The beneficial effects of the present application are:

[0071] An embodiment of the present application provides an image processing method, device, equipment, and storage medium. It mainly adopts an eyeball position editing scheme based on a generative adversarial network, only needing to obtain a single face image to be processed and the target position where the eyeball in the face image is to be moved; then, according to this single face image and the target position where the eyeball in the face image is to be moved, determine the original eye region image in the face image and the mask image where the eyeball in the original eye region image is in the target position; then, input the original eye region image in the face image and the mask image where the eyeball in the original eye region image is in the target position into a pre-trained target image generator model to obtain a target eye region image output by the target image generator model; finally, fuse the obtained target eye region image into the face image to obtain the processed target image; compared with the prior art, the solution proposed in the present application can achieve the editing of the eyeball position in the face image at any position within a reasonable range only relying on a single face image, effectively solving the technical problem in the prior art that it cannot adapt to the editing scenario of single-picture input; at the same time, the solution proposed in the present application focuses on the generation of the eye region image, thereby maintaining the consistency of other regions except the eyes with the original image, and further making the editing processing effect of the eyeball position in the face image more realistic while maximizing the retention of the original image features when adjusting the eye details. Description of the Drawings

[0072] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0073] Figure 1 Schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0074] Figure 2 Schematic flow diagram of an image processing method provided by an embodiment of the present application;

[0075] Figure 3 Schematic flow diagram of another image processing method provided by an embodiment of the present application;

[0076] Figure 4 Schematic flow diagram of yet another image processing method provided by an embodiment of the present application;

[0077] Figure 5 Schematic network structure diagram of an initial image generator model provided by an embodiment of the present application;

[0078] Figure 6 Flow framework of an image processing method provided by an embodiment of the present application Figure 1 ;

[0079] Figure 7 Schematic flow diagram of another image processing method provided by an embodiment of the present application;

[0080] Figure 8 Schematic network structure diagram of an image discriminator provided by an embodiment of the present application;

[0081] Figure 9 Flow framework of an image processing method provided by an embodiment of the present application Figure 2 ;

[0082] Figure 10 Schematic flow diagram of yet another image processing method provided by an embodiment of the present application;

[0083] Figure 11 Schematic diagram of a mask image in an image processing method provided by an embodiment of the present application;

[0084] Figure 12 Schematic flow diagram of another image processing method provided by an embodiment of the present application;

[0085] Figure 13 Schematic structural diagram of an image processing device provided by an embodiment of the present application. Detailed implementation manners

[0086] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.

[0087] In addition, the described embodiments are only some embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.

[0088] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated thereafter, but does not exclude adding other features.

[0089] First, before specifically describing the technical solutions provided by the present application, a brief description of the relevant background involved in the present application will be given.

[0090] Currently, in the existing eyeball position editing solutions for face images, mainly the eyeball tracking technology is used to pre-record the movement trajectory of the eyeball and the eyeball states at different positions. During actual editing, the eyeball states at the corresponding positions are taken out and pasted to the current eyeball position to obtain the target image after eyeball position editing.

[0091] However, in order to ensure that the eyeball position editing effect is close to the real state, the existing such eyeball position editing solutions need to obtain the eyeball states at different positions in advance, which results in the inability to adapt to the editing scenario of single-image input.

[0092] To solve the technical problems existing in the above-mentioned prior art, the present application proposes an image processing method. This method mainly adopts an eyeball position editing scheme based on a generative adversarial network, which only needs to obtain a single face image to be processed and the target position where the eyeballs in the face image are to be moved; then, based on this single face image and the target position where the eyeballs in the face image are to be moved, determine the original eye region image in the face image and the mask image where the eyeballs in the original eye region image are in the target position; then, input the original eye region image in the face image and the mask image where the eyeballs in the original eye region image are in the target position into a pre-trained target image generator model to obtain a target eye region image output by the target image generator model; finally, fuse the obtained target eye region image into the face image to obtain a processed target image. Compared with the prior art, the solution proposed in the present application can realize the editing of the eyeball position in a face image at any position within a reasonable range only relying on a single face image, effectively solving the technical problem in the prior art that it cannot adapt to the editing scenario of single-picture input.

[0093] At the same time, the solution proposed in the present application focuses on the generation of the eye region image, so as to maintain the consistency of other regions except the eyes with the original image, and further retain the original image features to the greatest extent while adjusting the eye details, making the editing effect of the eyeball position in the face image more realistic.

[0094] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application; this electronic device can be a processing device such as a computer or a server, etc., for implementing the image processing method provided by the present application. As Figure 1 shown, the electronic device includes: a processor 101 and a memory 102.

[0095] The processor 101 and the memory 102 are directly or indirectly electrically connected to realize data transmission or interaction. For example, electrical connection can be achieved through one or more communication buses or signal lines.

[0096] Among them, the processor 101 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 101 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0097] The memory 102 can be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), etc.

[0098] It can be understood that Figure 1 the described structure is only schematic, and the electronic device 100 may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 that shown in. Figure 1 Each component shown in can be implemented by hardware, software, or a combination thereof.

[0099] The memory 102 is used to store programs, and the processor 101 calls the programs stored in the memory 102 to execute the image processing method provided in the following embodiments.

[0100] The image processing method provided in this application and the corresponding beneficial effects will be described below through multiple embodiments.

[0101] Figure 2 FIG. is a schematic flowchart of an image processing method provided in an embodiment of this application. Optionally, the execution subject of this method can be an electronic device such as a server or a computer, which has data processing capabilities.

[0102] It should be understood that in other embodiments, the order of some steps of the image processing method can be interchanged according to actual needs, or some of the steps can also be omitted or deleted. As Figure 2 shown, this method includes:

[0103] S201. Obtain a face image to be processed and a target position input by the user.

[0104] Among them, the target position is the position where the eyeballs in the face image are to move to.

[0105] In this embodiment, the face image to be processed is an image containing complete face information, and it is necessary to adjust the current position of the eyeballs in the face image. For example, the current position of the eyeballs in the face image is point A, and the target position to be moved to is point B, that is, it is necessary to move the eyeballs in the face image from point A to point B to achieve the editing process of the position of the eyeballs in the face image.

[0106] S202. Determine a first image and a second image according to the face image and the target position.

[0107] Among them, the first image is the original eye region image in the face image, and the second image is the mask image in which the eyeballs in the original eye region image are at the target position.

[0108] It should be understood that in this embodiment, the first image is the original eye region image extracted from a complete face image, that is, only the eye region in the face image is processed specifically, so that the consistency of other regions except the eyes with the original image can be maintained.

[0109] In this embodiment, for example, according to the eye features in the face image, a local eye image, that is, the first image, can be cropped from the face image.

[0110] Optionally, the mask image in which the eyeballs in the original eye region image are at the target position, that is, the third image, can be obtained by combining the target position input by the user and the first image. Among them, the third image is the mask image in which the eyeballs in the original eye region image are at the target position.

[0111] S203. Input the first image and the second image into a pre-trained target image generator model to obtain a third image.

[0112] Among them, the third image is the target eye region image, and the eyeballs in the third image are at the target position.

[0113] Exemplarily, for example, the pre-trained target image generator model can be a network model obtained by training a pre-collected original sample image through a deep neural network model.

[0114] In this embodiment, specifically, the first image and the second image are input into the pre-trained target image generator model to adjust the current position of the eyeballs in the face image, that is, to move the eyeballs from the current position to the target position, and the third image output by the target image generator model is obtained. Among them, the third image is the target eye region image (that is, both the third image and the first image are local eye images), and the eyeballs in the third image are at the target position, realizing the editing process of the position of the eyeballs in the face image.

[0115] S204. Perform a fusion process on the eye region image in the face image according to the face image and the third image to obtain a processed target image.

[0116] In this embodiment, in order to obtain a complete image after editing the position of the eyeball, it is also necessary to fuse the above-obtained third image into the face image. Specifically, the eye region image in the face image can be combined with the third image to perform a fusion process on the eye region image in the face image to obtain a processed target image, and thus the final effect after editing the position of the eyeball can be obtained.

[0117] In summary, the embodiment of the present application provides an image processing method, which mainly adopts an eyeball position editing scheme based on a generative adversarial network. Only a single face image to be processed and the target position where the eyeball in the face image is to be moved need to be obtained; then, according to this single face image and the target position where the eyeball in the face image is to be moved, the original eye region image in the face image and the mask image of the eyeball in the original eye region image at the target position are determined; then, the original eye region image in the face image and the mask image of the eyeball in the original eye region image at the target position are input into a pre-trained target image generator model together to obtain a target eye region image output by the target image generator model; finally, the obtained target eye region image is fused into the face image to obtain a processed target image. Compared with the prior art, the scheme proposed in the present application can realize the editing of the eyeball at any position within a reasonable range in the face image only relying on a single face image, effectively solving the technical problem that the prior art cannot adapt to the editing scenario of single-picture input; at the same time, the scheme proposed in the present application focuses on the generation of the eye region image, so as to maintain the consistency of other regions except the eyes with the original image, and further make the editing effect of the eyeball position in the face image more real while maximizing the retention of the original image features when adjusting the eye details.

[0118] The following embodiments will specifically explain how to train the target image generator model.

[0119] Optionally, as shown in Figure 3 The above pre-trained target image generator model can be trained in the following manner:

[0120] S301. Obtain the original input images and original mask images of multiple face samples.

[0121] Among them, the original input image is the original eye region image in the face image, and the original mask image is the mask image of the actual position of the eyeball in the face image.

[0122] For example, the face samples include: the first face sample A and the second face sample B. The original input image of the first face sample is Input_imageA, and the original mask image is Guided_maskA. The original input image of the second face sample is Input_imageB, and the original mask image is Guided_maskB.

[0123] For example, the face image of the first face sample is imageA. Among them, the original input image Input_imageA of the first face sample is the eye region image in the first face image, and the original mask image Guided_maskA of the first face sample is the image of the actual position of the eyeball in the first face image.

[0124] S302. Based on the original input images and original mask images of multiple face samples, perform iterative training on the initial image generator model to obtain a target image generator model.

[0125] It should be noted that when performing iterative training on the initial image generator model, randomly sample two groups of original input images Input_image and original mask images Guided_mask in each batch. After swapping the two groups of original mask images Guided_mask, recombine them with the original input images Input_image as the input of the initial image generator model.

[0126] For example, take the original input image Input_imageA of the first face sample and the original mask image Guided_maskB of the second face sample as a set of paired data and input it into the initial image generator model; and take the original input image Input_imageB of the second face sample and the original input image Guided_maskA of the second face sample as another set of paired data, input it into the initial image generator model, perform iterative training on the initial image generator model, and calculate the loss value of the initial image generator model based on the output result of the initial image generator model. Continuously correct the parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a target image generator model.

[0127] The process of calculating the loss value of the initial image generator model is iteratively executed until the initial image generator model is gradually corrected to obtain a target image generator model.

[0128] The following embodiments will specifically explain how to perform iterative training on the initial image generator model based on the original input images and original mask images of multiple face samples to obtain a target image generator model.

[0129] Optionally, refer toFigure 4 , Figure 5 and Figure 6 As shown in ,

[0130] , the above step S302 includes:

[0130] S401. Input the original input image of the first face sample and the original mask image of the second face sample into the initial image generator model to obtain the generated image of the second face sample output by the initial image generator model.

[0131] Among them, the generated image of the second face sample is denoted as Gen_imageB.

[0132] Before training the initial image generator model, first introduce the network structure of the initial image generator model in combination with Figure 5

[0133] Figure 5 Refer to Figure 5As shown, in this embodiment, the image generator model includes twenty network layers. Among them, the first and second network layers have 64 channels and a stride of 2. The first network layer is a convolutional layer with a default convolutional kernel size of 3*3; the second network layer is an activation layer, and the leaky_relu activation function is selected; the third, fourth, and fifth network layers have 128 channels and a stride of 2. Among them, the third network layer is a convolutional layer with a default convolutional kernel size of 3*3; the fourth network layer is a batch normalization layer, and the batch normalization operation BN along the batch dimension is adopted; the fifth network layer is an activation layer, and the leaky_relu activation function is selected; the sixth, seventh, and eighth network layers have 256 channels and a stride of 2. The sixth network layer is a convolutional layer with a default convolutional kernel size of 3*3; the seventh layer is a batch normalization layer; the eighth network layer is an activation layer, and the leaky_relu activation function is selected; the ninth, tenth, eleventh, twelfth, and thirteenth network layers have 128 channels and a stride of 1. The ninth network layer is a convolutional layer with a default convolutional kernel size of 3*3, the tenth layer is a batch normalization layer, the eleventh network layer is an activation layer, and the Relu activation function is selected; the twelfth network layer is a convolutional layer with a default convolutional kernel size of 3*3, and the thirteenth network layer is a batch normalization layer; the fourteenth, fifteenth, and sixteenth network layers have 128 channels and a stride of 2. The fourteenth network layer is a transposed convolutional layer with a default convolutional kernel size of 3*3, the fifteenth layer is a batch normalization layer, and the sixteenth network layer is an activation layer, and the Relu activation function is selected; the seventeenth, eighteenth, and nineteenth network layers have 64 channels and a stride of 2. The sixteenth network layer is a transposed convolutional layer with a default convolutional kernel size of 3*3, the seventeenth layer is a batch normalization layer, and the eighteenth network layer is an activation layer, and the Relu activation function is selected; the twentieth network layer has 3 channels and a stride of 1. The twentieth network layer is a transposed convolutional layer with a default convolutional kernel size of 3*3; the twenty-first network layer is an activation layer with the Tanh activation function.

[0134] In addition, the size of the input feature map of the entire network is 6*128*128, the size of the output feature map of the second network layer is 64*128*128, the size of the output feature map of the fifth network layer is 128*64*64, the size of the output feature map of the eighth network layer is 256*32*32, the size of the output feature map of the thirteenth network layer is 256*32*32, and the size of the output feature map of the sixteenth network layer is 128*128*128; the size of the output feature map of the nineteenth network layer is 3*128*128, and the size of the output feature map of the entire network is 3*128*128.

[0135] Similarly, the network structure of the mask generator model is the same as that of the image generator model, and will not be elaborated here.

[0136] Refer to Figure 6 As shown, in the image-to-image cycle, the initial image generator model GI is different from the traditional encoder-decoder mode. The initial image generator model GI adopts a symmetric Unet structure, and the output of the corresponding downsampling layer of the input is concatenated before each convolutional layer in the upsampling stage of Unet, that is, a lateral skip connection structure is added, which is beneficial to capturing image features at different scales.

[0137] Continue to refer to Figure 6 As shown, the original input image of the first face sample is Input_imageA, and the original mask image of the second face sample is Guided_maskB. They can be concatenated to obtain the 6-channel input InputGI of the initial image generator model, and then input into the initial image generator model to obtain the generated image Gen_imageB of the second face sample output by the initial image generator model.

[0138] S402. Input the generated image of the second face sample and the original mask image of the first face sample into the initial image generator model to obtain the restored image of the first face sample output by the initial image generator model.

[0139] Among them, the restored image of the first face sample is denoted as Recover_imageA.

[0140] At the same time, the generated image Gen_imageB of the second face sample and the original mask image Guided_maskA of the first face sample are concatenated and then input into the initial image generator model to obtain the restored image Recover_imageA of the first face sample output by the initial image generator model.

[0141] S403. Calculate the loss value of the initial image generator model according to the original input image of the first face sample and the restored image of the first face sample, and correct the parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a new initial image generator model.

[0142] Among them, the loss value of the initial image generator model mainly uses l1-Loss to evaluate the difference between the generated result and the target image. The calculation method of l1-Loss is relatively simple, which is obtained by taking the absolute value of the difference between the pixel values at the corresponding positions.

[0143] Optionally, the loss value of the initial image generator model (i.e., the cycle consistency loss value of the initial image generator model, denoted as IC) can be calculated based on the original input image Input_imageA of the first face sample and the restored image Recover_imageA of the first face sample. Alternatively, the loss value of the initial image generator model can also be calculated based on the generated image Gen_imageB of the second face sample and the original input image Input_imageB of the second face sample.

[0144] Then, based on the loss value of the initial image generator model, the parameters of the loss value of the initial image generator model are corrected to obtain a new initial image generator model.

[0145] S404: Input the generated image of the second face sample and the restored image of the first face sample into the initial mask generator model respectively to obtain the restored mask image of the second face sample and the generated mask image of the first face sample.

[0146] It should be noted that during the process of editing the eye position in the face image, only the finally trained target image generator model is used. During the training process of the initial image generator model GI, the initial mask generator model GM is used to enable the target image generator model to better learn the correlation between the original input image Image and the mask image Mask, be able to better learn the guiding information in the mask image Mask, and make the training process more stable.

[0147] In this embodiment, the initial mask generator model GM and the initial image generator model GI have a similar structure. Since the input of the initial mask generator model GM is a 3-channel local eye image, the input channel parameter of the first-layer convolution is 3. The generated image Gen_imageB of the second face sample and the restored image Recover_imageA of the first face sample are respectively input into the initial mask generator model GM to obtain the restored mask image Recover_maskB of the second face sample and the generated mask image Gen_maskA of the first face sample.

[0148] S405: Calculate the loss value of the initial mask generator model based on the restored mask image of the second face sample and the original mask image of the second face sample, and correct the parameters of the initial mask generator model according to the loss value of the initial mask generator model to obtain a new initial mask generator model.

[0149] Optionally, according to the recovered mask image Recover_maskB of the second face sample and the original mask image Guided_maskB of the second face sample, the loss value of the initial mask generator model is calculated (i.e., the cycle consistency loss value of the initial mask generator model, denoted as MC). Alternatively, the loss value of the initial mask generator model can also be calculated according to the original mask image Guided_maskA of the first face sample and the generated mask image Gen_maskA of the first face sample.

[0150] Then, according to the loss value of the initial mask generator model, the parameters of the initial mask generator model are corrected to obtain a new initial mask generator model.

[0151] S406. Until the loss value of the initial image generator model meets the preset first loss threshold, and the loss value of the initial mask generator model meets the preset second loss threshold, the initial image generator model that meets the first loss threshold is used as the target image generator model, and the initial mask generator model that meets the second loss threshold is used as the target mask generator model.

[0152] Among them, the first loss threshold and the second loss threshold are respectively preset empirical values. For example, when the loss value of the initial mask generator model is less than the second loss threshold, it means that the accuracy of the newly obtained initial mask generator model cannot be significantly improved.

[0153] Optionally, by continuously iterating and executing steps S401 - S405, the parameters of the initial image generator model and the initial mask image generator model can be continuously adjusted. When the loss value of the newly obtained initial image generator model after update is less than the first loss threshold, and the loss value of the newly obtained initial mask image generator model after update is less than the second loss threshold, the correction of the initial image generator model and the initial mask image generator model can be ended. At this time, the initial mask image generator model is used as the target mask generator model.

[0154] The following embodiments will specifically explain how the loss value of the initial image generator model meets the preset first loss threshold, and the loss value of the initial mask generator model meets the preset second loss threshold, and the initial image generator model that meets the first loss threshold is used as the target image generator model, and the initial mask generator model that meets the second loss threshold is used as the target mask generator model.

[0155] Optionally, as shown in Figure 7 the above step S406 includes:

[0156] S701. If the loss value of the initial image generator model meets the first loss threshold and the loss value of the initial mask generator model meets the second loss threshold, then the original input image of the first face sample and the generated image of the second face sample are input into the image discriminator to determine whether the output result of the image discriminator is false.

[0157] In this embodiment, in order to ensure the accuracy of the output result of the finally trained target image generator model, an image discriminator is also required to identify the authenticity of the generated result output by the target image generator. Different from a general discriminator that only globally evaluates an image and directly outputs a global true / false discrimination result, the image discriminator DI adopted in this application outputs an N*N discrimination matrix. The feature map after the convolutional layer is not sent to the fully connected layer for direct classification, but is directly mapped to an N*N matrix. Each element in the matrix is different true / false, which actually represents a relatively large receptive field in the original image, that is, a patch corresponding to the original image. Adopting the discrimination method proposed in this application can pay more attention to the local generation effect of the image.

[0158] Reference Figure 8 shown, combined with Figure 8 to introduce the network structure of the image discriminator.

[0159] In this embodiment, the image discriminator model includes fourteen network layers. Among them, the channels of the first network layer and the second network layer are 64 and the stride is 2. The first network layer is a convolutional layer with a default convolutional kernel size of 3*3; the second network layer is an activation layer, and the leaky_relu activation function is selected; the channels of the third network layer, the fourth network layer, and the fifth network layer are 128 and the stride is 2. Among them, the third network layer is a convolutional layer with a default convolutional kernel size of 3*3, the fourth layer is a batch normalization layer, and the batch normalization operation BN along the batch dimension is adopted. The fifth network layer is an activation layer, and the leaky_relu activation function is selected; the channels of the sixth network layer, the seventh network layer, and the eighth network layer are 256 and the stride is 2. The sixth network layer is a convolutional layer with a default convolutional kernel size of 3*3, the seventh layer is a batch normalization layer, and the eighth network layer is an activation layer, and the leaky_relu activation function is selected; the channels of the ninth network layer, the tenth network layer, and the eleventh network layer are 512 and the stride is 2. The ninth network layer is a convolutional layer with a default convolutional kernel size of 3*3, the tenth layer is a batch normalization layer, and the eleventh network layer is an activation layer, and the leaky_relu activation function is selected; the channels of the twelfth network layer, the thirteenth network layer, and the fourteenth network layer are 512 and the stride is 1. The twelfth network layer is a convolutional layer with a default convolutional kernel size of 3*3, the thirteenth layer is a batch normalization layer, and the fourteenth network layer is an activation layer, and the leaky_relu activation function is selected.

[0160] In addition, the size of the feature map output by the second network layer is 64 * 128 * 128, the size of the feature map output by the fifth network layer is 128 * 64 * 64, the size of the feature map output by the eighth network layer is 256 * 32 * 32, the size of the feature map output by the eleventh network layer is 521 * 32 * 32, and the size of the feature map output by the fourteenth network layer is 1 * 32 * 32.

[0161] Similarly, the network structure of the mask discriminator is the same as that of the image discriminator, and will not be elaborated here.

[0162] It should be noted that for the labels of the discriminator, when the input is combined with the generated result, the corresponding output result is false, and when the input is combined with the real image, the corresponding output result is true. The initial image generator model GI and the image discriminator DI are alternately trained to improve the authenticity of the generated result of the initial image generator model GI.

[0163] S702. If so, use the initial image generator model as the target image generator model.

[0164] In this embodiment, as shown in Figure 9 When the loss value of the initial image generator model meets the first loss threshold and the loss value of the initial mask generator model meets the second loss threshold, the original input image Input_imageA of the first face sample and the generated image Gen_imageB of the second face sample are input into the image discriminator DI to determine whether the output result of the image discriminator DI is false; if the output result of the image discriminator DI is false, use the initial image generator model as the target image generator model; if the output result of the image discriminator DI is true, the initial image generator model still needs to be iteratively trained until the conditions are met.

[0165] S703. Input the generated image of the second face sample and the original mask image of the second face sample into the mask discriminator to determine whether the output result of the mask discriminator is false.

[0166] S704. If so, use the initial mask generator model as the target mask generator model.

[0167] Among them, for the mask discriminator DM, its structure is similar to that of the image discriminator DI and is used to identify the authenticity of the generation result of the initial mask generator model GM. Since the input image received by the mask discriminator DM is 6 channels after splicing, the input channel parameter of the first-layer convolution is 6. Similar to the training process of the image discriminator DI, when the input is combined with the generation result, the corresponding output result is false; when the input is combined with the real image, the corresponding output result is true. By alternately training the initial mask generator model GM and the mask discriminator DM, the authenticity of the generation result of the initial mask generator model GM is improved.

[0168] In this embodiment, continue to refer to Figure 9 As shown, when the loss value of the initial image generator model satisfies the first loss threshold and the loss value of the initial mask generator model satisfies the second loss threshold, the generated image Gen_imageB of the second face sample and the original mask image Mask_imageB of the second face sample are input to the mask discriminator DM to determine whether the output result of the mask discriminator DM is false; if the output result of the mask discriminator DM is false, the initial mask generator model is used as the target mask generator model; if the output result of the mask discriminator DM is true, the initial mask generator model needs to be iteratively trained until the condition is met.

[0169] In addition, in this embodiment, the loss function used in the training process of this solution will be further introduced.

[0170] img in the loss function shown below A , img B , mask A , mask B are the inputs of the model in the training stage, corresponding to the set composed of the original input image Input_imageA of the first face sample, the original input image Input_imageB of the second face sample, the original mask image Guided_maskA of the first face sample, and the original mask image Guided_maskB of the second face sample in the input data of each batch of models.

[0171] Among them, img A , img B , mask A , mask B There are two pairing methods in total: img A , img B , mask B and img A , img B , mask A .

[0172] For the combined img A , img B , mask B the loss function is given by the following formula (1):

[0173]

[0174] For another combination of img A , img B , mask A the loss function is given by the following formula (2):

[0175] Therefore, the loss value GAN-Loss of the image generator model G I and the image discriminator model D I is shown by the following formula (3):

[0176]

[0177] The following uses existing mathematical theories to explain the above formulas (1)-(3):

[0178] Assume that the amount of data fed into the network in each batch during the training process is N, then the input set can be expressed by the following formula (4):

[0179] img A , img B , mask B ={img a1 , img b1 , mask b1 , img a2 , img b2 , mask b2 …, img aN , img bN , mask bN} (4)

[0180] For the above input, it can be regarded as a discrete distribution, and for each group of inputs, corresponding discriminator outputs will also be obtained, and the outputs are also discrete distributions. In probability theory, cross-entropy can be used to describe the difference between two distributions. According to the following calculation formula (5) of cross-entropy:

[0181] H(P|Q) = -(p log q + (1 - p) log(1 - q)) (5)

[0182] In a GAN, p in formula (5) is the probability that all the inputs to the discriminator are sampled from the real dataset. There are only two types of input combinations. One is that all samples are from the real dataset, in which case p is 1. The other type includes the generated samples from the generator, in which case p is 0.

[0183] Taking formula (1) as an example, for the first half of formula (1), the combination img A , img B , mask B is a real sample, at which time p is 1, and the cross-entropy value is as shown in the following formula (6):

[0184] H1(P|Q) = -log q = -logD I ([img A , img B , mask B ) (6)

[0185] For the second half of formula (1), the discriminator input includes the generation result, p is 0, and the cross-entropy value is as shown in the following formula (7):

[0186] H2(P|Q) = -log(1 - q) = -log(1 - D I ([img A , mask B , G I (img A , mask B )]) (7)

[0187] Therefore, for the cross-entropy of the input and output distributions of a single batch of data in formula (1), it is as shown in the following formula (8):

[0188] H(P|Q) = -(log q + log(1 - q)) = -(logD I ([img A , img B , mask B ) + logD I ([img A , img B , mask B )) (8)

[0189] In the entire dataset, the cross-entropy of the input and output distributions is the expectation of the cross-entropy values of single-batch data, that is, the result in formula (1). For the combination of img A , img B , mask A in formula (2), the same principle applies.

[0190] The above formula (3) is carried out in two stages. In the first stage, the image discriminator D I hopes to maximize the above loss as much as possible and widen the difference between the input distribution and the output distribution to distinguish the generated result and the real label. In the second stage, the image generator model G I aims to minimize the above to make the generated result close to the label to deceive the image discriminator D I .

[0191] Meanwhile, in order to further improve the generation effect of the generator, for the generated result of the image generator model G I the cycle consistency loss IC is additionally adopted and it is supervised by the corresponding label, as shown in the following formulas (9)-(11):

[0192]

[0193]

[0194]

[0195] In the cycle consistency loss, the l1-Loss is mainly used to evaluate the difference between the generated result and the target image. The calculation method of l1-Loss is relatively simple, which is obtained by taking the absolute value of the difference of the pixel values at the corresponding positions.

[0196] For the mask generator model G M and the mask discriminator D M similar to the Loss in the above image generator model G I and the image discriminator D I it consists of the GAN-Loss and the cycle consistency loss MC, which will not be elaborated here.

[0197] The following embodiments will specifically explain how to determine the first image and the second image according to the face image in the above step S203.

[0198] Optionally, as shown in Figure 10 the above step S203 includes:

[0199] S1001. Extract the eye feature information in the face image.

[0200] Among them, the eye feature information includes: the center point of the eyeball, the width and height of the eye area, the upper and lower eyelid points, the upper and lower vertex points of the eyeball, and the radius of the eyeball.

[0201] In this embodiment, for example, an existing face key point localization model can be used to localize the face key points in the input face image, and the center point of the eyeball, the width and height of the eye area, and the upper and lower eyelid points are taken out from the face key points.

[0202] An affine transformation can be used to describe the mapping relationship between two images. The affine transformation matrix is a 2x3 matrix. For example, M2 in formulas (12)-(13) can translate the image, while the diagonal elements in M1 determine the scaling of the image, and the off-diagonal elements determine the rotation of the image.

[0203]

[0204]

[0205] Given the coordinates (x, y) of the original image pixels, after an affine transformation, the point (u, v) is obtained. The specific transformation is shown in the following formulas (14)-(15):

[0206]

[0207] That is,

[0208] S1002. Calculate the vertex coordinates of the eye bounding box based on the center point of the eyeball and the width and height of the eye region.

[0209] Among them, the vertex coordinates of the eye bounding box are the position information of each vertex of the eye region in the face image.

[0210] S1003. Calculate the first affine transformation matrix based on the vertex coordinates of the eye bounding box and a preset cropped eye image.

[0211] Among them, the first affine transformation matrix is the affine transformation matrix from the original image to the local image, and the image size of the cropped eye image is a preset value.

[0212] Exemplarily, for example, the size of the preset cropped eye image is 128*128.

[0213] S1004. Perform an affine transformation on the face image according to the first affine transformation matrix to obtain an initial image containing the eye region.

[0214] In this embodiment, the vertex coordinates of the eye bounding box can be calculated based on the center point of the eyeball and the width and height of the eye region in the eye feature information. Based on the vertex coordinates of the eye bounding box (i.e., the position information of the eye region in the face image) and the four vertex coordinates of the cropped image, the first affine transformation matrix is calculated. Using the first affine transformation matrix, a corresponding affine transformation is performed on the face image to extract the eye region, that is, an initial image containing the eye region is obtained.

[0215] S1005. Obtain the first image based on the upper and lower eyelid points and the initial image containing the eye region.

[0216] In this embodiment, quadratic curve fitting can be performed on the upper and lower eyelid points among the facial key points to obtain a fitting curve. Dense sampling of the upper and lower eyelid points is carried out on the fitting curve, and the point set obtained by sampling is connected to form a smooth eye closed area mask, that is, the first area in the mask image Guided_mask as shown in Figure 11 For the initial image containing the eye area, using the eye closed area mask (i.e., the first area in the mask image Guided_mask as shown in Figure 11 ), the irrelevant information outside the mask is set to zero, and the first image Input_image can be obtained.

[0217] S1006. Obtain a second image according to the target position and the position mapping relationship between the target position and the upper and lower vertex positions of the eyeball in the eye feature information.

[0218] Among them, the target position is the information obtained by user input.

[0219] In a feasible manner, for example, the information input by the user is the horizontal degree value of the eyeball, that is, the horizontal degree value of the eyeball can be mapped to the coordinates of the center point of the eyeball (i.e., the target position where the eyeball in the facial image is to move).

[0220] In this embodiment, since the target position where the eyeball in the facial image is to move needs to be obtained according to user input. Therefore, first, according to the upper and lower eyelid key points among the eye key points, function fitting is performed on the possible movement trajectories of the coordinates of the center point of the eyeball during the left and right movement of the eyeball, and the position mapping relationship between the coordinates of the center point of the eyeball and the upper and lower vertex positions of the eyeball is established in advance; after obtaining the target position where the eyeball in the facial image is to move, combined with the eyeball radius in the eye key points in the facial image and the position mapping relationship between the target position and the upper and lower vertex positions of the eyeball in the eye key points, approximate fitting is performed on the area where the eyeball is located, and the second area in the mask image Guided_mask as shown in Figure 11 can be obtained. Combining the first area in the mask image Guided_mask as shown in Figure 11 obtained above, and obtaining the final second image Guided_mask, that is, the mask image where the eyeball in the original eye area image is in the target position.

[0221] The following embodiments will specifically explain how to perform fusion processing on the eye area image in the facial image according to the facial image and the third image in step S204 to obtain the processed target image.

[0222] Optionally, as shown in reference Figure 12 , the above step S204 includes:

[0223] S1201. Use the eye segmentation model to perform eye segmentation on the first image and the third image respectively, to obtain the original eyeball region in the first image and the eyeball region in the third image.

[0224] S1202. Perform Poisson fusion on the original eyeball region in the first image and the eyeball region in the third image to obtain the processed third image.

[0225] In this embodiment, in order to further restore the highlights and pupil colors in the original eyeball, the present application proposes that an existing eye segmentation model can be used to perform eye segmentation on the first image and the third image respectively to obtain the eye segmentation results, that is, the original eyeball region in the first image and the eyeball region in the third image. Then, the original eyeball region in the first image is Poisson-fused with the eyeball region of the third image Gen_image to obtain the processed third image.

[0226] S1203. According to the face image and the processed third image, perform fusion processing on the eye region image in the face image to obtain the processed target image.

[0227] In this embodiment, after the generation of the local eye image in the new state is completed, in the fusion stage, it is necessary to fuse the generated third image (i.e., the local eye image) with the original eye region in the face image. Specifically, it is necessary to first calculate the second affine transformation matrix according to the face image and the third image, where the second affine transformation matrix is the affine transformation matrix from the local image to the original image; then, according to the second affine transformation matrix, determine the coordinate points of the third image in the face image; finally, according to the coordinate points of the third image in the face image and the third image, perform fusion processing on the eye region image in the face image, and then perform Poisson fusion on the face image and the affine-transformed local eye image to obtain the processed target image, and the final effect after eyeball position editing can be obtained.

[0228] Based on the same inventive concept, an image processing apparatus corresponding to the image processing method is further provided in the embodiments of the present application. Since the principle of solving problems by the apparatus in the embodiments of the present application is similar to that of the above-mentioned image processing method in the embodiments of the present application, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be described again.

[0229] Optionally, as shown in Figure 13 An image processing apparatus is further provided in the embodiments of the present application. The apparatus includes:

[0230] An acquisition module 1301, configured to acquire a face image to be processed and a target position input by a user, where the target position is the position where the eyeball in the face image is to be moved.

[0231] A processing module 1302, configured to determine a first image and a second image according to a face image and a target position, where the first image is an original eye region image in the face image, and the second image is a mask image of the eyeball at the target position in the original eye region image; input the first image and the second image into a pre-trained target image generator model to obtain a third image, where the third image is a target eye region image, and the eyeball in the third image is located at the target position;

[0232] A fusion module 1303, configured to perform a fusion process on the eye region image in the face image according to the face image and the third image to obtain a processed target image.

[0233] Optionally, the apparatus further includes:

[0234] An acquisition module, configured to acquire an original input image and an original mask image of multiple face samples, where the original input image is an original eye region image in the face image, and the original mask image is a mask image of the actual position of the eyeball in the face image;

[0235] A training module, configured to perform iterative training on an initial image generator model based on the original input images and the original mask images of multiple face samples to obtain a target image generator model.

[0236] Optionally, the training module is further configured to:

[0237] Input the original input image of the first face sample and the original mask image of the second face sample into the initial image generator model to obtain a generated image of the second face sample output by the initial image generator model;

[0238] Input the generated image of the second face sample and the original mask image of the first face sample into the initial image generator model to obtain a restored image of the first face sample output by the initial image generator model;

[0239] Calculate a loss value of the initial image generator model according to the original input image of the first face sample and the restored image of the first face sample, and correct the parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a new initial image generator model;

[0240] Input the generated image of the second face sample and the restored image of the first face sample into the initial mask generator model respectively to obtain a restored mask image of the second face sample and a generated mask image of the first face sample;

[0241] Calculate the loss value of the initial mask generator model based on the restored mask image of the second face sample and the original mask image of the second face sample, and correct the parameters of the initial mask generator model according to the loss value of the initial mask generator model to obtain a new initial mask generator model;

[0242] Loop through the above steps until the loss value of the initial image generator model meets a preset first loss threshold and the loss value of the initial mask generator model meets a preset second loss threshold. Take the initial image generator model that meets the first loss threshold as the target image generator model and the initial mask generator model that meets the second loss threshold as the target mask generator model.

[0243] Optionally, the training module is further configured to:

[0244] If the loss value of the initial image generator model meets the first loss threshold and the loss value of the initial mask generator model meets the second loss threshold, then input the original input image of the first face sample and the generated image of the second face sample into the image discriminator to determine whether the output result of the image discriminator is false;

[0245] If so, take the initial image generator model as the target image generator model;

[0246] Input the generated image of the second face sample and the original mask image of the second face sample into the mask discriminator to determine whether the output result of the mask discriminator is false;

[0247] If so, take the initial mask generator model as the target mask generator model.

[0248] Optionally, the processing module 1302 is further configured to:

[0249] Extract the eye feature information in the face image, where the eye feature information includes: the center point of the eyeball, the width and height of the eye area, the upper and lower eyelid points, the upper and lower vertices of the eyeball, and the radius of the eyeball;

[0250] Calculate the vertex coordinates of the eye bounding box based on the center point of the eyeball and the width and height of the eye area;

[0251] Calculate the first affine transformation matrix based on the vertex coordinates of the eye bounding box and a preset cropped eye image, where the first affine transformation matrix is the affine transformation matrix from the original image to the local image, and the image size of the cropped eye image is a preset value;

[0252] Perform an affine transformation on the face image according to the first affine transformation matrix to obtain an initial image containing the eye area;

[0253] Obtain a first image based on the upper and lower eyelid key points and the initial image containing the eye area;

[0254] A second image is obtained according to the target position and the position mapping relationship between the target position and the upper and lower vertices of the eyeball in the eye key points.

[0255] Optionally, the fusion module 1303 is further configured to:

[0256] Use an eye segmentation model to perform eye segmentation on the first image and the third image respectively, to obtain the original eyeball region in the first image and the eyeball region in the third image;

[0257] Perform Poisson fusion on the original eyeball region in the first image and the eyeball region in the third image to obtain the processed third image;

[0258] According to the face image and the processed third image, perform fusion processing on the eye region image in the face image to obtain the processed target image.

[0259] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0260] The above modules may be one or more integrated circuits configured to implement the above method. For example: one or more application specific integrated circuits (ASICs), or, one or more microprocessors (digital singnal processors, DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when the above certain module is implemented in the form of a processing element dispatching program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0261] Optionally, the present application further provides a program product, such as a computer-readable storage medium, including a program that is used to execute the above method embodiment when executed by a processor.

[0262] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0263] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0264] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0265] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks, or optical discs that can store program codes.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining a face image to be processed and a target position input by a user, where the target position is the position where the eyeballs in the face image are to move to; Determining a first image and a second image according to the face image and the target position, where the first image is an original eye region image in the face image, and the second image is a mask image in the original eye region image where the eyeballs are at the target position; Inputting the first image and the second image into a pre-trained target image generator model to obtain a third image, where the third image is a target eye region image, and the eyeballs in the third image are at the target position; Performing fusion processing on the eye region image in the face image according to the face image and the third image to obtain a processed target image; Among them, the target image generator model is trained in the following manner: Obtaining original input images and original mask images of multiple face samples, where the original input images are original eye region images in the face images, and the original mask images are mask images of the actual positions of the eyeballs in the face images; Performing iterative training on an initial image generator model based on the original input images and original mask images of the multiple face samples to obtain the target image generator model; Among them, performing iterative training on an initial image generator model based on the original input images and original mask images of the multiple face samples to obtain the target image generator model includes: Inputting the original input image of the first face sample and the original mask image of the second face sample into the initial image generator model to obtain a generated image of the second face sample output by the initial image generator model; Inputting the generated image of the second face sample and the original mask image of the first face sample into the initial image generator model to obtain a restored image of the first face sample output by the initial image generator model; Calculating a loss value of the initial image generator model according to the original input image of the first face sample and the restored image of the first face sample, and correcting parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a new initial image generator model; Inputting the generated image of the second face sample and the restored image of the first face sample into an initial mask generator model respectively to obtain a restored mask image of the second face sample and a generated mask image of the first face sample; Calculating a loss value of the initial mask generator model according to the restored mask image of the second face sample and the original mask image of the second face sample, and correcting parameters of the initial mask generator model according to the loss value of the initial mask generator model to obtain a new initial mask generator model; Execute the above steps cyclically until the loss value of the initial image generator model meets a preset first loss threshold, and the loss value of the initial mask generator model meets a preset second loss threshold. Then, use the initial image generator model that meets the first loss threshold as the target image generator model, and use the initial mask generator model that meets the second loss threshold as the target mask generator model.

2. The method according to claim 1, characterized in that, The loss value of the initial image generator model meets a preset first loss threshold, and the loss value of the initial mask generator model meets a preset second loss threshold. Using the initial image generator model that meets the first loss threshold as the target image generator model, and using the initial mask generator model that meets the second loss threshold as the target mask generator model includes: If the loss value of the initial image generator model meets the first loss threshold, and the loss value of the initial mask generator model meets the second loss threshold, then input the original input image of the first face sample and the generated image of the second face sample into the image discriminator to determine whether the output result of the image discriminator is false; If so, use the initial image generator model as the target image generator model; Input the generated image of the second face sample and the original mask image of the second face sample into the mask discriminator to determine whether the output result of the mask discriminator is false; If so, use the initial mask generator model as the target mask generator model.

3. The method according to claim 1, characterized in that, Determining the first image and the second image according to the face image includes: Extract the eye feature information in the face image, where the eye feature information includes: the center point of the eyeball, the width and height of the eye region, the upper and lower eyelid points, the upper and lower vertices of the eyeball, and the radius of the eyeball; Calculate the vertex coordinates of the eye bounding box according to the center point of the eyeball and the width and height of the eye region; Calculate a first affine transformation matrix according to the vertex coordinates of the eye bounding box and a preset post-cropping eye image, where the first affine transformation matrix is the affine transformation matrix from the original image to the local image, and the image size of the post-cropping eye image is a preset value; Perform an affine transformation on the face image according to the first affine transformation matrix to obtain an initial image containing the eye region; Obtain the first image according to the upper and lower eyelid points and the initial image containing the eye region; Obtain the second image according to the target position and the position mapping relationship between the target position and the upper and lower vertices of the eyeball in the eye feature information.

4. The method according to claim 1, characterized in that, Fusing the eye region images in the face image according to the face image and the third image to obtain a processed target image includes: Use an eye segmentation model to perform eye segmentation on the first image and the third image respectively to obtain the original eyeball region in the first image and the eyeball region in the third image; Perform Poisson fusion on the original eyeball region in the first image and the eyeball region in the third image to obtain a processed third image; Based on the face image and the processed third image, perform a fusion process on the eye region image in the face image to obtain a processed target image.

5. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire a face image to be processed and a target position input by a user, where the target position is the position where the eyeball in the face image is to move to; A processing module, configured to determine a first image and a second image according to the face image and the target position, where the first image is the original eye region image in the face image, and the second image is a mask image in which the eyeball in the original eye region image is at the target position; input the first image and the second image into a pre-trained target image generator model to obtain a third image, where the third image is a target eye region image and the eyeball in the third image is at the target position; A fusion module, configured to perform a fusion process on the eye region image in the face image according to the face image and the third image to obtain a processed target image; Wherein, the device further includes: An acquisition module, configured to acquire an original input image and an original mask image of multiple face samples, where the original input image is the original eye region image in the face image, and the original mask image is a mask image of the actual position of the eyeball in the face image; A training module, configured to perform iterative training on an initial image generator model based on the original input images and original mask images of the multiple face samples to obtain the target image generator model; The training module is specifically configured to: Input the original input image of the first face sample and the original mask image of the second face sample into the initial image generator model to obtain a generated image of the second face sample output by the initial image generator model; Input the generated image of the second face sample and the original mask image of the first face sample into the initial image generator model to obtain a restored image of the first face sample output by the initial image generator model; Calculate the loss value of the initial image generator model according to the original input image of the first face sample and the restored image of the first face sample, and correct the parameters of the initial image generator model according to the loss value of the initial image generator model to obtain a new initial image generator model; Input the generated image of the second face sample and the restored image of the first face sample into an initial mask generator model respectively to obtain a restored mask image of the second face sample and a generated mask image of the first face sample; Calculate the loss value of the initial mask generator model according to the restored mask image of the second face sample and the original mask image of the second face sample, and correct the parameters of the initial mask generator model according to the loss value of the initial mask generator model to obtain a new initial mask generator model; Repeat the above steps until the loss value of the initial image generator model meets a preset first loss threshold and the loss value of the initial mask generator model meets a preset second loss threshold. Then, use the initial image generator model that meets the first loss threshold as the target image generator model, and use the initial mask generator model that meets the second loss threshold as the target mask generator model.

6. An electronic device, characterized in that, Comprising: A processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus. The processor executes the machine-readable instructions to perform the steps of the method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium. When the computer program is run by the processor, it performs the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Image processing method and device and electronic equipment

    CN112533071A

  • Face image sight correction method and device, equipment and storage medium

    CN112733797A