An image processing method, apparatus and device

By merging segmented iris images with colored contact lens images, the problem of unstable colored contact lens trial images was solved, generating highly realistic colored contact lens trial images.

CN117275077BActive Publication Date: 2026-03-17HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210664959.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2026-03-17
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

Existing technologies generate inconsistent and distorted images of colored contact lenses, and cannot effectively incorporate the iris texture features and contact lens styles of different users.

Method used

By segmenting a local iris image from a facial image and performing a first fusion process with the colored contact lens image to preserve iris texture features, and then performing a second fusion process with the facial image, a colored contact lens trial image is generated.

Benefits of technology

The generated images of colored contact lenses in a trial setting retain the texture features of the iris area while incorporating the style of the colored contact lenses, achieving a highly realistic wearing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117275077B_ABST
    Figure CN117275077B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image processing method, device and equipment, and relate to the technical field of computers, in particular to the technical field of image processing. The specific implementation scheme is: obtaining a beauty lens image and a face image to be tried on with the beauty lens; segmenting an iris local image containing an iris region from the face image; performing first fusion processing on the iris local image and the beauty lens image to obtain a local beauty lens effect image; wherein the first fusion processing is used for integrating an image style of the beauty lens image into the iris local image and retaining texture features of the iris local image; performing second fusion processing on the local beauty lens effect image and the face image to obtain a beauty lens try-on image; wherein the second fusion processing is used for integrating the image style of the local beauty lens effect image into an eye region of the face image. It can be seen that, through the present scheme, a beauty lens try-on image with high simulation effect can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to the field of image processing technology, specifically to an image processing method, apparatus, and device. Background Technology

[0002] In recent years, colored contact lenses have become increasingly popular among consumers, with more and more young people choosing to wear them. However, because everyone's eye shape, iris size, texture, and color differ, the effect of colored contact lenses varies from person to person. Therefore, it is necessary to select the right size of colored contact lenses before wearing them to achieve the best wearing effect.

[0003] To improve the efficiency of users when choosing colored contact lenses, it is common practice to generate trial images of the lenses to assist users in their selection. However, in related technologies, trial images are usually generated through methods such as texturing, rendering, and overlaying. These methods result in inconsistent effects for different users and in different scenarios, easily leading to distortion and even undesirable phenomena such as colored contact lenses appearing in non-eye areas.

[0004] Therefore, how to generate highly realistic images of colored contact lenses being tried on has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of this invention is to provide an image processing method, apparatus, and device to generate highly realistic images of contact lens try-on. The specific technical solution is as follows:

[0006] In a first aspect, embodiments of the present invention provide an image processing method, the method comprising:

[0007] Acquire images of the colored contact lenses and facial images of the face to be tested with the lenses;

[0008] Segment the iris region from the facial image;

[0009] The partial iris image and the colored contact lens image are subjected to a first fusion process to obtain a partial colored contact lens effect image; wherein, the first fusion process is used to incorporate the image style of the colored contact lens image into the partial iris image while preserving the texture features of the partial iris image;

[0010] The partial contact lens effect image and the facial image are subjected to a second fusion process to obtain a contact lens trial image; wherein, the second fusion process is used to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image.

[0011] Optionally, segmenting the iris region from the facial image includes:

[0012] Identify key eye points in the facial image;

[0013] Using the key eye point information, the iris region in the facial image is determined;

[0014] The facial image is cropped to obtain a partial iris image.

[0015] Optionally, determining the iris region in the facial image using the eye key point information includes:

[0016] The eye contour in the facial image is fitted using the key eye point information to obtain an eye fitting curve; wherein, the eye contour is the contour that fits the eyelid.

[0017] Based on the eye fitting curve and the contour parameters of the iris contour, the facial image is subjected to foreground masking to obtain a mask image of the facial image; wherein, the foreground masking is a masking process that uses the iris region of the facial image as the foreground region; the contour parameters of the iris contour are generated based on specified key points in the eye key point information, and the specified key point information is key point information located on the contour of the outer circle of the iris.

[0018] Based on the mask image of the facial image, the iris region in the facial image is determined.

[0019] Optionally, the step of performing foreground masking processing on the facial image based on the eye fitting curve and the contour parameters of the iris contour to obtain a mask image of the facial image includes:

[0020] Based on the eye fitting curve and the contour parameters of the iris contour, a specified pixel in the facial image is identified; wherein the specified pixel is located within the area formed by the eye contour and within the area formed by the contour of the outer circle of the iris.

[0021] The pixel value of a specified pixel in the facial image is set to a first value, and the pixels in the facial image other than the specified pixel are set to a second value to obtain a mask image of the facial image.

[0022] Optionally, the second fusion processing of the partial contact lens effect image and the facial image to obtain a contact lens trial image includes:

[0023] Based on the partial contact lens effect image, a backup effect image is generated; wherein, the size of the backup effect image is the same as the size of the target eye image, and the target area in the backup effect image has the image information of the iris area of ​​the partial contact lens effect image, the target eye image is the eye area image of the face image, and the target area is the area at the same position as the iris area of ​​the eye area image;

[0024] Determine the fusion coefficient corresponding to each pixel in the backup effect image;

[0025] Using the determined fusion coefficient, the backup effect image and the eye area in the facial image are fused to obtain a colored contact lens trial image.

[0026] Optionally, generating a backup effect image based on the partial contact lens effect image includes:

[0027] The iris region in the partial contact lens effect image is identified as the area to be utilized;

[0028] The image information of the target region in the specified image is replaced with the image information of the region to be used to obtain a backup effect image; wherein, the specified image is the same size as the target eye image.

[0029] Optionally, determining the iris region in the partial contact lens effect image as the region to be utilized includes:

[0030] Determine the target mask image of the iris region; wherein the target mask image is a binarized image with the iris region as the foreground;

[0031] Multiply the pixel values ​​of each pixel in the local contact lens effect image with the pixel values ​​of the corresponding pixels in the target mask image;

[0032] Based on the pixel values ​​of the pixels in the image obtained after multiplication, the iris region in the local contact lens effect image is determined as the region to be utilized.

[0033] Optionally, the first fusion process of the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image includes:

[0034] Extract the style features of the colored contact lens image;

[0035] Extract the texture features of the local iris image;

[0036] The style features and texture features are fused to obtain the fused features corresponding to the local iris image;

[0037] An image with fused features corresponding to the local iris image is generated as a local contact lens effect image.

[0038] Optionally, the step of performing a first fusion process on the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image is implemented by a pre-trained fusion network;

[0039] The fusion network is a neural network trained on sample iris local images and sample contact lens images, used to output fusion results; the fusion result is the result of incorporating the image style of the sample contact lens image into the sample iris local image while retaining the texture features of the sample iris local image.

[0040] Optionally, the pre-trained fusion network is a generative network in a pre-trained generative adversarial network;

[0041] The training process of the generative adversarial network includes:

[0042] Obtain a partial image of the sample iris and an image of the sample contact lens;

[0043] The sample iris partial image and the sample colored contact lens image are input into the generation network to obtain a generated image as the fusion result;

[0044] The generated image and the sample iris image are respectively input into the adversarial network in the generative adversarial network to obtain a first discrimination result corresponding to the generated image and a second discrimination result corresponding to the sample iris image; wherein, the adversarial network is used to identify whether the input image is a real image;

[0045] Based on the difference between the first identification result and the corresponding true value, and the difference between the second identification result and the corresponding true value, the loss value of the adversarial network is calculated;

[0046] Based on the difference between the first identification result and the corresponding false value, the loss value of the generator network is calculated;

[0047] Based on the loss value of the adversarial network and the loss value of the generator network, it is determined whether the adversarial network and the generator network have reached a Nash equilibrium state; wherein, the Nash equilibrium state is used to characterize a stable state in which the loss values ​​of the adversarial network and the generator network fluctuate within a specified range.

[0048] If the adversarial network and the generative network do not reach Nash equilibrium, adjust the parameters of the generative adversarial network and return to the step of obtaining the sample iris local image and the sample contact lens image.

[0049] Optionally, before calculating the loss value of the generator network based on the difference between the first identification result and the corresponding false value, the method further includes:

[0050] The generated image and the sample iris image are respectively input into the feature extraction network, and a first loss value is calculated based on the difference between the output results of the feature extraction network for the generated image and the sample iris image.

[0051] The generated image and the sample contact lens image are respectively input into a style constraint network, and a second loss value is calculated based on the difference between the output results of the style constraint network for the generated image and the sample contact lens image; wherein, the style constraint network is a neural network used to identify image styles;

[0052] The step of calculating the loss value of the generator network based on the difference between the first identification result and the corresponding false value includes:

[0053] Based on the difference between the first identification result and the corresponding false value, as well as the first loss value and the second loss value, the loss value of the generator network is calculated.

[0054] Optionally, the generating network includes a style coding network, an encoder, and a decoder;

[0055] The step of inputting the sample iris local image and the sample contact lens image into the generation network to obtain the generated image as the fusion result includes:

[0056] The sample colored contact lens image is input into the style coding network to obtain the style features of the sample colored contact lens image;

[0057] The style features of the sample contact lens image and the sample iris partial image are input into the encoder for encoding to obtain the fused features corresponding to the sample iris partial image; the fused features corresponding to the sample iris partial image are feature images that fuse the style features of the sample contact lens image and the texture features of the sample iris partial image.

[0058] The fused features corresponding to the local iris image of the sample are input into the decoder for decoding to obtain the generated image as the fusion result.

[0059] In a second aspect, embodiments of the present invention provide an image processing apparatus, the apparatus comprising:

[0060] The acquisition module is used to acquire images of the colored contact lenses and facial images of the face to be tested with the colored contact lenses;

[0061] A segmentation module is used to segment a local iris image containing the iris region from the facial image;

[0062] The first fusion module is used to perform a first fusion process on the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image; wherein, the first fusion process is used to incorporate the image style of the colored contact lens image into the partial iris image and retain the texture features of the partial iris image;

[0063] The second fusion module is used to perform a second fusion process on the partial contact lens effect image and the facial image to obtain a contact lens trial image; wherein, the second fusion process is used to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image.

[0064] Optionally, the segmentation module includes:

[0065] The recognition submodule is used to recognize key eye point information in the facial image;

[0066] The first determining submodule is used to determine the iris region in the facial image using the eye key point information;

[0067] The cropping submodule is used to crop the iris region of the facial image to obtain a partial iris image.

[0068] Optionally, the determining submodule is specifically used for:

[0069] The eye contour in the facial image is fitted using the key eye point information to obtain an eye fitting curve; wherein, the eye contour is the contour that fits the eyelid.

[0070] Based on the eye fitting curve and the contour parameters of the iris contour, the facial image is subjected to foreground masking to obtain a mask image of the facial image; wherein, the foreground masking is a masking process that uses the iris region of the facial image as the foreground region; the contour parameters of the iris contour are generated based on specified key points in the eye key point information, and the specified key point information is key point information located on the contour of the outer circle of the iris.

[0071] Based on the mask image of the facial image, the iris region in the facial image is determined.

[0072] Optionally, the step of performing foreground masking processing on the facial image based on the eye fitting curve and the contour parameters of the iris contour to obtain a mask image of the facial image includes:

[0073] Based on the eye fitting curve and the contour parameters of the iris contour, a specified pixel in the facial image is identified; wherein the specified pixel is located within the area formed by the eye contour and within the area formed by the contour of the outer circle of the iris.

[0074] The pixel value of a specified pixel in the facial image is set to a first value, and the pixels in the facial image other than the specified pixel are set to a second value to obtain a mask image of the facial image.

[0075] Optionally, the second fusion module includes:

[0076] The first generation submodule is used to generate a backup effect image based on the partial contact lens effect image; wherein, the size of the backup effect image is the same as the size of the target eye image, and the target area in the backup effect image has image information of the iris area of ​​the partial contact lens effect image, the target eye image is the eye area image of the face image, and the target area is the area at the same position as the iris area of ​​the eye area image;

[0077] The second determining submodule is used to determine the fusion coefficient corresponding to each pixel in the backup effect image;

[0078] The second fusion submodule is used to fuse the backup effect image and the eye area in the facial image using the determined fusion coefficient to obtain a colored contact lens trial image.

[0079] Optionally, the first generation submodule is specifically used for:

[0080] The iris region in the partial contact lens effect image is identified as the area to be utilized;

[0081] The image information of the target region in the specified image is replaced with the image information of the region to be used to obtain a backup effect image; wherein, the specified image is the same size as the target eye image.

[0082] Optionally, determining the iris region in the partial contact lens effect image as the region to be utilized includes:

[0083] Determine the target mask image of the iris region; wherein the target mask image is a binarized image with the iris region as the foreground;

[0084] Multiply the pixel values ​​of each pixel in the local contact lens effect image with the pixel values ​​of the corresponding pixels in the target mask image;

[0085] Based on the pixel values ​​of the pixels in the image obtained after multiplication, the iris region in the local contact lens effect image is determined as the region to be utilized.

[0086] Optionally, the first fusion module includes:

[0087] The first extraction submodule is used to extract the style features of the colored contact lens image;

[0088] The second extraction submodule is used to extract the texture features of the local iris image;

[0089] The first fusion submodule is used to fuse the style features and the texture features to obtain the fused features corresponding to the local iris image;

[0090] The second generation submodule is used to generate an image with fused features corresponding to the local iris image, as a local contact lens effect image.

[0091] Optionally, the step of performing a first fusion process on the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image is implemented by a pre-trained fusion network; wherein, the fusion network is a neural network trained based on the sample partial iris image and the sample colored contact lens image, used to output the fusion result; wherein, the fusion result is the result of incorporating the image style of the sample colored contact lens image into the sample partial iris image, while retaining the texture features of the sample partial iris image.

[0092] Optionally, the pre-trained fusion network is a generative network in a pre-trained generative adversarial network;

[0093] The training process of the generative adversarial network includes:

[0094] Obtain a partial image of the sample iris and an image of the sample contact lens;

[0095] The sample iris partial image and the sample colored contact lens image are input into the generation network to obtain a generated image as the fusion result;

[0096] The generated image and the sample iris image are respectively input into the adversarial network in the generative adversarial network to obtain a first discrimination result corresponding to the generated image and a second discrimination result corresponding to the sample iris image; wherein, the adversarial network is used to identify whether the input image is a real image;

[0097] Based on the difference between the first identification result and the corresponding true value, and the difference between the second identification result and the corresponding true value, the loss value of the adversarial network is calculated;

[0098] Based on the difference between the first identification result and the corresponding false value, the loss value of the generator network is calculated;

[0099] Based on the loss value of the adversarial network and the loss value of the generator network, it is determined whether the adversarial network and the generator network have reached a Nash equilibrium state; wherein, the Nash equilibrium state is used to characterize a stable state in which the loss values ​​of the adversarial network and the generator network fluctuate within a specified range.

[0100] If the adversarial network and the generative network do not reach Nash equilibrium, adjust the parameters of the generative adversarial network and return to the step of obtaining the sample iris local image and the sample contact lens image.

[0101] Optionally, before calculating the loss value of the generator network based on the difference between the first identification result and the corresponding false value, the method further includes:

[0102] The generated image and the sample iris image are respectively input into the feature extraction network, and a first loss value is calculated based on the difference between the output results of the feature extraction network for the generated image and the sample iris image.

[0103] The generated image and the sample contact lens image are respectively input into a style constraint network, and a second loss value is calculated based on the difference between the output results of the style constraint network for the generated image and the sample contact lens image; wherein, the style constraint network is a neural network used to identify image styles;

[0104] The step of calculating the loss value of the generator network based on the difference between the first identification result and the corresponding false value includes:

[0105] Based on the difference between the first identification result and the corresponding false value, as well as the first loss value and the second loss value, the loss value of the generator network is calculated.

[0106] Optionally, the generating network includes a style coding network, an encoder, and a decoder;

[0107] The step of inputting the sample iris local image and the sample contact lens image into the generation network to obtain the generated image as the fusion result includes:

[0108] The sample colored contact lens image is input into the style coding network to obtain the style features of the sample colored contact lens image;

[0109] The style features of the sample contact lens image and the sample iris partial image are input into the encoder for encoding to obtain the fused features corresponding to the sample iris partial image; the fused features corresponding to the sample iris partial image are feature images that fuse the style features of the sample contact lens image and the texture features of the sample iris partial image.

[0110] The fused features corresponding to the local iris image of the sample are input into the decoder for decoding to obtain the generated image as the fusion result.

[0111] Thirdly, embodiments of the present invention provide an electronic device, including a camera, a display, a processor, and a memory;

[0112] The camera is used to capture facial images of the person being tested for contact lenses;

[0113] Memory, used to store computer programs;

[0114] A processor, when executing a program stored in memory, implements the steps of any of the image processing methods described above;

[0115] The display is used to display the colored contact lens trial image obtained by the processor after implementing the steps of any of the image processing methods described above.

[0116] Optionally, the electronic device is a computer, mobile phone, electronic mirror, or electronic photo album.

[0117] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the image processing methods described above.

[0118] Beneficial effects of the embodiments of the present invention:

[0119] The solution provided in this invention first segments a partial iris image, including the iris region, from a facial image. Then, it performs a first fusion process with a colored contact lens image to obtain a partial colored contact lens effect image. Finally, it performs a second fusion process with the facial image to obtain a colored contact lens trial image. Because the first fusion process integrates the image style of the colored contact lens image into the partial iris image while preserving the texture features of the partial iris image, the subsequent second fusion of the partial colored contact lens effect image with the facial image ensures that the trial image incorporates the colored contact lens style while retaining the texture features of the iris region in the facial image, thus achieving a highly realistic colored contact lens wearing effect. Therefore, this solution can generate highly realistic colored contact lens trial images.

[0120] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0121] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0122] Figure 1 A flowchart of an image processing method provided in an embodiment of the present invention;

[0123] Figure 2 This is another flowchart of the image processing method provided in an embodiment of the present invention;

[0124] Figure 3 This is a flowchart of training a generative adversarial network provided in an embodiment of the present invention;

[0125] Figure 4 This is another flowchart of the image processing method provided in the embodiments of the present invention;

[0126] Figure 5 A schematic diagram of the structure of a specific apparatus for implementing the image processing method provided in the embodiments of the present invention;

[0127] Figure 6 This is a schematic diagram of an eye key point template provided in an embodiment of the present invention;

[0128] Figure 7 This is a flowchart of generating a trial image of colored contact lenses provided in an embodiment of the present invention;

[0129] Figure 8 This is a schematic diagram of the structure of an image processing apparatus provided according to an embodiment of the present invention;

[0130] Figure 9 This is a block diagram of an electronic device used to implement the image processing method provided in the embodiments of the present invention. Detailed Implementation

[0131] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.

[0132] In related technologies, when generating images of colored contact lenses for trial wear, methods such as texturing, rendering, and overlay are usually used. However, the effects of these methods are unstable under different users and different scenarios, and the images are prone to distortion, or even produce colored contact lens effects in non-eye areas. For example, when generating images of colored contact lenses for trial use, a linear summation method is used to blend the colored contact lens material with the original image. During rendering, low weights and shadow effects are considered to reduce the distortion caused by the colored contact lenses adhering to the eyelids. However, some parameters involved here require manual intervention, especially when modifying transparency, which will affect the texture of the colored contact lenses and may result in significant differences from actual wearing. Alternatively, image generation techniques, such as CycleGAN (Cycle Generative Adversarial Networks), can be used to implement colored contact lens wearing and removal. By training two generators and combining recurrent loss and GAN loss, the goal is to maintain the original texture information as much as possible while wearing or removing colored contact lenses. However, the CycleGAN method can only learn the behavior of wearing and removing colored contact lenses and cannot specify specific colored contact lens styles, resulting in a single generation effect that cannot be applied to various colored contact lens materials for trial use. Furthermore, it does not constrain the colored contact lens generation area, which may result in incomplete colored contact lenses or lenses covering the eyelids, leading to distortion.

[0133] Based on the above, in order to generate highly realistic images of contact lens try-on, embodiments of the present invention provide an image processing method, apparatus, and device.

[0134] The following is a description of an image processing method, apparatus, and device provided by an embodiment of the present invention.

[0135] The image processing method provided in this embodiment of the invention can be applied to electronic devices. In specific applications, the electronic device can be a server or a terminal device, both of which are reasonable. In practical applications, the terminal device can be a mobile phone, tablet computer, desktop computer, etc.

[0136] Specifically, the entity executing this image processing method can be an image processing device. For example, when the image processing method is applied to a terminal device, the image processing device can be functional software running on the terminal device, such as image processing software; of course, the image processing device can also be a plugin in existing functional software, such as a plugin in image processing software for generating images of colored contact lenses trying on, or a plugin in a shopping client for generating images of colored contact lenses trying on. For example, when the image processing method is applied to a server, the image processing device can be a computer program running on the server, which can be used to generate images of colored contact lenses trying on.

[0137] The image processing method provided in this embodiment of the invention may include the following steps:

[0138] Acquire images of the colored contact lenses and facial images of the face to be tested with the lenses;

[0139] Segment the iris region from the facial image;

[0140] The partial iris image and the colored contact lens image are subjected to a first fusion process to obtain a partial colored contact lens effect image; wherein, the first fusion process is used to incorporate the image style of the colored contact lens image into the partial iris image while preserving the texture features of the partial iris image;

[0141] The partial contact lens effect image and the facial image are subjected to a second fusion process to obtain a contact lens trial image; wherein, the second fusion process is used to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image.

[0142] The solution provided in this invention first segments a partial iris image, including the iris region, from a facial image. Then, it performs a first fusion process with a colored contact lens image to obtain a partial colored contact lens effect image. Finally, it performs a second fusion process with the facial image to obtain a colored contact lens trial image. Because the first fusion process integrates the image style of the colored contact lens image into the partial iris image while preserving the texture features of the partial iris image, the subsequent second fusion of the partial colored contact lens effect image with the facial image ensures that the trial image incorporates the colored contact lens style while retaining the texture features of the iris region in the facial image, thus achieving a highly realistic colored contact lens wearing effect. Therefore, this solution can generate highly realistic colored contact lens trial images.

[0143] The image processing method provided by the embodiments of the present invention will be described below with reference to the accompanying drawings.

[0144] like Figure 1 As shown, the image processing method provided in this embodiment of the invention may include steps S101-S104:

[0145] S101, acquire the colored contact lens image and the facial image of the face to be tested with the colored contact lenses;

[0146] In this embodiment, the contact lens image and the facial image can be images pre-stored in the local memory of the electronic device, or images acquired in real time. For example, the contact lens image can be a contact lens image stored on the mobile phone, or a contact lens image downloaded from a contact lens image library on a website. For example, the facial image can be a human face image or an animal face image stored on the mobile phone, or a human face image or an animal face image captured by the user through relevant software functions on the electronic device.

[0147] It should be noted that the facial images in this embodiment may come from public datasets, or the acquisition of facial images may be authorized by the corresponding users; and the collection, storage, use, and processing of facial images comply with relevant laws and regulations and do not violate public order and good morals.

[0148] S102, Segment the iris region from the facial image;

[0149] It is understandable that, since colored contact lenses are worn on the iris region of the eye, to generate a trial image of the colored contact lenses with the desired effect, the image style of the colored contact lens image can be integrated into the iris region of the facial image. After obtaining the colored contact lens image and the facial image to be tested in step S101, to integrate the image style of the colored contact lens image into the iris region of the facial image, a partial iris image containing the iris region can first be segmented from the facial image for subsequent fusion processing. Furthermore, it should be noted that during the first fusion process between the partial iris image and the colored contact lens image, to ensure the information of each pixel in the colored contact lens image is fused with the information of corresponding pixels in the partial iris image, the partial iris image segmented from the facial image can be an image containing the iris region and the same size as the colored contact lens image.

[0150] For example, in one optional implementation, the method for segmenting the iris region from the facial image can be to first identify the iris region in the facial image, and then crop the facial image using the bounding rectangle of the iris region as the cropping area to obtain the iris region image. The method for identifying the iris region in the facial image can be manual identification, machine identification, etc., and this embodiment of the invention does not limit the specific method for identifying the iris region in the facial image.

[0151] S103, perform a first fusion process on the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image; wherein, the first fusion process is used to incorporate the image style of the colored contact lens image into the partial iris image and retain the texture features of the partial iris image;

[0152] It is understandable that after segmenting the iris region from the facial image through step S102, by incorporating the image style of the colored contact lens image into the iris region and preserving the texture features of the iris region, the iris region in the fused local colored contact lens effect image can have the iris texture features of the facial image and the image style of the colored contact lens image, thereby achieving a highly realistic colored contact lens wearing effect.

[0153] Optionally, in one implementation, the partial iris image and the colored contact lens image are subjected to a first fusion process to obtain a partial colored contact lens effect image, which may include steps A1-A4:

[0154] A1, extract the style features of the colored contact lens image;

[0155] A2, extract the texture features of the local iris image;

[0156] Understandably, in order to obtain a localized image of the contact lens effect that combines the image style of the contact lens image with the iris texture features of the facial image, the style features of the contact lens image and the texture features of the localized iris image can be extracted. Subsequently, the style features and texture features are fused together to obtain an image with the fused features corresponding to the localized iris image.

[0157] In this implementation, neural networks can be used to extract the style features of the contact lens image and the texture features of the local iris image, but this is not a limitation. For example, when extracting style features using a neural network, the neural network used to extract the style features of the contact lens image can be a style encoding network based on a convolutional neural network. The output layer of this style encoding network consists of two parallel fully connected layers, used to encode the image style of the contact lens image and output it in the form of mean and variance, that is, fitting the style feature distribution of the contact lens image with a Gaussian distribution. For example, when extracting texture features using a neural network, the neural network used to extract the texture features of the local iris image can be a feature extraction network based on a convolutional neural network. This feature extraction network mainly consists of fully connected layers, which can fully utilize the information of each pixel in the local iris image for feature extraction, obtaining the texture features of the local iris image.

[0158] A3, fuse the style feature with the texture feature to obtain the fused feature corresponding to the local image of the iris;

[0159] For example, the style feature and texture feature can be fused by multiplying each value of the texture feature by the variance of the style feature and adding the mean of the style feature to obtain the fused feature corresponding to the local iris image.

[0160] A4 generates an image with the fused features corresponding to the local iris image, which serves as the local contact lens effect image.

[0161] It is understandable that after fusing the style feature with the texture feature in step A3, the resulting image, which is composed of the fused features corresponding to the local iris image, is the local contact lens effect image. It retains the iris texture features in the facial image and has the image style of the contact lens image.

[0162] S104, perform a second fusion process on the partial contact lens effect image and the facial image to obtain a contact lens trial image; wherein, the second fusion process is used to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image.

[0163] It is understandable that after obtaining the partial contact lens effect image through step S103, in order to obtain the contact lens trial image corresponding to the facial image, the partial contact lens effect image and the facial image can be subjected to a second fusion process to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image, thereby obtaining the contact lens trial image.

[0164] For example, the way to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image can be to paste the partial contact lens image onto the corresponding position in the facial image, that is, to replace the pixel information of the corresponding position in the facial image with the pixel information of each pixel in the partial contact lens image, or to perform a weighted summation of each pixel in the partial contact lens effect image and each pixel in the corresponding position in the facial image, and so on.

[0165] The solution provided in this invention first segments a partial iris image, including the iris region, from a facial image. Then, it performs a first fusion process with a colored contact lens image to obtain a partial colored contact lens effect image. Finally, it performs a second fusion process with the facial image to obtain a colored contact lens trial image. Because the first fusion process integrates the image style of the colored contact lens image into the partial iris image while preserving the texture features of the partial iris image, the subsequent second fusion of the partial colored contact lens effect image with the facial image ensures that the trial image incorporates the colored contact lens style while retaining the texture features of the iris region in the facial image, thus achieving a highly realistic colored contact lens wearing effect. Therefore, this solution can generate highly realistic colored contact lens trial images.

[0166] Alternatively, in another embodiment of the invention, in Figure 1 Based on the illustrated embodiments, as Figure 2 As shown, the step S102 above, which involves segmenting the iris region from the facial image, may include steps S1021-S1023:

[0167] S1021, Identify key eye information in the facial image;

[0168] Optionally, in one implementation, the eye region of a facial image can first be identified, and then key eye point information can be determined from this region. When identifying the eye region of a facial image, a deep learning-based object detection model can be used for rapid eye region recognition, but this is not a limitation. For example, the object detection model can be YOLO, SSD, Faster R-CNN, etc. In practical applications, this object detection model can be trained on a large number of sample facial images, each with pre-labeled eye outlines. When using the trained object detection model to identify the eye region in a facial image, the model's input is the facial image, and the output is the coordinate information of the eye socket position in the facial image, thereby achieving the identification of the eye region from the facial image.

[0169] After identifying the eye region in a facial image, a keypoint localization model obtained through deep learning can be used to quickly predict the keypoints within that region, thus obtaining keypoint information. However, this method is not limited to this. For example, the keypoint localization model can be a regression model, a heatmap model, or other deep learning models. In practical applications, this keypoint localization model can be trained on a large number of sample eye images, each with pre-labeled keypoints. When using the trained keypoint localization model to predict keypoints, the model's input is the eye image from the facial image, and the output is the predicted keypoint information, thus obtaining the keypoint information of the eyes in the facial image.

[0170] Alternatively, in another implementation, the facial image can be input into a pre-trained eye keypoint recognition model to quickly obtain eye keypoint information from the facial image. The eye keypoint recognition model can be any deep learning model. In practical applications, the eye keypoint recognition model can be trained using a large number of sample facial images, each with pre-labeled eye keypoints. When using the trained keypoint recognition model to predict eye keypoints, the model's input is the facial image, and the output is the predicted eye keypoints, thus obtaining the eye keypoint information from the facial image.

[0171] It should be emphasized that the specific implementation of identifying key eye points in the facial image described above is merely an example and should not be construed as limiting the embodiments of the present invention.

[0172] Additionally, it should be noted that eye key points can include key points used to depict the contours of the upper and lower eyelids, iris, pupil, etc., within the eye region. The eye key point information can be the coordinates of the key point within the eye region, or the coordinates of the key point within the facial image, and so on. For example, such as... Figure 4 As shown, the key point information of the eye can include the coordinate information of multiple key points of the eye shown in the figure (0-15). Of course, in order to better utilize the key point information of the eye to fit the contour curves in the eye region, the key point information of the eye can also include the coordinate information of more key points. This embodiment of the invention does not limit the number of key points of the eye.

[0173] S1022, Using the eye key point information, determine the iris region in the facial image;

[0174] It is understandable that after obtaining the key eye information in the facial image through step S1021, since the key eye information includes multiple key points such as the upper eyelid, lower eyelid, iris, and pupil, the iris region in the facial image can be determined using this key eye information.

[0175] Optionally, in one implementation, using the eye key point information to determine the iris region in the facial image may include steps B1-B3:

[0176] B1. Using the key point information of the eye, fit the eye contour in the facial image to obtain the eye fitting curve; wherein, the eye contour is the contour that fits the eyelid.

[0177] For example, when fitting the eye contour in a facial image using eye keypoint information, the shapes of the upper and lower eyelid contours can be matched using parabolas. Specifically, the parabola matching the upper eyelid keypoint information is calculated as the upper eyelid fitting curve, and the parabola matching the lower eyelid keypoint information is calculated as the lower eyelid fitting curve. Then, the upper and lower eyelid fitting curves are combined to form the overall eye fitting curve. For example, as... Figure 6 As shown, the parabola corresponding to the upper eyelid can be calculated based on the coordinate information of the key eye points labeled 0, 2, 3, 4, 1, and the parabola corresponding to the lower eyelid can be calculated based on the coordinate information of the key eye points labeled 0, 5, 6, 7, 1, thus obtaining the eye fitting curve composed of the parabola corresponding to the upper eyelid and the parabola corresponding to the lower eyelid.

[0178] Additionally, it should be noted that when calculating the matching parabola using the key points of the upper and lower eyelids respectively, in order to achieve a better curve fitting effect, the upper eyelid parabola can be fitted piecewise using every three adjacent key points in the upper eyelid key points, and the lower eyelid parabola can be fitted piecewise using every three adjacent key points in the lower eyelid key points, thus obtaining the eye fitting curve.

[0179] B2. Based on the eye fitting curve and the contour parameters of the iris contour, a foreground masking process is performed on the facial image to obtain a mask image of the facial image; wherein, the foreground masking process is a masking process that uses the iris region of the facial image as the foreground region; the contour parameters of the iris contour are generated based on specified key points in the eye key point information, and the specified key point information is key point information located on the contour of the outer circle of the iris;

[0180] In this implementation, when generating the contour parameters of the iris contour using specified key point information, the iris contour is regarded as a perfect circle. The parameters of the matching circle are calculated using the specified key point information, and the contour parameters of the iris contour can be obtained.

[0181] B3. Based on the mask image of the facial image, determine the iris region in the facial image.

[0182] It is understandable that, given that some facial images have eyelids covering part of the iris region, directly using the aforementioned key points to fit the iris contour of the iris region in such facial images may result in the fitted iris contour exceeding the range of the eye contour. Consequently, when the local iris image is segmented from the facial image and then fused with the contact lens image for the first time, the resulting local contact lens effect image will be distorted.

[0183] To address the aforementioned distortion issue, it is necessary to accurately determine the iris region in the facial image. In this implementation, by combining the eye fitting curve and the contour parameters of the iris outline, a foreground mask is applied to the facial image to obtain a mask image. Then, based on this mask image, the iris region in the facial image is determined.

[0184] For example, in one specific implementation, based on the eye fitting curve and the contour parameters of the iris contour, a foreground masking process is performed on the facial image to obtain a mask image of the facial image. This includes: identifying a specified pixel in the facial image based on the eye fitting curve and the contour parameters of the iris contour; wherein the specified pixel is located within the region formed by the eye contour and within the region formed by the contour of the outer circle of the iris; setting the pixel value of the specified pixel in the facial image to a first value, and setting the pixel values ​​of all pixels in the facial image except the specified pixel to a second value, thereby obtaining the mask image of the facial image. For example, the pixel values ​​of pixels located inside the region formed by the eye fitting curve and within the region formed by the contour of the outer circle of the iris in the facial image can be assigned a value of 1, and the pixel values ​​of pixels at other locations can be assigned a value of 0, thereby obtaining the mask image of the facial image, and the region with a pixel value of 1 in the mask image is determined as the iris region. For example, as... Figure 6 As shown, the pixel values ​​of the pixels inside the eye fitting curve (i.e., the eye fitting curve fitted using the eye key point information labeled 0, 2, 3, 4, 1, 5, 6, 7) and inside the iris contour (i.e., the iris contour fitted using the eye key point information labeled 8, 9, 10, 11) are assigned to 1, and the pixel values ​​of the pixels in other positions are assigned to 0, thus obtaining the mask image of the face image.

[0185] S1023, perform partial iris cropping on the facial image to obtain a partial iris image.

[0186] After determining the iris region in the facial image through step S1022, the facial image can be cropped to obtain a partial iris image. The partial iris image cropping is used to crop out a partial image containing the iris region as the partial iris image.

[0187] As can be seen, by using this method, after identifying the key eye points in the facial image, the iris region in the facial image can be determined using the key eye points, thereby enabling quick and accurate cropping of the local iris region of the facial image to obtain a local iris image.

[0188] Optionally, in another embodiment of the present invention, the step of performing a first fusion process on the partial iris image and the colored contact lens image in step S103 to obtain a partial colored contact lens effect image can be implemented by a pre-trained fusion network; wherein, the fusion network is a neural network trained based on the sample partial iris image and the sample colored contact lens image, used to output the fusion result; wherein, the fusion result is the result of incorporating the image style of the sample colored contact lens image into the sample partial iris image, while retaining the texture features of the sample partial iris image.

[0189] In other words, a pre-trained fusion network is used to perform the first fusion process on the local iris image and the contact lens image. In practical applications, this fusion network can be a convolutional neural network, a fully connected neural network, etc.

[0190] Optionally, in one implementation, the pre-trained fusion network can be the generator network in a pre-trained generative adversarial network. As those skilled in the art will know, generative adversarial networks are one of the unsupervised learning methods on complex distributions in recent years. Generative adversarial networks produce fairly good outputs through the mutual game learning between two modules in the network framework: the generator network and the discriminator network.

[0191] Correspondingly, in this implementation method, such as Figure 3 As shown, the training process of this generative adversarial network may include steps S301-S307:

[0192] S301, acquire a partial image of the sample iris and an image of the sample contact lenses;

[0193] In this implementation, a sample face image and a sample contact lens image can be obtained first, and then a sample iris local image containing the iris region can be segmented from the sample face image. The specific segmentation process is similar to step S102 above, and will not be repeated here.

[0194] S302, input the sample iris local image and the sample colored contact lens image into the generation network to obtain the generated image as the fusion result;

[0195] In this implementation, the generator network can be composed of an encoder and a decoder based on a convolutional neural network. After the sample iris local image and the sample contact lens image are input into the generator network, the encoder first performs downsampling feature encoding on the input image, that is, downsampling and compressing the image information and obtaining multi-resolution feature encoding to obtain the encoding result. Then, the decoder upsamples according to the encoding result to restore the resolution of the encoding result to the size of the input image. The output of the decoder is the generated image as the fusion result.

[0196] S303, the generated image and the sample iris image are respectively input into the adversarial network in the generative adversarial network to obtain a first discrimination result corresponding to the generated image and a second discrimination result corresponding to the sample iris local image; wherein, the adversarial network is used to discriminate whether the input image is a real image;

[0197] In this implementation, the adversarial network can be composed of a convolutional network with a downsampling mechanism. It takes an image as input and outputs a discrimination result of whether the image is a real image. That is, the discrimination result can be a real image or a generated image.

[0198] S304, Based on the difference between the first identification result and the corresponding true value, and the difference between the second identification result and the corresponding true value, calculate the loss value of the adversarial network;

[0199] Understandably, since the training objective of adversarial networks is to effectively distinguish between generated and real images, the loss value of the adversarial network can be calculated based on the difference between the first discrimination result and its corresponding ground truth value, and the difference between the second discrimination result and its corresponding ground truth value. Since the generated image is generated based on a sample iris local image and a sample contact lens image, and the sample iris local image is cropped from a sample face image, the ground truth value for the first discrimination result is the generated image, and the ground truth value for the second discrimination result is the real image. For example, the differences between the first and second discrimination results and their corresponding ground truth values ​​can be calculated using a loss function, such as an L1 loss function, an L2 loss function, etc.

[0200] S305, Based on the difference between the first identification result and the corresponding false value, calculate the loss value of the generator network;

[0201] Understandably, since the generated image is based on sample iris images and sample contact lens images, the ground truth value of the first discrimination result is the generated image, and the false value is the real image. For the generative network, the training objective is to generate images that are as realistic as possible. Therefore, when training the generative network, it needs to be combined with an adversarial network to achieve this goal. That is, when the adversarial network identifies the generated image as a real image, it can be considered that the generative network has generated an image that approximates the real image. During the training process of the generative network, since the difference between the first discrimination result and its corresponding false value represents the loss value for the generated image to be identified as a real image, the loss value of the generative network can be calculated based on the difference between the first discrimination result and the corresponding false value.

[0202] S306, Based on the loss value of the adversarial network and the loss value of the generator network, determine whether the adversarial network and the generator network have reached a Nash equilibrium state; wherein, the Nash equilibrium state is used to characterize the stable state in which the loss values ​​of the adversarial network and the generator network fluctuate within a specified range.

[0203] Understandably, since training generative networks (GANs) and adversarial networks (AANs) are essentially contradictory processes, it's difficult for them to converge during training. The optimal state is generally a Nash equilibrium, where the generated images are as realistic as possible, and the AAN can no longer distinguish between real and generated images. Nash equilibrium represents a stable state where the loss values ​​of both the GAN and the AAN fluctuate within a specified range, which can be a range around a certain value. During GAN training, when the GAN's loss value decreases to a certain level and fluctuates slightly around a certain value, if the AAN's loss value also fluctuates slightly around a certain value as the GAN optimizes, then the GAN and GAN are considered to have reached a Nash equilibrium, and training can be stopped. Otherwise, the GAN has not reached the stopping condition and training needs to continue with parameter adjustments.

[0204] In this implementation, the loss value of the adversarial network represents the loss between the discrimination result output by the adversarial network and the corresponding ground truth value, that is, the probability value of the adversarial network misidentifying the input image. The loss value of the generator network represents the loss value of the generated image output by the generator network being identified as a real image. During the training process of the generator adversarial network, the optimization of the generator adversarial network can be analyzed by using the loss values ​​of the adversarial network and the generator network, thereby determining whether the adversarial network and the generator network have reached a Nash equilibrium state.

[0205] S307, If the adversarial network and the generator network do not reach Nash equilibrium, adjust the parameters of the generator adversarial network and return to the step of acquiring the sample iris local image and the sample contact lens image.

[0206] If step S306 determines that the adversarial network and the generator network have not reached a Nash equilibrium, then the generator adversarial network has not met the stopping condition. Therefore, based on the loss values ​​of the adversarial network and the generator network, the parameters of the adversarial network and the generator network can be adjusted respectively. Then, the process returns to the step of acquiring sample iris images and sample contact lens images to continue training the generator adversarial network until they reach a Nash equilibrium, at which point training stops. It should also be noted that other methods can be used to determine whether the generator adversarial network has reached the stopping condition. For example, when training the generator adversarial network, an iteration count can be set, and reaching the stopping condition can be achieved by reaching the set number of iterations. This embodiment does not limit the method used to determine whether the generator adversarial network has reached the stopping condition.

[0207] As can be seen, by using the pre-trained fusion network to perform the first fusion process on the local iris image and the contact lens image, the fused local contact lens effect image can be generated quickly and automatically.

[0208] Optionally, in another embodiment of the present invention, before calculating the loss value of the generator network based on the difference between the first discrimination result and the corresponding false value in step S305 above, the method may further include steps D1-D2:

[0209] D1: Input the generated image and the sample iris local image into the feature extraction network respectively, and calculate the first loss value based on the difference between the output results of the feature extraction network for the generated image and the sample iris image;

[0210] In this embodiment, the feature extraction network can be a feature encoding network mainly composed of fully connected layers, used to analyze the differences in texture and contour between the sample iris local image and the generated image during the encoding process. The first loss value can be the sum of the differences in the output results of each layer in the feature extraction network, that is, the differences in iris texture and eye contour between different layers and different locations of features generated during the encoding process of the sample iris local image and the generated image, or it can be the difference between the output results of the last layer of the feature extraction network. For example, the loss function for calculating this first loss value can be the cross-entropy loss function, the L1 loss function, etc.

[0211] D2, the generated image and the sample contact lens image are respectively input into the style constraint network, and a second loss value is calculated based on the difference between the output results of the style constraint network for the generated image and the sample contact lens image; wherein, the style constraint network is a neural network used to identify the style of the image;

[0212] For example, the style constraint network can be a pre-trained convolutional neural network, such as a VGG network, a fully connected neural network, etc. The generated image and the sample contact lens image are input into the style constraint network respectively. The style constraint network can extract style features from the different images, thereby calculating the difference in style features between the generated image and the sample contact lens image, and obtaining a second loss value. The loss function used to calculate the second loss value can be an L1 loss function, an L2 loss function, etc.

[0213] Accordingly, in this embodiment, step S305 above, calculating the loss value of the generator network based on the difference between the first discrimination result and the corresponding false value, may include:

[0214] Based on the difference between the first identification result and the corresponding false value, as well as the first loss value and the second loss value, the loss value of the generator network is calculated.

[0215] In this embodiment, the difference between the first identification result and the corresponding false value represents the loss value for the generated image to be identified as a real image. The smaller the difference, the closer the features of the generated image are to the features of the real image, that is, the closer they are to the features of the sample iris local image. The first loss value represents the difference in iris texture and eye contour between the generated image and the sample iris local image during the encoding process. The smaller the first loss value, the closer the texture features and eye contour of the generated image are to the features of the real image. The second loss value represents the difference in style features between the generated image and the sample colored contact lens image. The smaller the second loss value, the closer the style features of the generated image are to the style features of the real image. Therefore, in order for the image generated by the trained generative network to retain both the texture features of the sample iris local image and the style features of the sample colored contact lens image, the loss value of the generative network can be calculated based on the difference between the first identification result and the corresponding false value, the sum of the first loss value and the second loss value, so that the parameters of the generative network can be adjusted subsequently based on the loss value of the generative network.

[0216] Understandably, since the smaller the difference between the first discrimination result and the corresponding false value, the smaller the first loss value and the smaller the second loss value, the better the simulation effect of the generated image is when training the generative network, the parameters of the generative network can be adjusted based on the loss value of the generative network if the generative adversarial network does not meet the stop training condition after calculating the loss value of the generative network.

[0217] As can be seen, by combining feature extraction network and style constraint network to assist in training the generative network, the pre-trained generative network can learn the style of colored contact lenses in the colored contact lens image and retain the texture features in the local iris image. Thus, the trained generative network can be used to generate local colored contact lens effect images with better simulation effect.

[0218] Optionally, in another embodiment of the present invention, the generation network in step S302 above includes a style coding network, an encoder, and a decoder;

[0219] Accordingly, in this embodiment, the step S302 above, in which the sample iris local image and the sample contact lens image are input into the generative network to obtain the generated image as the fusion result, may include steps E1-E3:

[0220] E1, input the sample colored contact lens image into the style coding network to obtain the style features of the sample colored contact lens image;

[0221] Understandably, to enable the generated image output by the generative network to better learn the contact lens style from the sample contact lens image, the style features of the sample contact lens image can be introduced into the encoding process when encoding the sample contact lens image and the sample iris local image, thereby influencing the image style of the fused generated image. Since image style can be fitted using a Gaussian distribution, and the parameters used in the Gaussian distribution are the mean and variance, in this embodiment, a style encoding network can be used to process the sample contact lens image to obtain its mean and variance, which can then be used as the style features of the sample contact lens image. These style features can then be incorporated into the subsequent encoding process to influence the image style of the fused generated image.

[0222] For example, the style encoding network can be a convolutional neural network, the output layer of which consists of two parallel fully connected layers used to encode the sample contact lens image and output the encoding in the form of mean and variance. Of course, the style encoding network can also be other types of neural networks, such as deep neural networks, etc. The specific type of style encoding network is not limited in the embodiments of the present invention.

[0223] E2, the style features of the sample contact lens image and the sample iris local image are input into the encoder for encoding to obtain the fused features corresponding to the sample iris local image; the fused features corresponding to the sample iris local image are feature images that fuse the style features of the sample contact lens image and the texture features of the sample iris local image.

[0224] In this embodiment, the encoder can be a convolutional neural network, which can be composed of multiple network layers. In order to make the fused features obtained after encoding have a better color contact lens style in the sample color contact lens image, the style features of the sample color contact lens image can be input into the last few network layers of the encoder to modify the features in the sample iris local image. Thus, the style features of the sample color contact lens image are imported into the style features of the sample iris local image to obtain the fused features corresponding to the sample iris local image.

[0225] For example, if the style features of the colored contact lens image are fitted with a Gaussian distribution, the features in the local iris image of the sample can be modified using the following formula:

[0226] in, It is a characteristic after fusion, L i σ represents the features before fusion, i.e., the features in the local image of the sample iris; σ refers to the variance of the sample colored contact lens image, and μ refers to the mean of the sample colored contact lens image.

[0227] E3 inputs the fused features corresponding to the local iris image of the sample into the decoder for decoding, and obtains the generated image as the fusion result.

[0228] Understandably, in order to obtain an output image of the same size as the input image, after encoding and feature fusion of the sample iris local image and the sample contact lens image, the fused features corresponding to the sample iris local image can be input into the decoder for decoding to obtain the generated image.

[0229] As can be seen, by adding a style encoding network to the generative network, the style features of the sample contact lens images can be extracted through this scheme. Integrating the extracted style features into the sample iris local image allows the generated image to better learn the contact lens style of the sample contact lens image.

[0230] Alternatively, in another embodiment of the invention, in Figure 1 Based on the illustrated embodiments, as Figure 4 As shown, in step S104 above, the second fusion process is performed on the partial contact lens effect image and the facial image to obtain the contact lens trial image, which may include steps S1041-S1043:

[0231] S1041, Based on the partial contact lens effect image, generate a backup effect image; wherein, the size of the backup effect image is the same as the size of the target eye image, and the target area in the backup effect image has the image information of the iris area of ​​the partial contact lens effect image, the target eye image is the eye area image of the face image, and the target area is the area with the same position as the iris area of ​​the eye area image.

[0232] Understandably, to integrate the image style of a partial contact lens effect image into the eye area of ​​a facial image, a backup effect image with the same size as the eye area of ​​the facial image can be generated based on the partial contact lens effect image. In other words, the partial contact lens effect image is adjusted to the same size as the eye area of ​​the facial portrait for subsequent fusion. Then, the fusion coefficient corresponding to each pixel in the backup effect image is determined, and the backup effect image and the eye area of ​​the facial image are fused using this fusion coefficient to obtain a contact lens trial image corresponding to the facial image.

[0233] Optionally, in one implementation, generating a backup effect image based on the partial contact lens effect image may include steps F1-F2:

[0234] F1 identifies the iris region in the localized contact lens effect image as the area to be utilized.

[0235] F2 replaces the image information of the target region in the specified image with the image information of the region to be used, to obtain a backup effect image; wherein, the specified image is the same size as the target eye image.

[0236] Understandably, to ensure that the target area in the generated backup effect image contains image information of the iris region from the partial contact lens effect image, the iris region in the partial contact lens effect image can first be identified as the area to be used. Then, the information of the target area in the specified image is replaced with the image information of the area to be used, thereby obtaining the backup effect image. For example, the specified image can be a human eye image of the same size as the target eye image, or an image with all zeros, etc.

[0237] For example, in one specific implementation, determining the iris region in the local contact lens effect image as the region to be utilized in step F1 above may include steps F11-F13:

[0238] F11 determines the target mask image of the local iris image; wherein, the target mask image is a binarized image with the iris region as the foreground;

[0239] In this implementation, in order to determine the iris region in the local colored contact lens image, a target mask image of the local iris image can be determined first. That is, the local iris image is masked with the iris region as the foreground. In the obtained target mask image, the pixel value of the pixel corresponding to the iris region is 1, and the pixel value of the pixel corresponding to the pixel outside the iris region is 0.

[0240] F12 multiplies the pixel values ​​of each pixel in the local contact lens effect image with the pixel values ​​of the corresponding pixels in the target mask image;

[0241] F13 determines the iris region in the local contact lens effect image based on the pixel values ​​of the pixels in the image obtained after multiplication, and uses it as the region to be utilized.

[0242] Understandably, since the pixel values ​​of pixels outside the iris region in the target mask image are 0, and the pixel values ​​of pixels corresponding to the iris region are 1, multiplying the pixel values ​​of each pixel in the local contact lens effect image with the corresponding pixel values ​​in the target mask image results in the pixel values ​​of each pixel corresponding to the iris region in the local contact lens effect image, while the pixel values ​​of pixels outside the iris region are 0. Therefore, the region with non-zero pixels after multiplication can be identified as the iris region in the local contact lens effect image, i.e., the region to be utilized.

[0243] S1042, Determine the fusion coefficient corresponding to each pixel in the backup effect image;

[0244] Understandably, by determining a corresponding fusion coefficient for each pixel in the backup effect image, the fusion effect can be adjusted based on this coefficient when subsequently merging the backup effect image with the eye region in the facial image. For example, the fusion coefficient can be determined by relevant technical personnel based on experience, or it can be determined based on the distance of each pixel in the backup effect image from the target area.

[0245] S1043, using the determined fusion coefficient, the backup effect image and the eye area in the facial image are fused to obtain the contact lens trial image.

[0246] It is understandable that, after determining the fusion coefficients corresponding to each pixel in the backup effect image through step S1042, the backup effect image and the eye area in the facial image are fused together, so that the eye area in the facial image is incorporated into the features of the target area in the backup effect image, thereby obtaining a colored contact lens trial image with the effect of wearing colored contact lenses.

[0247] For example, in one specific implementation, the specified image is an image with all zeros, and the fusion coefficient is determined based on the distance between each pixel in the alternative effect image and the target region;

[0248] For example, the method for determining the fusion coefficient of each pixel in the backup effect image based on its distance from the target area can be: the closer a pixel is to the target area, the larger its fusion coefficient. This ensures that when the backup effect image is subsequently fused with the eye area in the facial image, the farther away from the iris region, the fewer features of the target area in the backup effect image are retained, thus achieving a smooth transition and reducing the visual texture appearance.

[0249] Accordingly, in this specific implementation, the determined fusion coefficient is used to fuse the backup effect image and the eye area in the facial image to obtain a contact lens trial image, including:

[0250] According to a predetermined calculation formula, using a determined fusion coefficient, the backup effect image and the eye area of ​​the facial image are fused to obtain a contact lens trial image; wherein, the predetermined calculation formula is:

[0251] C E =W E CM E +(1-W E E; where C E For the image of trying on the colored contact lenses, WE CM is the fusion coefficient. E For backup effect image, E is the eye area in the facial image.

[0252] It is understandable that, because the closer to the target area, the more W corresponds to each pixel in the backup effect image. E The larger the value, the closer it is to 1, the farther away it is from the target area. E The smaller the value, that is, the closer it is to 0, the more the eye area in the backup effect image and the facial image are fused using the above formula. In the resulting contact lens trial image, the features closer to the iris area are closer to the features of the target area in the backup effect image, and the features farther away from the iris area are closer to the features of the eye area in the facial image, thus making the edge of the iris area have a smooth transition effect.

[0253] As can be seen, by using this method, a corresponding fusion coefficient is determined for each pixel in the backup effect image. When the backup effect image is fused with the eye area in the facial image, the fusion effect can be adjusted according to the fusion coefficient.

[0254] To better understand the embodiments of the present invention, the following will be combined with... Figure 5 , Figure 6 and Figure 7 Specific examples of embodiments of the present invention will be described below.

[0255] like Figure 5 As shown, an apparatus for implementing the image processing method provided in the embodiments of the present invention is illustrated. The apparatus includes a facial image acquisition module, an eye detection module, an eye key point localization module, an iris segmentation module, a colored contact lens material acquisition module, a colored contact lens generation module, and an image output module. A specific example of the process for generating a colored contact lens trial image using this apparatus is as follows:

[0256] (1) Obtain a facial image of the user without colored contact lenses from the terminal device through the facial image acquisition module.

[0257] (2) Input the facial image into the eye detection module, which is used to detect the position of the eye region in the facial image and output the coordinate information of the facial image and all eye frame positions relative to the facial image.

[0258] The eye detection module can include a detection model trained on a large batch of labeled facial images and corresponding eye labels. The model takes a facial image as input and outputs the coordinates of the eye outline position. This trained detection model is then applied to the eye detection module to detect the eye outline position. The detection model can be a deep learning model such as YOLO (a detection model that redefines object detection as a regression problem), SSD (Single Shot MultiBoxDetector, a first-level object detection model), or Faster R-CNN (Faster Regions with CNN features, a fast end-to-end deep learning detection model).

[0259] (3) Input the facial image and the coordinate information of the eye socket position relative to the facial image output in step (2) into the eye key point localization module. The eye key point localization module is used to crop the eye area image from the facial image according to the coordinate information of the eye socket position relative to the facial image, and then use the key point detector to locate the coordinate information of the key points required for the eye, and output the facial image and all eye key point information from different positions relative to the facial image.

[0260] The eye keypoint localization module can include a keypoint localization model trained based on a large number of calibrated eye images and corresponding eye keypoint labels. The input of the model is the eye image, and the output is the keypoint prediction result. The trained keypoint localization model is then applied to the eye keypoint localization module to achieve eye keypoint prediction.

[0261] When training the keypoint localization model, the eye keypoint labels for the eye image can be assigned as follows: Figure 6 The key eye points are marked using the template shown. For example... Figure 6 As shown, numbers 0 and 1 represent the left and right corners of the eye, respectively; numbers 2, 3, and 4 represent the four equal division points of the upper eyelid; numbers 5, 6, and 7 represent the four equal division points of the lower eyelid; numbers 8, 9, 10, and 11 represent the left, upper, right, and lower four equal division points of the outer circle of the iris; and numbers 12, 13, 14, and 15 represent the left, upper, right, and lower four equal division points of the inner circle of the iris.

[0262] The keypoint localization model can be a deep learning-based regression model, which takes an eye image as input and outputs a regression offset Δp. In actual use, the regression offset is combined with the template point position information p0 to obtain the final keypoint localization result p1 = Δp + p0. Alternatively, the keypoint localization model can be a deep learning-based heatmap model, which takes an eye image as input and outputs a heatmap of keypoints corresponding to the eye image.

[0263] (4) Input the facial image and eye key point information output in step (3) into the iris segmentation module. The iris segmentation module fits the eye fitting curves of different eyes according to the input facial image and eye key point information, and draws the iris foreground mask of the corresponding facial image (corresponding to the mask image of the facial image above) according to the eye fitting curve and the contour parameters of the iris contour, and outputs the facial image, the iris foreground mask of the facial image and the eye key point information.

[0264] This scheme treats the iris contour as a perfect circle. The keypoint localization model predicts four equally spaced points within the circle, aligned horizontally or vertically with the center. These four points provide the circle's specific parameters. The scheme uses a parabola for template matching of the eyelid contour shape, fitting a segmented parabola using every three adjacent points. For example, for keypoints numbered 0-4 on the upper eyelid, the parabola for segment 0-2 is fitted using keypoints numbered 0, 2, and 3; the parabola for segment 2-3-4 is fitted using keypoints numbered 2, 3, and 4; and the parabola for segment 4-1 is fitted using keypoints numbered 3, 4, and 1. The selection and fitting strategy for the lower eyelid is similar. The number and types of keypoints mentioned above are merely examples; the number of keypoints can be increased to achieve optimal shape fitting.

[0265] After obtaining the fitted curve parameters of the iris contour, upper eyelid contour, and lower eyelid contour, the contour parameter of the iris contour is denoted as (c out ,r out ,R out ), where c out Let r be the x-coordinate of the center of the circle. out Let R be the ordinate of the center of the circle. out Let be the radius. The parameters of the upper eyelid contour are (a... u ,b u ,c u ), where a u b u c u Let be the coefficients of the parabola fitted to the upper eyelid. The parameters of the lower eyelid contour are (a...). d b d ,c d ), where a d b d c d The coefficients are used to fit the parabola of the lower eyelid. The iris foreground mask can be obtained from the following formula:

[0266]

[0267] Where Mask(i,j) is the iris foreground mask, i represents the horizontal axis coordinate, and j represents the vertical axis coordinate. This formula means that the pixel value is 1 in the area below the upper eyelid contour, above the lower eyelid contour, and within the iris contour area, and the pixel value is 0 in other areas.

[0268] (5) Obtain an image of a contact lens through the contact lens material acquisition module and output the image;

[0269] (6) Input the facial image and key eye information output in step (3) into the contact lens generation module. The contact lens generation module is used to first preprocess the facial image based on the facial image and key eye information, that is, to crop the iris region of the facial image with the circumscribed square of the iris outline to obtain a partial iris image; then input the partial iris image and the contact lens image output in step (5) into the contact lens generator to obtain and output a partial contact lens effect image.

[0270] The contact lens generator can be a deep learning-based fusion network. The input to this fusion network is a contact lens image and a partial iris image, and the output is a partial contact lens effect image. Specifically, this fusion network can be a generative network within a generative adversarial network (GAN), or other deep learning models such as VQ-VAE.

[0271] (7) Input the local contact lens effect image output in step (6) and the facial image, iris foreground mask and eye key point information output in step (4) into the image output module. The image output module is used to calculate the fusion weight matrix using the iris foreground mask. Based on the fusion weight matrix, the local contact lens effect image of the iris and the eye area in the facial image are fused to obtain the contact lens trial effect image. Based on the contact lens trial effect image, the final facial contact lens effect image is obtained.

[0272] like Figure 7 As shown, this illustrates how the image output module processes the eye region E and the iris foreground mask M in a facial image. E and key information about the eyes P E Partial view of the effect of colored contact lenses (C) iris The processing procedure includes the following steps:

[0273] Step 1: Based on the iris foreground mask M E Calculate the weight matrix W, which is the same size as the region. E The weight matrix has the following characteristics: the weight corresponding to the iris foreground region is 1; for non-iris foreground regions, the closer to the foreground, the larger the weight and the closer to 1; the farther away from the foreground region, the smaller the weight and the closer to 0. This weight matrix can be generated using the OpenCV distance matrix function from the open-source library, or a custom calculation method can be designed.

[0274] Step 2: Based on the iris foreground mask M E and key information about the eyes P E Determine the mask image M in the facial image that is the same size as the local image of the iris. iris The mask image is then multiplied with the local contact lens effect image to obtain the foreground contact lens effect image.

[0275] Step 3: In the image with all zeros in the size of the eye area (corresponding to the specified image above), restore the foreground contact lens effect image obtained in Step 2 to the original iris foreground position in the eye area to obtain the backup effect image;

[0276] Step 4: Combine the backup effect image obtained in Step 3 and the eye area in the facial image with a weighted sum according to the weight matrix obtained in Step 1 to obtain the contact lens trial image;

[0277] After obtaining the contact lens trial image in step 4, the contact lens trial image can be restored onto the facial image, for example, pasted back onto the facial image, to obtain a facial contact lens effect image that presents the overall facial effect of wearing contact lenses.

[0278] To better understand the principle behind generating localized contact lens effect images using a fusion network, the following section uses a generative adversarial network (GAN) as an example to introduce the process of generating localized contact lens effect images using this fusion network. This GAN consists of two stages: pre-training and practical application.

[0279] During the training phase, it is necessary to train the generative network and the adversarial network. Since this solution needs to learn the color contact lens style and keep the original iris texture features unchanged, it is necessary to add an additional style encoding network, feature extraction network and style analysis network as auxiliary.

[0280] The specific training phase includes the following steps:

[0281] Step 1: Input the sample iris partial image and the sample colored contact lens image into the generator network to obtain the generated image.

[0282] Step 2: Calculate the loss value D of the adversarial network. loss The adversarial network is then optimized based on this loss value. The training objective of the adversarial network is to effectively distinguish between generated and real images, that is, to make the features corresponding to the generated image as close to false as possible, and the features corresponding to the real image as close to true as possible. Specifically, a first loss value between generated and false images and a second loss value between real and true images can be calculated, and then these are weighted and summed to obtain the loss value D of the adversarial network. loss Then according to D loss Backpropagation updates the parameters of the adversarial network.

[0283] Step 3: Calculate the loss value G of the generator network. loss The generator network is then optimized based on this loss value. The training objective of the generator network is to make the generated images resemble real-world images while preserving the iris texture and eye contour features of the sample iris images, and also to achieve the look of colored contact lenses. This loss value mainly consists of three parts: adversarial loss L... GAN (Corresponding to the difference between the first identification result and the corresponding false value mentioned above), block-based noise adversarial estimation loss L PatchNCE (corresponding to the first loss value mentioned above) and style loss L Style (Corresponding to the second loss value mentioned above). Wherein, the block-based noise adversarial estimation loss L... PatchNCE The style loss L represents the differences in iris texture and eye contour features at different layers and locations during the encoding process between the sample iris local image and the generated image. Style To differentiate the style features between the generated image and the sample contact lens image, the loss value of the generative network can be calculated using the following formula:

[0284] G loss =L GAN +L PatchNCE +L Style

[0285] Then, according to G loss Backpropagation updates the generated network parameters.

[0286] During the usage phase, only the generator network needs to be used. The specific usage phase includes the following steps:

[0287] Step 1: Input the partial iris image and the colored contact lens image into the generative network;

[0288] Step 2: Generate a partial image of the effect of wearing colored contact lenses using a generative network. This partial image of the effect of wearing colored contact lenses has the iris texture and outline of the partial iris image, as well as the colored contact lens style of the colored contact lens image.

[0289] Step 3: Output the generated partial colored contact lens effect image.

[0290] As can be seen, this solution allows users selecting colored contact lenses to see the effect without actually wearing them, making the selection process more convenient, hygienic, and faster. By combining deep learning technology, this solution quickly obtains clear images of the lenses being worn, exhibits strong adaptability to different people and scenarios, and eliminates the need for manual parameter adjustments, resulting in a better and more intuitive user experience. Furthermore, the generated images effectively preserve the internal texture features of the iris in the facial image, leading to better simulation results. This solution also has significant reference value for future research on iris recognition technology obscured by colored contact lenses, or for eye-related research in colored contact lens scenarios.

[0291] Corresponding to the above method embodiments, this invention also provides an image processing apparatus, such as... Figure 8 As shown, the device includes:

[0292] The acquisition module 810 is used to acquire images of colored contact lenses and facial images of the face to be tested with colored contact lenses;

[0293] The segmentation module 820 is used to segment a partial iris image containing the iris region from the facial image;

[0294] The first fusion module 830 is used to perform a first fusion process on the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image; wherein, the first fusion process is used to incorporate the image style of the colored contact lens image into the partial iris image and retain the texture features of the partial iris image;

[0295] The second fusion module 840 is used to perform a second fusion process on the partial contact lens effect image and the facial image to obtain a contact lens trial image; wherein, the second fusion process is used to integrate the image style of the partial contact lens effect image into the eye area of ​​the facial image.

[0296] Optionally, the segmentation module includes:

[0297] The recognition submodule is used to recognize key eye point information in the facial image;

[0298] The first determining submodule is used to determine the iris region in the facial image using the eye key point information;

[0299] The cropping submodule is used to crop the iris region of the facial image to obtain a partial iris image.

[0300] Optionally, the determining submodule is specifically used for:

[0301] The eye contour in the facial image is fitted using the key eye point information to obtain an eye fitting curve; wherein, the eye contour is the contour that fits the eyelid.

[0302] Based on the eye fitting curve and the contour parameters of the iris contour, the facial image is subjected to foreground masking to obtain a mask image of the facial image; wherein, the foreground masking is a masking process that uses the iris region of the facial image as the foreground region; the contour parameters of the iris contour are generated based on specified key points in the eye key point information, and the specified key point information is key point information located on the contour of the outer circle of the iris.

[0303] Based on the mask image of the facial image, the iris region in the facial image is determined.

[0304] Optionally, the step of performing foreground masking processing on the facial image based on the eye fitting curve and the contour parameters of the iris contour to obtain a mask image of the facial image includes:

[0305] Based on the eye fitting curve and the contour parameters of the iris contour, a specified pixel in the facial image is identified; wherein the specified pixel is located within the area formed by the eye contour and within the area formed by the contour of the outer circle of the iris.

[0306] The pixel value of a specified pixel in the facial image is set to a first value, and the pixels in the facial image other than the specified pixel are set to a second value to obtain a mask image of the facial image.

[0307] Optionally, the second fusion module includes:

[0308] The first generation submodule is used to generate a backup effect image based on the partial contact lens effect image; wherein, the size of the backup effect image is the same as the size of the target eye image, and the target area in the backup effect image has image information of the iris area of ​​the partial contact lens effect image, the target eye image is the eye area image of the face image, and the target area is the area at the same position as the iris area of ​​the eye area image;

[0309] The second determining submodule is used to determine the fusion coefficient corresponding to each pixel in the backup effect image;

[0310] The second fusion submodule is used to fuse the backup effect image and the eye area in the facial image using the determined fusion coefficient to obtain a colored contact lens trial image.

[0311] Optionally, the first generation submodule is specifically used for:

[0312] The iris region in the partial contact lens effect image is identified as the area to be utilized;

[0313] The image information of the target region in the specified image is replaced with the image information of the region to be used to obtain a backup effect image; wherein, the specified image is the same size as the target eye image.

[0314] Optionally, determining the iris region in the partial contact lens effect image as the area to be utilized is specifically used for:

[0315] Determine the target mask image of the iris region; wherein the target mask image is a binarized image with the iris region as the foreground;

[0316] Multiply the pixel values ​​of each pixel in the local contact lens effect image with the pixel values ​​of the corresponding pixels in the target mask image;

[0317] Based on the pixel values ​​of the pixels in the image obtained after multiplication, the iris region in the local contact lens effect image is determined as the region to be utilized.

[0318] Optionally, the first fusion module includes:

[0319] The first extraction submodule is used to extract the style features of the colored contact lens image;

[0320] The second extraction submodule is used to extract the texture features of the local iris image;

[0321] The first fusion submodule is used to fuse the style features and the texture features to obtain the fused features corresponding to the local iris image;

[0322] The second generation submodule is used to generate an image with fused features corresponding to the local iris image, as a local contact lens effect image.

[0323] Optionally, the step of performing a first fusion process on the partial iris image and the colored contact lens image to obtain a partial colored contact lens effect image is implemented by a pre-trained fusion network; wherein, the fusion network is a neural network trained based on the sample partial iris image and the sample colored contact lens image, used to output the fusion result; wherein, the fusion result is the result of incorporating the image style of the sample colored contact lens image into the sample partial iris image, while retaining the texture features of the sample partial iris image.

[0324] Optionally, the pre-trained fusion network is a generative network in a pre-trained generative adversarial network;

[0325] The training process of the generative adversarial network includes:

[0326] Obtain a partial image of the sample iris and an image of the sample contact lens;

[0327] The sample iris partial image and the sample colored contact lens image are input into the generation network to obtain a generated image as the fusion result;

[0328] The generated image and the sample iris image are respectively input into the adversarial network in the generative adversarial network to obtain a first discrimination result corresponding to the generated image and a second discrimination result corresponding to the sample iris image; wherein, the adversarial network is used to identify whether the input image is a real image;

[0329] Based on the difference between the first identification result and the corresponding true value, and the difference between the second identification result and the corresponding true value, the loss value of the adversarial network is calculated;

[0330] Based on the difference between the first identification result and the corresponding false value, the loss value of the generator network is calculated;

[0331] Based on the loss value of the adversarial network and the loss value of the generator network, it is determined whether the adversarial network and the generator network have reached a Nash equilibrium state; wherein, the Nash equilibrium state is used to characterize a stable state in which the loss values ​​of the adversarial network and the generator network fluctuate within a specified range.

[0332] If the adversarial network and the generative network do not reach Nash equilibrium, adjust the parameters of the generative adversarial network and return to the step of obtaining the sample iris local image and the sample contact lens image.

[0333] Optionally, before calculating the loss value of the generator network based on the difference between the first identification result and the corresponding false value, the method further includes:

[0334] The generated image and the sample iris image are respectively input into the feature extraction network, and a first loss value is calculated based on the difference between the output results of the feature extraction network for the generated image and the sample iris image.

[0335] The generated image and the sample contact lens image are respectively input into a style constraint network, and a second loss value is calculated based on the difference between the output results of the style constraint network for the generated image and the sample contact lens image; wherein, the style constraint network is a neural network used to identify image styles;

[0336] The step of calculating the loss value of the generator network based on the difference between the first identification result and the corresponding false value includes:

[0337] Based on the difference between the first identification result and the corresponding false value, as well as the first loss value and the second loss value, the loss value of the generator network is calculated.

[0338] Optionally, the generating network includes a style coding network, an encoder, and a decoder;

[0339] The step of inputting the sample iris local image and the sample contact lens image into the generation network to obtain the generated image as the fusion result includes:

[0340] The sample colored contact lens image is input into the style coding network to obtain the style features of the sample colored contact lens image;

[0341] The style features of the sample contact lens image and the sample iris partial image are input into the encoder for encoding to obtain the fused features corresponding to the sample iris partial image; the fused features corresponding to the sample iris partial image are feature images that fuse the style features of the sample contact lens image and the texture features of the sample iris partial image.

[0342] The fused features corresponding to the local iris image of the sample are input into the decoder for decoding to obtain the generated image as the fusion result.

[0343] This invention also provides an electronic device, such as... Figure 9 As shown, it includes a camera 901, a display 902, a processor 903, and a memory 904;

[0344] The camera 901 is used to capture facial images of the person to be trying on colored contact lenses;

[0345] Memory 904 is used to store computer programs;

[0346] When the processor 903 executes the program stored in the memory 904, it implements the steps of any of the image processing methods described in the above embodiments.

[0347] The display 902 is used to display the colored contact lens trial image obtained by the processor 903 after implementing the steps of any of the image processing methods described above.

[0348] Optionally, the electronic device is a computer, mobile phone, electronic mirror, or electronic photo album.

[0349] The aforementioned electronic devices

[0350] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0351] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0352] In addition, the aforementioned electronic photo album can be a photo album device that enables functions such as taking photos and managing albums.

[0353] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the image processing methods described in the above embodiments.

[0354] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of any of the image processing methods described in the above embodiments.

[0355] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0356] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0357] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0358] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An image processing method, characterized by, The method comprises: obtaining a beauty lens image and a face image to be tried on with the beauty lens; segmenting an iris local image containing an iris region from the face image; performing first fusion processing on the iris local image and the beauty lens image to obtain a local beauty lens effect image; wherein the first fusion processing is used for blending the image style of the beauty lens image into the iris local image and retaining the texture features of the iris local image; the iris local image is an image containing an iris region and having the same size as the beauty lens image; the step of performing the first fusion processing on the iris local image and the beauty lens image to obtain the local beauty lens effect image is implemented by using a pre-trained fusion network; the fusion network is a neural network trained based on sample iris local images and sample beauty lens images and used for outputting a fusion result; wherein the fusion result is a result of blending the image style of the sample beauty lens image into the sample iris local image and retaining the texture features of the sample iris local image; performing second fusion processing on the local beauty lens effect image and the face image to obtain a beauty lens try-on image; wherein the second fusion processing is used for blending the image style of the local beauty lens effect image into the eye region of the face image; the step of performing the second fusion processing on the local beauty lens effect image and the face image to obtain the beauty lens try-on image comprises: generating a backup effect image based on the local beauty lens effect image; wherein the backup effect image has the same size as a target eye image, and a target region in the backup effect image has image information of the iris region of the local beauty lens effect image; the target eye image is an eye region image of the face image, and the target region is a region having the same position as the iris region of the eye region image; determining a fusion coefficient corresponding to each pixel point in the backup effect image; wherein the fusion coefficient is determined according to the distance of each pixel point in the backup effect image from the target region; performing fusion on the eye region in the backup effect image and the face image by using the determined fusion coefficient to obtain the beauty lens try-on image.

2. The method of claim 1, wherein, The step of segmenting the iris local image containing the iris region from the face image comprises: recognizing eye key point information in the face image; determining the iris region in the face image by using the eye key point information; performing iris local region cropping on the face image to obtain the iris local image.

3. The method of claim 2, wherein, The step of determining the iris region in the face image by using the eye key point information comprises: fitting an eye contour in the face image by using the eye key point information to obtain an eye fitting curve; wherein the eye contour is a contour adhering to the eyelid. The face image is foreground mask processed based on the eye fitting curve and the contour parameter of the iris contour, to obtain a mask image of the face image; wherein the foreground mask processing is a mask processing taking the iris region of the face image as a foreground region; the contour parameter of the iris contour is generated based on a specified key point in the eye key point information, and the specified key point information is key point information located on the contour of an outer circle of the iris; An iris region in the face image is determined based on the mask image of the face image.

4. The method of claim 3, wherein, The face image is foreground mask processed based on the eye fitting curve and the contour parameter of the iris contour, to obtain a mask image of the face image, including: A specified pixel point in the face image is identified based on the eye fitting curve and the contour parameter of the iris contour; wherein the specified pixel point is located in a region formed by an eye contour and is located in a region formed by a contour of an outer circle of the iris; A pixel value of the specified pixel point in the face image is set to a first numerical value, and a pixel value of a pixel point other than the specified pixel point in the face image is set to a second numerical value, to obtain a mask image of the face image.

5. The method of claim 1, wherein, The backup effect picture is generated based on the local beauty lens effect picture, including: An iris region in the local beauty lens effect picture is determined as a region to be used; Image information of the target region in the specified image is replaced with image information of the region to be used, to obtain a backup effect picture; wherein the specified image is the same size as the target eye image.

6. The method of claim 5, wherein, The iris region in the local beauty lens effect picture is determined as the region to be used, including: A target mask image of the iris local image is determined; wherein the target mask image is a binary image taking the iris region as the foreground; Pixel values of each pixel point of the local beauty lens effect picture are multiplied by pixel values of corresponding pixel points in the target mask image; Based on pixel values of pixel points in the image obtained after multiplication, the iris region in the local beauty lens effect picture is determined as the region to be used.

7. The method of claim 1, wherein, The first fusion processing is performed on the iris local image and the beauty lens image, to obtain a local beauty lens effect picture, including: Style features of the beauty lens image are extracted; Texture features of the iris local image are extracted; The style features and the texture features are fused, to obtain fused features corresponding to the iris local image; An image with the fused features corresponding to the iris local image is generated as a local beauty lens effect picture.

8. The method of claim 1, wherein, The pre-trained fusion network is a generation network in a pre-trained generative adversarial network; The training process of the generative adversarial network includes: Sample iris local images and sample beauty lens images are obtained; The sample iris local images and the sample beauty lens images are input into the generation network, to obtain a generated image as a fusion result; inputting the generated image and the sample iris image into an adversarial network in the generative adversarial network respectively to obtain a first discrimination result corresponding to the generated image and a second discrimination result corresponding to the sample iris local image; wherein the adversarial network is used to discriminate whether an input image is a real image; calculating a loss value of the adversarial network based on a difference between the first discrimination result and a corresponding true value and a difference between the second discrimination result and a corresponding true value; calculating a loss value of the generative network based on a difference between the first discrimination result and a corresponding false value; judging whether the adversarial network and the generative network reach a Nash equilibrium state based on the loss value of the adversarial network and the loss value of the generative network; wherein the Nash equilibrium state is used to represent a stable state in which the loss values of the adversarial network and the generative network fluctuate within a specified range; if the adversarial network and the generative network do not reach the Nash equilibrium state, adjusting parameters of the generative adversarial network, and returning to the step of obtaining the sample iris local image and the sample beauty lens image.

9. The method of claim 8, wherein, Before the calculating the loss value of the generative network based on the difference between the first discrimination result and the corresponding false value, the method further includes: inputting the generated image and the sample iris local image into a feature extraction network respectively, and calculating a first loss value based on a difference between output results of the generated image and the sample iris image by the feature extraction network; inputting the generated image and the sample beauty lens image into a style constraint network respectively, and calculating a second loss value based on a difference between output results of the generated image and the sample beauty lens image by the style constraint network; wherein the style constraint network is a neural network used to discriminate image styles; The calculating the loss value of the generative network based on the difference between the first discrimination result and the corresponding false value includes: calculating the loss value of the generative network based on the difference between the first discrimination result and the corresponding false value, and the first loss value and the second loss value.

10. The method of claim 8, wherein, The generative network includes a style encoding network, an encoder and a decoder; The inputting the sample iris local image and the sample beauty lens image into the generative network to obtain a generated image as a fusion result includes: inputting the sample beauty lens image into the style encoding network to obtain a style feature of the sample beauty lens image; inputting the style feature of the sample beauty lens image and the sample iris local image into the encoder to encode and obtain a fusion feature corresponding to the sample iris local image; the fusion feature corresponding to the sample iris local image is a feature image in which the style feature of the sample beauty lens image and a texture feature of the sample iris local image are fused; inputting the fusion feature corresponding to the sample iris local image into the decoder to decode and obtain the generated image as the fusion result.

11. An image processing apparatus characterized by comprising: The device includes: an acquisition module configured to acquire a beauty lens image and a face image to be tried on with a beauty lens; a segmentation module configured to segment an iris local image containing an iris region from the face image; The first fusion module is configured to perform first fusion processing on the iris local image and the beauty lens image to obtain a local beauty lens effect image. The first fusion processing is configured to incorporate the image style of the beauty lens image into the iris local image and retain the texture features of the iris local image. The iris local image is an image containing an iris region and having the same size as the beauty lens image. The first fusion processing on the iris local image and the beauty lens image to obtain the local beauty lens effect image is implemented by using a pre-trained fusion network. The fusion network is a neural network trained based on sample iris local images and sample beauty lens images and configured to output a fusion result. The fusion result is a result of incorporating the image style of the sample beauty lens image into the sample iris local image and retaining the texture features of the sample iris local image. The second fusion module is configured to perform second fusion processing on the local beauty lens effect image and the face image to obtain a beauty lens try-on image. The second fusion processing is configured to incorporate the image style of the local beauty lens effect image into the eye region of the face image. The second fusion processing on the local beauty lens effect image and the face image to obtain the beauty lens try-on image includes: generating a backup effect image based on the local beauty lens effect image. The backup effect image has the same size as a target eye image, and a target region in the backup effect image has image information of the iris region of the local beauty lens effect image. The target eye image is an eye region image of the face image, and the target region is a region having the same position as the iris region of the eye region image. determining a fusion coefficient corresponding to each pixel point in the backup effect image. The fusion coefficient is determined according to the distance of each pixel point in the backup effect image from the target region. performing fusion on the eye region in the backup effect image and the face image by using the determined fusion coefficient to obtain the beauty lens try-on image.

12. An electronic device, comprising: The electronic device includes a camera, a display, a processor, and a memory. The camera is configured to capture a face image to be used for beauty lens try-on. The memory is configured to store a computer program. The processor is configured to execute the program stored in the memory to implement the steps of the image processing method according to any one of claims 1-10. The display is configured to display the beauty lens try-on image obtained after the processor implements the image processing method according to any one of claims 1-10.

13. The electronic device of claim 12, wherein, The electronic device is a computer, a mobile phone, an electronic mirror, or an electronic photo album.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN107818305A

  • Face image processing method and device, equipment and computer readable storage medium

    CN113642364A

  • KR20220015200A