A virtual makeup method, system, device and storage medium

By decomposing the user's facial structure and target makeup information using GAN neural networks, this technology solves the problems of poor makeup overlay effects and training material in existing virtual makeup technology, achieving efficient and natural makeup synthesis to meet the rapidly changing needs of the Internet.

CN115936796BActive Publication Date: 2026-08-25BEIJING MOMO INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111169098.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-02
Publication Date
2026-08-25
Estimated Expiration
2041-10-02

AI Technical Summary

Technical Problem

Existing virtual makeup-changing technology struggles to achieve arbitrary makeup changes and cannot effectively remove the user's original makeup, especially when the user is already wearing heavy makeup. Furthermore, the difficulty in collecting training materials leads to unsatisfactory model training results.

Method used

Using a GAN neural network, the user's facial structure information and target makeup information are decomposed and extracted, and used as independent input conditions to reconstruct the makeup. During the training process, a single image with makeup is used for self-decomposition and perturbation processing to generate a synthetic image.

Benefits of technology

It achieves high-quality virtual makeup effects with natural and realistic makeup, reduces reliance on training materials, improves model training efficiency and the accuracy of synthesized images, and adapts to the rapidly changing needs of the Internet.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115936796B_ABST
    Figure CN115936796B_ABST
Patent Text Reader

Abstract

The application discloses a virtual makeup changing method, comprising the following steps: 1) a user uploads a picture to be changed; 2) the user selects or uploads a target makeup picture; 3) model features of the picture to be changed and the target makeup picture are disassembled and extracted; 4) the model features comprise facial structure information of the picture to be changed and facial makeup information of the target makeup picture; 5) the processed makeup picture and the extracted features of the facial structure information of the user picture are input into a GAN neural network as conditions; and 6) the GAN neural network outputs a synthesized makeup changing picture according to the conditions. The facial structure information and the makeup information are used as input conditions of the GAN neural network, and a more realistic makeup changing effect is obtained. Meanwhile, an improved neural network training method is used, a single picture is used to disassemble two required conditions, and a training process is carried out, so that it is possible to obtain large-scale one-to-one paired data required for previous neural network training, and compared with a traditional texture material method, the reality and fusion degree of image generation are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of virtual makeup change, specifically relating to a method, system, device, and storage medium for virtual makeup change using neural networks, and in particular, a method for generating virtual makeup change images using a GAN neural network model trained by decomposing and processing a single image to obtain model features. Background Technology

[0002] With the development of internet technology, new models such as online shopping, live streaming, online interactive entertainment, and online dating are becoming increasingly popular. Compared to shopping in physical stores or experiencing offline shops, online shopping has advantages such as a wider selection, more diverse products, and time and effort savings. Online interactive entertainment and dating offer advantages such as convenience, speed, and a large audience. However, buying goods online also presents some unresolved problems, the most significant being the inability to directly view the actual effect of the product. This problem is most prominent among all product categories, especially for beauty products such as eyeshadow, eyeliner, blush, lipstick, and foundation. Unlike physical store shopping where customers can change makeup and view the product's effect in real time, online beauty shopping cannot provide images tailored to the consumer's face shape and skin. It only offers pictures of models trying the product on, and sometimes there are no trial images at all. Consumers cannot directly and intuitively understand how well the cosmetics match their own face shape and skin, causing considerable confusion.

[0003] Furthermore, in some interactive entertainment settings, users also have a need to experience the effects of different makeup looks. For ordinary users, seeing a makeup look they like online often leads them to want it realistically recreated on their own faces to see if it suits them. Existing technology has already developed some makeup-changing techniques to address this need. Virtual makeup changing typically involves transferring makeup from a target image onto the original image. Given an image to be changed and a target makeup image, the algorithm needs to transfer the makeup from the target image to the image to be changed. This method allows users to see the effect of makeup on themselves without spending time and money to actually buy cosmetics or apply makeup, and it can be widely used in the internet industry.

[0004] Existing technology 1 discloses a robust makeup transfer method based on reference images, including the following specific steps: Step 1: Define the makeup transfer problem as: where x is the image to be transferred, y is the reference image, and the mapping G takes x and y as input and outputs the image after makeup transfer, which is the same person as x and has the makeup of the reference image y; Step 2: Extract the makeup matrix of the reference image y, and use a makeup extraction network to extract the first makeup matrix γ and the second makeup matrix β from the reference image; Step 3: Calculate the similarity between each pixel in the image to be transferred and each pixel in the reference image; Step 4: Perform deformation processing on the makeup matrix extracted in Step 2 using the pixel similarity to obtain an adaptive makeup matrix; Step 5: Transfer the makeup of the visual feature map of the image to be transferred using the adaptive makeup matrix; Step 6: Upsample the visual feature map to obtain the makeup-transferred image. However, although this makeup-changing method can complete the corresponding makeup operation with simple processing, it is mainly completed by rendering and pasting materials after locating key points. The main operation is effect blending, which is heavily dependent on key point positioning and material operation. It cannot achieve arbitrary makeup changing and has no effect on removing existing makeup in the image. In cases where some users are already wearing heavy makeup, the superimposed effect cannot meet the high requirements of users.

[0005] In the current field of virtual makeup replacement, using the target makeup to cover the user's original makeup is the mainstream approach due to its low computational cost and simplicity. However, with the rapid development of neural network model technology in recent years, methods using neural network models for virtual makeup replacement have gradually emerged in this field. Therefore, the quality of the final output image of the neural network model—that is, the degree of similarity between the original image and the target image—has become the most critical factor. Especially for deep learning neural networks, how to train them and achieve optimal training results and efficiency has become a pressing issue that needs to be addressed. Summary of the Invention

[0006] To align with the development trends of the internet industry and meet the demands of virtual makeup application—allowing users to freely try on their favorite looks and experience personalized effects—we need a virtual makeup application method that requires simple input, minimizes computational load on terminal devices, and produces results close to real-life makeup. To address these issues, this invention provides a makeup application method that overcomes the aforementioned problems, and also includes a neural network training method.

[0007] This invention provides a virtual makeup-changing method, the method comprising: 1) a user uploading an image to be changed to; 2) a user selecting or uploading a target makeup image; 3) disassembling and extracting model features from the image to be changed and the target makeup image; 4) the model features include facial structure information of the image to be changed and facial makeup information of the target makeup image; 5) using the features extracted from the processed makeup image and the user's facial structure information as conditions input into a GAN neural network; 6) the GAN neural network outputting a synthesized makeup-changing image based on the conditions.

[0008] Furthermore, in step 2), users can select a target makeup image provided by the system or upload a partial image of the target makeup.

[0009] Furthermore, in step 3), the terminal uploads the image information to the server to complete the decomposition and extraction of model features, or directly completes the decomposition and extraction of model features on the terminal and then uploads the feature information to the server.

[0010] Furthermore, the steps for decomposing and extracting the features of the image model to be changed include: 1) obtaining a human head contour image based on the user's two-dimensional image processing; 2) inputting the two-dimensional human head contour image into the first neural network after deep learning for key point regression; 3) obtaining the key point information of the user's head; 4) obtaining semantic segmentation maps of each part of the head; 5) perturbing the image with color, and removing facial makeup information by using lab, brightness, contrast, and light and shadow perturbation methods; 6) extracting features from the perturbated image through an encoder to obtain representative user image facial structure information.

[0011] Furthermore, the steps for deconstructing and extracting the model features of the target makeup image also include: 1) converting the 3D facial image into a 2D image; 2) performing WARP stretching and deformation processing on the image according to certain preset rules; 3) fixing and unfolding the image into a picture of a specified part in a fixed position according to preset key point information; 4) this image WARP to a fixed position carries the facial makeup information of the target makeup image and can be directly input into the GAN neural network.

[0012] Furthermore, this step also includes: 1) setting 5 preset key points, with the specified locations including the corners of the eyes, the tip of the nose, and the corners of the mouth; 2) in addition to the preset key points, perturbing other points by adding color, adding or subtracting RGB values, changing contrast, brightness, and lighting; 3) retaining makeup information including the 5 key points and the lipstick, eyeshadow, and color information within a certain range around them.

[0013] Furthermore, before the model features are used as conditional inputs to the GAN neural network model, there is also a process of training the GAN neural network. The training process is completed using a single image with makeup through a self-disassembly method.

[0014] Furthermore, the training process includes: 1) removing perturbations from the facial makeup information in the image and extracting facial structure information from the image; 2) removing perturbations from the facial structure information in the image and extracting facial makeup information from the image; 3) inputting the extracted facial structure information into the encoder to extract features; 4) inputting the obtained features and facial makeup information as input conditions into the GAN neural network model; 5) using the image itself as the ground truth for training the model.

[0015] Furthermore, when removing facial makeup information from the original image, in addition to random perturbation, the following perturbation operation is performed: the colors of specified parts extracted from other facial images are overlaid onto specified parts of the original image using histogram matching, so that the makeup information at specified locations in the original image is completely stripped away.

[0016] Furthermore, this invention also provides a system for implementing a virtual makeup change method, comprising: 1) an image acquisition module for acquiring user-uploaded images of the makeup to be changed, wherein the user selects or uploads a target makeup image; 2) a decomposition and extraction module for acquiring model features of the image of the makeup to be changed and the target makeup image; wherein the model features include facial structure information of the image of the makeup to be changed and facial makeup information of the target makeup image; 3) an image synthesis output module, comprising a GAN neural network for inputting the processed makeup image and features extracted from the user's facial structure information as conditions into the GAN neural network and outputting the synthesized makeup change image.

[0017] A computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the steps described above.

[0018] An electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor, when executing the program stored in the memory, implements any of the steps described above.

[0019] The beneficial effects of this invention are:

[0020] 1. High Realism. This invention utilizes a GAN network to directly generate synthetic images. Unlike existing GAN networks, we creatively propose two completely independent and non-interfering model input conditions: user facial structure information and target facial makeup information. In this mode, the input conditions of the conditional generative adversarial network are simplified to two unrelated conditions: one is the user's own facial structure information, and the other is the target facial makeup information to be changed. This is equivalent to reconstructing makeup information based on facial information that is even more thorough than bare face. Compared with the masking virtual makeup method, although the processing steps are increased, it avoids the need to rely solely on rendering and pasting materials after key point localization to complete the fusion operation. It overcomes the defect that the superimposed effect cannot meet the user's requirements when the original image itself is heavily made up. At the same time, our solution has lower requirements for user images, better makeup effect, and allows users to choose any target makeup image without relying on pre-made operational materials, achieving a personalized look for each user.

[0021] 2. Simple acquisition of training materials and good training results. Following typical model design principles, training a GAN for virtual makeup tasks should involve providing the user's bare face photo, the target makeup photo, and a photo of the user with the makeup applied (referencing the target makeup). The photo with the makeup applied is used as the model's ground truth for training. However, in practice, it's difficult to collect complete sets of photos, resulting in fewer makeup types available for training, which struggles to meet the demands of internet entertainment and interaction. Therefore, we creatively utilize specific information perturbation methods to decompose and extract facial structure and makeup information from the photos. These two pieces of information replace the user's bare face photo and the target makeup photo as input conditions, and the made-up photo itself is used as the model's output ground truth for training. In this case, any made-up photo collected from various channels can be used as training material, significantly reducing the requirements for training materials. While ensuring training with massive amounts of data, the GAN neural network can produce seamless synthetic images. The virtual makeup replacement method provided by this invention adopts an improved neural network training method, which can complete relatively accurate training using only a single image. This greatly improves the parameter accuracy of the deep learning neural network and significantly accelerates the convergence speed, thereby greatly improving the consistency and fidelity of the synthesized image generation.

[0022] 3. Utilizing Deep Learning Neural Networks. This invention fully leverages the advantages of deep learning networks, enabling high-precision reconstruction of facial structure and makeup information in various complex scenarios. Different neural networks are used for different purposes, employing neural network models with varying input conditions and training methods to achieve accurate contour separation, semantic segmentation of different parts of the human body, and determination of key points and joints against complex backgrounds. This eliminates the influence of head posture, hairstyle, etc., achieving the closest possible approximation to the real appearance of the head or face. While existing technologies also use neural network models, differences in input conditions, input parameters, and training methods lead to significant variations in the functions and effects of these models.

[0023] 4. Simple User Operation. This invention provides a method for generating composite images of original facial information and target makeup information using a GAN neural network. Virtual makeup change can be quickly achieved with just two photos, perfectly adapting to the characteristics and trends of the internet age—it's simple and fast. Users don't need any preparation; simply uploading a photo and selecting a makeup look completes the process. Applying this invention to entertainment mini-programs or online shopping scenarios will greatly enhance user experience and engagement. It eliminates the need for measuring the actual shape of the human body and face, providing broad application scenarios for various industries such as interactive entertainment and online shopping. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a processing embodiment;

[0026] Figure 2 This is a flowchart of a processing embodiment;

[0027] Figure 3 This is a schematic diagram of the structural information of condition A in a neural network model of one embodiment;

[0028] Figure 4 A schematic diagram of makeup information under condition B of a neural network model in one embodiment;

[0029] Figure 5 This is a schematic diagram of a neural network model training method according to one embodiment;

[0030] Figure 6 This is a schematic diagram of the system of the present invention. Detailed Implementation

[0031] The features and exemplary embodiments of various aspects of the present invention will now be described in detail. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be practiced without some of these specific details.

[0032] It should also be noted that, for ease of description, only the parts relevant to this application are shown in the accompanying drawings, not all of them. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. The process may correspond to a method, function, procedure, subroutine, subprogram, etc.

[0033] like Figure 1-2 As shown, this invention provides a virtual makeup-changing method, the method comprising: 1) a user uploading an image to be changed to; 2) a user selecting or uploading a target makeup image; 3) disassembling and extracting model features from the image to be changed and the target makeup image; 4) the model features include facial structure information of the image to be changed and facial makeup information of the target makeup image; 5) using the features extracted from the processed makeup image and the user's facial structure information as conditions input into a GAN neural network; 6) the GAN neural network outputting a synthesized makeup-changing image according to the conditions.

[0034] Here, after uploading an image, the user doesn't actually see the subsequent steps until the synthesized makeup-altered image is sent back to the user's device. After obtaining the image to be altered and the target makeup image, these two images need to be processed separately using sub-neural networks. The user image retains only facial structure information as input condition A, and the makeup image retains only makeup information as input condition B. Conditions A and B are then input into the GAN network. (See attached diagram) Figure 3 and 4Conditions A and B can be understood as images lacking makeup information (especially color information) and structural information, respectively. Essentially, in the image of the user to be re-made up, the outline of the face can still be seen, meaning the person can be basically identified, but the original color information is completely lost. Similarly, in the target makeup image, the makeup information of the face can still be seen, but the user's facial contour features are completely indistinguishable, meaning the person in the image cannot be identified. The trained main GAN network, after receiving these two input conditions, autonomously superimposes and merges the two independent conditions A (structural information) and B (makeup information) to generate a composite image with the changed makeup. Because new makeup information is superimposed on a facial structure image that has no makeup information, there is no interference from the original makeup, significantly improving the makeup effect and making the changed makeup more natural and realistic.

[0035] Furthermore, in step 2), users can select a target makeup image provided by the system or upload a partial image of the target makeup. This demonstrates the user's selection process. If the user is satisfied with the makeup photos provided by the system, they can directly select the system-provided image as the target makeup image, saving the user the trouble of uploading their own makeup image. Moreover, since the template photos provided by the system are all post-processed images, they perform well in terms of integration and realism.

[0036] Furthermore, in step 3), the terminal uploads the image information to the server to complete the disassembly and extraction of model features, or directly completes the disassembly and extraction of model features on the terminal and then uploads the feature information to the server. With privacy protection receiving increasing attention from users and regulators, how to protect collected customer information is a problem that must be properly addressed. In addition to implementing strict data compliance measures on the server side, the computing power and storage space of the user terminal can be utilized to complete some model calculations and raw data storage functions. In this operating mode, the server will no longer need to disassemble and process the images uploaded by users, and since the received image information is processed according to the prescribed format and size, some network resources and server bandwidth resources are saved; at the same time, the user's raw data is directly stored locally, and only the image information that has been perturbed is uploaded, reducing the risk of privacy leakage.

[0037] Furthermore, the steps of disassembling and extracting the features of the image model to be changed also include: 1) obtaining a human head contour image based on the user's two-dimensional image processing; 2) inputting the two-dimensional human head contour image into a first neural network after deep learning for key point regression; 3) obtaining key point information of the user's head; 4) obtaining semantic segmentation maps of each part of the head; 5) perturbing the image with color and removing facial makeup information by using LAB, brightness, contrast, and light and shadow perturbation methods; 6) extracting features from the perturbated image through an encoder to obtain representative user image facial structure information.

[0038] This section primarily involves processing the acquired original user images to obtain the structural parameter information needed to generate the makeover images, such as... Figure 2 As shown. Previously, the selection of these key points was usually done manually, but this method is inefficient and unsuitable for the fast-paced demands of the internet age. Therefore, with the widespread use of neural networks today, using deep learning-based neural networks to replace manual key point selection has become a trend. However, how to efficiently utilize neural networks is a question that requires further research.

[0039] The acquisition of a two-dimensional facial contour image utilizes an object detection algorithm, which is a target region fast generation network based on a convolutional neural network. Before inputting the two-dimensional facial image into the first neural network model, a process of training the neural network is included. The training samples include standard images with annotated original keypoint positions, which are manually annotated with high accuracy on the two-dimensional image. Here, the target image is first acquired, and the object detection algorithm is used to detect human structural information in the target image. Human detection does not involve using measuring instruments to detect a real human body; in this invention, it actually refers to any given image, typically a two-dimensional photograph containing sufficient information, such as an image including facial features and head information. Then, a certain strategy is used to search the given image to determine whether it contains a face. If the given image contains a face, structural parameters such as the location and size of facial organs are provided. In this embodiment, before obtaining the key points of the facial structure in the target image, it is necessary to perform face detection on the target image to obtain the bounding boxes marking the positions of the faces in the target image. Since the input image can be any image, it is inevitable that there will be some non-human backgrounds, such as tables, chairs, trees, cars, buildings, etc. These useless backgrounds need to be removed using some mature algorithms.

[0040] Simultaneously, we also need to perform keypoint detection and edge detection, using a neural network to generate a keypoint map of the face. Optionally, the object detection algorithm can quickly generate a network for the target region based on a convolutional neural network. This first neural network requires extensive data training. Keypoints are manually annotated on collected photos and then input into the neural network for training. After deep learning, the neural network can basically obtain keypoints with the same accuracy and effect as manually annotated keypoints immediately after inputting photos, while being tens or even hundreds of times more efficient. In this invention, obtaining the location of facial keypoints in a photo is only the first step, obtaining 1D point information. 2D surface information must then be generated from this 1D point information, such as obtaining the position of eyebrows, the corners of the eyes, the contours of the eyes, and the contours of the lips through keypoints. These tasks can be accomplished using neural network models and mature algorithms in existing technologies.

[0041] Furthermore, the steps of disassembling and extracting the target makeup image model features also include: 1) converting the three-dimensional facial image into a two-dimensional image; 2) performing WARP stretching and deformation processing on the image according to certain preset rules; 3) fixing and unfolding the image into a picture of a specified part at a fixed position according to preset key point information; 4) this image WARP to a fixed position carries the facial makeup information of the target makeup image, which can be directly input into the GAN neural network.

[0042] This step primarily involves removing structural information while completely preserving makeup details. This includes using traditional WARP image affine transformations to stretch and deform the originally three-dimensional facial features. See the attached image for the final result. Figure 3 The original shape of the face is no longer discernible. During the stretching and deformation process, based on the key point information obtained in the previous step, several key points are preset as fixed points. This allows for accurate matching of the same parts in different original images. For example, in different original images, the position of the corner of the eye will be fixed at a certain coordinate point after processing. In this way, the makeup information of the corner of the eye can be accurately transferred to the corner of the eye in the target image.

[0043] Furthermore, this step also includes: 1) setting the preset key points to 5 points, with the specified locations including the corners of the eyes, the tip of the nose, and the corners of the mouth; 2) in addition to the preset key points, perturbing other points by adding color, adding or subtracting RGB values, changing contrast, brightness, and light and shadow; 3) retaining makeup information including the 5 key points and the lipstick, eyeshadow, and color information within a certain range around them.

[0044] This step addresses special cases that may be encountered in the output of the GAN model. While the synthesis output task in this invention can be accomplished using a common GAN (Generative Adversarial Network) model, after processing a certain number of samples, we found that mutual interference of makeup information (color, brightness, etc.) between different parts or regions still occurs, leading to uneven or inconsistent colors after synthesis. Our initial assessment is that the perturbation of the structural (ID) information subtly affects the makeup information between different regions, resulting in unnatural transitions. Therefore, we attempted to fix the makeup information at several designated key points and their surrounding areas, setting area or pixel thresholds for these areas, and further perturbing the makeup information in other non-key areas. The entire perturbation process and method are similar to those in decomposition condition A; of course, some additional methods will be added as needed. After this processing, the color accuracy of the makeup information during the transfer process is higher, and the transition is more natural.

[0045] Furthermore, before the model features are used as conditional inputs to the GAN neural network model, a process of training the GAN neural network is also included, as described above. Figure 5 The training process uses a single image with makeup applied and employs a self-deconstruction method. The training process includes: 1) removing perturbations from the facial makeup information in the image and extracting facial structure information; 2) removing perturbations from the facial structure information in the image and extracting facial makeup information; 3) inputting the extracted facial structure information into an encoder to extract features; 4) inputting the obtained features and facial makeup information as input conditions into a GAN neural network model; and 5) using the image itself as the ground truth for training the model.

[0046] The core of the training method in this invention lies in the fact that the entire training process can be completed using only a single image with makeup. As is well known, there are numerous algorithms and parameters for neural networks, and various loss functions (LOSS) vary depending on project requirements. Only through continuous training and iteration can the most accurate "answer" be obtained. According to the most standard design approach, training a GAN for virtual makeup tasks should provide a user's bare face photo, a photo of the target makeup, and a photo of the user with the target makeup as a reference. Training the model with real photos of the user after makeup application yields better results, and the more samples, the more accurate the model output. However, in actual testing, we found that this training strategy places too high demands on the training materials. Finding photos of bare faces, makeup, and after makeup is very difficult for most companies, resulting in a limited number of training materials and a small number of interchangeable makeup looks. Various model parameters and LOSS functions are difficult to adjust, leading to lower realism and less seamless integration of the results, especially with some complex transformations exhibiting noticeable inconsistencies.

[0047] To overcome this deficiency, we use any acquired photo of a user with makeup applied as training material. We extract and deconstruct the facial structure and makeup information from the photo, replacing the user's bare-faced photo and the target makeup photo as input. This makeup-applied photo is then used as the model's ground truth output for training. Because we train the neural network with the most accurate answer, and the training material becomes readily available instead of difficult to obtain, we significantly reduce the model's dependence on the training data. With training on massive amounts of data, we can greatly improve the accuracy of parameters and the reasonableness of loss. The GAN neural network model converges very quickly, and after training, using the model to synthesize images directly yields synthesized images with excellent realism and integration. The structural information of the face in the image closely matches the input image. Figure 1 Therefore, the overlay of makeup information is more natural and realistic due to the absence of interference from the original makeup. This invention employs an improved self-disassembly method for training the neural network, enabling relatively accurate training using only a single image. This significantly reduces the difficulty of collecting training sets, allowing us to obtain a large amount of data at a lower cost. It also improves the network's generalization ability, resulting in a significantly faster convergence speed for the GAN neural network, which greatly enhances the accuracy and realism of synthesized image generation.

[0048] The method we use in training the GAN model is consistent with the actual computation process of the GAN model. Compared to the actual makeup transformation process which uses two images as the data sources for conditions A and B, the training method of this invention uses only one original image as the training material to decompose conditions A and B, simply for convenience and accuracy. It cleverly utilizes the fact that the source image already contains information about conditions A and B, and decomposes and extracts it accordingly. Simultaneously, the decomposition method used in the model training process and the actual processing process is consistent: makeup information is removed when decomposing and extracting condition A; structural information is removed when decomposing and extracting condition B. The specific processing techniques used are also consistent, including various perturbation methods, without increasing the difficulty of the system performing these operations. The same set of operations can be used to train the model and also to complete the output of synthesized images in the main process. This effectively ensures the relevance and continuity of the training, and the model iterations will increasingly approach user needs.

[0049] Furthermore, when removing facial makeup information from the original image, in addition to random perturbation, the following perturbation operation is performed: the colors of specified areas extracted from other facial images are overlaid onto specified areas of the original image using histogram matching, so that the makeup information at specified locations in the original image is completely stripped away. During model training, various conditions will have varying degrees of impact on the final training results. We found that in the decomposition process of condition A, if only random perturbation of various parameters is performed, the makeup information stripping effect is often incomplete, failing to completely remove elements such as lip color and eyeshadow color. This may further lead to unnatural-looking makeup after synthesis and overlay. We aim to completely hide lip color and other elements in the original image, especially when the makeup color is dark and the area is large. To address this, we designed a simple yet effective method: completely replace the makeup of a specified area in the original image with the makeup of the same specified area in our pre-made image. Because the makeup image copied from other images differs greatly from the original image's makeup, it is easy to achieve a complete removal effect during perturbation. This greatly enhances the perturbation effect on the specified area or range, making the final image more natural and less affected by the original image.

[0050] We re-input the original makeup image as the ground truth into the neural network for training, so that the neural network can know what the most accurate match between the input conditions and the output model is. There are many methods for training neural networks, but the basic process and ideas are similar; simply put, it's about teaching the neural network what is right and what is wrong. One of the commonly used deep learning neural network training methods usually includes: (1) data preprocessing; (2) inputting data into the neural network (each neuron first inputs values, weights are accumulated, and then inputs the activation function as the output value of the neuron) and propagating forward to obtain a score or result; (3) inputting the "score" or "result" into the error function (regularization penalty to prevent overfitting), comparing it with the expected value to obtain the error, and summing multiple errors, and judging the degree of recognition through the error (the smaller the loss value, the better); (4) determining the gradient vector through backpropagation (backward differentiation, the error function and each activation function in the neural network are required, the ultimate goal is to minimize the error); (5) finally adjusting each weight through the gradient vector to adjust the trend of the error tending to 0 or convergence towards the "score" or "result"; (6) repeating the above process until the set number of times or the average value of the loss no longer decreases (the lowest point); (7) training is completed.

[0051] As we can see, most neural network training processes are inherently based on large amounts of data. Initially, the neural network model performs poorly in many scenarios, so a large number of bad cases need to be labeled. These labeled bad cases are then added to the training set, allowing the neural network to learn what the true values ​​of these bad cases should be. Once the network learns this, it can accurately predict similar scenes. However, this training method is not very efficient. Therefore, the training process is essentially iterative. If the model can be trained to identify what is correct and good using results that are almost identical to the standard answer, fewer bad cases will appear later, the convergence speed will increase rapidly, and the model performance will improve.

[0052] In addition, combined Figures 1 to 5 The resume screening method described in this embodiment of the invention can be implemented by a corresponding electronic device. Figure 6 This is a schematic diagram illustrating a hardware structure 300 according to an embodiment of the present invention.

[0053] This invention also discloses a system for implementing a virtual makeup change method, comprising: 1) an image acquisition module for acquiring user-uploaded images of the makeup to be changed, wherein the user selects or uploads a target makeup image; 2) a decomposition and extraction module for acquiring model features of the image of the makeup to be changed and the target makeup image; wherein the model features include facial structure information of the image of the makeup to be changed and facial makeup information of the target makeup image; 3) an image synthesis output module, comprising a GAN neural network for inputting the processed makeup image and features extracted from the user's facial structure information as conditions into the GAN neural network and outputting the synthesized makeup change image.

[0054] And, an apparatus characterized in that it comprises: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to perform the resume screening method described in any of the preceding claims.

[0055] And a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, performs the resume screening method as described in any of the preceding items.

[0056] The device 300 implementing the present invention in this embodiment includes: a processor 301, a memory 302, a communication interface 303, and a bus 310, wherein the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication with each other.

[0057] Specifically, the processor 301 may include a central processing unit (CPU), an ASIC, or one or more integrated circuits that can be configured to implement embodiments of the present invention.

[0058] In other words, device 300 can be implemented as including: processor 301, memory 302, communication interface 303, and bus 310. Processor 301, memory 302, and communication interface 303 are connected via bus 310 and communicate with each other. Memory 302 is used to store program code; processor 301 reads the executable program code stored in memory 302 to run a program corresponding to the executable program code, so as to execute the method in any embodiment of the present invention, thereby realizing the method and apparatus described in conjunction with the accompanying drawings.

[0059] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0060] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A virtual makeup change method, characterized in that, The method includes: 1) Users upload photos of themselves to be dressed up; 2) The user selects or uploads a picture of the target makeup look; 3) Deconstruct and extract the model features of the image to be changed and the target makeup image; The process of disassembling and extracting the model features of the image to be re-made includes: Keypoint detection and semantic segmentation of facial images; and Facial makeup information is removed by perturbation of color, brightness, contrast and light and shadow. The perturbation-processed image is then processed by an encoder to extract features and obtain user image facial structure information. The facial structure information is then input into the encoder to extract structural features. Deconstructing and extracting features from the target makeup image model includes: Convert a 3D facial image into a 2D image; The image is subjected to WARP stretching and deformation processing according to preset rules; the image is fixed and unfolded into an image with the specified parts located in a fixed position according to preset key point information; the lipstick information, eyeshadow information and color number information of the preset key points and their surrounding areas are retained; 4) Input the features extracted from the facial structure information of the processed makeup image and user image into the GAN neural network as input conditions; 5) The GAN neural network is trained using the face image with makeup as the ground truth. After training, the facial structure features corresponding to the image to be remade and the facial makeup information corresponding to the target makeup image are input into the trained GAN neural network, and the synthesized image after makeup is output.

2. The method according to claim 1, characterized in that, In step 2), users can select a target makeup image provided by the system or upload a partial image of the target makeup.

3. The method according to claim 1, characterized in that, In step 3), the terminal uploads the image information to the server to complete the decomposition and extraction of model features, or the terminal directly completes the decomposition and extraction of model features and then uploads the feature information to the server.

4. The method according to claim 1, characterized in that, It also includes: 1) the preset key points are set to 5 points, and the specified parts include the corners of the eyes, the tip of the nose and the corners of the mouth; 2) in addition to the preset key points, other points are disturbed by adding color, adding or subtracting RGB values, changing contrast, brightness and light and shadow.

5. The method according to claim 1, characterized in that, Before the model features are used as conditional inputs to the GAN neural network model, the process of training the GAN neural network is also included. The training process is completed using a single image with makeup through a self-disassembly method.

6. The method according to claim 5, characterized in that, The training process includes: 1) removing perturbations from the facial makeup information in the image and extracting facial structure information from the image; 2) removing perturbations from the facial structure information in the image and extracting facial makeup information from the image; 3) inputting the extracted facial structure information into the encoder to extract features; 4) inputting the obtained features and facial makeup information as input conditions into the GAN neural network model; 5) using the image itself as the ground truth of the model for training.

7. The method according to claim 6, characterized in that, When removing facial makeup information from the original image, in addition to random perturbation, the following perturbation operation is performed: the colors of specified parts extracted from other facial images are overlaid on specified parts of the original image using histogram matching, so that the makeup information at specified locations in the original image is completely stripped away.

8. A virtual makeup change system for implementing virtual makeup change methods, characterized in that, include: The image acquisition module is used to acquire images of the makeup to be changed uploaded by the user. The user selects or uploads an image of the target makeup look. The disassembly and extraction module is used to obtain model features of the image to be changed and the target makeup image; the model features include facial structure information of the image to be changed and facial makeup information of the target makeup image; wherein, disassembling and extracting the model features of the image to be changed includes: Keypoint detection and semantic segmentation of facial images; and Facial makeup information is removed by perturbation of color, brightness, contrast and light and shadow. The perturbation-processed image is then processed by an encoder to extract features and obtain user image facial structure information. The facial structure information is then input into the encoder to extract structural features. Deconstructing and extracting features from the target makeup image model includes: Convert a 3D facial image into a 2D image; The image is subjected to WARP stretching and deformation processing according to preset rules; the image is fixed and unfolded into an image with the specified parts located in a fixed position according to preset key point information; the lipstick information, eyeshadow information and color number information of the preset key points and their surrounding areas are retained; The image synthesis output module includes a GAN neural network, which is used to train the GAN neural network by using the face image with makeup as the ground truth. After training, the facial structure features corresponding to the image to be remade and the facial makeup information corresponding to the target makeup image are input into the trained GAN neural network, and the synthesized image after makeup is output.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-7.

10. An electronic device, characterized in that, The system includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor, when executing the program stored in the memory, implements the steps of the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Shadow robust makeup migration system and method based on decoupling representation

    CN113362422A