Method for generating training data, makeup try-on method, electronic device, and storage medium
By acquiring and fusing face images and obstructing images, generating masked makeup area data and training segmentation models, the problem of difficulty in obtaining mask data in makeup parts is solved, and a more natural virtual makeup trial effect is achieved.
Patent Information
- Application Number
- CN202111674777.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In the prior art, it is difficult to obtain relevant data that the makeup area is blocked by objects, and the cost is high, which affects the authenticity of the virtual makeup trial.
By acquiring the first face image and the occlusion image, a second face image is generated in which the upper makeup part is covered by the occlusion object, and a third segmented mask image of the upper makeup area is generated using the segmented mask image for training the segmented model.
A large amount of blocked makeup area segmentation data was obtained, which reduced the cost of obtaining training data and improved the effect and authenticity of virtual makeup trials.
Smart Images

Figure CN114387285B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly relates to a method for generating training data, a makeup try-on method, an electronic device, and a storage medium. Background Art
[0002] In recent years, the development of Internet technology has brought many conveniences to people's lives. For example, it is possible to purchase beauty products such as lipsticks and foundations online. Since beauty products are relatively special and users need to actually try them before deciding whether to purchase, it is very necessary to perform virtual makeup try-on based on the beauty products selected by users online.
[0003] In the prior art, a virtual makeup try-on algorithm based on key points is used for makeup try-on. When the makeup application area of a user is blocked by other objects, the blocked position will also be virtually made up. Taking virtual lipstick application as an example, when a user's lips are blocked by a finger, the try-on algorithm will still apply lipstick to the blocked position of the finger, affecting the authenticity of virtual makeup application.
[0004] In order to improve the authenticity of virtual makeup application, relevant data where the makeup application area is blocked by an object can be used for model training, and the trained model is used to predict the makeup application area of a user. However, currently, the acquisition of the above data is difficult and costly. Summary of the Invention
[0005] Embodiments of the present application provide a method for generating training data, a makeup try-on method, an electronic device, and a storage medium to solve the technical problem in the prior art that it is difficult and costly to obtain relevant data where the makeup application area is blocked by an object.
[0006] According to a first aspect of the present application, a method for generating training data is disclosed, and the method includes:
[0007] Obtain a first face image and a first occluder image, where the makeup application area in the first face image is not blocked, and the first occluder image is an image corresponding to an occluding object used to block the makeup application area;
[0008] Obtain a first segmentation mask map corresponding to the makeup application area in the first face image, and obtain a second segmentation mask map corresponding to the occluding object in the first occluder image;
[0009] Use the second segmentation mask map to fuse the first face image and the first occluder image to generate a second face image in which the makeup application area is covered by the occluding object;
[0010] Generate a third segmentation mask map corresponding to the makeup area on the second face image according to the first segmentation mask map and the second segmentation mask map, where the makeup area includes: the area in the area where the makeup application part is located and not covered by the occlusion object.
[0011] According to a second aspect of the present application, a makeup trial method is disclosed. The method includes:
[0012] Obtain a target face image to be made up and the color value of a target makeup product;
[0013] Input the target face image into a target segmentation model for processing to obtain a fifth segmentation mask map corresponding to the target makeup area in the target face image, where the target segmentation model is trained based on the training data generated by the method according to the first aspect;
[0014] Generate a target makeup color value corresponding to the target makeup area according to a target color generation strategy, the initial color value corresponding to the target makeup area, and the color value of the target makeup product, where the target color generation strategy includes: for the area in the target makeup area where the initial color value is not higher than a second value, generate the corresponding target makeup color value in a multiply mode, and for the area higher than the second value, generate the corresponding target makeup color value in a screen mode;
[0015] Generate a makeup trial face image according to the fifth segmentation mask map and the target makeup color value corresponding to the target makeup area.
[0016] According to a third aspect of the present application, a device for generating training data is disclosed. The device includes:
[0017] A first acquisition module for acquiring a first face image and a first occlusion image, where the makeup application part in the first face image is not occluded, and the first occlusion image is an image corresponding to an occlusion object for occluding the makeup application part;
[0018] A second acquisition module for acquiring a first segmentation mask map corresponding to the makeup application part in the first face image, and acquiring a second segmentation mask map corresponding to the occlusion object in the first occlusion image;
[0019] A first generation module for using the second segmentation mask map to fuse the first face image and the first occlusion image to generate a second face image in which the makeup application part is covered by the occlusion object;
[0020] A second generation module, configured to generate a third segmentation mask map corresponding to the make-up area on the second face image according to the first segmentation mask map and the second segmentation mask map, where the make-up area includes: an area in the area where the make-up part is located and not covered by the occlusion object.
[0021] According to a fourth aspect of the present application, a make-up trial device is disclosed, and the device includes:
[0022] A third acquisition module, configured to acquire a target face image to be made up and the color value of a target make-up product;
[0023] A processing module, configured to input the target face image into a target segmentation model for processing to obtain a fifth segmentation mask map corresponding to the target make-up area in the target face image, where the target segmentation model is trained based on the training data generated by the method according to the first aspect;
[0024] A third generation module, configured to generate a target make-up color value corresponding to the target make-up area according to a target color generation strategy, the initial color value corresponding to the target make-up area, and the color value of the target make-up product, where the target color generation strategy includes: for areas in the target make-up area where the initial color value is not higher than a second value, the target make-up color value is generated in a multiply mode, and for areas higher than the second value, the target make-up color value is generated in a screen mode;
[0025] A fourth generation module, configured to generate a make-up trial face image according to the fifth segmentation mask map and the target make-up color value corresponding to the target make-up area.
[0026] According to a fifth aspect of the present application, an electronic device is disclosed, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the method for generating training data as in the first aspect, or the processor executes the computer program to implement the make-up trial method as in the second aspect.
[0027] According to a sixth aspect of the present application, a computer-readable storage medium is disclosed, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the method for generating training data as in the first aspect is implemented, or when the computer program / instructions are executed by a processor, the make-up trial method as in the second aspect is implemented.
[0028] According to a seventh aspect of the present application, a computer program product is disclosed, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the method for generating training data as in the first aspect is implemented, or when the computer program / instructions are executed by a processor, the make-up trial method as in the second aspect is implemented.
[0029] In the embodiments of the present application, existing face images, occluder images, and corresponding mask data can be obtained. Through image fusion, the occluder is covered on the makeup area of the face image, thereby obtaining a large amount of segmented data and annotation information of the makeup area with occlusion for the training of the segmentation model to improve the user's makeup try-on experience based on the trained segmentation model. Compared with the prior art of collecting and annotating segmented data of the makeup area with occlusion, in the embodiments of the present application, a larger amount of data can be obtained, the annotation difficulty can be reduced, and thus the acquisition cost of training data can be reduced.
[0030] In the embodiments of the present application, when performing virtual makeup try-on, a segmentation model trained based on the aforementioned generated training data is used to predict the makeup area. Based on the predicted makeup area and the color value of the makeup product, a more refined makeup template image is generated, and then the template image is fused with the original image to be made up, making the makeup try-on effect more natural and delicate. Description of the Drawings
[0031] Figure 1 is a flowchart of a method for generating training data according to an embodiment of the present application;
[0032] Figure 2 is an example diagram of dense point annotation of lips according to an embodiment of the present application;
[0033] Figure 3 is an example diagram of a first occluder image according to an embodiment of the present application;
[0034] Figure 4 is a flowchart of a method for training a segmentation model according to an embodiment of the present application;
[0035] Figure 5 is a flowchart of a makeup try-on method according to an embodiment of the present application;
[0036] Figure 6 is an example diagram of a fifth segmentation mask in a makeup try-on method according to an embodiment of the present application;
[0037] Figure 7 is a schematic structural diagram of a device for generating training data according to an embodiment of the present application;
[0038] Figure 8 is a schematic structural diagram of a makeup try-on device according to an embodiment of the present application;
[0039] Figure 9 is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments
[0040] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] It should be noted that for method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.
[0042] In recent years, important progress has been made in the research of technologies such as computer vision, deep learning, machine learning, image processing, and image recognition based on artificial intelligence. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. The discipline of artificial intelligence is a comprehensive discipline, involving many technical categories such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. As an important branch of artificial intelligence, computer vision specifically enables machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, character recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning, etc. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, face access, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart home, wearable devices, driverless, autonomous driving, intelligent medical care, face payment, face unlocking, fingerprint unlocking, person-card verification, smart screen, smart TV, cameras, mobile Internet, webcasting, beauty, makeup, medical beauty, intelligent temperature measurement, etc.
[0043] Taking the beauty makeup field as an example, more and more people choose to buy beauty makeup products such as lipsticks and foundations online. However, beauty makeup products are relatively special and people need to actually try them before they can better decide whether to buy them. Therefore, a virtual makeup try-on algorithm for beauty makeup products is needed. An excellent virtual makeup try-on algorithm can truly reflect the actual use effect of beauty makeup products and increase users' desire to purchase.
[0044] In the prior art, a virtual makeup algorithm based on key points is used for virtual makeup. Specifically, different face images are collected, key points are marked on the makeup areas of the face images, and a model for predicting the makeup areas on the face images is trained based on the marked face images. During actual use, the face image to be made up is input into the model, and the predicted makeup areas are output, and virtual makeup is performed on the predicted makeup areas.
[0045] However, when there is a certain occlusion in the makeup area of the face image to be made up, since the makeup areas of the face images used for model training in the prior art are not occluded by other objects, the model will still output the predicted point coordinates of the area where the makeup area is located. Therefore, the occluded position of the makeup area will also be made up, affecting the authenticity of virtual makeup.
[0046] In order to train a makeup area prediction model based on segmentation, relevant segmentation data is necessarily required. However, compared with the key point model, the acquisition cost of the training data for the segmentation model is higher, especially the segmentation data with occlusion. Taking the lip segmentation data as an example, not only different people's lips need to be collected, but also different occluders need to be collected. In addition, the annotation cost of the segmentation data is significantly higher than the key point annotation cost. For example, for key points, the annotation price of a face image is about 0.5 yuan, while the annotation cost of a segmentation image of the same size is about 5 times that. It can be seen that the acquisition of the training data for training the segmentation model in the prior art is difficult and costly.
[0047] To solve the above technical problems, the embodiments of the present application provide a method for generating training data, a makeup method, an electronic device, and a storage medium.
[0048] First, a method for generating training data provided by the embodiments of the present application will be introduced below.
[0049] Figure 1 is a flowchart of the method for generating training data according to an embodiment of the present application. As Figure 1 shown, the method may include the following steps: Step 101, Step 102, Step 103, and Step 104, where
[0050] In Step 101, a first face image and a first occluder image are obtained, where the makeup area in the first face image is not occluded, and the first occluder image is an image corresponding to the occluding object used to occlude the makeup area.
[0051] In the embodiments of the present application, first face images in various different states such as different skin colors, different scenarios, and different lighting conditions, as well as first occluder images covering different types of occluding objects, can be used to generate training data. In practical applications, the first occlusion image can be sourced from the COCO dataset or obtained by taking a photo of the occluding object.
[0052] In the embodiments of the present application, the makeup application areas may include: lips, eyebrows, eyes, nose, ears, or the face, etc.
[0053] For example, when training a segmentation model for predicting the makeup application area on the lips, the makeup application area in the first face image is the lips;
[0054] When training a segmentation model for predicting the makeup application area on the eyebrows, the makeup application area in the first face image is the eyebrows;
[0055] When training a segmentation model for predicting the makeup application area on the eyes, the makeup application area in the first face image is the eyes;
[0056] When training a segmentation model for predicting the makeup application area on the nose, the makeup application area in the first face image is the nose;
[0057] When training a segmentation model for predicting the makeup application area on the ears, the makeup application area in the first face image is the ears;
[0058] When training a segmentation model for predicting the makeup application area on the face, the makeup application area in the first face image is the face.
[0059] In the embodiments of the present application, the first face image is required to have as little occlusion as possible. The reason for not selecting face images with occlusion is that face data with occlusion is difficult to accurately annotate using key points.
[0060] In the embodiments of the present application, the occluding object can be a mask, a backpack, a hand, or an eye mask, etc.
[0061] When generating training data for a segmentation model for predicting the makeup application area, a large number of first face images and first occluder images are required. For the sake of easy understanding and description, in the embodiments of the present application, it is taken as an example that one first face image and one first occluder image generate a set of training data (including an image and corresponding annotation data). The process of generating other sets of training data based on other first face images and other first occluder images is similar to the process of generating a set of training data described above, and will not be elaborated here.
[0062] In step 102, obtain a first segmentation mask map corresponding to the makeup application area in the first face image, and obtain a second segmentation mask map corresponding to the occluding object in the first occluder image.
[0063] In the embodiment of the present application, the first segmentation mask map is the mask of the area where the makeup area is located in the first face image. The first segmentation mask map is a binary image, and its image size is the same as that of the first face image. The color of the area where the makeup area is located in the first segmentation mask map is white, and the other areas are black.
[0064] In the embodiment of the present application, the second segmentation mask map is the mask of the area where the occluding object is located in the first occluding object image. The second segmentation mask map is also a binary image, and its image size is the same as that of the first occluding object image. The color of the area where the occluding object is located in the second segmentation mask map is white, and the other areas are black.
[0065] For ease of understanding, hereinafter, taking the makeup area as the lips as an example, the generation process of the training data will be described.
[0066] In the embodiment of the present application, after obtaining the first face image, the face key points and lip dense points in the first face image can be labeled. For example, Figure 2 shows the positions of the lip dense points in the labeled face image. Through the lip dense points, the contour information of the lips can be outlined, so as to obtain the mask of the lip area, that is, the first segmentation mask map.
[0067] When performing the above labeling, the face key points of the first face image can be obtained by performing key point detection or directly obtained.
[0068] In the embodiment of the present application, after obtaining the first segmentation mask map, the segmentation data of the lips has been obtained to a certain extent. However, since the collected data does not contain occlusion information, the segmentation model trained with this data cannot handle occlusion problems. Therefore, it is necessary to obtain the segmentation data of the occluding object.
[0069] In the embodiment of the present application, after obtaining the first occluding object image, the first occluding object image can be input into an existing object segmentation model for processing to obtain the mask of the occluding object, that is, the second segmentation mask image. Considering that the mask of the occluding object is usually included in the publicly available object segmentation data sets on the network, it can be directly obtained from the publicly available object segmentation data sets to save a certain amount of labeling cost. For example, Figure 3 shows the image of a backpack in the publicly available object segmentation data set.
[0070] In step 103, using the second segmentation mask map, the first face image and the first occluding object image are fused to generate a second face image in which the makeup area is covered by the occluding object.
[0071] In the embodiments of the present application, a second segmentation mask image can be used to perform Alpha blending on the first face image and the first occluder image to generate a second face image in which the made-up part is covered by the occluded object. The Alpha blending process can be represented by the following formula: output = foreground * mask + background * (1 - mask), where foreground is the pixel value of each pixel point in the foreground image, background is the pixel value of each pixel point in the background image, mask is the pixel value of each pixel point in the mask image, and output is the pixel value of each pixel point in the new image.
[0072] When using the above formula, first, according to the above formula, the pixel values of the corresponding position pixel points in the foreground image, the background image, and the mask image are calculated to obtain the pixel values of each pixel point in the new image. Then, based on the pixel values of each pixel point in the new image obtained by the calculation, the above new image is generated.
[0073] In the embodiments of the present application, the first face image is used as the background image, the first occluder image is used as the foreground image, and the second segmentation mask image is used as the mask image. According to the above formula, the pixel values of the corresponding position pixel points in the second segmentation mask image, the first face image, and the first occluder image are calculated to obtain the pixel values of each position pixel point. Based on the pixel values of each position pixel point obtained by the calculation, a second face image, that is, a face image with occlusion, is generated.
[0074] In an example, the first face image is a lip image, and the first occluder image is a human hand image. The lip image is used as the background image, and the human hand image is used as the foreground image. The mask of the human hand can be obtained through step 102. Using the mask of the human hand, the lip image and the human hand image are subjected to Alpha blending to obtain a blended image in which the lips are occluded by the human hand.
[0075] In an embodiment provided by the present application, in order to make the occlusion effect more realistic and natural, the color of the foreground object can be appropriately adjusted so that the foreground object and the background are more matched in the color dimension. Therefore, before covering the occluded object to the made-up part of the first face image, the color of the occluded object can be adjusted first. After the color adjustment is completed, it is then covered to the made-up part of the first face image. At this time, before the above step 103, the following step (not shown in the figure) can also be added: step 105, where
[0076] In step 105, according to the color value of the skin area in the first face image, the color value of the occluded object in the first occluder image is adjusted to obtain a second occluder image, where the difference between the color value of the occluded object in the second occluder image and the color value of the skin area in the first face image is less than the first value.
[0077] Among them, the specific value of the above first numerical value can be set according to actual needs, and the embodiments of the present application do not limit the specific value of the above first numerical value.
[0078] In the embodiments of the present application, the color value of the occluding object can be adjusted according to the color value of the skin area in the first face image, so that the color effect of the occluding object in the second occluding object image is closer to the color effect of the skin area in the first face image, and the occlusion effect is more natural during image fusion.
[0079] In some embodiments, the above step 105 may specifically include the following steps (not shown in the figure): step 1051, step 1052, and step 1053, where
[0080] In step 1051, obtain the first average color value of the skin area in the first face image and the second average color value of the area where the occluding object is located in the first occluding object image.
[0081] In the embodiments of the present application, the color value of the image can be divided into color values of the R, G, and B channels. In this case, the first average color value of the skin area in the first face image includes: the R average value, G average value, and B average value of the skin area; the second average color value of the area where the occluding object is located in the first occluding object image includes: the R average value, G average value, and B average value of the area where the occluding object is located.
[0082] In step 1052, determine the target color value corresponding to each pixel point in the area where the occluding object is located according to the first average color value, the second average color value, and the initial color value corresponding to each pixel point in the area where the occluding object is located.
[0083] In the embodiments of the present application, when the first average color value includes: the R average value, G average value, and B average value of the skin area, and the second average color value includes: the R average value, G average value, and B average value of the area where the occluding object is located, the target color value corresponding to each pixel point in the area where the occluding object is located includes: the R value, G value, and B value of the target color of the pixel point.
[0084] For the convenience of understanding, the color value of one channel is taken as an example for description. For example, to calculate the R value of the target color, first obtain the R average value R of the skin area f and the R average value R of the area where the occluding object is located b , for each pixel point A in the area where the occluding object is located, the R value of its initial color is x, and using the single mapping function f(x), calculate the R value of the target color corresponding to pixel point A, where the R value of the target color corresponding to A = f(x).
[0085]
[0086] Similarly, the G value and B value of the target color corresponding to pixel point A can be calculated.
[0087] In step 1053, the initial color values corresponding to the pixel points in the area where the occluding object is located in the first occluding object image are replaced with the corresponding target color values to obtain a second occluding object image.
[0088] In the embodiments of the present application, compared with the first occluding object image, the color effect of the occluding object in the second occluding object image is closer to the color effect of the skin area in the first face image.
[0089] It can be seen that in the embodiments of the present application, the authenticity of the occlusion effect can be improved by adjusting the color value of the occluding object to make the color effect of the occluding object closer to the color effect of the skin area in the first face image.
[0090] Considering that according to the second segmentation mask image, when the occluding object is directly covered on the makeup area of the first face image, there will be an obvious edge of the occluding object, that is, an edge transition band, in the obtained second face image, and the fusion effect is not natural. To solve the above problems, in another embodiment provided by the present application, the above step 103 may specifically include the following steps:
[0091] Perform edge smoothing processing on the contour of the occluding object in the second segmentation mask image to obtain a fourth segmentation mask image; use the fourth segmentation mask image to perform Alpha fusion on the first face image and the first occluding object image to generate a second face image in which the makeup area is covered by the occluding object.
[0092] In the embodiments of the present application, the contour of the occluding object in the second segmentation mask image can be first subjected to edge smoothing processing to make the edge of the occluding object smoother, obtaining a fourth segmentation mask image, and then the fourth segmentation mask image is used to cover the occluding object on the makeup area of the first face image, making the fusion of the two images more natural, thereby improving the authenticity of the occlusion effect.
[0093] In the embodiments of the present application, edge smoothing processing can be performed on the contour of the occluding object in the second segmentation mask image by using Gaussian blur or mean blur. Taking Gaussian blur as an example, the Gaussian blur method is as follows:
[0094] First, determine the Gaussian kernel according to the radius size and variance. The calculation formula of the two-dimensional Gaussian kernel is
[0095] After that, for each center point, calculate the pixel weights in the neighborhood according to the Gaussian kernel and sum them to obtain the pixel value; finally, repeat this process for all points to obtain the Gaussian-blurred image.
[0096] In step 104, according to the first segmentation mask image and the second segmentation mask image, a third segmentation mask image corresponding to the make-up area on the second face image is generated, where the make-up area includes: the area in the area where the make-up part is located that is not covered by the occluding object.
[0097] In the embodiment of the present application, when using the second segmentation mask image to perform Alpha fusion on the first face image and the first occluding object image, the pixel values of the corresponding pixels in the first segmentation mask image and the second segmentation mask image can be subtracted to obtain the third segmentation mask image corresponding to the make-up area on the second face image.
[0098] In the embodiment of the present application, when using the fourth segmentation mask image to perform Alpha fusion on the first face image and the first occluding object image, the pixel values of the corresponding pixels in the first segmentation mask image and the fourth segmentation mask image can be subtracted to obtain the third segmentation mask image corresponding to the make-up area on the second face image.
[0099] In the embodiment of the present application, the second face image and the corresponding third segmentation mask image form a set of training data. When performing model training, the second face image is used as the input of the model, and the corresponding third segmentation mask image is used as the output target for model training.
[0100] It can be seen that in the embodiment of the present application, through steps 101 to 104, based on the existing key point annotations of the make-up part and the publicly available object mask data, the occluding object can be covered on the known make-up part by means of image fusion, so as to obtain a large amount of segmented data of the make-up part with occlusion.
[0101] As can be seen from the above embodiments, in this embodiment, the existing face image, occluding object image and corresponding mask data can be obtained, and the occluding object can be covered on the make-up part of the face image by means of image fusion, so as to obtain a large amount of segmented data and annotation information of the make-up area with occlusion for the training of the segmentation model, so as to improve the user's makeup trial experience based on the trained segmentation model. Compared with the prior art of collecting and annotating the segmented data of the make-up area with occlusion, in the embodiment of the present application, a larger amount of data can be obtained, the annotation difficulty can be reduced, and thus the acquisition cost of the training data can be reduced.
[0102] Figure 4 is a flowchart of a method for training a segmentation model according to an embodiment of the present application. In the embodiment of the present application, based on Figure 1 the training data generated in the shown embodiment, a segmentation model for predicting the make-up area is trained, as Figure 4 shown, the method may include the following steps: step 401 and step 402, where,
[0103] In step 401, an initial segmentation model is constructed.
[0104] In the embodiments of the present application, an existing convolutional neural network algorithm or other deep learning algorithms can be used to construct an initial segmentation model for predicting the makeup area.
[0105] In step 402, multiple different second face images are respectively input into the initial segmentation model for processing, and corresponding predicted segmentation mask images are output. Based on the third segmentation mask images and the predicted segmentation mask images corresponding to the multiple different second face images, the initial segmentation model is trained until the model converges, and a target segmentation model is obtained.
[0106] In the embodiments of the present application, the second face image is input into the constructed model to obtain the segmentation mask image predicted by the model. The predicted makeup area is marked in the predicted segmentation mask image. The loss is calculated between the predicted makeup area and the makeup area marked in the third segmentation mask image, and the parameters of the model are updated according to the loss. Then, the second face image is input into the model again, and the above process is repeated until the model converges. Finally, a trained segmentation model, that is, the target segmentation model, is obtained.
[0107] As can be seen from the above embodiments, in this embodiment, a face image with the makeup area occluded is used for training the segmentation model, and the obtained target segmentation model can be used in the makeup trial scenario to make the makeup trial effect more realistic and natural.
[0108] Figure 5 It is a flowchart of a makeup trial method according to an embodiment of the present application. In the embodiments of the present application, based on Figure 4 the segmentation model trained in the shown embodiment, virtual makeup trial is performed. As Figure 5 shown, the method may include the following steps: step 501, step 502, step 503, and step 504, where
[0109] In step 501, a target face image to be made up and the color value of the target makeup product are obtained.
[0110] In the embodiments of the present application, the target face image to be made up can be an image previously taken by the user or an image taken by the user in real time.
[0111] In the embodiments of the present application, the target makeup product is the makeup product selected by the user.
[0112] Taking virtual lipstick makeup as an example, the target makeup product is a lipstick of a certain brand and a certain color number selected by the user.
[0113] In step 502, the target face image is input into the target segmentation model for processing to obtain a fifth segmentation mask image corresponding to the target makeup area in the target face image.
[0114] In the embodiment of the present application, the target makeup application area is the makeup application area predicted by the target segmentation model.
[0115] Still taking the virtual lipstick makeup application as an example, the target face image (where the target face image is an image with the lips blocked by a finger) is input into the target segmentation model for processing, and the Figure 6 fifth segmentation mask image shown is obtained. Among them, the white area in the fifth segmentation mask image corresponds to the target makeup application area in the target face image, that is, the white area is the area where the lips are not blocked by the finger.
[0116] In step 503, according to the target color generation strategy, the initial color value corresponding to the target makeup application area, and the color value of the target makeup product, the target makeup color value corresponding to the target makeup application area is generated. Among them, the target color generation strategy includes: for the area where the initial color value in the target makeup application area is not higher than the second value, the color is generated by using the multiply method to generate the corresponding target makeup color value, and for the area higher than the second value, the screen method is used to generate the corresponding target makeup color value.
[0117] In the embodiment of the present application, the color value of the image area can be divided into the color values of the R, G, and B channels. In this case, the initial color value corresponding to the target makeup application area includes: the R value, G value, and B value of the initial color of each pixel point in the target makeup application area, and the color value of the target makeup product includes: the R value, G value, and B value of the color of the target makeup product, and the target makeup color value corresponding to the target makeup application area: the R value, G value, and B value of the target makeup color corresponding to each pixel point in the target makeup application area.
[0118] In practical applications, the second value can be 50%, of course, it can also be other values. The specific value of the above second value can be set according to the actual application scenario, and the embodiment of the present application does not limit the specific value of the second value.
[0119] In the embodiment of the present application, the target makeup color values corresponding to the color values of the three channels are respectively generated. For the sake of easy understanding, an example of the color value of one channel is used for description. For example, to calculate the R value of the target makeup color corresponding to each pixel point i in the target makeup application area, first obtain the R value R of the initial color of pixel point i i and the R value R of the color of the target makeup product z , and then taking the virtual lipstick try-on as an example, the formula g(R i ) can be used to calculate the R value of the target makeup color corresponding to pixel point i. Among them, the R value of the target makeup color corresponding to pixel point i = g(R i ).
[0120]
[0121] Similarly, the G value and B value of the target makeup color corresponding to each pixel point i in the target makeup area can be calculated.
[0122] In the embodiment of the present application, using the target color generation strategy can achieve the following effects: the colors that change are mainly in the medium brightness area, and the colors in the bright and dark areas (i.e., the highlight and shadow areas) basically remain unchanged, making the makeup try-on effect more natural.
[0123] In step 504, according to the fifth segmentation mask map and the target makeup color value corresponding to the target makeup area, a makeup try-on face image is generated.
[0124] In an embodiment provided by the present application, the above step 504 may include the following steps (not shown in the figure): step 5041 and step 5042, where
[0125] In step 5041, according to the fifth segmentation mask map and the target makeup color value corresponding to the target makeup area, a makeup try-on template map for applying makeup to the target makeup area is generated.
[0126] In the embodiment of the present application, the size and contour of the makeup try-on template map are the same as those of the target makeup area, and the color value of each pixel point in the makeup try-on template map is the same as the target makeup color value corresponding to the pixel point at the corresponding position in the target makeup area.
[0127] In step 5042, the makeup try-on template map is covered on the target makeup area of the target face image to obtain a makeup try-on face image.
[0128] In the embodiment of the present application, the makeup try-on template map can be directly covered on the target makeup area of the target face image for image fusion to obtain a makeup try-on face image. Or a transparency coefficient can be used to simulate the thick and thin application effects of makeup. At this time, the above step 5042 may include the following steps (not shown in the figure): step 50421, step 50422, and step 50423, where
[0129] In step 50421, a transparency coefficient for adjusting the depth effect of the makeup color is obtained;
[0130] In the embodiment of the present application, the transparency coefficient can be selected by the user, or can be automatically recommended for the user according to the user's behavior habits.
[0131] In step 50422, according to the transparency coefficient, the transparency of the makeup try-on template map and the target makeup area of the target face image is adjusted respectively.
[0132] In step 50423, the adjusted trial makeup template image is overlaid on the target makeup application area of the adjusted target face image to obtain a trial makeup face image.
[0133] In the embodiments of the present application, the following formula can be used: T = k * C + (1 - k) * B to adjust the transparency of the trial makeup template image and the target makeup application area of the target face image, where k is the transparency coefficient, C is the trial makeup template image, B is the target makeup application area, and T is the effect image presented after the fusion of the two images.
[0134] As can be seen from the above embodiments, in this embodiment, when performing virtual trial makeup, a target segmentation model is used to predict the makeup application area, and the predicted makeup application area is the area in the makeup application part area that is not blocked by the blocking object. Based on the predicted makeup application area and the color of the makeup product, a more refined makeup template image is generated, and then the template image is fused with the original image to be made up, making the trial makeup effect more natural and delicate.
[0135] Figure 7 It is a schematic structural diagram of a training data generation device according to an embodiment of the present application. As Figure 7 shown, the training data generation device 700 may include: a first acquisition module 701, a second acquisition module 702, a first generation module 703, and a second generation module 704, where
[0136] The first acquisition module 701 is configured to acquire a first face image and a first occluder image, where the makeup application part in the first face image is not occluded, and the first occluder image is an image corresponding to an occluding object used to occlude the makeup application part;
[0137] The second acquisition module 702 is configured to acquire a first segmentation mask image corresponding to the makeup application part in the first face image, and acquire a second segmentation mask image corresponding to the occluding object in the first occluder image;
[0138] The first generation module 703 is configured to use the second segmentation mask image to fuse the first face image and the first occluder image to generate a second face image in which the makeup application part is covered by the occluding object;
[0139] The second generation module 704 is configured to generate a third segmentation mask image corresponding to the makeup application area in the second face image according to the first segmentation mask image and the second segmentation mask image, where the makeup application area includes: the area in the makeup application part area that is not covered by the occluding object.
[0140] As can be seen from the above embodiments, in this embodiment, existing face images, occluder images, and corresponding mask data can be obtained. By means of image fusion, the occluder is covered on the makeup area of the face image, so as to obtain a large amount of segmented data and annotation information of the makeup area with occlusion, which is used for the training of the segmentation model to improve the user's makeup try-on experience based on the trained segmentation model. Compared with collecting and annotating the segmented data of the makeup area with occlusion in the prior art, in the embodiments of the present application, a larger amount of data can be obtained, the annotation difficulty can be reduced, and thus the acquisition cost of the training data can be reduced.
[0141] Optionally, as an embodiment, the training data generation device 700 may further include:
[0142] A fifth generation module, configured to adjust the color value of the occluding object in the first occluder image according to the color value of the skin area in the first face image, to obtain a second occluder image, where the difference between the color value of the occluding object in the second occluder image and the color value of the skin area in the first face image is less than a first value.
[0143] Optionally, as an embodiment, the fifth generation module may include:
[0144] An acquisition sub-module, configured to acquire a first average color value of the skin area in the first face image and a second average color value of the area where the occluding object is located in the first occluder image;
[0145] A determination sub-module, configured to determine the target color value corresponding to each pixel point in the area where the occluding object is located according to the first average color value, the second average color value, and the initial color value corresponding to each pixel point in the area where the occluding object is located;
[0146] A replacement sub-module, configured to replace the initial color value corresponding to each pixel point in the area where the occluding object is located in the first occluder image with the corresponding target color value, to obtain a second occluder image.
[0147] Optionally, as an embodiment, the first generation module 703 may include:
[0148] A processing sub-module, configured to perform edge smoothing processing on the contour of the occluding object in the second segmentation mask map, to obtain a fourth segmentation mask map;
[0149] A fusion sub-module, configured to use the fourth segmentation mask map to perform Alpha fusion on the first face image and the first occluder image, to generate a second face image in which the makeup area is covered by the occluding object;
[0150] The second generation module 704 may include:
[0151] An operator module, configured to perform a subtraction operation on the first segmentation mask map and the fourth segmentation mask map to obtain a third segmentation mask map corresponding to the makeup area on the second face image.
[0152] Optionally, as an embodiment, the training data generation device 700 may further include:
[0153] A training module, configured to construct an initial segmentation model; input a plurality of different second face images into the initial segmentation model for processing, output corresponding predicted segmentation mask maps, and train the initial segmentation model based on the third segmentation mask maps and the predicted segmentation mask maps corresponding to the plurality of different second face images until the model converges to obtain a target segmentation model.
[0154] Any step and the specific operation in any step in the embodiments of the training data generation method provided in this application can be completed by the corresponding module in the training data generation device. The process of the corresponding operations completed by each module in the training data generation device refers to the process of the corresponding operations described in the embodiments of the training data generation method.
[0155] Figure 8 It is a schematic structural diagram of a makeup try-on device according to an embodiment of the present application. As Figure 8 shown, the makeup try-on device 800 may include: a third acquisition module 801, a processing module 802, a third generation module 803, and a fourth generation module 804, where
[0156] The third acquisition module 801 is configured to acquire a target face image to be made up and the color value of a target makeup product.
[0157] The processing module 802 is configured to input the target face image into a target segmentation model for processing to obtain a fifth segmentation mask map corresponding to the target makeup area on the target face image, where the target segmentation model is generated based on the training data generation method in the first aspect.
[0158] The third generation module 803 is configured to generate a target makeup color value corresponding to the target makeup area according to a target color generation strategy, the initial color value corresponding to the target makeup area, and the color value of the target makeup product, where the target color generation strategy includes: using the multiply mode to generate the corresponding target makeup color value for the area where the initial color value in the target makeup area is not higher than a second value, and using the screen mode to generate the corresponding target makeup color value for the area higher than the second value.
[0159] The fourth generation module 804 is configured to generate a trial makeup face image according to the fifth segmentation mask image and the target makeup color value corresponding to the target makeup area.
[0160] As can be seen from the above embodiments, in this embodiment, when performing virtual makeup trial, a segmentation model trained using the training data generated as described above is used to predict the makeup area. Based on the predicted makeup area and the color value of the makeup product, a more refined makeup template image is generated, and then the template image is fused with the original image to be made up, making the makeup trial effect more natural and delicate.
[0161] Optionally, as an embodiment, the fourth generation module 804 may include:
[0162] A generation sub-module, configured to generate a trial makeup template image for applying makeup to the target makeup area according to the fifth segmentation mask image and the target makeup color value corresponding to the target makeup area;
[0163] An overlay sub-module, configured to overlay the trial makeup template image on the target makeup area of the target face image to obtain a trial makeup face image.
[0164] Optionally, as an embodiment, the overlay sub-module may include:
[0165] An acquisition unit, configured to acquire a transparency coefficient for adjusting the depth effect of the makeup color;
[0166] An adjustment unit, configured to adjust the transparency of the trial makeup template image and the target makeup area of the target face image respectively according to the transparency coefficient;
[0167] An overlay unit, configured to overlay the adjusted trial makeup template image on the adjusted target makeup area of the target face image to obtain a trial makeup face image.
[0168] Any step in the embodiments of the makeup trial method provided in this application and the specific operations in any step can be completed by the corresponding modules or units in the makeup trial device. The process of the corresponding operations completed by each module or unit in the makeup trial device refers to the process of the corresponding operations described in the embodiments of the makeup trial method.
[0169] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.
[0170] Figure 9It is a structural block diagram of an electronic device according to an embodiment of the present application. The electronic device includes a processing component 922, which further includes one or more processors, and memory resources represented by a memory 932 for storing instructions executable by the processing component 922, such as application programs. The application programs stored in the memory 932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute instructions to perform the above method.
[0171] The electronic device may further include a power component 926 configured to perform power management of the electronic device, a wired or wireless network interface 950 configured to connect the electronic device to a network, and an input / output (I / O) interface 958. The electronic device may operate based on an operating system stored in the memory 932, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0172] According to still another embodiment of the present application, the present application also provides a computer-readable storage medium, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the steps in the method for generating training data described in any one of the above embodiments are implemented, or when the computer programs / instructions are executed by a processor, the steps in the method for trying on makeup described in any one of the above embodiments are implemented.
[0173] According to still another embodiment of the present application, the present application also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps in the method for generating training data described in any one of the above embodiments are implemented, or when the computer programs / instructions are executed by a processor, the steps in the method for trying on makeup described in any one of the above embodiments are implemented.
[0174] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0175] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0176] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0178] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0179] Finally, it should also be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or terminal device including the said element.
[0180] The above has introduced in detail a method for generating training data, a makeup try-on method, an electronic device, and a storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for generating training data, characterized in that, The method includes: Obtaining a first face image and a first occluder image, wherein the made-up part in the first face image is not occluded, and the first occluder image is an image corresponding to an occluding object for occluding the made-up part; Obtaining a first segmentation mask map corresponding to the made-up part in the first face image, and obtaining a second segmentation mask map corresponding to the occluding object in the first occluder image; Using the second segmentation mask map to fuse the first face image and the first occluder image to generate a second face image in which the made-up part is covered by the occluding object; Generating a third segmentation mask map corresponding to the made-up area in the second face image according to the first segmentation mask map and the second segmentation mask map, wherein the made-up area includes: an area in the area where the made-up part is located and not covered by the occluding object, and the third segmentation mask map is obtained by subtracting the pixel values of the corresponding position pixel points in the first segmentation mask map and the second segmentation mask map; the second face image and the third segmentation mask map form a set of training data.
2. The method according to claim 1, characterized in that, Before the step of using the second segmentation mask map to fuse the first face image and the first occluder image to generate a second face image in which the made-up part is covered by the occluding object, it further includes: Adjusting the color value of the occluding object in the first occluder image according to the color value of the skin area in the first face image to obtain a second occluder image, wherein the difference between the color value of the occluding object in the second occluder image and the color value of the skin area in the first face image is less than a first value.
3. The method according to claim 2, wherein The adjusting the color value of the occluding object in the first occluder image according to the color value of the skin area in the first face image to obtain a second occluder image includes: Obtaining a first average color value of the skin area in the first face image and a second average color value of the area where the occluding object is located in the first occluder image; Determining the target color value corresponding to each pixel point in the area where the occluding object is located according to the first average color value, the second average color value, and the initial color value corresponding to each pixel point in the area where the occluding object is located; Replacing the initial color value corresponding to each pixel point in the area where the occluding object is located in the first occluder image with the corresponding target color value to obtain a second occluder image.
4. The method according to any one of claims 1 to 3, characterized in that, The using the second segmentation mask map to fuse the first face image and the first occluder image to generate a second face image in which the made-up part is covered by the occluding object includes: Performing edge smoothing processing on the contour of the occluding object in the second segmentation mask map to obtain a fourth segmentation mask map; Using the fourth segmentation mask map to perform Alpha fusion on the first face image and the first occluder image to generate a second face image in which the made-up part is covered by the occluding object; Generating a third segmentation mask map corresponding to the make-up area in the second face image according to the first segmentation mask map and the second segmentation mask map includes: Performing a subtraction operation on the first segmentation mask map and the fourth segmentation mask map to obtain a third segmentation mask map corresponding to the make-up area in the second face image.
5. The method according to claim 1, characterized in that, The method further includes: Constructing an initial segmentation model; Inputting multiple different second face images into the initial segmentation model for processing, outputting corresponding predicted segmentation mask maps, and training the initial segmentation model based on the third segmentation mask maps and the predicted segmentation mask maps corresponding to the multiple different second face images until the model converges to obtain a target segmentation model.
6. A makeup try-on method, characterized in that, The method includes: Obtaining a target face image to be made up and the color value of a target make-up product; Inputting the target face image into the target segmentation model for processing to obtain a fifth segmentation mask map corresponding to the target make-up area in the target face image, where the target segmentation model is trained based on the training data generated by the method according to any one of claims 1-5; Generating a target make-up color value corresponding to the target make-up area according to a target color generation strategy, the initial color value corresponding to the target make-up area, and the color value of the target make-up product, where the target color generation strategy includes: using the multiply method to generate a corresponding target make-up color value for the area where the initial color value in the target make-up area is not higher than a second value, and using the screen method to generate a corresponding target make-up color value for the area higher than the second value; Generating a trial make-up face image according to the fifth segmentation mask map and the target make-up color value corresponding to the target make-up area.
7. The method according to claim 6, wherein Generating a trial make-up face image according to the fifth segmentation mask map and the target make-up color value corresponding to the target make-up area includes: Generating a trial make-up template map for making up the target make-up area according to the fifth segmentation mask map and the target make-up color value corresponding to the target make-up area; Covering the trial make-up template map on the target make-up area of the target face image to obtain a trial make-up face image.
8. The method according to claim 7, wherein Covering the trial make-up template map on the target make-up area of the target face image to obtain a trial make-up face image includes: Obtaining a transparency coefficient for adjusting the depth effect of the make-up color; Adjusting the transparency of the trial make-up template map and the target make-up area of the target face image according to the transparency coefficient; Covering the adjusted trial make-up template map on the adjusted target make-up area of the target face image to obtain a trial make-up face image.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-5, or the processor executes the computer program to implement the method according to any one of claims 6-8.
10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-5 is implemented, or when the computer program / instructions are executed by the processor, the method according to any one of claims 6-8 is implemented.
11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the method according to any one of claims 1-5 is implemented, or when the computer program / instructions are executed by a processor, the method according to any one of claims 6-8 is implemented.
Citation Information
Patent Citations
Human face makeup method, device, equipment and medium
CN110689479A
Information processing device, information processing method, and program
JP2020009308A