Model Training Method and Apparatus, Image Processing Method, Electronic Device, Medium
By using the initial generative model in the makeup transfer technology for two-way makeup transfer and adjusting the model according to the target loss value, the problem that the quality of the training set in the prior art affects the stability of the model is solved, and a more robust and flexible makeup transfer effect is achieved.
Patent Information
- Application Number
- CN202210908744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The existing makeup transfer technology relies on accurately paired bare face pictures and makeup pictures as training sets, resulting in the quality of the training set affecting the model stability and the makeup transfer effect is poor.
By obtaining the sample picture of bare face and the sample picture of makeup on the face, using the initial generated model for two-way makeup transfer, calculating the target loss value and adjusting the model, a model that can achieve two-way makeup transfer is built.
Without relying on an accurate pairing training set, a robust makeup transfer model is built by digging out the potential feature relationship between bare face pictures and makeup face pictures, which improves the flexibility and stability of makeup transfer.
Smart Images

Figure CN115222583B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method and apparatus, an image processing method, an electronic device, and a medium. Background Art
[0002] Human face makeup transfer is a common technology in the field of computer vision. By extracting the makeup style from a reference image, it can be applied to the face images of different people. The existing makeup transfer methods usually need to collect paired plain face images and makeup images as the training set, so as to use the training set to train a supervised model for makeup transfer. However, this method depends on an accurately paired training set. Therefore, the quality of the training set is likely to affect the stability of model training, resulting in poor makeup transfer effects. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a model training method and apparatus, an image processing method, an electronic device, and a medium, which can achieve stable makeup transfer effects.
[0004] To achieve the above object, a first aspect of the embodiments of this application proposes a model training method for training a makeup transfer model. The method includes:
[0005] Obtain a plain face sample image and a made-up face sample image. The plain face sample image does not include makeup information, and the made-up face sample image includes makeup information; input the plain face sample image into an initial generation model for processing to obtain a first generation result, where the first generation result includes a first made-up face image converted from the plain face sample image and a first plain face image converted from the first made-up face image; input the made-up face sample image into the initial generation model for processing to obtain a second generation result, where the second generation result includes a second plain face image converted from the made-up face sample image and a second made-up face image converted from the second plain face image; calculate a target loss value according to the first generation result and the second generation result, and use the target loss value to adjust the initial generation model to obtain a makeup transfer model, where the makeup transfer model is used to perform makeup transfer processing on images.
[0006] In some optional embodiments, the inputting the plain face sample image into an initial generation model for processing to obtain a first generation result includes:
[0007] Input the natural face sample image into the initial generation model; through the initial generation model, perform diffusion processing on the natural face sample image to obtain a first original image that satisfies the Gaussian distribution, and perform denoising processing on the first original image to obtain a first made-up face image; through the initial generation model, perform diffusion processing on the first made-up face image to obtain a second original image that satisfies the Gaussian distribution, and perform denoising processing on the second original image to obtain a first natural face image.
[0008] The processing of inputting the made-up face sample image into the initial generation model to obtain a second generation result includes:
[0009] Input the made-up face sample image into the initial generation model; through the initial generation model, perform diffusion processing on the made-up face sample image to obtain a third original image that satisfies the Gaussian distribution, and perform denoising processing on the third original image to obtain a second natural face image; through the initial generation model, perform diffusion processing on the second natural face image to obtain a fourth original image that satisfies the Gaussian distribution, and perform denoising processing on the fourth original image to obtain a second made-up face image.
[0010] In some optional embodiments, the target loss value includes a spatial similarity loss value. The calculating the target loss value according to the first generation result and the second generation result includes:
[0011] Obtain a first reference image and a second reference image. Both the first reference image and the second reference image satisfy the standard normal distribution condition. The dimension of the first reference image is the same as that of the natural face sample image, and the dimension of the second reference image is the same as that of the made-up face sample image; calculate a first spatial loss according to the first reference image and the first original image; calculate a second spatial loss according to the second reference image and the second original image; perform a fusion operation according to the first spatial loss and the second spatial loss to obtain a spatial similarity loss value.
[0012] In some optional embodiments, the target loss value includes a cycle consistency loss value; the calculating the target loss value according to the first generation result and the second generation result includes:
[0013] Calculate a first transfer loss according to the natural face sample image and the first natural face image; calculate a second transfer loss according to the made-up face sample image and the second made-up face image; perform a fusion operation according to the first transfer loss and the second transfer loss to obtain a cycle consistency loss value.
[0014] In some optional embodiments, the target loss value includes a makeup loss value; the calculating the target loss value according to the first generation result and the second generation result includes:
[0015] Perform mask annotation processing on the original face sample image, and obtain at least one feature region and the first mask information corresponding to the feature region from the original face sample image; perform mask annotation processing on the first made-up face image, and obtain the second mask information corresponding to each feature region from the first made-up face image; calculate the first makeup loss corresponding to each feature region according to the first mask information and the second mask information; perform mask annotation processing on the made-up face sample image, and obtain the third mask information corresponding to each feature region from the made-up face sample image; perform mask annotation processing on the second original face image, and obtain the fourth mask information corresponding to each feature region from the second original face image; calculate the second makeup loss corresponding to each feature region according to the third mask information and the fourth mask information; calculate the target makeup loss corresponding to each feature region according to the first makeup loss and the second makeup loss; perform a fusion operation on the target makeup losses corresponding to all the feature regions to obtain a makeup loss value.
[0016] In some optional embodiments, the adjusting the initial generation model by using the target loss value to obtain a makeup transfer model includes:
[0017] If the initial generation model does not meet the training completion condition, adjust the initial generation model by using the target loss value to obtain an adjusted initial generation model, and continue to perform the step of obtaining the original face sample image and the made-up face sample image; if the initial generation model meets the training completion condition, determine the makeup transfer model according to the initial generation model; where the training completion condition includes at least any one of the following: the target loss value is less than a preset loss value; the loss value of the initial generation model increases a preset number of times.
[0018] To achieve the above object, a second aspect of the embodiments of the present application provides an image processing method, the method includes:
[0019] Obtain a face image; if the face image includes makeup information, input the face image into the makeup transfer model for makeup transfer processing to obtain a target original face image, and the target original face image does not include makeup information; if the face image does not include makeup information, input the face image into the makeup transfer model for makeup transfer processing to obtain a target made-up face image, and the target made-up face image includes makeup information; where the makeup transfer model is trained by the model training method according to any one of the embodiments of the first aspect.
[0020] To achieve the above object, a third aspect of the embodiments of the present application provides a model training device, the device includes:
[0021] A first acquisition module, configured to acquire a natural face sample image and a made-up face sample image, where the natural face sample image does not include makeup information, and the made-up face sample image includes makeup information;
[0022] A first generation module, configured to input the natural face sample image into an initial generation model for processing to obtain a first generation result, where the first generation result includes a first made-up face image converted from the natural face sample image and a first natural face image converted from the first made-up face image;
[0023] A second generation module, configured to input the made-up face sample image into the initial generation model for processing to obtain a second generation result, where the second generation result includes a second natural face image converted from the made-up face sample image and a second made-up face image converted from the second natural face image;
[0024] A calculation module, configured to calculate a target loss value according to the first generation result and the second generation result;
[0025] An adjustment module, configured to adjust the initial generation model by using the target loss value to obtain a makeup transfer model, where the makeup transfer model is used to perform makeup transfer processing on an image.
[0026] To achieve the above object, a fourth aspect of the embodiments of the present disclosure provides an electronic device, including at least one memory;
[0027] At least one processor;
[0028] At least one computer program;
[0029] The computer program is stored in the memory, and the processor executes the at least one computer program to implement:
[0030] The model training method according to any one of the embodiments of the first aspect; or
[0031] The image processing method according to the embodiment of the second aspect.
[0032] To achieve the above object, a fifth aspect of the embodiments of the present disclosure further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute:
[0033] The model training method according to any one of the embodiments of the first aspect; or
[0034] The image processing method according to the embodiment of the second aspect.
[0035] The model training method, device, image processing method, electronic device, and medium proposed in the embodiments of the present application input a plain face sample image into an initial generation model. Through the initial generation model, the plain face sample image can be converted into a first made-up face image and the first made-up face image can be converted into a first plain face image, obtaining a first generation result and realizing two-way transfer of makeup. Similarly, inputting a made-up face sample image into the initial generation model for two-way makeup transfer to obtain a second generation result. Then, using the first generation result and the second generation result to train the initial generation model to obtain a makeup transfer model. Therefore, it does not rely on an accurately paired training set. Instead, by exploring the potential feature relationship between the plain face image and the made-up face image, a model that can achieve two-way makeup transfer is constructed. The training process is more efficient and robust, which is beneficial to improving the flexibility of the actual use of the makeup transfer model and achieving a stable makeup transfer effect. Description of the Drawings
[0036] Figure 1 is a flowchart of the model training method provided by the embodiments of the present application;
[0037] Figure 2 is Figure 1 a specific flowchart of step S102 in
[0038] Figure 3 is a schematic diagram of an application process for generating a first plain face image according to a plain face sample image in the embodiments of the present application;
[0039] Figure 4 is Figure 1 a specific flowchart of step S104 in
[0040] Figure 5 is a flowchart of the image processing method provided by the embodiments of the present application;
[0041] Figure 6 is a block diagram of the model training device provided by the embodiments of the present application;
[0042] Figure 7 is a block diagram of the image processing device provided by the embodiments of the present application;
[0043] Figure 8 is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed Embodiments
[0044] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0045] It should be noted that although functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division from that in the device or a different order from that in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0047] First, some nouns involved in this application are parsed as follows:
[0048] Artificial intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.
[0049] Natural language processing (NLP): NLP uses computers to process, understand and apply human languages (such as Chinese, English, etc.). NLP belongs to a branch of artificial intelligence and is an interdisciplinary subject of computer science and linguistics, and is often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information retrieval, information extraction and filtering, text classification and clustering, public opinion analysis and opinion mining, etc. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing and linguistic research related to language computing, etc.
[0050] Face makeup transfer is a common technique in the field of computer vision. By extracting the makeup style from a reference image, it can be applied to the face images of different people. Existing makeup transfer methods usually require collecting paired plain face images and made-up face images as a training set, so as to use the training set to train a supervised model for makeup transfer. However, this method depends on an accurately paired training set. Therefore, the quality of the training set is likely to affect the stability of model training, resulting in poor makeup transfer effects.
[0051] Based on this, the embodiments of the present application provide a model training method and device, an image processing method, an electronic device, and a medium, which can achieve stable makeup transfer effects.
[0052] The embodiments of the present application provide a model training method and device, an image processing method, an electronic device, and a medium, which are specifically described through the following embodiments. First, the model training method and the image processing method in the embodiments of the present application are described.
[0053] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0054] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0055] The model training method and the image processing method provided by the embodiments of the present application relate to the field of artificial intelligence technology. The model training method or the image processing method provided by the embodiments of the present application can be applied to a terminal, or can be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, or a smart watch, etc.; the server can be an independent server, or can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms; the software can be an application that implements the model training method or the image processing method, etc., but is not limited to the above forms.
[0056] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0057] In a first aspect, please refer to Figure 1 , Figure 1 which is a flowchart of a model training method provided by an embodiment of this application for training a makeup transfer model. The model training method includes steps S101 to S105. It should be understood that the model training method of the embodiment of this application includes but is not limited to steps S101 to S105. The following will be introduced in detail with reference to Figure 1 .
[0058] Step S101: Obtain a plain face sample image and a made-up face sample image.
[0059] In the embodiment of this application, the plain face sample image does not include makeup information. Specifically, the plain face sample image can include a human face object without makeup. The made-up face sample image includes makeup information. Specifically, the made-up face sample image can include a human face object without makeup. It can be understood that the human face object in the plain face sample image and the human face object in the made-up face sample image can be the same or different, and no specific limitation is made.
[0060] Step S102: Input the plain face sample image into an initial generation model for processing to obtain a first generation result. The first generation result includes a first made-up face image converted from the plain face sample image and a first plain face image converted from the first made-up face image.
[0061] That is to say, in step S102, through the initial generation model, the natural face sample image can be first converted into the first made-up face image, and then the first made-up face image can be converted into the first natural face image to achieve two-way transfer of makeup. Among them, the initial generation model can be the makeup transfer model obtained from the previous training or the initially constructed generation model, without specific limitation. Specifically, the initial generation model can at least include a first transfer module and a second transfer module. The first transfer module is used to convert the natural face image into a made-up face image, and the second transfer module is used to convert the made-up face image into a natural face image. The first transfer module and the second transfer module form a cycle to realize the internal loop and connection between the made-up face image space and the natural face image space. In some optional implementation manners, the first transfer module and the second transfer module can adopt a denoising diffusion probabilistic model (DDPM), which does not require the supervision of paired images and can fully mine the feature information of natural face images and made-up face images, and further explore the potential relationship between the two different image spaces. In some other optional implementation manners, the first transfer module and the second transfer module can also adopt other unsupervised models such as a generative adversarial network model.
[0062] Step S103: Input the made-up face sample image into the initial generation model for processing to obtain a second generation result, where the second generation result includes a second natural face image converted from the made-up face sample image and a second made-up face image converted from the second natural face image.
[0063] Specifically, in step S103, the made-up face sample image can be first converted into a second natural face image by the second transfer module, and then the second natural face image can be converted into a second made-up face image by the first transfer module, which also realizes two-way transfer of makeup.
[0064] Step S104: Calculate a target loss value according to the first generation result and the second generation result.
[0065] In the embodiment of the present application, the target loss value is used to estimate the degree of inconsistency between the prediction result and the real result of the model. Therefore, the smaller the target loss value is, the better the training effect of the model is. Specifically, in step S104, a preset loss function can be obtained, and the target loss value can be obtained by substituting the first generation result and the second generation result into the loss function. Among them, the loss function can be adjusted according to actual needs. For example, the loss function can adopt any one of a 0-1 loss function, an absolute value loss function, a logarithmic loss function, an exponential loss function, and a Hinge loss function, without specific limitation.
[0066] In some optional implementations, the loss function may include at least one of a cycle consistency loss function, a spatial similarity loss function, and a makeup loss function. The cycle consistency loss function is used to estimate the accuracy of the model's description of the internal relationship between different image spaces, the spatial similarity loss function is used to estimate the accuracy of the description of the internal features of the same image space, and the makeup loss function is used to estimate the makeup transfer effect of a certain area in the image. It can be understood that the cycle consistency loss function and the spatial similarity loss function belong to global losses, while the makeup loss function belongs to regional losses.
[0067] Step S105: Use the target loss value to adjust the initial generation model to obtain a makeup transfer model.
[0068] Specifically, in step S105, the loss function of the initial generation model can be optimized according to the target loss value, the model loss of the loss function can be back-propagated, and the model parameters can be continuously adjusted until the training completion conditions are met, and the optimization of the initial generation model is stopped to obtain a makeup transfer model that meets the requirements.
[0069] It can be understood that the training steps of the makeup transfer model are implemented through steps S101 to S105, so there is no need to rely on an accurately matched training set. Instead, by mining the potential feature relationship between the bare face image and the makeup face image, a model that can achieve two-way makeup transfer is constructed. The training process is more efficient and robust, which is conducive to improving the flexibility of the actual use of the makeup transfer model and achieving a stable makeup transfer effect.
[0070] See also Figure 2 , Figure 2 yes Figure 1 In some embodiments, step S102 may include but is not limited to the following steps S201 to S203.
[0071] Step S201: Input the bare face sample image into the initial generation model.
[0072] Step S202: performing diffusion processing on the bare face sample image through the initial generation model to obtain a first original image satisfying the Gaussian distribution, and performing denoising processing on the first original image to obtain a first makeup face image.
[0073] Specifically, the diffusion processing of the bare face sample image may include T 1 step diffusion process, where T 1 represents the number of adjustable steps, and T 1 is a positive integer. In practical applications, we can first 1 The value is 1000, and then adjusted accordingly according to the model performance to ensure T 1The value of is within a reasonable range, which can avoid the problem of under-fitting in sampling during denoising, resulting in poor image quality, and the problem of over-fitting due to excessively detailed denoising, resulting in long image generation time and high computing power requirements. The i-th step diffusion process q 1 (x i |x i-1 )satisfy:
[0074] i∈i1,T 1 ], x 0 This is a sample picture without makeup, x i (i≥1) represents the image obtained through the i-th diffusion process, then is the first original graph mentioned above, which conforms to the Gaussian distribution with mean 0 and variance as the unit matrix I; β i is an adjustable constant, optional, β 1 To β T1 It can be a geometric progression determined from i0, 1]. Correspondingly, the denoising process for the first original image can also include T 1 The denoising process is a step-by-step denoising process. Both the denoising process and the diffusion process have Markov properties.
[0075] In some optional implementations, diffusion processing is performed on the bare face sample image through the initial generation model, which may include but is not limited to the following steps: obtaining T from the standard normal distribution interval through the initial generation model 1 parameter values, and according to T 1 The first noise vector is calculated based on the parameter value to realize the noise sampling based on the standard normal distribution. Based on this, the first noise vector and the first reference image are used to perform noise processing on the plain sample image to obtain the first original image, which is conducive to generating more diverse images.
[0076] The first reference image satisfies the standard normal distribution condition, that is, a Gaussian distribution with a mean of 0 and a variance of the unit matrix I, and the dimension of the first reference image is consistent with the dimension of the plain sample image. Optionally, T is obtained from the standard normal distribution interval i0, 1]. 1 parameter value β 1 to Then the first noise vector Can satisfy: t∈i1,T 1 ], the image obtained by the t-th step diffusion process can be ε is the first reference graph.
[0077] Step S203: performing diffusion processing on the first face image with makeup through the initial generation model to obtain a second original image satisfying Gaussian distribution, and performing denoising processing on the second original image to obtain a first face image without makeup.
[0078] For ease of understanding, please refer to Figure 3 , Figure 3 which is a schematic diagram of an application process for generating a first natural face image based on a natural face sample image in an embodiment of the present application. As Figure 3 shown, through an initial generation model, the natural face sample image x 0 is subjected to diffusion processing q 1 to obtain a first original image q 1 (x 0 ). The first original image q 1 (x 0 ) is subjected to denoising processing p 1 to obtain a first made-up face image q 1 p 1 (x 0 ). The first made-up face image q 1 p 1 (x 0 ) is subjected to diffusion processing q 2 to obtain a second original image q 1 p 1 q 2 (x 0 ). The second original image q 1 p 1 q 2 (x 0 ) is subjected to denoising processing p 2 to obtain a first natural face image q 1 p 1 q 2 p 2 (x 0 ), thus constituting a two-way diffusion-denoising processing process.
[0079] It can be understood that in step S203, the description of the diffusion processing of the first made-up face image can refer to the description of the diffusion processing of the natural face sample image in step S202, and the description of the denoising processing of the second original image can refer to the description of the denoising processing of the first original image in step S202, which will not be elaborated here.
[0080] Similarly, step S103 may specifically include but is not limited to the following steps: inputting a made-up face sample image into the initial generation model. Through the initial generation model, the made-up face sample image is subjected to diffusion processing to obtain a third original image that satisfies a Gaussian distribution, and the third original image is subjected to denoising processing to obtain a second natural face image. Through the initial generation model, the second natural face image is subjected to diffusion processing to obtain a fourth original image that satisfies a Gaussian distribution, and the fourth original image is subjected to denoising processing to obtain a second made-up face image.
[0081] In some optional implementations, diffusion processing is performed on the makeup face sample image through the initial generation model, which may include but is not limited to the following steps: obtaining T from the standard normal distribution interval through the initial generation model 2 parameter values, and according to T 2 The second noise vector is calculated by using the parameter values, where T 2 is a positive integer, T 2 Can be used with T 1 Same as T 1 The second noise vector and the second reference image are used to perform noise processing on the makeup face sample image to obtain a third original image. The second reference image satisfies the standard normal distribution condition, and the dimension of the second reference image is consistent with the dimension of the makeup face sample image. Optionally, T is obtained from the standard normal distribution interval i0, 1] 2 parameter value β 1 'to Then the second noise vector Can satisfy: t∈i1,T 2 ], the image obtained by the t-th step diffusion process can be Then ε′ is the second reference graph, x 0 ′ is a sample picture of a face with makeup applied.
[0082] For further information, see Figure 4 , Figure 4 yes Figure 1 In step S104 of some embodiments, if the target loss value includes a spatial similarity loss value, step S104 may specifically include but is not limited to the following steps S401 to S404.
[0083] Step S401: Acquire a first reference image and a second reference image.
[0084] Among them, both the first reference image and the second reference image meet the standard normal distribution conditions, the dimension of the first reference image is consistent with the dimension of the bare face sample image, and the dimension of the second reference image is consistent with the dimension of the makeup face sample image.
[0085] Step S402: Calculate a first spatial loss according to the first reference image and the first original image.
[0086] Step S403: Calculate a second spatial loss according to the second reference image and the second original image.
[0087] Step S404: performing a fusion operation according to the first spatial loss and the second spatial loss to obtain a spatial similarity loss value.
[0088] In the embodiments of the present application, the spatial similarity loss value is used to represent the distance between a noise vector that satisfies the standard normal distribution and the image data. The smaller the spatial similarity loss value, the closer or even overlapping the final distribution of the diffusion process is to the standard normal distribution. It can be seen that by focusing on the data features within the same image space, effective image information can be mined to promote the generation ability of the model.
[0089] In some optional implementation manners, the formula for calculating the first spatial loss can be:
[0090]
[0091] wherein, represents the first original image.
[0092] The formula for calculating the second spatial loss can be:
[0093]
[0094] wherein, represents the second original image.
[0095] Based on this, the spatial similarity loss value
[0096] In step S104 of some embodiments, if the target loss value includes the cycle consistency loss value, step S104 may further include but is not limited to the following steps: calculating a first transfer loss according to the plain face sample image and the first plain face image; calculating a second transfer loss according to the made-up face sample image and the second made-up face image; and performing a fusion operation on the first transfer loss and the second transfer loss to obtain the cycle consistency loss value.
[0097] Further, in some optional implementation manners, the formula for calculating the first transfer loss can be: p 2 q 2 p 1 q 1 (x 0 ) represents the first plain face image. The formula for calculating the second transfer loss can be: represents the second made-up face image. Based on this, the cycle consistency loss value It can be seen that by constraining the cycle consistency loss, it is ensured that the model can capture and describe the internal feature information of the image data, thereby improving the image generation ability based on two-way makeup transfer.
[0098] In step S104 of some embodiments, if the target loss value includes a makeup loss value, step S104 may further include but is not limited to the following steps:
[0099] First, perform mask annotation processing on the bare face sample image to obtain at least one feature region and the corresponding first mask information for the feature region from the bare face sample image. The feature region may include but is not limited to at least one of the following: cheek region, eye region, and mouth region. Perform mask annotation processing on the first made-up face image to obtain the corresponding second mask information for each feature region from the first made-up face image. Calculate the first makeup loss corresponding to each feature region according to the first mask information and the second mask information.
[0100] Optionally, in the case where there are multiple bare face sample images, for each bare face sample image: the initial mask information corresponding to each feature region in the bare face sample image can be obtained first. The initial mask information may include but is not limited to the mask dimension and the mask box position information. The mask dimension is used to represent the region size of the feature region. For example, the feature region is 10×10, and the mask box position information includes the positions of multiple positioning points on the mask box corresponding to the bare face sample image. By averaging the mask dimensions corresponding to all feature regions, the target mask dimension is obtained. According to the mask box position information, calculate the mask center position. For example, if the mask box position information includes (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), and (3,3), then the coordinate values of different dimensions are averaged respectively to obtain the mask center position, that is, (2,2). Determine the target mask information corresponding to the corresponding feature region in the bare face sample image based on the target mask dimension and the mask center position. Based on this, sum up the target mask information corresponding to each feature region in different bare face sample images to obtain the first mask information corresponding to the feature region. Similarly, the above calculation method for obtaining the first mask information is also applicable to the second mask information, the following third mask information, and the fourth mask information, which will not be elaborated.
[0101] After that, perform mask annotation processing on the made-up face sample image to obtain the corresponding third mask information for each feature region from the made-up face sample image. Perform mask annotation processing on the second bare face image to obtain the corresponding fourth mask information for each feature region from the second bare face image. Calculate the second makeup loss corresponding to each feature region according to the third mask information and the fourth mask information.
[0102] Finally, calculate the target makeup loss corresponding to each feature region according to the first makeup loss and the second makeup loss, and then perform a fusion operation on the target makeup losses corresponding to all feature regions to obtain the makeup loss value. Optionally, makeup loss value = first makeup loss + second makeup loss.
[0103] For example, if three feature regions are pre-determined, namely the cheek region, the eye region, and the mouth region, then the makeup loss value is the first makeup loss corresponding to the eye region, is the second makeup loss corresponding to the eye region, is the first makeup loss corresponding to the cheek region, is the second makeup loss corresponding to the cheek region, is the first makeup loss corresponding to the mouth region, is the second makeup loss corresponding to the mouth region.
[0104] It can be seen that by extracting the feature of different regions of the face image, it can be accurate to specific feature regions, so that the makeup transfer model meets the makeup loss condition in the local image region, thus improving the makeup transfer ability of the model locally.
[0105] In some implementation manners, if the target loss value includes at least two of the cycle consistency loss value, the spatial similarity loss value, and the makeup loss value, then in step S105, all the loss values included in the target loss value can also be fused and calculated to obtain a fusion result, so as to adjust the initial generation model by using the fusion result to obtain the makeup transfer model. Exemplarily, if the target loss value includes the cycle consistency loss value, the spatial similarity loss value, and the makeup loss value, then the fusion result Loss = Loss cycle + Loss sim + Loss make-up . At this time, through the constraint of the loss function in multiple directions, the generation direction of the initial generation model is broken, so that the model can achieve the makeup style transfer in the constraint direction. In addition, even if the makeup transfer model does not depend on the paired images, it can also achieve the paired comparison in the supervised task, and then more accurately and efficiently achieve the makeup transfer task.
[0106] In step S105 of some embodiments, step S105 may specifically include but is not limited to the following steps:
[0107] On the one hand, if the initial generation model does not meet the training completion condition, then the initial generation model is adjusted by using the target loss value to obtain the adjusted initial generation model, and step S101 is continued. In practical applications, the parameters of the initial generation model can be adjusted by using the stochastic gradient descent (SGD) method.
[0108] On the other hand, if the initial generation model meets the training completion condition, then the makeup transfer model is determined according to the initial generation model.
[0109] Among them, the training completion conditions include at least any one of the following: the target loss value is less than the preset loss value; the loss value of the initial generation model increases by a preset number of times, or the loss value of the initial generation model increases continuously by a preset number of times. The preset loss value and the preset number of times can be determined and adjusted according to human experience. For example, the preset number of times can be set to 10 times, without specific limitation. It can be seen that through the above steps, the overfitting problem of model training can be avoided, thereby ensuring that the model achieves better generalization performance.
[0110] In a second aspect, please refer to Figure 5 , Figure 5 which is a flowchart of the image processing method provided by the embodiments of the present application. The image processing method includes but is not limited to step S501 and step S503. The following will be introduced in detail with reference to Figure 5 for a detailed introduction.
[0111] Step S501: Obtain a face image.
[0112] Step S502: If the face image includes makeup information, input the face image into the makeup transfer model for makeup transfer processing to obtain a target natural face image, which does not include makeup information.
[0113] Step S503: If the face image does not include makeup information, input the face image into the makeup transfer model for makeup transfer processing to obtain a target made-up face image, which includes makeup information.
[0114] In the embodiments of the present application, the makeup transfer model is trained according to the model training method shown in the method embodiments of the first aspect above.
[0115] It can be seen that by implementing the above steps S501 to S503, only one makeup transfer model is required, which can not only achieve the makeup effect on natural face images but also achieve the makeup removal effect on made-up face images. Therefore, the flexibility of using the makeup transfer model is improved, and a stable makeup transfer effect is achieved.
[0116] Please refer to Figure 6 , Figure 6 which is a block diagram of the model training device provided by the embodiments of the present application. In some embodiments, the model training device includes a first acquisition module 601, a first generation module 602, a second generation module 603, a calculation module 604, and an adjustment module 605.
[0117] The first acquisition module 601 is used to acquire a natural face sample image and a made-up face sample image. The natural face sample image does not include makeup information, and the made-up face sample image includes makeup information.
[0118] The first generation module 602 is configured to input a natural face sample image into an initial generation model for processing to obtain a first generation result, where the first generation result includes a first made-up face image converted from the natural face sample image and a first natural face image converted from the first made-up face image.
[0119] The second generation module 603 is configured to input a made-up face sample image into the initial generation model for processing to obtain a second generation result, where the second generation result includes a second natural face image converted from the made-up face sample image and a second made-up face image converted from the second natural face image.
[0120] The calculation module 604 is configured to calculate a target loss value according to the first generation result and the second generation result.
[0121] The adjustment module 605 is configured to adjust the initial generation model by using the target loss value to obtain a makeup transfer model, where the makeup transfer model is used to perform makeup transfer processing on an image.
[0122] The model training device of the embodiment of the present application does not need to rely on an accurately paired training set. Instead, by mining the potential feature relationship between a natural face image and a made-up face image, a model that can achieve bidirectional makeup transfer is constructed. The training process is more efficient and robust, which is beneficial to improving the flexibility of actually using the makeup transfer model and achieving a stable makeup transfer effect.
[0123] It should be noted that the model training device of the embodiment of the present application corresponds to the foregoing model training method. For the specific training process, please refer to the foregoing model training method, which will not be elaborated here one by one.
[0124] Please refer to Figure 7 , Figure 7 is a block diagram of modules of an image processing device provided by an embodiment of the present application. In some embodiments, the image processing device includes a second acquisition module 701 and a processing module 702.
[0125] The second acquisition module 701 is configured to acquire a face image;
[0126] The processing module 702 is configured to, when the face image includes makeup information, input the face image into the makeup transfer model for makeup transfer processing to obtain a target natural face image that does not include makeup information; and when the face image does not include makeup information, input the face image into the makeup transfer model for makeup transfer processing to obtain a target made-up face image that includes makeup information; where the makeup transfer model is trained according to the model training method shown in the method embodiment of the first aspect above.
[0127] It should be noted that the image processing device in the embodiments of the present application corresponds to the aforementioned image processing method. For specific image processing steps, please refer to the aforementioned image processing method and will not be elaborated here one by one.
[0128] The embodiments of the present application further provide an electronic device, including:
[0129] At least one memory;
[0130] At least one processor;
[0131] At least one program;
[0132] The program is stored in the memory, and the processor executes at least one program to implement the model training method or the image processing method described above in the present disclosure. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0133] The following Figure 8 introduces the electronic device in the embodiments of the present application in detail.
[0134] As Figure 8 , Figure 8 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0135] A processor 801, which can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0136] A memory 802, which can be implemented in the form of a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute the model training method or the image processing method in the embodiments of the present application;
[0137] An input / output interface 803, which is used to implement information input and output;
[0138] A communication interface 804 for implementing communication interactions between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0139] A bus 805 for transmitting information between various components of the device (such as a processor 801, a memory 802, an input / output interface 803, and a communication interface 804);
[0140] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 achieve communication connections with each other inside the device through the bus 805.
[0141] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the above-mentioned model training method or image processing method.
[0142] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0143] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0144] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine some steps, or different steps.
[0145] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0147] As used in the description of the present application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0148] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0149] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0150] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] In addition, each functional unit in various embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store programs.
[0153] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of this application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. A model training method for training a makeup transfer model, characterized in that, the method includes: Obtain a plain face sample image and a made-up face sample image. The plain face sample image does not include makeup information, and the made-up face sample image includes makeup information; Input the plain face sample image into an initial generation model for processing to obtain a first generation result. The first generation result includes a first made-up face image converted from the plain face sample image and a first plain face image converted from the first made-up face image; Input the made-up face sample image into the initial generation model for processing to obtain a second generation result. The second generation result includes a second plain face image converted from the made-up face sample image and a second made-up face image converted from the second plain face image; Calculate a target loss value according to the first generation result and the second generation result, and use the target loss value to adjust the initial generation model to obtain a makeup transfer model, which is used for performing makeup transfer processing on images; wherein, the target loss value includes a makeup loss value; calculating the target loss value according to the first generation result and the second generation result includes: Perform mask annotation processing on the plain face sample image to obtain at least one feature region and first mask information corresponding to the feature region from the plain face sample image; Perform mask annotation processing on the first made-up face image to obtain second mask information corresponding to each of the feature regions from the first made-up face image; Calculate a first makeup loss corresponding to each of the feature regions according to the first mask information and the second mask information; Perform mask annotation processing on the made-up face sample image to obtain third mask information corresponding to each of the feature regions from the made-up face sample image; Perform mask annotation processing on the second plain face image to obtain fourth mask information corresponding to each of the feature regions from the second plain face image; Calculate a second makeup loss corresponding to each of the feature regions according to the third mask information and the fourth mask information; Calculate a target makeup loss corresponding to each of the feature regions according to the first makeup loss and the second makeup loss; Perform a fusion operation on the target makeup losses corresponding to all the feature regions to obtain a makeup loss value.
2. The method according to claim 1, characterized in that, the inputting the plain face sample image into an initial generation model for processing to obtain a first generation result includes: Input the plain face sample image into the initial generation model; Through the initial generation model, perform diffusion processing on the plain face sample image to obtain a first original image that satisfies a Gaussian distribution, and perform denoising processing on the first original image to obtain a first made-up face image; Through the initial generation model, perform diffusion processing on the first made-up face image to obtain a second original image that satisfies a Gaussian distribution, and perform denoising processing on the second original image to obtain a first plain face image; the inputting the made-up face sample image into the initial generation model for processing to obtain a second generation result includes: Inputting the makeup face sample image into the initial generation model; The makeup face sample image is diffused by the initial generation model to obtain a third original image satisfying Gaussian distribution, and the third original image is denoised to obtain a second bare face image; The second bare-face image is diffused by the initial generation model to obtain a fourth original image satisfying Gaussian distribution, and the fourth original image is denoised to obtain a second makeup-applied face image.
3. The method according to claim 2, It is characterized in that The target loss value includes a spatial similarity loss value, and the calculating the target loss value according to the first generation result and the second generation result includes: Obtaining a first reference image and a second reference image, wherein both the first reference image and the second reference image satisfy a standard normal distribution condition, the dimension of the first reference image is consistent with the dimension of the bare face sample image, and the dimension of the second reference image is consistent with the dimension of the makeup face sample image; Calculating a first spatial loss according to the first reference image and the first original image; Calculating a second spatial loss according to the second reference image and the second original image; A fusion operation is performed according to the first spatial loss and the second spatial loss to obtain a spatial similarity loss value.
4. The method according to claim 1, It is characterized in that The target loss value includes a cycle consistency loss value; The calculating the target loss value according to the first generation result and the second generation result includes: Calculating a first transfer loss according to the plain face sample image and the first plain face image; Calculating a second transfer loss according to the makeup face sample image and the second makeup face image; A fusion operation is performed according to the first transfer loss and the second transfer loss to obtain a cycle consistency loss value.
5. The method according to any one of claims 1 to 4, It is characterized in that The step of adjusting the initial generation model by using the target loss value to obtain a makeup transfer model includes: If the initial generation model does not meet the training completion condition, the initial generation model is adjusted using the target loss value to obtain an adjusted initial generation model, and the step of obtaining the bare face sample image and the makeup face sample image is continued; If the initial generation model satisfies the training completion condition, determining a makeup transfer model according to the initial generation model; Among them, the training completion condition package includes at least any one of the following: the target loss value is less than the preset loss value; the loss value of the initial generation model rises a preset number of times.
6. An image processing method, It is characterized in that The method comprises: Get face image; If the face image includes makeup information, input the face image into a makeup transfer model to perform makeup transfer processing to obtain a target bare face image, wherein the target bare face image does not include makeup information; If the face image does not include makeup information, input the face image into the makeup transfer model to perform makeup transfer processing to obtain a target makeup face image, wherein the target makeup face image includes makeup information; Among them, the makeup transfer model is trained according to the method described in any one of claims 1 to 5.
7. A model training device Characterized in that The device includes: A first acquisition module, configured to acquire a natural face sample image and a made-up face sample image, where the natural face sample image does not include makeup information, and the made-up face sample image includes makeup information; A first generation module, configured to input the natural face sample image into an initial generation model for processing to obtain a first generation result, where the first generation result includes a first made-up face image converted from the natural face sample image and a first natural face image converted from the first made-up face image; A second generation module, configured to input the made-up face sample image into the initial generation model for processing to obtain a second generation result, where the second generation result includes a second natural face image converted from the made-up face sample image and a second made-up face image converted from the second natural face image; A calculation module, configured to calculate a target loss value according to the first generation result and the second generation result; An adjustment module, configured to adjust the initial generation model by using the target loss value to obtain a makeup transfer model, where the makeup transfer model is used to perform makeup transfer processing on an image; Among them, the target loss value includes a makeup loss value; the calculation module, configured to calculate a target loss value according to the first generation result and the second generation result, includes: Performing mask annotation processing on the natural face sample image to obtain at least one feature region and first mask information corresponding to the feature region from the natural face sample image; Performing mask annotation processing on the first made-up face image to obtain second mask information corresponding to each of the feature regions from the first made-up face image; Calculating a first makeup loss corresponding to each of the feature regions according to the first mask information and the second mask information; Performing mask annotation processing on the made-up face sample image to obtain third mask information corresponding to each of the feature regions from the made-up face sample image; Performing mask annotation processing on the second natural face image to obtain fourth mask information corresponding to each of the feature regions from the second natural face image; Calculating a second makeup loss corresponding to each of the feature regions according to the third mask information and the fourth mask information; Calculating a target makeup loss corresponding to each of the feature regions according to the first makeup loss and the second makeup loss; Performing a fusion operation on the target makeup losses corresponding to all the feature regions to obtain a makeup loss value.
8. An electronic device Characterized in that It includes: At least one memory; At least one processor; At least one computer program; The computer program is stored in the memory, and the processor executes the at least one computer program to implement: The model training method described in any one of claims 1 to 5; or The image processing method described in claim 6.
9. A computer-readable storage medium Characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute: The model training method according to any one of claims 1 to 5; or The image processing method according to claim 6.
Citation Information
Patent Citations
Real-time skin makeup migration method and device, electronic equipment and readable storage medium
CN111815534A
Image processing method and device, storage medium and electronic equipment
CN113822793A