Style modulation controllable makeup migration method based on dense deformation alignment and related equipment

Through the combination of dense deformation alignment module and style modulation module, the problem of posture misalignment and semantic region mismatch in makeup transfer is solved, the fine control of makeup style and the controllability of local makeup is achieved, and the alignment accuracy and fidelity of makeup migration is improved.

CN120510653APending Publication Date: 2025-08-19SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500447.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

When existing makeup transfer technologies deal with posture misalignment and semantic region mismatch, it is difficult to achieve the transfer of fine makeup styles, and lack control over local makeup and fine control of makeup styles.

Method used

The dense deformation alignment module is used to construct semantic alignment using KL divergence and unbalanced optimal transmission theory, and fine style control is carried out in combination with the style modulation module, and the makeup style and identity characteristics are integrated through the makeup fusion module.

Benefits of technology

It significantly improves the alignment accuracy and fidelity of makeup styles, achieves fine control of local makeup and natural integration of makeup styles, and maintains the consistency of facial identity characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510653A_ABST
    Figure CN120510653A_ABST
Patent Text Reader

Abstract

The invention discloses a style modulation controllable makeup migration method based on dense deformation alignment and related equipment. The method comprises the following steps: acquiring a to-be-processed image; inputting the to-be-processed image into a pre-trained makeup migration model, and outputting an image with a corresponding makeup style; according to the makeup migration model, a dense deformation alignment module adopts KL divergence as a constraint, a feature alignment problem is solved by using an unbalanced optimal transmission theory, and a makeup migration region is obtained; the style modulation module combines makeup migration areas from the dense deformation alignment module to realize fine regulation and control of a normalization process; and the makeup fusion module is used for fusing the features obtained by the dense deformation alignment module and the style modulation module to generate a target image. The dense deformation alignment module explicitly constructs the dense semantic correspondence between the reference image and the source image, and can significantly improve the alignment precision and style fidelity of makeup migration under the conditions of inconsistent processing attitudes and significant structural difference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer image processing, and in particular to a style-modulated controllable makeup migration method based on dense deformation alignment and related equipment. Background Art

[0002] Makeup transfer technology aims to transfer the makeup style from a reference image to a target image while preserving the target image's identity. With the advancement of virtual makeup technology, makeup transfer has become an important research direction in many application areas, particularly in beauty, entertainment, social media, and augmented reality. Users desire to enhance their appearance through virtual makeup, and even customize their makeup experience for different application scenarios or styles. However, achieving this goal faces several technical challenges, particularly in handling pose misalignment and finely controlling makeup style transfer.

[0003] At present, the mainstream methods of makeup transfer technology include those based on generative adversarial networks (GANs), which usually achieve makeup style transfer by constructing a cyclic training process. Specifically, one network transfers the makeup style in the source image to the target image, and the other network is responsible for removing the transferred makeup style. Although GAN can achieve makeup style transfer, it has some limitations. First, these methods usually treat different semantic areas of the face as equal. This processing method makes it impossible to finely control facial features when transferring makeup. For example, the eyes, mouth and other areas of the face have different expressions when applying makeup, so more detailed adjustments are required during the transfer process. Secondly, since the poses of the reference image and the source image may be inconsistent, how to perform effective transfer in this case becomes an important issue.

[0004] While some approaches have attempted to address these issues, such as semantic matching by aligning features using cosine similarity, this approach is prone to many-to-one matching issues when semantic regions do not match, affecting the accuracy of the transfer effect. Therefore, effectively aligning the semantics between the reference and target images in the presence of pose misalignment and achieving fine control during makeup transfer remain key challenges in current research. Summary of the Invention

[0005] In order to at least partially solve one of the technical problems existing in the prior art, the present invention aims to provide a style-modulated controllable makeup migration method based on dense deformation alignment and related equipment.

[0006] The first technical solution adopted by the present invention is:

[0007] A style-modulated controllable makeup transfer method based on dense deformation alignment includes the following steps:

[0008] Get the image to be processed;

[0009] Input the image to be processed into the pre-trained makeup transfer model and output an image with the corresponding makeup style;

[0010] The makeup transfer model includes a dense deformation alignment module, a style modulation module and a makeup fusion module;

[0011] The dense deformation alignment module adopts KL divergence as a constraint and uses the unbalanced optimal transfer theory to solve the feature alignment problem and obtain the makeup migration area;

[0012] The style modulation module dynamically combines the rough makeup transfer regions from the dense deformation alignment module to enhance the sharpness and details of the transferred style and achieve fine control of the normalization process;

[0013] The makeup fusion module is used to fuse the features obtained by the dense deformation alignment module and the style modulation module to generate a target image.

[0014] Furthermore, the input of the dense deformation alignment module is the reference image y r and semantic segmentation map of the source image

[0015] The dense deformation alignment module works as follows:

[0016] Two feature extractors are used to extract the source image x s and the reference image y r Extract local feature vectors and construct corresponding feature sets f x ={s1,…,s n} and f y ={r1,…,r n}, where n represents the number of eigenvectors; introducing the quality variable γ i and δ j Respectively represent the assignment to s i and r j The mass of the two sets is and

[0017] Let the distance matrix C ij Indicates that the mass γ i From the feature i Transfer to feature r j Corresponding mass δ j the price;

[0018] KL divergence is used as a soft constraint term, and the entropy regularization term is introduced:

[0019]

[0020] Where P represents the transmission matrix, P ij Indicates that in γ i and δ j The quality of transmission between

[0021] Use the transfer matrix P to warp the reference image y r , obtain the intermediate representation W that combines makeup style and structure yr :

[0022] W yr =y r ·P

[0023] Reference semantic graph Perform the same transformation to extract W yr The specific deformation area is denoted as (where k = {lip, skin, eyes}):

[0024]

[0025] Extract filtered makeup migration area

[0026]

[0027] Where ⊙ is the Hadamard product;

[0028] The deformed semantic map not only preserves the structural information of the source identity, but also effectively captures the fine-grained makeup style in the reference image.

[0029] Furthermore, the final expression of the optimal transmission objective function is as follows:

[0030]

[0031] Where τ is the regularization parameter of the KL divergence term, η is the regularization coefficient that adjusts the smoothness and dispersion of the transmission plan; T is the transpose.

[0032] Furthermore, the expression of the optimal transmission objective function is transformed into the Fenchel-Legendre dual form:

[0033]

[0034] The Sinkhorn algorithm is used to solve the problem and obtain the optimal transmission plan P. P is encoded by the dual vectors u and v and is expressed as follows:

[0035]

[0036] Furthermore, the quality term is dynamically calculated using the following formula:

[0037]

[0038] Distance matrix C ij Defined as cosine distance, the expression is as follows:

[0039]

[0040] Where T is the transpose.

[0041] Furthermore, the style modulation module works as follows:

[0042] From the deformed semantic graph Extract the style information of the corresponding area and concatenate them to form the style matrix ST;

[0043] Migrate makeup to the area And the style matrix ST inputs independent convolution channels, extracting two sets of normalized modulation parameters: At each normalized layer i, let its input activation be Among them, B and C i 、H i 、W i They represent the batch size, number of channels, height, and width of the feature map of the i-th layer respectively; the mean and standard deviation in the channel direction are and The normalized output is as follows:

[0044]

[0045] Where, and It is a spatially sensitive learnable weight that is adaptively learned based on the style map and deformation results.

[0046] Furthermore, the makeup fusion module consists of an identity encoder, two upsampling layers, five fusion blocks, and a decoder;

[0047] The makeup fusion module works as follows:

[0048] First, use the identity encoder to get the source image x s Multi-scale identity features are extracted from the fusion block; these features are then fed into the fusion block for processing;

[0049] Each fusion block receives the affine transformation parameters generated by the dense deformation alignment module and uses them to modulate the normalization process to achieve an adaptive combination of style and structure;

[0050] To gradually restore the image resolution, two upsampling layers are inserted between the fusion blocks to expand the spatial size of the feature map; finally, the decoder receives the output of the last fusion block and generates the target image While maintaining the identity characteristics of the source image, the image effectively integrates the makeup style of the reference image, achieving a natural and detailed makeup transfer effect.

[0051] Furthermore, the loss function during the training of the makeup transfer model is as follows:

[0052]

[0053] Where λ1,λ2,λ3,λ4,λ5 are the weight hyperparameters of each loss term;

[0054] Among them, the perceptual loss The expression is:

[0055]

[0056] Where, φ l (·) represents the features after the ReLU activation layer, is the generated target image, |·|2 represents the L2 norm;

[0057] Makeup loss The expression is:

[0058]

[0059] Where, PGT(x s ,y r ) and PGT(y r ,x s ) represent the pseudo labels of makeup transfer and reverse transfer respectively; f(x s ,y r ) represents the source image x s and the reference image y r The makeup migration image generated by input, f(y r ,x s ) represents the reference image y r and the source image x s The reverse makeup transfer image generated as input;

[0060] Cycle consistency loss The expression is:

[0061]

[0062] Where |·|1 represents the L1 norm;

[0063] identity loss The expression is:

[0064]

[0065] The losses of the discriminator D and the generator G in the adversarial loss are expressed as follows:

[0066]

[0067] Where E represents the mathematical expectation, and h() represents the function used to regularize the discriminator.

[0068] The second technical solution adopted by the present invention is:

[0069] An electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the above-mentioned style-modulated controllable makeup transfer method based on dense deformation alignment.

[0070] The third technical solution adopted by the present invention is:

[0071] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-mentioned style-modulated controllable makeup transfer method based on dense deformation alignment.

[0072] The fourth technical solution adopted by the present invention is:

[0073] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, so that the computer device performs the above method.

[0074] The beneficial effects of the present invention include:

[0075] (1) The dense deformation alignment module proposed in this paper is based on the imbalanced optimal transfer theory and explicitly constructs a dense semantic correspondence between the reference image and the source image. It can significantly improve the alignment accuracy and style fidelity of makeup transfer when dealing with inconsistent postures and significant structural differences.

[0076] (2) The style modulation module of the present invention achieves fine control of makeup style in semantic areas and spatial consistency fusion by introducing a region-aware style normalization mechanism.

[0077] (3) The makeup fusion module of the present invention integrates identity feature encoding, style modulation and multi-level decoding strategy, aiming to achieve high-quality makeup migration while maintaining identity consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.

[0079] Figure 1 is a schematic diagram of the results of the dense deformation alignment module in an embodiment of the present invention;

[0080] Figure 2 is a schematic diagram of the results of the style modulation module in an embodiment of the present invention;

[0081] Figure 3 2 is a schematic diagram of the result of the makeup fusion module in an embodiment of the present invention. DETAILED DESCRIPTION

[0082] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application. For the step numbers in the following embodiments, they are provided only for the convenience of explanation and are not intended to limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0083] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as setting, installing, and connecting should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0084] In the description of this application, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on this application.

[0085] In the description of this application, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The terms "first" and "second" are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or as implicitly specifying the number or order of the technical features indicated.

[0086] In the description of this application, "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0087] After research and analysis by the inventors, the existing technical solutions have the following technical problems:

[0088] (1) Posture misalignment and region mismatch: Since there are often significant posture differences between the source image and the reference image (such as head rotation, expression changes, etc.), the positions of the corresponding semantic regions (such as eye shadow, lipstick, blush, etc.) between the two are inconsistent. Traditional methods often rely on matching mechanisms based on feature similarity, which is prone to semantic mismatch or many-to-one matching, and cannot achieve accurate alignment and migration.

[0089] (2) Controllability of local and degree control: Existing methods often treat different semantic areas of the face (such as eyes, lips, and facial contours) equally, ignoring the differences in makeup expression in different areas, making it difficult to achieve independent control of local makeup (such as changing only eye makeup or adjusting lipstick color). In addition, users' expectations for makeup transfer are usually adjustable, such as applying only a certain part of the makeup effect or adjusting the makeup intensity. However, existing methods lack a unified control mechanism, making it difficult to flexibly adjust the scope and degree of transfer.

[0090] (3) Compensation for feature loss during alignment: In the process of deforming the reference image to achieve semantic alignment, the makeup style information is easily lost, resulting in incomplete or distorted makeup effects and a lack of realism and details.

[0091] Based on this, the present invention provides a style modulation controllable makeup transfer scheme based on dense deformation alignment. In the makeup transfer task, X and Y represent the domains of the image without makeup and with makeup, respectively. Given the original image x s ∈X and the reference image y r ∈Y, the goal is to learn a mapping function in With y r Same makeup style, while retaining x s In addition, since makeup removal is a special case of makeup transfer, a mapping function needs to be learned in With x s Same makeup style while keeping y r facial identity features.

[0092] Example 1

[0093] This embodiment provides a style-modulated controllable makeup transfer method based on dense deformation alignment, including the following steps:

[0094] S1, obtaining the image to be processed;

[0095] S2. Input the image to be processed into a pre-trained makeup transfer model, and output an image with the corresponding makeup style; wherein the makeup transfer model includes a dense deformation alignment module, a style modulation module, and a makeup fusion module;

[0096] The dense deformation alignment module adopts KL divergence as a constraint and uses the unbalanced optimal transfer theory to solve the feature alignment problem and obtain the makeup migration area;

[0097] The style modulation module dynamically combines the rough makeup transfer regions from the dense deformation alignment module to enhance the sharpness and details of the transferred style and achieve fine control of the normalization process;

[0098] The makeup fusion module is used to fuse the features obtained by the dense deformation alignment module and the style modulation module to generate a target image.

[0099] Each module is explained in detail below with reference to the accompanying drawings and specific embodiments.

[0100] (1) Dense deformation alignment module

[0101] In order to retain the fine makeup style with spatial contextual semantic information and effectively deal with the problem of feature misalignment caused by inconsistent posture, this embodiment proposes an image deformation method that explicitly constructs dense correspondences. This method uses unbalanced optimal transfer theory to solve the feature alignment problem and effectively alleviates the makeup transfer distortion caused by the differences between the reference image and the source image in terms of posture, semantic area, etc. Optimal transfer aims to determine a transfer plan that transfers samples from one distribution to another with minimal cost. However, when the source image and the reference image have different poses (such as front and side images), the total mass of the two distributions is usually not equal, which often affects the effect of makeup transfer. Therefore, this module uses unbalanced optimal transfer with divergence metric constraints to solve this problem.

[0102] See also Figure 1 , the module starts from the reference image y r and semantic segmentation map of the source image As input. For example, the segmentation map The existing face parsing network can be used to extract the source image x s Get. Use two feature extractors from x s and y r Extract local feature vectors and construct corresponding feature sets f x ={s1,…,s n} and f y ={r1,…,r n}, where n represents the number of eigenvectors. Introducing the quality variable γ i and δ j Respectively represent the assignment to s i and r j The total mass of the two sets is and

[0103] As an optional implementation, in order to adaptively assign feature weights based on the similarity between feature sets, the quality term is dynamically calculated using the following formula:

[0104]

[0105] Distance matrix C ij Indicates that the mass γ i From the feature i Transfer to feature r j Corresponding mass δ j The cost is defined as the cosine distance:

[0106]

[0107] To address the issue of unequal quality of source and reference feature distributions, this embodiment introduces the Kullback-Leibler (KL) divergence as a soft constraint:

[0108]

[0109] In addition, since the optimal transport problem has linear objectives and constraints, it is not differentiable everywhere. In order to make it strictly convex and differentiable, and to improve the numerical stability of the optimal transport problem, we introduce the entropy regularization term:

[0110]

[0111] Where P represents the transmission matrix, P ij Indicates that in γ i and δ j The quality of transmission between.

[0112] The final optimal transmission objective function can be expressed as:

[0113]

[0114] Where τ is the regularization parameter of the KL divergence term, and η is the regularization coefficient that adjusts the smoothness and dispersion of the transmission plan.

[0115] In order to solve the problem efficiently, the above formula is transformed into the Fenchel-Legendre dual form:

[0116]

[0117] in:

[0118]

[0119] The Sinkhorn algorithm is used to solve the problem in the above formula to obtain the optimal transmission plan P. P is encoded by the dual vectors u and v and can be expressed as follows:

[0120]

[0121] In order to achieve controllable makeup migration, representative local areas of the face (such as lips, eyes and skin) are further extracted to generate intermediate results. Specifically, the reference image y is deformed using the transfer matrix P r , obtain the intermediate representation W that combines makeup style and structure yr :

[0122] W yr =y r ·P

[0123] At the same time, the reference semantic graph Perform the same transformation to extract W yr The specific deformation area is denoted as (where k = {lip, skin, eyes}):

[0124]

[0125] Combine the above results to extract the filtered makeup migration area

[0126]

[0127] The deformed semantic map not only preserves the structural information of the source identity, but also effectively captures the fine-grained makeup style in the reference image.

[0128] The dense deformation alignment module based on unbalanced transfer can handle global correspondences and is particularly suitable for processing large-scale, non-rigid makeup areas (such as blush). The deformation results of unbalanced optimal transfer not only help transfer images smoother, but also make the deformation results contain more fine-grained makeup styles, with greater detail fidelity and local style consistency.

[0129] (2) Style modulation module

[0130] To further enhance the sharpness and detail of the transferred makeup, this example proposes a style modulation module. This module leverages the locality and structural independence of makeup styles within semantic regions, dynamically combining the rough makeup regions from the dense deformation alignment module to enhance the sharpness and detail of the transferred style and achieve fine-grained control of the normalization process.

[0131] See also Figure 2 In this module, the style is considered as a high-dimensional embedding vector independent of the shape, which is used to adjust the affine transformation parameters in the normalization layer. First, the deformed image W output from the dense deformation alignment module is dynamically fused. yr , and regional style encoding extracted from the reference image to preserve the spatial context information while improving the detail quality of the generated results. Specifically, we do not directly r Instead of re-extracting features from the CNN, the reference features from the alignment module are reused, which maintains the consistency and continuity of style representation and avoids the potential inconsistency introduced by external features.

[0132] The method uses regional average pooling operation to deform the map from the semantic mask Extract the style information of the corresponding area. The three main facial areas (lips, skin and eyes) are spliced together to form a style matrix Where s=3 represents the number of region categories.

[0133] By deforming the semantic graph As a guide, the style matrix is broadcast to the corresponding semantic region after deformation to generate the corresponding target style map. This style map can be used independently for normalization control or combined with the deformation result from the dense deformation alignment module. Further dynamic fusion is performed to achieve style modulation with spatial context consistency.

[0134] To this end, the The convolution channel is independent of the ST input, and two sets of normalized modulation parameters are extracted. At each normalized layer i, let its input activation be Among them, B and C i 、H i 、W i Represent the batch size, number of channels, height and width of the feature map of the i-th layer respectively. The mean and standard deviation in the channel direction are and The normalized output is as follows:

[0135]

[0136] in, and The spatial sensitivity learnable weights are adaptively learned based on the style map and deformation results. They are derived from the features of the convolution output and the deformation results after combination. and the weighted sum of the style matrix ST, which is defined as follows:

[0137]

[0138] Among them, θ α and are dynamically learned regional adaptation coefficients, which are used to weightedly fuse the features from the two sources.

[0139] By introducing style maps and semantic masks as spatial priors, this module introduces a region-level spatial alignment mechanism in the normalization stage, achieving position-dependent adjustment of activation values, thereby improving the controllability and precision of makeup style transfer.

[0140] (3) Makeup fusion module

[0141] The makeup fusion module consists of an identity encoder, two upsampling layers, five fusion blocks, and a decoder. Unlike generative methods that use random noise as input, our method uses the source image as a starting point and aims to transfer the reference makeup style while preserving the source identity.

[0142] First, a face identity encoder is used to generate the face from the source image x sMulti-scale identity features are extracted from the fusion blocks. These features are then fed into a series of fusion blocks for processing. Each fusion block receives the affine transformation parameters generated by the dense deformation alignment module and is used to modulate the normalization process to achieve an adaptive combination of style and structure. To gradually restore the image resolution, two upsampling layers are inserted between the fusion blocks to expand the spatial size of the feature map. Finally, the decoder receives the output of the last fusion block and generates the target image. While maintaining the identity characteristics of the source image, the image effectively integrates the makeup style of the reference image, achieving a natural and detailed makeup transfer effect.

[0143] (4) Optimization goal

[0144] Two types of loss functions are designed for the imbalanced optimal transfer plan in the dense deformation alignment module: one is the domain alignment loss, which is used to ensure that the feature embeddings generated by two independent feature extractors are in a consistent representation space; the other is the cycle consistency loss, which restores the deformed image to its original domain by applying the reverse transfer plan and encourages the image content before and after the transformation to remain consistent.

[0145] In addition, this embodiment jointly optimizes the training process of the style modulation module and the makeup fusion module, and the overall loss function Defined as:

[0146]

[0147] Among them, λ1, λ2, λ3, λ4, and λ5 are the weight hyperparameters of each loss term.

[0148] (4.1) Perceptual loss: The input source image x is extracted using the trained VGG-19 network s and generate images Perceptual features. Since the two are not aligned at the pixel level, only the feature representation φ after the ReLU activation layer is used l (·). Perceptual loss The definition is as follows:

[0149]

[0150] where |·|2 represents the L2 norm.

[0151] (4.2) Makeup loss: Some methods propose to generate pseudo ground truth based on thin plate spline (TPS) deformation and histogram matching. However, TPS deformation with low degrees of freedom is not sufficient to preserve complex makeup styles under large-scale geometric changes. To address this limitation, we improve the generation of pseudo ground truth by replacing TPS deformation with optimal transfer matching, because optimal transfer matching is able to establish dense correspondences between distributions. Specifically, we use histogram matching to align the color distribution of the lip, skin and eye regions in the source image to be consistent with the color distribution of the reference image. This ensures that the transferred makeup style is similar to the desired appearance in color and intensity. We then extract the corresponding regions from the deformation results generated by optimal transfer matching and fuse them with the output of histogram matching to generate the final pseudo labels. Let PGT(x,y) and PGT(y,x) denote the pseudo labels for makeup transfer and inverse transfer, respectively. Makeup loss It can be expressed as:

[0152]

[0153] (4.3) Cycle consistency loss: To encourage the consistency of the generated model in the bidirectional transfer process, we introduce a global cycle loss The form is as follows:

[0154]

[0155] (4.4) Identity loss: To maintain the consistency of the image identity features, we introduce identity loss, which requires the input image to be reconstructed into the original image after self-mapping. Identity loss The definition is as follows:

[0156]

[0157] (4.5) Adversarial loss: Introducing an adversarial training mechanism, using the discriminator D to model the latent space of the output image to enhance the generated image Realism. Discriminator and generators The losses are expressed as:

[0158]

[0159] Among them, the hinge function h(t)=min(0,-1+t) is used to regularize the discriminator.

[0160] Example 2

[0161] An embodiment of the present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a style-modulated controllable makeup migration method based on dense deformation alignment as described in Example 1.

[0162] It is understood that the memory may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the server, etc.

[0163] The processor may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various components within the server. It executes various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, as well as accessing data stored in memory. Optionally, the processor may be implemented using at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may integrate one or a combination of a central processing unit (CPU) and a modem. The CPU primarily processes the operating system and application programs, while the modem handles wireless communications. It is understood that the modem may not be integrated into the processor and may be implemented separately via a single chip.

[0164] Since the electronic device is an electronic device corresponding to a style-modulated controllable makeup transfer method based on dense deformation alignment in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0165] Example 3

[0166] An embodiment of the present invention also provides a computer-readable storage medium, which stores at least one instruction, at least one program, a code set, or an instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a style-modulated controllable makeup migration method based on dense deformation alignment as described in Example 1.

[0167] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0168] Since the storage medium is a storage medium corresponding to a style-modulated controllable makeup migration method based on dense deformation alignment in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0169] Example 4

[0170] In some possible implementations, various aspects of the method of the embodiments of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to cause the computer device to execute the steps of the style-modulated controllable makeup transfer method based on dense deformation alignment according to various exemplary embodiments of the present application described above in this specification. The executable computer program code or "code" used to execute various embodiments may be written in a high-level programming language such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0171] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0172] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0173] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.

Claims

1. A style-modulated controllable makeup transfer method based on dense deformation alignment, characterized by: The following steps are involved: Get the image to be processed; Input the image to be processed into the pre-trained makeup transfer model and output an image with the corresponding makeup style; The makeup transfer model includes a dense deformation alignment module, a style modulation module and a makeup fusion module; The dense deformation alignment module adopts KL divergence as a constraint and uses the unbalanced optimal transfer theory to solve the feature alignment problem and obtain the makeup migration area; The style modulation module combines the makeup migration regions from the dense deformation alignment module to achieve fine-grained control of the normalization process; The makeup fusion module is used to fuse the features obtained by the dense deformation alignment module and the style modulation module to generate a target image.

2. The style-modulated controllable makeup transfer method based on dense deformation alignment according to claim 1, characterized in that: The input of the dense deformation alignment module is the reference image y r and semantic segmentation map of the source image The dense deformation alignment module works as follows: From the source image x s and the reference image y r Extract local feature vectors and construct corresponding feature sets f x ={s1,…,s n } and f y ={r1,…,r n }, where n represents the number of eigenvectors; introducing the quality variable γ i and δ j Respectively represent the assignment to s i and r j The mass of the two sets is and Let the distance matrix C ij Indicates that the mass γ i From the feature i Transfer to feature r j Corresponding mass δ j the price; KL divergence is used as a soft constraint term, and the entropy regularization term is introduced: Where P represents the transmission matrix, P ij Indicates that in γ i and δ j The quality of transmission between Use the transfer matrix P to warp the reference image y r , obtain the intermediate representation W that combines makeup style and structure yr : W yr =y r ·P Reference semantic graph Deform to extract W yr The specific deformation area is denoted as Extract makeup migration areas Where ⊙ is the Hadamard product.

3. The style-modulated controllable makeup transfer method based on dense deformation alignment according to claim 2, characterized in that: The final expression of the optimal transmission objective function is as follows: Where τ is the regularization parameter of the KL divergence term, η is the regularization coefficient that adjusts the smoothness and dispersion of the transmission plan; T is the transpose.

4. The style-modulated controllable makeup transfer method based on dense deformation alignment according to claim 3, characterized in that: Convert the expression of the optimal transport objective function into the Fenchel-Legendre dual form: The Sinkhorn algorithm is used to solve the problem and obtain the optimal transmission plan P. P is encoded by the dual vectors u and v and is expressed as follows:

5. The method for style-modulated controllable makeup transfer based on dense deformation alignment according to claim 2, characterized in that: The mass term is dynamically calculated using the following formula: Distance matrix C ij Defined as cosine distance, the expression is as follows: Where T is the transpose.

6. The style-modulated controllable makeup transfer method based on dense deformation alignment according to claim 1, characterized in that: The style modulation module works as follows: From the deformed semantic graph Extract the style information of the corresponding area and concatenate them to form the style matrix ST; Migrate makeup to the area And the style matrix ST inputs independent convolution channels, extracting two sets of normalized modulation parameters: At each normalized layer i, let its input activation be Among them, B and C i 、H i 、W i They represent the batch size, number of channels, height, and width of the feature map of the i-th layer respectively; the mean and standard deviation in the channel direction are and The normalized output is as follows: Where, and It is a spatially sensitive learnable weight that is adaptively learned based on the style map and deformation results.

7. The style-modulated controllable makeup transfer method based on dense deformation alignment according to claim 1, characterized in that: The makeup fusion module consists of an identity encoder, two upsampling layers, five fusion blocks, and a decoder; The makeup fusion module works as follows: First, use the identity encoder to get the source image x s Multi-scale identity features are extracted from the fusion block; these features are then fed into the fusion block for processing; Each fusion block receives the affine transformation parameters generated by the dense deformation alignment module and uses them to modulate the normalization process to achieve an adaptive combination of style and structure; To gradually restore the image resolution, two upsampling layers are inserted between the fusion blocks to expand the spatial size of the feature map; Finally, the decoder receives the output of the last fusion block and generates the target image The image effectively incorporates the makeup style of the reference image while maintaining the identity of the source image.

8. The style-modulated controllable makeup transfer method based on dense deformation alignment according to claim 1, characterized in that: The loss function during the training of the makeup transfer model as follows: Where λ1,λ2,λ3,λ4,λ5 are the weight hyperparameters of each loss term; Among them, the perceptual loss The expression is: Where, φ l (·) represents the features after the ReLU activation layer, is the generated target image, ||2 represents the L2 norm; Makeup loss The expression is: Where, PGT(x s ,y r ) and PGT(y r ,x s ) represent the pseudo labels of makeup transfer and reverse transfer respectively; f(x s ,y r ) represents the source image x s and the reference image y r The makeup migration image generated by input, f(y r ,x s ) represents the reference image y r and the source image x s The reverse makeup transfer image generated as input; Cycle consistency loss The expression is: Where, ||1 represents the L1 norm; identity loss The expression is: Fighting Losses The losses of the discriminator D and the generator G are expressed as follows: Where E represents the mathematical expectation, and h() represents the function used to regularize the discriminator.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.