An image makeup transfer method and system based on diffusion type converter and feature decoupling injection
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHAOYI INFORMATION TECH (SHANGHAI) CO LTD
- Filing Date
- 2026-06-10
- Publication Date
- 2026-08-07
AI Technical Summary
1、辅助控制模块带来的误差累积与系统臃肿:现有的妆容迁移技术(无论是基于GAN还是早期的扩散模型),为了将妆容对齐到目标人脸,高度依赖外部的辅助面部控制模块(如3D面部形变模型、106关键点检测或面部语义分割网络)和损失函数
1、无需辅助面部控制模块,系统架构精简高效:彻底摆脱了对人脸关键点检测、3D形变模型或面部语义分割等额外算法的依赖,从根本上避免了因辅助模块识别偏差或遮挡造成的误差累积,显著提高了系统的整体鲁棒性与运行效率。
Smart Images

Figure CN122529960A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and artificial intelligence (AI), and in particular to an image makeup transfer method and system based on a diffusion-type transformer and feature decoupling injection. Background Technology
[0002] Image makeup transfer, an innovative technology applied to facial feature processing and computational aesthetics, aims to accurately transfer makeup styles (such as color distribution, makeup details, and local textures) from a given reference image to the target user's bare face, while preserving the user's original facial features and structure. This technology has enormous commercial and application value in fields such as virtual makeup try-on, film and television post-production editing, interactive entertainment, digital human generation, and beauty-enhanced social media.
[0003] Currently, the technical solutions for achieving makeup transfer can be mainly divided into two categories: methods based on traditional image processing and methods based on unsupervised deep learning.
[0004] Traditional makeup transfer methods largely rely on low-level image processing algorithms, such as color space conversion, histogram matching, or Poisson blending. These methods extract color statistics from the facial regions of a reference image and rigidly overlay them onto the source image. However, these static rule-based algorithms typically only handle simple global color transformations (such as changing overall skin tone or lip color), lacking the ability to perceive local spatial details and struggling to adapt to facial geometric deformations caused by changes in posture, expression, and lighting. Especially when faced with various complex and extreme makeup looks in reality, traditional methods often produce severe edge artifacts and unnatural appearances. They usually rely on human perception, simply overlaying existing beauty software's skin smoothing and makeup techniques using a template, requiring a large amount of manpower and failing to meet diverse and personalized high-precision transfer needs.
[0005] With the development of deep learning models such as Generative Adversarial Networks (GANs) and Diffusion Models, makeup transfer technology based on neural networks has gradually become mainstream. This type of technology leverages the powerful feature extraction and generation capabilities of networks to significantly improve the naturalness of makeup transfer. However, current mainstream research and industrial applications still have the following shortcomings: 1. Error Accumulation and System Bulkiness Caused by Auxiliary Control Modules: Existing makeup transfer technologies (whether based on GANs or earlier diffusion models) heavily rely on external auxiliary facial control modules (such as 3D facial deformation models, 106 keypoint detection, or facial semantic segmentation networks) and loss functions to align makeup to the target face. This architecture has a fatal flaw: if the source image contains a profile, exaggerated expressions, or partial occlusion, the auxiliary module will produce recognition bias. These prior errors are amplified layer by layer in the network, eventually leading to unreasonable images such as misaligned makeup. At the same time, the complex external modules greatly increase the computational cost and deployment difficulty of the system.
[0006] 2. Identity Structure Disruption and "Over-Alignment" Leading to a Texture-like Feel: When using deep learning for transfer learning, existing large models (such as mainstream diffusion model fine-tuning schemes) typically directly concatenate the source image and the reference makeup image in the latent space or at the input. This fusion method easily leads to "feature confusion" in the network: the model not only extracts the makeup information from the reference image but also incorrectly creates a dependency, directly using the reference image as the output and forcibly covering the features of the source image. This phenomenon is known in the industry as "over-alignment," which directly causes the generated image to lose the user's original identity, producing a very unnatural and stiff "texture-like feel." Summary of the Invention
[0007] To address the aforementioned shortcomings of existing technologies, this invention provides an image makeup transfer method and system based on a diffusion-type transformer and feature decoupling injection.
[0008] The present invention provides an image makeup transfer method based on a diffusion-type transformer and feature decoupling injection, comprising the following steps: S1. Data Input and Preprocessing: Obtain an image of the target face with bare skin or basic makeup, as well as an image of the target makeup. Perform preprocessing to obtain the source image. Compared with reference image ; S2. Latent space coding: The source image... Encoding as native conditional features The reference image Encoding as reference latent features ; S3. Makeup Feature Decoupling Injection: The reference latent features are decoupled and injected using a low-rank adaptive projection mechanism. Perform a dimension reduction-up linear projection to generate a makeup reference key matrix for injection. Makeup Reference Value Matrix ; S4. Diffused Transformer (DiT) Denoising Generation: Acquiring As the native conditional input to the backbone generation network of the diffusion transformer, a native joint sequence is constructed to generate the query matrix. ,Will , The original bond matrices are asymmetrically concatenated to the DiT self-attention layer. Original value matrix Maintain the query matrix With the length remaining constant, the target latent features are output after multiple rounds of denoising iterations. ;as well as S5. Image Decoding: Decoding the target's latent features Decoding and reconstructing into a high-fidelity image with makeup. And output it.
[0009] This invention provides an image makeup transfer system based on a diffusion-type transformer and feature decoupling injection, which uses the aforementioned image makeup transfer method, characterized by comprising: The data input and preprocessing module is used to acquire images of the target face with bare skin or basic makeup, as well as images of the target makeup, and to obtain the source image after preprocessing. Compared with reference image ; Latent space coding module, used to encode the source image Encoding as native conditional features The reference image Encoding as reference latent features ; The makeup feature decoupling injection module is used to inject the reference latent features through a low-rank adaptive projection mechanism. Perform a dimension reduction-up linear projection to generate a makeup reference key matrix for injection. Makeup Reference Value Matrix ; The diffused converter backbone generation module, composed of multiple converter blocks connected in series, is used to obtain... As the original condition input, a native joint sequence is constructed to generate the query matrix. ,Will , The original bond matrices are asymmetrically concatenated to the DiT self-attention layer. Original value matrix Maintain the query matrix With the length remaining constant, the target latent features are output after multiple rounds of denoising iterations. ;as well as Image decoding module, used to extract the target's hidden features Decoding and reconstructing into a high-fidelity image with makeup. And output it.
[0010] The present invention has the following beneficial effects: 1. No auxiliary facial control module required, streamlined and efficient system architecture: It completely eliminates the reliance on additional algorithms such as facial key point detection, 3D deformation model or facial semantic segmentation, fundamentally avoiding the accumulation of errors caused by recognition deviation or occlusion of auxiliary modules, and significantly improving the overall robustness and operating efficiency of the system.
[0011] 2. Effectively prevents "over-alignment" and ensures high fidelity in makeup transfer: Utilizing an innovative low-rank feature injection module, the model possesses the superior characteristic of "extracting only makeup styles without copying facial structures." Whether it's everyday makeup or extremely exaggerated creative makeup, it can be naturally and seamlessly rendered on the target person's face, completely eliminating the harsh "texture look."
[0012] 3. Ultimate identity consistency and lossless background preservation: With the native conditional input architecture, the generated image can perfectly preserve the user's original facial proportions, facial topology and complex image background, without the background distortion, edge splicing or facial distortion that are common in traditional generative models.
[0013] 4. Strong industrial applicability and scalability: Combining the powerful scalability of the diffusion-type Transformer, this invention is easily deployed in various high-concurrency cloud image editing, virtual makeup, and film and television post-production systems. Attached Figure Description
[0014] Figure 1 This is a flowchart of an image makeup transfer method based on a diffusion-type transformer and feature decoupling injection, according to a preferred embodiment of the present invention.
[0015] Figure 2 This is a block diagram of an image makeup transfer system based on a diffusion-type transformer and feature decoupling injection, according to a preferred embodiment of the present invention. Detailed Implementation
[0016] The present invention will be further illustrated below through examples, the purpose of which is only to better understand the research content of the present invention and not to limit the scope of protection of the present invention.
[0017] A preferred embodiment of the present invention provides an image makeup transfer method based on a diffusion-type transformer and feature decoupling injection. This method does not rely on external auxiliary modules such as facial landmark detection, semantic segmentation, or 3D deformation models. Through native conditional input, makeup feature decoupling, low-rank adaptive feature injection, asymmetric attention concatenation, and stream matching denoising generation, it achieves high-fidelity, identity-consistent, and background-loss-free makeup transfer. Figure 1 As shown, the overall process of the preferred embodiment of the present invention includes the following steps S1 to S5.
[0018] Step S1 is the data input and preprocessing step. In this step, the source image and reference image to be processed are obtained. The source image is a user's face image with bare skin or basic makeup, and the reference image is a face image containing the target makeup style. Only basic preprocessing operations are performed on the source image and reference image, specifically including face detection, center cropping, and size scaling. Both images are uniformly adjusted to a standard resolution of 512×512 pixels, and the image pixel values are normalized to the [-1,1] range to obtain the preprocessed source image. and reference image This step does not introduce any external face alignment algorithms, thus avoiding the accumulation of errors caused by recognition bias or occlusion in the auxiliary module from the source, and ensuring the robustness of subsequent processing.
[0019] Step S2 is the latent space encoding step. In this step, a trained variational autoencoder (VAE) with frozen weights and shared parameters is used to encode the preprocessed source image. and reference image Perform latent space mapping to obtain source sequence features. and reference sequence features VAE is used to compress image data in pixel space and map it to a low-dimensional latent space.
[0020] Specifically, and After inputting into the frozen VAE encoder, first output the dimensions as follows: The source latent space feature map and the reference latent space feature map, where denoted as the number of channels in the latent space feature map. Let be the height and width of the latent space feature map; further, flatten the two types of feature maps into the sequence features required by the backbone generation network of the diffusion transformer, obtaining a dimension of . source sequence features and reference sequence features ,in The length of the encoded image sequence. For feature dimensions. Here, source sequence features. It is directly used as the native conditional input to the backbone network of the diffusion transformer, fully inheriting the background environment, light and shadow distribution, and facial topology of the source image; reference sequence features This serves as the foundational data for subsequent makeup feature decoupling and injection.
[0021] Step S3 is the makeup feature decoupling and injection step. In this step, a low-rank adaptive projection mechanism is used to decouple the features of the reference sequence. Feature decoupling is performed to eliminate interference from facial structure information in the reference image, retaining only makeup-related color, texture, and detail features. Preferably, this invention introduces a trainable dimensionality reduction matrix. and the increasing dimension matrix ,in It is a low-rank dimension and satisfies , For the overall embedding dimension of features, For attention feature dimensions, in the current architecture, there are [number] for each layer. This preserves the feature embedding dimension; a residual projection method is used to calculate the makeup reference key matrix. Makeup Reference Value Matrix The calculation formula is: In the above formula, , These are the frozen key projection matrix and value projection matrix in the backbone network of the diffusion converter. The remaining matrices are trainable parameter matrices learned through data-driven training. , , , The final output dimension is Makeup reference key matrix Makeup Reference Value Matrix This achieves effective decoupling of makeup features from facial structural features. , It will be injected into the attention layer of the backbone network in subsequent steps.
[0022] Step S4 is the denoising generation step using the diffusion converter. In this step, the source sequence features are obtained. As the native conditional input to the backbone generation network of the diffusion transformer, a native joint sequence is constructed to generate the query matrix. Makeup reference key matrix Makeup Reference Value Matrix The original key matrix and value matrix are asymmetrically concatenated to the self-attention layer of DiT, and asymmetric key-value injection and flow-matched denoising operations are performed in the Diffusion Transformer (DiT).
[0023] Regarding native joint sequences, based on the basic architecture of the diffusion model, the system generates and Pure Gaussian noise latent variables with completely uniform dimensions ,in For the current time step, during initialization This indicates that the current latent variable is pure Gaussian noise, and the fixed text prompt "Makeup the person" is extracted into a text feature sequence by a text encoder. ,in The length of the encoded text sequence. The feature dimension is defined as follows. In each Transformer block of the backbone network DiT, the text sequence, source feature sequence, and noise sequence are concatenated along the length direction. The length of the structure is The input for the native join query condition (native join sequence) is the sequence length. , The sequence lengths are all Text features The sequence length is The total length of the original joint sequence after splicing The feature embedding dimension of each sequence is unified as .
[0024] The following section details asymmetric key-value injection and Transformer denoising generation. First, layer normalization and initial feature projection are performed. The original joint sequence is first processed by the current time step... The control layer is normalized (LN), and then its own query matrix is mapped through a linear query, key, and value projection matrix. Key matrix Sum matrix .at this time , , All dimensions are .
[0025] Then, asymmetric key-value concatenation is performed to combine the makeup reference key matrix. Makeup Reference Value Matrix They are concatenated along the sequence dimension to the original key matrix of the backbone network. Original value matrix Then, an enhanced bond matrix is formed. and enhancement value matrix After splicing, and The dimension becomes In this step, the query matrix is performed. It remains absolutely isolated and does not participate in the splicing process; its dimensions are constant. .
[0026] Then, the original join query conditions are linearly projected to obtain a query matrix of constant length. ,use Key matrix expanded by makeup features Value matrix Perform joint self-attention computation to obtain the attention output sequence. The calculation formula is as follows: In the formula This is a normalized exponential function used to probabilistically normalize the similarity results between the query and the key. (Due to the query matrix...) The sequence length is fixed, and the attention output sequence obtained after self-attention operation is... The sequence length is still equal to the length of the original joint sequence. ,Right now In this step, the latent variables to be denoised are matched with the source image features for identity and facial topology information based on the query matrix, and the makeup color and texture of the reference image are matched based on the expanded key-value pairs, thus completing the accurate alignment and rendering of makeup features in the latent space. This solution does not construct the query end with reference makeup features, eliminating the erroneous dependency relationship of the network learning the reference face structure, and effectively avoiding the problem of the generated result replacing the facial features of the source person.
[0027] Next, the denoising process is iteratively executed using a stream-matching sampling algorithm according to a preset number of sampling steps (e.g., 25 steps). In each iteration, the attention layer outputs... After residual connections and layer normalization, the input is fed into a feed-forward network (FFN) to complete feature nonlinear activation and dimension mapping. The output is then fed into the next diffuse transformer (DiT) block via residual connections. After 25 iterations, the initial pure Gaussian noise latent variable... The target latent features are gradually denoised into clean ones. This process uses an asymmetric key-value injection mechanism to mathematically prevent the reference face structure from covering the source face, ensuring accurate makeup rendering while fully preserving the identity features of the source image.
[0028] Finally, there is the image decoding output step S5. In this step, the target latent features generated by denoising are... The image is input to a variational autoencoder decoder, and after upsampling and pixel reconstruction, the final RGB image with makeup is generated. And output. Because the system strictly implements the "constant query sequence length" and "residual asymmetric injection" mechanisms in the underlying Transformer tensor operations, the output image with makeup is... It fully preserves the facial proportions, facial topology, facial features, and complex background information of the source image, while naturally presenting the makeup colors, textures, and details of the reference image, without edge artifacts or harsh textures, achieving a high-fidelity makeup transfer effect.
[0029] The method of this invention also includes a corresponding training process, using a pairwise dataset of bare-faced and made-up individuals with the same identity for supervised learning. During the training phase, all parameters of the variational autoencoder and the diffusion transformer backbone network are frozen, and only the low-rank projection matrix parameters in the makeup feature decoupling injection module are opened for end-to-end fine-tuning. A flow matching target loss function is used to calculate the mean square error between the network's predicted noise and the actual added noise, and gradient backpropagation and parameter updates are completed through a supervised optimization module. Simultaneously, a course learning scheduling strategy is introduced. In the early stages of training, data pairs with small structural differences and simple makeup are selected for training to help the model quickly establish basic color mapping logic. As the training loss gradually converges, data pairs with large structural differences and complex makeup are dynamically unlocked and introduced, effectively improving the model's robustness in complex scenarios such as side profiles, exaggerated expressions, and creative heavy makeup, ensuring the model's stability and adaptability in practical applications.
[0030] A preferred embodiment of the present invention also provides an image makeup transfer system based on a diffusion-type transformer and feature decoupling injection. For example... Figure 2 As shown, it includes: a data input and preprocessing module 21, used to acquire an image of the target face with bare skin or basic makeup and an image of the target makeup, and to obtain the source image after preprocessing. Compared with reference image Latent space coding module 22 is used to encode the source image Encoding as native conditional features The reference image Encoding as reference latent features The makeup feature decoupling injection module 23 is used to perform makeup feature decoupling injection on the reference latent features through a low-rank adaptive projection mechanism. Perform a dimension reduction-up linear projection to generate a makeup reference key matrix for injection. Makeup Reference Value Matrix The diffused converter backbone generation module 24, composed of multiple converter blocks connected in series, is used to obtain... As the original condition input, a native joint sequence is constructed to generate the query matrix. ,Will , The original bond matrices are asymmetrically concatenated to the DiT self-attention layer. Original value matrix Maintain the query matrix With the length remaining constant, the target latent features are output after multiple rounds of denoising iterations. ; and image decoding module 25, used to decode the target hidden features Decoding and reconstructing into a high-fidelity image with makeup. And output. Preferably, the latent space coding module 22 is composed of a variational autoencoder with shared weights, and the image decoding module 25 is composed of a variational autoencoder decoder.
[0031] Preferably, the image makeup transfer system of the present invention further includes a supervised optimization module 26, which is used to calculate the flow matching target loss based on the same identity "bare face-with makeup" binary pair dataset, freeze the parameters of the VAE and DiT backbone networks, and only perform gradient backpropagation update on the low-rank projection matrix of the makeup feature decoupling injection module 23. Specifically, the supervised optimization module 26 mainly optimizes the learnable parameters in the system end-to-end through a data-driven approach. This system abandons the alignment defects caused by traditional weakly supervised learning, and directly obtains a large number of high-quality "bare face-target makeup" pair datasets as supervision benchmarks. During training, the system freezes most of the parameters of the variational autoencoder module and the DiT backbone network, and only opens the makeup feature decoupling injection module 23 (low-rank projection matrix, i.e., dimensionality reduction matrix). and the increasing dimension matrix The parameters are updated using gradients. Furthermore, the mean squared error between the predicted noise in the network output and the actual added noise is calculated using flow matching target loss, forcing the injection module to learn how to extract and map makeup features. This training strategy not only significantly reduces the memory and time costs of model training but also ensures that the strong background preservation and identity consistency capabilities of the original backbone network are not compromised.
[0032] Obviously, those skilled in the art should recognize that the above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Any changes or modifications to the above embodiments that are within the essential spirit of the present invention will fall within the scope of the claims of the present invention.
Claims
1. An image makeup transfer method based on a diffusion-type transformer and feature decoupling injection, characterized in that, Includes the following steps: S1. Data Input and Preprocessing: Obtain an image of the target face with bare skin or basic makeup, as well as an image of the target makeup. Perform preprocessing to obtain the source image. Compared with reference image ; S2. Latent space coding: The source image... Encoding as native conditional features The reference image Encoding as reference latent features ; S3. Makeup Feature Decoupling Injection: The reference latent features are decoupled and injected using a low-rank adaptive projection mechanism. Perform a dimension reduction-up linear projection to generate a makeup reference key matrix for injection. Makeup Reference Value Matrix ; S4. Diffused Transformer (DiT) Denoising Generation: Acquiring As the native conditional input to the backbone generation network of the diffusion transformer, a native joint sequence is constructed to generate the query matrix. ,Will , The original bond matrices are asymmetrically concatenated to the DiT self-attention layer. Original value matrix Maintain the query matrix With the length remaining constant, the target latent features are output after multiple rounds of denoising iterations. ; as well as S5. Image Decoding: Decoding the target's latent features Decoding and reconstructing into a high-fidelity image with makeup. And output it.
2. The method according to claim 1, characterized in that, In step S1, the preprocessing involves performing face recognition, cropping, and resizing on the source image and reference image, and unifying them to a preset resolution.
3. The method according to claim 1, characterized in that, Step S2 also includes: , Flattening is a sequence feature, with all dimensions being [missing information]. ,in For sequence length, For feature dimensions.
4. The method according to claim 3, characterized in that, In step S2, the native joint sequence is composed of a text sequence. Source feature sequence With noise sequence It is assembled by splicing along the length direction.
5. The method according to claim 4, characterized in that, In step S3, the low-rank adaptive projection mechanism employs a dimension reduction matrix. and the increasing dimension matrix right Perform dimensionality reduction-up linear projection calculations to obtain the makeup reference key matrix. Makeup Reference Value Matrix : , , In the above formula, , The first matrix represents the key-value projection matrix frozen by DiT, while the remaining matrices are trainable parameter matrices obtained through data-driven training. , , , ,in To unify the feature embedding dimension, For attention, the single-head projection output dimension, It is a low-rank dimension and satisfies low-rank constraints. .
6. The method according to claim 5, characterized in that, In step S4, the enhanced bond matrix is obtained after asymmetric splicing. and enhancement value matrix Query matrix It maintains a constant length and does not participate in splicing.
7. The method according to claim 6, characterized in that, In step S4, a query matrix of constant length is used. The key matrix after makeup feature expansion Value matrix Perform joint self-attention computation to obtain the attention output sequence: ; The denoising iteration employs a stream-matched sampling algorithm with 25 iterations and an initial time step of T=1000; in each iteration, the attention layer outputs a sequence... After residual connection and layer normalization, the input is fed forward neural network to complete feature nonlinear activation and dimension mapping. Then, after residual connection, the output is sent to the next diffusion transformer block. After 25 iterations, the initial pure Gaussian noise latent variable Z is transformed. T Denoising to clean target latent features .
8. The method according to claim 1, characterized in that, It also includes training steps: using a binary pair dataset with the same identity "bare face - made-up face", freezing the backbone network parameters, updating only the low-rank adaptive projection matrix through gradient backpropagation, and using the flow matching target loss as the loss function.
9. An image makeup transfer system based on a diffusion-type transformer and feature decoupling injection, which uses the method of any one of claims 1-8, characterized in that, include: The data input and preprocessing module is used to acquire images of the target face with bare skin or basic makeup, as well as images of the target makeup, and to obtain the source image after preprocessing. Compared with reference image ; Latent space coding module, used to encode the source image Encoding as native conditional features The reference image Encoding as reference latent features ; The makeup feature decoupling injection module is used to inject the reference latent features through a low-rank adaptive projection mechanism. Perform a dimension reduction-up linear projection to generate a makeup reference key matrix for injection. Makeup Reference Value Matrix ; The diffused converter backbone generation module, composed of multiple converter blocks connected in series, is used to obtain... As the original condition input, a native joint sequence is constructed to generate the query matrix. ,Will , The original bond matrices are asymmetrically concatenated to the DiT self-attention layer. Original value matrix Maintain the query matrix With the length remaining constant, the target latent features are output after multiple rounds of denoising iterations. ; as well as Image decoding module, used to extract the target's hidden features Decoding and reconstructing into a high-fidelity image with makeup. And output it.
10. The system according to claim 9, characterized in that, It also includes a supervised optimization module, which uses a binary pair dataset with the same identity ("bare face - made up"), freezes the backbone network parameters, updates the gradient backpropagation only on the low-rank adaptive projection matrix, and uses the flow matching target loss as the loss function.