Imitation portrait editing system based on diffusion model
Patent Information
- Application Number
- CN202610813521.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-15
AI Technical Summary
然而,现有基于示例的编辑方法主要针对风格迁移或大范围视觉变化任务,对于人像结构性编辑所涉及的细微变化缺乏足够的敏感性,难以准确提取并迁移面部结构层面的微小差异
1、本发明通过引入基于示例图像对的模仿式编辑机制,有效克服了自然语言在描述精细化人像结构调整方面的表达局限。通过直接从原始参考图像与修图后参考图像之间提取修图操作特征,系统能够准确捕捉面部五官位置、轮廓及比例等细微结构变化,并将该变化迁移至新的查询图像,从而避免了传统文本引导方法中常见的语义歧义问题,使编辑结果更加自然、精确且符合预期。
Smart Images

Figure CN122760740A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and artificial intelligence technology, and more specifically, to a mimicry-based portrait editing system based on a diffusion model. Background Technology
[0002] With the rapid development of artificial intelligence and computer vision technologies, deep learning-based image generation and editing techniques have been widely researched and applied. Among these, generative models, represented by diffusion models, have made significant progress in image generation quality, diversity, and detail representation, gradually becoming an important technical approach in the current image editing field. Building upon this foundation, researchers have proposed various text-guided image editing methods that modify image content by inputting natural language descriptions, achieving a degree of controllable image content generation.
[0003] However, in the field of portrait image editing, especially when it comes to fine-grained structural editing tasks such as adjusting facial features, optimizing contours, and fine-tuning proportions, existing technologies still have significant limitations. First, natural language itself is difficult to accurately describe fine-grained spatial structural changes. For example, information such as minute displacements, deformation amplitudes, and directions of local facial regions is difficult to express accurately through text. This leads to misunderstandings in text-based editing models, resulting in under-editing or over-editing, and even damaging the original identity features of the person, making the generated results unnatural or distorted.
[0004] To overcome the limitations of textual representation, example-based image editing methods have gained increasing attention. These methods provide "before-edit" and "after-edit" example image pairs, allowing the model to learn the implicit editing operations and transfer them to new images. However, existing example-based editing methods primarily target style transfer or tasks involving large-scale visual changes, lacking sufficient sensitivity to the subtle changes involved in structural editing of portraits, and struggling to accurately extract and transfer minute differences at the facial structure level. Furthermore, in practical applications, obtaining strictly aligned data of different individuals under the same retouching operation is extremely difficult, causing the model to easily rely on fixed spatial locations or specific sample distributions during training, thus limiting its cross-identity generalization ability.
[0005] Furthermore, traditional image editing methods often operate directly in pixel space, making it difficult to balance structural consistency and generated quality. While existing diffusion models offer advantages in generated quality, they lack a dedicated mechanism for "operation transfer," failing to leverage editing information from example image pairs for fine-grained control. Therefore, accurately extracting, effectively representing, and stably transferring example retouching operations within the diffusion model framework has become a critical issue that urgently needs to be addressed in current portrait editing technology.
[0006] Therefore, there is an urgent need for a diffusion-based, imitative portrait editing system to solve these problems. Summary of the Invention
[0007] The purpose of this invention is to solve the technical problems mentioned in the background section and to provide a diffusion-based mimicry portrait editing system.
[0008] The above-mentioned objective of the present invention is achieved as follows: The mimicry-based portrait editing system based on the diffusion model includes: an image retouching operation feature extraction module, a feature pre-training module, a diffusion model operation transfer module, and a data self-reinforcement training module. The image retouching operation feature extraction module receives the original reference image and the corresponding retouched reference image, inputs the two images into a mask autoencoder to extract image block-level feature sequences, and inputs the feature sequences into a Transformer structure. By introducing a learnable query sequence and interacting with the image block-level feature sequences through attention, difference information is extracted between the two images to form a feature sequence characterizing the image retouching operation. The feature pre-training module is used in the early stage of model training to map the image editing operation feature sequence into an edit embedding by constructing an auxiliary training process based on the reconstruction task. After fusing it with the image features of the query image, the embedding is input into the decoding network for image reconstruction, so that the image editing operation features can stably represent changes in image structure. The diffusion model operation transfer module is used to map the image retouching operation feature sequence to the conditional input information of the diffusion model through a connector, and inject the conditional input information into the denoising process of the diffusion Transformer model. During the gradual denoising process of the latent representation of the query image, it guides the corresponding structural changes to be generated, thereby generating a target image consistent with the image retouching operation. The data self-enhancing training module is used to synchronously apply a consistent random spatial transformation to the original reference image and the retouched reference image to generate a query image and a corresponding target image. This allows the model to decouple the correspondence between spatial location and retouching operation during training, thereby improving the transferability of retouching operation between different people.
[0009] Furthermore, the mask autoencoder is a parameter-frozen visual coding network that divides the input image into multiple fixed-size image blocks and encodes each image block to generate a feature sequence with spatial location information.
[0010] Furthermore, the Transformer structure is an R-Former network for extracting image retouching operations. This network introduces multiple sets of learnable query sequences at the input, enabling the query sequences to interact with the original reference image features and the retouched reference image features in the same feature space through multi-layer attention, and outputs features corresponding to the query sequence positions as representations of the image retouching operations.
[0011] Furthermore, the image retouching operation feature sequence is an information representation that can reflect changes in the local structure of a human figure. It is used to describe subtle adjustments to the structure of a human figure, such as the position of facial features, facial contours, body contours, and limb proportions, rather than changes in overall style or color.
[0012] Furthermore, in the feature pre-training module, the edit embedding is obtained by dimensionality reduction mapping of the image editing operation feature sequence, and is superimposed on the image features of the query image in a block-by-block manner, so that the decoding network can reconstruct the corresponding edited image based on the edit embedding.
[0013] Furthermore, in the diffusion model operation transfer module, the query image is first encoded into a latent representation by a variational autoencoder, and then gradually restored to an image space representation during the diffusion denoising process. The image editing operation features are converted into conditional control information through the connector and participate in the denoising process.
[0014] Furthermore, the connector is a mapping network used to map the features of the image retouching operation to the conditional space of the diffusion model. Its output is consistent with the feature dimension inside the diffusion Transformer model, so as to perform feature fusion during the denoising process.
[0015] Furthermore, the diffusion Transformer model keeps the original backbone network parameters frozen during training, and only performs joint training on the low-rank adaptive module embedded in the attention layer, as well as the connector and image editing operation extraction module, in order to reduce training costs and maintain generation stability.
[0016] Furthermore, in the diffusion model operation transfer module, an attention mechanism is used to realize the information interaction between the latent representation of the query image, visual features, and image retouching operation features, so that the image retouching operation can adaptively act on the human image structure at different spatial locations.
[0017] Furthermore, in the data self-enhancement training module, the random spatial transformation includes a combination of at least two operations among scaling, cropping, translation, and rotation, and the same transformation is applied to both the original reference image and the edited reference image to reduce the model's dependence on fixed spatial position correspondence, thereby improving its generalization ability for cross-identity portrait editing tasks.
[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention effectively overcomes the limitations of natural language in describing refined human facial structure adjustments by introducing a mimicry-based editing mechanism based on example image pairs. By directly extracting editing operation features from the original reference image and the edited reference image, the system can accurately capture subtle structural changes in facial features such as position, contour, and proportion, and transfer these changes to the new query image. This avoids the semantic ambiguity problems common in traditional text-guided methods, making the editing results more natural, accurate, and as expected.
[0019] 2. This invention constructs an image retouching operation extraction mechanism composed of a masked autoencoder and an R-Former network, and combines it with a pre-training strategy to enable the model to perceive and stably represent minute differences in example image pairs with high sensitivity. Simultaneously, by introducing a connector structure to map the image retouching operation features to the conditional space of the diffusion model, and combining it with a low-rank adaptive module to efficiently fine-tune the diffusion model, precise control of structured editing operations is achieved while maintaining the original generation capabilities, thereby significantly improving the quality and stability of portrait editing.
[0020] 3. This invention employs a data self-reinforcement training mechanism, constructing training samples by applying a consistent spatial transformation to example images. This reduces the model's reliance on fixed spatial locations during training, allowing it to learn the structural changes inherent in the image editing operations themselves. This significantly improves the model's transferability and generalization performance across different individuals. This method does not require large-scale, strictly aligned datasets, reducing data acquisition costs and enhancing the system's applicability and robustness in practical applications. Attached Figure Description
[0021] Figure 1 This is a diagram showing the overall system architecture and module interaction of the mimicry portrait editing system based on the diffusion model in this embodiment of the invention. Figure 2 This is a schematic diagram of the structure of a diffusion-based mimicry portrait editing system. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with embodiments and appendices. Figure 1-2 The present invention will be further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0023] Example 1: This example provides a diffusion-based mimicry-based portrait editing system. This system can be deployed on servers, image processing terminals, or devices with artificial intelligence computing capabilities. It is used to perform structural portrait editing on new query images based on the retouching effects in example image pairs. During operation, the system uses the original reference image and its corresponding retouched reference image as example inputs, combined with a query image to be processed. It automatically extracts and transfers the retouching operations from the examples to generate the corresponding target portrait image.
[0024] In this embodiment, the system first receives the original reference image and the retouched reference image, and then inputs both images into a pre-trained mask autoencoder with frozen parameters for processing. The mask autoencoder divides the input image into multiple fixed-size image blocks and encodes each image block to obtain an image block feature sequence with spatial location information. In this way, fine-grained feature representations of each local region can be extracted while preserving the overall structural information of the image.
[0025] After obtaining the feature sequences of the original reference image and the retouched reference image, the system inputs both, along with a set of pre-set learnable query sequences, into the R-Former network. The R-Former network is a feature interaction model built on the Transformer structure. Through a multi-layer attention mechanism, it enables the query sequences and the two sets of image features to interact in the same feature space, thereby automatically focusing on regions where there are differences between the two images. After multi-layer feature interaction, the system extracts the output features corresponding to the query sequences as retouching operation features. These features characterize the portrait structure adjustment information contained in the example image pair, including changes in facial contours, fine adjustments to the position and proportion of facial features, and subtle structural changes such as body contours and limb proportions.
[0026] To improve the expressive power of image retouching operation features, this embodiment introduces a feature pre-training process in the early stages of system training. In this process, the system maps the extracted image retouching operation features to edit embeddings, and then fuses these edit embeddings with the image features of the query image before inputting them into the decoding network for image reconstruction. By continuously comparing the differences between the reconstructed results and the corresponding target image, the image retouching operation feature extraction module is optimized to ensure it can stably and accurately express the structural change information corresponding to the image retouching operation. After pre-training is complete, the decoding network used for reconstruction is removed, and only the image retouching operation feature extraction module is retained for the subsequent training and inference process of the diffusion model.
[0027] After completing the feature extraction for image retouching, the system enters the diffusion model operation transfer stage. First, the query image is input into a variational autoencoder to map it into a latent space representation, reducing data dimensionality and improving computational efficiency. Simultaneously, the system utilizes a visual feature extraction network to encode semantic information into the query image, obtaining feature representations that characterize the overall structure and semantic information of the human image. The visual feature extraction network is preferably implemented using a visual language model, such as the Qwen2.5 VL model, but is not limited to this model.
[0028] Subsequently, the image retouching operation features are input into the connector structure and converted into conditional information that the diffusion model can accept through mapping. This conditional information is consistent with the internal features of the diffusion Transformer model in terms of feature dimension and representation space, thus enabling it to effectively participate in feature fusion during subsequent denoising.
[0029] In the diffusion model, the system employs a diffusion Transformer as the backbone network, introducing a two-stream structure within it to handle the latent representation and conditional information of the image separately. During training, the system injects noise into the latent representation of the target image at different time steps to construct noisy samples. These noisy latent representations, the semantic features of the query image, and the editing conditional information are then input into the diffusion Transformer model. During the model's progressive denoising process, an attention mechanism facilitates information interaction between the latent representation, semantic features, and editing operations. This allows the model to progressively apply structural editing operations from example image pairs while restoring the image content, thereby generating a target image consistent with the example results.
[0030] To reduce training costs while maintaining generative capabilities, this embodiment freezes the parameters of the original backbone network during the diffusion model training process, training only the image editing operation feature extraction module, the connector structure, and the low-rank adaptive module embedded in the diffusion Transformer attention layer. This approach allows for efficient adaptation to specific portrait editing tasks while fully utilizing the generative capabilities of the pre-trained diffusion model.
[0031] Furthermore, to address the challenge of obtaining strictly aligned training data, this embodiment employs a data self-enhancement training strategy. When constructing training samples, the system applies the same random spatial transformations—including scaling, cropping, translation, and rotation—to both the original and retouched reference images, generating new query images and corresponding target images. This approach ensures that the retouching operation remains consistent across different spatial locations while altering its specific positional distribution within the image. This reduces the model's reliance on fixed spatial correspondences during training, allowing it to focus more on learning the structural changes inherent in the retouching operation itself. Consequently, it enhances the model's transferability and generalization performance across different individuals.
[0032] In practical applications, when the system receives a new query image, it only needs to input the original reference image and the retouched reference image as examples. The system can automatically extract the retouching operation features and gradually edit the query image during the diffusion denoising process, ultimately outputting a target image that retains the person's identity features and matches the example retouching effect. Through the above process, the system provided in this embodiment can accurately extract and stably transfer fine-grained portrait structure editing operations without the need for text description, demonstrating good application effects and practical value.
[0033] Example 2: This example illustrates the specific execution flow of the present invention during the training and inference phases. The system takes the original reference image, the edited reference image, and the query image as input. By extracting the editing operation features and injecting them into a diffusion model, it achieves the transfer of portrait structure editing from the query image.
[0034] In this embodiment, training samples are first constructed. Given an original reference image... Reference image after retouching The system applies the same random spatial transformation to both simultaneously to obtain the query image. With target image : ; in, Represents the original reference image; This refers to the reference image after retouching. Indicates a query image; Represents the target image; Represents a shared-space transformation operator; This represents a set of transformation parameters, including at least two of scaling, cropping, translation, and rotation. This step causes a spatial change in the image retouching operation, thereby reducing the model's dependence on a fixed spatial position and enabling the model to learn structural change patterns.
[0035] The system then proceeds to the feature extraction stage for image retouching. The system inputs the original reference image and the retouched reference image into a frozen mask autoencoder to obtain the image patch feature sequence: ; in, and All are of length Feature dimension is Image patch feature sequences; Indicates a frozen mask autoencoder encoder; Indicates the number of image patches; The dimension of each image patch feature is represented, and then a learnable query sequence is introduced. Its dimensions are ,in Indicates the length of the query sequence. , and The sequences are concatenated along the sequence dimension and then fed into the R-Former network for feature interaction. ; ; in, Represents an R-Former network; Indicates the output feature sequence; This represents the feature sequence of image retouching operations, with dimension 1. Sel(·) indicates that the output feature corresponding to the position of the query sequence is selected.
[0036] This retouching operation feature sequence is used to characterize local structural changes in a portrait. To ensure the retouching operation features have good expressive power, the system pre-trains them. First, the retouching operation feature sequence is compressed into an edit embedding vector using a projection network: ; in, The dimension is This represents a projection network.
[0037] because Dimensions With query image feature dimensions Inconsistency, therefore, dimension alignment mapping is introduced: ; in, The dimension is This indicates a dimension-aligned mapping network.
[0038] The query image is then input into the mask autoencoder encoder: ; in, To query the image patch feature sequence of an image, with dimension 1 .
[0039] Broadcast the edited embeddings and overlay them onto the query image features: ; in, This represents the fused query feature sequence; Represents the broadcast operator, which will Expand to Same sequence length.
[0040] Subsequently, the fused features are input into the decoder for reconstruction: ; And it is optimized using the following loss function: ; in, and These are the weighting coefficients; Indicates the mean square error loss; This represents perceptual loss. Through this process, the image editing operation features can accurately express structural editing information.
[0041] After pre-training, the system enters the joint training phase of the diffusion model. First, a variational autoencoder maps the query image and the target image into the latent space: ; in, and All are latent representations, with dimensions of ; Indicates the length of the potential sequence; This represents the dimension of the latent features. To ensure consistent distribution in the latent space, a diffusion path is constructed: ; in, Represents a noisy latent representation; and For time step The coefficient is used to control the ratio of data to noise; It is standard Gaussian noise, and its dimension is... same.
[0042] The corresponding true velocity field is defined as: ; in, and These are the time-step-related derivative coefficients, used to describe the rate of change of the diffusion path. Subsequently, the image retouching operation features are input into the connector: ; in, This represents a sequence of image editing conditions, with dimension 1. And by mapping it internally to the feature dimensions of the diffuse Transformer model Maintain consistency.
[0043] Simultaneously, extract semantic features from the query image: ; Finally, the noisy latent representation, the query image latent representation, the time step, semantic features, and image retouching conditions are input into the diffusion model: ; in, Represents the diffusion Transformer model; Indicates the freeze parameter; This indicates that trainable parameters are introduced into the model using a low-rank adaptive structure: ; in, This is the original weight matrix; and It is a low-rank matrix, and its rank is less than 1. The dimension is used to achieve efficient fine-tuning of the model optimization objective as follows: ; During the reasoning phase, only input... , and The image retouching operation features are extracted through the above process, and a diffusion denoising process is performed to obtain the final retouched image: ; in, This represents the potential representation after denoising.
[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A diffusion-based mimicry portrait editing system, characterized in that, include: The system includes an image retouching operation feature extraction module, a feature pre-training module, a diffusion model operation transfer module, and a data self-enhancement training module. The image retouching operation feature extraction module receives the original reference image and the corresponding retouched reference image, inputs the two images into a mask autoencoder to extract image block-level feature sequences, and inputs the feature sequences into a Transformer structure. By introducing a learnable query sequence and interacting with the image block-level feature sequences through attention, the system extracts difference information between the two images to form a feature sequence characterizing the image retouching operation. The feature pre-training module is used in the early stage of model training to map the image editing operation feature sequence into an edit embedding by constructing an auxiliary training process based on the reconstruction task. After fusing it with the image features of the query image, the embedding is input into the decoding network for image reconstruction, so that the image editing operation features can stably represent changes in image structure. The diffusion model operation transfer module is used to map the image retouching operation feature sequence to the conditional input information of the diffusion model through a connector, and inject the conditional input information into the denoising process of the diffusion Transformer model. During the gradual denoising process of the latent representation of the query image, it guides the corresponding structural changes to be generated, thereby generating a target image consistent with the image retouching operation. The data self-enhancing training module is used to synchronously apply a consistent random spatial transformation to the original reference image and the retouched reference image to generate a query image and a corresponding target image. This allows the model to decouple the correspondence between spatial location and retouching operation during training, thereby improving the transferability of retouching operation between different people.
2. The system according to claim 1, characterized in that, The mask autoencoder is a parameter-frozen visual coding network that divides the input image into multiple fixed-size image blocks and encodes each image block to generate a feature sequence with spatial location information.
3. The system according to claim 1, characterized in that, The Transformer structure is an R-Former network used to extract image retouching operations. This network introduces multiple sets of learnable query sequences at the input, allowing the query sequences to interact with the original reference image features and the retouched reference image features in the same feature space through multi-layer attention interactions, and outputs features corresponding to the query sequence positions as representations of the image retouching operations.
4. The system according to claim 1, characterized in that, The image retouching operation feature sequence is an information representation that can reflect changes in the local structure of a human portrait. It is used to describe subtle adjustments to the structure of a human portrait, such as the position of facial features, facial contours, body contours, and limb proportions, rather than changes in overall style or color.
5. The system according to claim 1, characterized in that, In the feature pre-training module, the editing embedding is obtained by dimensionality reduction mapping of the image editing operation feature sequence and superimposed on the image features of the query image in a block-by-block manner, so that the decoding network can reconstruct the corresponding edited image based on the editing embedding.
6. The system according to claim 1, characterized in that, In the diffusion model operation transfer module, the query image is first encoded into a latent representation by a variational autoencoder and then gradually restored to an image space representation during the diffusion denoising process. The image editing operation features are converted into conditional control information through the connector and participate in the denoising process.
7. The system according to claim 1, characterized in that, The connector is a mapping network used to map the features of the image retouching operation to the conditional space of the diffusion model, and its output has the same feature dimension as the internal feature dimension of the diffusion Transformer model.
8. The system according to claim 1, characterized in that, The diffusion Transformer model keeps the original backbone network parameters frozen during training, and only performs joint training on the low-rank adaptive module embedded in the attention layer, as well as the connector and image editing operation extraction module.
9. The system according to claim 1, characterized in that, In the diffusion model operation transfer module, the information interaction between the latent representation of the query image, visual features, and image retouching operation features is realized through the attention mechanism, so that the image retouching operation can adaptively act on the human image structure at different spatial locations.
10. The system according to claim 1, characterized in that, In the data self-enhancing training module, the random space transformation includes a combination of at least two operations among scaling, cropping, translation, and rotation, and the same transformation is applied to both the original reference image and the retouched reference image.