Efficient and structure-fidelity personalized image generation method
Optimizing personalized image generation through global semantic appearance space and loss function solves the problems of insufficient appearance quality and low computing efficiency, and realizes efficient and structurally fidelity personalized image generation.
Patent Information
- Application Number
- CN202510505209.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
AI Technical Summary
The existing personalized text-to-image generation methods are inadequate in appearance quality and low computational efficiency, especially during the generation process, which is prone to damage the image structure and lead to instability.
By constructing a global semantic appearance space (GSAS), it performs feature dimensionality reduction, and combines cosine loss and gram loss, guides the generation process, optimizes the appearance characteristics of personalized images, maintains the image structure while improving computing efficiency.
It significantly improves the appearance quality and visual reality of personalized images, while reducing computational complexity, achieving more efficient image generation.
Smart Images

Figure CN120339441A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of personalized text-to-image generation, and particularly relates to an efficient and structure-preserving personalized image generation method. Background Art
[0002] Text-to-Image Personalization aims to generate images that meet the personalized needs of users according to the given text descriptions. With the rapid development of Generative Adversarial Networks (GANs) and Diffusion Models, significant progress has been made in text-to-image generation technology, which is widely applied in fields such as art creation, virtual reality, advertising design, and personalized recommendation. Especially in personalized image generation, how to generate personalized images based on users' specific preferences, styles, or object features has become a key problem to be solved urgently.
[0003] In addition, personalized text-to-image generation also faces challenges such as dataset bias, ambiguity of text expression, and instability of generation results. In order to make the generated images more accurately reflect users' personalized needs, researchers need to design more efficient network architectures and training methods, and introduce new technical frameworks to strengthen the understanding and control of personalized features by the generation model.
[0004] In recent years, many personalized text-to-image generation methods, such as DreamBooth, CustomDiffusion, ViCo, Mix-of-Show, Cones 2, MuDI, etc., usually ignore the appearance quality in personalized image generation. Although these methods have made certain progress in text-to-image generation, they are insufficient in ensuring the visual consistency and detail accuracy of images. In addition, the calculation processes of methods such as DreamMatcher and MasaCtrl are relatively complex, requiring recalculation of the self-attention maps of the model, and may damage the structure of the target image to a certain extent. Especially when performing personalized appearance enhancement, complex calculations and additional processing steps may lead to instability of the generation results, affecting the naturalness and authenticity of the images.
[0005] Therefore, how to improve the appearance quality of personalized text-to-image generation while maintaining computational efficiency and the structural stability of the generated images has become a major challenge in current research. Summary of the Invention
[0006] The object of the present invention is to provide an efficient and structure-preserving personalized image generation method. By improving the computational efficiency and enhancing the appearance features of personalized objects, this method improves the user's visual experience. By precisely controlling the personalized features during the generation process, the present invention can improve the detail performance and visual realism of the generated images while ensuring the image quality, solving the problems of insufficient appearance quality and low computational efficiency in the prior art.
[0007] To achieve the object of the present invention, the present invention provides an efficient and structure-preserving personalized image generation method, comprising the following steps:
[0008] Step 1: Collect images of specific personalized objects or people to form a training sample, and preprocess the training sample;
[0009] Step 2: Sample random noise, predict the noise according to the noise prediction network U-Net in the Diffusion network to generate a random image; construct a loss function between the preprocessed training sample and the generated image, and train the noise prediction network in the U-Net to make the generated image gradually approach the real training image;
[0010] Step 3: After training, optimize the neural network parameters to obtain a model that can generate personalized object images according to text after training;
[0011] Step 4: In order to efficiently improve the appearance quality of the generated personalized objects, select a corresponding training image as a condition to guide the trained model to generate more realistic and personalized feature images.
[0012] An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the above method are implemented.
[0013] A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the steps of the above method.
[0014] A computer program product, comprising computer program instructions, which when running on a computer, cause the computer to execute the steps of the above method.
[0015] Compared with the prior art, the significant improvements of the present invention are as follows: (1) The present invention proposes an efficient appearance enhancement guidance method, which effectively improves the calculation efficiency by performing dimensionality reduction on appearance features; combined with cosine loss and Gram loss, it can accurately enhance the appearance performance of the target image, enabling the appearance features to be optimized while retaining the original image structure; (2) The present invention innovatively proposes a technique for independently enhancing the appearance of an image without destroying the structure of the target image. By effectively decoupling the structure and appearance of the target image, the separation control between the two is successfully achieved, further improving the flexibility and accuracy of image enhancement; (3) The present invention conducts a comprehensive comparison with existing personalized algorithm models, covering multiple SOTA (State-of-the-Art) methods such as DreamBooth, CustomDiffusion, ViCo, Mix-of-Show, Cones 2, MuDI, MasaCtrl, etc.; the experimental results show that the technology proposed by the present invention is superior to existing methods in terms of appearance enhancement effect and calculation efficiency, and has better practicability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0017] Figure 1 is the overall architecture diagram of the present invention.
[0018] Figures 2 to 5 is the qualitative result comparison between the present invention and existing methods.
[0019] Figure 6 is the qualitative result analysis of the present invention on multiple objects. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] An efficient text-to-image personalized appearance enhancement method based on the Diffusion model of the present invention, combined with Figure 1 , includes the following steps:
[0022] Step 1: Collect specific personalized object or person images to form a training sample, and preprocess the training sample;
[0023] Step 1-1: Collect an image dataset in a specific field (such as pets, architecture, toys, etc.), and use the dataset as training samples.
[0024] Step 1-2: Preprocess the training samples, specifically perform data augmentation by means of image flipping, translation, etc., and at the same time adjust the image resolution to 512×512 by cropping.
[0025] Step 2: Generate random images by sampling random noise and predicting the noise according to the noise prediction network U-Net in the Diffusion network; construct a loss function between the preprocessed training samples and the generated images, and train the noise prediction network in U-Net to make the generated images gradually approximate the real training images.
[0026] Step 2-1: Randomly sample noise z from a Gaussian distribution t , where the time is t;
[0027] Step 2-2: Input the preprocessed training samples into the neural network, and implement the noise addition process by adding noise ∈ at each moment t ; then, train the U-Net noise prediction network and the text encoder (Text encoder, T e ), regard the text-to-image personalized generation task as a reconstruction task, and use the mean square error (MSE) as the loss function for model training; the specific training process is as follows:
[0028]
[0029] Among them, c is the text conditional input, ∈ θ () is the U-Net noise prediction network, w t is the weight represented by the loss, T e () represents the text encoder, represents calculating the expected average value of the variables during the training process.
[0030] Step 3: After training, optimize the neural network parameters to obtain a model that can generate personalized object images according to text after training.
[0031] Step 3-1: Save the model weights of the trained Diffusion in Step 2, including the noise prediction network U-Net (∈ θ ) and the text encoder T e ;
[0032] Step 3-2: Use the trained Diffusion model weights to generate images corresponding to personalized objects in the way of generating images from text, and this generation process can include different scenarios.
[0033] Step 3-3: Repeat the training process in Step 2 so that the model can generate different types of personalized objects, and save the model weights corresponding to each type of personalized object.
[0034] Step 4: To efficiently improve the appearance quality of the generated personalized objects, select a corresponding training image as a condition to guide the trained model to generate more realistic and personalized feature images.
[0035] Step 4-1: Collect images of different types of personalized objects, such as pets, toys, flowers, buildings, etc., and extract the appearance features V t :
[0036]
[0037] where V t is taken from the Value of the self-attention layer in the U-Net network, represents the sum of the appearance features of these personalized images, N represents the number of images of personalized objects;
[0038] Step 4-2: To improve the computational efficiency, reduce the dimensionality of the appearance features and construct a Global Semantic Appearance Space (GSAS); construct GSAS by calculating the mean and basis vectors, and the specific calculation process is as follows:
[0039]
[0040] where represents the mean of GSAS, represents the centered appearance features, U t represents the left singular matrix in the singular value decomposition, δ t represents the singular value matrix in the singular value decomposition, represents the transpose of the right singular matrix in the singular value decomposition, represents the matrix after dimensionality reduction by the transpose of the right singular matrix, k represents the dimension after dimensionality reduction, and finally the basis vectors of GSAS
[0041] Step 4-3: Take an image in the training samples as a reference image and input it into the trained model, and extract the appearance features of the reference image at each moment
[0042] Step 4-4: Sample random noise, use the trained model to generate personalized target images, and extract the appearance features at each moment
[0043] Step 4-5: Project the appearance features of the reference image and the appearance features of the personalized target image into the global semantic appearance space (GSAS) to achieve dimensionality reduction of the features. The specific process is as follows:
[0044]
[0045] Among them, represents the appearance features of the reference image after projection dimensionality reduction, represents the appearance features of the target image after projection dimensionality reduction;
[0046] Step 4-6: To more precisely locate the personalized object whose appearance needs to be enhanced, use the token mask of the cross-attention layer in U-Net to and further process the dimensionality-reduced semantic appearance features:
[0047]
[0048] Among them, represents the appearance features of the dimensionality-reduced reference image processed by the reference token mask , represents the appearance features of the dimensionality-reduced target image processed by the target token mask ;
[0049] Step 4-7: After reducing the dimensionality of the appearance features and improving the computational efficiency, use the features of the dimensionality-reduced reference image to guide the appearance features of the personalized target image
[0050]
[0051] so that the target image can enhance its appearance through the reference image during the generation process. This patent adopts two effective appearance guidance strategies to achieve this goal: cos where g and denotes the cosine loss between , denotes the Gram matrix of the reference image, gm denotes and denotes the loss of the corresponding Gram matrices of
[0052] Step 4-8: Combine g cos and g gmAs a loss function, it is used to guide the personalized target image during the denoising process and enhance its appearance. The specific process is as follows:
[0053]
[0054] where s is the weight of the unclassified guidance, and α cos and β gm are the weights of the cosine loss and the Gram loss respectively, represents the noise variable after being guided by the cosine loss and the Gram loss.
[0055] Based on the trained network model and the input reference image, the present invention systematically evaluates the personalized image generation results of the model and comprehensively compares them with multiple current state-of-the-art SOTA (State-of-the-Art) methods. Combining Figures 2 to 5 , Figure 2 and Figure 3 show the qualitative comparison results of the method of the present invention with the prior art, including Cones 2, Mix-of-Show, and DreamMatcher, on multiple tasks. Figure 4 and Figure 5 show the qualitative performance after integrating the present invention into existing mainstream personalized text-to-image generation frameworks such as DreamBooth and CustomDiffusion. The experimental results show that the present invention significantly improves the deficiencies of the existing methods in terms of appearance quality in the text-to-image personalized generation task, enhances the detail performance and visual realism of the images, and makes the generated results highly similar to the reference image in appearance.
[0056] Figure 6 show the appearance optimization results of the present invention on multiple personalized objects. Experiments show that the present invention is not only applicable to the appearance optimization of a single personalized object, but also can effectively optimize the appearance features of multiple personalized objects, verifying the generality and applicability of the method.
[0057] Table 1 shows the evaluation results of the correlation between the generated target image and the reference image and the text of the present invention and the existing methods Cones 2, Mix-of-Show, DreamMatcher, and MuDI.
[0058] Table 1 Quantitative comparison of the generated results of the present invention and the existing methods
[0059]
[0060] Specifically, quantitative indicators such as I_DINO, I_CLIP, and T_CLIP are used for comparative analysis. The experimental data shows that the present invention significantly outperforms the existing SOTA methods in both of the two core indicators of I_DINO and T_CLIP, fully demonstrating the technical advantages of the present invention in terms of image generation quality and text relevance.
[0061] Table 2 shows the comparison results of the inference time between the present invention and the existing methods DreamMatcher, ViCo, and MasaCtrl.
[0062] Table 2 Comparison of the calculation time between the present invention and the existing methods
[0063]
[0064] Experiments show that the present invention reduces the dimensionality of features by constructing a global semantic appearance space, and directly guides the diffusion process by combining cosine loss and Gram loss without recalculating the feature attention map, thereby significantly reducing the complexity of inference calculation. The results show that compared with the existing methods, the present invention has a significant optimization effect in terms of inference time, making the calculation more efficient and improving the feasibility of practical applications.
[0065] In summary, the present invention not only makes a breakthrough in the appearance quality of personalized image generation, but also is superior to the existing SOTA methods in terms of computational efficiency, and has higher practical value and promotion potential.
[0066] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0067] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An efficient and structure-preserving personalized image generation method, which generally includes the following steps: Step 1: Collect specific personalized object or person images to form a training sample, and preprocess the training sample; Step 2: Sample random noise, predict the noise according to the noise prediction network U-Net in the Diffusion network to generate a random image; construct a loss function between the preprocessed training sample and the generated image, and train the noise prediction network in U-Net to make the generated image gradually approach the real training image; Step 3: After training, optimize the neural network parameters to obtain a model that can generate personalized object images according to text after training; Step 4: Select a corresponding training image as a condition to guide the trained model to generate a more real and personalized feature image.
2. An efficient and structure-fidelity personalized image generation method according to claim 1, characterized in that The specific steps of Step 1 include the following steps: Step 1-1: Collect an image dataset in a set field, including pets, buildings, toys, and use the dataset as a training sample; Step 1-2: Preprocess the training sample, specifically perform data augmentation through image flipping and translation, and at the same time adjust the image resolution to 512×512 by cropping.
3. An efficient and structure-preserving personalized image generation method according to claim 1, characterized in that The specific steps of Step 2 include the following steps: Step 2-1: Randomly sample noise z from a Gaussian distribution t , where the time is t; Step 2-2: Input the preprocessed training samples into the neural network, and add noise ∈ at each moment t to implement the noise addition process; then, train the U-Net noise prediction network and the text encoder T e , regard the text-to-image personalized generation task as a reconstruction task, and use the mean square error as the loss function for model training; the specific training process is as follows: where c is the text condition input, ∈ θ () is the U-Net noise prediction network, w t is the weight represented by the loss, T e () represents the text encoder, denotes the calculation of the expected average value of the variables during the training process.
4. An efficient and structure-preserving personalized image generation method according to claim 3, characterized in that The specific steps of Step 3 include the following steps: Step 3-1: Save the model weights of the Diffusion trained in Step 2, including the noise prediction network U-Net (∈ θ ) and the text encoder T e ; Step 3-2: Use the trained Diffusion model weights to generate an image corresponding to the personalized object in the way of generating an image from text, and this generation process includes different scenarios; Step 3-3: Repeat the training process in Step 2 so that the model can generate different types of personalized objects, and save the model weights corresponding to each personalized object.
5. An efficient and structure-preserving personalized image generation method according to claim 4, characterized in that, The specific steps of Step 4 include the following steps: Step 4-1: Collect personalized object images of different types, including pets, toys, flowers, and buildings, and extract the appearance features V in these personalized images t : Among them, N represents the number of images of personalized objects, and V t Value taken from the self-attention layer in the U-Net network, represents the sum of the appearance features of these personalized images; Step 4-2: Reduce the dimension of the appearance features to construct a global appearance semantic space GSAS; construct GSAS by calculating the mean value and basis vectors, and the specific calculation process is as follows: Among them, represents the mean value of GSAS, represents the decentralized appearance feature, U t represents the left singular matrix in the singular value decomposition, δ t represents the singular value matrix in the singular value decomposition, represents the transpose of the right singular matrix in the singular value decomposition, represents the matrix after dimensionality reduction by the transpose of the right singular matrix. k represents the dimension after dimensionality reduction, and finally the basis vectors of GSAS are obtained Step 4-3: Use an image in the training samples as a reference image and input it into the trained model to extract the appearance features of the reference image at each moment. Step 4-4: Sample random noise, generate a personalized target image using the trained model, and extract the appearance features at each moment Step 4-5: Project the appearance features of the reference image and the appearance features of the target image onto the global semantic space GSAS to achieve dimensionality reduction of the features. The specific process is as follows: Among them, represents the appearance feature of the reference image after projection dimensionality reduction, represents the appearance feature of the target image after projection dimensionality reduction; Steps 4-6: To locate the personalized object whose appearance needs to be enhanced, the token mask of the cross-attention layer in the U-Net is used to further process the downsampled semantic appearance features and as follows: Among them, represents the appearance feature of the dimensionality-reduced reference image processed by the reference token mask ; represents the appearance feature of the dimensionality-reduced target image processed by the target token mask ; Step 4-7: Use the features after dimensionality reduction of the reference image to guide the appearance features of the personalized target image such that the target image can enhance its appearance through the reference image during the generation process; two appearance guidance strategies are adopted to achieve this goal: where g cos represents and the cosine loss between, represents the Gram matrix of the reference image, represents the Gram matrix of the target image, and g gm represents and the loss of the corresponding Gram matrix; Step 4-8: Take g cos and g gm as the loss function to guide the personalized target image during the denoising process and enhance its appearance. The specific process is as follows: where s is the weight of the unclassified guidance, α cos and β gm are the weights of the cosine loss and the Gram loss respectively, represents the noise variable guided by the cosine loss and the Gram loss.
6. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 5.
7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the method according to any one of claims 1 to 5.
8. A computer program product, including computer program instructions, when the computer program instructions run on a computer, cause the computer to execute the method according to any one of claims 1 to 5.