A new object AI generation method based on image and text fusion
Through the adaptive text image harmony method (ATIH), using the scale factor α and injection step i to balance text and image features in cross attention and self-attention, the problem of imbalance in image and text fusion in the prior art is solved, and high-quality and flexible new object image generation is achieved.
Patent Information
- Application Number
- CN202411265710.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-09-10
AI Technical Summary
Existing diffusion models are prone to imbalance problems in the fusion process of images and text, resulting in the generated images being biased towards text or image features, and it is impossible to effectively achieve harmonious integration of object text and images.
An adaptive text image harmony method (ATIH) is proposed, by introducing a scale factor α and injection step i, the text and image features are balanced in cross attention and self-attention, and a novel similarity loss function and golden segmentation search algorithm are used to adaptively adjust the parameters, ensuring that the similarity between the generated image and the input text/image is optimally balanced.
The balanced fusion of object text and image is realized. The generated new object images not only have high similarity, but also retain the fidelity and editability of the original image, which significantly improves the quality and flexibility of image generation.
Smart Images

Figure CN119338947B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual design, and in particular to an AI generation method for a new object by integrating images and texts. Background Art
[0002] Image synthesis using diffusion models (such as Stable Diffusion, SDXL, and DALL·E3), especially text-based or image-based synthesis, has attracted widespread attention in recent years. These models are highly regarded for their excellent generation capabilities and practical applications (including editing and reversal). Many methods focus on object-centric diffusion, manipulating objects in images through text descriptions, such as combining, adding, deleting, replacing, moving, and adjusting size, shape, action, and pose. However, the object synthesis task we study aims to create a new object image by combining object text and object image. For example, combining "kingfisher" (image) and "hound" (text) to generate a new harmonious hound-shaped kingfisher object.
[0003] To achieve object-text-image fusion, most diffusion models (such as SDXL-Turbo) usually use cross attention to integrate the input text and image. However, cross attention often leads to unbalanced results. When inputting "Axolotl" (image) and "Toucan" (text), SDXL-Turbo only generates an image of a toucan, showing a bias towards the toucan text. In contrast, when inputting "rooster" (image) and "iron" (text), it generates an image that is very similar to the original rooster image, indicating a bias towards image features. These observations suggest that during the diffusion process, text (or image) features often suppress the influence of image (or text) features, resulting in fusion failure.
[0004] To alleviate the image degradation problem, plug-and-play methods can inject guiding image features into self-attention. However, even with the best reversal editing methods such as PnPinv, we still observe similar imbalance. This raises an important question: how to balance the integration of object text and image? To address this issue, we propose an adaptive text-image harmony method (ATIH) for novel object synthesis. First, during the reversal diffusion process, we introduce a scaling factor α to balance text and image features in cross-attention, and introduce an injection step i to preserve image information in self-attention for adaptive adjustment. The reversal noise maps follow the statistical properties of uncorrelated Gaussian white noise, which reduces editability. However, they are more suitable for approximating feedforward noise maps, which enhances fidelity. To better integrate object text and image, we consider the sampling noise as a parameter to design a balanced loss function that strikes a balance between reconstruction and Gaussian white noise approximation, ensuring optimal editability and fidelity of object images. We propose a novel similarity loss that takes i and α into account. This loss function not only maximizes the similarity between the generated object image and the input text / image, but also balances these two similarities to achieve a harmonious integration of text and image. In addition, we adopt the golden section search algorithm to quickly find the optimal parameter α. Therefore, our ATIH method is able to generate novel object combinations. Summary of the invention
[0005] The purpose of the present invention is to provide a new object AI generation method for image-text fusion, balance the fidelity and editability of implicit features in the reverse diffusion process, and achieve balanced fusion of object text and image by adaptively adjusting the scaling factor and injection step; introduce a novel similarity scoring function, which combines the scaling factor and injection step to balance and maximize the similarity between the generated image and the input text / image, thereby achieving a harmonious fusion of text and image.
[0006] The technical solution to achieve the purpose of the present invention is: a new object AI generation method of image-text fusion, which generates a new object image by fusing the original object image and the object text, and comprises the following steps:
[0007] Step 1: Input image / text encoder ε(·), object image O I , object text O T , empty text O N , inversion step number T, self-attention injection step length i, sampling noise ∈ t , the pre-trained U-Net model ∈ θ (·), scale factor α;
[0008] Step 2: Feature extraction: use encoder ε(·) to extract OI The implicit feature z0=ε(O I ), empty text O N Text encoding τ N =ε(O N ), and object text O T The text encoding τ=ε(O T );
[0009] Step 3, through the formula: right Perform inversion to obtain in represents the final noise implicit feature of the t-th step obtained by inversion, is the variance of the t-th step defined by the sampler;
[0010] Step 4: Balance fidelity and editability. As the starting point of the noise implicit feature, through the formula
[0011] Perform denoising and use l2 norm loss optimization ∈ t+1 Make Approximate inversion process Improve fidelity and use KL divergence loss to measure ∈ t+1 The distance from the Gaussian distribution, dealing with both fidelity and editability, to obtain the final optimized sequence where t∈{0,…,T}; It is a pre-trained U network structure, which includes a self-attention module. is in∈ θ (·) The derived query, key, and value feature vectors, d is the dimension of the value feature vector;
[0012] Step 5: Define the fusion denoising process. As the starting point of the noise implicit feature, through the formula Denoising is performed to obtain the fused image O(α, i), where α is the scale factor for scaling the cross-attention layer and i is the number of self-attention injection steps;
[0013] Step 6, ∈ θ (z t , t, τ, α, i) is a pre-trained U network structure, which contains the injection self-attention layer and the scale cross-attention layer; the self-attention injection step i is introduced, 0≤i≤T, and the injection self-attention layer is defined as in
[0014]
[0015] in is in∈ θ (·) from z t The derived query and key feature vectors, d is the dimension of the value vector, During normal denoising Self-attention map; a scale factor α∈[0, 2] is introduced to control the fusion process, and the scale cross attention layer is defined as in is the inquiry feature, is the text encoding τ=ε(O T ) derived query and key feature vectors, d is the dimension of the value vector;
[0016] Step 7: Define the similarity distance calculation as follows: Image O I The similarity between the image O(α, i) generated by fusion in step 5 is expressed as I sim (α, i) = d(O I ,O(α,i)); text O T The similarity between the image O(α, i) generated by fusion in step 5 is expressed as T sim (α, i) = d(O T , O(α, i)), where d(·, ·) is the similarity distance function of image / text;
[0017] Step 8, using the definition in step 7, when α=1, I sim (i) = d(O I , O(α=1,i)); Adaptively adjust the number of injection steps of self-attention injection in step 6: Initialization The update method is as follows:
[0018]
[0019] in and is a soft threshold set according to observation; after updating, i = i*;
[0020] Step 9: Based on the i=i* obtained in step 8, use the definition in step 7 to simplify to, I sim (α) = d(O I , O(α, i=i * )) and T sim (α) = d(O T , O(α, i=i * )), where d(·,·) represents the similarity distance between text / image and image; the fusion score function F(α) is defined as follows:
[0021]
[0022] Where β is the weight factor, and parameter k is used to alleviate the high I caused by the difference between text and image modalities. sim (α) and low T sim The scales of (α) are not consistent, thus ensuring that their scales are balanced;
[0023] Step 10: Using the fusion score function defined in step 9 as the objective function, and using the golden section search method based on the fusion denoising process defined in step 5 to determine the optimal value of α, α=α*;
[0024] Step 11: Find the optimal i* and α* through steps 8 and 10, and input the Wensheng graph model to generate the optimal fusion visual image O(α*, i*).
[0025] Furthermore, the optimization method for improving fidelity in step 4 first sets ∈t as the optimization parameter, and then
[0026]
[0027] Optimize ∈t so that the denoising process Approximation to the normal inversion process For fidelity.
[0028] Furthermore, the editability measurement method in step 4 defines the distance between ∈t and the Gaussian distribution to measure editability.
[0029]
[0030] When ∈ t The closer the distance to the Gaussian distribution, the higher the editability, and Does not participate in gradient calculation.
[0031] Further, the balance between fidelity and editability in step 4 defines the factor Factor to weigh the relationship between fidelity and editability, and set the final balance loss for
[0032]
[0033] λ represents balance and The weight of .
[0034] Furthermore, λ=125 is set.
[0035] Furthermore, the method of controlling the object fusion process in step 6 is to control the object image OI , target object text O T , which are fused in the cross-attention layer of the Wensheng graph model
[0036]
[0037] in is the query feature output from the self-attention layer, and From O T The key and value features obtained by embedding τ of the text; the object image O is adjusted by the scaling factor α I With the target object text O T The degree of fusion is as follows:
[0038]
[0039] α∈[0, 2] controls the fusion process.
[0040] Furthermore, the adaptive self-attention injection step adjustment method in step 8 makes the object fusion process smoother through adaptive step adjustment. and are set to 0.45 and 0.85 respectively; I sim (i) = d(O I , O(α=1,i)) means that when α=1, the generated image O(α=1,i) is fused with the original image O under the condition of self-attention injection i steps. I The similarity.
[0041] Furthermore, in the fusion score function in step 9, k=2.3 and β=1 are set.
[0042] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned new object AI generation method of image-text fusion is implemented.
[0043] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned new object AI generation method by integrating image and text.
[0044] Compared with the prior art, the present invention has the following significant advantages: (1) The present invention can operate based on objects that exist in reality and perform fusion starting from the objects that exist in reality. It is not limited to virtual or synthetic data, but can start from real object images. This method is closer to the real world, making the generated images more realistic and practical; (2) The present invention does not require additional training or fine-tuning of the generative model, and can directly use the generative model as a black box to generate images with text fusion; (3) Since the present invention allows users to start from real objects and combine them with text descriptions for fusion, it provides users with extremely high flexibility and creative space; users can freely explore the possibilities of different object and text combinations to create image content that is both novel and meets their needs.
[0045] The present invention is described in further detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic diagram of the process of the present invention.
[0047] Figure 2 It is a framework diagram of the NOS of the present invention.
[0048] Figure 3 These are some different fusion result diagrams of the present invention.
[0049] Figure 4 This is a comparison chart of the effects of the present invention on image object fusion and other existing image editing methods.
[0050] Figure 5 This is a comparison chart of the effects of the present invention on image object fusion and other existing object fusion methods.
[0051] Figure 6 This is an ablation experiment diagram of the effects of each module in the present invention.
[0052] Figure 7 This is a graph showing changes in similarity during the object fusion process of the present invention.
[0053] Figure 8 This is a graph corresponding to the similarity threshold in the object fusion process of the present invention. DETAILED DESCRIPTION
[0054] The present invention proposes a new object AI generation method of image-text fusion, which aims to combine object text with object image to generate different object images. The input of the present invention is an original object image and an object text name to be fused, and the final output is a new object image in which the two objects are smoothly fused. The fused object not only has the layout or posture of the original object image, but also contains the feature information of the text object. This method is called Novel Object Synthesis (NOS), which consists of an image editability balance module (Balance Editing, BE) and an adaptive text-image harmony module (Adaptive Text-Image Harmony, ATIH). The purpose of the BE module is to achieve better fusion of object text and object image by balancing the fidelity and editability of the original object image during the image inversion process. We designed a balanced loss function with a noise parameter to ensure the editability and fidelity of the object image. The purpose of the ATIH module is to make the fusion process of the object image and the object text smoother and find a harmonious fused image. We first introduce a scale factor to balance text and image features in the cross-attention module, and an injection step factor to preserve the information of the original object image in the self-attention module during the text-image inversion diffusion process. Secondly, in order to adaptively adjust these parameters, we propose a novel similarity score function that not only maximizes the similarity between the generated object image and the input text / image, but also balances these similarities to coordinate the fusion of text and image.
[0055] The AI generation method of a new object by integrating image and text of the present invention specifically comprises the following steps:
[0056] Step 1: Input image / text encoder ε(·), object image O I , object text O T , empty text O N , inversion step number T, self-attention injection step length i, sampling noise ∈ t , the pre-trained U-Net model ∈ θ (·), scaling factor α.
[0057] Step 2: Feature extraction: use encoder ε(·) to extract O I The implicit feature z0=ε(O I ), empty text O N Text encoding τ N =ε(O N ), and object text O T The text encoding τ=ε(O T ).
[0058] Step 3, through the formula: right Perform inversion to obtain in represents the final noise implicit feature of the t-th step obtained by inversion, is the t-th step variance defined by the sampler.
[0059] Step 4: Balance fidelity and editability. As the starting point of the noise implicit feature, through the formula Perform denoising and use l2 norm loss optimization ∈ t+1 Make Approximate inversion process Improve the fidelity and propose KL divergence loss to measure ∈ t+1 The distance from the Gaussian distribution, dealing with both fidelity and editability, to obtain the final optimized sequence where t∈{0,…,T}. It is a pre-trained U network structure, which includes a self-attention module. is in∈ θ (·) The derived query, key, and value feature vectors, where d is the dimension of the value feature vector.
[0060] Step 5: Define the fusion denoising process. As the starting point of the noise implicit feature, through the formula Denoising is performed to obtain the fused image O(α, i), where α is the scale factor for scaling the cross-attention layer and i is the number of self-attention injection steps.
[0061] Step 6, ∈ θ (z t , t, τ, α, i) is a pre-trained U network structure, which contains the injection self-attention layer and the scale cross-attention layer. The self-attention injection step i (0≤i≤T) is introduced, and the injection self-attention layer is defined as in
[0062]
[0063] in is in∈ θ (·) from z t The derived query and key feature vectors, d is the dimension of the value vector, During normal denoising A scale factor α∈[0, 2] is introduced to control the fusion process, and the scale cross attention layer is defined as in is the query feature (injected into the output of the self-attention module), is the text encoding τ=ε(O T ) The derived query and key feature vectors, d is the dimension of the value vector
[0064] Step 7: Define the similarity distance calculation as follows: Image O I The similarity between the image O(α, i) generated by fusion in step 5 is expressed as I sim (α, i) = d(O I ,O(α,i)); text O T The similarity between the image O(α, i) generated by fusion in step 5 is expressed as T sim (α, i) = d(O T , O(α, i)), where d(·, ·) is the similarity distance function of image / text.
[0065] Step 8: To simplify the description, use the definition in step 7. When α = 1, I sim (i) = d(O I , O(α=1,i)). Adaptively adjust the number of injection steps of self-attention injection in step 6: Initialization The update method is as follows:
[0066]
[0067] in and is a soft threshold set according to observation. After updating, we get i=i*.
[0068] Step 9: Based on the i=i* obtained in step 8, use the definition in step 7 to simplify to, I sim (α) = d(O I , O(α, i=i * )) and T sim (α) = d(O T , O(α, i=i * )). We define the fusion score function F(α) as follows:
[0069]
[0070] Step 10: Using the fusion score function defined in step 9 as the objective function, and based on the fusion denoising process defined in step 5, use the golden section search method to determine the optimal value of α, α=α*.
[0071] Step 11: Find the optimal i* and α* through steps 8 and 10, and input the Wensheng graph model to generate the optimal fusion visual image O(α*, i*).
[0072] Preferably, in step 1, the pre-trained large model image generator and its text encoder are directly used, thereby eliminating the most time-consuming and computationally demanding training process;
[0073] Preferably, step 3 is a general text graph model sampler S(.), so that the method can be used in any text graph model.
[0074] Preferably, the method of step 4 that does not require optimization parameters of the neural network can greatly reduce the time complexity and reduce the amount of data transferred by the gradient, thereby reducing the calculation time. On the other hand, it makes the fidelity and editability controllable.
[0075] Preferably, the novel in step 4 is obtained by comparing the denoising process Compared with the normal inversion process The difference between The ratio of the distance to the Gaussian distribution indirectly evaluates editability and fidelity. This method cleverly achieves a balance, that is, while maintaining image fidelity, it also ensures that the image has a certain degree of editability.
[0076] Preferably, in step 6, the object fusion process is controlled. Since the traditional model lacks precise control over the object fusion process, we introduce a scaling factor α to adjust the object image O I With the target object text O T This method makes the fusion between two different modalities highly controllable, so that the fusion result can be fine-tuned as needed. Through the scale factor α, the user can effectively control the relative contribution of image and text information in the fusion process, and achieve a continuous transition from fully retaining the original image features to fully following the text description. This flexible control mechanism not only improves the interactivity and user experience of the fusion process, but also greatly expands the application scope of the text-based graph model, enabling it to better adapt to diverse generation needs.
[0077] Preferably, the adaptive self-attention injection step adjustment strategy in step 8 has the core advantage of being able to achieve a balance between fidelity and editability. By dynamically adjusting the number of self-attention injection steps, the method can ensure a smooth image fusion process and avoid sudden changes or unnatural transitions during the fusion process. This adaptive adjustment mechanism allows the system to flexibly control the intensity and timing of self-attention according to the image content and fusion requirements, thereby gradually introducing new text-guided modifications while maintaining the original features of the image.
[0078] Preferably, the new fusion scoring function in step 9 is designed to maximize the similarity between the fused image and the original object image and the target object text while ensuring the harmony of the object fusion. The key to this scoring function is that it can effectively determine the degree of object fusion, thereby providing a quantitative evaluation standard for the fusion process.
[0079] Preferably, the golden section search method is used in step 10 to determine the optimal value of α. By using the golden section algorithm to optimize this scoring function, we can quickly find the optimal scale factor α, thereby maximizing the fusion scoring function. This combination of fusion scoring function and optimization algorithm not only ensures the harmony and naturalness of the fusion process, but also improves the automation and user-friendliness of the entire system, making it faster and more reliable to generate high-quality fused images.
[0080] The present invention fuses object images and object texts, and the final output is a stunning and novel image of the two objects. These fused images are not just a simple copy or splicing of the original objects, but a brand new creation. They retain the layout structure of the original object images, while harmoniously integrating the characteristics of the two objects to form a unique visual presentation. The fused objects show characteristics that go beyond the original object types. They are an artistic transformation of the original input, which not only retains their respective identity characteristics, but also creates a new visual experience on this basis.
[0081] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0082] Example
[0083] like Figure 1 As shown in the figure, a new object AI generation method of image-text fusion is given an arbitrary text-based image model sampler S(·), image text / encoder ε(·), object image O I , target object text O T , empty text O N , inversion step number T, self-attention injection step length i, sampling noise ∈ t , the pre-trained U-Net model ∈ θ (·), the scale factor α, and the encoder ε(·) is used to extract O I The implicit feature z0=ε(O I ), empty text O N Text encoding τ N =ε(O N ), and the target text O T The text encoding τ=ε(O T );Use sampler S(·) to obtain by inverting z0 where z Trepresents the noise implicit feature of the Tth step obtained by inversion. Set ∈ t To optimize the parameters, we use the 2-norm loss optimization ∈ t In the denoising process Approximation to the normal inversion process Improve fidelity and propose KL divergence loss to measure ∈ t-1 The distance from the Gaussian distribution, dealing with both fidelity and editability, ultimately resulting in an optimized noise collection Next use In the fusion denoising process, the number of self-attention injection steps is adjusted adaptively to initialize the number of self-attention steps Calculate the similarity I between the generated image O(i) and the original image sim (i) = d(O I , O(i)), adjust i so that After getting the optimal number of self-attention injection steps i * Finally, the scale factor α is used to control O I With O T The fusion between them is performed, and the golden section algorithm is set to find the optimal score α by fusing the score function. The optimal i and α are input into the Wensheng graph model to generate the optimal fusion visual image. The specific steps are as follows:
[0084] Step 1: Given any text-based graph model sampler S(·), image text / encoder ε(·), object image O I , target object text O T , empty text O N , inversion step number T, self-attention injection step length i, sampling noise ∈ t , the pre-trained U-Net model ∈ θ (·), scale factor α. In this embodiment, the text graph model is the Stable-Diffusion xl turbo model, the text encoder is the CLIP model using ViT-bigG / 14, the image encoder is the Dinov2 model, the fused object text prompt word is "[word]", and the words in [word] are object categories.
[0085] Step 2: Feature extraction: use encoder ε(·) to extract O I The implicit feature z0=ε(O I ), empty text O N Text encoding τ N =ε(O N ), and object text O T The text encoding τ=ε(O T ), where T, T N ∈R 77×1024 , z0∈R4×64×64 .
[0086] Step 3: Use the sampler S(·) by the formula: Inverting z0 yields where z T represents the noise implicit feature of the final T-th step obtained by inversion, is the t-th step variance defined by the sampler.
[0087] Step 4: Figure 2 As shown, a balance between fidelity and editability is achieved. T As the starting point of the noise implicit feature, use the sampler S(·) through the formula To denoise, set ∈t as the optimization parameter during the denoising process.
[0088]
[0089] Optimize t Make Approximate inversion process Improve fidelity. Also define ∈ t The distance from the Gaussian distribution is used to measure editability. The formula is as follows
[0090]
[0091] When ∈ t The closer the distance to the Gaussian distribution, the higher the editability. And Does not participate in the gradient calculation. We define the factor Factor to measure the relationship between fidelity and editability. And set the final balance loss to
[0092]
[0093] λ represents balance and In this example, we set λ = 125 based on experimental observations to obtain the final optimized
[0094] Step 5: Define the fusion denoising process, with z T As the starting point of the noise implicit feature, through the formula Denoising is performed to obtain the fused image (α,i), where α is the scale factor for scaling the cross-attention layer and i is the number of self-attention injection steps.
[0095] Step 6, ∈ θ (z t, t, τ, α, i) is a pre-trained U network structure, which contains the injection self-attention layer and the scale cross-attention layer. The self-attention injection step i (0≤i≤T) is introduced, and the injection self-attention layer is defined as in
[0096]
[0097] in is in∈ θ (·) from z t The derived query and key feature vectors, d is the dimension of the value vector, During normal denoising A scale factor α∈[0, 2] is introduced to control the fusion process, and the scale cross attention layer is defined as in is the query feature (injected into the output of the self-attention module), is the text encoding τ=ε(O T ) are the derived query and key feature vectors, and d is the dimension of the value vector.
[0098] Step 7: Define the similarity distance calculation as follows: Image O I The similarity between the image O(α, i) generated by fusion in step 5 is expressed as I sim (α, i) = d(O I ,O(α,i)), where d(O T , O(α, i)) is the co-distinction distance function of its DINO feature; text O T The similarity between the image O(α, i) generated by fusion in step 5 is expressed as T sim (α, i) = d(O T ,O(α,i)), where d(O T , O(α, i)) is the residual distance function of its CLIP feature.
[0099] Step 8: To simplify the description, use the definition in step 7. When α = 1, I sim (i) = d(O I , O(α=1,i)). Adaptively adjust the number of injection steps of self-attention injection in step 6: Initialization The update method is as follows:
[0100]
[0101] in and In this example, the soft thresholds obtained by human observation are set to 0.45 and 0.85 respectively. The object fusion process can be made smoother by adjusting the number of adaptive steps. After updating, i=i* is obtained.
[0102] Step 9: Based on the i=i* obtained in step 8, use the definition in step 7 to simplify to, I sim (α) = d(O I , O(α, i=i * )) and T sim (α) = d(O T , O(α, i=i * )). We define the fusion score function F(α) as follows:
[0103]
[0104] Where β is the weight factor, and the parameter k is introduced to alleviate the high I caused by the difference between text and image modalities. sim (α) and low T sim The scales of (α) are not consistent, so as to ensure their scale balance. In this example, through experimental observation, we set k=2.3 and β=1.
[0105] Step 10: Use the golden section search method to determine the optimal value of α. The key steps of the golden section search algorithm are summarized as follows:
[0106]
[0107] Where φ (about 1.618) is the golden ratio, and a and b are the current search ranges for α. During each iteration, we compare F(α1) and F(α2) and adjust the search ranges accordingly:
[0108] if F(α1)>F(α2)thenb=α2else a=α1.
[0109] This process continues until the length of the search interval |ba| is less than a predefined difference, indicating convergence to a local maximum.
[0110] Step 11: Find the optimal i and α through steps 7 and 10, and input the text graph model to generate the optimal fusion visual image.
[0111] Table 1 User survey and comparison of image editing methods.
[0112] method Ours MasaCtr InstructPix2Pix InfEdit Vote↑ 422 16 48 84
[0113] Table 2 Comparison of user survey and fusion methods.
[0114] method Ours MagicMix ConceptLab Vote↑ 453 47 70
[0115] Tables 1 and 2 show the results of the user studies. We conducted two user studies to evaluate how our model compares with image editing methods and fusion methods in terms of fusion image harmony and novelty. Each participant evaluated 6 image editing sets and 6 fusion sets. These studies received a total of 570 votes from 95 participants. Our method received the highest score in both studies, receiving 74.03% and 79.47% of the total votes, respectively. Among the image editing methods, InfEdit received 14.7% of the votes for its excellent editing performance, while InstructPix2Pix and MasaCtrl only received 8% and 2.8%, respectively. In the fusion category, ConceptLab received 12.28% of the votes, while MagicMix received 8%.
[0116] Table 3 Comparison of user survey and fusion methods.
[0117] method Ours MagicMix InfEdit MasaCtrl InstructPix2Pix Dino-I↑ 0.756 0.587 0.817 0.815 0.384 Clip-T 0.296 0.328 0.255 0.234 0.394 AES↑ 6.124 5.786 6.080 5.684 5.881 HPS↑ 0.383 0.373 0.367 0.343 0.375 Fscore↑ 1.362 1.174 1.173 1.077 0.768 Bsim↓ 0.075 0.167 0.230 0.277 0.522
[0118] Table 3 shows the quantitative comparison of different methods in the object fusion task. We comprehensively evaluate the methods by using aesthetic score (AES), CLIP text-image similarity (CLIP-T), Dinov2 image similarity (Dino-I), human preference score (HPS), scoring function score (Fscore) and balanced similarity (Bsim). The results show that the proposed method performs well in text-image fusion, especially in improving visual appeal and meeting human preferences. In addition, our method achieves the best results in both scoring function score and balanced similarity, indicating that our method can achieve more harmonious object fusion.
[0119] Figure 3 The fusion image generation effect of the present invention is demonstrated, which embodies the novelty. The present invention uses the central image and the text input around it to generate these combined object images, such as the glass jar (image) and porcupine (text) in the left image, and the horse (image) and bald eagle (text) in the right image.
[0120] Figure 4The comparison chart of the image object fusion effect of the present invention and other existing image editing methods is shown. It is observed that MasaCtrl and InfEdit can usually better preserve the details of the original image during the editing process, while InstructPix2Pix is more inclined to modify the image significantly according to the text description. Secondly, different methods show different degrees of distortion when fusing two objects, and the method of the present invention performs better in maintaining image harmony and high quality. Thirdly, the method of the present invention has obvious advantages in enhancing the editability of images, and can effectively fuse two objects harmoniously, greatly improving the editing effect and operability.
[0121] Figure 5 The comparison chart of the effects of the present invention on image object fusion and other existing object fusion methods is shown. It is observed that MagicMix and ConceptLab tend to favor one category in object synthesis, while the present invention method achieves a more harmonious balance between features. In addition, the fused images of MagicMix often have insufficient smooth features, for example, the facial features of the rabbit almost disappear in the fusion of the rabbit and penguin. In contrast, the present invention method seamlessly fuses the facial features of the penguin and the rabbit, retaining the main features of each.
[0122] Figure 6 The ablation experiment diagrams showing the effects of each module in the present invention are shown. It is observed that PnPinv for direct inversion and fast editing leads to some distortion and blurring. Compared with PnPinv, the balanced loss significantly improves image fidelity, improving details, textures, and editability. Adaptive injection enables smooth transitions between objects in the image. Without this injection, the transition is too abrupt and lacks a seamless fusion process. Finally, adaptive selection of α achieves a balanced image, harmoniously blending the original and target features.
[0123] Figure 7 The graph of the change of similarity in the object fusion process of the present invention is shown. We use Stable-Diffusion xl turbo as the basic model, set the range of i to [0,4], and for each i value, α is iterated from 0 to 2.2 with a step size of 0.02 to observe the change of the fused image. The average experimental results produce two smooth curves. As α increases, the image similarity decreases and the text similarity increases. Based on human observation, the optimal range of k is determined to be between [2.1,2.7]. And based on human observation, we set the value of k to 2.3.
[0124] Figure 8The corresponding graph of the similarity threshold in the object fusion process of the present invention is shown. We visualized several specific node images generated during the change of different α factor values. When the similarity between the image and the original image exceeds 0.85, the images become too similar. For example, in the dog-zebra fusion experiment, the texture of the dog remains basically unchanged, and the characteristics of the zebra cannot be seen. In contrast, when the image similarity is lower than 0.45, the image is too consistent with the text description. In this case, the entire head of the image becomes a zebra, representing an over-transformation phenomenon. Based on these observations, we set the minimum similarity threshold Set to 0.45, the maximum similarity threshold Set to 0.85. This range helps us achieve a good balance between preserving the original image information and fusing text features.
[0125] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A new object AI generation method for image-text fusion, characterized in that: According to the original object image and the object text, a new object image is generated by fusion. The method includes the following steps: Step 1: Input image / text encoder ε(·), object image O I , object text O T , empty text O N , inversion step number T, self-attention injection step length i, sampling noise ∈ t , the pre-trained U-Net model ∈ θ (·), scale factor α; Step 2: Feature extraction: use encoder ε(·) to extract O I The implicit feature z0=ε(O I ), empty text O N Text encoding τ N =ε(O N ), and object text O T The text encoding τ=ε(O T ); Step 3, through the formula: right Perform inversion to obtain in represents the final noise implicit feature of the t-th step obtained by inversion, is the variance of the t-th step defined by the sampler; Step 4: Balance fidelity and editability. As the starting point of the noise implicit feature, through the formula Perform denoising and use l2 norm loss optimization ∈ t+1 Make Approximate inversion process Improve fidelity and use KL divergence loss to measure ∈ t+1 The distance from the Gaussian distribution, dealing with both fidelity and editability, to obtain the final optimized sequence where t∈{0,…,T}; It is a pre-trained U network structure, which includes a self-attention module. is in∈ θ (·) The derived query, key, and value feature vectors, d is the dimension of the value feature vector; Step 5: Define the fusion denoising process. As the starting point of the noise implicit feature, through the formula Denoising is performed to obtain the fused image O(α,i), where α is the scale factor for scaling the cross-attention layer and i is the number of self-attention injection steps; Step 6, ∈ θ (z t ,t,τ,α,i) is a pre-trained U network structure, which contains the injection self-attention layer and the scale cross-attention layer; the self-attention injection step i is introduced, 0≤i≤T, and the injection self-attention layer is defined as in in is in∈ θ (·) from z t The derived query and key feature vectors, d is the dimension of the value vector, During normal denoising Self-attention map; a scale factor α∈[0,2] is introduced to control the fusion process, and the scale cross attention layer is defined as in is the inquiry feature, is the text encoding τ=ε(O T ) derived query and key feature vectors, d is the dimension of the value vector; Step 7: Define the similarity distance calculation as follows: Image O I The similarity between the image O(α,i) generated by fusion in step 5 is expressed as I sim (α,i)=d(O I ,O(α,i)); textO T The similarity between the image O(α,i) generated by fusion in step 5 is expressed as T sim (α,i)=d(O T ,O(α,i)), where d(·,·) is the similarity distance function of image / text; Step 8: Use the definition in step 7. When α = 1, I sim (i) = d(O I ,O(α=1,i)); Adaptively adjust the number of injection steps for self-attention injection in step 6: Initialization The update method is as follows: in and is a soft threshold set according to observation; after updating, i=i * ; Step 9: i=i obtained in step 8 * , using the definition in step 7, and then simplified to, I sim (α) = d(O I ,O(α,i=i * )) and T sim (α) = d(O T ,O(α,i=i * )), where d(·,·) represents the similarity distance between text / image and image; the fusion score function F(α) is defined as follows: Where β is the weight factor, and parameter k is used to alleviate the high I caused by the difference between text and image modalities. sim (α) and low T sim The scales of (α) are not consistent, thus ensuring that their scales are balanced; Step 10: Using the fusion score function defined in step 9 as the objective function, the golden section search method is used to determine the optimal value of α based on the fusion denoising process defined in step 5. * ; Step 11: Find the optimal i through steps 8 and 10 * With α * , input the Wensheng graph model to generate the optimal fusion visual image O(α * ,i * ).
2. The method for generating new objects by AI based on image-text fusion according to claim 1, characterized in that: The fidelity optimization method in step 4 first sets ∈ t To optimize the parameters, Optimize t In the denoising process Approximation to the normal inversion process For fidelity.
3. The method for generating new objects by AI by integrating images and texts according to claim 1, characterized in that: The method for measuring editability in step 4 is defined as ∈ t Distance from Gaussian distribution to measure editability When ∈ t The closer the distance to the Gaussian distribution, the higher the editability, and Does not participate in gradient calculation.
4. The method for generating new objects by AI by integrating image and text according to claim 1, characterized in that: The Fidelity vs. Editability Balancing Approach in Step 4, Defining Factors Factor to weigh the relationship between fidelity and editability, and set the final balance loss for λ represents balance and The weight of .
5. The method for generating new objects by AI by integrating images and texts according to claim 1, characterized in that: Set λ=125.
6. The method for generating new objects by AI by integrating images and texts according to claim 1, characterized in that: The method of controlling the object fusion process in step 6 is as follows: I , target object text O T , which are fused in the cross-attention layer of the Wensheng graph model in is the query feature output from the self-attention layer, and From O T The key and value features obtained by embedding the text τ; Adjust the object image O by the scale factor α I With the target object text O T The degree of fusion is as follows: α∈[0,2] controls the fusion process.
7. The method for generating new objects by AI by integrating image and text according to claim 1, characterized in that: The adaptive self-attention injection step adjustment method in step 8 makes the object fusion process smoother through adaptive step adjustment. and are set to 0.45 and 0.85 respectively; I sim (i) = d(O I , O(α=1,i)) means that when α=1, the generated image O(α=1,i) is fused with the original image O under the condition of self-attention injection i steps. I The similarity.
8. The method for generating new objects by AI by integrating image and text according to claim 1, characterized in that: In the fusion score function in step 9, k=2.3 and β=1 are set.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the new object AI generation method of image-text fusion as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for generating a new object by AI using image and text fusion as described in any one of claims 1 to 8 is implemented.