Interactive image synthesis method based on visual language model
Through the chain reasoning and filtering strategy of visual language modeling, the challenge of foreground and background interaction modeling in image synthesis is solved, high-quality image synthesis results are generated, and seamless fusion in commercial use cases is achieved.
Patent Information
- Application Number
- CN202510365220.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
While maintaining the details of the foreground, existing image synthesis technologies are difficult to effectively model the physical interaction between the foreground and the background, resulting in a lack of realistic and consistent image generated, especially in commercial use cases, which is difficult to achieve seamless fusion.
An interactive image synthesis method based on visual language model is adopted, and through technical means such as element-by-element replacement, chain reasoning, similarity redistribution, Fourier domain filter update and attention injection, the physical interaction between modeling prospects and backgrounds is clarified, and interaction concepts are introduced into the diffusion model to ensure appearance consistency.
It realizes that high-quality image synthesis results are generated without training, maintaining the physical interaction between the foreground and the background, and improving the alignment and distribution smoothness of the generated results.
Smart Images

Figure CN120259098A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and generative artificial intelligence image synthesis, and specifically to an interactive image synthesis method based on a vision language model. Background Art
[0002] In the fields of computer vision and artificial intelligence image synthesis, significant progress has been made in image generation and editing technologies, especially in the area of diffusion models. These technologies have produced impressive results in various tasks. For example, when creating composite posters in fields such as commercial advertising, precise control of appearance and object position becomes crucial. However, although existing personalized methods have been successful in maintaining subject consistency and appearance fidelity, they often lack the background control required for commercial use cases.
[0003] Some research has addressed the problem of controllable image synthesis through training on specific datasets, but these methods have weak generalization capabilities outside the training domain. Other methods have explored approaches for untrained inference using prior knowledge encoded in generative models. However, there is still a gap between real-world images and the priors inherent in these models, which requires careful selection of inversion techniques and latent space operations to achieve seamless integration.
[0004] Although existing models are able to produce natural subject fusions, they often overlook the complex interactions between the inserted object and its surrounding environment. Realistic image synthesis is not just a simple appearance "paste", but also requires synthetic scene-aware "interactions", including visual dynamics, texture deformation, and the creation of new object concepts. These physical phenomena are crucial for realistic image synthesis but remain under-explored. Maintaining a balance between appearance consistency and realistic interaction effects remains a major challenge.
[0005] To address the challenge of modeling interactions while preserving foreground details, the present invention proposes a scene-object interaction-aware image synthesis method based on a diffusion model. This method explicitly models the physical interaction between the foreground and the background while ensuring appearance consistency for both, with the aim of achieving production-level image synthesis effects. Summary of the Invention
[0006] The purpose of the present invention is to provide an interactive image synthesis method based on a vision language model to solve the problems in the prior art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions: The present invention provides an interactive image synthesis method based on a vision language model, including the following steps:
[0008] S1. Input four pictures, a foreground picture and a background picture I f 、I b, and the segmentation map M of the foreground image seg and the map indicating the synthesis position are first at position M p using I f pixels of I b for element-wise replacement to obtain the reference image I ref ;
[0009] As shown in the following formula:
[0010] I ref = I b ⊙(1 - M seg ) + I f ⊙M seg
[0011] where ⊙ represents element-wise dot product, for I f , I b , I ref perform the diffusion inverse process to obtain the corresponding most accurate latent variable vector z ref T , z b T , z f T , as the initial noise for the next inference stage, through the vision-language model (VLM) with Chain of Thought (CoT), to identify the foreground and background, and generate a special text condition: PromotionPrompt, describing the most detailed actions between them, and select several tokens most relevant to the interaction as the key tokens T for Similarity Redistribution in the next inference stage;
[0012] S2. The z of the reference batch b t , z f t is the latent variable gradually denoised from the initial noise z b T , z f T ; the z of the output batch out t is denoised from the initial noise z ref T ;
[0013] S3. z out t is denoised under the text condition, using the VLM to obtain the Promotion Prompt that emphasizes the interaction, and then using Promotion Prompt: as the condition in the generation process to promote the interaction;
[0014] S4. Apply Similarity Redistribution (SR) to the corresponding markers in the cross-attention layer to enhance the interaction between the two images;
[0015] S5. Adopt the Fourier Domain Filter Updating Strategy (FDFUS) to optimally maintain the key information of the foreground and merge it with the generated new information into the output;
[0016] S6. Implement Attention Injection in the attention layer to maintain semantics at a finer-grained level. Under the combined effect of all the above features, image I can be obtained out .
[0017] Preferably, after providing the relevant image, the original prompt y, and a series of CoT instructions in S3, the VLM needs to output a new prompt, describing and introducing a new concept of the dynamic effect between the two images, called the promotion prompt
[0018] Preferably, the formula of S3 is:
[0019]
[0020] This process starts with providing the mask M f of the foreground image I seg to highlight the target object. Gemini will generate a detailed description of the foreground object, emphasizing its key features and potential behaviors. Subsequently, the background image I b is input into the VLM, which analyzes the scene and automatically segments it into different parts. The background mask M p indicates the position of the foreground object in the background. Gemini is responsible for reasoning about the objects in the M seg region and predicting its interaction with the background elements. To improve the coherence of the output, Chain of Thought (CoT) is introduced to gradually reason and optimize the model output, ensuring that the interaction between the foreground and background elements is considered. Despite the inherent uncertainty of the language model, additional instructions can refine the output of the VLM, focusing on the interaction prompt and selecting the key marker T. Finally, the output is ensured to be consistent with the task requirements under the designed CoT framework;
[0021] VLM instr is a vision-language model under the constraints of the designed Chain of Thought instructions, and y is the original prompt. The output is used as the condition for the reasoning process.
[0022] Preferably, in S4, Similarity Redistribution is specifically as follows: aiming to strengthen M seg ⊕M P region response, Similarity Redistribution (SR) is applied to the cross-attention layer. SR is applied to the cross-attention layer to generate the similarity map M sim , whose value represents the probability of the token at the element position. The goal of SR is to keep the distribution of the current token T in M sim stable, while changing the position of the strong probability response. For all elements in the similarity map MT of the key token T, the following operations are performed:
[0023]
[0024] μ' m p , S' m p , μ' 1-m p , S' 1-m p are the mean and variance of the elements ν' p , 1 - M p at the corresponding downsampling scale in M m p , ν' 1-m p ; μ' o , S' o are the mean and variance of M sim before the SR operation. Under these two constraints, t and k can be calculated explicitly. However, due to the normalization of the similarity map, the contribution of maintaining the mean is relatively small. Therefore, an intensity control value sc is introduced to adjust the final formula. After this operation, the response of the specific token T will be enhanced.
[0025] Preferably, in S5, the Fourier Domain Filter Updating Strategy (FDFUS) is used to maintain the consistency between the foreground and the background;
[0026] In the step-by-step iterative diffusion process, the high-frequency components are relatively less noisy. Therefore, accurately describing the proportion of the useful high-frequency part in the current step t and injecting high-frequency components are beneficial. The Fourier domain filter update algorithm is proposed. The input is two additional masks mseg and mp, and the result is divided into three regions: the background region bg that remains unchanged, the interaction region pad where interaction is expected, and the foreground region seg that maintains the high-frequency part of the foreground;
[0027] In the initial stage of denoising (i < u), the excitation model maintains the original background in the pad area; after initial denoising (i > u), the potential of the model is fully exploited. Throughout the process, high-frequency foreground information is retained, low-frequency content is provided at a ratio of rseg, foreground colors are maintained, and artifacts are avoided. Usually, rseg is set small (<0.3), and u is set within 15% of the total inference steps.
[0028] Preferably, the attention value injection Attention Injection in S6 is used to maintain a fine-grained understanding of the background semantics. Specifically, simply injecting the background in the latent space cannot achieve fine-grained semantic understanding. Therefore, attention injection (Attention Injection) is designed to interact with other elements, maintain the colors of the foreground and some other non-high-frequency elements, and at the same time keep the overall structure from undergoing rapid or non-background semantic changes. Since there are two batches, for z b t , corresponding Q can also be obtained in layer l t b(l) and K t b(l) . In the attention injection, since Q t (l) is always obtained from z t , whether it is a self-attention layer or a cross-attention layer, replacing Q t b(l) with Q t out(l) , and then performing the attention operation, so the output features of z out t after attention injection can be formalized as:
[0029]
[0030] This operation only changes the output batch z out t , and the reference batch z b t , z f t then performs a typical attention operation. It should be noted that during the inference process, two additional user input thresholds τs, τc are set as control intensities. When these thresholds are exceeded, it means that this mechanism will no longer be applied;
[0031] to interact with other elements, maintain foreground colors and certain non-high-frequency elements, and at the same time avoid rapid or non-background semantic changes in the overall structure.
[0032] The framework constructed by the present invention can be easily applied to any pre-trained diffusion model and achieve universal and plug-and-play synthesis capabilities.
[0033] To this end, first, provide I f and its mask M seg to highlight the target object. The VLM generates a detailed description of the foreground object and analyzes the background image, automatically segmenting the scene and labeling the background mask of the foreground object's position. Optimize the model output through Chain of Thought (CoT) to ensure consideration of the interaction between the foreground and the background, and finally generate prompts and key tokens T.
[0034] After that, apply SR to the cross-attention layer to generate the similarity map M sim , whose value represents the probability of the token at the element position. SR aims to maintain the stable distribution of token T in M sim while adjusting the positions of strong probability responses. By introducing the intensity control value sc, enhance the response of specific token T.
[0035] In addition, introduce FDFUS to maintain the consistency of the foreground. During the diffusion process, accurately describe the proportion of the useful high-frequency part in the current step t to inject high-frequency components. The proposed algorithm divides the result into the background region bg, the interaction region pad, and the foreground region seg. In the initial stage of denoising, keep the original background, and in the later stage, give full play to the potential of the model, retaining the foreground high-frequency information and providing low-frequency content.
[0036] Finally, use Attention Injection to inject fine-grained background. Attention Injection is designed for fine-grained semantic understanding, ensuring the preservation of the foreground color and non-high-frequency elements while avoiding rapid changes in the overall structure, thereby enhancing the interaction effect.
[0037] The constructed framework can effectively synthesize pictures that emphasize the interaction between the foreground and the background without any training process.
[0038] The present invention has at least the following beneficial effects:
[0039] The interactive image synthesis method based on the vision language model provided by the present invention constructs a control framework without training, which can synthesize the input foreground and background, and explicitly model the physical interaction between the foreground and the background while ensuring the appearance consistency between the two. This framework is plug-and-play without any additional training.
[0040] The present invention expands the vision language model, combines the chain of reasoning, and introduces a new concept of "interaction". These concepts are strengthened in the denoising step, thereby improving the alignment of the generated results while maintaining the smoothness of the distribution.
[0041] A filtering strategy based on Fourier transform is proposed and applied in the denoising process. This strategy ensures that the pre-trained diffusion model can generate high-quality results while retaining the texture and shape details of the foreground, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 The overall framework flowchart provided by the present invention;
[0043] Figure 2 The schematic diagram provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs. The terms used in this specification in the description of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0045] Embodiment
[0046] An interactive image synthesis method based on a vision-language model includes the following steps:
[0047] S1. Input four pictures, a foreground picture and background pictures I f 、I b , and a segmentation map M of the foreground picture seg and a picture representing the synthesis position. First, at position M p use the pixels of I f to perform element-wise replacement of I b to obtain a reference image I ref ;
[0048] As shown in the following formula:
[0049] I ref =I b ☉(1 - M seg ) + I f ☉M seg
[0050] where ⊙ represents element-wise multiplication. Perform the diffusion inverse process on I f , I b , I ref to obtain the corresponding most accurate latent variable vectors z ref T , z b T , z f T, as the initial noise for the next reasoning stage, a vision - language model (VLM) with Chain of Thought (CoT) is used to identify the foreground and background, and generate a special text condition: PromotionPrompt, which describes the most detailed actions between them. Select several tags most relevant to the interaction as the key tags T for Similarity Redistribution in the next reasoning stage;
[0051] S2. The z of the reference batch b t , z f t , is the latent variable gradually denoised from the initial noise z b T , z f T ; The z of the output batch out t , is denoised from the initial noise z ref T ;
[0052] S3. z out t is denoised under the text condition, using the VLM to obtain the Promotion Prompt that emphasizes the interaction, and then using the Promotion Prompt: as a condition in the generation process to promote the interaction;
[0053] S4. Apply Similarity Redistribution (SR) in the corresponding tokens of the cross - attention layer to enhance the interaction between the two images;
[0054] S5. Adopt the Fourier Domain Filter Updating Strategy (FDFUS) to best maintain the key information of the foreground and merge it with the newly generated information into the output;
[0055] S6. Implement Attention Injection in the attention layer to maintain semantics at a finer - grained level. Under the combined action of all the above features, the image I out can be obtained.
[0056] Preferably, after providing the relevant image, the original prompt y, and a series of CoT instructions in S3, the VLM is required to output a new prompt that describes and introduces a new concept of the dynamic effect between the two images, called the promotion prompt
[0057] Preferably, in this module, after providing the relevant images, the original prompt y, and a series of CoT instructions, we need the VLM to output a new prompt that describes and introduces a new concept of the dynamic effect between the two images, called the facilitation prompt. Based on this prompt, some words strongly related to the interaction will also be selected as the key tokens T for the next similarity redistribution.
[0058] We use Gemini as the implemented VLM. The process starts with providing the foreground image I f 's mask M seg to highlight the target object in the image. Gemini uses this mask to generate a detailed description of the foreground object, emphasizing its key features and potential future behaviors or trends. Subsequently, we input the background image I b into the VLM, instructing it to analyze the entire scene and automatically segment it into different parts. The prominent details and objects in each part and their spatial relationships will be evaluated. The background mask M p also indicates the position of the foreground object in the background. Gemini is responsible for reasoning about the objects in the masked region M seg and predicting how the foreground object might interact with these background elements. A prompt that emphasizes the potential interaction between the foreground and the background is generated. Although this is reasonable in most cases, due to the inherent uncertainty of the language model and the tendency of the vision-language model to mainly focus on the input images after fine-tuning, in some cases, the VLM may be more inclined to describe the images rather than generate the required interaction prompt. To address this issue, we introduce chain-of-thought (CoT).
[0059] Chain-of-thought optimizes the output of the model through step-by-step reasoning. This method enhances the generation of coherent multimodal prompts by ensuring that foreground and background elements are considered in the output. It improves the contextual accuracy of the model's response and makes it consistent with the task requirements. However, due to the inherent uncertainty of the language model and its dependence on the training dataset, the generated interaction may still be ambiguous. With the introduction of additional instructions specifying the interaction content, the output of the VLM will be refined into a prompt that focuses on the interaction and the key tokens T are selected as shown in the following formula:
[0060]
[0061] VLM instr is a vision-language model under the constraints of the designed chain-of-thought instructions, and y is the original prompt. The output is used as the condition for the reasoning process, and the entire process is represented in Figure 1
[0062] Preferably, Similarity Redistribution (SR) is applied to the cross-attention layer, aiming to strengthen M seg ⊕M P responses in the region, where ⊕ is the exclusive OR operation.
[0063] Consider a typical attention operation. If the input feature of z out t at the time step t in layer l is f out t , and the input feature of the prompt condition is τθ(y), then the Q value and K value can be calculated from the following formula:
[0064]
[0065] where WQ and WK are both learnable matrices. After that, the Attn operation can be performed:
[0066]
[0067] Similarly, WV is also a learnable matrix, and dk is the dimension of Q and K. Among them
[0068] The result of is usually regarded as the similarity map M sim , whose value indicates the probability of the corresponding token at the position of this element. SR aims to keep the distribution of the current token in the entire M sim from changing drastically, while changing the position of the stronger probability response. We perform the following operation on all elements at the corresponding positions in the similarity map M T of the key token T:
[0069]
[0070] μ' m p , S' m p , μ' 1-m p , S' 1-m p are the mean and variance of the elements ν' p , 1 - M p at the corresponding downsampling scale in the position of M m p , ν' 1-m p ; μ' o , S' o are before the SR operation in M simThe mean and variance. Under these two constraints, t and k can be explicitly calculated. However, these constraints are indeed too strict, and since the similarity graph is always normalized to keep the sum equal to 1, the contribution of maintaining equal means is small. Therefore, we add a strength control value for the degree of control, and the final formula can be written as:
[0071]
[0072] For consistency, the values of t and k are the same as in the previous equation, and s c is the control strength. After this operation, the response of a specific marker T will be greater, as Figure 1 shown.
[0073] Preferably, in S5, the Fourier Domain Filter Updating Strategy (FDFUS) is used to maintain the consistency between the foreground and the background;
[0074] In the step-by-step iterative diffusion process, the high-frequency components are relatively less noisy. Therefore, accurately describing the proportion of the useful high-frequency part in the current step t and injecting high-frequency components is beneficial. A Fourier domain filter update algorithm is proposed. The input consists of two additional masks, mseg and mp, and the result is divided into three regions: the background region bg that remains unchanged, the interaction region pad where interaction is desired, and the foreground region seg that retains the high-frequency part of the foreground;
[0075] In the initial stage of denoising (i < u), the excitation model maintains the original background in the pad region; after the initial denoising (i > u), the potential of the model is fully exerted. Throughout the process, the high-frequency information of the foreground is retained, and low-frequency content is provided in proportion rseg to maintain the foreground color and avoid artifacts. Usually, rseg is set to be small (<0.3), and u is set to be within 15% of the total inference steps.
[0076] Preferably, in S6, the Attention Injection is used to maintain a fine-grained understanding of the background semantics. Specifically, simply injecting the background in the latent space cannot achieve a fine-grained semantic understanding. Therefore, the Attention Injection is designed to interact with other elements and maintain the color of the foreground and some other non-high-frequency elements, while keeping the overall structure from undergoing rapid or non-background semantic changes. Since there are two batches, for z b t ,, the corresponding Q t b(l) and K t b(l) . In the attention injection, since Q t(l) is always obtained from z t in both self-attention and cross-attention layers, replacing Q with Q t b(l) and then performing the attention operation. Thus, the output features of z after attention injection can be formalized as: t out(l) This operation only changes the output batch z out t and the reference batch z
[0077]
[0078] out t b t f t , z f seg then performs the typical attention operation. It should be noted that during the inference process, two additional user input thresholds τs and τc are set as control intensities. When these thresholds are exceeded, it means that this mechanism will no longer be applied;
[0079] to interact with other elements and maintain the foreground color and certain non-high-frequency elements, while avoiding rapid or non-background semantic changes in the overall structure.
[0080] The framework constructed by the present invention can be easily applied to any pre-trained diffusion model and achieve universal and plug-and-play synthesis capabilities.
[0081] For this purpose, first, provide I f and its mask M seg to highlight the target object. The VLM generates a detailed description of the foreground object and analyzes the background image, automatically segmenting the scene and labeling the background mask of the foreground object's position. Optimize the model output through chain-of-thought (CoT) reasoning to ensure that the interaction between the foreground and background is considered, and finally generate a prompt focusing on the interaction and the key token T.
[0082] After that, apply SR to the cross-attention layer to generate the similarity map M sim whose value represents the probability of the token at the element position. SR aims to maintain the stable distribution of the token T in M sim while adjusting the position of the strong probability response. By introducing the intensity control value sc, enhance the response of the specific token T.
[0083] In addition, FDFUS is introduced to maintain the consistency of the foreground. During the diffusion process, the proportion of the useful high-frequency part in the current step t is accurately described to inject high-frequency components. The proposed algorithm divides the result into a background region bg, an interaction region pad, and a foreground region seg. In the initial stage of denoising, the original background is maintained, and in the later stage, the potential of the model is fully exploited to retain the high-frequency information of the foreground and provide low-frequency content.
[0084] Finally, Attention Injection is used to inject fine-grained background. Attention Injection is designed for fine-grained semantic understanding, ensuring the retention of foreground colors and non-high-frequency elements while avoiding rapid changes in the overall structure, thereby enhancing the interaction effect.
[0085] The constructed framework can effectively synthesize pictures with emphasized interaction between the foreground and background without any training process.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above. For the sake of brevity, they are not provided in detail; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An interactive image synthesis method based on a vision-language model, characterized in that, Including the following steps: S1. Input four images, the foreground image and background image I f 、I b , and the segmentation map M of the foreground image seg and the map indicating the synthesis position. First, at position M p use the pixels of I f to perform element-by-element replacement of I b to obtain the reference image I ref ; As shown in the following formula: I ref = I b ⊙(1 - M seg ) + I f ⊙M seg For I f , I b , I ref Execute the diffusion inverse process to obtain the corresponding most accurate latent variable vector z ref T , z b T , z f T , as the initial noise for the next inference stage, through a vision - language model (VLM) with Chain of Thought (CoT), identify the foreground and background, and generate a special text condition: PromotionPrompt, which describes the most detailed actions between them, select several tokens most relevant to the interaction as key tokens T for SimilarityRedistribution in the next inference stage; S2, z of the reference batch b t , z f t , is the latent variable that is gradually denoised from the initial noise z b T , z f T ; z of the output batch out t , is denoised from the initial noise z ref T . S3, z out t Denoise under text conditions, use VLM to obtain a Promotion Prompt that emphasizes interaction, and then use the Promotion Prompt: as a condition during the generation process to promote interaction; S4. Apply Similarity Redistribution (SR) to the corresponding markers in the cross-attention layer to enhance the interaction between the two images; S5. Adopt the Fourier Domain Filter Updating Strategy (FDFUS) to optimally maintain the key information of the foreground and merge it with the generated new information into the output; S6. Implement AttentionInjection in the attention layer to maintain semantics at a finer-grained level. Under the combined action of all the above features, the image I can be obtained out .
2. The interactive image synthesis method based on a vision-language model according to claim 1, wherein After providing relevant images, the original prompt y, and a series of CoT instructions in S3, the VLM is required to output a new prompt that describes and introduces a new concept of the dynamic effect between two images, called the facilitating prompt.
3. The interactive image synthesis method based on a vision-language model according to claim 2, wherein The formula of S3 is:
4. The interactive image synthesis method based on a vision-language model according to claim 1, wherein The Similarity Redistribution in S4 is specifically as follows: It aims to enhance the response of the M seg ⊕M P region, and the formula is:
5. The interactive image synthesis method based on a vision-language model according to claim 1, wherein In S5, the Fourier Domain Filter Updating Strategy (FDFUS) is used to maintain the consistency between the foreground and the background.
6. The interactive image synthesis method based on a vision-language model according to claim 1, wherein, In S6, the Attention Injection for attention value injection is used to maintain the fine-grained understanding of the background semantics, and the formula is: