Diffusion model-based image fusion method for reinforcement learning optimization
Through the diffusion model-based image fusion method optimized by reinforcement learning, the agent automatically selects the fusion location, solving the problem of time-consuming and labor-consuming user manual operation in the existing methods, and improving the image fusion effect and user experience.
Patent Information
- Application Number
- CN202510212614.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing image fusion method based on diffusion model relies too much on user manual operations when selecting fusion locations, which consumes a lot of time and energy, and the generated results are not ideal and the user experience is poor.
Using reinforcement learning optimization method, the agent selects the appropriate fusion position in the action space, uses the binary mask and image evaluation function LPIPS for feedback training, and automatically selects the position with better fusion effect in the picture.
Automatically selecting the fusion location is realized, reducing the time and energy of user manual operations, improving the effect and user experience of image fusion, and not damaging the prior knowledge of the diffusion model.
Smart Images

Figure CN120013778A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image fusion method for optimizing a diffusion model, and more specifically to an image fusion method based on a diffusion model optimized by reinforcement learning. Background Art
[0002] In recent years, the continuous development of science and technology has promoted the rapid progress of deep learning technology. The text-driven diffusion model has accelerated its development and matured rapidly in the rapid development of social media technology and the increasing market demand for high-quality and highly playable images. In order to respond to user needs, the diffusion model can already support various image editing tasks and specific instance generation tasks. For example, image fusion, text-driven modification of images, low-order adaptation model Lora of large language models, and Dreambooth training of fine-tuned diffusion models. Existing image editing tasks usually involve very expensive optimization based on specific instances, or require pre-training models on given data sets. These methods will destroy the prior knowledge of the model to a certain extent, resulting in a decrease in the model generation ability and scalability. In addition, Lora training and Dreambooth training of fine-tuned diffusion models will cause the scalability of the model to decrease, resulting in the generated images failing to achieve the expected effect. Therefore, a training-free cross-domain image fusion method based on a diffusion model has emerged - TF-ICON (TF-ICON: Diffusion-Based Training-FreeCross-Domain Image Composition). TF-ICON does not require any training for the diffusion model. It only needs to make some improvements to the structure to achieve training-free image fusion. This method does not require training, is low-cost, and does not damage the prior knowledge of the diffusion model. In addition, TF-ICON also supports the fusion of two images in different style domains, which can meet the user's many needs for playability. However, TF-ICON is very sensitive to the choice of image fusion position when fusing. In order to achieve a better fusion effect, users need to constantly try to select different fusion positions, which takes a lot of time and effort, and the final synthesis result may not be ideal, and the user experience is not friendly enough.
[0003] This paper optimizes and improves the existing TF-ICON algorithm, which can automatically select the position with better fusion effect in the picture, avoiding the situation where the fusion effect is poor when the user blindly chooses it. It saves a lot of human resources and achieves better fusion effect through prompt word optimization technology. In addition, this paper applies the image fusion method based on diffusion model optimized by reinforcement learning to the actual application of Lora, which solves the shortcoming of poor scalability of Lora in actual scenes, so that the generated pictures can better meet the requirements of users and bring users a richer experience. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides an image fusion method based on a diffusion model and optimized by reinforcement learning.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a reinforcement learning optimized image fusion method based on a diffusion model. The image fusion method based on a diffusion model is a fusion method that is sensitive to fusion position. The selection of fusion position greatly affects the fusion effect. The reinforcement learning optimized fusion method includes the following steps: 1. Process the background image and the subject image to be integrated; obtain the aspect ratio of the subject image, pre-process the background image and the subject image to be integrated, and obtain the action space of the intelligent agent based on the pre-processed background image and subject image. 2. Extract the action of the agent; the agent makes a selection in the action space of step 1 and obtains action action = (x, y, w, h), where x represents the horizontal coordinate position of the selection, y represents the vertical coordinate position of the selection, w represents the width of the selection area, and h represents the height of the selection area. 3. Obtain a binary mask; obtain a binary mask image of the background image fusion position according to the action of the intelligent agent. 4. Extract the fused intermediate image; input the background image in step 1, the main image to be fused, and the binary mask image in step 3 into a training-free cross-domain image fusion method based on a diffusion model, referred to as TF-ICON, to obtain a fused intermediate image. 5. Obtain the final fused guiding prompt words; use a picture-text multimodal network BLIP model to parse the intermediate image of step 4, obtain the description of the image, input the image and image prompt words into the generative large model based on the Transformer transformer, further expand and weight the image description words, and obtain the prompt words that guide image generation. 6. Use the image evaluation function to evaluate the fused image; input the prompt word in step 5 into the TF-ICON algorithm as a conditional parameter to guide image synthesis, obtain the fused image, and use the image evaluation method LPIPS to evaluate the image and obtain the evaluation score. 7. Map the evaluation scores; process the scores obtained in step 6, obtain the features and position information of the currently selected image region, feed back the agent, and proceed to step 2 to continue training until the reinforcement learning agent converges. Obtain the converged reinforcement learning model and proceed to step 8. 8. The trained intelligent agent obtained by the reinforcement learning mechanism obtains the image fusion area according to the background image and the main image, and guides the final fusion.
[0006] In step 3, the action action=(x, y, w, h) selected by the agent is used to obtain the selected area on the background image and convert it into a binary mask image as input of step 4.
[0007] The processing in the input TF-ICON of step 4 is to scale the main image of step 1 according to the ratio of the binary mask image of step 3, adjust the ratio of the main image to be integrated, and place it in a manner that the binary mask image is aligned with the center position of the background image, so as to obtain the required intermediate image.
[0008] In the processing of step 6, the perceptual loss LPIPS image difference between the main image area of the fused image and the main image area before fusion is evaluated by using an image evaluation function, and the evaluation score is recorded; the self-attention map and the cross-attention map in the generation process of the background image and the main image to be integrated are obtained from the fusion method TF-ICON, and then the two are spliced to form a new attention map; at the same time, the prompt word of step 5 is received as the text condition of the diffusion model image generation, and the fusion of the fused image is jointly guided.
[0009] The processing of step 7 is to facilitate the reinforcement learning agent to guide the agent to learn better fusion area selection based on the image perception loss LPIPS score obtained in step 6, and to perform linear mapping on the perception loss LPIPS score; the lower the perception loss LPIPS score, the higher the score obtained by the agent. Because the perception loss LPIPS score is the degree of difference between the two images, the lower the score means that the characteristics of the main image have not changed much after fusion and before fusion, which means the fusion effect is better, so linear mapping is performed here.
[0010] The lower the LPIPS score, the higher the score the agent gets. Specifically, when the LPIPS score is less than or equal to 0.5, the agent is given a larger reward score of reward = 10; the higher the LPIPS score, the worse the fusion effect, and the action is penalized for the poor effect. Specifically, when the LPIPS score is greater than or equal to 0.65, the agent is given a reward = -1. Negative scores make the agent try more in areas with high scores.
[0011] The present invention has the following beneficial technical effects: the present invention proposes an image fusion method based on a diffusion model optimized by reinforcement learning. First, the background image and the subject image to be integrated are preprocessed, and the action space of the intelligent agent is obtained; the binary mask on the background image is obtained according to the action taken by the intelligent agent; the binary mask, the preprocessed background image and the subject image are input into the TF-ICON method to obtain the intermediate image in the fusion process, that is, the spliced image of the background image and the subject image; the image is parsed through the graphic multimodal network BLIP model to obtain the image description; the description is input into the generative large model based on the Transformer transformer for further weighting and expansion to obtain the final prompt word; the final prompt word is used to guide the image fusion, the fused image is obtained, and the image is evaluated by the image evaluation method LPIPS; the LPIPS evaluation score is processed, and the image features and position information of the selected area of the current intelligent agent are fed back to the intelligent agent for learning; the trained intelligent agent is obtained and finally fused. The present invention can provide users with a convenient and efficient image fusion method, which can bring excellent fusion effects while freeing users' hands and reducing users' trial and error costs; the present invention can meet the diverse image needs of more users and bring rich, convenient and better image synthesis experience to more people. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of a flow chart of an embodiment of the present invention.
[0013] Figure 2 It is a complete flow chart of an embodiment of the present invention.
[0014] Figure 3 It is a flowchart of prompt word generation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention; in addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. The present invention will be described in more detail below with reference to the accompanying drawings. In each of the accompanying drawings, the same elements are represented by similar figure numerals. For clarity, the various parts in the accompanying drawings are not drawn to scale.
[0016] The present invention proposes an image fusion method based on a diffusion model and optimized by reinforcement learning. Figure 1 As shown, the method includes: S1. Process the background image and the subject image to be integrated. Preprocess the background image Xbg and the subject image to be processed Xfg. The background image is processed into a standard 512x512 pixel Xbg'. The subject image is segmented into the subject part Xfg' using the semantic segmentation method SAC; the aspect ratio w / h of the subject image Xfg is obtained, where w represents the width of the background image Xbg' and h represents the height of Xbg'. Then, the action space of the agent is initialized with this aspect ratio so that the action of the agent conforms to the aspect ratio of the subject image for better integration. S2. Extract the action of the agent. The agent selects in the action space of S1 and obtains the action action = (x, y, w, h) according to the current state space, where x represents the horizontal coordinate position in the background image Xbg', y represents the vertical coordinate position in the background image Xbg', w represents the width of the selected area, and h represents the height of the selected area. S3, obtain the binary mask. According to the action selected by the agent in S2, obtain the area selected by the action in the background image, and then convert it into a binary image, with the selected area as the white area and the rest of the background image as the black area. S4, extract the fused intermediate image. The background image Xbg' in S1, the main image Xfg' after semantic segmentation, and the binary mask in S3 are input into the TF-ICON algorithm for splicing and combination. Specifically, Xfg' is scaled to fit the size of the white area in the binary mask, and then covered with the background image area to achieve the purpose of combination. S5, obtaining the final fused guiding prompt words. Figure 3 As shown, the intermediate image obtained in S4 is input into the image-text multimodal network BLIP model to obtain a description of the intermediate image, and then the description and the text prompt are input into the Transformer-based generative large model to obtain the final guiding prompt words. S6, using the image evaluation function to evaluate the fused image; Figure 2 As shown, the final guiding prompt word is input into the TF-ICON algorithm to guide the continued fusion, and the fused image X is obtained, and the image evaluation function LPIPS is used to evaluate the image X. S61, extracting the main part of the fused image X according to the binary mask obtained in S3, and evaluating the image evaluation method LPIPS perceptual loss with Xfg in S1 to measure the similarity between the two images. The higher the similarity, the closer the main image features after fusion are to the original image, and the better the fusion effect; S62, LPIPS perceptual loss is an image quality evaluation indicator based on deep learning, which can more accurately simulate human perception of image quality. Using this as an evaluation indicator can make the fused image as close to human perception as possible. S7, mapping the evaluation score; in order to enable the agent to learn according to the fusion effect, so as to achieve the effect of selecting a better fusion position, the evaluation score is used as the agent's reward. Since the lower the score, the better the fusion effect, the perceptual loss LPIPS evaluation score in S6 needs to be mapped. S71. The obtained perceptual loss LPIPS score is lpips_socre. A linear mapping is performed to obtain the score score = (1-lpips_socre)*10, so that the perceptual loss LPIPS score is proportional to the reward of the agent, while expanding the difference, further promoting the agent to select a better fusion area; S72. In order to guide the agent to better integrate the learning of region selection, a larger reward is given when the action taken results in a lower LPIPS score, and a negative reward is given when the action taken results in a higher LPIPS score, thereby further promoting the agent to prefer regions with lower LPIPS scores. Specifically, when the LPIPS score is less than or equal to 0.5, the agent is given a reward score of reward = 10, and when the LPIPS score is greater than or equal to 0.65, the agent is given a score of reward = -1. By further increasing the difference between good actions and bad actions, the agent is encouraged to learn better; S73, return the score, the features of the current image extracted by the ResNet residual network, and the position information of the area selected by the agent to step S2 as new state information to guide the agent's choice of the next action, thereby guiding the agent to achieve better learning, until the reinforcement learning agent converges and obtains the converged model to proceed to step S8. S8. Obtain a trained intelligent agent and provide guidance for fusion. After the training is completed, obtain the trained intelligent agent, provide a background image and a subject image to be integrated, and finally obtain a fused image.
[0017] Obviously, the described embodiments are only some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
Claims
1. A reinforcement learning optimized diffusion model-based image fusion method, characterized in that: It includes the following steps: Step 1: Process the background image and the subject image to be integrated; obtain the aspect ratio of the subject image, pre-process the background image and the subject image to be integrated, and obtain the action space of the intelligent agent according to the pre-processed background image and subject image; Step 2, extract the action of the agent; the agent selects in the action space of step 1, and obtains action action = (x, y, w, h), where x represents the horizontal coordinate position of the selection, y represents the vertical coordinate position of the selection, w represents the width of the selection area, and h represents the height of the selection area; Step 3, obtaining a binary mask; obtaining a binary mask image of the background image fusion position according to the action of the agent; Step 4: extract the fused intermediate image; input the background image in step 1 and the subject image to be fused, and the binary mask image in step 3 into a training-free cross-domain image fusion method based on a diffusion model, referred to as TF-ICON, to obtain a fused intermediate image; Step 5: Obtain the final fused guiding prompt words; use the image-text multimodal network BLIP model to parse the intermediate image of step 4, obtain the description of the image, input the image and image prompt words into the generative large model based on the Transformer transformer, further expand and weight the image description words, and obtain the prompt words that guide image generation; Step 6: Use the image evaluation function to evaluate the fused image; input the prompt word in step 5 into the TF-ICON algorithm as a conditional parameter to guide image synthesis, obtain the fused image, and use the image evaluation method LPIPS to evaluate the image and obtain the evaluation score; Step 7: Map the evaluation scores; process the scores obtained in step 6, obtain the features and position information of the currently selected image area, feed back the agent, and continue training in step 2 until the reinforcement learning agent converges, and transfer the obtained reinforcement learning model to step 8; Step 8: Using the trained agent obtained by the reinforcement learning mechanism, obtain the image fusion area according to the background image and the subject image, and guide the final fusion.
2. The image fusion method based on diffusion model and optimized by reinforcement learning according to claim 1, characterized in that: In step 3, the action action=(x, y, w, h) selected by the agent is used to obtain the selected area on the background image and convert it into a binary mask image as input of step 4.
3. The image fusion method based on diffusion model and optimized by reinforcement learning according to claim 1, characterized in that: The processing in the input TF-ICON of step 4 is to scale the main image of step 1 according to the ratio of the binary mask image of step 3, adjust the ratio of the main image to be integrated, and place it in a manner that the binary mask image is aligned with the center position of the background image, so as to obtain the required intermediate image.
4. The image fusion method based on diffusion model and optimized by reinforcement learning according to claim 1, characterized in that: In the processing of step 6, the perceptual loss LPIPS image difference between the main image area of the fused image and the main image area before fusion is evaluated by using an image evaluation function, and the evaluation score is recorded; the self-attention map and the cross-attention map in the generation process of the background image and the main image to be integrated are obtained from the fusion method TF-ICON, and then the two are spliced to form a new attention map; at the same time, the prompt word of step 5 is received as the text condition of the diffusion model image generation, and the fusion of the fused image is jointly guided.
5. The image fusion method based on diffusion model and optimized by reinforcement learning according to claim 1, characterized in that: The processing of step 7 is to facilitate the reinforcement learning agent to guide the agent to learn better fusion area selection based on the image perception loss LPIPS score obtained in step 6, and to perform linear mapping on the perception loss LPIPS score; when the perception loss LPIPS score is lower, the score obtained by the agent is higher.
6. The image fusion method based on diffusion model and optimized by reinforcement learning according to claim 5, characterized in that: When the LPIPS score is less than or equal to 0.5, the agent is given a larger reward score of reward = 10; the higher the LPIPS score, the worse the fusion effect, and the action is penalized for the poor effect. Specifically, when the LPIPS score is greater than or equal to 0.65, the agent is given a reward = -1. Negative scores make the agent try more in areas with high scores.
Citation Information
Patent Citations
Multi-modal scene fusion method based on large diffusion model
CN119338940A
Real scene image editing method based on hierarchically classified text guidance
US20250005825A1
Cited By
Infrared image synthesis method and system based on cross attention and reinforcement learning
CN121353458A
Infrared image synthesis method and system based on cross attention and reinforcement learning
CN121353458B