Multi-scene image semantic transformation and enhancement method and system based on diffusion model

By employing a multi-scene image semantic transformation and enhancement method based on a diffusion model, and utilizing visual language models and segmentation models for image editing, this method solves the problem of unnatural image generation in existing technologies, achieving high-precision and realistic image editing applicable to diverse scenarios.

CN121860902APending Publication Date: 2026-04-14HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2025-12-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing image editing technologies struggle to generate high-precision and realistic images in diverse real-world scenarios, particularly in terms of complex semantic understanding and detail generation, making it difficult to maintain the overall harmony and realism of the image.

Method used

A multi-scene image semantic transformation and enhancement method based on diffusion model parses user intent through visual language model, locates target region by combining segmentation model, and performs image inpainting and content generation using diffusion model and attention constraints, thereby achieving accurate semantic transformation and enhancement of target region or global image.

Benefits of technology

It significantly improves the realism and quality of image editing, enabling the generation of high-precision, naturally blended images in diverse real-world application scenarios. It supports multi-round interactive editing, allowing for high-precision, controllable editing based on time, season, and weather, and naturally inserting objects that fit the festive atmosphere of the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860902A_ABST
    Figure CN121860902A_ABST
Patent Text Reader

Abstract

The invention provides a multi-scene image semantic transformation and enhancement method and system based on a diffusion model, and relates to the technical field of artificial intelligence and computer vision. Identifying a user intention and generating a structured editing plan containing editing type labels and candidate region descriptions; a target area is positioned through a segmentation model, and an initial mask is generated; extracting surrounding environment information based on the mask boundary as a context basis for the diffusion model to carry out image restoration or content generation; and an attention constraint condition is set in combination with an editing type label, and the constraint is introduced in a diffusion process, so that accurate semantic transformation and enhancement of a target area or a global image are realized. The method solves the problem that in the prior art, a high-precision and naturally-fused image is difficult to generate in a complex scene, the authenticity and quality of image editing are remarkably improved, and the method is suitable for diversified practical application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and computer vision, and more specifically, to a method and system for multi-scene image semantic transformation and enhancement based on a diffusion model. Background Technology

[0002] With the rapid development of computer vision and deep learning technologies, image generation and editing techniques have been widely applied in digital media creation, advertising design, film and television special effects, and other fields. Traditional image editing methods typically rely on manual operations or rule-based pixel-level adjustments, making it difficult to achieve semantic-level modifications in complex scenes. In recent years, the rise of deep learning, especially Generative Adversarial Networks (GANs) and Diffusion Models, has provided new technical pathways for image-to-image conversion, making it possible to automatically modify image content based on semantic intent.

[0003] In existing technologies, image generation methods driven by semantic segmentation maps, sketches, or text descriptions can achieve a certain degree of semantic editing, such as changing object categories, adjusting scene layouts, or transferring artistic styles. However, they still have significant limitations when facing multi-scene, multi-level semantic transformation tasks. For example, when multiple regions need to be modified simultaneously or new objects are introduced, it is often difficult to maintain the overall harmony and realism of the image. Although some methods can maintain the global structure, they lack effective control over detail textures and style consistency, leading to problems such as artifacts, discontinuous edges, or style breaks in the generated results. In addition, existing models lack flexibility in the mapping relationship between user intent and actual output, especially performing poorly in application scenarios such as cross-domain semantic replacement or complex environment fusion.

[0004] In summary, existing image editing technologies still have shortcomings in complex semantic understanding and detail generation, making it difficult to generate high-precision and realistic images in diverse real-world scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-scene image semantic transformation and enhancement method and system based on a diffusion model, in order to solve the technical problem in the prior art that it is difficult to generate high-precision and realistic images in diverse real-world scenarios.

[0006] In a first aspect, embodiments of the present invention provide a multi-scene image semantic transformation and enhancement method based on a diffusion model. The method includes: semantically parsing the input original image and editing instructions based on a visual language model to identify user intent and generate a structured editing plan; the editing plan includes: an editing type label and a candidate region range description corresponding to the target object; based on the candidate region range description, calling a segmentation model to locate the target region corresponding to the target object in the original image and generating an initial mask covering the target region; extracting environmental information within the region not covered by the initial mask according to the boundary of the initial mask, and using the environmental information as input to a generation model to perform image restoration or content generation within the target region; determining corresponding attention constraints based on the editing type label, and introducing the attention constraints during the generation process of the diffusion model to perform semantic transformation or enhancement on the target region or the global image.

[0007] In some optional implementations, the above editing instructions include text prompts and keywords; semantic parsing of the editing instructions based on a visual language model includes: identifying and classifying user intents, the types of which include at least one of the following: global semantic transformation, local style enhancement, object insertion, and object removal; outputting editing type labels based on the types of the above user intents; wherein, the above global semantic transformations include changes in time, season, or weather; Extract potential target objects related to the aforementioned user intent, determine the location of the potential target objects based on the image semantic layout, and generate the aforementioned candidate region range description.

[0008] In some optional implementations, the above editing instructions also include a reference image; the above method further includes: extracting reference elements from the reference image, the reference elements including the visual style to be transferred, object features, and key structural information to be retained; determining the target scene based on the location or target region of the target object; and establishing a cross-modal association between the reference elements and the target scene based on the above visual language model.

[0009] In some optional implementations, environmental information is extracted from the area not covered by the initial mask based on the boundary of the initial mask, and this environmental information is used as input to the generative model to perform image inpainting or content generation within the target area. This includes: extracting contextual information from the area not covered by the initial mask based on the boundary defined by the initial mask; the contextual information includes texture, lighting, and color distribution features; encoding the contextual information into a conditional vector; and inputting the conditional vector into a context-consistent generative model to perform image inpainting or content generation within the target area.

[0010] In some optional implementations, the aforementioned attention constraints are introduced into the generation process of the diffusion model, including: when the aforementioned edit type label is a global semantic transformation, a spatial attention mask is generated based on the aforementioned candidate region range description, and the aforementioned spatial attention mask is input into the cross-attention layer of the diffusion model to mask the response weights of the key vectors and value vectors corresponding to non-target regions in multiple generation time steps; when the aforementioned edit type label is a local style enhancement, the spatial region where the target object is located is determined based on the aforementioned initial mask, an object-level attention guidance signal for the aforementioned spatial region is generated, and the aforementioned object-level attention guidance signal is fused with the text prompt word embedding vector and input into the self-attention module of the diffusion model to regulate the semantic expression of local regions during the generation process.

[0011] In some optional implementations, when the above-mentioned edit type label is object insertion and the object to be inserted has cultural attributes, the above method further includes: selecting the corresponding target generation model according to the cultural attributes of the object to be inserted; inputting the above-mentioned initial mask, the original image and the text prompt containing the festival keyword into the above-mentioned target generation model, and generating an image in the region corresponding to the above-mentioned initial mask; when the above-mentioned object to be inserted is a high-precision object with cultural identity, loading the pre-trained LoRA model parameters and superimposing and fusing them with the weights of the above-mentioned target generation model, and generating an image in the region corresponding to the above-mentioned initial mask.

[0012] In some optional implementations, the above method further includes: saving state information after each image editing is completed, including: the current intermediate result, the confirmed target region mask, the applied attention constraints, and the user interaction record; and reusing the previously saved state information as the basic constraints for the next round of image generation when responding to a new editing instruction.

[0013] Secondly, embodiments of the present invention provide a multi-scene image semantic transformation and enhancement system based on a diffusion model. The system includes: a semantic parsing module, used to perform semantic parsing on the input original image and editing instructions based on a visual language model, identify user intent, and generate a structured editing plan; the editing plan includes: an editing type label and a candidate region range description corresponding to the target object; a region localization and mask generation module, used to, based on the candidate region range description, call a segmentation model to locate the target region corresponding to the target object in the original image, and generate an initial mask covering the target region; an image restoration and generation module, used to, based on the boundary of the initial mask, extract environmental information within the area not covered by the initial mask, and use the environmental information as input to a generation model to perform image restoration or content generation within the target region; and an image enhancement module, used to, based on the editing type label, determine the corresponding attention constraints, and introduce the attention constraints during the generation process of the diffusion model to perform semantic transformation or enhancement on the target region or the global image.

[0014] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the steps of the method described in any of the first aspects above.

[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to perform the method described in any of the first aspects above.

[0016] This invention provides a method and system for multi-scene image semantic transformation and enhancement based on a diffusion model. The method first uses a visual language model to parse the input image and editing instructions, identify user intent, and generate a structured editing plan containing editing type labels and candidate region descriptions. Then, it locates the target region and generates an initial mask using a segmentation model. Based on the mask boundaries, it extracts surrounding environmental information as context for image inpainting or content generation using the diffusion model. Finally, it sets attention constraints based on the editing type labels and introduces these constraints during the diffusion process, achieving accurate semantic transformation and enhancement of the target region or the entire image. This invention solves the problem of existing technologies struggling to generate high-precision, naturally blended images in complex scenes, significantly improving the realism and quality of image editing, and is suitable for diverse practical application scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a multi-scene image semantic transformation and enhancement method based on a diffusion model, provided in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a multi-scene image semantic transformation and enhancement system based on a diffusion model provided in an embodiment of the present invention; Figure 3 This is an application diagram of a multi-scene image semantic transformation and enhancement method based on a diffusion model provided in an embodiment of the present invention; Figure 4 This is an example image illustrating the effect of inserting objects to create a festive atmosphere (Christmas) according to an embodiment of the present invention. Figure 5 This is an example image illustrating the effect of seasonal transformation (spring to winter) using a diffusion-based multi-scene image semantic transformation and enhancement method, provided by an embodiment of the present invention. Figure 6 An example image of the seasonal transformation (spring to autumn) effect provided by an embodiment of the present invention, applying a multi-scene image semantic transformation and enhancement method based on a diffusion model; Figure 7 An example image illustrating the effect of time transformation (morning to evening, outdoors) using a diffusion-based multi-scene image semantic transformation and enhancement method provided in an embodiment of the present invention; Figure 8 This is an example image illustrating the effect of time transformation (morning to evening, indoors) using a diffusion-based multi-scene image semantic transformation and enhancement method provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Because semantic transformations related to time, season, or festival involve multiple factors such as global lighting changes, the addition or removal of local objects, and cultural context, existing editing methods often struggle to simultaneously optimize realism, structural consistency, and style integration during the process. Therefore, existing technologies frequently encounter problems such as distorted details in generated areas, unnatural object boundaries, disordered lighting relationships, or inserted elements that do not conform to the cultural context of the original scene when performing complex semantic editing.

[0021] Based on this, the present invention provides a method and system for multi-scene image semantic transformation and enhancement based on a diffusion model, in order to solve the above-mentioned technical problems existing in the prior art.

[0022] To facilitate understanding of this embodiment, a detailed description of a multi-scene image semantic transformation and enhancement method based on a diffusion model disclosed in this invention will be provided first. (See [link to relevant documentation]). Figure 1 The diagram shows a flowchart of a multi-scene image semantic transformation and enhancement method based on a diffusion model. This method can be executed by an electronic device and mainly includes the following steps S102 to S108: Step S102: Based on the visual language model, perform semantic parsing on the input original image and editing instructions to identify user intent and generate a structured editing plan; the editing plan includes: editing type labels and a description of the candidate region range corresponding to the target object.

[0023] The editing instructions include text prompts and keywords. Semantic parsing of the editing instructions is performed based on a visual language model, including: identifying and classifying user intents, where the types of user intents include at least one of the following: global semantic transformation, local style enhancement, object insertion, and object removal; then, editing type labels are output based on the type of user intent; where global semantic transformations include changes in time, season, or weather; finally, potential target objects related to the user intent are extracted, and the location of the potential target objects is determined based on the image semantic layout, generating a candidate region range description.

[0024] In another embodiment, the above editing instructions also include a reference image; further, the above method may also include: extracting reference elements from the reference image, the reference elements including the visual style to be transferred, object features, and key structural information to be retained; determining the target scene based on the location of the target object or the target region; and establishing a cross-modal association between the reference elements and the target scene based on a visual language model.

[0025] Specifically, the aforementioned editing instructions can be text prompts, keywords, natural language editing instructions that include text prompts and keywords, or composite inputs that include reference images.

[0026] Text prompts refer to descriptive statements with a certain structure that conform to the common format of generative models. They are usually concise but complete sentences, designed to directly drive the generative model. For example: "Snow scene at night, warm light shining through the windows" or "Remove people and fill the background naturally." Keywords refer to scattered, incomplete words or phrases entered by the user, usually nouns or verbs, such as: "night," "snow scene," "remove people," "add lanterns." Natural language editing instructions are usually complex, multi-layered instructions entered by the user in the form of everyday conversational language, including subject-verb-object structures, conditional judgments, logical relationships, and even rhetorical expressions. For example: "I want to change this daytime photo to a winter night, but don't change the shape of the house, just put a Christmas tree in the yard, and don't make the tree too bright so it doesn't overshadow the subject." Preferably, the above editing plan may also include: priority information and risk control items, which can be used to limit the intensity of editing of sensitive areas, such as faces, brand logos, or key compositional elements.

[0027] Specifically, the implementation of step 102 above may include: first, receiving the original image and editing instructions input by the user, which may be text prompts, keywords, or composite input containing a reference image; then, invoking a Visual Language Model (VLM) to perform semantic understanding and reasoning on the editing instructions, the reasoning content may include: recognizing the user's intent, such as performing a global seasonal change, local object modification, object insertion or removal; determining the target object or region to be edited, for example, "changing daytime to nighttime" involves global illumination, "inserting a Christmas tree on the table" involves local addition; inferring the possible editing range and generating candidate region descriptions. Through the above process, different editing types can be automatically distinguished and a structured editing plan can be output.

[0028] After identifying the editing type, segmentation models (such as Grounded-SAM) can be used to accurately locate the target area.

[0029] Step S104: Based on the candidate region range description, the segmentation model is called to locate the target region corresponding to the target object in the original image and generate an initial mask covering the target region.

[0030] In one embodiment, after locating the target region in the original image by calling the segmentation model according to the candidate region range description in the structured editing plan, a fine mask after confidence fusion can be further generated by combining geometric constraints with a user-provided manual mask.

[0031] In another embodiment, in an object insertion scenario, the perspective relationship and spatial position of the inserted object can be inferred based on the geometric and depth information of the original image to generate a mask consistent with the background perspective.

[0032] Based on the inference results of the Visual Language Model (VLM) described above, the potential target location or object is first determined. Then, the Grounded-SAM segmentation model is used to perform segmentation operations, generating a fine-grained mask to delineate the area to be edited. For addition tasks, the mask for the insertion area can be reasonably generated based on semantics and spatial layout to ensure correct perspective proportions and spatial relationships.

[0033] When the Grounded-SAM model cannot accurately locate the area to be edited, users can manually specify the editing range, i.e., user-provided manual masks are also supported for greater controllability. The output of this step is a mask that is highly matched to the editing task, which is used for subsequent generation or repair.

[0034] Step S106: Based on the boundary of the initial mask, extract the environmental information in the area not covered by the initial mask, and use the environmental information as the input of the generation model to perform image inpainting or content generation in the target area.

[0035] After obtaining the mask of the target region, the image inpainting and generation stage begins. At this point, a generative model with contextual consistency (such as BrushNet) can be invoked to perform content generation or inpainting within the masked area. This model can infer semantically reasonable filling based on environmental information outside the mask boundaries, thus ensuring that the generated result blends naturally with the surrounding area in terms of texture, color, and lighting.

[0036] Specifically, when the task is to remove an object, the background can be reconstructed within the masked area using a diffusion model, so that the removed area is seamlessly connected with the original scene, maintaining overall visual coherence.

[0037] In one embodiment, the specific implementation of step S106 above may be as follows: when the edit type label is object removal, contextual information in the area not covered by the initial mask can be extracted based on the boundary defined by the initial mask; the contextual information may include: texture, lighting and color distribution features; then the contextual information is encoded into a conditional vector, and the conditional vector is input into the context-consistent generation model to perform image repair or content generation in the target area.

[0038] Specifically, within the area indicated by the mask, the repair or generation model can be invoked to perform content completion, replacement, or insertion operations. The repair or generation model can infer the content to be filled based on the external context information of the mask, so as to achieve a natural fusion with the original image in terms of texture, lighting, and color.

[0039] In another embodiment, when the edit type label is object insertion, the appropriate spatial landing point and perspective ratio can be generated based on the candidate region range description, and the generation model is used to generate filling content that conforms to semantic logic within the initial mask; the filling content may include virtual objects with reasonable three-dimensional pose, material texture and light effect, and its generation process is controlled by geometric constraints, which may include vanishing line direction, horizontal reference plane and coarse depth estimation.

[0040] Step S108: Determine the corresponding attention constraints based on the edit type label, and introduce the attention constraints during the generation process of the diffusion model to perform semantic transformation or enhancement on the target region or global image.

[0041] When global transformations such as seasonal, temporal, and weather changes are required, or style enhancements are needed in a localized area, attention ensemble models (such as InfEdit) can be invoked. By injecting attention constraints during the diffusion generation process, the core structural features of the edited object remain unchanged, while semantic style adjustments are achieved. In global transformations, effects such as day-night cycles, spring-autumn transitions, and weather changes can be implemented, generating details that conform to semantic logic (such as snow accumulation, fallen leaves, and changes in light and shadow). In local enhancements, the style, color, or texture of specified elements can be adjusted while maintaining the original spatial relationships and outlines. Through the attention constraint mechanism, the risk of the image deviating from the original image due to global regeneration is avoided.

[0042] In step S108 above, during the iterative generation of the diffusion model, the attention constraint can be introduced by applying a spatial mask to the cross-attention map, so that the generated result retains the main outline and spatial layout of the original image while performing semantic transformation, so as to achieve dynamic semantic transformation or style enhancement of the target region or global image while maintaining the core structural features of the original image.

[0043] In this embodiment, introducing attention constraints during the generation of the diffusion model may include: When the edit type label is a global semantic transformation, a spatial attention mask is generated based on the candidate region range description, and the spatial attention mask is input into the cross attention layer of the diffusion model to mask the response weights of the key vector and value vector corresponding to the non-target region in multiple generation time steps. When the editing type label is local style enhancement, the spatial region where the target object is located is determined based on the initial mask, an object-level attention guidance signal for the spatial region is generated, and the object-level attention guidance signal is fused with the text prompt word embedding vector and then input into the self-attention module of the diffusion model to regulate the semantic expression of the local region during the generation process.

[0044] Existing image transformation techniques also have significant shortcomings in inserting objects to create a festive atmosphere. Most applications currently rely on manual design or texture overlay, while some attempt automatic insertion using generated models. However, these methods often fail to accurately determine the object's placement and perspective proportions, resulting in a lack of spatial coherence in the generated images. Furthermore, inconsistencies between the inserted elements and the background style can create jarring effects, hindering the creation of a natural festive atmosphere. In summary, classic methods often remain at the surface level of color changes, leading to unnatural generated images; existing methods lack the ability to generate complex semantics and details, resulting in a lack of consistency and rationality in object insertion and removal. These issues prevent current technologies from meeting the demands for high-precision, highly controllable, and naturally integrated image editing in scenarios involving time, seasonal changes, and the insertion of objects to evoke a festive atmosphere.

[0045] In another embodiment, when the edit type label is "object insertion" and the object to be inserted has cultural attributes, the above method may further include: Select the appropriate target generation model based on the cultural attributes of the object to be inserted; then input the initial mask, the original image, and the text prompt containing festival keywords into the target generation model, and generate the image within the area corresponding to the initial mask, that is, generate the object to be inserted containing festival decorations.

[0046] Preferably, for inserting objects related to foreign festivals, the Flux-Kontext-Dev model can be used to generate Christmas trees, fairy lights, or roses; for inserting objects related to domestic festivals, the Qwen-Image-Edit model can be used to generate lanterns, Spring Festival couplets, zongzi (sticky rice dumplings), or mooncakes, etc.

[0047] In another embodiment, when the object to be inserted is a high-precision object with cultural identity, the pre-trained LoRA model parameters can be loaded and superimposed and fused with the weights of the target generation model to generate an image within the region corresponding to the initial mask, thereby improving the realism and recognizability of its shape, texture and material. During the generation process, the candidate landing points, occlusion relationships and lighting estimates obtained in the above steps are combined to constrain the scale, pose and hierarchy of the inserted object, and the shadow modeling and reflection synthesis algorithms are used to make it optically consistent with the original scene.

[0048] Specifically, the above-mentioned loading of LoRA model parameters and superimposing and fusing them with the weights of the generative model may include: inserting the low-rank matrix of LoRA into the query and value branches of the cross-attention layer of the generative model, guiding the local texture to evolve toward the predetermined cultural features in multiple time steps of the diffusion process.

[0049] In summary, when inserting objects to create a festive atmosphere, different editing methods can be selected based on the type of festival to ensure that the generated results conform to cultural characteristics and aesthetic habits. Specifically, for foreign festivals such as Christmas and Valentine's Day, the system uses the Flux-Kontext-Dev model, which is based on cross-modal editing capabilities, to generate objects that conform to the semantics of the festival within the masked area, such as Christmas trees, colored lights, or roses, maintaining perspective and style consistency with the original scene. For domestic festivals such as Spring Festival, Dragon Boat Festival, and Mid-Autumn Festival, the system calls the Qwen-Image-Edit model to perform the insertion operation, generating typical festival elements such as lanterns, couplets, zongzi (sticky rice dumplings), and mooncakes, ensuring that the image has authenticity and naturalness within the cultural context.

[0050] For certain festival objects with strong cultural significance, customized models trained on LoRa can be further employed to enhance the accuracy and detail of the generated data. For example, in the Mid-Autumn Festival scenario, a trained LoRa model can generate mooncakes with more realistic shapes and textures; in the Dragon Boat Festival scenario, it can generate zongzi (sticky rice dumplings) with clearer texture details. This approach not only improves the recognizability of the generated objects but also makes the results more in line with the semantic requirements of the festival atmosphere.

[0051] Through the above mechanisms, in different editing tasks such as object removal, insertion of holiday objects, and local modifications, image results that are highly coordinated with the surrounding content can be generated, thereby ensuring that the overall visual effect of the image is natural and smooth while meeting functional requirements.

[0052] In another embodiment, the method may further include: saving state information after each image editing is completed, the state information including: the current intermediate result, the confirmed target region mask, the applied attention constraints, and the user interaction record; and reusing the saved previous state information as the basic constraints for the next round of image generation when responding to a new editing instruction.

[0053] The method provided in this embodiment allows users to continue entering new editing commands after completing an initial edit, enabling multi-round interactive editing. During this process, not only is the overall structure and style of the original image maintained, but global transformations, local modifications, and object insertions can also be flexibly applied according to task requirements. Especially in editing tasks related to holiday atmospheres, users can first change the time or season, then insert holiday objects, or continue to adjust the style, position, and quantity of holiday objects after insertion.

[0054] To meet the needs of different cultural contexts, the system can automatically call upon appropriate editing models during multi-round interactions: for foreign holiday scenes, the Flux-Kontext-Dev model is prioritized to generate objects such as Christmas trees and jack-o'-lanterns; for domestic holiday scenes, the Qwen-Image-Edit model is used to generate elements such as lanterns, Spring Festival couplets, zongzi (sticky rice dumplings), and mooncakes; for specific high-precision objects, a dedicated model trained with LoRa can also be loaded to achieve precise enhancement at the detail level. In this way, during continuous multi-round editing, the system ensures the natural integration of holiday-themed objects with the background while maintaining the overall coherence and controllability of the image, avoiding distortion and abruptness caused by repeated generation.

[0055] In summary, the multi-scene image semantic transformation and enhancement method based on a diffusion model provided in this invention parses user commands using a visual language model, performs precise localization using a segmentation model, achieves context-consistent image generation using a repair and generation model, and completes global or local semantic enhancement through an attention integration mechanism, ultimately supporting multi-round interactive editing. Compared with existing technologies, this invention can achieve high-precision controllable editing of time, season, and weather, naturally insert objects with a festive atmosphere that fits the scene, accurately perform local modifications and object removal, and maintain the overall consistency and stability of the image in multiple rounds of operation. This method can be applied to multiple fields such as film and television post-production, advertising and marketing, e-commerce promotion, social entertainment, smart cities, and digital twins, and has significant practical value and industrial promotion significance.

[0056] Based on the same inventive concept, this invention also provides a multi-scene image semantic transformation and enhancement system based on a diffusion model, see [link to relevant documentation]. Figure 2 As shown, the system mainly includes the following parts: The semantic parsing module 210 is used to perform semantic parsing on the input raw image and editing instructions based on the visual language model, identify user intent, and generate a structured editing plan; the editing plan includes: editing type labels and a description of the candidate region range corresponding to the target object; Specifically, the aforementioned semantic parsing module can also be used to receive the original image, reference image, and text prompts or editing instructions input by the user, and parse the input information through a visual language model. This parsing process includes: identifying the user's editing intent, distinguishing different types of tasks such as global seasonal or time changes, local style enhancement, object insertion or removal, etc.; extracting potential target objects and candidate editing regions; semantically decomposing complex natural language instructions and generating a structured editing plan, which includes editing type labels, region range descriptions, and priority information, thereby providing semantic guidance for subsequent positioning and generation.

[0057] The region localization and mask generation module 220 is used to locate the target region corresponding to the target object in the original image based on the candidate region range description and call the segmentation model to generate an initial mask covering the target region. Specifically, the aforementioned region localization and mask generation module can call the segmentation model to accurately locate the target region according to the editing plan and generate a mask that matches the editing task. In object insertion scenarios, the module further combines image geometric information, vanishing lines, and single-view depth estimation to infer reasonable perspective relationships and spatial positions to ensure the consistency between the inserted object and the background. The module also supports user-provided manual masks and performs confidence fusion with automatically generated masks to improve user control while maintaining automation efficiency. The output mask boundaries are refined to ensure that smooth and natural editing ranges are generated in high-texture areas or at edges.

[0058] The image inpainting and generation module 230 is used to extract environmental information in areas not covered by the initial mask based on the boundaries of the initial mask, and use the environmental information as input to the generation model to perform image inpainting or content generation in the target area. The aforementioned image restoration and generation module can call a generative model with contextual consistency capabilities to perform restoration or generation within the masked area. For object removal tasks, the module uses environmental information outside the mask to reconstruct the background within the masked area, making the removed area seamlessly connected with the surrounding scene. For object replacement or local enhancement tasks, the module generates new target objects or style details while maintaining overall lighting, color, and material consistency. The module further uses a fusion algorithm to perform boundary transition and color alignment between the generated patch and the original image to avoid "patchwork" or local abruptness, thereby improving the overall naturalness and realism.

[0059] Image enhancement module 240 is used to determine the corresponding attention constraints based on the editing type label and introduce attention constraints during the generation process of the diffusion model to perform semantic transformation or enhancement on the target region or global image.

[0060] When performing global or local semantic transformations, attention constraints are introduced during the diffusion generation process to preserve core structural features and enhance semantic style. In global editing scenarios, this module can automatically execute effects such as day-night transitions, seasonal changes, and weather switching according to the editing plan, and generate details that conform to semantic logic in the results, such as winter snow scenes, autumn leaves, or evening light and shadow. In local editing scenarios, this module can adjust the style, material, or color of only specified objects or regions, while ensuring that their spatial position and shape are not changed. By limiting the generation changes of non-target regions, this module avoids structural shifts or semantic drifts in multiple editing sessions.

[0061] In another embodiment, the system may further include a festival atmosphere insertion module 250. This module is used to select a corresponding target generation model based on the cultural attributes of the object to be inserted when the editing type label is object insertion and the object to be inserted has cultural attributes; input the initial mask, the original image, and text prompts containing festival keywords into the target generation model, and generate an image within the area corresponding to the initial mask; when the object to be inserted is a high-precision object with cultural identity, load the pre-trained LoRA model parameters and superimpose and fuse them with the weights of the target generation model, and generate an image within the area corresponding to the initial mask.

[0062] When editing festival scenes, different generation models are automatically selected based on the festival type: for foreign festivals, the Flux-Kontext-Dev model is used to generate festive elements such as Christmas trees, colored lights, and roses within the masked area, while maintaining their perspective consistency and lighting coordination with the scene; for domestic festivals, the Qwen-Image-Edit model is used to generate elements such as lanterns, Spring Festival couplets, zongzi (sticky rice dumplings), and mooncakes, ensuring that the inserted objects conform to the Chinese cultural context; for objects with high detail requirements, this module can further call a customized model trained on LoRA to provide higher realism and recognizability in terms of shape, texture, and material; this module also uses lighting reconstruction and occlusion inference to make the inserted objects blend naturally with the original scene without producing an abrupt feeling.

[0063] In another embodiment, the system may further include a multi-round interactive editing module 260, which can be used to save state information after each image editing is completed, and reuse the saved previous state information as the basic constraint condition for the next round of image generation when responding to a new editing instruction.

[0064] Specifically, the aforementioned multi-round interactive editing module supports users in continuing to input new editing commands after completing one edit, and maintains image consistency during continuous editing. This module saves the intermediate results of each round of editing, the generated mask, and attention constraints, and reuses historical constraints for the next round of generation, thereby avoiding result distortion or accumulated errors in multiple iterations. This module allows users to flexibly switch between global editing (such as seasonal changes) and local insertion (such as adding holiday decorations), and also allows for secondary adjustments to the position, size, quantity, or style of inserted objects. In this way, the system can ensure the coherence, stability, and natural integration of the overall image in long chains and multiple rounds of interaction.

[0065] To facilitate understanding, this invention also provides an application example of a multi-scene image semantic transformation and enhancement method based on a diffusion model, see [link to relevant documentation]. Figure 3 The diagram illustrates an application of a multi-scene image semantic transformation and enhancement method based on a diffusion model. This method mainly includes: (1) Input and multi-round editing control; The system receives input images and editing instructions from the user and proceeds through a multi-round editing process. The editing instructions and images are first input into a visual language VLM model (such as GPT-4o) for parsing and decision-making.

[0066] (2) Editing type judgment and branching; For global editing, the visual language model provides source and target cue pairs and performs inference editing to generate target cue and the corresponding mask image.

[0067] If it is a partial edit or removal, a segmentation hint is provided, and the corresponding segmentation mask is generated through a grounded segmentation model.

[0068] If it is an add operation, a random mask is provided, or a user mask manually specified by the user.

[0069] (3) Mask generation and image restoration; During local editing / removal or addition processes, the grounded segmentation model generates accurate masks based on prompts.

[0070] The mask is combined with the original image to form a mask image, which is then input into a repair model (such as BrushNet) for content generation or modification.

[0071] (4) Output results: The image processed by the repair model is output as the final editing result, completing one round of editing. The system supports multiple iterations to achieve multi-round interactive image editing.

[0072] The method provided in this embodiment understands the editing intent through a visual language model, combines a segmentation model to achieve automatic or interactive mask generation, and relies on a repair model to achieve high-quality image editing, taking into account the flexibility of global and local editing.

[0073] As a specific example, the preferred implementation methods in each step of the above example will be described in detail below.

[0074] S301, Input and semantic parsing, generating an editing plan.

[0075] First, it receives the original image input by the user, supporting single image input as well as composite input with reference images, and also receives text prompts, keywords or natural language editing instructions.

[0076] To avoid the chain reaction of "ambiguous instructions—unclear editing scope—unstable results," the system invokes a multimodal large language model (such as Qwen-2.5VL) with visual-language alignment capabilities to parse instructions. This includes: intent recognition (global seasonal / time / weather changes, local style enhancement, object insertion or removal), target object or region extraction, and preliminary inference of candidate editing scope. The parsing results are organized into a structured editing plan, including: editing type labels, candidate region descriptions, priority and risk control items (e.g., default tightening of editing scope for sensitive areas such as the subject's face and brand logo).

[0077] In scenarios where reference images exist, the system can simultaneously complete cross-modal association between "reference elements and target scene," clarifying which style or object features need to be transferred and which structures need to be preserved, thereby reducing the risk of subsequent offsetting from the source.

[0078] S302, target region localization and mask generation, maintaining geometric / semantic consistency.

[0079] After clarifying the editing type and candidate range, the system uses a text-driven detection-segmentation linkage mechanism to accurately locate the target area and obtain the initial mask; For "insertion" type tasks, the system additionally evaluates possible landing points, feasible placement areas and their perspective relationships, and prioritizes generating masks in locations with more reasonable spatial relationships.

[0080] To improve boundary quality and spatial consistency, the system offers three capabilities: First, multi-scale mask refinement ensures that boundaries remain smooth and do not swallow details even in areas with dense texture. Second, geometrically constrained (horizontal lines, vanishing lines, coarse depth) assisted region correction avoids unreasonable scale or perspective. Third, interactive correction allows users to revise automatic masks with brushes or import manual masks; the system fuses both with confidence levels to achieve a balance between "automation efficiency" and "human controllability." The final output is a mask that highly matches the task and has perspective consistency, providing reliable boundaries for subsequent repair / generation and fusion.

[0081] S303, mask-based repair, generation, and seamless integration; enables removal / completion / detail generation.

[0082] The system enters the content generation or repair phase within the mask range: When the task is "remove", the reconstructed background generated by the diffusion model is based on the context of the unoccluded area to ensure consistency in texture, lighting and color levels; When the task is "complete / replace", semantically consistent details are generated within the mask, and the material and color transition naturally with the surrounding environment.

[0083] To further reduce common issues such as "visible edges" and "patch-like appearance," the system performs multi-level fusion processing (e.g., gradient-based or multi-scale pyramid fusion) on the generated patch and the original image, and makes global color / brightness fine-tuning to make the edited area visually consistent with the original image.

[0084] For tasks involving global refinement (such as unifying white balance or contrast so that local changes do not disrupt the overall atmosphere), the system performs only lightweight statistical alignment to avoid unnecessary impact on unedited areas, thereby maintaining overall structural stability.

[0085] S304, Attention-integrated global / local semantic transformation (structure-preserving diffusion editing).

[0086] When users request significant semantic changes such as season, time, or weather, or enhancements to style, color, and material in a localized area, the system uses an attention integration mechanism to constrain the shape and hierarchy of key objects during the generation process, avoiding "structural drift caused by semantic transfer." This can be achieved by calling an attention integration model (such as InfEdit).

[0087] In global transformations (such as day-night switching, spring-autumn transition, and sunny / snowy replacement), the system prioritizes maintaining the main outline and spatial relationships, injecting only semantically relevant details (such as snow cover in winter, leaf color changes in autumn, and the direction and color temperature of light and shadow in the evening). This allows the edited image to achieve an "atmosphere switch" in terms of visual appeal while remaining structurally faithful to the original image.

[0088] In local enhancements (such as adjusting the material and color of specified clothing or replacing the style of a home furnishing object), the system ensures that the shape, boundaries, and spatial relationship with the target are not changed. Controllable editing is achieved by transferring style / color within a limited range, allowing for "changing only the temperament without changing the structure".

[0089] S305, Insertion of festive atmosphere objects (model splitting and specialization enhancement).

[0090] For holiday scenarios, the system selects the appropriate editing backend based on the cultural context and the type of festival: for foreign holidays (such as Christmas, Valentine's Day, and Halloween), Flux-Kontext-Dev, which has cross-modal insertion and style consistency capabilities, is prioritized to generate objects such as Christmas trees, colored lights, pumpkin lanterns, and roses within a mask; for domestic holidays (such as Spring Festival, Dragon Boat Festival, and Mid-Autumn Festival), Qwen-Image-Edit is used for insertion and partial editing to generate elements such as lanterns, Spring Festival couplets, zongzi (sticky rice dumplings), and mooncakes, and automatically completes style coordination and lighting consistency with the background.

[0091] For objects with distinct cultural symbols and high requirements for the realism of details (such as the texture of zongzi leaves and the pattern of mooncakes), the system supports loading a customized model obtained through lightweight dedicated training to improve the credibility of the object in terms of form, texture and feel.

[0092] In terms of spatial layout, the system uses the candidate landing points and occlusion relationships obtained by S302 to constrain the scale, posture and hierarchical relationship of the inserted object, and processes shadows, reflections and brightness during the final compositing to make the insertion result fit the scene "closely without being a sticker", avoiding jarring and incongruous results.

[0093] Figure 4 The system demonstrates how Christmas elements can be automatically generated in indoor scenes and blend naturally with ambient lighting effects. Similarly, the system can generate elements with regional cultural styles in traditional Chinese festival scenes while maintaining overall unity.

[0094] S306, multi-round interactive editing and historical consistency maintenance (traceable and reversible).

[0095] The system supports entering new instructions after one editing session, enabling multiple rounds of interactive editing without having to start over; each round of editing adds new changes while preserving the existing structure and style constraints, avoiding cumulative distortion.

[0096] To this end, the system maintains a status record for each round of editing, including: the original image and the current intermediate result, confirmed masks and regions, applied style and structural constraints, and the user's selection / rollback history.

[0097] When moving to the next round, the system automatically adjusts the new editing strategy based on historical constraints: for example, first complete the global atmosphere transformation from "summer to evening", then insert local elements such as "birthday cake and balloons", or further adjust their position, size and quantity after "inserting lanterns". The system will retain and reuse the geometric / style consistency constraints of the previous step to ensure that the overall effect remains coherent and natural after multiple rounds of overlay.

[0098] In addition, the system provides reversible historical versions, reproducible random number seeds, and parameter snapshots, which facilitate repeated comparisons and rapid iterations during the production process. Figure 5 , Figure 6 , Figure 7 and Figure 8 The illustrations demonstrate the natural migration process of "seasonal change" and "time change" in different scenarios (outdoor / indoor), consistent with the coherent performance when multiple layers are superimposed.

[0099] In summary, this invention, through a chain of "semantic parsing, precise positioning, context-consistent repair / generation, structure-preserving semantic transformation, cultural context-aware festival insertion, and consistency maintenance through multi-round interactions," achieves high-precision, highly controllable, and naturally integrated multi-scenario image semantic editing capabilities for applications such as film and television post-production, advertising and e-commerce marketing, social entertainment, and digital twins. It can also be integrated with... Figures 3-8 The process and the results shown corroborate each other.

[0100] Based on the same inventive concept, this embodiment of the invention also provides an electronic device, which includes: a processor, a memory, and a bus. The memory stores machine-readable instructions that can be executed by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the multi-scene image semantic transformation and enhancement method based on the diffusion model described above.

[0101] Specifically, the aforementioned memory and processor can be general-purpose memory and processor, without any specific limitations. When the processor runs the computer program stored in the memory, it can execute the aforementioned multi-scene image semantic transformation and enhancement method based on the diffusion model.

[0102] The processor may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above methods can be completed by integrated logic circuits in the processor's hardware or by software instructions. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0103] Corresponding to the above-described multi-scene image semantic transformation and enhancement method based on the diffusion model, this embodiment of the invention also provides a computer-readable storage medium storing machine-executable instructions. When the machine-executable instructions are called and run by a processor, the machine-executable instructions cause the processor to perform the steps of the above-described multi-scene image semantic transformation and enhancement method based on the diffusion model.

[0104] The multi-scene image semantic transformation and enhancement device based on a diffusion model provided in this invention embodiment can be specific hardware on a device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this invention embodiment are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.

[0105] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0106] For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] In addition, the functional units in the embodiments provided by the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0109] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0110] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0111] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0112] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0113] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-scene image semantic transformation and enhancement method based on a diffusion model, characterized in that, include: Based on a visual language model, semantic parsing is performed on the input raw image and editing instructions to identify user intent and generate a structured editing plan; The editing plan includes: an editing type label and a description of the candidate region range corresponding to the target object; Based on the candidate region range description, the segmentation model is invoked to locate the target region in the original image corresponding to the target object, and an initial mask covering the target region is generated; Based on the boundaries of the initial mask, environmental information in areas not covered by the initial mask is extracted, and the environmental information is used as input to the generation model to perform image inpainting or content generation in the target area. Based on the edit type label, the corresponding attention constraints are determined, and these attention constraints are introduced during the generation of the diffusion model to perform semantic transformation or enhancement on the target region or global image.

2. The multi-scene image semantic transformation and enhancement method based on the diffusion model according to claim 1, characterized in that, The editing instructions include text prompts and keywords; Semantic parsing of editing instructions based on a visual language model includes: Identify and classify user intents, wherein the types of user intents include at least one of the following: global semantic transformation, local style enhancement, object insertion, and object removal; The type of output is an editable label based on the user's intent; wherein, the global semantic transformation includes changes in time, season, or weather; Extract potential target objects related to the user intent, determine the location of the potential target objects based on the image semantic layout, and generate the candidate region range description.

3. The multi-scene image semantic transformation and enhancement method based on the diffusion model according to claim 2, characterized in that, The editing instructions also include a reference image; the method further includes: Reference elements are extracted from the reference image, including the visual style to be transferred, object features, and key structural information to be retained; The target scene is determined based on the location of the target object or the target area; Based on the visual language model, a cross-modal association relationship is established between the reference elements and the target scene.

4. The multi-scene image semantic transformation and enhancement method based on the diffusion model according to claim 1, characterized in that, Based on the boundaries of the initial mask, environmental information is extracted from areas not covered by the initial mask, and this environmental information is used as input to the generative model to perform image inpainting or content generation within the target area, including: Based on the boundary defined by the initial mask, extract the contextual information of the area not covered by the initial mask; the contextual information includes: texture, lighting and color distribution features; The context information is encoded into a condition vector; The conditional vector is input into a context-consistent generative model to perform image inpainting or content generation within the target area.

5. The multi-scene image semantic transformation and enhancement method based on the diffusion model according to claim 2, characterized in that, The attention constraint is introduced during the generation of the diffusion model, including: When the edit type label is a global semantic transformation, a spatial attention mask is generated based on the candidate region range description, and the spatial attention mask is input into the cross attention layer of the diffusion model to mask the response weights of the key vector and value vector corresponding to the non-target region in multiple generation time steps. When the edit type label is local style enhancement, the spatial region where the target object is located is determined based on the initial mask, an object-level attention guidance signal for the spatial region is generated, and the object-level attention guidance signal is fused with the text prompt word embedding vector and then input into the self-attention module of the diffusion model to regulate the semantic expression of the local region during the generation process.

6. The multi-scene image semantic transformation and enhancement method based on the diffusion model according to claim 2, characterized in that, When the edit type label is "object insertion" and the object to be inserted has cultural attributes, the method further includes: Select the appropriate target generation model based on the cultural attributes of the object to be inserted; The initial mask, the original image, and the text prompt containing holiday keywords are input into the target generation model, and the image is generated within the area corresponding to the initial mask. When the object to be inserted is a high-precision object with cultural significance, the pre-trained LoRA model parameters are loaded and superimposed and fused with the weights of the target generation model, and the image is generated within the region corresponding to the initial mask.

7. The multi-scene image semantic transformation and enhancement method based on the diffusion model according to claim 1, characterized in that, The method further includes: After each image editing is completed, state information is saved, including: the current intermediate result, the confirmed target region mask, the applied attention constraints, and the user interaction record; When responding to a new editing instruction, the saved state information from the previous iteration is reused as the basis for constraints in the next round of image generation.

8. A multi-scene image semantic transformation and enhancement system based on a diffusion model, characterized in that, include: The semantic parsing module is used to perform semantic parsing on the input raw image and editing instructions based on the visual language model, identify user intent, and generate a structured editing plan; The editing plan includes: an editing type label and a description of the candidate region range corresponding to the target object; The region localization and mask generation module is used to locate the target region corresponding to the target object in the original image based on the candidate region range description, and generate an initial mask covering the target region; The image restoration and generation module is used to extract environmental information in areas not covered by the initial mask based on the boundary of the initial mask, and use the environmental information as input to the generation model to perform image restoration or content generation in the target area; The image enhancement module is used to determine the corresponding attention constraints based on the edit type label, and to introduce the attention constraints during the generation process of the diffusion model to perform semantic transformation or enhancement on the target region or global image.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.