Cognitive reasoning-based instruction type image editing method and system
By employing a two-stage editing method based on cognitive reasoning and a filtering process for locating and modifying prompts, the problem of inaccurate instruction understanding and region localization in instruction-based image editing systems is solved, achieving image editing with high-level semantic reasoning and visual consistency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing instruction-based image editing systems suffer from poor accuracy in understanding natural language instructions, inaccurate localization of editing areas, and insufficient controllability of local operations, especially in high-level semantic reasoning and visual consistency.
A two-stage editing method based on cognitive reasoning is adopted, including the location cognitive process (LCP) and the modification cognitive process (MCP). By generating location and modification prompts and combining them with reflective system prompts, the optimal result is selected, simulating human cognitive decision-making steps and improving the accuracy of instruction understanding and editing areas.
It improves the understanding of editing instructions and the accuracy of editing area positioning, ensuring visual consistency and local controllability, reducing error accumulation in multiple editing processes, and achieving precise and interpretable image editing.
Smart Images

Figure CN121982152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of instruction-based image editing in computer vision. Background Technology
[0002] Instruction-based image editing aims to achieve precise modifications to images using natural language. This paradigm allows users to express abstract or high-level editing intentions through natural language, thus bridging the gap between human imagination and visual manipulation.
[0003] In recent years, the rapid development of large-scale multimodal models has driven progress in imperative image editing. However, existing methods still face several long-standing challenges, especially in scenarios requiring high-level semantic reasoning or fine visual consistency. When editing results do not meet expectations, it is often difficult to determine whether the problem stems from a misunderstanding of the instruction or insufficient generation control. On the one hand, most models lack a structured reasoning mechanism, typically mapping abstract user-provided instructions directly to pixel-level generation or transformation processes of the entire image. This makes it difficult to break down complex instructions into clear, executable editing steps or plans, resulting in limited generalization ability across different levels of abstraction or diverse visual contexts. On the other hand, existing methods usually process the entire image holistically, lacking explicit region isolation mechanisms. This leads to poor accuracy in instruction understanding, inaccurate editing region positioning, and insufficient local controllability of editing operations, easily causing unnecessary modifications to areas unrelated to the instructions. This is particularly fatal in images containing text or precise visual layouts that are highly sensitive to detail. In continuous multi-round editing scenarios, these problems accumulate, ultimately making the editing results unacceptable. Summary of the Invention
[0004] The purpose of this invention is to solve the problems of poor accuracy in understanding natural language instructions, inaccurate positioning of editing areas, and insufficient local controllability of editing operations in existing instruction-based image editing systems. This invention provides an instruction-based image editing method and system based on cognitive reasoning.
[0005] Instructional image editing methods based on cognitive reasoning include:
[0006] Phase 1, Locational Cognitive Process (LCP):
[0007] Location system prompts Under the constraints of editing instructions semantic information and input image Edit and plan to generate a set of location tips consisting of multiple location tips. And combined with reflective system prompts Filter out the best location prompts ;
[0008] Based on optimal location prompts and input image Generate a candidate localization mask set consisting of multiple candidate binary localization masks. And combined with reflective system prompts Filter out the optimal positioning mask ;
[0009] The second stage involves modifying cognitive processes (MCP):
[0010] Modified system prompts and optimal positioning mask Under the constraints of editing instructions semantic information of the input image The modification plan is modified, generating a set of modification prompts consisting of multiple modification prompts. And combined with reflective system prompts Filter out the best modification suggestions ;
[0011] In optimal modification prompts and optimal positioning mask Under the constraints, a candidate edited image set consisting of multiple edited images is generated. And combined with reflective system prompts Select the best edited image .
[0012] Preferably, the system prompts... Under the constraints of editing instructions semantic information and input image Edit and plan to generate a set of location tips consisting of multiple location tips. The implementation method is as follows:
[0013] Location system prompts Under the constraints, large-scale multimodal models are based on editing instructions. semantic information and input image Perform cross-modal alignment to determine the relationship with the edit instructions. The relevant target area is identified, and a textual semantic description of that target area is provided to generate a location cue. The process of generating location prompts is repeated multiple times to obtain a set of location prompts. .
[0014] Preferably, the optimal location suggestion is selected. The implementation method is as follows:
[0015] In reflective system prompts Under the constraints, large-scale multimodal models provide localization cue sets. Location prompts in With the input image and editing instructions The matching degree between the constructed multimodal contexts is scored, and a score is obtained accordingly. The above matching degree scoring process is repeated multiple times to obtain multiple scores, and the positioning prompt with the highest score is selected as the optimal positioning prompt. .
[0016] Preferably, based on the optimal location prompt and input image The implementation method for generating multiple candidate binary location masks is as follows:
[0017] Instruction segmentation model for input image With optimal positioning hints Perform semantic matching and output an initial binary localization mask. Then, a morphological dilation operation is applied to it to obtain the corresponding candidate binary localization mask. The process of obtaining candidate binary localization masks is repeated multiple times to obtain a candidate localization mask set consisting of multiple candidate binary localization masks. .
[0018] Preferably, the optimal positioning mask is selected. The implementation method is as follows:
[0019] In reflective system prompts Under constraints, large-scale multimodal models combined with optimal localization hints and input image The candidate localization mask set was evaluated from two dimensions: spatial localization accuracy and semantic relevance. The binary positioning masks in the dataset are evaluated comprehensively, and the binary positioning mask with the highest score is selected as the optimal positioning mask. .
[0020] Preferably, in the modified system prompt and optimal positioning mask Under the constraints of editing instructions semantic information of the input image The modification plan is modified, generating a set of modification prompts consisting of multiple modification prompts. The implementation methods include:
[0021] Modified system prompts Hints, constraints, and optimal positioning mask Under the provided spatial constraints, large multimodal models respond to editing instructions With input image Perform cross-modal alignment to determine a modification hint for the target region. Repeat the above process of generating modification suggestions multiple times to obtain a set of modification suggestions. 7. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that the optimal modification prompt is selected. The implementation method is as follows:
[0022] In reflective system prompts Under the constraints, large-scale multimodal models are combined with editing instructions Input image and optimal positioning mask The modification prompt set was analyzed from two dimensions: semantic consistency and visual context. Modification prompts in Perform a comprehensive evaluation and select the highest-scoring modification suggestion as the optimal modification suggestion. .
[0023] Preferably, in the optimal modification prompt and optimal positioning mask Under the constraints, a candidate edited image set consisting of multiple edited images is generated. The implementation methods include:
[0024] Based on the optimal positioning mask Determine the input image The target area, and at the same time, the optimal modification prompt. Under constraints, the image restoration model processes the input image. The target region is conditionally generated to produce an edited image. The process of generating edited images is repeated multiple times to obtain a candidate edited image set consisting of multiple edited images. ;in,
[0025] The noise parameters of the image restoration model are randomized each time an edited image is generated.
[0026] Preferably, the best edited image is selected. The implementation methods include:
[0027] In reflective system prompts Under the constraints, large-scale multimodal models are combined with editing instructions and input image The candidate editable image set was evaluated based on three aspects: semantic fidelity, visual quality, and consistency with the unedited region. The edited images are evaluated comprehensively, and the edited image with the highest score is selected as the best edited image. .
[0028] The cognitive reasoning-based instruction-based image editing system includes a storage device, a processor, and a computer program stored in the storage device and executable on the processor. The processor executes the computer program to implement the cognitive reasoning-based instruction-based image editing method as described above.
[0029] The beneficial effects of this invention are:
[0030] This invention simulates human cognitive decision-making steps, breaking down abstract instruction-based image editing into two sub-tasks: localization and modification. In the LCP and MCP stages, multiple localization and modification prompts are generated, and these are scored using multimodal context. The result with the highest score is selected, effectively avoiding localization offset and modification intent drift, thus improving the accuracy of understanding editing instructions. Furthermore, regarding editing region localization, this invention differs from traditional end-to-end mapping methods. Instead, it first obtains the optimal localization prompt and generates multiple binary localization masks under its guidance. These masks are then scored using multimodal context, and the highest-scoring optimal localization mask is selected, achieving precise localization of the editing region and giving the entire system local controllability over editing operations.
[0031] This invention achieves local controllability of editing operations while ensuring accurate understanding of instructions and accurate positioning of the editing area. This enables the system to reliably execute sophisticated image editing driven by complex natural language instructions, while maintaining consistency in high-level semantic reasoning and visual consistency.
[0032] This invention includes a two-stage cognitive framework of location cognitive process (LCP) and modification cognitive process (MCP), and introduces reflective systemic cues. This improves the model's robustness when instructions are ambiguous and effectively reduces the accumulation of errors during multiple rounds of editing.
[0033] This invention is based on open-source instruction-based segmentation models, image inpainting models, and large-scale multimodal models. It can achieve accurate, interpretable, and visually consistent image editing without relying on task-specific datasets or high-cost fine-tuning. Attached Figure Description
[0034] Figure 1 This is a schematic diagram illustrating the principle of the instruction-based image editing method based on cognitive reasoning described in this invention. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0037] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0038] Specific Implementation Method 1: Combination Figure 1 This embodiment describes a cognitive reasoning-based instruction-based image editing method, which includes:
[0039] Phase 1, Locational Cognitive Process (LCP):
[0040] Location system prompts Under the constraints of editing instructions semantic information and input image Edit and plan to generate a set of location tips consisting of multiple location tips. And combined with reflective system prompts Filter out the best location prompts ;
[0041] Based on optimal location prompts and input image Generate a candidate localization mask set consisting of multiple candidate binary localization masks. And combined with reflective system prompts Filter out the optimal positioning mask ;
[0042] The second stage involves modifying cognitive processes (MCP):
[0043] Modified system prompts and optimal positioning mask Under the constraints of editing instructions semantic information of the input image The modification plan is modified, generating a set of modification prompts consisting of multiple modification prompts. And combined with reflective system prompts Filter out the best modification suggestions ;
[0044] In optimal modification prompts and optimal positioning mask Under the constraints, a candidate edited image set consisting of multiple edited images is generated. And combined with reflective system prompts Select the best edited image .
[0045] This invention designs a two-stage cognitive architecture comprising a Localization Cognition Process (LCP) and a Modification Cognition Process (MCP). The collaboration between the two stages enables semantic understanding and precise execution of editing instructions, improving the accuracy of high-level semantic reasoning and ensuring visual consistency. In the LCP stage, optimal localization hints are generated and selected through editing planning, and then a precise target region mask is generated and determined based on these hints. In the MCP stage, the optimal localization mask... Under the constraints of the optimal positioning mask and the optimal modification prompts, the modification plan is modified, and the optimal modification prompts are generated and selected. Finally, under the constraints of the optimal positioning mask and the optimal modification prompts, the optimal modification prompts are determined based on the input image. Multiple candidate results are generated, and the highest quality edited image is selected. This invention's method, by simulating human cognitive decision-making steps, decomposes the complex editing task into ordered localization and modification sub-tasks, effectively improving the depth of understanding of natural language instructions, the accuracy of editing region localization, and the semantic rationality and visual quality of the final image modification.
[0046] Location-based system prompts Reflective system prompts and modified system prompts All prompts are given in text form.
[0047] Furthermore, in the system prompt Under the constraints of editing instructions semantic information and input image Edit and plan to generate a set of location tips consisting of multiple location tips. The implementation method is as follows:
[0048] Location system prompts Under the constraints, large-scale multimodal models are based on editing instructions. semantic information and input image Perform cross-modal alignment to determine the relationship with the edit instructions. The relevant target area is identified, and a textual semantic description of that target area is provided to generate a location cue. The process of generating location prompts is repeated multiple times to obtain a set of location prompts. .
[0049] ;
[0050] in, This refers to large multimodal models, which are LMMs (large multimodal models).
[0051] In this preferred embodiment, a system prompt is provided. Under constraints, the model is based on receiving editing instructions. semantic information and input image The multimodal input is subjected to cross-modal alignment, and multiple candidate localization cues are output to form a localization cue set. This allows for the generation of multiple candidate textual location suggestions for image regions related to editing commands. Furthermore, the optimal location suggestion is selected. The implementation method is as follows:
[0052] In reflective system prompts Under the constraints, large-scale multimodal models provide localization cue sets. Location prompts in With the input image and editing instructions The matching degree between the constructed multimodal contexts is scored, and a score is obtained accordingly. The above matching degree scoring process is repeated multiple times to obtain multiple scores, and the positioning prompt with the highest score is selected as the optimal positioning prompt. Optimal positioning suggestion Used to semantically describe the target region related to the instructions, guiding the subsequent segmentation process, and .
[0053] This preferred embodiment selects the optimal location prompt. During the process, in reflective system prompts Under the constraints, the large multimodal model quantifies the matching degree between each prompt in the set of location prompts and the multimodal context composed of images and instructions by evaluating them one by one. The location prompt corresponding to the highest score is selected from all the scores, thereby filtering out the location prompt with the best fit with the editing intention and visual content.
[0054] Furthermore, based on the optimal positioning prompts and input image The implementation method for generating multiple candidate binary location masks is as follows:
[0055] Instruction segmentation model for input image With optimal positioning hints Perform semantic matching and output an initial binary localization mask. Then, a morphological dilation operation is applied to it to obtain the corresponding candidate binary localization mask. The process of obtaining candidate binary localization masks is repeated multiple times to obtain a candidate localization mask set consisting of multiple candidate binary localization masks. .
[0056] , ;
[0057] in, This represents the instruction segmentation model. This indicates the application of a morphological dilation operation.
[0058] In this preferred embodiment, the instruction segmentation model is based on optimal positioning prompts. semantic information and input image Perform matching, generating a different initial binary mask each time. Subsequently, morphological dilation is used to enhance the spatial coherence and context coverage of the mask, effectively mitigating issues such as mask fragmentation or omission of relevant regions, thereby generating a more reliable set of candidate localization masks. This provides a multi-granular and scalable spatial positioning foundation for subsequent editing operations.
[0059] Furthermore, the optimal positioning mask is selected. The implementation method is as follows:
[0060] In reflective system prompts Under constraints, large-scale multimodal models combined with optimal localization hints and input image The candidate localization mask set was evaluated from two dimensions: spatial localization accuracy and semantic relevance. The binary positioning masks in the dataset are evaluated comprehensively, and the binary positioning mask with the highest score is selected as the optimal positioning mask. .
[0061] In this embodiment, large-scale multimodal models provide feedback in reflective systems. Guided by this principle, and taking into account both the accuracy of spatial positioning and the fit of semantic association, the candidate positioning mask set is... A dual-dimensional automated evaluation and quantitative comparison is performed to select those that match the optimal positioning prompt in both geometric boundaries and conceptual meaning. and input image The best-matching mask serves as the basis for the final spatial operations.
[0062] Furthermore, in the modified system prompts and optimal positioning mask Under the constraints of editing instructions semantic information of the input image The modification plan is modified, generating a set of modification prompts consisting of multiple modification prompts. The implementation methods include:
[0063] Modified system prompts Hints, constraints, and optimal positioning mask Under the provided spatial constraints, large multimodal models respond to editing instructions With input image Perform cross-modal alignment to determine a modification hint for the target region. Repeat the above process of generating modification suggestions multiple times to obtain a set of modification suggestions. .
[0064] In this preferred embodiment, under dual constraints, the large multimodal model generates multiple modification suggestions through cross-modal alignment to form a modification suggestion set. Diversity allows for modification of the target region from different semantic perspectives and feature dimensions.
[0065] Furthermore, the optimal modification suggestions are selected. The implementation method is as follows:
[0066] In reflective system prompts Under the constraints, large-scale multimodal models are combined with editing instructions Input image and optimal positioning mask The modification prompt set was analyzed from two dimensions: semantic consistency and visual context. Modification prompts in Perform a comprehensive evaluation and select the highest-scoring modification suggestion as the optimal modification suggestion. ,and .
[0067] In this preferred embodiment, the optimal modification suggestion Used to guide subsequent editing processes, in reflective system prompts Guided by these factors, the large-scale multimodal model comprehensively evaluates various modification prompts and editing instructions. Editing intent semantic fit and its relationship with the optimal localization mask Input image under spatial constraints The consistency between the resulting visual contexts is used to quantitatively compare and select the optimal final modification scheme.
[0068] Furthermore, in the optimal modification prompt and optimal positioning mask Under the constraints, a candidate edited image set consisting of multiple edited images is generated. The implementation methods include:
[0069] Based on the optimal positioning mask Determine the input image The target area, and at the same time, the optimal modification prompt. Under constraints, the image restoration model processes the input image. The target region is conditionally generated to produce an edited image. The process of generating edited images is repeated multiple times to obtain a candidate edited image set consisting of multiple edited images. The noise parameters of the image restoration model are randomized during each generation of the edited image.
[0070] Indicates the input image Optimal mask And optimal modification tips Image restoration model under certain conditions.
[0071] In this preferred embodiment, the optimal modification prompt With optimal positioning mask Under the dual constraints, the target region is conditionally generated multiple times through an image inpainting model, thereby obtaining a set of candidate editable images that maintains semantic intent and spatial accuracy while also possessing visual diversity and randomness. .
[0072] Furthermore, the best edited image is selected. The implementation methods include:
[0073] In reflective system prompts Under the constraints, large-scale multimodal models are combined with editing instructions and input image The candidate editable image set was evaluated based on three aspects: semantic fidelity, visual quality, and consistency with the unedited region. The edited images are evaluated comprehensively, and the edited image with the highest score is selected as the best edited image. .
[0074] This preferred embodiment provides a reflective system prompt. Under the constraints, large-scale multimodal models comprehensively evaluate candidate images in three dimensions: semantic fidelity, visual quality, and consistency of unedited regions. They perform quantitative scoring and comparison to select the images that best match the editing intent and are most consistent with the original image (i.e., the input image). The final edited image that is most context-coordinated and of the highest quality.
[0075] Principle Analysis: The instruction-based image editing method and system based on cognitive reasoning adopts the dual-process theory in cognitive science, viewing image editing as a complex task requiring System 2 (i.e., a logical and deliberate reasoning process), rather than simply an intuitive, direct pixel transformation dominated by System 1. A structured cognitive reasoning mechanism is designed, decomposing the image editing process into two collaborative cognitive stages: "what to edit" and "how to edit." This ensures the editing process has good interpretability and maintains stable generalization performance in complex or ambiguous scenarios. Therefore, given an image and its corresponding editing instructions, the system achieves accurate and highly visually consistent image editing through a dual-cognitive-stage architecture encompassing perception, planning, and reflection.
[0076] In practical applications, download the open-source large-scale multimodal model Qwen2.5-VL-72B-Instruct, the language-guided instruction segmentation model LISA-13B, and the diffusion-based image inpainting model Flux-Inpainting. These models can all be replaced by other functionally equivalent open-source or closed-source models. Simultaneously, select an input image to be edited. and provide the corresponding editing instructions. In this embodiment of the invention, an image to be edited and editing instructions are required; the large multimodal model Qwen2.5-VL-72B-Instruct, the instruction segmentation model LISA-13B, and the diffusion-based image inpainting model Flux-Inpainting are downloaded and deployed; if the above models cannot be obtained, open-source or closed-source models that are equivalent in function and input / output form can be used as replacements, or the corresponding models can be trained from scratch.
[0077] Specific Implementation Method Two: The instruction-based image editing system based on cognitive reasoning described in this implementation method includes a storage device, a processor, and a computer program stored in the storage device and executable on the processor. The processor executes the computer program to implement the instruction-based image editing method based on cognitive reasoning as described.
[0078] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.
Claims
1. A command-based image editing method based on cognitive reasoning, characterized in that, include: The first stage, positioning the cognitive process: Location system prompts Under the constraints of editing instructions semantic information and input image Edit and plan to generate a set of location tips consisting of multiple location tips. And combined with reflective system prompts Filter out the best location prompts ; Based on optimal location prompts and input image Generate a candidate localization mask set consisting of multiple candidate binary localization masks. And combined with reflective system prompts Filter out the optimal positioning mask ; The second stage involves modifying the cognitive process: Modified system prompts and optimal positioning mask Under the constraints of editing instructions semantic information of the input image The modification plan is modified, generating a set of modification prompts consisting of multiple modification suggestions. And combined with reflective system prompts Filter out the best modification suggestions ; In optimal modification prompts and optimal positioning mask Under the constraints, a candidate edited image set consisting of multiple edited images is generated. And combined with reflective system prompts Select the best edited image .
2. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, In system prompts Under the constraints of editing instructions semantic information and input image Edit and plan to generate a set of location tips consisting of multiple location tips. The implementation method is as follows: Location system prompts Under the constraints, large-scale multimodal models are based on editing instructions. semantic information and input image Perform cross-modal alignment to determine the relationship with the edit instructions. The relevant target area is identified, and a textual semantic description of that target area is provided to generate a location cue. The process of generating location prompts is repeated multiple times to obtain a set of location prompts. .
3. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, Filter out the best location prompts The implementation method is as follows: In reflective system prompts Under the constraints, large-scale multimodal models provide localization cue sets. Location prompts in With the input image and editing instructions The matching degree between the constructed multimodal contexts is scored, and a score is obtained accordingly. The above matching degree scoring process is repeated multiple times to obtain multiple scores, and the positioning prompt with the highest score is selected as the optimal positioning prompt. .
4. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, Based on optimal location prompts and input image The method for generating multiple candidate binary location masks is as follows: Instruction segmentation model for input image With optimal positioning hints Perform semantic matching and output an initial binary localization mask. Then, a morphological dilation operation is applied to it to obtain the corresponding candidate binary localization mask. The process of obtaining candidate binary localization masks is repeated multiple times to obtain a candidate localization mask set consisting of multiple candidate binary localization masks. .
5. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, Filter out the optimal positioning mask The implementation method is as follows: In reflective system prompts Under constraints, large-scale multimodal models combined with optimal localization hints and input image The candidate localization mask set was evaluated from two dimensions: spatial localization accuracy and semantic relevance. The binary positioning masks in the dataset are evaluated comprehensively, and the binary positioning mask with the highest score is selected as the optimal positioning mask. .
6. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, Modified system prompts and optimal positioning mask Under the constraints of editing instructions semantic information of the input image The modification plan is modified, generating a set of modification prompts consisting of multiple modification suggestions. The implementation methods include...
7. In the modified system prompt Hints, constraints, and optimal positioning mask Under the provided spatial constraints, large multimodal models respond to editing instructions With input image Perform cross-modal alignment to determine a modification hint for the target region. Repeat the above process of generating modification suggestions multiple times to obtain a set of modification suggestions. . The instruction-based image editing method based on cognitive reasoning according to claim 1 is characterized in that, Filter out the best modification suggestions The implementation method is as follows: In reflective system prompts Under the constraints, large-scale multimodal models are combined with editing instructions Input image and optimal positioning mask The modification prompt set is analyzed from two dimensions: semantic consistency and visual context. Modification prompts in Perform a comprehensive evaluation and select the highest-scoring modification suggestion as the optimal modification suggestion. .
8. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, In optimal modification prompts and optimal positioning mask Under the constraints, a candidate edited image set consisting of multiple edited images is generated. The implementation methods include: Based on the optimal positioning mask Determine the input image The target area, and at the same time, the optimal modification prompt. Under constraints, the image restoration model processes the input image. The target region is conditionally generated to produce an edited image. The process of generating edited images is repeated multiple times to obtain a candidate edited image set consisting of multiple edited images. ;in, The noise parameters of the image restoration model are randomized each time an edited image is generated.
9. The instruction-based image editing method based on cognitive reasoning according to claim 1, characterized in that, Select the best edited image The implementation methods include: In reflective system prompts Under the constraints, large-scale multimodal models are combined with editing instructions and input image The candidate editable image set was evaluated based on three aspects: semantic fidelity, visual quality, and consistency with the unedited region. The edited images are evaluated comprehensively, and the edited image with the highest score is selected as the best edited image. .
10. A cognitive reasoning-based instruction-based image editing system, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that, The processor executes a computer program to implement the instruction-based image editing method based on cognitive reasoning as described in any one of claims 1 to 9.