Closed-loop emotion alignment image editing method and system based on multi-modal large model
By employing a closed-loop emotion alignment image editing method based on a multimodal large model, and utilizing emotion evaluators, reversible patches, and semantic-attention masks, the instability problem of emotional image editing in existing technologies is solved. This achieves controllability and content fidelity in emotion alignment, providing a highly reliable, controllable, and deployable multimodal emotional image editing system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHIPU PILOT TECHNOLOGY CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-10
AI Technical Summary
Existing emotional image editing technologies struggle to simultaneously satisfy the accuracy of emotion alignment, content fidelity, and controllability of the editing process. They lack stable closed-loop mechanisms and reversible semantic-attention constraints, resulting in unstable and unpredictable editing results.
We employ a closed-loop emotion alignment image editing method based on a multimodal large model. Through a feedback-driven closed-loop mechanism of emotion evaluator, combined with reversible patching and semantic-attention mask constraints, we achieve automatic correction of emotion bias and fine control of the editing range, thus constructing a reversible and pluggable cross-emotion semantic transformation capability.
It achieves controllable iteration and stable convergence of emotion editing, reduces user interaction costs, improves the hit rate of emotion editing, reduces identity drift and background mis-modification, and ensures content consistency and engineerable deployment.
Smart Images

Figure CN122368253A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology and relates to a closed-loop emotion alignment image editing method and system, especially a closed-loop emotion alignment image editing method and system based on a multimodal large model. Background Technology
[0002] With the rapid application of Multimodal Large Language Models (MLLM) and Diffusion Models in image understanding, image editing, advertising creativity, film and television post-production, social content production, and intelligent design, user demand for visual editing that "imbues emotions, atmosphere, and subjective biases" has increased significantly. Compared to traditional objective attribute editing such as "changing backgrounds, objects, and styles," emotional image editing emphasizes the control of subjective dimensions such as the tension of facial expressions, the relaxation of postures, lighting, color tone, atmosphere, and narrative bias of scenes. Examples include "changing anger to relaxation," "making the image less oppressive," "making it warmer but not too exaggerated," and "changing only the expression without altering the background." These demands are characterized by significant subjectivity, continuity, and iterability, placing three simultaneous hard requirements on image editing: accurate emotional alignment, stable non-emotional content, and a controllable and convergent editing process.
[0003] However, most existing emotional image editing systems are either "cue-driven single-round generation" or "weakly constrained attention guidance," making it difficult to simultaneously achieve accurate emotional alignment and high-fidelity content preservation, and they lack a stable, engineered, and deployable closed-loop mechanism. Specifically, existing emotional image editing technologies mainly include: 1. One-shot Editing solution.
[0004] Current mainstream image editing methods typically adopt the paradigm of "one input instruction → one generation result": the user provides a text prompt (e.g., "make her relaxed"), and the system directly generates the editing result using diffusion generation models such as img2img, inversion, or inpainting. This type of approach works for objective editing, but it has obvious inherent defects in emotion editing: (1) lack of precise definition and calibration of subjective goals: the same "relaxation" means different intensities and different expression carriers (e.g., facial expressions, postures, tones, background atmosphere, etc.) for different users, and it is difficult to hit the mark in one generation; (2) lack of convergent feedback mechanism: when the result is "not relaxed enough / too exaggerated / too cold", the user can only repeatedly change the prompt words to try and make mistakes, and the system cannot automatically "understand the deviation and correct it"; (3) lack of stopping criteria and protection mechanism: multiple attempts can easily lead to more and more deviations, identity drift, and background being mistakenly changed, making it impossible to form a stable and controllable editing process.
[0005] Therefore, single-round editing driven by single-round instructions often manifests as "usable but uncontrollable, capable of producing images but difficult to align" in emotional scenarios.
[0006] 2. Emotion guidance strategies based on prompts or light controls struggle to balance "emotional alignment" and "content authenticity".
[0007] Some solutions attempt to enhance emotional control through stronger prompt engineering, negative prompts, guidance scales, or simple local editing (such as inpainting). However, these methods usually have typical problems: (1) Emotional semantics are difficult to map stably to visual changes: Emotion is not a single visual attribute, but is often a combination of multiple factors (e.g., facial muscle groups + posture + lighting tone + contextual objects). Relying solely on text guidance can easily "change the atmosphere but not the expression" or "change the expression but not the overall feeling"; (2) Prompt words are not generalized stably to different models / different images: The same prompt has huge differences in effect on different characters, different lighting, and different compositions; (3) Lack of strong constraints on content preservation: Increasing the intensity of guidance often leads to greater structural drift, resulting in the mis-modification of identity, background objects, clothing details, etc.
[0008] Therefore, such solutions often struggle to achieve a controllable balance between "making the emotions more authentic" and "distorting the content."
[0009] 3. A general solution based on the diffusion editing framework can perform editing, but it does not provide sufficient support for subjective goals such as "emotion".
[0010] In the field of general image editing, there are several mature technical routes, such as: (1) Instruction-based editing: using instruction-to-image editing models (such as the instruct-pix2pix approach) to complete "editing the image according to instructions"; (2) Attention replacement / reweighting editing (Prompt-to-Prompt type): controlling semantic changes through cross-attention replacement / injection; (3) Conditional control editing (ControlNet, etc.): using conditions such as edge, pose, and depth to constrain the structure; (4) Differential editing (DiffEdit, etc.): locating the modification area through differential mask.
[0011] The closest point between these schemes and the present invention is that they all attempt to achieve controllable editing through "diffusion generation + conditional constraints". However, their core contradiction with emotion editing has not been fully resolved, specifically in the following ways: (1) Emotion goals cannot be accurately evaluated and converged: most schemes do not take "whether emotions are aligned" as the closed-loop optimization goal; (2) The editing region is unstable: attention heatmaps or differential masks are prone to drift on subjective emotion tasks; (3) They lack the ability to stabilize cross-emotion semantic transformation: the model may misunderstand "relaxation" as "fatigue / indifference", or be unstable in its performance on specific emotion pairs (anger → satisfaction).
[0012] Therefore, these general editing frameworks are more like "editing toolboxes," but lack systematic alignment and safeguards for emotional goals.
[0013] 4. The latest attempt at emotional editing.
[0014] In recent years, some new research attempts have emerged, which introduce multimodal large models into image editing and emphasize the control of emotions or subjective attributes. Existing research attempts usually have the following common characteristics: (1) Introducing MLLM as an "understanding and planning machine": first use multimodal large models to understand instructions and generate more specific editing descriptions, and then hand them over to the diffusion model for generation; (2) Introducing attention interaction or adapter fine-tuning: trying to better inject "emotion-related cues" into the diffusion model to reduce irrelevant content changes; (3) Constructing emotion evaluation indicators: using emotion recognition models or MLLM review to evaluate whether the generated result is correct in emotion.
[0015] These approaches are very close to the present invention in form because they also explore the combination of "multimodal large model + diffusion model" and begin to focus on the separation of "emotion-related / irrelevant content". However, they still have the following problems: (1) lack of closed-loop self-correction driven by emotion bias: even if there is emotion evaluation, the evaluation often stays at offline indicators or one-time selection, and does not form a convergent closed loop of "evaluation → bias → automatic correction → regeneration"; (2) lack of reversible and pluggable cross-emotion semantic conversion enhancement mechanism: usually adopts overall fine-tuning or fixed editing method, which makes it difficult to achieve "the ability to load different emotions on demand and the ability to roll back without polluting the base after use"; (3) lack of explicit double mask constraint system of semantic boundary + attention positioning: many methods mainly rely on soft attention guidance, which makes it difficult to provide stable engineering guarantees for strong constraints such as "do not change the background / do not change the identity / only change the local face".
[0016] Therefore, although these solutions are similar in research direction, they still cannot achieve the system-level capabilities of "convergence, rollback, and precise control" emphasized in this invention.
[0017] In summary, existing emotional image editing systems generally suffer from three fundamental deficiencies in engineering implementation: (1) The editing process is one-off and trial-and-error, lacking a closed-loop convergence mechanism driven by emotional bias feedback; (2) The ability to perform cross-emotional semantic conversion is not modularly managed, lacking pluggable and rollback-able enhancement methods, which can easily lead to generalization instability or model side effects; (3) The control of the editing area relies on soft guidance, lacking the joint constraint of "semantic boundary + emotional attention", making it difficult to ensure that emotionally irrelevant content such as identity, background, and objects are not mistakenly modified.
[0018] Given the aforementioned deficiencies in existing technologies, there is an urgent need for an emotional image editing method, system, device, and storage medium based on multimodal large model-based emotion bias closed-loop alignment, reversible patch-style model editing, and semantic-attention mask constraints. Summary of the Invention
[0019] To address the problems existing in the prior art, this invention proposes a closed-loop emotion alignment image editing method and system based on a multimodal large model. It achieves a comprehensive upgrade of emotional image editing from "single-round prompt-driven" to "closed-loop convergent alignment," and establishes an engineering-feasible balance mechanism between cross-emotional semantic transformation and content fidelity. It can automatically select and load appropriate emotion transformation patches based on the user's target emotion and constraints, and finely control the diffusion editing range through semantic-attention masks. Simultaneously, it utilizes emotion deviation feedback for continuous self-correction, ultimately generating edited images that stably match the target emotion while ensuring content consistency. This provides a new technical path for building a highly reliable, controllable, and deployable multimodal emotional image editing system.
[0020] To achieve the above objectives, the present invention provides the following technical solution: A closed-loop emotion alignment image editing method based on a multimodal large model, characterized by the following steps: S1: Receive input image With user instructions ; S2: Transfer the input image With user instructions Input a large multimodal model and output a structured editing plan, which includes: target emotion. Editing carrier Editing intensity and mandatory constraint set ; S3: Use an emotion evaluator to evaluate the input image. Perform an emotion assessment and output the current emotion expression. And based on the target emotion Editing carrier and current mood expression Constructing an index for mood transition ; S4: Retrieve the index related to the stated mood transition from the patch library. If the corresponding reversible patch is hit, the reversible patch is loaded into the multimodal large model; if it is not hit, a reversible patch is generated online and written into the patch library, and then the reversible patch is loaded into the multimodal large model, resulting in an enhanced multimodal model. S5: For the input image Perform semantic segmentation to obtain a semantic mask. ; S6: Extract attention heatmaps of lexical terms related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. And based on the attention heatmap Obtain attention mask ; S7: Merge the semantic mask With attention mask , to obtain the final edit mask and non-editable region mask ; S8: In the final edit mask and non-editable region mask Under the constraints of the input image Perform diffusion editing to generate candidate images ; S9: Evaluate the candidate images And based on the evaluation results and the target emotion Calculating emotional bias and based on the non-editable region mask Calculate content drift score ; S10: Based on the aforementioned emotional bias and content drift Determine whether the stopping criterion is met. If the stopping criterion is not met, perform self-correction and return to step S6 after self-correction. Continue until the stopping criterion is met or the number of iterations is reached to obtain the edited image.
[0021] Preferably, the semantic mask in step S5 Including facial region mask Body area mask and background area mask .
[0022] Preferably, step S6 specifically includes: S61: Extract attention heatmaps of lexical terms related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. ; S62: Regarding the attention heatmap The original pixel values at each coordinate are normalized to obtain the pixel values at each coordinate of the normalized heatmap:
[0023] In the formula, This represents the pixel value at coordinates (x, y) of the normalized heatmap. This is the attention heatmap. The original pixel value at coordinates (x, y). This is the attention heatmap. The minimum value of all original pixel values. This is the attention heatmap. The maximum value of all original pixel values; S63: Threshold the pixel values at each coordinate of the normalized heatmap to obtain binary results at each coordinate of the normalized heatmap:
[0024] In the formula, It is the binary result at coordinates (x, y) of the normalized heatmap; It is the set binarization segmentation threshold; S64: Binary results at each coordinate of the normalized heatmap By fusing the results, we obtain the attention mask. .
[0025] Preferably, step S7 specifically includes: S71: Employ intersection fusion to obtain the facial region mask. and attention mask , to obtain the final edit mask :
[0026] S72: Based on the final edit mask Obtain the mask of the non-editable region :
[0027] In the formula, This represents a mask matrix consisting entirely of 1s.
[0028] Preferably, step S9 specifically includes: S91: Use an emotion evaluator to evaluate the candidate images. Conduct situation assessment and output emotion expression. ; S92: Based on the aforementioned emotion expression and the target emotion Calculating emotional bias :
[0029] In the formula, This is a distance metric function for the emotional space. S93: Based on the non-editable region mask Calculate content drift score :
[0030] In the formula, This represents element-wise multiplication. It is a structural similarity index.
[0031] Preferably, the stopping criterion in step S10 includes a stopping condition and a convergence condition. The stopping criterion is considered satisfied when either the stopping condition or the convergence condition is met. The stopping condition is: , The convergence condition is: , In the formula, The preset emotion alignment threshold, The preset content drift limit, The preset deviation convergence threshold, .
[0032] Preferably, the self-correction process in step S10 specifically includes: adjusting the editing intensity, adjusting the binarization segmentation threshold, and... Reduce the adjusted editing intensity or increase the adjusted binarization segmentation threshold.
[0033] Furthermore, this invention also provides a closed-loop emotion alignment image editing system based on a multimodal large model, characterized in that it includes: The data receiving module is used to receive input images. With user instructions ; The instruction parsing module is used to process the input image. With user instructions Input a large multimodal model and output a structured editing plan, which includes: target emotion. Editing carrier Editing intensity and mandatory constraint set ; The current emotion assessment and emotion transformation index construction module is used to assess the input image using an emotion evaluator. Perform an emotion assessment and output the current emotion expression. And based on the target emotion Editing carrier and current mood expression Constructing an index for mood transition ; A retrieval and reversible patch loading module is used to retrieve the sentiment conversion index from the patch library. If the corresponding reversible patch is hit, the reversible patch is loaded into the multimodal large model; if it is not hit, a reversible patch is generated online and written into the patch library, and then the reversible patch is loaded into the multimodal large model, resulting in an enhanced multimodal model. A semantic mask generation module is used to generate a semantic mask from the input image. Perform semantic segmentation to obtain a semantic mask. ; An attention mask generation module is used to extract attention heatmaps of lexical units related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. And based on the attention heatmap Obtain attention mask ; The mask fusion module is used to fuse the semantic masks. With attention mask , to obtain the final edit mask and non-editable region mask ; Candidate image generation module, which is used in the final edit mask and non-editable region mask Under the constraints of the input image Perform diffusion editing to generate candidate images ; Candidate image evaluation module, which is used to evaluate the candidate images And based on the evaluation results and the target emotion Calculating emotional bias and based on the non-editable region mask Calculate content drift score ; The stopping criterion judgment module is used to determine the emotional deviation. and content drift Determine whether the stopping criteria are met; The self-correction module is used to perform self-correction when the stopping criteria are not met.
[0034] Furthermore, the present invention also provides a closed-loop emotion alignment image editing device based on a multimodal large model, characterized in that it includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the closed-loop emotion alignment image editing method based on a multimodal large model as described above.
[0035] Finally, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the program is executed by a processor, it implements the steps of the closed-loop emotion alignment image editing method based on a multimodal large model as described above.
[0036] Compared with existing sentiment-based image editing schemes that rely on single-round instruction-driven diffusion editing, soft-guided attention control, and overall fine-tuning or pure prompt-word parameter tuning, the closed-loop sentiment alignment image editing method and system proposed in this invention, based on a multimodal large model, has the following significant advantages in terms of sentiment alignment accuracy, content fidelity and stability, controllability and rollback capability, and engineering deployability: (1) Upgraded from one-time generation to emotional bias closed-loop alignment, significantly improving emotional controllability and stable convergence ability.
[0037] Traditional emotional image editing typically employs a single-round process of "one input prompt → one result generation." When encountering deviations such as "not relaxed enough / too exaggerated / the atmosphere is wrong," it relies solely on the user repeatedly modifying the prompts for trial and error. This lack of measurable alignment signals and automatic correction mechanisms leads to unstable and unpredictable results. This invention introduces an emotion evaluator after generation to assess the current emotional representation and calculate its deviation from the target emotion. This deviation serves as feedback to automatically adjust the next round of editing conditions (editing intensity, mask threshold, iteration strategy, etc.), while setting explicit stopping criteria (emotional target achieved, deviation convergence, content drift protection, maximum iteration limit, etc.).
[0038] This mechanism represents a fundamental shift from "prompt-based trial-and-error editing" to "feedback-enabled, convergent closed-loop alignment editing." It can automatically approach the target and stabilize under subjective emotional goals, significantly reducing user interaction costs and improving the hit rate of emotional editing.
[0039] (2) Upgrade from overall capability drift to reversible patch-type emotion transformation enhancement, realizing pluggable and rollback of cross-emotion semantic transformation.
[0040] Existing solutions either rely on pure prompts, leading to unstable cross-emotion semantic conversion capabilities, or require end-to-end fine-tuning, which can easily result in degradation of general capabilities, interference between different emotion tasks, and high version maintenance costs. This invention encapsulates cross-emotion semantic conversion capabilities as reversible patches (sparse / low-rank incremental form) and establishes a patch library managed by indexing "emotion pairs + editing carriers": during inference, patches are loaded on demand to enhance the corresponding emotion conversion capabilities, and after the task is completed, they are unloaded and rolled back to restore the base model parameters.
[0041] This mechanism can quickly switch between different emotion transition tasks and maintain the stability of the base capabilities, achieving the engineering deployment advantage of "capability enhancement and controllable side effects": it not only improves the consistency of transitions between specific emotion pairs (such as anger → relaxation, pleasure → disgust), but also avoids the uncontrollable drift and maintenance complexity caused by long-term fine-tuning.
[0042] (3) The soft guidance region control is upgraded to semantic-attention dual mask hard constraint, which significantly reduces the risk of identity drift and background mismodification.
[0043] Existing diffusion editing methods often rely on attention or implicit guidance to determine the modification area. When faced with strong constraints such as "change only the expression, not the background," "keep the identity unchanged," or "don't move a certain object," errors can easily occur: character identity features shift, background elements change, clothing details are altered, and emotional expressions are applied to the wrong areas. This invention introduces both semantic and attention masks: the semantic mask provides structured constraints on object boundaries and uneditable areas (faces / people / backgrounds / objects, etc.), while the attention mask provides the location of the target emotion-related lexical token's effective area. The two are fused to obtain the final editing mask. During the diffusion process, freezing or strong consistency constraints are applied to the area outside the mask, while emotional conditional injection is enhanced into the area inside the mask.
[0044] This mechanism achieves "explicit control over the editing scope," ensuring that "what should be changed is changed, and what should not be changed remains unchanged," significantly reducing identity drift and background mis-modification, and improving the interpretability, consistency, and commercial usability of the results. It is especially suitable for high-fidelity scenarios such as portrait retouching, advertising materials, and brand visuals.
[0045] Through the above innovations, this invention achieves a holistic technological upgrade across three levels: editing workflow organization, cross-emotional semantic capability management, and spatial constraint control. The editing workflow has shifted from one-time generation to closed-loop convergence driven by emotion bias; the emotion transformation capability has changed from uncontrollable overall fine-tuning to pluggable and rollback-enabled patch-based enhancement; and the editing scope control has shifted from soft guidance to hard control through joint constraints of semantic boundaries and emotional attention. With these improvements, this invention enables the emotional image editing system to enhance emotion alignment while maintaining content fidelity, stability, controllability, and ease of deployment and maintenance, providing a new technical path for building a highly reliable and implementable multimodal emotional visual editing system. Attached Figure Description
[0046] Figure 1 This is a flowchart of the closed-loop emotion alignment image editing method based on a multimodal large model according to the present invention.
[0047] Figure 2 This is a schematic diagram of the structure of the closed-loop emotion alignment image editing system based on a multimodal large model according to the present invention.
[0048] Figure 3 This is a structural block diagram of the closed-loop emotion alignment image editing device based on a multimodal large model according to the present invention. Detailed Implementation
[0049] Before detailing any embodiment of the invention, it should be understood that the invention, in its application, is not limited to the details of the construction and arrangement of the components set forth in the following description or illustrated in the following figures. The invention can have other embodiments and can be practiced or carried out in various ways. Furthermore, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered limiting. The use of “comprising” or “having” and variations thereof in this invention is intended to cover the items set forth below and their equivalents, as well as any additional items. Unless otherwise specified or limited, the terms “installation,” “connection,” “support,” and “linkage,” and variations thereof are used broadly and cover both direct and indirect installation, connection, support, and linking. Moreover, “connection” and “linkage” are not limited to physical or mechanical connections or links.
[0050] Furthermore, firstly, in the disclosure of this invention, the terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting this invention. Secondly, the term "a" should be understood as "at least one" or "one or more," that is, in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple. The term "a" should not be construed as a limitation on the quantity.
[0051] To address the problems existing in the prior art, this invention proposes a closed-loop emotion alignment image editing method and system based on a multimodal large model. By making the emotion alignment process closed-loop, modularizing and rollbackable emotion conversion capabilities, and making the editing area constraints explicit, it upgrades the emotion editing from a "single-round, uncontrollable, and drift-prone" system to a "multi-round convergent, pluggable, and precisely controllable" system.
[0052] To this end, the present invention introduces three key technological innovations in its system: (1) Closed-loop alignment and self-correction mechanism driven by emotional bias.
[0053] This invention introduces an emotion evaluator after each generation to assess the emotion of the result and calculate the deviation between the target emotion and the current emotion. This deviation is then used as a feedback signal to automatically update the editing conditions for the next round (editing intensity, mask threshold, iteration strategy, etc.). At the same time, executable stopping criteria are set (emotion target achievement, deviation convergence, content drift protection, maximum iteration limit, etc.), thereby achieving controllable iteration and stable convergence of emotion editing, significantly reducing the cost of repeated trial and error for users.
[0054] (2) A mechanism for enhancing cross-emotional semantic transformation through reversible patch-based model editing.
[0055] This invention constructs an emotion conversion patch library, extracting cross-emotion semantic conversion capabilities from the basic multimodal model in the form of "sparse / low-rank incremental patches," enabling on-demand loading and post-use unloading. This mechanism improves the stability of emotion conversions such as "anger → relaxation" and "pleasure → disgust" while avoiding long-term side effects on the general capabilities of the basic model, supporting rapid switching and maintainable deployment across multiple emotions and scenarios.
[0056] (3) Diffusion editing region constraint mechanism of semantic-attention dual mask fusion.
[0057] This invention generates semantic masks through semantic segmentation to determine object boundaries (such as faces, people, backgrounds, objects, etc.). At the same time, it generates attention masks by utilizing the cross-attention response of target emotion-related tokens to locate the area where the emotion should be applied. The two are then fused to obtain the final editing mask. During the diffusion editing process, the area outside the mask is frozen or subject to strong consistency constraints, while the area inside the mask is enhanced with emotion condition injection. This mechanism ensures that "what should be changed is changed, and what should not be changed remains unchanged," reducing the risk of identity drift and background mis-modification, and improving the interpretability and controllability of the editing results.
[0058] Through the aforementioned technological innovations, this invention achieves a comprehensive upgrade of emotional image editing from "single-round prompt-driven" to "closed-loop convergent alignment," and establishes an engineerable balance mechanism between cross-emotional semantic transformation and content fidelity. This invention can automatically select and load appropriate emotion transformation patches based on the user's target emotion and constraints, and uses semantic-attention masks to finely control the diffusion editing range. Simultaneously, it utilizes emotion deviation feedback for continuous self-correction, ultimately generating edited images that stably match the target emotion while ensuring content consistency. This provides a new technical path for building a highly reliable, controllable, and deployable multimodal emotional image editing system.
[0059] Figure 1 A flowchart of the closed-loop emotion alignment image editing method based on a multimodal large model of the present invention is shown. Figure 1 As shown, the closed-loop emotion alignment image editing method based on a multimodal large model of the present invention includes the following steps: S1: Data reception.
[0060] In this invention, the first step is to receive an input image. With user instructions The input image This is an image that requires emotional image editing. The user instructions... This expresses the user's need for emotional editing, such as, "Change this person from angry to relaxed, but don't change the background."
[0061] Among them, emotional image editing refers to editing images according to user instructions while maintaining the main semantic content and structure of the original image. For input image Emotional image editing is the process of adjusting the emotional expression or atmosphere of an image. It includes both modifying local emotional carriers such as facial expressions and posture tension, and modifying overall emotional carriers such as lighting, tone, contrast, and background atmosphere. The goal is to make the output image more in line with the user's subjective emotional expectations, while avoiding accidental modification of content unrelated to emotion (such as the person's identity, background objects, clothing textures, etc.).
[0062] S2: Instruction parsing.
[0063] The input image With user instructions Input a large multimodal model, output a structured editing plan The structured editing plan Includes: target emotions Editing carrier Editing intensity and mandatory constraint set .
[0064] For the user instruction exemplified in step S1 The target emotion obtained For "relaxation", editing carrier "Focusing on facial expressions" and editing intensity For "medium" set of mandatory constraints The settings are "background frozen, identity preserved, and clothing preserved".
[0065] S3: Construction of current emotion assessment and emotion transition index.
[0066] The input image is evaluated using an emotion evaluator. Perform an emotion assessment and output the current emotion expression. And based on the target emotion. Editing carrier and current mood expression Constructing an index for mood transition The emotion transition index Used for reversible patch retrieval and loading.
[0067] The user instructions exemplified in step S1 It can be seen that the emotion evaluator is used to evaluate the input image. The assessment outputs the current sentiment representation. For "anger".
[0068] In this invention, the emotion evaluator can be a pre-trained multimodal large model, whose input is an image with emotion and whose output is the emotion representation in the image.
[0069] S4: Retrieval and reversible patch loading.
[0070] Based on the emotion transition index Retrieve the sentiment transition index from the patch library. Corresponding reversible patch If a match is found, that is, if a match is found with the aforementioned emotion transition index... Corresponding reversible patch If a match is found, the reversible patch is loaded. If no match is found, a reversible patch is generated online, written to the patch library, and then loaded into the multimodal large model, resulting in an enhanced multimodal model. The patch loading process is reversible; at the end of the process, the patch can be unloaded to restore the multimodal model to its original state.
[0071] The reversible patch is typically a sparse or low-rank parameter increment (e.g., LoRA / adapter / sparse increment matrix) used to improve the semantic conversion consistency of specific emotion pairs under specific editing vectors. Reversibility means that the multimodal model can recover its original parameter state after the reversible patch is unloaded, thus avoiding the drift of general capabilities caused by long-term fine-tuning and facilitating rapid switching and engineering maintenance between multiple emotion tasks.
[0072] In this invention, the cross-emotion semantic conversion capability is encapsulated as a reversible patch (sparse / low-rank incremental form), and a patch library is established and managed by indexing "emotion pair + editing carrier": during inference, the patch is loaded on demand to enhance the corresponding emotion conversion capability, and after the task is completed, it is unloaded and rolled back to restore the base model parameters.
[0073] This invention constructs a patch library that extracts cross-emotional semantic conversion capabilities from the basic multimodal model in the form of "sparse / low-rank incremental patches," enabling on-demand loading and post-use unloading. This mechanism improves the stability of emotion conversions such as "anger → relaxation" and "pleasure → disgust" while avoiding long-term side effects on the general capabilities of the basic model, supporting rapid switching and maintainable deployment across multiple emotions and scenarios.
[0074] S5: Semantic mask generation.
[0075] The input image is processed using a semantic segmentation model. Perform semantic segmentation to obtain a semantic mask. .in, For facial area masking, Body area mask, for Used as a mask for the background area.
[0076] According to the exemplary set of mandatory constraints in step S2 "Background frozen, identity preserved, clothing preserved" can be tagged. This is a non-editable area.
[0077] Semantic masks, in this context, refer to region masks generated by semantic segmentation or object detection models. They are used to represent the boundaries and categories of different semantic objects or regions in an input image, such as faces, bodies, backgrounds, and key objects. In this invention, semantic masks are used to clearly define the boundaries between "editable regions" and "non-editable regions," particularly for achieving spatial control under strong constraints such as "keeping the background unchanged," "maintaining the identity unchanged," and "not modifying a certain object."
[0078] S6: Attention mask generation.
[0079] Extracting the target emotion from the cross-attention mechanism of the enhanced multimodal model. Attention heatmap of related lexical tokens And based on the attention heatmap Obtain attention mask .
[0080] Among them, attention masking refers to targeting emotion. Region masks generated by cross-attention responses of relevant text lexical tokens in a multimodal large model to represent target sentiment. The attention mask represents the area to be affected in the image space. It reflects the distribution of attention within the model to the "location where emotion should be modified," and can be obtained through attention heatmap normalization, thresholding, etc. In this invention, the attention mask is used to locate key regions of emotional expression (e.g., facial expression areas or background atmosphere areas) and is fused with a semantic mask to improve localization stability.
[0081] In this invention, attention mask generation specifically includes: 1. Extracting the target emotion from the cross-attention mechanism of the enhanced multimodal model. Attention heatmap of related lexical terms .
[0082] 2. Regarding the attention heatmap The original pixel values at each coordinate are normalized to obtain the pixel values at each coordinate of the normalized heatmap:
[0083] In the formula, This represents the pixel value at coordinates (x, y) of the normalized heatmap. This is the attention heatmap. The original pixel value at coordinates (x, y). This is the attention heatmap. The minimum value of all original pixel values. This is the attention heatmap. The maximum value of all original pixel values.
[0084] 3. Threshold the pixel values at each coordinate of the normalized heatmap to obtain binary results at each coordinate of the normalized heatmap:
[0085] In the formula, It is the binary result at coordinates (x, y) of the normalized heatmap; It is the set binarization segmentation threshold.
[0086] 4. Binary results for each coordinate point of the normalized heatmap. By fusing the results, we obtain the attention mask. .
[0087] S7: Masking Fusion of the semantic mask With attention mask , to obtain the final edit mask and non-editable region mask .
[0088] In this invention, mask fusion specifically includes: 1. Since the editing medium primarily focuses on facial expressions and is subject to mandatory constraints including background freezing, the facial region mask is obtained through intersection fusion. and attention mask , to obtain the final edit mask :
[0089] 2. Based on the final edit mask Obtain the mask of the non-editable region :
[0090] In the formula, This represents a mask matrix consisting entirely of 1s.
[0091] S8: Candidate image generation.
[0092] In the final edit mask and non-editable region mask Under the constraints of the input image Perform diffusion editing to generate candidate images The constraint method can be: [to...] Apply stronger consistency constraints or latent locking to the corresponding regions; Enhance emotional conditioning in the corresponding areas.
[0093] In this invention, a diffusion-based image editing model can be used to edit the input image. Emotional image editing is performed to generate the candidate images. The diffusion-based image editing model refers to a model that implements image editing based on a diffusion generation framework, including forms such as img2img, inversion+regeneration, and inpainting. This model generates output images through an iterative denoising process and can receive control signals such as textual conditions, multimodal conditions, and spatial masks. In this invention, the diffusion-based image editing model generates candidate images under the editing conditions output by the multimodal large model, reversible patch enhancement capabilities, and final editing mask constraints.
[0094] S9: Candidate image evaluation.
[0095] Evaluate the candidate images And based on the evaluation results and the target emotion Calculating emotional bias and based on the non-editable region mask Calculate content drift score .
[0096] In this invention, candidate image evaluation includes: 1. Evaluate the candidate images using an emotion evaluator. Conduct situation assessment and output emotion expression. .
[0097] 2. Based on the aforementioned emotion expression and the target emotion Calculating emotional bias :
[0098] In the formula, This is a distance metric function for the emotional space (e.g., Euclidean distance or cosine distance).
[0099] 3. Based on the aforementioned non-editable region mask Calculate content drift score :
[0100] In the formula, This represents element-wise multiplication. It is a structural similarity index.
[0101] The content drift Used to express a measure of difference in non-editable regions.
[0102] S10: Stopping criterion judgment.
[0103] Based on the aforementioned emotional bias and content drift Determine whether the stopping criterion is met. If the stopping criterion is met, output the candidate image as the edit image. If the stopping criterion is not met, perform self-correction and return to step S6 after self-correction. Continue until the stopping criterion is met or the number of iterations is reached, and output the candidate image generated in the last iteration as the edit image.
[0104] In this invention, the stopping criterion includes a stopping condition and a convergence condition. The stopping criterion is considered satisfied when either the stopping condition or the convergence condition is met. The stopping condition is: , The convergence condition is: , In the formula, The preset emotion alignment threshold, The preset content drift limit, The preset deviation convergence threshold, .
[0105] If none of the above conditions are met and the iteration rounds are... Less than the maximum number of iterations (In this invention, If it can be set to 3), then self-correction will be performed.
[0106] In this invention, the self-correction includes: 1. Adjust editing intensity:
[0107] In the formula, To generate candidate images The editing intensity at that time, that is, the editing intensity obtained by parsing the instruction in step S2. ; Adjust the step size for editing intensity; To map the deviation to a monotonic function of the edit intensity adjustment.
[0108] 2. Adjust the binarization segmentation threshold (expand or shrink the mask):
[0109] In the formula, To generate candidate images The binarization segmentation threshold at that time, i.e., the binarization segmentation threshold set in step S6. ; This is the adjustment step size for the binarization segmentation threshold; This is a mapping function that maps the bias to the binarization segmentation threshold adjustment.
[0110] 3. If the content drift is too large, a protection rollback will be triggered: That is, if Then reduce or increase .
[0111] Based on the adjusted and Proceed to the next round and regenerate from step six. , integration and Generate candidate images and calculate
[0112] If satisfied or If the iteration stops, the candidate image is output. Edit the image; otherwise, continue iterating until the number of iterations reaches [number]. Or it may satisfy the convergence condition.
[0113] After exporting the edited image, you can uninstall the reversible patch. This allows the multimodal model to be restored from the enhanced multimodal large model to the original multimodal large model.
[0114] Figure 2 A schematic diagram of the closed-loop emotion alignment image editing system based on a multimodal large model according to the present invention is shown. Figure 2 As shown, the closed-loop emotion alignment image editing system based on a multimodal large model of the present invention includes: 1. Data receiving module.
[0115] The data receiving module is used to receive input images. With user instructions .
[0116] 2. Instruction parsing module.
[0117] The instruction parsing module is used to process the input image. With user instructions Input a large multimodal model and output a structured editing plan, which includes: target emotion. Editing carrier Editing intensity and mandatory constraint set .
[0118] 3. Current emotion assessment and emotion transformation index construction module.
[0119] The current emotion assessment and emotion transition index construction module is used to assess the input image using an emotion evaluator. Perform an emotion assessment and output the current emotion expression. And based on the target emotion Editing carrier and current mood expression Constructing an index for mood transition .
[0120] 4. Retrieval and reversible patch loading module.
[0121] The retrieval and reversible patch loading module is used to retrieve the emotion transition index from the patch library. If the corresponding reversible patch is hit, the reversible patch is loaded into the multimodal large model; if it is not hit, a reversible patch is generated online and written into the patch library, and then the reversible patch is loaded into the multimodal large model, resulting in an enhanced multimodal model.
[0122] 5. Semantic mask generation module.
[0123] The semantic mask generation module is used to process the input image. Perform semantic segmentation to obtain a semantic mask. .
[0124] 6. Attention mask generation module.
[0125] The attention mask generation module is used to extract attention heatmaps of lexical units related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. And based on the attention heatmap Obtain attention mask .
[0126] 7. Mask fusion module.
[0127] The mask fusion module is used to fuse the semantic mask. With attention mask , to obtain the final edit mask and non-editable region mask .
[0128] 8. Candidate image generation module.
[0129] The candidate image generation module is used in the final edit mask. and non-editable region mask Under the constraints of the input image Perform diffusion editing to generate candidate images .
[0130] 9. Candidate Image Evaluation Module.
[0131] The candidate image evaluation module is used to evaluate the candidate images. And based on the evaluation results and the target emotion Calculating emotional bias and based on the non-editable region mask Calculate content drift score .
[0132] 10. Stop Criterion Judgment Module.
[0133] The stopping criterion judgment module is used to determine the emotional deviation. and content drift Determine whether the stopping criteria are met.
[0134] 11. Self-correcting module.
[0135] The self-correction module is used to perform self-correction when the stopping criteria are not met.
[0136] Furthermore, this invention also provides a closed-loop emotion alignment image editing device based on a multimodal large model. For example... Figure 3 As shown, the closed-loop emotion alignment image editing device based on a multimodal large model of the present invention includes: a memory 11 for storing one or more programs; one or more processors 12; when the one or more programs are executed by the one or more processors 12, the one or more processors 12 implement the closed-loop emotion alignment image editing method based on a multimodal large model of the present invention.
[0137] Finally, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the closed-loop emotion alignment image editing method based on a multimodal large model in the present invention.
[0138] The computer-readable storage medium includes both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device. As defined in this invention, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0139] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0140] The steps of the methods or algorithms described in conjunction with the embodiments disclosed in this invention can be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. Those skilled in the art can modify or make equivalent substitutions to the technical solutions of the present invention based on the concept of the present invention, without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A closed-loop emotion alignment image editing method based on a multimodal large model, characterized in that, Includes the following steps: S1: Receive input image With user instructions ; S2: Transfer the input image With user instructions Input a large multimodal model and output a structured editing plan, which includes: target emotion. Editing carrier Editing intensity and mandatory constraint set ; S3: Use an emotion evaluator to evaluate the input image. Perform an emotion assessment and output the current emotion expression. And based on the target emotion Editing carrier and current mood expression Constructing an index for mood transition ; S4: Retrieve the index related to the stated mood transition from the patch library. If the corresponding reversible patch is hit, the reversible patch is loaded into the multimodal large model; if it is not hit, a reversible patch is generated online and written into the patch library, and then the reversible patch is loaded into the multimodal large model, resulting in an enhanced multimodal model. S5: For the input image Perform semantic segmentation to obtain a semantic mask. ; S6: Extract attention heatmaps of lexical terms related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. And based on the attention heatmap Obtain attention mask ; S7: Merge the semantic mask With attention mask , to obtain the final edit mask and non-editable region mask ; S8: In the final edit mask and non-editable region mask Under the constraints of the input image Perform diffusion editing to generate candidate images ; S9: Evaluate the candidate images And based on the evaluation results and the target emotion Calculating Emotional Bias and based on the non-editable region mask Calculate content drift score ; S10: Based on the aforementioned emotional bias and content drift Determine whether the stopping criterion is met. If the stopping criterion is not met, perform self-correction and return to step S6 after self-correction. Continue until the stopping criterion is met or the number of iterations is reached to obtain the edited image.
2. The closed-loop emotion alignment image editing method based on a multimodal large model according to claim 1, characterized in that, The semantic mask in step S5 Including facial region mask Body area mask and background area mask .
3. The closed-loop emotion alignment image editing method based on a multimodal large model according to claim 2, characterized in that, Step S6 specifically includes: S61: Extract attention heatmaps of lexical terms related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. ; S62: Regarding the attention heatmap The original pixel values at each coordinate are normalized to obtain the pixel values at each coordinate of the normalized heatmap: In the formula, This represents the pixel value at coordinates (x, y) of the normalized heatmap. This is the attention heatmap. The original pixel value at coordinates (x, y). This is the attention heatmap. The minimum value of all original pixel values. This is the attention heatmap. The maximum value of all original pixel values; S63: Threshold the pixel values at each coordinate of the normalized heatmap to obtain binary results at each coordinate of the normalized heatmap: In the formula, It is the binary result at coordinates (x, y) of the normalized heatmap; It is the set binarization segmentation threshold; S64: Binary results at each coordinate of the normalized heatmap By fusing the results, we obtain the attention mask. .
4. The closed-loop emotion alignment image editing method based on a multimodal large model according to claim 3, characterized in that, Step S7 specifically includes: S71: Employ intersection fusion to obtain the facial region mask. and attention mask , to obtain the final edit mask : S72: Based on the final edit mask Obtain the mask of the non-editable region : In the formula, This represents a mask matrix consisting entirely of 1s.
5. The closed-loop emotion alignment image editing method based on a multimodal large model according to claim 4, characterized in that, Step S9 specifically includes: S91: Use an emotion evaluator to evaluate the candidate images. Conduct situation assessment and output emotion expression. ; S92: Based on the aforementioned emotion expression and the target emotion Calculating Emotional Bias : In the formula, This is a distance metric function for the emotional space. S93: Based on the non-editable region mask Calculate content drift score : In the formula, This represents element-wise multiplication. It is a structural similarity index.
6. The closed-loop emotion alignment image editing method based on a multimodal large model according to claim 5, characterized in that, The stopping criteria in step S10 include stopping conditions and convergence conditions. The stopping criteria are considered satisfied when either the stopping condition or the convergence condition is met. The stopping condition is: , The convergence condition is: , In the formula, The preset emotion alignment threshold, The preset content drift limit, The preset deviation convergence threshold, .
7. The closed-loop emotion alignment image editing method based on a multimodal large model according to claim 6, characterized in that, The self-correction process in step S10 specifically includes: adjusting the editing intensity, adjusting the binarization segmentation threshold, and... Reduce the adjusted editing intensity or increase the adjusted binarization segmentation threshold.
8. A closed-loop emotion alignment image editing system based on a multimodal large model, characterized in that, include: The data receiving module is used to receive input images. With user instructions ; The instruction parsing module is used to process the input image. With user instructions Input a large multimodal model and output a structured editing plan, which includes: target emotion. Editing carrier Editing intensity and mandatory constraint set ; The current emotion assessment and emotion transformation index construction module is used to assess the input image using an emotion evaluator. Perform an emotion assessment and output the current emotion expression. And based on the target emotion Editing carrier and current mood expression Constructing an index for mood transition ; A retrieval and reversible patch loading module is used to retrieve the sentiment conversion index from the patch library. If the corresponding reversible patch is hit, the reversible patch is loaded into the multimodal large model; if it is not hit, a reversible patch is generated online and written into the patch library, and then the reversible patch is loaded into the multimodal large model, resulting in an enhanced multimodal model. A semantic mask generation module is used to generate a semantic mask from the input image. Perform semantic segmentation to obtain a semantic mask. ; An attention mask generation module is used to extract attention heatmaps of lexical units related to the target emotion from the cross-attention mechanism of the enhanced multimodal model. And based on the attention heatmap Obtain attention mask ; The mask fusion module is used to fuse the semantic masks. With attention mask , to obtain the final edit mask and non-editable region mask ; Candidate image generation module, which is used in the final edit mask and non-editable region mask Under the constraints of the input image Perform diffusion editing to generate candidate images ; Candidate image evaluation module, which is used to evaluate the candidate images And based on the evaluation results and the target emotion Calculating Emotional Bias and based on the non-editable region mask Calculate content drift score ; The stopping criterion judgment module is used to determine the emotional deviation. and content drift Determine whether the stopping criteria are met; The self-correction module is used to perform self-correction when the stopping criteria are not met.
9. A closed-loop emotion alignment image editing device based on a multimodal large model, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the closed-loop emotion alignment image editing method based on a multimodal large model as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the closed-loop emotion alignment image editing method based on a multimodal large model as described in any one of claims 1-7.