Multi-agent automatic picture retouching system based on content analysis

Through a multimodal large model and multi-agent collaborative framework, we achieve a deep understanding of images and a closed-loop process for automatic image retouching, solving the cost and efficiency issues of high-frequency, high-consistency image updates for small and medium-sized enterprises, and improving image quality and update efficiency.

CN120612259AActive Publication Date: 2025-09-09信华信(大连)软件服务股份有限公司

Patent Information

Application Number
CN202511105777.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-09
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve high-frequency and high-consistency image updates, which is especially costly in small and medium-sized enterprises. Furthermore, there is a lack of comprehensive understanding of image content and semantic levels, as well as automatic image retouching strategies, resulting in insufficient accuracy and automation levels in image restoration systems.

Method used

A multimodal large model is used for joint semantic-visual analysis, combined with a multi-agent collaboration framework, including a semantic analysis agent, a strategy selection agent, and a path planning agent. This realizes a closed-loop process from image semantic understanding to restoration execution, generates structured retouching strategies, and plans the retouching task sequence through a multi-agent collaboration mechanism, supporting local scoring feedback and dynamic task adjustment.

Benefits of technology

It achieves deep understanding of image content and extraction of semantic structure, automatically generates targeted image retouching strategies, reduces dependence on professional photography and post-processing, improves image quality and update efficiency, and is suitable for industry scenarios with batch image processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612259A_ABST
    Figure CN120612259A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent automatic image retouching system based on content analysis, and the method comprises the steps: 1, receiving the input of an original image, calling a finely-adjusted multi-mode large model to carry out the combined semantic-visual analysis of the image, extracting key semantic elements in the image, and carrying out the recognition of the key semantic elements; comprising but not limited to background coordination degree, illumination condition, figure hair style definition, weather quality and composition layout rationality content information. Based on these information, the system performs comprehensive image quality scoring on the image, which combines image style, aesthetic level, definition and subjective quality multi-dimensional evaluation criteria. Meanwhile, based on a multi-dimensional image quality scoring result and understanding of image semantic content, the system automatically generates a repair and optimization strategy set for specific defects so as to guide a subsequent image quality enhancement process. The invention relates to the field of multi-modal model and multi-agent cooperation, and can meet the increasing requirements of rapid optimization and high-quality propagation of image contents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal models and multi-agent collaboration, and in particular to a multi-agent automatic image retouching system based on content analysis. Background Art

[0002] With the increasing popularity of digital media, images, as core vehicles for information dissemination, brand promotion, and user engagement, are widely used in a variety of scenarios, including e-commerce platforms, social media, brochures, and video covers. This is particularly true in service industries like catering, hairdressing, beauty salons, and photography studios, where image quality and visual expression play a crucial role in attracting customers and demonstrating service quality. Consequently, these industries generally require periodic updates to promotional images.

[0003] The traditional image production process typically involves shooting, screening, retouching and color grading, and aesthetic optimization, relying heavily on the experience and skills of professional photographers and post-production retouchers. For small and medium-sized businesses, outsourcing or building their own retouching teams is costly, inefficient, and difficult to achieve frequent and consistent image updates. Even with the use of AI-powered filter tools or automated beautification apps, these tools often suffer from single-processing, uncontrollable results, fragmented image styles, and a lack of specificity, failing to truly meet the dual requirements of "aesthetics and practicality" for commercial applications.

[0004] While current automated image restoration and enhancement technologies have made significant algorithmic progress, most remain limited to single-point improvements to local image attributes (such as exposure, contrast, and noise reduction), lacking a comprehensive understanding of image content, semantic hierarchy, and target scene. For example, when capturing a barbershop hairstyle display, the system must not only enhance brightness and contrast but also identify facial areas, hairstyle details, and background clutter, and accordingly determine whether to apply background blur, skin texture optimization, or position adjustments. Current systems generally lack the ability to reason about this "content-based, goal-oriented" automatic image retouching strategy.

[0005] The emergence of large multimodal models provides a technical foundation for addressing these issues. By jointly learning images and text, these models possess cross-modal understanding and expression capabilities, capable of identifying complex semantic information from images and generating analysis results and optimization recommendations in natural language. However, these recommendations are often unstructured text, requiring manual interpretation, screening, and manipulation, and cannot directly drive automated execution of the image restoration process.

[0006] At the same time, recent advances in intelligent agent systems and multi-agent collaborative frameworks have demonstrated a high level of intelligence in task scheduling, path planning, and module orchestration. Introducing intelligent agent systems into image retouching scenarios enables structured decomposition and automated decision-making of retouching tasks, standardizing and proceduralizing processes that previously relied on manual experience, significantly reducing labor costs and improving the overall intelligence and controllability of the system.

[0007] From an industry perspective, many small and medium-sized businesses are looking to leverage AI to improve their image production capabilities and achieve a "low-cost, high-frequency, and controllable" image update process. For example, businesses like barbershops, restaurants, and nail salons often need to introduce new menus, styling displays, or environmental photos for different seasons, promotions, and holidays. If a system could automatically identify problem areas based on the current image content and recommend or perform retouching operations, users only need to confirm or fine-tune the image. This would significantly reduce labor investment, improve image quality, and enhance marketing effectiveness.

[0008] However, there is currently a lack of a complete system that can integrate multimodal content understanding, multi-agent task reasoning, modular image restoration, and iterative optimization with quality feedback. This makes it difficult to build a closed-loop workflow from "image analysis → image restoration strategy generation → execution path planning → automatic restoration → feedback optimization." This leads to many limitations in the accuracy, automation level, and practical usability of existing systems. Therefore, there is an urgent need to build a new type of automatic image restoration system that has a deep understanding of image content and the ability to extract semantic structure, automatically generate structured image restoration strategies based on recommended content, plan the order and execution path of image restoration tasks through a multi-agent collaboration mechanism, support local scoring feedback and dynamic task adjustment, and ultimately achieve high automation, low cost, and strong adaptability of image optimization.

[0009] The "Automatic Photo Retouching Optimization Suggestion Generation and Path Planning Assistance System Based on Content Analysis and Multi-agent Collaboration" proposed in this invention is precisely to solve the above-mentioned problems. The system intelligently scores and generates suggestions for images through a joint multimodal large model, and introduces multiple functional components such as semantic analysis agent, strategy selection agent and path planning agent to achieve a complete closed-loop process from image semantic understanding, strategy formulation, path reasoning to repair execution and effect feedback. The system is particularly suitable for industry scenarios with batch image processing needs, such as barber shops, restaurants, beauty, clothing, e-commerce displays, etc. It can be widely used in scenarios such as intelligent photo retouching platforms, content generation services, corporate promotional image management, etc., and has significant practical value and commercial promotion prospects. Summary of the Invention

[0010] The present invention aims to provide an automatic image retouching optimization suggestion generation and path planning execution system, which is to meet the growing demand for rapid optimization and high-quality dissemination of image content, especially for a wide range of scenarios in the daily publicity map updates of service industries such as barbershops, catering, and beauty makeup.

[0011] To achieve the above object, the technical solution of the present invention is a multi-agent automatic image retouching system based on content analysis, including the following steps:

[0012] Step 1: Receive the input of the original image, call the fine-tuned multi-modal large model to perform joint semantic-visual analysis on the image, extract key semantic elements, perform a comprehensive image quality score on the image, and output a set of repair and optimization strategies for image defects;

[0013] Step 2: Judge whether the image quality score in Step 1 is lower than the preset threshold T0. If the image quality score S < T0, trigger the intelligent image retouching process to enter the next step; if the score S ≥ T0, the image already meets the quality requirements, directly output the score and suggestion results, and no further image retouching is required;

[0014] Step 3: Start the multi-agent collaboration framework, and call the semantic analysis Agent, strategy selection Agent, and path planning Agent respectively to complete three phased tasks of semantic analysis, strategy selection, and path planning, and generate an executable structured image retouching strategy and process;

[0015] Step 4: According to the image retouching planning path generated in Step 3, call each image repair module in turn to execute the image retouching subtasks. After each image retouching subtask is completed, call the corresponding local score and repair suggestion generation model for this effect to score this effect, and compare it with the set threshold T a If the local score S a < T a , then according to the optimization suggestions of this model, the Agent performs the second repair to achieve local iterative optimization;

[0016] Step 5: After the current module's repair task, if the local score S obtained by the system judgment a ≥ T a , the system automatically enters the execution of the next module task;

[0017] Step 6: After all image retouching modules are executed, the system calculates the total score S according to the preset weighting rule based on the local score results of each module final , and compares it with the preset threshold T final . If the total score S final ≥ T final , output the final repaired image; if the total score S final < T final, it is determined that the next step needs to be entered;

[0018] Step 7: If the total score is S final <T final , the system calls the multimodal large model in step 1 again, generates the image quality score result and optimization strategy based on the current image retouching result, and enters the next round of judgment and optimization iteration process. The above steps can be executed repeatedly until the image quality meets the set requirements and satisfies the output conditions.

[0019] Preferably, the multimodal large model used in step 1 is fine-tuned based on the basic capabilities using a dataset with a core structure of "image-subjective evaluation-score" triples, specifically including the following steps:

[0020] S1 Data Preparation: Construct a training dataset for fine-tuning. The data format is a triple: {"image":<image encoding>,"text":<subjective evaluation text>,"score":<quality score (0-10)>}. The "evaluation text" is used to describe the current subjective quality issues of the image and the improvement direction. The score value comes from the user's subjective score or market research sampling results. The system selects one of them as the training input based on the image's application scenario and data sampling strategy, aiming to accurately reflect the image's quality acceptance and aesthetic preferences among the target user group. Formally, the training dataset can be represented as a data set consisting of an image, text, and score triple:

[0021] ;

[0022] in:

[0023] I i : the i-th image;

[0024] T i : subjective evaluation text corresponding to the image;

[0025] S i : Corresponding quality score;

[0026] Based on the triple data, two types of training samples are constructed:

[0027] Subtask 1: Text generation task: input image and output subjective evaluation text, the modeling method is image-text instruction generation;

[0028] Subtask 2: Rating regression task: input image, output subjective rating, modeled by multimodal regression prediction;

[0029] S2 multi-task modeling design: The model adopts a joint training framework, combining the two subtasks in S1 to achieve multi-objective learning. The text generation task is optimized by language modeling loss, and the score regression task is optimized by minimizing the mean squared error loss. The weighted sum of the two loss functions guides the overall gradient update.

[0030] Implementation of the S3 fine-tuning strategy: Parameters of the image encoder portion of the pre-trained multimodal model are frozen. In the language generation module, an efficient fine-tuning strategy combining low-bit quantization and low-rank parameter injection is adopted to achieve targeted optimization of the model. This enables the fine-tuned multimodal model to accurately identify image quality issues, generate repair suggestions in market-style language, and output subjective quality scores consistent with the target scoring criteria.

[0031] Preferably, in step 3:

[0032] The functional goal of the semantic analysis agent is to extract a semantic map based on the image content structure, identify image quality issues, and output the image's structured semantic map, object recognition information, and problem area attribution;

[0033] The functional goal of the strategy selection agent is to formulate a retouching strategy and task list based on the semantic analysis results and optimization suggestions provided by the multimodal model, output the retouching task list, and clearly map it to a specific repair module that can be called by the system, including the module name, target area, and parameters;

[0034] The functional goal of the path planning agent is to combine the task list with semantic information, plan the execution order and method of the photo editing task, generate the execution path, and output the module execution order, dependency logic, and execution strategy.

[0035] Preferably, the local scoring and repair suggestion generation model in step 4 is a lightweight multimodal model for the photo retouching subtask, which is distilled from the multimodal large model in step 1. Its capabilities focus on: single-dimensional quality evaluation + optimization suggestion generation. Its construction and training steps are:

[0036] S1 Data Preparation: Construct a batch of structured training data pairs for the image editing subtask. These data pairs are obtained by inputting images into the multimodal model in step 1 and having it output quality scores and optimization suggestions.

[0037] S2 distillation training: A local scoring model specifically built for the photo retouching subtask is designed as the student model. The model receives three inputs: the original image, the restored image, and the corresponding ROI, and outputs two types of results: a numerical local score and a textual optimization suggestion. During training, a multi-task loss function is used for joint optimization.

[0038] Preferably, in step 6, the total score is calculated according to a preset weighting rule. The dimensions of the weight design include: semantic importance weight, task impact weight, and local score credibility weight. The specific weighting calculation method is as follows: ;

[0039] Where: S i : local score of the i-th retouching module;

[0040] W i : The comprehensive weight of the module is calculated by three factors: semantics, task impact, and confidence; W i The specific calculation formula is as follows:

[0041] ;

[0042] The three sub-weights satisfy .

[0043] Preferably, the repair module includes exposure correction, distortion correction, removal of reflection and image ghosting, enhancement of rainy day effect, changing rainy day to sunny day, white balance adjustment, noise removal, filter effect, correction of warm and cold colors, hand shake correction, intelligent removal of unnecessary objects in the background, adjustment of close-up photos to distant views, background blur, background replacement, adjustment of subject position to center, and picture angle adjustment function modules.

[0044] A multi-agent automatic image retouching system based on content analysis, developed using the technical solutions of this invention, utilizes a multimodal large model as its core, integrating semantic understanding with image quality assessment mechanisms. This system intelligently extracts key semantic elements from images, such as lighting, composition, and subject clarity, and performs multi-dimensional image quality scoring and defect identification. Based on the scoring results, the system automatically generates targeted retouching strategies. Using a multi-agent collaborative framework, it completes semantic analysis, strategy formulation, and retouching path planning, achieving a structured and automated image optimization process. During the retouching execution phase, the system incorporates a local scoring and suggestion generation module, enabling module-level quality feedback and local iterative optimization, ensuring that each retouching step is controllable and evaluable. This technology is designed to meet the needs for rapid image content optimization and high-quality dissemination, and is particularly well-suited for frequent image updates in service industries such as barbershops, restaurants, and beauty salons. Users simply upload their images, and the system automatically identifies issues and performs optimizations, significantly reducing the need for professional photography and post-processing. This allows even non-professional users to easily achieve high-quality image results, thereby improving content update efficiency and helping small businesses and individuals achieve effective marketing and brand building. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flow chart of a multi-agent automatic photo editing system based on content analysis according to the present invention;

[0046] Figure 2 This is an image processing flow chart of a multi-agent collaborative framework in a multi-agent automatic photo editing system based on content analysis according to the present invention;

[0047] Figure 3 This is a flowchart of the implementation steps of fine-tuning a multimodal large model of a multi-agent automatic photo retouching system based on content analysis according to the present invention;

[0048] Figure 4 This is a flowchart of a local scoring and repair suggestion generation model for a multi-modal model that distills needles for specific tasks using a multi-agent automatic photo retouching system based on content analysis described in the present invention. DETAILED DESCRIPTION

[0049] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1-4 As shown in FIG, a multi-agent automatic image editing system based on content analysis specifically includes the following steps:

[0050] Step 1: This step receives the original image input, calls the fine-tuned multimodal large model to perform joint semantic-visual analysis on the image, and extracts key semantic elements in the image, including but not limited to background coordination, lighting conditions, character hairstyle clarity, weather quality, composition layout rationality and other content information. Based on this information, the system performs a comprehensive image quality score on the image, which combines multi-dimensional evaluation criteria such as image style, aesthetic level, clarity and subjective quality. At the same time, based on the multi-dimensional image quality scoring results and the understanding of the semantic content of the image, the system automatically generates a set of repair and optimization strategies for specific defects, such as insufficient lighting, blurred subjects, and unbalanced composition, to guide the subsequent image quality enhancement process;

[0051] Step 2: The system determines whether the image quality score in step 1 is lower than the preset threshold T0. If the image quality score S < T0, the intelligent image retouching process is triggered to enter the next step. If the score S ≥ T0, the image has met the quality requirements and the score and recommended results are directly output without further retouching.

[0052] Step 3: The system starts the multi-agent collaboration framework and calls three functional Agent components to complete the three phased tasks of semantic analysis, strategy selection and path planning. The core purpose of this step is to integrate the unstructured repair suggestions output by the multimodal large model with the image structure semantics to generate executable structured photo retouching strategies and processes. The functional goal of the semantic analysis agent is to extract semantic maps based on the image content structure, such as "portrait + backlight + complex background + complete face" and other information, and identify image quality problems such as underexposure, cluttered background, and lack of character prominence. The functional goal of the strategy selection agent is to formulate photo retouching strategies and task lists based on the semantic analysis results and the optimization suggestions provided by the multimodal model. The functional goal of the path planning agent is to combine the task list with semantic information to plan the execution order and method of the photo retouching tasks and generate an execution path, such as "light compensation → background optimization → skin quality enhancement";

[0053] Step 4: The system calls each image restoration module in turn to perform the restoration subtask according to the restoration planning path generated in step 3. After completing each restoration subtask, such as distortion correction, the system calls the corresponding local scoring and restoration suggestion generation model for effect A, such as distortion correction (an evaluation model focused on effect A obtained by distillation of the multimodal model) to score the effect and compare it with the preset threshold T a For comparison, if the local score S a < T a , then the optimization suggestions of the model are generated according to the local scores and repair suggestions for effect A, and the Agent performs the second repair to achieve local iterative optimization focusing on effect A.

[0054] Step 5: If the current module repairs the task, the system determines the local score S a ≥ T a , the system automatically proceeds to the next module task and continues to call the next restoration module, such as applying a filter. This process continues, and each step includes a local scoring and opinion generation mechanism, forming an adaptive image retouching chain based on quality feedback.

[0055] Step 6: After all the editing modules are completed, the system calculates the total score S based on the local scoring results of each module according to the preset weighting rules. final and the preset threshold T final Compare. If the total score S final ≥ T final , then the final repaired image is output; if the total score S final < T final , it is determined that the next step needs to be taken.

[0056] Step 7: If the total score is not S final < Tfinal , the system re-calls the multimodal large model from step 1 and generates an image quality score and optimization strategy based on the current retouching results, entering the next round of judgment and optimization iterations. The above steps can be repeated until the image quality meets the set requirements and satisfies the output conditions.

[0057] In step 1, to enable the system to accurately identify quality defects in images and generate evaluation text and ratings in a language that aligns with the target market's aesthetic language, the multimodal large model used is fine-tuned using a dataset with a core structure of "image-subjective evaluation-rating" triples, building on its basic capabilities. This fine-tuning process aims to give the model dual capabilities: first, generating repair suggestion text with a professional expression style, and second, outputting quality ratings that meet the user's subjective aesthetic standards. The specific fine-tuning steps are as follows:

[0058] S1: Data Preparation: Construct a training dataset for fine-tuning. The data format is a triple: {"image":<image encoding>, "text":<subjective evaluation text>, "score":<quality score (0-10)>}. The "evaluation text" is used to describe the current subjective quality issues of the image and improvement directions. The score value comes from user subjective ratings or market research sampling results. The system selects one of these as training input based on the image's application scenario and data sampling strategy, aiming to accurately reflect the image's quality acceptance and aesthetic preferences among the target user group. Formally, the training dataset can be represented as a data set consisting of an image, text, and score triple:

[0059] ;

[0060] in:

[0061] I i : the i-th image;

[0062] T i : subjective evaluation text corresponding to the image;

[0063] S i : The corresponding quality score.

[0064] Based on the triple data, two types of training samples are constructed:

[0065] Subtask 1: Text generation task: input image, output subjective evaluation text, modeling method is image-text instruction generation;

[0066] Subtask 2: Rating regression task: input image, output subjective rating, modeling method is multimodal regression prediction.

[0067] S2: Multi-task Modeling Design: This model utilizes a joint training framework, combining two subtasks to achieve multi-objective learning. The text generation task is optimized using a language modeling loss, while the score regression task is optimized by minimizing the mean squared error loss. A weighted sum of these two loss functions guides the overall gradient update, achieving a synergistic improvement in image content understanding and evaluation output capabilities.

[0068] S3: Fine-tuning Strategy Implementation: To achieve efficient fine-tuning and flexible module deployment while ensuring the model's perception capabilities and semantic generation quality, this system employs structural freezing and low-rank adaptation mechanisms during fine-tuning. Specifically, the image encoder component of the pre-trained multimodal model is parameter-frozen, meaning that the visual perception weights of this component are excluded from training updates. This preserves the general image feature extraction capabilities learned during large-scale visual pre-training and ensures the model's visual robustness across a wide range of image inputs.

[0069] At the same time, an efficient fine-tuning strategy combining low-bit quantization and low-rank parameter injection is employed in the language generation module. This strategy first quantizes the language model's main weights to 4 bits, using a near-optimal asymmetric quantization format to significantly reduce video memory and computing resource consumption. Subsequently, a trainable low-rank matrix is ​​introduced into the Transformer's multi-head attention mechanism to achieve the fine-tuning goal of updating only a small number of parameters.

[0070] This hybrid fine-tuning approach not only retains most of the model's original parameters while also enabling targeted optimization through a small number of learnable weights, but is particularly well-suited for fine-tuning tasks involving image-text-rating triples. This allows the system to build a multimodal model capable of subjective aesthetic understanding, rating regression, and natural language suggestion generation, meeting the requirements for rapid deployment and evaluation in large-scale, multi-process image restoration tasks.

[0071] The resulting fine-tuned multimodal model has the following capabilities:

[0072] (1) Accurately identify image quality issues and generate repair suggestions in market-style language;

[0073] (2) It can output subjective quality scores that are consistent with the target scoring standards, providing scoring references and execution triggers for subsequent image editing processes.

[0074] Regarding the set of repair and optimization suggestions, the structured suggestions are as follows:

[0075] Image subjective score: 0.58

[0076] Semantic label set: including multiple visual description labels such as "dark figure", "messy background", "portrait on the left", "tilted picture", "rainy day", and "cold colors";

[0077] Optimization suggestion set: includes several executable optimization instructions, each consisting of a "problem" and its "suggestion", for example:

[0078] Problem: "Underexposure", Suggestion: "Perform local exposure correction";

[0079] Problem: "The subject is not centered", suggestion: "Adjust the subject's position to the center";

[0080] …

[0081] The semantic analysis agent's functional goal in step 3 is to combine the set of recommendations output by the multimodal large model with the structural information of the original image to extract a structured semantic map. This identifies each major object in the image and its corresponding problem area, providing an operational semantic foundation for subsequent strategy and path reasoning. The semantic analysis agent's input is the original image, the quality analysis output by the multimodal large model, and the set of recommendations (from step 1). The output includes the image's structured semantic map, object recognition information, and problem area attribution. The following structured results are generated:

[0082] The overall visual description label of the image includes:

[0083] Portrait

[0084] Dark

[0085] rain

[0086] tilt

[0087] Cool colors

[0088] Complex background

[0089] Regions and attributes (an example of one is shown below):

[0090] Object Type: Portrait

[0091] Area coordinates and size: x=80, y=90, width=250, height=320

[0092] Attributes: dark, left-aligned, cool

[0093] The following quality issues were identified and optimization suggestions were given (an example of one of them is shown below):

[0094] Area: Portrait

[0095] Problem: Underexposure

[0096] Recommendation: Perform local exposure correction

[0097] The strategy selection agent's functional goal is to infer which modules need to be invoked and what tasks are pending from the semantic map and suggestion set, ultimately generating a structured list of retouching strategies. This agent integrates a small language model to accomplish this semantic understanding task. The agent's input includes the semantic map (from Agent 1), the set of repair suggestions (from Step 1), and the historical rating distribution / experience knowledge base. The agent's output is a list of retouching tasks, explicitly mapped to specific retouching modules that can be invoked by the system. This list includes the module name, target area, and parameters. An example of the generated structure is shown below:

[0098] Task list (an example of one is below):

[0099] Task module: exposure correction

[0100] Target area: Portrait

[0101] Problem: Underexposure

[0102] Processing parameters: partial face, medium intensity

[0103] The functional goal of the path planning agent is to infer the order of photo editing from the task list, forming an "execution path" that conforms to the dependency relationship and image processing logic, and improving the stability of the photo editing process and the final effect. The input includes the photo editing task list (from Agent 2), the image structured semantic information (from Agent 1), and the historical photo editing planning process knowledge base (experience base). The order of photo editing needs to be considered. For example, the order of background blur and aesthetic filter application should be placed at the end, because background blur should be performed after the character is complete to avoid background processing accidentally damaging the subject's outline, and too late will affect the unity of the filter. The aesthetic filter application module has the feature of "global tonality" and must be performed after all structural repairs to avoid interfering with intermediate processing. The final output includes the module execution order, dependency logic, and execution strategy. The generated structured example is as follows:

[0104] Execution paths (an example of one is shown below):

[0105] Step: 1

[0106] Task module: exposure correction

[0107] Target area: Portrait

[0108] Planning Description:

[0109] First note: Perform exposure correction first to ensure that the image brightness is standardized.

[0110] Second note: Color temperature adjustment must be performed after correct exposure to avoid color cast.

[0111] …

[0112] The local scoring and repair suggestion generation model in step 4 refers to a lightweight multimodal model for photo retouching subtasks (such as exposure correction, clarity enhancement, lighting enhancement, etc.), which specifically evaluates the repair quality of a certain subtask and scores it, and can generate further optimization suggestions. This model is distilled from the multimodal large model in step 1, and its capabilities focus on "single-dimensional quality evaluation + optimization suggestion generation", so it has lower inference costs, higher focus, and faster response speed. Although the original multimodal large model can complete the global analysis of the image, it has a high cost per call and is not suitable for frequent calls within the photo retouching module (such as scoring immediately after the subtask is repaired in step 4), and it cannot focus on only one quality dimension (such as only evaluating whether the skin color is natural). The global multi-factor scoring may interfere with a single judgment. Therefore, we "migrated" its "capabilities" to multiple local lightweight models that focus on subtasks. The specific steps of distillation training are:

[0113] S1: Data preparation: Construct a batch of structured training data pairs for photo editing subtasks, such as repairing exposure. This data pair is obtained by inputting images into the multimodal model in step 1 and letting it output quality scores and optimization suggestions.

[0114] S2: Distillation Training: Design a local scoring model specifically built for the image retouching subtask as the student model. This model receives three inputs: the original image, the restored image, and the corresponding ROI (region of interest). It outputs two types of results:

[0115] Numerical local scores are used to evaluate the quality of image restoration tasks and simulate the effectiveness of multimodal models in this task.

[0116] Text-based optimization suggestions generate local repair suggestions based on the repair effect, simulating the optimization suggestion generation function of the multimodal model.

[0117] During training, a multi-task loss function is jointly optimized, including the mean squared error loss term for regressing the score value and the cross-entropy language modeling loss term for generating recommended text.

[0118] Regarding the restoration module, it currently includes exposure correction, distortion correction, removal of reflections and image ghosting, enhancement of rainy day effects, turning rainy days into sunny days, white balance adjustment, noise removal, filter effects, correction of warm and cold colors, hand shake correction, intelligent removal of unnecessary objects in the background, adjustment of close-up photos to distant views, background blur, background replacement, adjustment of subject position to center, picture angle adjustment and other functions.

[0119] In step 6, the total score is calculated using pre-set weighting rules. The weighting should reflect the importance of different retouching dimensions in the overall perception of image quality. The weighting dimensions include semantic importance, task impact, and local score credibility.

[0120] The semantic importance weight indicates that the image regions corresponding to different restoration modules have different visual attention importance. For example, the semantic importance of the face region is higher than that of the background region, and the semantic importance of facial clarity is higher than that of clothing texture details. For this part, we have a specified region semantic label-weight comparison table for weighting. For example:

[0121] Face area: includes the face, eyes, mouth, etc., with a weight value of 0.9, indicating that this area is more important in image quality perception.

[0122] Facial clarity: refers to whether the face in the image is clear. The weight value is 0.8, indicating that facial clarity has a greater impact on the overall quality perception.

[0123] Hue: refers to the color effect of the image, with a weight value of 0.7, showing the impact of this visual feature on the overall effect.

[0124] Lighting and Shadows: This involves the lighting effects and shadows of the image, with a weight value of 0.7, indicating that this feature has a greater impact on quality perception.

[0125] Background area: includes scenery, walls, etc., with a weight value of 0.3, indicating that this area has little impact on image quality.

[0126] …

[0127] The local scoring confidence weight indicates instability under certain conditions (such as scoring bias in highlight areas). Therefore, we also use the confidence score in the scoring model output as part of the weighting factor to avoid areas with high uncertainty when weighting.

[0128] Regarding the specific weighted calculation method, our image total score as follows:

[0129] ;

[0130] in:

[0131] S i : The local score of the i-th retouching module (e.g. 0.8);

[0132] W i : The comprehensive weight of this module is calculated by three factors: semantics, task impact, and confidence. i The specific calculation formula is as follows:

[0133] ;

[0134] The three sub-weights satisfy .

[0135] The above technical solutions only reflect the preferred technical solutions of the technical solutions of the present invention. Any changes that may be made to certain parts thereof by those skilled in the art all reflect the principles of the present invention and fall within the scope of protection of the present invention.

Claims

1. A multi-agent automatic photo editing system based on content analysis, characterized by: It includes the following steps: Step 1: Receive the original image input, call the fine-tuned multimodal large model to perform joint semantic-visual analysis on the image, extract key semantic elements, perform comprehensive image quality scoring on the image, and output a set of repair and optimization strategies for image defects; Step 2: Judge whether the image quality score in Step 1 is lower than the preset threshold T0. If the image quality score S < T0, trigger the intelligent image repair process to enter the next step; If the score S ≥ T0, the image already meets the quality requirements, directly output the score and recommended results, and no further image repair is required; Step 3: Start the multi-agent collaboration framework, and call the semantic analysis Agent, strategy selection Agent, and path planning Agent respectively to complete the three phased tasks of semantic analysis, strategy selection, and path planning, and generate an executable structured image repair strategy and process; Step 4: Based on the retouching planning path generated in step 3, call each image restoration module in turn to perform the retouching subtask. After completing each retouching subtask, call the corresponding local scoring and restoration suggestion generation model for the effect to score the effect and compare it with the threshold T. a For comparison, if the local score S a <T a , then according to the optimization suggestions of the model, the Agent performs the second repair to achieve local iterative optimization; Step 5: If the current module repairs the task, the system determines the local score S a ≥T a , the system automatically enters the execution of the next module task; Step 6: After all the editing modules are completed, the system calculates the total score S based on the local scoring results of each module according to the preset weighting rules. final and the preset threshold T final For comparison, if the total score S final ≥T final , then the final repaired image is output; if the total score S final <T final , it is determined that the next step needs to be entered; Step 7: If the total score is S final <T final , the system calls the multimodal large model in step 1 again, generates the image quality score result and optimization strategy based on the current image retouching result, and enters the next round of judgment and optimization iteration process. The above steps can be executed repeatedly until the image quality meets the set requirements and satisfies the output conditions.

2. The multi-agent automatic photo editing system based on content analysis according to claim 1 is characterized in that: The multimodal large model used in Step 1 is fine-tuned on the basis of its basic capabilities using a dataset with the core structure of "image-subjective evaluation-score" triple. The specific steps are as follows: S1 Data preparation: Construct a training dataset for fine-tuning. The data format is a triple: {"image": <image encoding>, "text": <subjective evaluation text>, "score": <quality score (0-10)>}, where the "evaluation text" is used to describe the current subjective quality problems and improvement directions of the image. The score value comes from user subjective scoring or market research sampling results. The system selects one of them as the training input according to the application scenario of the image and the data sampling strategy, aiming to accurately reflect the quality acceptance and aesthetic preference of the image in the target user group. Formally, this training dataset can be represented as a data set composed of image, text, and score triples: ; Among them: I i : the i-th image; T i : subjective evaluation text corresponding to the image; S i : Corresponding quality score; Based on this triple data, construct two types of training samples: Sub-task 1: Text generation task: Input an image and output the subjective evaluation text. The modeling method is graphic instruction generation; Sub-task 2: Score regression task: Input an image and output the subjective score. The modeling method is multimodal regression prediction; S2 Multi-task modeling design: The model adopts a joint training framework, combines the two sub-tasks in S1 to achieve multi-objective learning. The text generation task is optimized through the language modeling loss, and the score regression task is optimized by minimizing the mean square error loss. The two loss functions are weighted and summed to guide the overall gradient update; S3 Fine-tuning strategy implementation: Perform parameter freezing on the image encoder part of the pre-trained multimodal model. In the language generation module, adopt an efficient fine-tuning strategy that combines low-bit quantization and low-rank parameter injection to achieve targeted optimization of the model, so that the fine-tuned multimodal model has the ability to accurately identify image quality problems, generate repair suggestions in market style language, and output subjective quality scores consistent with the target scoring standard.

3. The multi-agent automatic photo editing system based on content analysis according to claim 1 is characterized in that: In Step 3: The functional goal of the semantic analysis Agent is to extract the semantic map based on the image content structure, identify image quality problems, and output the structured semantic map of the image, object recognition information, and problem area attribution; The functional goal of the strategy selection agent is to formulate a retouching strategy and task list based on the semantic analysis results and optimization suggestions provided by the multimodal model, output the retouching task list, and clearly map it to a specific repair module that can be called by the system, including the module name, target area, and parameters; The functional goal of the path planning agent is to combine the task list with semantic information, plan the execution order and method of the photo editing task, generate the execution path, and output the module execution order, dependency logic, and execution strategy.

4. The multi-agent automatic photo editing system based on content analysis according to claim 1 is characterized in that: The local scoring and repair suggestion generation model in step 4 is a lightweight multimodal model for the photo retouching subtask. It is distilled from the large multimodal model in step 1 and focuses on single-dimensional quality evaluation and optimization suggestion generation. Its construction and training steps are as follows: S1 Data Preparation: Construct a batch of structured training data pairs for the image editing subtask. These data pairs are obtained by inputting images into the multimodal model in step 1 and having it output quality scores and optimization suggestions. S2 distillation training: A local scoring model specifically built for the photo retouching subtask is designed as the student model. The model receives three inputs: the original image, the restored image, and the corresponding ROI, and outputs two types of results: a numerical local score and a textual optimization suggestion. During training, a multi-task loss function is used for joint optimization.

5. The multi-agent automatic photo editing system based on content analysis according to claim 1 is characterized in that: In step 6, the total score is calculated according to the preset weighting rules. The dimensions of weight design include: semantic importance weight, task impact weight, and local score credibility weight. The specific weighting calculation method is as follows: ; Where: S i : local score of the i-th retouching module; W i : The comprehensive weight of the module is calculated by three factors: semantics, task impact, and confidence; W i The specific calculation formula is as follows: ; The three sub-weights satisfy .

6. The multi-agent automatic photo editing system based on content analysis according to claim 1 is characterized in that: The repair module includes exposure correction, distortion correction, removal of reflections and image ghosting, rainy day effect enhancement, rainy day to sunny day change, white balance adjustment, noise removal, filter effect, warm and cold color correction, hand shake correction, intelligent removal of unnecessary objects in the background, close-up photo adjustment to distant view, background blur, background replacement, adjustment of subject position to center, and picture angle adjustment function modules.

Citation Information

Patent Citations

  • Aesthetics quality evaluation model and method based on multi-modal learning

    CN115601772A

  • Personalized complex report generation method based on multi-agent system

    CN118569237A

  • Harmful image detection and recognition system and method based on agent collaboration

    CN120070906A

  • Low-light image enhancement method based on reinforcement learning and aesthetic evaluation

    WO2023236565A1

  • Multi-agent system-based information processing method and multi-agent system

    WO2025148684A1

Cited By

  • Image restoration method based on automatic evaluation and dynamic optimization and related equipment

    CN120876322A

  • Image inpainting method based on automated evaluation and dynamic optimization and related devices

    CN120876322B

  • Newspaper picture color optimization intelligent model construction method

    CN121437762A

  • Image data integration method based on multi-point image acquisition

    CN121600361A

  • Intelligent inspection method and system based on AI image recognition

    CN121686252A