A high-fidelity e-commerce photo generation method

By using a multimodal planning agent and a closed-source image generation model, combined with database feature matching, we have achieved automatic assessment, detection, and repair of clothing differences. This solves the problems of high fidelity and automation in e-commerce image generation, reducing costs and improving efficiency.

CN122265460APending Publication Date: 2026-06-23ANOTHER ME (BEIJING) VIRTUAL TECH DEV CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANOTHER ME (BEIJING) VIRTUAL TECH DEV CO LTD
Filing Date
2026-03-25
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-fidelity, controllable e-commerce commercial images, resulting in issues such as distorted details, uncontrollable generation, disconnect from marketing, and high costs.

Method used

By employing a multimodal planning agent and a closed-source image generation model, combined with database feature matching and image retrieval, the system automatically assesses, detects, and repairs clothing differences. It also generates selling point copy through a planning text model, achieving a fully automated process across the entire supply chain.

Benefits of technology

It enables the generation of high-fidelity e-commerce business photos, reduces human intervention, ensures brand consistency and controllability in the generated style, and improves generation efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265460A_ABST
    Figure CN122265460A_ABST
Patent Text Reader

Abstract

The application discloses a kind of high fidelity e-commerce commercial photo generation methods.The present application relates to the technical field of artificial intelligence.It includes: pre-processing and verifying clothing images;Through multi-modal intelligent agent, data is extracted and matching images are retrieved;The original image, attribute, search image and pose description are input into the closed model to generate commercial intermediate image;Using multi-modal evaluation model to monitor score, automatically positioning repair difference area for substandard image, combining manual check to output repair image;Finally, through the text model, the selling point copy is generated, and the layout intelligent agent instance is segmented in the non-occlusion area for automatic layout, and the final commercial photo is output.The present application can realize the automatic evaluation, detection matching and repair of clothing differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a method for generating high-fidelity e-commerce commercial images. Background Technology

[0002] The shortcomings of the existing technology are as follows: The problem of detail distortion: Traditional AI-generated images (AIGC) cannot guarantee that the clothing style and texture are 100% consistent with the original image (i.e., "AI illusion").

[0003] Uncontrollable generation: Randomly generated models often exhibit distortions in their limbs and facial features, and the background blending is unnatural. The poses are also monotonous.

[0004] Marketing disconnect: The generated images and the subsequent copywriting and layout of selling points are independent of each other, lacking an integrated automated process.

[0005] The necessity of solving this problem: Cost reduction and efficiency improvement: Traditional photography involves rental, models, and retouching, which are costly and time-consuming.

[0006] Business compliance and accuracy: E-commerce images must accurately reflect the product; any discrepancies in details may lead to an increased return rate.

[0007] Existing Solution 1: The advantage of traditional real-life commercial photography is its highest degree of realism. The disadvantages are its extreme cost, slow response time, and inability to produce on a large scale quickly.

[0008] Existing Solution 2: Pure text-driven AI-generated images (such as native Midjourney / SD)

[0009] Its advantages are high speed and beautiful graphics.

[0010] The downside is that the clothing details are poorly reproduced, and it is a "random creation" rather than a "product display", so it cannot be directly used in a rigorous e-commerce environment.

[0011] Therefore, designing a high-fidelity e-commerce commercial image generation method that can automatically assess, detect, and match clothing differences, and automatically repair differences is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0012] In view of this, the present invention provides a method for generating high-fidelity e-commerce commercial images, which can automatically evaluate clothing differences, automatically detect and match differences, and automatically repair differences.

[0013] To achieve the above objectives, the present invention adopts the following technical solution: A method for generating high-fidelity e-commerce commercial images includes: Acquire the original image of the clothing, perform preprocessing, and verify the preprocessed image; The verified clothing image and preset prompts are input into the multimodal planning agent to extract multi-dimensional data. At the same time, the database is called to perform feature matching and image retrieval to obtain the retrieved image. Input the original clothing image, attribute description, retrieved image, pose description text and preset prompt words into the closed-source image generation model, and generate intermediate images of commercial auction results through pose fission. A multimodal evaluation model is used to monitor and score the intermediate images of commercial auction results in multiple dimensions. For substandard images, the feasibility of repair is first determined. For repairable images, the difference area is located and matched through an automatic verification algorithm, and then pixel-level automatic repair is performed. Combined with manual secondary verification, the repaired images of commercial auction results are output. The text model is used to generate clothing selling point copy. The planning and layout intelligent agent performs instance segmentation on the commercial photography result repair image, completes the automatic layout of the selling point copy in the non-obscured area, and outputs the final commercial photography image.

[0014] Preferably, it also includes: performing facial sensitivity information and infringement detection on the generated final commercial image, and suspending the infringement warning image for storage and re-verification.

[0015] Preferably, the database includes a model face database, a scene database, and an outfit database. When performing feature matching and image retrieval, the database can be updated in real time. The update process includes: crawling text and image data of popular scenes or outfits from social media operation platforms via the Internet, and supplementing them into the corresponding database after structural parsing.

[0016] Preferably, the multimodal planning agent is based on the langGraph multi-agent collaboration framework of langChain, which combines agents for different tasks into a total agent; Specifically, this includes: generating attitude fission descriptions for adapted scenarios based on closed-source Gemini3-VL multimodal models; Based on the finely tuned open-source qwen3-vl multimodal model, it outputs the original attribute description of clothing, the feature description of clothing adaptation, and the basic selling points of clothing.

[0017] Preferably, the closed-source image generation model is the Google Nano Banana Pro model. The process of generating intermediate images of commercial photography results through pose fission includes: analyzing all categories of clothing in the apparel industry, outputting a 1+n prompt word scheme, where 1 is the basic prompt word framework and n represents the pose subdivision requirements of each clothing category, and completing the image generation work by adjusting the prompt word scheme to obtain the image generation result.

[0018] Preferably, the process of determining the feasibility of repair includes: The difference target region is automatically located and segmented by sam3. The mask of the correct target region is extracted from the original clothing image and denoted as target mask1. The corresponding difference region in the commercial photo is locally erased and denoted as target_loca1. The target mask1 is scaled to match the target_loca1 region. The scaling is completed by determining the threshold through IOU. The center of the two regions is calculated and hard-fitted to obtain the intermediate image mid_pic1. Repeat the above extraction, erasing, scaling, and pasting steps for all remaining paired regions to obtain the image final_pic after full-region pasting; Input the final_pic and the preset repair prompts into the Google Nano Banana Pro model to complete all local pixel-level automatic repairs and obtain the repaired image of the commercial auction result.

[0019] Preferably, the planning text model is a copywriting generation module based on deepseek and RAG, which generates copywriting for clothing selling points that is adapted to different e-commerce platforms by combining RAG retrieval and the deepseek model based on the basic selling points of clothing. The automated typesetting steps of the planning and typesetting intelligent agent include: The sam3 was used to segment the main figures and clothing in the restored image of the commercial photography result, and the unobstructed areas that are neither figures nor clothing were obtained. Based on the Google Nano Banana Pro model, the selling points are generated into white background images, and then the RGBA images of the white background images are obtained through sam3 and denoted as str_png series images; The system automatically filters target areas where text can be placed, scales the str_png series images, calculates the mIOU between the text mask and the person / clothing mask, and stops scaling once a predetermined standard is reached. Place the str_png series images that meet the predetermined standards into the corresponding target areas to complete the automated layout of all selling point texts.

[0020] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a method for generating high-fidelity e-commerce commercial images, with the following beneficial effects: The unique advantages of this solution: High-fidelity verification (core advantage): A dedicated "automatic verification algorithm" is introduced to identify the differences between the generated image and the original clothing, solving the pain point of "seemingly similar but actually different" AI-generated images.

[0021] End-to-end automation: From planning, image generation, evaluation, image retouching to copywriting and layout, it is a near-complete closed loop, reducing human intervention.

[0022] Resource library integration: By searching the model library and scene library, brand consistency and controllability of the generated style are ensured.

[0023] Lowering the technical barrier: High-quality marketing materials can be obtained quickly and efficiently through intelligent agents, reducing the barrier to entry. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0025] Figure 1 The method flowchart provided by the present invention.

[0026] Figure 2 Figure a shows the effect provided by the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] This invention discloses a method for generating high-fidelity e-commerce commercial images, including: Acquire the original image of the clothing, perform preprocessing, and verify the preprocessed image; The verified clothing image and preset prompts are input into the multimodal planning agent to extract multi-dimensional data. At the same time, the database is called to perform feature matching and image retrieval to obtain the retrieved image. Input the original clothing image, attribute description, retrieved image, pose description text and preset prompt words into the closed-source image generation model, and generate intermediate images of commercial auction results through pose fission. A multimodal evaluation model is used to monitor and score the intermediate images of commercial auction results in multiple dimensions. For substandard images, the feasibility of repair is first determined. For repairable images, the difference area is located and matched through an automatic verification algorithm, and then pixel-level automatic repair is performed. Combined with manual secondary verification, the repaired images of commercial auction results are output. The text model is used to generate clothing selling point copy. The planning and layout intelligent agent performs instance segmentation on the commercial photography result repair image, completes the automatic layout of the selling point copy in the non-obscured area, and outputs the final commercial photography image.

[0029] Specifically, this also includes: performing facial sensitivity information and infringement detection on the generated final commercial images, and suspending and re-verifying infringement warning images after they have been stored in the database.

[0030] Preferably, the database includes a model face database, a scene database, and an outfit database. When performing feature matching and image retrieval, the database can be updated in real time. The update process includes: crawling text and image data of popular scenes or outfits from social media operation platforms via the Internet, and supplementing them into the corresponding database after structural parsing.

[0031] Specifically, the multimodal planning agent is based on the langGraph multi-agent collaboration framework of langChain, which combines agents for different tasks into a total agent; Specifically, this includes: generating attitude fission descriptions for adapted scenarios based on closed-source Gemini3-VL multimodal models; Based on the finely tuned open-source qwen3-vl multimodal model, it outputs the original attribute description of clothing, the feature description of clothing adaptation, and the basic selling points of clothing.

[0032] Specifically, the closed-source image generation model is the Google Nano Banana Pro model. The process of generating intermediate images of commercial photography results through pose fission includes: analyzing all categories of clothing in the apparel industry, outputting a 1+n prompt word scheme, where 1 is the basic prompt word framework and n represents the pose subdivision requirements of each clothing category, and completing the image generation work by adjusting the prompt word scheme to obtain the image result.

[0033] Specifically, the process of determining the feasibility of the repair includes: The difference target region is automatically located and segmented by sam3. The mask of the correct target region is extracted from the original clothing image and denoted as target mask1. The corresponding difference region in the commercial photo is locally erased and denoted as target_loca1. The target mask1 is scaled to match the target_loca1 region. The scaling is completed by determining the threshold through IOU. The center of the two regions is calculated and hard-fitted to obtain the intermediate image mid_pic1. Repeat the above extraction, erasing, scaling, and pasting steps for all remaining paired regions to obtain the image final_pic after full-region pasting; Input the final_pic and the preset repair prompts into the Google Nano Banana Pro model to complete all local pixel-level automatic repairs and obtain the repaired image of the commercial auction result.

[0034] Specifically, the planning text model is a copywriting generation module based on deepseek and RAG. According to the basic selling points copywriting of clothing, combined with RAG retrieval and deepseek model, it generates clothing selling points copywriting adapted to different e-commerce platforms; The automated typesetting steps of the planning typesetting agent include: Use sam3 to perform instance segmentation on the repaired image of the commercial shooting result for the human body and clothing body, and obtain the unobstructed areas that are neither human nor clothing; Based on the google nano banana pro model, generate white-background images from the selling points copywriting, and then use sam3 to obtain the RGBA images of the white-background images, denoted as the str_png series of images; Use an automated program to screen the target areas where the copywriting can be placed, scale the str_png series of images, calculate the mIOU of the copywriting mask and the human / clothing body mask, and stop scaling after reaching the predetermined standard; Place the str_png series of images that meet the predetermined standard into the corresponding target areas to complete the automated typesetting of all selling points copywriting.

[0035] In a specific embodiment of the present invention, as Figure 1 shown, it includes: Input and planning stage: Obtain single (requiring a front view) or multiple (front view, 45° to the left of the front, 45° to the right of the front, 45° to the left of the back, 45° to the right of the back, etc.) original clothing images taken by the user, perform preprocessing verification on the clothing image size. If the clothing image size is too small (resolution less than 1k) and does not meet the subsequent image generation function, feedback an alarm and prompt to use a higher-resolution image. The clothing input is based on the "big" shape standard.

[0036] Input the single or multiple clothing images and the preset prompt words into the multi-modal planning agent to extract multi-dimensional data.

[0037] First, use the fine-tuned multi-modal model to obtain the attribute description of the original clothing and output it in json format.

[0038] According to the clothing attribute json and the clothing image data, assign them to different open-source fine-tuning and closed-source multi-modal models, and output the following adapted to the current clothing: face model feature description, scene feature description, dressing feature description, selling point feature description.

[0039] Based on the above face model, scene, and dressing feature descriptions, perform retrieval and matching in their respective databases, and output single or multiple images with relatively high similarity. Enter the next image generation link.

[0040] Related database construction. The image data in the database was pre-stored through generation and crawling in the early stages.

[0041] Model Face Database: Faces are sensitive data. This solution obtains actual model face data from different clothing categories, generates new faces, and then uses an internet face sensitivity detection solution to detect sensitive information and prevent infringement.

[0042] Scene Database: By acquiring commercial photos of different clothing categories, removing people and creating partial raw images, images that meet the specified clothing category scene are acquired in batches using the image-to-image method.

[0043] Outfit Database: By acquiring commercial photos of different clothing categories, we extract all outfit images and text descriptions of different styles and colors and store them in the database.

[0044] Database network search solution: This process intelligently crawls data from social operation platforms via the Internet, keeping track of current popular scenes, outfits, and other text and image data in real time. This data is used for subsequent image generation and for structural analysis of the collected images, which are then stored in the corresponding database.

[0045] Raw image stage: Based on the existing image generation model with the best closed-source performance, the original clothing image and attribute description, model face image, scene image, pose description text, outfit image and preset prompt words are simultaneously input into the image generation model to obtain a single output result.

[0046] Pose Fission: A single scene can meet the requirements of multi-pose commercial photography. A multimodal model is used to analyze the default poses of a single scene, and multi-pose text descriptions are generated by combining preset prompts with different clothing categories, and then applied to the raw image model.

[0047] The algorithm model used in this stage is the closed-source model that currently has the best effect on maintaining clothing consistency and the scene atmosphere effect is more in line with commercial photography. Therefore, this stage does not involve model fine-tuning.

[0048] Assessment and Repair Phase

[0049] Based on the intermediate results output from the raw image stage, quality monitoring is conducted by combining the original clothing image, preset prompts, and a fine-tuned multimodal evaluation model.

[0050] The main evaluation criteria include: clothing fidelity score, whether the posture has changed and the posture score, and the integration score between the person, the object (clothing), and the scene.

[0051] First, the scores across all dimensions are used to determine if all standards are met. If all standards are met, the garment is considered to require no repair and proceeds directly to the automated marketing layout stage. If any score fails to meet the standard, the process moves to the next stage to determine if repair is possible.

[0052] If the following are encountered: the pose is still in a starfish shape, the pose is strange, the texture is too strong, the physical structure of the clothing is incorrect (adding hats, fur collars, etc. out of thin air), the face swap fails, the facial expression is strange, or there are physical errors (the accessories and clothing do not match logically), it is determined that the initial screening cannot be repaired and the image is discarded directly.

[0053] Other minor issues, such as logo errors or zipper errors, will be addressed by inputting the text and images into the next stage for verification and repair.

[0054] Automatic verification function: Based on the current raw image results, the original clothing image, and the text description output during the evaluation phase, the target regions of the generated image and the original image are located by combining the self-supervised algorithm and the fine-tuned target detection algorithm.

[0055] Based on the extracted regional features from different images, dimensional distance is used to determine the closest regions in different images, and then they are paired one by one.

[0056] Automatic repair and manual secondary verification and repair: Based on the target region matching results output by the verification process, automated local repair is performed using preset prompts and closed-source image editing schemes.

[0057] The repair results require a second manual assessment to determine whether the repair is satisfactory and whether there are any unrepaired details. If not, the repair is directly output to the automated copywriting and layout workflow; if so, manual repair is performed on the affected areas.

[0058] Automated Marketing Stage

[0059] Based on the selling point information extracted from the original clothing (including clothing category, etc.), the selling point copy is generated through the planning text model, and the pre-set prompts are sent to the automatic planning intelligent agent.

[0060] Planning and layout intelligent agent

[0061] By analyzing commercial photography scenes and the selling points of clothing brands, a text style that matches the clothing brand is generated and applied to the planning text content.

[0062] The system segments the people in the commercial photos to obtain the non-people areas. The intelligent agent controls the size of the planning text to avoid overlap between the selling points text and the clothing.

[0063] The scaled-up selling points are automatically placed into the target area by analyzing the size of the commercial photos.

[0064] Full-process face infringement detection

[0065] Given the randomness of image generation and editing models, facial features may change to varying degrees during the raw image and automated editing processes. To avoid this change from causing facial infringement issues, facial infringement detection is performed after each of the four stages.

[0066] If an infringement warning is detected, the raw image result will be temporarily stored in the database pending manual review, while the process will be restarted.

[0067] In a specific embodiment of the present invention, the core intelligent agent framework is as follows: The intelligent agent in this solution is a multi-agent collaborative framework based on langChain - langGraph, which combines intelligent agents for different tasks into a total intelligent agent.

[0068] Planning Agent: Model: The closed-source Gemini3-VL multimodal model describes the pose fission based on the generated scene. Gemini3 has achieved commercial photography standards for general image content extraction, but the closed-source model has low adaptability for the target tasks of domestic e-commerce platforms, requiring the use of relevant domestic models and other aspects of work.

[0069] This section fine-tunes the open-source qwen3-vl multimodal model to output original clothing attribute descriptions, model face descriptions adapted to the clothing, styling descriptions, scene descriptions, clothing selling point descriptions, and clothing selling point copy. While the original qwen3-vl model can meet general text content extraction needs in China to a certain extent, its granularity for extracting specific clothing category-related content is insufficient, requiring further fine-tuning. Furthermore, for the same category of clothing, standards and seasonal indicators vary across different e-commerce platforms in China. These factors necessitate further adjustments to qwen3-vl.

[0070] Combining multiple multimodal models is one of the key requirements of this invention.

[0071] Input: Preset prompts, clothing images.

[0072] Output: Output according to the standard JSON data structure.

[0073] Raw image workflow: Model: This is a raw image model based on the closed-source Google Nano Banana Pro. Currently, this raw image model performs best compared to all open-source and closed-source algorithms in the commercial image field.

[0074] This solution is based on a comprehensive comparison and completes the image generation work at this stage by adjusting the prompt words.

[0075] Single-scene multi-pose image generation solution

[0076] Since commercial photos are only retained for a short time in e-commerce systems, a large number of commercial photos are needed as backups. However, there are certain limitations in the selection and generation of scenes that are suitable for specific clothing. Therefore, this invention designs a scheme that can perform pose fission based on different clothing and scenes.

[0077] There are significant differences between categories like baby sleeping bags and down jackets, and their requirements for commercial photography are completely different. Therefore, designing only one set of universally applicable pose-based prompts cannot meet the requirements of commercial photography for each category. This invention analyzes almost all categories of clothing in the apparel industry and outputs a "1+n" prompt scheme. "1" represents the basic prompt framework, including but not limited to positive prompts such as maintaining clothing consistency and blending clothing lighting and background, and negative prompts such as textured images and flat figures. "n" represents the specific pose requirements for each clothing category. For example, baby sleeping bags should be presented in a pose that conveys a sense of peaceful sleep, down jackets should be presented in an urban setting with a sophisticated and capable pose, and down jackets should be presented in a sporty pose in a snowy mountain setting.

[0078] Among them, the 1+n attitude fission scheme is one of the key requirements of this invention.

[0079] Input: Different types of images, and prompt text.

[0080] Output: Intermediate image of the commercial auction results.

[0081] Assessment and Repair Workflow

[0082] Model: Fine-tuned multimodal evaluation model: This model takes the generated result image and the original clothing image as input, and uses a multimodal model fine-tuned based on qwen3-vl to comprehensively evaluate the two types of images, outputting a comprehensive scoring index and detailed problem descriptions.

[0083] Automated Area Detection and Matching Scheme: SAM3 automatically locates and segments the target area based on the problem description. For small targets such as buttons, the target detection algorithm YOLO v10 + ResNet feature extraction and matching algorithm, combined with Euclidean distance, is used for target area matching. The correct physical target area on the paired original image is extracted by segmentation masking, denoted as target mask1. Simultaneously, the target swapping area on the commercial image is also partially erased using a mask, target_loca1. The same process is used to extract and erase the remaining paired areas.

[0084] The Google Nano Banana Pro-based local automatic repair solution utilizes the industry's best image editing algorithm. Mask1 is automatically scaled to approximately the same size as target_loca1. If the IOU exceeds 90%, the scaling is considered complete, and the centers of the two regions are calculated for hard merging. This image is denoted as mid_pic1. Based on mid_pic1, target_loca2 and the original mask2 are obtained. This process is repeated until all regions are merged, denoted as final_pic. final_pic and preset repair prompts are simultaneously input into the Google Nano Banana Pro algorithm model to complete all local repairs.

[0085] The automatic assessment, automatic detection and matching, and automatic repair of clothing differences at this stage are among the key requirements of this invention.

[0086] Input: intermediate image of the commercial auction result, original clothing photo, and prompt text.

[0087] Output: Repaired image of the commercial auction result.

[0088] Automatic layout intelligent agent

[0089] Model: Planning Text Model: A planning copy generation module based on Deepseek and RAG. It generates commercial images of clothing selling points that are suitable for different platforms by using RAG and Deepseek based on the clothing selling point copy.

[0090] Automated Marketing Agent: Combining SAM3 and Google Nano Banana Pro solutions. First, SAM3 segments the main subject and clothing theme, extracting non-person and non-clothing areas to avoid obscuring selling point text. Next, Nano Pro generates white-background images of the selling points, then SAM3 extracts RGBA images of the main subject area, denoted as str_png1, str_png2, etc. (multiple sets of selling point text). Finally, an automated program selects multiple areas suitable for text placement. Based on these areas, it scales different str_png images, calculates the mIOU between the text mask and the main subject / clothing mask, and stops scaling once a predetermined standard is met. The str_png is then placed in the target area of ​​the commercial image, and this process is repeated to complete the automatic text layout.

[0091] Intelligent sales copy generation and automated layout are among the key requirements of this invention.

[0092] Input: Intermediate image of the commercial auction result, marketing copy highlighting key selling points, and text prompts.

[0093] Output: Final commercial photo. (e.g.) Figure 2 (As shown).

[0094] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0095] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating high-fidelity e-commerce commercial images, characterized in that, include: Acquire the original image of the clothing, perform preprocessing, and verify the preprocessed image; The verified clothing image and preset prompts are input into the multimodal planning agent to extract multi-dimensional data. At the same time, the database is called to perform feature matching and image retrieval to obtain the retrieved image. Input the original clothing image, attribute description, retrieved image, pose description text and preset prompt words into the closed-source image generation model, and generate intermediate images of commercial auction results through pose fission. A multimodal evaluation model is used to monitor and score the intermediate images of commercial auction results in multiple dimensions. For substandard images, the feasibility of repair is first determined. For repairable images, the difference area is located and matched through an automatic verification algorithm, and then pixel-level automatic repair is performed. Combined with manual secondary verification, the repaired images of commercial auction results are output. The text model is used to generate clothing selling point copy. The planning and layout intelligent agent performs instance segmentation on the commercial photography result repair image, completes the automatic layout of the selling point copy in the non-obscured area, and outputs the final commercial photography image.

2. The method for generating high-fidelity e-commerce commercial images according to claim 1, characterized in that, Also includes: The generated final commercial images are subjected to facial sensitivity information and infringement detection. Infringement warning images are suspended and re-verified after being added to the database.

3. The method for generating high-fidelity e-commerce commercial images according to claim 1, characterized in that, The database includes a model face database, a scene database, and an outfit database. When performing feature matching and image retrieval, the database can be updated in real time. The update process includes: crawling text and image data of popular scenes or outfits from social media platforms, and then supplementing them into the corresponding databases after structural parsing.

4. The method for generating high-fidelity e-commerce commercial images according to claim 1, characterized in that, The multimodal planning agent is based on the langGraph multi-agent collaboration framework of langChain, which combines agents for different tasks into a total agent. Specifically, this includes: generating attitude fission descriptions for adapted scenarios based on closed-source Gemini3-VL multimodal models; Based on the finely tuned open-source qwen3-vl multimodal model, it outputs the original attribute description of clothing, the feature description of clothing adaptation, and the basic selling points of clothing.

5. The method for generating high-fidelity e-commerce commercial images according to claim 1, characterized in that, The closed-source image generation model is the Google Nano Banana Pro model. The process of generating intermediate images of commercial photography results through pose fission includes: analyzing all categories of clothing in the apparel industry, outputting a 1+n prompt word scheme, where 1 is the basic prompt word framework and n is the pose subdivision requirements of each clothing category, and completing the image generation work by adjusting the prompt word scheme to obtain the image result.

6. The method for generating high-fidelity e-commerce commercial images according to claim 5, characterized in that, The process for determining the feasibility of the repair includes: The difference target region is automatically located and segmented by sam3. The mask of the correct target region is extracted from the original clothing image and denoted as target mask1. The corresponding difference region in the commercial photo is locally erased and denoted as target_loca1. The target mask1 is scaled to match the target_loca1 region. The scaling is completed by determining the threshold through IOU. The center of the two regions is calculated and hard-fitted to obtain the intermediate image mid_pic1. Repeat the above extraction, erasing, scaling, and pasting steps for all remaining paired regions to obtain the image final_pic after full-region pasting; Input the final_pic and the preset repair prompts into the Google Nano Banana Pro model to complete all local pixel-level automatic repairs and obtain the repaired image of the commercial auction result.

7. The method for generating high-fidelity e-commerce commercial images according to claim 6, characterized in that, The planning text model is a copywriting generation module based on deepseek and RAG. It generates copywriting for clothing selling points that is adapted to different e-commerce platforms by combining RAG retrieval and the deepseek model with the basic selling points of clothing. The automated typesetting steps of the planning and typesetting intelligent agent include: The sam3 was used to segment the main figures and clothing in the restored image of the commercial photography result, and the unobstructed areas that are neither figures nor clothing were obtained. Based on the Google Nano Banana Pro model, the selling points are generated into white background images, and then the RGBA images of the white background images are obtained through sam3 and denoted as str_png series images; The system automatically filters target areas where text can be placed, scales the str_png series images, calculates the mIOU between the text mask and the person / clothing mask, and stops scaling once a predetermined standard is reached. Place the str_png series images that meet the predetermined standards into the corresponding target areas to complete the automated layout of all selling point texts.