A commodity image background replacement method, device, equipment and medium

By identifying product categories using Qwen-VL Max and Kontext models, and combining them with the SAM model for segmentation and proportion calculation, the accuracy and consistency issues of product image background replacement were resolved, achieving efficient and accurate background replacement results.

CN122155966APending Publication Date: 2026-06-05XIAMEN ZIXUN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN ZIXUN INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-01-31
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately replace the background of product images, exhibiting issues such as insufficient product recognition accuracy, low background generation matching degree, low product segmentation accuracy, and a lack of effect verification mechanisms.

Method used

The Qwen-VL Max model is used for product category recognition, and the Kontext model is used to generate background replacement images. The SAM model is used for segmentation and scale calculation, and the preset mapping library and magnification are used for adaptive adjustment to achieve consistency verification.

Benefits of technology

The accuracy of product image background replacement has been improved, with product recognition accuracy increased to 98.5%, background matching accuracy increased to 95%, and segmentation accuracy reaching 97.2%, ensuring the consistency of the generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122155966A_ABST
    Figure CN122155966A_ABST
Patent Text Reader

Abstract

The application provides a commodity image background replacement method, device, equipment and medium, the method comprises the following steps: S10, a commodity white background picture is acquired, the commodity white background picture is input into a recognition model to obtain a commodity category; S20, an initial commodity white background picture and a description scene prompt word are generated according to the commodity category and a preset prompt word template; S30, the initial commodity white background picture and the description scene prompt word are input into a background generation fusion model to generate a background replacement image; S40, an image segmentation model is used to segment the commodity white background picture and the background replacement image, proportion calculation and consistency verification are performed, and a final background replacement image is output; through accurate commodity recognition, intelligent background generation and consistency verification, the accuracy and effect of commodity image background replacement are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, device, and medium for replacing the background of a product image. Background Technology

[0002] In e-commerce product image processing, existing technologies struggle to efficiently and accurately replace the background of product images, primarily due to the following issues: 1. Insufficient product recognition accuracy: Traditional image recognition models have difficulty accurately identifying product types, resulting in the inability to perform targeted background replacement based on product type, which affects the final result. 2. Low background generation matching: The scene generated by the existing background generation model does not match the product, or the ratio of the product to the background is not coordinated, making the product look unnatural in the new background. 3. Low product segmentation accuracy: Existing segmentation models have low segmentation accuracy when processing products with different backgrounds, which affects the fusion effect between products and background. 4. Lack of effect verification mechanism: Existing technology lacks effective methods to verify the consistency of generated effects, and cannot ensure that the proportion of the product in the new background is the same as the original. Figure 1 To. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method, apparatus, device and medium for replacing the background of a product image, which improves the accuracy and effect of replacing the background of a product image through accurate product recognition, intelligent background generation and consistency verification. In a first aspect, the present invention provides a method for replacing the background of a product image, comprising the following steps: S10. Obtain a white background image of the product, input the white background image of the product into the recognition model, and obtain the product category; S20. Generate an initial white background image of the product and descriptive scene prompts based on the product type and preset prompt template; S30. Input the initial product white background image and the descriptive scene prompts into the background generation fusion model to generate a background replacement image; S40. Use an image segmentation model to segment the product white background image and the background replacement image, perform ratio calculation and consistency verification, and output the final background replacement image. Secondly, the present invention provides a product image background replacement device, comprising the following modules: The product category identification module acquires a white background image of the product, inputs the white background image of the product into the identification model, and obtains the product category. The initial generation module generates an initial white background image of the product and descriptive scene prompts based on the product type and preset prompt word template; The background replacement module inputs the initial product white background image and the descriptive scene prompts into the background generation and fusion model to generate a background replacement image; The verification output module uses an image segmentation model to segment the product white background image and the background replacement image, performs ratio calculation and consistency verification, and outputs the final background replacement image. Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect. Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect. One or more technical solutions provided by this invention have at least the following technical effects or advantages: 1. Accurate Product Recognition: Product category recognition is achieved through the Qwen-VL Max model, with an accuracy rate of 98.5%, providing accurate input for subsequent processing. 2. Intelligent background generation: Generates high-quality recommended scenes based on product type and preset prompt word templates, improving the background-product matching accuracy to 95%. 3. Adaptive scaling adjustment: Through a preset mapping library and magnification calculation, the scaling of the product in the new background is adaptively adjusted, reducing manual intervention. 4. High-precision segmentation: By using the SAM model combined with prior product information, the segmentation accuracy reaches 97.2%. 5. Consistency Verification: By calculating the difference in product proportions, the consistency of the generated effect is effectively verified, avoiding the problem of inconsistent proportions. The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description The present invention will be further described below with reference to the accompanying drawings and embodiments. Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the device in Embodiment 2 of the present invention. Detailed Implementation This application provides a method, apparatus, device, and medium for replacing the background of product images, thereby solving the problem of inaccurate background replacement of product images in the prior art. The overall concept of the technical solution in this application is as follows: A method for replacing the background of a product image based on Qwen-VL Max and Kontext models includes: obtaining a white background image of the product; inputting the white background image of the product into a Qwen-VL Max model for product type recognition; obtaining the corresponding magnification from a preset mapping library according to the product type; determining whether the product needs to be placed on a larger white background image; generating recommended scene prompts by Qwen-VL Max based on the product type and a preset prompt template; inputting the white background image of the product and the generated prompts into a Kontext LoRA model and a Kontext Dev model to generate a product image with a replaced background; segmenting the product using a SAM model; calculating the difference between the proportion of the product in the original white background image after generation and the proportion of the product in the original white background image before generation; if the difference exceeds a preset threshold, it is determined to be inconsistent. Example 1 like Figure 1 As shown, this embodiment provides a method for replacing the background of a product image, including the following steps: S10. Obtain a white background image of the product, input the white background image of the product into the recognition model, and obtain the product category; S11. Resize the product white background image to 1024×1024 pixels; S12. Input the white background image of the product into the Qwen-VL Max model to identify the product category; S20. Generate an initial white background image of the product and descriptive scene prompts based on the product type and preset prompt template; S21. Obtain the magnification factor from the preset mapping library according to the product type: the preset mapping library contains magnification factors corresponding to different product types; S22. Determine whether the product in the white background image needs to be replaced with a larger white background image based on the magnification factor. If it is necessary, replace it; otherwise, retain the original product white background image and output the initial product white background image. S23. Input the preset prompt word template and product category into the Qwen-VL Max model to generate prompt words describing the scenario; S30. Input the initial product white background image and the descriptive scene prompts into the background generation fusion model to generate a background replacement image; S31. Training the Kontext LoRA model: The Kontext LoRA model is trained using image data; the image data includes a product white background image, an initial product white background image, and a target image. The Kontext Dev model is used as the base model, and a pre-trained Kontext LoRA model is concatenated to construct the Kontext model. S32. Input the initial product white background image and the descriptive scene prompts into the Kontext model to generate a background image and output the background image. S33. Input the background image and the initial white background image of the product into the Kontext model for fusion, and output the background replacement image; S40. Use an image segmentation model to segment the product white background image and the background replacement image, perform ratio calculation and consistency verification, and output the final background replacement image; S41. The image segmentation model segments the white background image of the product to obtain the main image of the product before generation; the image segmentation model segments the background replacement image to obtain the main image of the product after generation. S42. Calculate the first proportion of the generated product main image in the product white background image, and calculate the second proportion of the product main image in the product white background image before generation. S43. Determine whether the difference between the first ratio and the second ratio exceeds a preset threshold; if the difference does not exceed the preset threshold, output the background replacement image as the final background replacement image; if the difference exceeds the preset threshold, adjust the parameters and return to step S30 to regenerate until the difference does not exceed the preset threshold, then output the final background replacement image. The adjustment parameters are specifically as follows: S431. If the difference exceeds a preset threshold, then calculate the third ratio, whereby the third ratio is the proportion of the initial product white background image to the product white background image. S432. Compare the third ratio with the second ratio. If the deviation between the third ratio and the second ratio is greater than the set value, adjust the magnification of the product type. If the deviation between the third ratio and the second ratio is less than or equal to a set value, then the scene description prompts are adjusted by deleting the scaling description in the scene description prompts and adding the ratio constraint in the scene description prompts. S433, Re-execute S30. Based on the same inventive concept, this application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2. Example 2 like Figure 2 As shown, this embodiment provides a product image background replacement device, including the following modules: The product category identification module acquires a white background image of the product, inputs the white background image of the product into the identification model, and obtains the product category. The initial generation module generates an initial white background image of the product and descriptive scene prompts based on the product type and preset prompt word template; The background replacement module inputs the initial product white background image and the descriptive scene prompts into the background generation and fusion model to generate a background replacement image; The verification output module uses an image segmentation model to segment the product white background image and the background replacement image, performs ratio calculation and consistency verification, and outputs the final background replacement image. The category identification module specifically includes: The size unit resizes the white background image of the product to 1024×1024 pixels. The identification unit inputs the white background image of the product into the Qwen-VL Max model to identify the product category; The initial generation module specifically includes: The magnification unit retrieves the magnification from a preset mapping library based on the product type: the preset mapping library contains magnifications corresponding to different product types; Replace the white background image unit. Based on the magnification factor, determine whether it is necessary to replace the product in the white background image with a larger white background image. If it is necessary, replace it. If it is not necessary, retain the original white background image and output the initial white background image. The prompt word unit inputs the preset prompt word template and product category into the Qwen-VL Max model to generate prompt words describing the scenario. The background replacement module specifically includes: The model training unit trains the Kontext LoRA model using image data; the image data includes a product white background image, an initial product white background image, and a target image. The Kontext Dev model is used as the base model, and a pre-trained Kontext LoRA model is concatenated to construct the Kontext model. The background generation unit inputs the initial product white background image and the descriptive scene prompts into the Kontext model to generate a background image and outputs the background image. The image fusion unit inputs the background image and the initial white background image of the product into the Kontext model for fusion and outputs a background replacement image. The verification output module specifically includes: The segmentation unit performs segmentation on the white background image of the product to obtain the main image of the product before generation; the image segmentation model also performs segmentation on the background replacement image to obtain the main image of the product after generation. The percentage calculation unit calculates the first percentage of the product's main image in the product's white background image after generation, and the second percentage of the product's main image in the product's white background image before generation. The difference judgment unit determines whether the difference between the first ratio and the second ratio exceeds a preset threshold; if the difference does not exceed the preset threshold, the background replacement image is output as the final background replacement image; if the difference exceeds the preset threshold, the parameters are adjusted and the process returns to step S30 to regenerate until the difference does not exceed the preset threshold, and then the final background replacement image is output. The adjustment parameters in the difference judgment unit are specifically as follows: The third proportion subunit calculates the third proportion if the difference exceeds a preset threshold. The third proportion is the proportion of the initial product white background image to the product white background image. Compare the sub-units, compare the third ratio with the second ratio, and if the deviation between the third ratio and the second ratio is greater than a set value, then adjust the magnification of the product type. If the deviation between the third ratio and the second ratio is less than or equal to a set value, then the scene description prompts are adjusted by deleting the scaling description in the scene description prompts and adding the ratio constraint in the scene description prompts. Return to the sub-unit and re-execute the background replacement module. Since the apparatus described in Embodiment 2 of the present invention is an apparatus used to implement the method of Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the apparatus based on the method described in Embodiment 1 of the present invention, and therefore will not be described again here. All apparatuses used in the method of Embodiment 1 of the present invention fall within the scope of protection of the present invention. Example 3 The overall approach of this application is as follows: Input a white background image of the product → Qwen-VL Max product recognition → Obtain magnification from the mapping library → Generate recommended prompts → Generate background using the Kontext model → Segment the product using SAM → Calculate and verify the proportions → Output the product image after background replacement. The specific steps include: (1) Product identification and preprocessing: Image input: Obtain a white background image of the product and resize it to 1024×1024 pixels. Product recognition: Input the white background image of the product into the Qwen-VL Max model to identify the product category (such as mobile phones, clothing, cosmetics, etc.). Preset mapping library: The preset mapping library contains standard magnification ratios for different product categories, such as 1.2x for mobile phones, 1.5x for clothing, and 1.1x for cosmetics. (2) Background generation preparation: Magnification application: Based on the product type, retrieve the magnification from the dictionary in the preset mapping library memory, and determine whether the product needs to be placed on a larger white background image. Prompt template design: The preset prompt templates include elements such as [product category], [scene type], [style description], [lighting conditions], and [composition requirements]. Prompt generation: Qwen-VL Max generates descriptive prompts based on product type and preset prompt templates, such as "mobile phone - modern home - minimalist style - natural light - central composition". (3) Background generation and product integration: Kontext Model Input: Input the product white background image and the generated prompts into the Kontext LoRA model and the Kontext Dev model. The Kontext LoRA model prepares rich image data pairs. Each data pair includes a product image, a corresponding descriptive word describing the scene (the descriptive word describing the target image with the changed background), and the target image with the changed background. These data pairs are used as training data for the Kontext LoRA model, allowing the model to learn how to generate a background image that meets the requirements based on the descriptive word and the product image. This model is pre-trained. That is: Kontext Dev → (as the base model) → Kontext LoRA. Background generation: The Kontext model generates high-quality background images that match the product based on the prompt words. Background blending: The product is blended with the generated background to form a preliminary background replacement image. (4) Commodity segmentation and proportion verification: Product segmentation: The SAM model is used to segment product regions, providing prior information based on known product types to improve segmentation accuracy. Prior information: (In the Qwen-VL Max model in the first step) Identify product type → Use type as index to retrieve the preset "product category - visual features"; map data to obtain prior information such as the shape / texture / proportion of the product. Proportion Calculation: Calculate the proportion of the generated product to the original white background image (A) and the proportion of the product to the original white background image before generation (B). Consistency verification: Calculate |AB|. If it exceeds a preset threshold (e.g., 5%), it is determined to be inconsistent and the background needs to be regenerated. (5) Output and Optimization Output: If the consistency verification passes, output the product image with the background replaced; if it fails, adjust the prompt text or magnification, and regenerate the background. The ratio difference exceeds the threshold → Calculate the third ratio (prepare the ratio of the product in the white background image to the original white background image) → Compare the third ratio with the second ratio: If the deviation is large, adjust the magnification (fine-tune according to the category standard). If the deviation is small, adjust the prompt (delete the scaling description + add a proportional constraint). → Re-execute S20 / S30 → Verify the ratio again. Based on the same inventive concept, this application provides an electronic device embodiment corresponding to Embodiment 1, as detailed in Embodiment 4. Example 4 This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement any of the implementation methods in Embodiment 1. Since the electronic device described in this embodiment is the device used to implement the method in Embodiment 1 of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in Embodiment 1 of this application. Therefore, how the electronic device implements the method in the embodiment of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiment of this application falls within the scope of protection of this application. Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 5. Example 5 This embodiment provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it can implement any of the implementation methods in Embodiment 1. The technical solutions provided in this application embodiment have at least the following technical effects or advantages: 1. Accurate Product Recognition: Product category recognition is achieved through the Qwen-VL Max model, with an accuracy rate of 98.5%, providing accurate input for subsequent processing. 2. Intelligent background generation: Generates high-quality recommended scenes based on product type and preset prompt word templates, improving the background-product matching accuracy to 95%. 3. Adaptive scaling adjustment: Through a preset mapping library and magnification calculation, the scaling of the product in the new background is adaptively adjusted, reducing manual intervention. 4. High-precision segmentation: By using the SAM model combined with prior product information, the segmentation accuracy reaches 97.2%. 5. Consistency Verification: By calculating the difference in product proportions, the consistency of the generated effect is effectively verified, avoiding the problem of inconsistent proportions. Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes. While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for replacing the background of a product image, characterized in that: Includes the following steps: S10. Obtain a white background image of the product, input the white background image of the product into the recognition model, and obtain the product category; S20. Generate an initial white background image of the product and descriptive scene prompts based on the product type and preset prompt template; S30. Input the initial product white background image and the descriptive scene prompts into the background generation fusion model to generate a background replacement image; S40. Use an image segmentation model to segment the product white background image and the background replacement image, perform ratio calculation and consistency verification, and output the final background replacement image.

2. The method for replacing the background of a product image according to claim 1, characterized in that: Specifically, S10 is: S11. Resize the product white background image to 1024×1024 pixels; S12. Input the white background image of the product into the Qwen-VL Max model to identify the product category; Specifically, S20 is: S21. Obtain the magnification factor from the preset mapping library according to the product type: the preset mapping library contains magnification factors corresponding to different product types; S22. Determine whether the product in the white background image needs to be replaced with a larger white background image based on the magnification factor. If it is necessary, replace it; otherwise, retain the original product white background image and output the initial product white background image. S23. Input the preset prompt word template and product category into the Qwen-VL Max model to generate prompt words describing the scene.

3. The method for replacing the background of a product image according to claim 1, characterized in that: Specifically, S30 is: S31. Training the Kontext LoRA model: The Kontext LoRA model is trained using image data; the image data includes a product white background image, an initial product white background image, and a target image. The Kontext Dev model is used as the base model, and a pre-trained Kontext LoRA model is concatenated to construct the Kontext model. S32. Input the initial product white background image and the descriptive scene prompts into the Kontext model to generate a background image and output the background image. S33. Input the background image and the initial white background image of the product into the Kontext model for fusion, and output the background replacement image.

4. The method for replacing the background of a product image according to claim 1, characterized in that: Specifically, S40 is: S41. The image segmentation model segments the white background image of the product to obtain the main image of the product before generation; the image segmentation model segments the background replacement image to obtain the main image of the product after generation. S42. Calculate the first proportion of the generated product main image in the product white background image, and calculate the second proportion of the product main image in the product white background image before generation. S43. Determine whether the difference between the first ratio and the second ratio exceeds a preset threshold. If the difference does not exceed the preset threshold, the background replacement image is output as the final background replacement image; If the difference exceeds the preset threshold, the parameters are adjusted and the process returns to step S30 to regenerate until the difference does not exceed the preset threshold, at which point the final background replacement image is output. The adjustment parameters are specifically as follows: S431. If the difference exceeds a preset threshold, then calculate the third ratio, whereby the third ratio is the proportion of the initial product white background image to the product white background image. S432. Compare the third ratio with the second ratio. If the deviation between the third ratio and the second ratio is greater than the set value, adjust the magnification of the product type. If the deviation between the third ratio and the second ratio is less than or equal to a set value, then the scene description prompts are adjusted by deleting the scaling description in the scene description prompts and adding the ratio constraint in the scene description prompts. S433, Re-execute S30.

5. A product image background replacement device, characterized in that: Includes the following modules: The product category identification module acquires a white background image of the product, inputs the white background image of the product into the identification model, and obtains the product category. The initial generation module generates an initial white background image of the product and descriptive scene prompts based on the product type and preset prompt word template; The background replacement module inputs the initial product white background image and the descriptive scene prompts into the background generation and fusion model to generate a background replacement image; The verification output module uses an image segmentation model to segment the product white background image and the background replacement image, performs ratio calculation and consistency verification, and outputs the final background replacement image.

6. The product image background replacement device according to claim 5, characterized in that: The category identification module specifically includes: The size unit resizes the white background image of the product to 1024×1024 pixels. The identification unit inputs the white background image of the product into the Qwen-VL Max model to identify the product category; The initial generation module specifically includes: The magnification unit retrieves the magnification from a preset mapping library based on the product type: the preset mapping library contains magnifications corresponding to different product types; Replace the white background image unit. Based on the magnification factor, determine whether it is necessary to replace the product in the white background image with a larger white background image. If it is necessary, replace it. If it is not necessary, retain the original white background image and output the initial white background image. The prompt word unit inputs the preset prompt word template and product category into the Qwen-VL Max model to generate prompt words describing the scenario.

7. A product image background replacement device according to claim 5, characterized in that: The background replacement module specifically includes: The model training unit trains the Kontext LoRA model using image data; the image data includes a product white background image, an initial product white background image, and a target image. The Kontext Dev model is used as the base model, and a pre-trained Kontext LoRA model is concatenated to construct the Kontext model. The background generation unit inputs the initial product white background image and the descriptive scene prompts into the Kontext model to generate a background image and outputs the background image. The image fusion unit inputs the background image and the initial white background image of the product into the Kontext model for fusion and outputs a background replacement image.

8. A product image background replacement device according to claim 5, characterized in that: The verification output module specifically includes: The segmentation unit performs segmentation on the white background image of the product to obtain the main image of the product before generation; the image segmentation model also performs segmentation on the background replacement image to obtain the main image of the product after generation. The percentage calculation unit calculates the first percentage of the product's main image in the product's white background image after generation, and the second percentage of the product's main image in the product's white background image before generation. The difference judgment unit determines whether the difference between the first ratio and the second ratio exceeds a preset threshold; if the difference does not exceed the preset threshold, the background replacement image is output as the final background replacement image; if the difference exceeds the preset threshold, the parameters are adjusted and the process returns to step S30 to regenerate until the difference does not exceed the preset threshold, and then the final background replacement image is output. The adjustment parameters in the difference judgment unit are specifically as follows: The third proportion subunit calculates the third proportion if the difference exceeds a preset threshold. The third proportion is the proportion of the initial product white background image to the product white background image. Compare the sub-units, compare the third ratio with the second ratio, and if the deviation between the third ratio and the second ratio is greater than a set value, then adjust the magnification of the product type. If the deviation between the third ratio and the second ratio is less than or equal to a set value, then the scene description prompts are adjusted by deleting the scaling description in the scene description prompts and adding the ratio constraint in the scene description prompts. Return to the sub-unit and re-execute the background replacement module.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 4.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 4.