BAGEL model-based commodity scene automatic supplement method, apparatus and device, and medium

By combining BiRefNet and BAGEL models, the problems of scene adaptation and edge blending in the generation of product images were solved, and the coordination between products and scenes and the visual effects were improved.

CN121600101APending Publication Date: 2026-03-03FUJIAN ZIXUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511450470.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing product image scene generation technologies suffer from insufficient scene adaptability, ambiguous semantic expression, and edge blending defects, resulting in inconsistencies between the scene and product style, logical contradictions, and poor visual naturalness.

Method used

The BiRefNet model is used to process the main image of the product, a preset prompt word template is constructed, the BAGEL model is used to generate scene description, and Gaussian filter kernel parameters and high contrast preservation technology are combined to achieve consistent lighting, spatial logic adaptation and natural transition between the product and the scene.

Benefits of technology

It improves the balance of the proportion of goods in the scene and the coordination of composition, enhances the adaptability of the generated image to the actual needs, achieves natural integration of lighting and color tone, smooth boundary transition, and improves the clarity of image details and visual naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600101A_ABST
    Figure CN121600101A_ABST
Patent Text Reader

Abstract

The invention provides an automatic commodity scene supplementing method and device based on a BAGEL model, equipment and a medium, and the method comprises the steps: carrying out the processing of a set commodity image through employing a BiRefNet model, and obtaining a commodity main body image; constructing a preset cue word template, inputting the commodity main body graph into the multi-modal model, and generating a corresponding scene description cue word according to the preset cue word template; inputting a commodity main body graph and the scene description prompt word into a BAGEL model, wherein the BAGEL model generates a required scene commodity graph; determining Gaussian filtering kernel parameters according to the size of the scene commodity image, converting the scene commodity image into a scene commodity grey-scale image, and performing Gaussian blur on the scene commodity grey-scale image through the Gaussian filtering kernel parameters to obtain a scene commodity fuzzy image; and mixing the commodity graph and the scene commodity fuzzy graph according to a soft light mode to obtain a scene commodity final graph, so that the proportion of the commodity in the scene is balanced, and the composition coordination is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of technology, and in particular to a method, apparatus, device, and medium for automatically supplementing product scenarios based on a BAGEL model. Background Technology

[0002] Current product image scene generation technology suffers from the following core bottlenecks: Insufficient scene adaptability: Traditional methods rely on manually designed scene templates or fixed style libraries, making it difficult to dynamically generate matching scenes based on product characteristics, leading to inconsistencies between the scene and product style (e.g., material reflection, lighting direction conflicts). Ambiguous semantic expression: Cue word generation relies on generic templates, lacking deep semantic analysis of product materials and environmental requirements, resulting in logical contradictions in the generated scene (e.g., a water cup conflicting with a desert background). Edge blending defects: The boundary transition between the product and the generated scene is abrupt, with color differences or shadow misalignment issues, affecting the overall visual naturalness. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for automatic supplementation of commodity scenes based on BAGEL model, so as to balance the proportion of commodities in the scene and improve the composition coordination.

[0004] In a first aspect, the present invention provides a method for automatically supplementing product scenarios based on the BAGEL model, comprising the following steps:

[0005] Step 1: Process the given product image using the BiRefNet model to obtain the main product image;

[0006] Step 2: Construct a preset prompt word template, input the main image of the product into the multimodal model, and generate corresponding scene description prompt words according to the preset prompt word template;

[0007] Step 3: Input the main product image and the scene description prompts into the BAGEL model, and the BAGEL model will generate the required scene product image;

[0008] Step 4: Determine the Gaussian filter kernel parameters based on the dimensions of the scene product image. Then, convert the scene product image into a grayscale image. Next, apply Gaussian blur to the grayscale image using the Gaussian filter kernel parameters to obtain a blurred image. Then, perform high contrast calculation on the blurred image and the grayscale image to obtain a high contrast preserved image. Finally, blend the product image and the high contrast preserved image in a soft light mode to obtain the final scene product image.

[0009] Secondly, the present invention provides an automatic product scene replenishment device based on the BAGEL model, comprising:

[0010] The module for obtaining the main product image uses the BiRefNet model to process the given product image and obtain the main product image.

[0011] A prompt word module is constructed, a preset prompt word template is built, the main image of the product is input into the multimodal model, and corresponding scene description prompt words are generated according to the preset prompt word template;

[0012] The generation and synthesis module inputs the main product image and the scene description prompts into the BAGEL model, and the BAGEL model generates the required scene product image.

[0013] The image processing module determines the Gaussian filter kernel parameters based on the size of the scene product image, then converts the scene product image into a scene product grayscale image, then applies Gaussian blur to the scene product grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image, then performs high contrast calculation on the blurred scene product image and the scene product grayscale image to obtain a high contrast preserved image; finally, the product image and the high contrast preserved image are blended in a soft light mode to obtain the final scene product image.

[0014] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect.

[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0016] One or more technical solutions provided by this invention have at least the following technical effects or advantages:

[0017] This invention uses a proportion adjustment mechanism to ensure that the products occupy a balanced proportion in the scene, thereby improving the compositional harmony.

[0018] This invention significantly enhances the correlation between the generated prompt words and the product feature scene, making the prompt words better adapt to the multimodal model and making the generated images more in line with actual needs;

[0019] This invention achieves natural blending of lighting and color tone through the BAGEL model, resulting in smooth boundary transitions. Furthermore, it utilizes high-contrast preservation technology with dynamic parameters and multi-stage processing to enhance the detail clarity of images of different sizes while maintaining the naturalness of the image and highlighting core visual information.

[0020] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention;

[0023] Figure 2 This is a schematic diagram of the device in Embodiment 2 of the present invention. Detailed Implementation

[0024] This application provides a method, apparatus, device, and medium for automatically replenishing product scenarios based on a BAGEL model.

[0025] The overall concept of the technical solution in this application is as follows:

[0026] The Janus-Pro-7B model is a unified multimodal understanding and generation model launched by DeepSeek-AI. Built upon the DeepSeek-LLM-7B base model, it decouples the understanding and generation tasks by separating the visual encoding path, employing a SigLIP-L visual encoder and a generation tagger with a downsampling rate of 16. It supports tasks such as text-to-image generation.

[0027] BiRefNet (Bilateral Reference Network) is a high-resolution binary image segmentation model proposed by Nankai University and other institutions. It is used to solve the task of fine segmentation of foreground and background, and supports applications such as background removal, mask generation, camouflaged object detection (COD), and salient object detection (SOD).

[0028] BAGEL-7B-MoT model: A multimodal foundational model developed by ByteDance-Seed. It adopts the Mixture-of-Transformer-Experts architecture, possesses powerful cross-modal understanding and generation capabilities, supports image editing, scene generation, and visual element matching, and can achieve natural integration of products with new scenes.

[0029] High contrast preservation: an image enhancement technique that enhances the details in areas of strong contrast between light and dark in an image while smoothing low-contrast areas, thereby improving the image's sense of depth and the recognizability of key information, making the main product stand out more in the scene.

[0030] Specifically, it includes the following:

[0031] 1. Precise image cutout stage of the product subject

[0032] The BiRefNet high-resolution segmentation model is used to refine the input product images, performing pixel-level semantic segmentation on the product's outline and texture (especially complex edges such as transparent materials, plush materials, and mesh textures) to generate a foreground mask containing soft edge information. The main product is extracted through mask operations, and the original background is removed, ensuring that the product's edge details (such as hair strands and glass reflection boundaries) are fully preserved, laying a high-precision foundation for subsequent scene fusion.

[0033] 2. Adjustment phase of commodity proportion

[0034] Calculate the pixel percentage of the foreground mask (product subject) in the original image after image matting:

[0035] If the ratio is ≤70%: directly retain the main image of the product after cutout and proceed to the next stage;

[0036] If the ratio is greater than 70%, a white background image with a size 1.8 times that of the original image will be automatically generated. At this ratio, the subsequent BAGEL model can generate the required image better. The main product image is placed in the center on the white background image. By expanding the background canvas, the visual proportion of the main product is reduced, and the product proportion is avoided in the subsequent generated scene.

[0037] 3. Automatic generation stage of scene prompts

[0038] Based on the Janus-Pro-7B multimodal model, a preset prompt word template is constructed (template format: "product category + core features (material / color / shape) + scene style (e.g., natural / urban / retro) + environmental elements (e.g., vegetation / building / props) + lighting conditions (e.g., sidelight / backlight / soft light)"). The model receives the product image after it has been cut out, and through visual feature analysis (e.g., identifying the product as a "black leather handbag" and extracting features such as "matte texture and square outline"), it automatically generates scene description prompt words that are highly adapted to the product by combining the template logic (example: "The black leather handbag is placed on the bar counter of a retro coffee shop, with warm yellow light shining obliquely, and the background has wooden furniture and retro murals, with a brownish-yellow warm color tone").

[0039] 4. Scene fusion and generation stage

[0040] The product image obtained through image matting and the automatically generated prompts are input into the BAGEL-7B-MoT model. The model leverages its Mixture-of-Transformer-Experts architecture's cross-modal fusion capabilities to achieve the following when generating new scenes:

[0041] Visual element matching: Adjust the light and shadow distribution and color tone of the scene according to the lighting direction and color tone of the main product to ensure that the lighting effects of the product and the scene are consistent;

[0042] Spatial logic adaptation: Based on the product's posture and outline, generate a scene space that conforms to the principles of perspective (such as matching the product's placement angle with the slope of the scene's ground).

[0043] Seamless integration: Through the model's built-in edge optimization mechanism, the sense of separation between the product and the new scene is eliminated, achieving a natural transition.

[0044] 5. High-contrast retention optimization stage

[0045] Detail enhancement is achieved through dynamic parameter adjustment and multi-stage processing. The specific process is as follows:

[0046] Dynamic parameter calculation: The Gaussian filter kernel parameters (kernel size and standard deviation σ) are determined based on the image size using a function.

[0047] High-contrast feature extraction:

[0048] The image generated in step 4 is converted to grayscale using the "maximum and minimum average method" (formula: value = (max_pixel + min_pixel) / 2);

[0049] Gaussian blur is applied to the grayscale image based on the Gaussian filter kernel parameters, and the high-pass image is calculated (formula: highpass = grayscale image - blurred image + 0.5).

[0050] Multi-level soft light blending:

[0051] First blending: Blend the product image from step 1 with the high-contrast image using the soft light mode;

[0052] Second fusion: Repeat high-contrast extraction and soft-light blending of the first fusion result to enhance detail levels;

[0053] Sharpness adjustment: The product image from step 1 is weighted and mixed with the secondary fusion result according to the α value (range [-1.0, 1.0]) (formula: final = original image * (1-α_abs) + secondary fusion image * α_abs). The final image is output after color gamut cropping (0-255).

[0054] Example 1

[0055] like Figure 1 As shown, this embodiment provides a method for automatically supplementing product scenarios based on the BAGEL model, including the following steps:

[0056] Step 1: Process the given product image using the BiRefNet model to obtain the main product image;

[0057] Step 2: Construct a preset prompt word template, input the main image of the product into the multimodal model, and generate corresponding scene description prompt words according to the preset prompt word template;

[0058] Step 3: Input the main product image and the scene description prompts into the BAGEL model, and the BAGEL model will generate the required scene product image;

[0059] Step 4: Determine the Gaussian filter kernel parameters based on the dimensions of the scene product image. Then, convert the scene product image into a grayscale image. Next, apply Gaussian blur to the grayscale image using the Gaussian filter kernel parameters to obtain a blurred image. Then, perform high contrast calculation on the blurred image and the grayscale image to obtain a high contrast preserved image. Finally, blend the product image and the high contrast preserved image in a soft light mode to obtain the final scene product image.

[0060] In this embodiment, preferably, step 1 specifically involves: processing the set product image using the BiRefNet model, performing pixel-level semantic segmentation on the outline and texture of the product in the product image, and generating a foreground mask containing soft edge information; then extracting the main body of the product through mask operation to obtain the main body image of the product.

[0061] In this embodiment, preferably, step 1 specifically involves: processing the set product image using the BiRefNet model to obtain the main product image;

[0062] Calculate the pixel percentage of the main product image in the product image. If the percentage is less than or equal to a set threshold, then retain the main product image.

[0063] If the proportion is less than the set threshold, a white background image with a size that is a multiple of the product image is generated, and the main product image is centered on the white background image to obtain a new product theme image.

[0064] In this embodiment, preferably, step 4 specifically involves: determining the Gaussian filter kernel parameters based on the size of the scene product image; then converting the scene product image into a scene product grayscale image; then performing Gaussian blurring on the scene product grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image; performing high contrast calculation on the blurred scene product image and the product grayscale image to obtain a high contrast preserved image; and mixing the product image and the high contrast preserved image in a soft light mode to obtain a first scene product mixed image.

[0065] The first scene product image is Gaussian blurred according to the Gaussian filter kernel parameters, then high contrast is calculated with the scene product grayscale image, and then the product image is mixed with the product image in soft light mode to obtain the second scene product mixed image.

[0066] The clarity of the mixed product image in the second scene will be adjusted as follows:

[0067] Based on the weight threshold α, according to the formula: final = product image P1 n *(1-α)+P2 n *α, P1 n P2 is the pixel value of the nth pixel in the product image. n The value of the nth pixel in the second scene product blending image is given. The set threshold range is [-1.0, 1.0]. The second scene product blending image and the original image are calculated, and the calculation result is cropped through the color gamut to obtain the final scene product image.

[0068] Based on the same inventive concept, this application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0069] Example 2

[0070] like Figure 2 As shown, this embodiment provides an automatic product scene replenishment device based on the BAGEL model, including:

[0071] The module for obtaining the main product image uses the BiRefNet model to process the given product image and obtain the main product image.

[0072] A prompt word module is constructed, a preset prompt word template is built, the main image of the product is input into the multimodal model, and corresponding scene description prompt words are generated according to the preset prompt word template;

[0073] The generation and synthesis module inputs the main product image and the scene description prompts into the BAGEL model, and the BAGEL model generates the required scene product image.

[0074] The image processing module determines the Gaussian filter kernel parameters based on the size of the scene product image, then converts the scene product image into a scene product grayscale image, then applies Gaussian blur to the scene product grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image, then performs high contrast calculation on the blurred scene product image and the scene product grayscale image to obtain a high contrast preserved image; finally, the product image and the high contrast preserved image are blended in a soft light mode to obtain the final scene product image.

[0075] In this embodiment, preferably, the module for obtaining the main body of the product specifically involves: processing the set product image using the BiRefNet model, performing pixel-level semantic segmentation on the outline and texture of the product in the product image, and generating a foreground mask containing soft edge information; then extracting the main body of the product through mask operation to obtain the main body image of the product.

[0076] In this embodiment, preferably, the module for obtaining the main product image specifically involves: processing the set product image using a BiRefNet model to obtain the main product image;

[0077] Calculate the pixel percentage of the main product image in the product image. If the percentage is less than or equal to a set threshold, then retain the main product image.

[0078] If the proportion is less than the set threshold, a white background image with a size that is a multiple of the product image is generated, and the main product image is centered on the white background image to obtain a new product theme image.

[0079] In this embodiment, preferably, the image processing module specifically performs the following steps: determining the Gaussian filter kernel parameters based on the size of the scene product image; then converting the scene product image into a scene product grayscale image; then performing Gaussian blurring on the scene product grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image; performing high contrast calculation on the blurred scene product image and the product grayscale image to obtain a high contrast preserved image; and finally, mixing the product image and the high contrast preserved image in a soft light mode to obtain a first scene product mixed image.

[0080] The first scene product image is Gaussian blurred according to the Gaussian filter kernel parameters, then high contrast is calculated with the scene product grayscale image, and then the product image is mixed with the product image in soft light mode to obtain the second scene product mixed image.

[0081] The clarity of the mixed product image in the second scene will be adjusted as follows:

[0082] Based on the weight threshold α, according to the formula: final = product image P1 n *(1-α)+P2 n *α, P1 n P2 is the pixel value of the nth pixel in the product image. n The value of the nth pixel in the second scene product blending image is given. The set threshold range is [-1.0, 1.0]. The second scene product blending image and the original image are calculated, and the calculation result is cropped through the color gamut to obtain the final scene product image.

[0083] Since the apparatus described in Embodiment 2 of the present invention is an apparatus used to implement the method of Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the apparatus based on the method described in Embodiment 1 of the present invention, and therefore will not be described again here. All apparatuses used in the method of Embodiment 1 of the present invention fall within the scope of protection of the present invention.

[0084] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to Embodiment 1, as detailed in Embodiment 3.

[0085] Example 3

[0086] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement any of the implementation methods in Embodiment 1.

[0087] Since the electronic device described in this embodiment is the device used to implement the method in Embodiment 1 of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in Embodiment 1 of this application. Therefore, how the electronic device implements the method in the embodiment of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiment of this application falls within the scope of protection of this application.

[0088] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.

[0089] Example 4

[0090] This embodiment provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it can implement any of the implementation methods in Embodiment 1.

[0091] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0093] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0094] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0095] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for automatically supplementing product scenarios based on the BAGEL model, characterized in that: Includes the following steps: Step 1: Process the given product image using the BiRefNet model to obtain the main product image; Step 2: Construct a preset prompt word template, input the main image of the product into the multimodal model, and generate corresponding scene description prompt words according to the preset prompt word template; Step 3: Input the main product image and the scene description prompts into the BAGEL model, and the BAGEL model will generate the required scene product image; Step 4: Determine the Gaussian filter kernel parameters based on the dimensions of the scene product image. Then, convert the scene product image into a grayscale image. Next, apply Gaussian blur to the grayscale image using the Gaussian filter kernel parameters to obtain a blurred image. Then, perform high contrast calculation on the blurred image and the grayscale image to obtain a high contrast preserved image. Finally, blend the product image and the high contrast preserved image in a soft light mode to obtain the final scene product image.

2. The method for automatic product scene replenishment based on the BAGEL model according to claim 1, characterized in that: Step 1 specifically involves: using the BiRefNet model to process the set product image, performing pixel-level semantic segmentation on the product's outline and texture to generate a foreground mask containing soft edge information; then extracting the product subject through mask operations to obtain the product subject image.

3. The method for automatic product scene replenishment based on the BAGEL model according to claim 1, characterized in that: Step 1 specifically involves: processing the set product image using the BiRefNet model to obtain the main product image; Calculate the pixel percentage of the main product image in the product image. If the percentage is less than or equal to a set threshold, then retain the main product image. If the proportion is less than the set threshold, a white background image with a size that is a multiple of the product image is generated, and the main product image is centered on the white background image to obtain a new product theme image.

4. The method for automatic product scene replenishment based on the BAGEL model according to claim 1, characterized in that: Step 4 specifically involves: determining the Gaussian filter kernel parameters based on the size of the scene product image; then converting the scene product image into a grayscale image; subsequently, applying Gaussian blur to the grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image; performing high-contrast calculations on the blurred scene product image and the grayscale image to obtain a high-contrast preserved image; and finally, blending the product image and the high-contrast preserved image in a soft light mode to obtain a first scene product mixed image. The first scene product image is Gaussian blurred according to the Gaussian filter kernel parameters, then high contrast is calculated with the scene product grayscale image, and then the product image is mixed with the product image in soft light mode to obtain the second scene product mixed image. The clarity of the mixed product image in the second scene will be adjusted as follows: Based on the weight threshold α, according to the formula: final = product image P1 n *(1-α)+P2 n *α, P1 n P2 is the pixel value of the nth pixel in the product image. n The value of the nth pixel in the second scene product blending image is given. The set threshold range is [-1.0, 1.0]. The second scene product blending image and the original image are calculated, and the calculation result is cropped through the color gamut to obtain the final scene product image.

5. An automatic product scene replenishment device based on a BAGEL model, characterized in that: include: The module for obtaining the main product image uses the BiRefNet model to process the given product image and obtain the main product image. A prompt word module is constructed, a preset prompt word template is built, the main image of the product is input into the multimodal model, and corresponding scene description prompt words are generated according to the preset prompt word template; The generation and synthesis module inputs the main product image and the scene description prompts into the BAGEL model, and the BAGEL model generates the required scene product image. The image processing module determines the Gaussian filter kernel parameters based on the size of the scene product image, then converts the scene product image into a scene product grayscale image, then applies Gaussian blur to the scene product grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image, then performs high contrast calculation on the blurred scene product image and the scene product grayscale image to obtain a high contrast preserved image; finally, the product image and the high contrast preserved image are blended in a soft light mode to obtain the final scene product image.

6. The automatic product scene replenishment device based on the BAGEL model according to claim 5, characterized in that: The module for obtaining the main product image specifically involves: processing the given product image using the BiRefNet model, performing pixel-level semantic segmentation on the product's outline and texture in the product image, and generating a foreground mask containing soft edge information; then extracting the main product image through mask operations.

7. The automatic product scene replenishment device based on the BAGEL model according to claim 5, characterized in that: The module for obtaining the main product image specifically involves processing the given product image using a BiRefNet model to obtain the main product image. Calculate the pixel percentage of the main product image in the product image. If the percentage is less than or equal to a set threshold, then retain the main product image. If the proportion is less than the set threshold, a white background image with a size that is a multiple of the product image is generated, and the main product image is centered on the white background image to obtain a new product theme image.

8. The automatic product scene replenishment device based on the BAGEL model according to claim 5, characterized in that: The image processing module specifically comprises: determining the Gaussian filter kernel parameters based on the size of the scene product image; converting the scene product image into a scene product grayscale image; then applying Gaussian blur to the scene product grayscale image using the Gaussian filter kernel parameters to obtain a blurred scene product image; performing high contrast calculation on the blurred scene product image and the product grayscale image to obtain a high contrast preserved image; and mixing the product image and the high contrast preserved image in a soft light mode to obtain a first scene product mixed image. The first scene product image is Gaussian blurred according to the Gaussian filter kernel parameters, then high contrast is calculated with the scene product grayscale image, and then the product image is mixed with the product image in soft light mode to obtain the second scene product mixed image. The clarity of the mixed product image in the second scene will be adjusted as follows: Based on the weight threshold α, according to the formula: final = product image P1 n *(1-α)+P2 n *α, P1 n P2 is the pixel value of the nth pixel in the product image. n The value of the nth pixel in the second scene product blending image is given. The set threshold range is [-1.0, 1.0]. The second scene product blending image and the original image are calculated, and the calculation result is cropped through the color gamut to obtain the final scene product image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 4.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 4.