Commodity map generation method, device and equipment for eliminating article suspension and medium
By cutting pictures, processing images, generating descriptive words and background prompt words, obtaining line drawings and depth infographics, and inputting them into the diffusion model, the problem of item suspension during product images is solved, and the realism and quality of the image is improved.
Patent Information
- Application Number
- CN202510108225.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to accurately understand and process the characteristics of items in masks, resulting in misjudgment of the masks when the product image is generated, resulting in suspended visual effects, affecting the quality and sense of reality of the image.
The main image of the product is obtained by cutting the image and the masked grayscale image is generated, and the processed into a white background image is used to generate descriptive words and background prompt words using a large language model. The line drawing and depth infographic images are obtained by combining the Canny algorithm and the depth estimation model, and input them to the diffusion model to generate the final image.
It significantly improves the success rate of the product placement plane, optimizes the product display effect, and enhances the realism and quality of the image.
Smart Images

Figure CN120070621A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of generating pictures, and particularly relates to a method, device, equipment and medium for generating product pictures that eliminate the suspension of objects. Background Art
[0002] In e-commerce platforms, product pictures need to be displayed. Some existing e-commerce platforms use software to automatically generate pictures. The generation process of existing product images involves many complex and subtle elements. Among them, a particularly challenging task is to find and determine a suitable placement plane to fit various products, so as to avoid giving people an unrealistic or floating visual effect in the product images. The reason for this problem is that it is difficult for the depth diffusion model to accurately understand and process the features of objects in the mask. This leads to the model may misjudge the mask of the object when trying to find the plane where the object should be placed, and then produce incorrect positions in the generated image, thus affecting the overall quality and realism of the product image. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for generating product pictures that eliminate the suspension of objects, which can better help products find a suitable plane, not only significantly improve the success rate of determining the placement plane of products, but also optimize the display effect of products.
[0004] In a first aspect, the present invention provides a method for generating product pictures that eliminate the suspension of objects, including the following steps:
[0005] Step 1: Cut out the product picture to obtain the main body picture of the product, and generate the corresponding grayscale mask picture of the product.
[0006] Step 2: Process the main body picture of the product to obtain a main body white-background picture with a white background.
[0007] Step 3: Input the main body white-background picture into a large language model to obtain a main body description word; input the main body description word and the background description word into the large language model to generate the required background description prompt word.
[0008] Step 4: Obtain the line drawing of the main body white-background picture through the Canny algorithm; obtain the depth information picture of the main body white-background picture through the depth estimation model.
[0009] Step 5: Input the line drawing, depth information picture, background description prompt word, main body white-background picture and grayscale mask picture of the product into the diffusion model to generate the required product generation picture.
[0010] In a second aspect, the present invention provides a device for generating product pictures that eliminate the suspension of objects, including:
[0011] Extract the main body module, perform matte extraction on the product image to obtain the product main body image, and generate the corresponding grayscale product matte image;
[0012] White background processing module, process the product main body image to obtain the main body white background image with a white background;
[0013] Prompt generation module, input the main body white background image into the large language model to obtain the main body description words; input the main body description words and the background description words into the large language model to generate the required background description prompts;
[0014] Guidance information module, obtain the line drawing of the main body white background image through the Canny algorithm; obtain the depth information map of the main body white background image through the depth estimation model;
[0015] Image generation module, input the line drawing, depth information map, background description prompts, main body white background image, and grayscale product matte image into the diffusion model to generate the required product generation image.
[0016] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.
[0018] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0019] The technical solution of the present invention can better help the product find a suitable plane, not only significantly improving the success rate of determining the product placement plane, but also optimizing the display effect of the product, contributing to improving the competitiveness of the enterprise and enabling more users to use the enterprise's mapping products;
[0020] The present invention adopts the visual function based on the large language model to generate variable and user-customizable prompts. First, it can greatly improve the generation quality of the background prompts and ensure their close correlation with the actual background scene. Second, it optimizes the expression effect of the language model, enabling it to generate more realistic language content. Effectively generate background prompts to adapt to the StableDiffusion model, realize more practical images, and meet user needs.
[0021] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings
[0022] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.
[0023] Figure 1 It is a flowchart of the method in the first embodiment of the present invention;
[0024] Figure 2 It is a schematic structural diagram of the device in the second embodiment of the present invention. Detailed Embodiments
[0025] The embodiments of the present application provide a method, device, equipment and medium for generating a product image that eliminates the suspension of objects, and solve the technical problem in the prior art of how to find and determine a suitable placement plane to adapt to the product.
[0026] The overall idea of the technical solution in the embodiments of the present application is as follows:
[0027] Use a matte extraction algorithm to extract the product main body from the original image and generate a matte grayscale image; then, fill the background with a white background image; then, extract the product description based on Microsoft's Florence2 model, and combine with LLAMA to generate a suitable background description prompt. Subsequently, use the Canny operator to extract the edge information of the white background image, and generate a depth map through a depth estimation model. Finally, redraw the background of the composite image based on an open-source diffusion model (such as Stable Diffusion), and combine the edge map, depth map and prompt to generate the final image, thereby improving the visual effect of the product image. The implementation steps are as follows:
[0028] Step 1: Extract the product main body from the original product image and generate a corresponding product matte grayscale image.
[0029] Step 2: Process the background of the product image and fill it with a white background image.
[0030] Among them, the calculation formula for filling the white background can be expressed as:
[0031] S = S 1 *a + S 2 *(1 - a)
[0032] Among them, S is the filled image, S 1 is the original product image, S 2It is a white background image, and a is the grayscale image of the commodity mask in Step 1. This formula generates the final filling effect by weighted combining the original image and the white background image according to the grayscale values.
[0033] Step 3: Based on the open-source Florence2 model of Microsoft, accurately extract the description of the commodity main body.
[0034] Microsoft's Florence2 model is an advanced multi-task vision lightweight model. It adopts a prompt-based architecture and a sequence-to-sequence architecture, is based on Transformer, and is trained using the large FLD-5B dataset. It integrates multiple functions such as picture text recognition, object detection, semantic segmentation, image description generation, and visual question answering. It can execute tasks through simple text prompts and performs excellently in zero-shot and fine-tuning settings. Compared with large models of the same type, Florence2 is faster, takes up less space, has higher efficiency, and provides a fine-tunable version and multiple modes, which can better meet the needs of different tasks and can be applied to multiple fields such as e-commerce, security, and document processing.
[0035] In the embodiment of this application, the Florence2 model and related resources are obtained through the Hugging Face platform. After downloading and installing, initialization and other operations are performed to prepare the running environment. The selected version of the model is large-ft. The selected task is dense_region_caption. The image to be detected is input, and after the model runs, the detection result will be output, and the result will display the basic element description information and position in the image, etc.
[0036] Step 4: Based on large language models such as LLAMA, combine the main body description in Step 3 with large language prompt words to generate corresponding background description prompt words.
[0037] Exemplarily, in an embodiment of the present application, the subject is described as "A man carrying a bag". The large language prompt is: "I want you to act as a photographer. I will provide you with a picture. You need to create a suitable background for the items in this picture." The output result is: "A man carrying a bag standing on the street, shops and storefronts, pedestrians walking by, street signs, natural light, front view, detailed texture, High-Resolution Image."
[0038] Step Five: Use the Canny operator of image gradient in OpenCV to extract the Canny line drawing of the white-background image in Step Two.
[0039] Canny edge detection can enhance the edge features of the image, thus better controlling the structure and contour of the image during the generation process. By extracting clear edge information, this method makes the generated image more accurate in details and shape.
[0040] Step Six: Use the depth estimation model to extract the depth information map of the white-background image in Step Two.
[0041] The depth estimation model is mainly reflected in the generation and enhancement of depth information. By performing depth estimation on a single image, the depth estimation model can provide a more abundant sense of space and hierarchy for the generated image. This depth information can help Stable Diffusion better understand the three-dimensional structure of the scene during image synthesis, thereby generating a more realistic and three-dimensional image effect.
[0042] Running steps:
[0043] 1. Make necessary adjustments to the input image, such as scaling, cropping, etc., to meet the model input requirements.
[0044] 2. Normalize the image pixel values to improve the processing efficiency of the model.
[0045] 3. Load the depth estimation model, extract image features through the convolutional layer, and generate feature maps.
[0046] 4. The model generates a depth map based on the extracted feature maps and outputs the depth values of each pixel.
[0047] 5. Post-process the depth map, such as smoothing and denoising, to improve the quality of depth estimation.
[0048] In the depth estimation model, multiple adaptations are made to traditional methods. First, the model is trained with a white background image. This background selection can effectively reduce interference, enabling Stable Diffusion to generate a more free background. Second, the model has the ability to automatically identify the subject and generate a planar depth map for it. This optimization enables the model to better solve the problem of objects floating in the air, thereby enhancing the stability of background generation.
[0049] Step 7: Redraw the background of the composite image based on an open-source diffusion model (such as StableDiffusion).
[0050] Exemplarily, in an embodiment of the present application, the model of Stable Diffusion is selected as: RealisticVisionV60B1_v51VAE, and the input includes: 1. The Canny line drawing in Step 5. 2. The depth image in Step 6. 3. The prompt in Step 4. 4. The product image in Step 2. 5. The grayscale product mask image in Step 1. Parameter settings: The redrawing amplitude is 1, the control intensity of Canny is 0.35, and the control intensity of the depth map is 0.65. Among them, the Canny line drawing and the depth image are mainly used to control the edges of the product subject in the generated image, while the prompt serves as a guide for generating background elements. The product image can specify the size of the generated image, and the grayscale product mask image controls that the product subject part does not participate in the generation.
[0051] Embodiment 1
[0052] As Figure 1 shown, this embodiment provides a method for generating a product image that eliminates the floating of objects, including the following steps:
[0053] Step 1: Cut out the product image to obtain the product subject image and generate the corresponding grayscale product mask image;
[0054] Step 2: Process the product subject image to obtain a subject white background image with a white background;
[0055] Step 3: Input the subject white background image into the large language model to obtain the subject description word; input the subject description word and the background description word into the large language model to generate the required background description prompt;
[0056] Step 4: Obtain the line drawing of the subject white background image through the Canny algorithm; obtain the depth information map of the subject white background image through the depth estimation model;
[0057] Step 5: Input the line drawing, depth information map, background description prompt, main body white background map, and commodity mask grayscale map into the diffusion model to generate the required commodity generation map.
[0058] In this embodiment, preferably, the specific operation of step 1 is as follows: Determine the pixels of the commodity map. If the pixels of the commodity map are less than 2000×2000, directly perform the matting operation through the Visual Intelligence Open Platform to obtain the required commodity main body map. If the pixels of the commodity map are greater than or equal to 2000×2000, use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the commodity map proportionally to obtain a scaled map. Perform the matting operation on the scaled map through the Visual Intelligence Open Platform to obtain the first mask map. Use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to restore the first mask map to the original size proportionally to obtain the second mask map. Read the alpha channel of each pixel point in the commodity map to obtain the first matrix, read the alpha channel of each pixel point in the second mask map to obtain the second matrix, and call Core.min to merge the first matrix and the second matrix to obtain the third matrix. Replace the alpha channel in the commodity map with the third matrix to obtain the required commodity main body map.
[0059] Due to the matting size limit of the Alibaba Cloud Visual Intelligence Open Platform, when the longest side exceeds 2000 pixels, it is necessary to perform proportional scaling. Call Imgproc.resize with the Imgproc.INTER_LANCZOS4 algorithm to perform image scaling.
[0060] Call Imgproc.resize with the Imgproc.INTER_LANCZOS4 algorithm to perform image scaling. This algorithm can reduce the impact of artifacts while maintaining the edge sharpness.
[0061] Obtain the matting result according to the black and white map + original image:
[0062] a. Take out the alpha channel of the original image. Call Core.min, which will take the minimum value of the transparency channels of the mask and the original image, ensuring that the original information of the transparency channel is retained. Through this step of processing, the lines of the obtained matting can be made more perfect.
[0063] b. Remove the alpha channel of the original image and add the black and white map as the new alpha channel for layer merging.
[0064] c. Set the color value of the transparent area to black. This step is to reduce the size of the matting result image and save storage costs.
[0065] The core code is as follows:
[0066] / / Create a fully transparent matrix the same size as the original image for comparison;
[0067] Mat compareAlpha = new Mat(outImg.size(), CvType.CV_8UC1, Scalar.all(0.0));
[0068] / / Used to save the comparison result;
[0069] Mat compareResult = new Mat();
[0070] / / Compare the alpha channel of the original image with alpha. The positions with a value of 0 in the resulting mask are the transparent positions in the original image;
[0071] Core.compare(outPlanes.get(3), compareAlpha, compareResult, Core.CMP_EQ);
[0072] / / Create a fully black matrix the same size as the original image;
[0073] Mat black = new Mat(outImg.size(), outImg.type(), Scalar.all(0));
[0074] / / Copy the black color in black to the corresponding positions in outImg according to the mask;
[0075] Core.bitwise_and(black, outImg, outImg, compareResult);
[0076] Generate the corresponding grayscale mask image of the commodity based on the main image of the commodity.
[0077] In this embodiment, preferably, step 3 is specifically as follows: Input the main body white background image into the large language model to obtain the main body description words; Input the set scene prompt words into the large language model, and then input the user-defined scene and the main body description words into the large language model to obtain the background description words; Input the set position prompt words into the large language model, and then input the background description words and the main body description words into the large language model to obtain the position description words; Input the set lighting perspective prompt words into the large language model, and then input the main body white background image into the large language model to obtain the lighting perspective description words; Input the set background prompt words into the large language model, and then input the main body description words, background description words, position description words, and lighting perspective description words into the large language model to obtain the background description prompt words.
[0078] In this embodiment, preferably, step 5 is specifically as follows: Input the line drawing, depth information map, background description prompt words, main body white background image, and commodity mask grayscale image into the diffusion model to generate the required commodity generation image; Set the control intensity of Canny in the diffusion model to 0.35, the control intensity of the depth map to 0.65, and the redrawing amplitude to 1.
[0079] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.
[0080] Embodiment 2
[0081] As Figure 2 shown, in this embodiment, a commodity image generation apparatus for eliminating the suspension of items is provided, including:
[0082] A main body extraction module that performs matte extraction on the commodity image to obtain the commodity main body image and generates the corresponding commodity mask grayscale image;
[0083] A white background processing module that processes the commodity main body image to obtain a main body white background image with a white background;
[0084] A prompt word generation module that inputs the main body white background image into the large language model to obtain the main body description words; Inputs the main body description words and the background description words into the large language model to generate the required background description prompt words;
[0085] A guidance information module that obtains the line drawing of the main body white background image through the Canny algorithm; Obtains the depth information map of the main body white background image through the depth estimation model;
[0086] A generated image module that inputs the line drawing, depth information map, background description prompt words, main body white background image, and commodity mask grayscale image into the diffusion model to generate the required commodity generation image.
[0087] In this embodiment, preferably, the main body extraction module is specifically as follows: judge the pixels of the product image. If the pixels of the product image are less than 2000×2000, directly perform the image extraction operation through the visual intelligence open platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image proportionally to obtain a scaled image; perform the image extraction operation on the scaled image through the visual intelligence open platform to obtain a first mask image; use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to restore the first mask image proportionally to the original size to obtain a second mask image of the original size; read the alpha channel of each pixel point in the product image to obtain a first matrix, read the alpha channel of each pixel point in the second mask image to obtain a second matrix, and call Core.min to merge the first matrix and the second matrix to obtain a third matrix; replace the alpha channel in the product image with the third matrix to obtain the required product main body image;
[0088] Generate a corresponding product mask grayscale image according to the product main body image.
[0089] In this embodiment, preferably, the prompt word generation module is specifically as follows: input the main body white background image into the large language model to obtain the main body description word; input the set scene prompt word into the large language model, and then input the user-defined scene and the main body description word into the large language model to obtain the background description word; input the set position prompt word into the large language model, and then input the background description word and the main body description word into the large language model to obtain the position description word; input the set lighting perspective prompt word into the large language model, and then input the main body white background image into the large language model to obtain the lighting perspective description word; input the set background prompt word into the large language model, and then input the main body description word, background description word, position description word, and lighting perspective description word into the large language model to obtain the background description prompt word.
[0090] In this embodiment, preferably, the image generation module is specifically as follows: input the line drawing, depth information map, background description prompt word, main body white background image, and product mask grayscale image into the diffusion model to generate the required product generated image; set the control intensity of Canny in the diffusion model to 0.35, the control intensity of the depth map to 0.65, and the redrawing amplitude to 1.
[0091] Since the device described in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method described in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, and thus will not be elaborated herein. Any device adopted for the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0092] Based on the same inventive concept, this application provides an embodiment of an electronic device corresponding to the first embodiment. For details, see the third embodiment.
[0093] The third embodiment
[0094] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.
[0095] Since the electronic device described in this embodiment is the device adopted for implementing the method in the first embodiment of this application, based on the method described in the first embodiment of this application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, how this electronic device implements the method in the embodiments of this application will not be introduced in detail herein. Any device adopted by those skilled in the art for implementing the method in the embodiments of this application falls within the scope of protection of this application.
[0096] Based on the same inventive concept, this application provides a storage medium corresponding to the first embodiment. For details, see the fourth embodiment.
[0097] The fourth embodiment
[0098] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in the first embodiment can be realized.
[0099] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0103] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than limiting the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of the claims of the present invention.
Claims
1. A method for generating a commodity image to eliminate hanging items, characterized in that: The steps include: Step 1: Cut out the product image, obtain the main product image, and generate the corresponding product mask grayscale image; Step 2: Process the main image of the product to obtain a main image with a white background; Step 3: Input the subject white background image into the large language model to obtain subject description words; input the subject description words and background description words into the large language model to generate the required background description prompt words; Step 4: Obtain a line drawing of the main white background image by using a Canny algorithm; Acquire a depth information map of the subject white background map through a depth estimation model; Step 5: Input the line drawing, depth information map, background description prompt words, main body white background map and product mask grayscale map into the diffusion model to generate the required product generation map.
2. A method for generating a commodity image to eliminate hanging items according to claim 1, characterized in that: The step 1 is specifically as follows: judging the pixels of the product image, if the pixels of the product image are less than 2000×2000, directly performing a cutout operation through the visual intelligence open platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, using Imgproc.resize in OpenCV and using the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image in proportion to obtain a scaled image; performing a cutout operation on the scaled image through the visual intelligence open platform to obtain a first mask image; using Imgproc.resize in OpenCV and using the Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the first mask image to a second mask image of the original size; reading out the transparent channel of each pixel in the product image to obtain a first matrix, reading out the transparent channel of each pixel in the second mask image to obtain a second matrix, calling Core.min to merge the first matrix with the second matrix to obtain a third matrix; The third matrix is used to replace the transparent channel in the product image to obtain the required product main image; Generate a corresponding product mask grayscale image based on the product main image.
3. The method for generating a commodity image to eliminate hanging items according to claim 1, characterized in that: The step 3 is specifically as follows: inputting the subject white background image into the large language model to obtain subject description words; inputting the set scene prompt words into the large language model, and then inputting the user-defined scene and the subject description words into the large language model to obtain background description words; inputting the set location prompt words into the large language model, and then inputting the background description words and the subject description words into the large language model to obtain location description words; inputting the set lighting perspective prompt words into the large language model, and then inputting the subject white background image into the large language model to obtain lighting perspective description words; inputting the set background prompt words into the large language model, and then inputting the subject description words, background description words, location description words and lighting perspective description words into the large language model to obtain background description prompt words.
4. The method for generating a commodity image to eliminate hanging items according to claim 1, characterized in that: The step 5 is specifically as follows: inputting the line drawing, depth information map, background description prompt words, main body white background map and product mask grayscale map into the diffusion model to generate the required product generation map; setting the control strength of Canny in the diffusion model to 0.35, the control strength of the depth map to 0.65, and the redraw amplitude to 1.
5. A commodity image generation device for eliminating suspended items, characterized in that: include: The main body cutout module cuts out the product image, obtains the main body image of the product, and generates the corresponding product mask grayscale image; A white background processing module processes the main image of the product to obtain a main white background image with a white background; The prompt word generation module inputs the subject white background image into the large language model to obtain the subject description words; the subject description words and background description words are input into the large language model to generate the required background description prompt words; A guidance information module, which obtains a line drawing of the main white background image through a Canny algorithm; Acquire a depth information map of the subject white background map through a depth estimation model; The image generation module inputs the line drawing, depth information map, background description prompt words, main body white background map and product mask grayscale map into the diffusion model to generate the required product generation map.
6. The device for generating a commodity image to eliminate hanging items according to claim 5, characterized in that: The cutout main body module is specifically as follows: the pixels of the product image are judged. If the pixels of the product image are less than 2000×2000, the cutout operation is directly performed through the visual intelligence open platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, the product image is proportionally scaled using the Imgproc.resize in OpenCV and the Imgproc.INTER_LANCZ0S4 algorithm to obtain a scaled image; the scaled image is cutout through the visual intelligence open platform to obtain a first mask image; the first mask image is proportionally restored to a second mask image of the original size through the Imgproc.resize in OpenCV and the Imgproc.INTER_LANCZ0S4 algorithm; the transparent channel of each pixel in the product image is read out to obtain a first matrix, the transparent channel of each pixel in the second mask image is read out to obtain a second matrix, and Core.min is called to merge the first matrix with the second matrix to obtain a third matrix; The third matrix is used to replace the transparent channel in the product image to obtain the required product main image; Generate a corresponding product mask grayscale image based on the product main image.
7. The device for generating a commodity image to eliminate hanging items according to claim 5, characterized in that: The prompt word generation module is specifically as follows: inputting the main body white background image into the large language model to obtain the main body description words; inputting the set scene prompt words into the large language model, and then inputting the user-defined scene and the main body description words into the large language model to obtain the background description words; inputting the set location prompt words into the large language model, and then inputting the background description words and the main body description words into the large language model to obtain the location description words; inputting the set lighting perspective prompt words into the large language model, and then inputting the main body white background image into the large language model to obtain the lighting perspective description words; inputting the set background prompt words into the large language model, and then inputting the main body description words, background description words, location description words and lighting perspective description words into the large language model to obtain the background description prompt words.
8. The device for generating a commodity image to eliminate hanging items according to claim 5, characterized in that: The image generation module specifically includes: inputting the line drawing, depth information map, background description prompt words, main body white background map and product mask grayscale map into the diffusion model to generate the required product generation map; setting the Canny control strength in the diffusion model to 0.35, the depth map control strength to 0.65, and the redraw amplitude to 1.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.