Proportional failure eliminating commodity graph generation method and device, equipment and medium
Through technical means such as cutting and depth estimation models, the problem of imbalance in the composition ratio of the product graph and inaccurate determination of the placement plane is solved, and better product display effect and competitiveness are achieved.
Patent Information
- Application Number
- CN202510108224.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the composition ratio of the product map is prone to imbalance, resulting in the proportion of the main part of the product being too large or too small, affecting the product display effect, and it is difficult for the depth diffusion model to accurately understand and process the characteristics of the items in the mask, resulting in inaccurate determination of the product placement plane.
The main image of the product is obtained by cutting the image and generating a masked grayscale image. The process is processed into a white background image and input a large language model to obtain a descriptive word. Combined with the Canny algorithm and the depth estimation model to obtain a line drawing and depth infographic, input the diffusion model to generate a product generation image, and optimize the background and composition ratio.
It significantly improves the success rate of the product placement plane, optimizes the product display effect, makes the composition ratio of the product map more balanced, and improves the competitiveness of the enterprise.
Smart Images

Figure CN120070620A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to a method, device, equipment and medium for generating a commodity image by eliminating ratio failure. Background Art
[0002] In e-commerce platforms, product images need to be displayed. Some existing e-commerce platforms use software to automatically generate product images. The existing product image generation process involves many complex and subtle elements. Among them, the problem that often occurs in the generated images is that the composition ratio of the product image is unbalanced, resulting in the proportion of the main part of the product being too large or too small, making the proportion of the product image in the overall composition uncoordinated, resulting in poor product display effect.
[0003] In addition, the existing technology also needs to find and determine a suitable placement plane to adapt to various products when generating images, so as to avoid the product image giving people an unreal or suspended visual effect; the reason for this problem is that the deep diffusion model is difficult to accurately understand and process the characteristics of objects in the mask, which results in the model misjudging the mask of the object when trying to find the plane where the object should be placed, and then producing an incorrect position in the generated image, thereby affecting the overall quality and realism of the product image. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for generating a product image that eliminates proportion failure, which can ensure the balance of the composition proportion of the product image and place the product on a suitable plane, which not only significantly improves the success rate of determining the product placement plane, but also optimizes the display effect of the product.
[0005] In a first aspect, the present invention provides a method for generating a product image by eliminating ratio failure, comprising the following steps:
[0006] Step 1: Cut out the product image, obtain the main product image, and generate the corresponding product mask grayscale image;
[0007] Step 2: Process the main image of the product to obtain a main image with a white background;
[0008] Step 3: Input the subject white background image into the large language model to obtain subject description words; input the subject description words and background description words into the large language model to generate the required background description prompt words;
[0009] Step 4: obtaining a line drawing of the subject white background image by using a Canny algorithm; obtaining a depth information image of the subject white background image by using a depth estimation model;
[0010] Step 5: Input the line drawing, depth information map, background description prompt, white-background main body map, and grayscale commodity mask map into the diffusion model to generate the required commodity generation map.
[0011] In a second aspect, the present invention provides a commodity map generation device for eliminating proportion failure, including:
[0012] A main body extraction module that performs matte extraction on the commodity map to obtain the commodity main body map and generates a corresponding grayscale commodity mask map;
[0013] A white-background processing module that processes the commodity main body map to obtain a white-background main body white-background map;
[0014] A prompt generation module that inputs the white-background main body map into a large language model to obtain a main body description word; inputs the main body description word and the background description word into the large language model to generate the required background description prompt;
[0015] A guidance information module that obtains the line drawing of the white-background main body map through the Canny algorithm; obtains the depth information map of the white-background main body map through a depth estimation model;
[0016] A picture generation module that inputs the line drawing, depth information map, background description prompt, white-background main body map, and grayscale commodity mask map into the diffusion model to generate the required commodity generation map.
[0017] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in the first aspect.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.
[0019] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0020] The technical solution of the present invention can evenly place the proportion of the commodity main body in the commodity map in the generated picture; and can place the commodity main body on an appropriate plane, which not only significantly improves the success rate of determining the commodity placement plane, but also optimizes the display effect of the commodity, helps to improve the competitiveness of the enterprise, and enables more users to use the enterprise's mapping products;
[0021] The present invention adopts a vision function based on a large language model to generate variable and user-customizable prompt words. First, it can greatly improve the generation quality of background prompt words and ensure their close relevance to the actual background scene. Second, it optimizes the expression effect of the language model so that it can generate more realistic language content. It effectively generates background prompt words to adapt to the StableDiffusion model, realizes images that are more in line with actual use, and meets user needs.
[0022] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific implementation manners of the present invention. Brief Description of the Drawings
[0023] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.
[0024] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;
[0025] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Description of the Invention
[0026] The embodiments of the present application solve the technical problem in the prior art of how to find and determine a suitable placement plane to adapt to the commodity by providing a method, device, equipment and medium for generating a commodity image that eliminates proportional failure.
[0027] The overall idea of the technical solution in the embodiments of the present application is as follows:
[0028] An image matting algorithm is used to extract the commodity main body from the original image and generate a mask grayscale image; then, the background is filled with a white background image; then, the commodity description is extracted based on Microsoft's Florence2 model, and suitable background description prompt words are generated in combination with LLAMA. Subsequently, the edge information of the white background image is extracted using the Canny operator, and a depth map is generated through a depth estimation model. Finally, the background of the composite image is redrawn based on an open-source diffusion model (such as Stable Diffusion), and the final image is generated by combining the edge map, depth map and prompt words, thereby improving the visual effect of the commodity image. The implementation steps are as follows:
[0029] Step 1: Extract the commodity main body from the original commodity image and generate a corresponding commodity mask grayscale image.
[0030] Step 2: Process the background of the commodity image and fill it with a white background image.
[0031] Among them, the calculation formula for filling the white background can be expressed as:
[0032] S = S 1 *a + S 2 *(1 - a)
[0033] Among them, S is the filled image, S 1 is the original product image, S 2 is the white background image, and a is the grayscale image of the product mask in the first step. This formula generates the final filling effect by weighted combination of the original image and the white background image according to the grayscale values.
[0034] Step 3: Based on the Microsoft open-source Florence2 model, accurately extract the description of the product main body.
[0035] Microsoft's Florence2 model is an advanced multi-task vision lightweight model. It adopts a prompt-based architecture and a sequence-to-sequence architecture, is based on Transformer, and is trained using the large FLD-5B dataset. It integrates multiple functions such as picture text recognition, object detection, semantic segmentation, image description generation, and visual question answering. It can execute tasks through simple text prompts and performs well in zero-shot and fine-tuning settings. Compared with similar large models, Florence2 is faster, occupies less space, has higher efficiency, and provides a fine-tunable version and multiple modes, which can better meet different task requirements and can be applied to multiple fields such as e-commerce, security, and document processing.
[0036] In the embodiment of this application, the Florence2 model and related resources are obtained through the Hugging Face platform. After downloading and installing, initialization and other operations are performed to prepare the running environment. The version of the model selected is large-ft. The task selected is dense_region_caption. The image to be detected is input. After the model runs, the detection results will be output, and the results will display the basic element description information and positions in the image, etc.
[0037] Step 4: Based on large language models such as LLAMA, combine the main body description in Step 3 with large language prompt words to generate corresponding background description prompt words.
[0038] Exemplarily, in an embodiment of the present application, the subject is described as A man carrying abag. The large language prompt is: I want you to act as a photographer. I will provide you with a picture. You need to create a suitable background for the items in this picture. The output result is: A man carrying abag standing on the street, shops and storefronts, pedestrians walking by, street signs, natural light, front view, detailed texture, High-Resolution Image.
[0039] Step 5: Use the Canny operator of the image gradient in OpenCV to extract the Canny line drawing of the white background image in Step 2.
[0040] Canny edge detection can enhance the edge features of the image, thus better controlling the structure and contour of the image during the generation process. By extracting clear edge information, the generated image is more accurate in terms of details and shape.
[0041] Step 6: Use the depth estimation model to extract the depth information map of the white background image in Step 2.
[0042] The depth estimation model (MiDaS algorithm) is mainly reflected in the generation and enhancement of depth information. By performing depth estimation on a single image, the depth estimation model can provide a more abundant sense of space and hierarchy for the generated image. This depth information can help Stable Diffusion better understand the three-dimensional structure of the scene during image synthesis, thus generating a more realistic and three-dimensional image effect.
[0043] Running steps:
[0044] 1. Make necessary adjustments to the input image, such as scaling, cropping, etc., to meet the model input requirements.
[0045] 2. Normalize the image pixel values to improve the processing efficiency of the model.
[0046] 3. Load the depth estimation model, extract image features through the convolutional layer, and generate a feature map.
[0047] 4. The model generates a depth map based on the extracted feature map and outputs the depth value of each pixel.
[0048] 5. Post-process the depth map, such as smoothing and denoising, to improve the quality of depth estimation.
[0049] First, the depth estimation model is trained using white background images. This background selection can effectively reduce interference, enabling Stable Diffusion to generate backgrounds more freely. Second, the depth estimation model has the ability to automatically identify the subject and generate a planar depth map for it. This optimization enables the depth estimation model to better solve the problem of objects floating in the air, thereby enhancing the stability of background generation.
[0050] Step 7: Scale the product image proportionally to ensure that the long side resolution of the scaled image is 512 pixels.
[0051] Among them, scaling the image to 512 is because: The Stable Diffusion 1.5 version is mainly trained using images with a resolution of 512×512, so it has the best effect when generating images of this size. Adapting to the image size of the model training can effectively alleviate the problem of disproportionate composition.
[0052] Step 8: Redraw the background of the composite image based on an open-source diffusion model (such as stablediffusion).
[0053] Exemplarily, in an embodiment of the present application, the model of Stable Diffusion is selected as: RealisticVisionV60B1_v51VAE, and the inputs include: 1. The canny line drawing in Step 5. 2. The depth image in Step 6. 3. The prompt in Step 4. 4. The product image in Step 7. 5. The grayscale product mask in Step 1. Parameter settings: The redrawing amplitude is 1, the control intensity of canny is 0.65, and the control intensity of the depth map is 0.35. Among them, the Canny line drawing and the depth image are mainly used to control the edges of the product subject in the generated image, while the prompt serves as a guide for generating background elements. The product image can specify the size of the generated image, and the grayscale product mask controls that the product subject part does not participate in the generation.
[0054] Step 9: Scale the image generated in Step 8 to the size of the original product image.
[0055] Step 10: Redraw and perform high-definition magnification based on an open-source diffusion model (such as stablediffusion).
[0056] Exemplarily, in one embodiment of the present application, the model selected for Stable Diffusion is: RealisticVisionV60B1_v51VAE, and the input includes: 1. High-definition zoom prompt words 2. The generated image of step nine. 3. The product mask grayscale image of step one. Parameter setting: The redrawing amplitude is 0.3. Among them, the high-definition zoom prompt words are, for example: best quality, masterpiece, (photorealistic: 1), ultrahigh res, highres, illustration.media, delicate, 8k wallpaper, etc. The redrawing amplitude is 0.3 to avoid being large enough to make the background change too much after redrawing, which will cause the composition ratio in the image to be out of balance again.
[0057] Embodiment 1
[0058] like Figure 1 As shown, this embodiment provides a method for generating a product image to eliminate ratio failure, comprising the following steps:
[0059] A method for generating a product image to eliminate ratio failure, characterized in that it comprises the following steps:
[0060] Step 1: Cut out the product image, obtain the main product image, and generate the corresponding product mask grayscale image;
[0061] Step 2: Process the main image of the product to obtain a main image with a white background;
[0062] Step 3: Input the subject white background image into the large language model to obtain subject description words; input the subject description words and background description words into the large language model to generate the required background description prompt words;
[0063] Step 4: obtaining a line drawing of the subject white background image by using a Canny algorithm; obtaining a depth information image of the subject white background image by using a depth estimation model;
[0064] Step 5: Process the main white background image to obtain a main scaled image with a long side resolution of 512 pixels, input the line drawing image, depth information image, background description prompt words, main scaled image and product mask grayscale image into the diffusion model to generate the required product generation image;
[0065] Step 6: Restore the product generated image to the product Figure 1 The size of the sample is obtained to obtain the product generation zoom image, and then the high-definition zoom prompt words, the product generation zoom image and the product mask grayscale image are input into the diffusion model to generate the product final draft image; the redrawing amplitude is 0.3.
[0066] In this embodiment, preferably, step 1 is specifically as follows: Determine the pixels of the product image. If the pixels of the product image are less than 2000×2000, directly perform a matting operation through the Visual Intelligence Open Platform to obtain the required product main image. If the pixels of the product image are greater than or equal to 2000×2000, use Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to perform proportional scaling on the product image to obtain a scaled image. Perform a matting operation on the scaled image through the Visual Intelligence Open Platform to obtain a first mask image. Use Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the first mask image to the original size to obtain a second mask image. Read the alpha channel of each pixel point in the product image to obtain a first matrix, read the alpha channel of each pixel point in the second mask image to obtain a second matrix, and call Core.min to merge the first matrix and the second matrix to obtain a third matrix. Replace the alpha channel in the product image with the third matrix to obtain the required product main image;
[0067] Due to the matting size limit of the Alibaba Cloud Visual Intelligence Open Platform, when the longest side exceeds 2000 pixels, proportional scaling is required. Call Imgproc.resize to use the Imgproc.INTER_LANCZOS4 algorithm for image proportional scaling;
[0068] Call Imgproc.resize to use the Imgproc.INTER_LANCZOS4 algorithm for image proportional scaling. This algorithm can reduce the impact of artifacts while maintaining edge sharpness.
[0069] Obtain the matting result based on the black and white image + original image:
[0070] a. Take out the alpha channel of the original image. Call Core.min, which will take the minimum value of the transparency channels of the mask and the original image, ensuring that the original information of the transparency channel is retained. Through this step of processing, the lines of the obtained matting can be made more perfect;
[0071] b. Remove the alpha channel of the original image and add the black and white image as the new alpha channel for layer merging;
[0072] c. Set the color value of the transparent area to black. This step is to reduce the size of the matting result image and save storage costs.
[0073] The core code is as follows:
[0074] / / Create a fully transparent matrix with the same size as the original image for comparison;
[0075] Mat compareAlpha = new Mat(outImg.size(), CvType.CV_8UC1, Scalar.all(0.0));
[0076] / / Used to save the comparison result;
[0077] Mat compareResult = new Mat();
[0078] / / Compare the alpha channel of the original image with alpha. The positions with a value of 0 in the obtained mask are the transparent positions in the original image;
[0079] Core.compare(outPlanes.get(3), compareAlpha, compareResult, Core.CMP_EQ);
[0080] / / Create a completely black matrix with the same size as the original image;
[0081] Mat black = new Mat(outImg.size(), outImg.type(), Scalar.all(0));
[0082] / / Copy the black color in black to the corresponding positions in outImg according to the mask;
[0083] Core.bitwise_and(black, outImg, outImg, compareResult);
[0084] Generate the corresponding commodity mask grayscale image according to the main commodity image.
[0085] In this embodiment, preferably, step 3 is specifically as follows: Input the main body white background image into the large language model to obtain the main body description words; Input the set scene prompt words into the large language model, and then input the user-defined scene and the main body description words into the large language model to obtain the background description words; Input the set position prompt words into the large language model, and then input the background description words and the main body description words into the large language model to obtain the position description words; Input the set lighting perspective prompt words into the large language model, and then input the main body white background image into the large language model to obtain the lighting perspective description words; Input the set background prompt words into the large language model, and then input the main body description words, background description words, position description words, and lighting perspective description words into the large language model to obtain the background description prompt words.
[0086] In this embodiment, preferably, step 5 is specifically as follows: process the main body white background image to obtain a scaled main body image with a long side resolution of 512 pixels. Input the line drawing image, depth information image, background description prompt, scaled main body image, and commodity mask grayscale image into the diffusion model to generate the required commodity generation image; set the control intensity of Canny in the diffusion model to 0.35, the control intensity of the depth map to 0.65, and the redrawing amplitude to 1.
[0087] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.
[0088] Embodiment 2
[0089] As Figure 2 shown, in this embodiment, a commodity image generation apparatus for eliminating scale failure is provided, including:
[0090] Subject extraction module: perform matte extraction on the commodity image to obtain the commodity main body image and generate the corresponding commodity mask grayscale image;
[0091] White background processing module: process the commodity main body image to obtain a main body white background image with a white background;
[0092] Prompt word generation module: input the main body white background image into the large language model to obtain the main body description word; input the main body description word and the background description word into the large language model to generate the required background description prompt;
[0093] Guidance information module: obtain the line drawing image of the main body white background image through the Canny algorithm; obtain the depth information image of the main body white background image through the depth estimation model;
[0094] Image generation module: process the main body white background image to obtain a scaled main body image with a long side resolution of 512 pixels. Input the line drawing image, depth information image, background description prompt, scaled main body image, and commodity mask grayscale image into the diffusion model to generate the required commodity generation image;
[0095] Scale restoration module: restore the commodity generation image to the Figure 1 same size as the commodity to obtain a scaled commodity generation image. Then, input the set high-definition magnification prompt, scaled commodity generation image, and commodity mask grayscale image into the diffusion model to generate the final commodity image; the redrawing amplitude is 0.3.
[0096] In this embodiment, preferably, the main body extraction module is specifically as follows: judge the pixels of the product image. If the pixels of the product image are less than 2000×2000, directly perform the image extraction operation through the Visual Intelligence Open Platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, use Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image proportionally to obtain a scaled image; perform the image extraction operation on the scaled image through the Visual Intelligence Open Platform to obtain the first mask image; use Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the first mask image to the second mask image of the original size; read the alpha channel of each pixel point in the product image to obtain the first matrix, read the alpha channel of each pixel point in the second mask image to obtain the second matrix, and call Core.min to merge the first matrix and the second matrix to obtain the third matrix; replace the alpha channel in the product image with the third matrix to obtain the required product main body image;
[0097] Generate a corresponding product mask grayscale image according to the product main body image.
[0098] In this embodiment, preferably, the prompt generation module is specifically as follows: input the main body white background image into the large language model to obtain the main body description word; input the set scene prompt word into the large language model, and then input the user-defined scene and the main body description word into the large language model to obtain the background description word; input the set position prompt word into the large language model, and then input the background description word and the main body description word into the large language model to obtain the position description word; input the set lighting perspective prompt word into the large language model, and then input the main body white background image into the large language model to obtain the lighting perspective description word; input the set background prompt word into the large language model, and then input the main body description word, background description word, position description word, and lighting perspective description word into the large language model to obtain the background description prompt word.
[0099] In this embodiment, preferably, the image generation module is specifically as follows: process the main body white background image to obtain a main body scaled image with a long side resolution of 512 pixels, and input the line drawing image, depth information image, background description prompt word, main body scaled image, and product mask grayscale image into the diffusion model to generate the required product generated image; set the control intensity of Canny in the diffusion model to 0.35, the control intensity of the depth map to 0.65, and the redrawing amplitude to 1.
[0100] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated herein. Any device adopted for the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0101] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as detailed in the third embodiment.
[0102] Embodiment Three
[0103] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.
[0104] Since the electronic device introduced in this embodiment is the device adopted for implementing the method in the first embodiment of this application, based on the method introduced in the first embodiment of this application, those skilled in the art can understand the specific implementation manner and various variations of the electronic device in this embodiment. Therefore, the details of how this electronic device implements the method in the embodiments of this application will not be introduced in detail here. Any device adopted by those skilled in the art for implementing the method in the embodiments of this application falls within the scope of protection of this application.
[0105] Based on the same inventive concept, this application provides a storage medium corresponding to the first embodiment, as detailed in the fourth embodiment.
[0106] Embodiment Four
[0107] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in the first embodiment can be realized.
[0108] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0109] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0112] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for generating a product image to eliminate ratio failure, characterized in that: The steps include: Step 1: Cut out the product image, obtain the main product image, and generate the corresponding product mask grayscale image; Step 2: Process the main image of the product to obtain a main image with a white background; Step 3: Input the subject white background image into the large language model to obtain subject description words; input the subject description words and background description words into the large language model to generate the required background description prompt words; Step 4: Obtain a line drawing of the main white background image by using a Canny algorithm; Acquire a depth information map of the subject white background map through a depth estimation model; Step 5: Process the main white background image to obtain a main scaled image with a long side resolution of 512 pixels, input the line drawing image, depth information image, background description prompt words, main scaled image and product mask grayscale image into the diffusion model to generate the required product generation image; Step 6: restore the product generation image to the same size as the product image to obtain the product generation zoom image, and then input the set high-definition zoom prompt words, the product generation zoom image and the product mask grayscale image into the diffusion model to generate the product final draft image.
2. A method for generating a product image with ratio failure eliminated according to claim 1, characterized in that: The step 1 is specifically as follows: judging the pixels of the product image, if the pixels of the product image are less than 2000×2000, directly performing a cutout operation through the visual intelligence open platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, using Imgproc.resize in OpenCV and using the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image in proportion to obtain a scaled image; performing a cutout operation on the scaled image through the visual intelligence open platform to obtain a first mask image; using Imgproc.resize in OpenCV and using the Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the first mask image to a second mask image of the original size; reading out the transparent channel of each pixel in the product image to obtain a first matrix, reading out the transparent channel of each pixel in the second mask image to obtain a second matrix, calling Core.min to merge the first matrix with the second matrix to obtain a third matrix; The third matrix is used to replace the transparent channel in the product image to obtain the required product main image; Generate a corresponding product mask grayscale image based on the product main image.
3. A method for generating a product image with ratio failure eliminated according to claim 1, characterized in that: The step 3 is specifically as follows: inputting the subject white background image into the large language model to obtain subject description words; inputting the set scene prompt words into the large language model, and then inputting the user-defined scene and the subject description words into the large language model to obtain background description words; inputting the set location prompt words into the large language model, and then inputting the background description words and the subject description words into the large language model to obtain location description words; inputting the set lighting perspective prompt words into the large language model, and then inputting the subject white background image into the large language model to obtain lighting perspective description words; inputting the set background prompt words into the large language model, and then inputting the subject description words, background description words, location description words and lighting perspective description words into the large language model to obtain background description prompt words.
4. The method for generating a product image with eliminating ratio failure according to claim 1, characterized in that: The step 5 is specifically as follows: processing the main white background image to obtain a main scaled image with a long side resolution of 512 pixels, inputting the line drawing image, depth information image, background description prompt words, main scaled image and product mask grayscale image into the diffusion model to generate the required product generation image; setting the control strength of Canny in the diffusion model to 0.35, the control strength of the depth map to 0.65, and the redrawing amplitude to 1.
5. A commodity image generation device for eliminating ratio failure, characterized in that: include: The main body cutout module cuts out the product image, obtains the main body image of the product, and generates the corresponding product mask grayscale image; A white background processing module processes the main image of the product to obtain a main white background image with a white background; The prompt word generation module inputs the subject white background image into the large language model to obtain the subject description words; the subject description words and background description words are input into the large language model to generate the required background description prompt words; A guidance information module, which obtains a line drawing of the main white background image through a Canny algorithm; Acquire a depth information map of the subject white background map through a depth estimation model; The image generation module processes the main white background image to obtain a main scaled image with a long side resolution of 512 pixels, and inputs the line drawing image, depth information image, background description prompt words, main scaled image and product mask grayscale image into the diffusion model to generate the required product generation image; The scale restoration module restores the product generation image to the same size as the product image to obtain the product generation zoom image, and then inputs the set high-definition zoom prompt words, the product generation zoom image and the product mask grayscale image into the diffusion model to generate the product final draft image.
6. The device for generating a product image with ratio invalidation eliminated according to claim 5, characterized in that: The cutout main body module is specifically as follows: the pixels of the product image are judged. If the pixels of the product image are less than 2000×2000, the cutout operation is directly performed through the visual intelligence open platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, the product image is proportionally scaled using the Imgproc.resize in OpenCV and the Imgproc.INTER_LANCZ0S4 algorithm to obtain a scaled image; the scaled image is cutout through the visual intelligence open platform to obtain a first mask image; the first mask image is proportionally restored to a second mask image of the original size through the Imgproc.resize in OpenCV and the Imgproc.INTER_LANCZ0S4 algorithm; the transparent channel of each pixel in the product image is read out to obtain a first matrix, the transparent channel of each pixel in the second mask image is read out to obtain a second matrix, and Core.min is called to merge the first matrix with the second matrix to obtain a third matrix; The third matrix is used to replace the transparent channel in the product image to obtain the required product main image; Generate a corresponding product mask grayscale image based on the product main image.
7. The device for generating a product image with ratio invalidation eliminated according to claim 5, characterized in that: The prompt word generation module is specifically as follows: inputting the main body white background image into the large language model to obtain the main body description words; inputting the set scene prompt words into the large language model, and then inputting the user-defined scene and the main body description words into the large language model to obtain the background description words; inputting the set location prompt words into the large language model, and then inputting the background description words and the main body description words into the large language model to obtain the location description words; inputting the set lighting perspective prompt words into the large language model, and then inputting the main body white background image into the large language model to obtain the lighting perspective description words; inputting the set background prompt words into the large language model, and then inputting the main body description words, background description words, location description words and lighting perspective description words into the large language model to obtain the background description prompt words.
8. The device for generating a product image with ratio invalidation eliminated according to claim 5, characterized in that: The image generation module specifically includes: processing the main white background image to obtain a main scaled image with a long side resolution of 512 pixels, inputting the line drawing image, depth information image, background description prompt words, main scaled image and product mask grayscale image into the diffusion model to generate the required product generation image; setting the control strength of Canny in the diffusion model to 0.35, the control strength of the depth map to 0.65, and the redrawing amplitude to 1.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.