Method, device and equipment for generating commodity graph through limited style background and medium
The product plan is generated through cutout and algorithm processing, and combined with style characteristics, and the result diagram is generated using the diffusion model, which solves the problems of uncontrollable background and inefficiency in the existing technology, and achieves product map generation with high stability and close to reality.
Patent Information
- Application Number
- CN202510108236.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
When generating product maps with set backgrounds, the background is uncontrollable, and it takes multiple generations to generate similar maps, and there are problems such as hanging, resulting in inefficiency.
The main image of the product is obtained by cutting the image, processed into a gray base image, and processed through the Depth and Lineart algorithms, and passed into Controlnet for drawing to generate the product floor plan. Then, the set style reference diagram is input to the ipadapter model, the style features are extracted, combined with the product plan, and drawn through the diffusion model to generate the result diagram.
It realizes the rapid generation of background composite images with a unified style, improves the stability of generated images, avoids the problems of hanging and background mismatch, makes the generated images closer to actual life, and improves user usage rate.
Smart Images

Figure CN120070623A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of generating pictures, and particularly relates to a method, device, equipment and medium for generating product pictures with a limited style background. Background Art
[0002] In the fields of e-commerce and advertising, the style consistency of product background pictures is crucial for brand image and user experience. However, the existing technology is to select the same background picture or a series of background pictures of the same style through art processing, and then combine the product with a specific background, which results in very low efficiency. When a large number of pictures need to be processed, it is necessary to recruit a large number of artists or outsource.
[0003] In the prior art, generating product pictures with a set background is achieved by using StableDiffusion, which requires technicians to write specific prompts. The disadvantage of this method is that the background of the generated pictures is uncontrollable, and it is necessary to generate multiple times to possibly obtain similar generated pictures; and there are problems such as suspension in the generated pictures. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for generating product pictures with a limited style background, which can adapt to the perspective of the product to generate pictures, making the generated pictures closer to reality and facilitating direct use by users.
[0005] In a first aspect, the present invention provides a method for generating product pictures with a limited style background, including the following steps:
[0006] Step 1: Cut out the product picture to obtain the main product picture.
[0007] Step 2: Process the main product picture to obtain a first gray-bottom picture with a gray background.
[0008] Step 3: Process the first gray-bottom picture through the Depth algorithm to obtain a first depth image; process the first gray-bottom picture through the Lineart algorithm to obtain a first line drawing.
[0009] Step 4: Input the first depth image and the first line drawing into ControlNet, set the redrawing amplitude to 0.85, and set the reference weights of the lineart module and the depth module in ControlNet to 0.5 and 0.6 respectively. Then, draw through the diffusion model to generate a product plan view.
[0010] Step 5: Input the set style reference image into the ipadapter model to extract the style features, and input the style features and the product floor plan into the diffusion model for drawing to generate the result image.
[0011] In a second aspect, the present invention provides an apparatus for generating a product image with a defined style background, including:
[0012] An acquisition main body module that performs matte extraction on the product image to obtain the product main body image;
[0013] A background setting module that processes the product main body image to obtain a first gray background image with a gray background;
[0014] An acquisition guidance module that processes the first gray background image through the Depth algorithm to obtain a first depth image; the first gray background image is processed through the Lineart algorithm to obtain a first line drawing;
[0015] A generation guidance module that inputs the first depth image and the first line drawing into Controlnet, sets the redrawing amplitude to 0.85, and sets the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively, and then performs drawing through the diffusion model to generate the product floor plan;
[0016] A generation result module that inputs the set style reference image into the ipadapter model to extract the style features, and inputs the style features and the product floor plan into the diffusion model for drawing to generate the result image.
[0017] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in the first aspect.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.
[0019] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0020] Regardless of different products or different angles of the same product, the present invention can quickly generate a background composite image with a unified style, which is greatly convenient for users to use;
[0021] And the generation method of the present invention greatly improves the stability of the generated image, so that there is no problem of suspension or mismatch between the product and the background, making the generated image closer to real life and further improving the utilization rate of the images generated by users.
[0022] The optimized method of the present invention can effectively solve the problem of insufficient light and shadow that easily occurs in traditional light and shadow synthesis algorithms, and this method can greatly improve work efficiency.
[0023] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings
[0024] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.
[0025] Figure 1 It is the flowchart of the method in Embodiment 1 of the present invention;
[0026] Figure 2 It is the structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments
[0027] Embodiment 1
[0028] As Figure 1 shown, this embodiment provides a method for generating a product picture with a defined style background, including the following steps:
[0029] Step 1, perform matting on the product picture to obtain the product main body picture;
[0030] Step 2, process the product main body picture to obtain the first gray-bottom picture with a gray background;
[0031] Step 3, process the first gray-bottom picture through the Depth algorithm to obtain the first depth image; the first gray-bottom picture is processed through the Lineart algorithm to obtain the first line drawing;
[0032] Step 4, input the first depth image and the first line drawing into Controlnet, set the redrawing amplitude to 0.85, and set the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively, and then draw through the diffusion model to generate a product plan view; this diffusion model is the Stable Diffusion model;
[0033] Step 5, input the set style reference picture into the ipadapter model to extract the style features, and input the style features and the product plan view into the diffusion model for drawing to generate the result picture.
[0034] In this embodiment, preferably, it further includes step 6: processing the result image to obtain a background-removed image with only the commodity and a background image with only the background; enhancing the background-removed image through Gamma transformation, increasing the brightness of the commodity main body in the background-removed image by a first set value to obtain an enhanced image; mapping the background image to the HSV space, reducing the V space in the HSV space by a second set value, and then mapping it back to the RGB space to obtain a background-darkened image; synthesizing the enhanced image and the background-darkened image to obtain a first synthesized image, and extracting first latent information from the first synthesized image through VAE decoding; synthesizing the background-removed image and the background image to obtain a second synthesized image, and extracting first set information from the second synthesized image by using the ControlNet model; sending the first latent information and the first set information to the diffusion model to generate a first light and shadow guidance image; extracting second latent information from the first light and shadow guidance image, and sending the second latent information and the first set information to the diffusion model to generate a second light and shadow guidance image; migrating the texture information of the first light and shadow guidance image to the second light and shadow guidance image through a high-contrast retention algorithm to obtain an intermediate image, and then migrating the color information of the background image to the intermediate image by using the welsh algorithm to obtain a final optimized image.
[0035] In this embodiment, preferably, step 2 is specifically: synthesizing the area other than the commodity main body part in the commodity main body image with a grayscale image with a grayscale value of 127 to obtain a commodity grayscale-bottom image; processing the commodity grayscale-bottom image to obtain a first grayscale-bottom image with a resolution of 640*640.
[0036] In this embodiment, preferably, step 5 is specifically: processing the generated commodity plan view through the depth algorithm to obtain a second depth image; processing the commodity grayscale-bottom image through the lineart algorithm to obtain a second line drawing; sending the second depth image and the second line drawing to Controlnet respectively, setting the redrawing amplitude to 1.0, setting the reference weights of the lineart module and the depth module in Controlnet to 0.45 and 0.55 respectively, and then performing drawing through the diffusion model to generate a guidance image; inputting the set style reference image into the ipadapter model to extract style features, and inputting the style features and the guidance image into the diffusion model for drawing to generate a result image.
[0037] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.
[0038] Embodiment 2
[0039] As Figure 2 shown, in this embodiment, an apparatus for generating a commodity image with a defined style background is provided, including:
[0040] Obtain the main body module, perform matte painting on the product image to obtain the product main body image;
[0041] Set the background module, process the product main body image to obtain the first gray background image with a gray background;
[0042] Obtain the guidance module, process the first gray background image through the Depth algorithm to obtain the first depth image; the first gray background image is processed through the Lineart algorithm to obtain the first line drawing;
[0043] Generate the guidance module, input the first depth image and the first line drawing into Controlnet, set the redrawing amplitude to 0.85, and set the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively. Then, perform drawing through the diffusion model to generate the product plan view; this diffusion model is the Stable Diffusion model;
[0044] Generate the result module, input the set style reference image into the ipadapter model to extract the style features, and input the style features and the product plan view into the diffusion model for drawing to generate the result image.
[0045] In this embodiment, preferably, it further includes an optimized result module, which processes the result image to obtain a cutout image with only the product and a background image with only the background; enhance the cutout image through Gamma transformation, increase the brightness of the product main body in the cutout image by the first set value to obtain the enhanced image; map the background image to the HSV space, reduce the V space in the HSV space by the second set value, and then map it back to the RGB space to obtain the background darkening image; synthesize the enhanced image and the background darkening image to obtain the first composite image, and extract the first latent information from the first composite image through VAE decoding; synthesize the cutout image and the background image to obtain the second composite image, and extract the first set information from the second composite image using the ControlNet model; send the first latent information and the first set information to the diffusion model to generate the first light and shadow guidance image; extract the second latent information from the first light and shadow guidance image, and send the second latent information and the first set information to the diffusion model to generate the second light and shadow guidance image; transfer the texture information of the first light and shadow guidance image to the second light and shadow guidance image through the high-contrast retention algorithm to obtain the intermediate image, and then transfer the color information of the background image to the intermediate image using the welsh algorithm to obtain the final optimized image.
[0046] In this embodiment, preferably, the background setting module is specifically as follows: The area of the product main body in the product main body image is combined with a grayscale image with a grayscale value of 127 to obtain a product grayscale bottom image; the product grayscale bottom image is processed to obtain a first grayscale bottom image with a resolution of 640*640.
[0047] In this embodiment, preferably, the result generation module is specifically as follows: The generated product plan view is processed by the depth algorithm to obtain a second depth image; the product grayscale bottom image is processed by the lineart algorithm to obtain a second line drawing; the second depth image and the second line drawing are respectively sent to Controlnet, the redrawing amplitude is set to 1.0, and the reference weights of the lineart module and the depth module in Controlnet are set to 0.45 and 0.55 respectively. Then, drawing is performed through the diffusion model to generate a guidance image; the set style reference image is input into the ipadapter model to extract style features, and the style features and the guidance image are input into the diffusion model for drawing to generate a result image.
[0048] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device adopted for the method of the first embodiment of the present invention belongs to the scope protected by the present invention.
[0049] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as detailed in the third embodiment.
[0050] Embodiment Three
[0051] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.
[0052] Since the electronic device introduced in this embodiment is the device adopted for implementing the method in the first embodiment of this application, based on the method introduced in the first embodiment of this application, those skilled in the art can understand the specific implementation manner and various variations of the electronic device in this embodiment. Therefore, how the electronic device implements the method in the embodiments of this application will not be introduced in detail here. Any device adopted by those skilled in the art for implementing the method in the embodiments of this application belongs to the scope protected by this application.
[0053] Based on the same inventive concept, this application provides a storage medium corresponding to the first embodiment, as detailed in the fourth embodiment.
[0054] Embodiment Four
[0055] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in the first embodiment can be realized.
[0056] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0057] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0058] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for realizing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0060] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for generating a product image with a limited style background, characterized in that: The steps include: Step 1: Cut out the product image to obtain the main image of the product; Step 2: Process the main image of the product to obtain a first gray background image with a gray background; Step 3: Process the first gray background image by using a Depth algorithm to obtain a first depth image; process the first gray background image by using a Lineart algorithm to obtain a first line drawing; Step 4: The first depth image and the first line drawing are passed into Controlnet, the redrawing amplitude is set to 0.85, and the reference weights of the lineart module and the depth module in Controlnet are set to 0.5 and 0.6 respectively. Then, the diffusion model is used for drawing to generate a product plan view. Step 5: Input the set style reference image into the iPad Adapter model to extract the style features, input the style features and the product plan view into the diffusion model for drawing, and generate a result image.
2. The method for generating a product image with a limited style background according to claim 1, characterized in that: The method also includes step 6, processing the result image to obtain a background-removed image of only the product and a background image of only the background; enhancing the background-removed image by Gamma transformation, increasing the brightness of the main body of the product in the background-removed image by a first set value, and obtaining an enhanced image; mapping the background image to the HSV space, reducing the V space in the HSV space by a second set value, and then mapping it back to the RGB space to obtain a background darkened image; synthesizing the enhanced image and the background darkened image to obtain a first synthesized image, and extracting first latent information from the first synthesized image by VAE decoding; The background-removed image and the background image are synthesized to obtain a second synthesized image, and the first setting information is extracted from the second synthesized image using the ControlNet model; the first latent information and the first setting information are sent to the diffusion model to generate a first light and shadow guidance map; the second latent information is extracted from the first light and shadow guidance map, and the second latent information and the first setting information are sent to the diffusion model to generate a second light and shadow guidance map; the texture information of the first light and shadow guidance map is migrated to the second light and shadow guidance map through a high-contrast retention algorithm to obtain an intermediate image, and then the Welsh algorithm is used to migrate the color information of the background image to the intermediate image to obtain the final optimized image.
3. The method for generating a product image with a limited style background according to claim 1, characterized in that: The step 2 specifically includes: synthesizing the area of the product main body image except the product main body with a grayscale image with a grayscale value of 127 to obtain a product gray background image; processing the product gray background image to obtain a first gray background image with a resolution of 640*640.
4. The method for generating a product image with a limited style background according to claim 1, characterized in that: The step 5 is specifically as follows: the generated product plan view is processed by the depth algorithm to obtain a second depth image; the product gray background image is processed by the lineart algorithm to obtain a second line drawing; the second depth image and the second line drawing are sent to Controlnet respectively, the redrawing amplitude is set to 1.0, and the reference weights of the lineart module and the depth module in Controlnet are set to 0.45 and 0.55 respectively, and then drawn by the diffusion model to generate a guidance map; the set style reference map is input into the iPad adapter model to extract the style features, and the style features and the guidance map are input into the diffusion model for drawing to generate a result map.
5. A device for generating a product image with a limited style background, characterized in that: include: Get the main module, cut out the product image, and get the main product image; Setting a background module to process the main image of the product to obtain a first gray background image with a gray background; The acquisition guidance module processes the first gray background image by using a Depth algorithm to obtain a first depth image; the first gray background image is processed by using a Lineart algorithm to obtain a first line drawing; Generate a guidance module, pass the first depth image and the first line drawing to Controlnet, set the redrawing amplitude to 0.85, and set the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively, and then draw through the diffusion model to generate a product plan view; The result generation module inputs the set style reference image into the iPad Adapter model to extract the style features, inputs the style features and the product plan into the diffusion model for drawing, and generates a result image.
6. The device for generating a product image with a limited style background according to claim 5, characterized in that: The method also includes an optimization result module, which processes the result image to obtain a background-removed image of only the product and a background image of only the background; enhances the background-removed image by Gamma transformation, increases the brightness of the main body of the product in the background-removed image by a first set value, and obtains an enhanced image; maps the background image to the HSV space, reduces the V space in the HSV space by a second set value, and then maps it back to the RGB space to obtain a background darkened image; synthesizes the enhanced image and the background darkened image to obtain a first synthesized image, and extracts first latent information from the first synthesized image by VAE decoding; The background-removed image and the background image are synthesized to obtain a second synthesized image, and the first setting information is extracted from the second synthesized image using the ControlNet model; the first latent information and the first setting information are sent to the diffusion model to generate a first light and shadow guidance map; the second latent information is extracted from the first light and shadow guidance map, and the second latent information and the first setting information are sent to the diffusion model to generate a second light and shadow guidance map; the texture information of the first light and shadow guidance map is migrated to the second light and shadow guidance map through a high-contrast retention algorithm to obtain an intermediate image, and then the Welsh algorithm is used to migrate the color information of the background image to the intermediate image to obtain the final optimized image.
7. The device for generating a product image with a limited style background according to claim 5, characterized in that: The background setting module specifically comprises: synthesizing the area of the product main body image except the product main body with a grayscale image with a grayscale value of 127 to obtain a product gray background image; processing the product gray background image to obtain a first gray background image with a resolution of 640*640.
8. The device for generating a product image with a limited style background according to claim 5, characterized in that: The generation result module is specifically as follows: the generated product plan map is processed by the depth algorithm to obtain a second depth image; the product gray background map is processed by the lineart algorithm to obtain a second line drawing; the second depth image and the second line drawing are sent to Controlnet respectively, the redrawing amplitude is set to 1.0, and the reference weights of the lineart module and the depth module in Controlnet are set to 0.45 and 0.55 respectively, and then drawn by the diffusion model to generate a guidance map; the set style reference map is input into the iPad adapter model to extract the style features, and the style features and the guidance map are input into the diffusion model for drawing to generate a result map.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.