Method, device and equipment for replacing background of commodity and medium

Through the cutout algorithm and large language model to generate scene descriptors, combined with the diffusion model to generate result diagrams, the problems of low efficiency and high cost of product replacement in the existing technology are solved, efficient and personalized product and background synthesis are achieved, and picture quality and display effect are improved.

CN120070630APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108350.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is inefficient and costly when replacing the background of the product, especially when processing high-definition pictures, the pictures need to be compressed to meet the platform requirements, resulting in a decline in picture quality.

Method used

The cutout algorithm is used to obtain the main image of the product, and the scene descriptor is generated through a large language model, and the result image is generated in combination with the diffusion model to achieve rapid synthesis of products and backgrounds.

Benefits of technology

It improves the efficiency of image generation and meets users' personalized needs. The generated images are naturally integrated with high pixels, enhances product display effects, and stimulates consumers' desire to buy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070630A_ABST
    Figure CN120070630A_ABST
Patent Text Reader

Abstract

The invention provides a method, device and equipment for replacing the background of a commodity and a medium, and the method comprises the steps: obtaining a commodity main body image through employing an image matting algorithm according to a commodity image uploaded by a user; setting an application scene, inputting the application scene into a large language model through a set prompt word format according to the commodity main body graph and the reference scene, and generating scene description words by the large language model; and inputting the commodity main body and the scene description word into the diffusion model to generate a result graph, thereby realizing rapid commodity background replacement, greatly improving efficiency, and reducing enterprise cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, equipment and medium for replacing the background of a commodity. Background Art

[0002] When replacing the background of a commodity, it is impossible to take a posed photo in some cases; for the venue requirements, merchants cannot meet them; if there is no product with a background, it will be difficult to sell; for example, for a sofa, it needs to be displayed in different living rooms so that users can make comparisons. For manufacturers, if they build exhibition halls for each living room, it will greatly increase the cost for users.

[0003] Moreover, in today's e-commerce field, when online sellers promote products, they often need to go through a series of complex steps to ensure the online display effect of the commodity. First of all, they must hire professional photographers to take pictures of the products. This process is not only time-consuming but also costly. Photographers need to have professional skills and equipment to capture the best angles and details of the products to ensure that the pictures can attract the attention of consumers.

[0004] Some existing enterprises also perform matte painting by artists, then search for background images through the Internet, and then combine the matte painting with the background images. However, the operation of artists requires a lot of time and high professional requirements for artists, which leads to an increase in the cost of enterprises; in the existing technology, the Alibaba Cloud Visual Intelligence Open Platform provides a segmentation and matte painting function, and this service does have certain limitations when processing images. Specifically, the platform requires that the size of the input image must be less than 2000×2000 pixels, that is, the longest side does not exceed 1999 pixels. This limitation has caused many users to be unable to directly use this function to process the high-definition images in their hands because most of the existing image resolutions exceed this limit; users have to take some compromise measures, such as first compressing the image to a size that meets the platform requirements, and then using the segmentation and matte painting function of the platform for processing. However, this approach will inevitably sacrifice the quality of the image. Especially, some details and clarity may be lost during the compression process. After that, when combining the matte painting with the searched background image, it is still done by artists, with low efficiency and reduced image quality. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for replacing the background of a commodity, which can realize fast replacement of the background of a commodity, greatly improve the efficiency, and reduce the cost of enterprises.

[0006] In a first aspect, the present invention provides a method for replacing the background of a commodity, including the following steps:

[0007] Step 1: Based on the product image uploaded by the user, a cutout algorithm is used to obtain the main image of the product;

[0008] Step 2: Set the application scenario, input the set prompt word format into the large language model according to the main picture of the product and the reference scenario, and the large language model generates the scenario description word;

[0009] Step 3: Input the product entity and scene description words into the diffusion model to generate a result graph.

[0010] In a second aspect, the present invention provides a device for changing the background of a commodity, comprising:

[0011] The cutout module uses a cutout algorithm to obtain the main image of the product based on the product image uploaded by the user;

[0012] Generate a description word module, set an application scenario, and input the set prompt word format into the large language model according to the product main image and the reference scenario, and the large language model generates a scenario description word;

[0013] The graph generation module inputs the product body and scene description words into the diffusion model to generate the result graph.

[0014] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0016] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0017] Through the method of the present invention, different backgrounds can be provided for different commodities, and combined with the diffusion model, a combination map of commodities and backgrounds can be quickly generated. This method can greatly improve the efficiency of generating required pictures and meet the needs of users. When merchants need to promote products, they can quickly generate pictures with different backgrounds, and then upload the pictures to their online shops for users to view.

[0018] Moreover, the picture has high pixels, and the product and the background picture are naturally integrated together, which makes it easier to display the product effect and more easily trigger buyers to place orders.

[0019] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the content of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Description of the Drawings

[0020] The present invention will be further described below with reference to the accompanying drawings in conjunction with the embodiments.

[0021] Figure 1 It is a flowchart of the method in the first embodiment of the present invention;

[0022] Figure 2 It is a structural schematic diagram of the device in the second embodiment of the present invention. Specific Embodiments

[0023] In the embodiments of the present application, by providing a method, device, equipment and medium for replacing the background of commodities, customized backgrounds are tailored for various commodities, and the diffusion model technology is skillfully integrated to realize the rapid synthesis of commodity and background images. The present invention greatly improves the efficiency of image generation and accurately meets the personalized needs of users. When merchants promote products, they can quickly produce pictures with various background styles and upload these high-definition and high-pixel pictures to the online store for consumers to browse. The natural fusion of commodities and backgrounds not only enhances the display effect of commodities, but also effectively stimulates consumers' desire to purchase, thus promoting sales conversion.

[0024] Embodiment 1

[0025] As Figure 1 shown, this embodiment provides a method for replacing the background of commodities, including the following steps:

[0026] Step 1: According to the commodity picture uploaded by the user, use a matting algorithm to obtain the main body picture of the commodity;

[0027] Step 2: Set the application scenario, and input the main body picture of the commodity and the reference scenario into the large language model in the set prompt word format, and the large language model generates a scene description word;

[0028] Step 3: Input the main body of the commodity and the scene description word into the diffusion model to generate a result picture.

[0029] In this embodiment, preferably, step 2 is specifically: set a scene library, the scene library includes multiple background pictures, set corresponding background prompt words for each background picture, select a background picture, and input the main body picture of the commodity and the background prompt words corresponding to the background picture into the large language model in the set prompt word format, and the large language model generates a scene description word.

[0030] In this embodiment, preferably, step 1 is specifically as follows:

[0031] Step 11: Judge the pixels of the product image uploaded by the user. If the pixels of the product image are less than 2000×2000, directly perform a matting operation through the Visual Intelligence Open Platform to obtain the required product main image; if the pixels of the image are greater than or equal to 2000×2000, proceed to step 12;

[0032] Step 12: Scale the image proportionally through a scaling algorithm to obtain a scaled image;

[0033] Step 13: Perform a matting operation on the scaled image through the Visual Intelligence Open Platform to obtain a first mask image;

[0034] Step 14: Restore the first mask image to the second mask image of the original size by using the scaling algorithm;

[0035] Step 15: Combine the second mask image with the image uploaded by the user to obtain the required product main image.

[0036] In this embodiment, preferably, step 12 is specifically as follows: Use the Imgproc.INTER_LANCZ0S4 algorithm in OpenCV's Imgproc.resize to scale the image proportionally to obtain a scaled image;

[0037] Step 14 is specifically as follows: Use the Imgproc.INTER_LANCZ0S4 algorithm in OpenCV's Imgproc.resize to restore the first mask image to the second mask image of the original size proportionally;

[0038] Step 15 is specifically as follows: Read the alpha channel of each pixel point in the image uploaded by the user to obtain a first matrix, read the alpha channel of each pixel point in the second mask image to obtain a second matrix, and call Core.min to combine the first matrix and the second matrix to obtain a third matrix; Replace the alpha channel in the image uploaded by the user with the third matrix to obtain the required image;

[0039] Due to the matting size limit of the Alibaba Cloud Visual Intelligence Open Platform, when the longest side exceeds 2000 pixels, proportional scaling is required, and the Imgproc.resize is called to use the Imgproc.INTER_LANCZOS4 algorithm for image scaling;

[0040] Call Imgproc.resize to use the Imgproc.INTER_LANCZOS4 algorithm for image scaling. This algorithm can reduce the impact of artifacts while maintaining edge sharpness.

[0041] The matte extraction result is obtained based on the black-and-white image + the original image:

[0042] a. Extract the alpha channel of the original image. Call Core.min, which will take the minimum value of the mask and the alpha channel of the original image to ensure that the original information of the alpha channel is retained. Through this step of processing, the lines of the obtained matte can be made more perfect;

[0043] b. Remove the alpha channel of the original image and add the black-and-white image as the new alpha channel, and perform layer merging;

[0044] c. Set the color value of the transparent area to black. This step is to reduce the size of the matte result image and save storage costs.

[0045] The core code is as follows:

[0046] / / Create a fully transparent matrix with the same size as the original image for comparison

[0047] Mat compareAlpha = new Mat(outImg.size(), CvType.CV_8UC1, Scalar.all(0.0));

[0048] / / Used to save the comparison result

[0049] Mat compareResult = new Mat();

[0050] / / Compare the alpha channel of the original image and alpha. The positions with a value of 0 in the obtained mask are the transparent positions in the original image

[0051] Core.compare(outPlanes.get(3), compareAlpha, compareResult, Core.CMP_EQ);

[0052] / / Create a fully black matrix with the same size as the original image

[0053] Mat black = new Mat(outImg.size(), outImg.type(), Scalar.all(0));

[0054] / / Copy the black color in black to the corresponding positions in outImg according to the mask

[0055] Core.bitwise_and(black, outImg, outImg, compareResult);

[0056] In this embodiment, preferably, step 3 is specifically as follows: Process the product main image to obtain a first gray-bottom image with a gray background; Process the first gray-bottom image through the Depth algorithm to obtain a first depth image; Process the first gray-bottom image through the Lineart algorithm to obtain a first line drawing; Input the first depth image and the first line drawing into Controlnet, set the redrawing amplitude to 0.85, and set the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively. Then, perform drawing through the diffusion model to generate a product floor plan; Process the generated product floor plan through the depth algorithm to obtain a second depth image; Process the product gray-bottom image through the lineart algorithm to obtain a second line drawing; Send the second depth image and the second line drawing to Controlnet respectively, set the redrawing amplitude to 1.0, set the reference weights of the lineart module and the depth module in Controlnet to 0.45 and 0.55 respectively. Then, input the scene description words into the diffusion model and guide the drawing through Controlnet to generate a result image.

[0057] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.

[0058] Embodiment 2

[0059] As Figure 2 shown, in this embodiment, an apparatus for replacing the background of a product is provided, including:

[0060] A matte extraction module, which obtains the product main image by using a matte extraction algorithm according to the product image uploaded by the user;

[0061] A description word generation module, which sets the application scenario, and inputs the product main image and the reference scenario into a large language model in a set prompt word format. The large language model generates scene description words;

[0062] A graph generation module, which inputs the product main body and the scene description words into the diffusion model to generate a result image.

[0063] In this embodiment, preferably, the description word generation module is specifically as follows: Set a scene library, which includes multiple background images. Set corresponding background prompt words for each background image. Select a background image, and input the product main image and the background prompt words corresponding to the background image into a large language model in a set prompt word format. The large language model generates scene description words.

[0064] In this embodiment, preferably, the matte extraction module is specifically as follows:

[0065] A judgment unit that judges the pixels of the product image uploaded by the user. If the pixels of the product image are less than 2000×2000, the image is directly cropped through the Visual Intelligence Open Platform to obtain the required product main image. If the pixels of the image are greater than or equal to 2000×2000, go to step 12;

[0066] A scaling unit that scales the image proportionally through a scaling algorithm to obtain a scaled image;

[0067] An operation unit that crops the scaled image through the Visual Intelligence Open Platform to obtain a first mask image;

[0068] A restoration unit that restores the first mask image to the second mask image of the original size by using a scaling algorithm;

[0069] A cropping unit that merges the second mask image with the image uploaded by the user to obtain the required product main image.

[0070] In this embodiment, preferably, the scaling unit is specifically: using Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the image proportionally to obtain a scaled image;

[0071] The restoration unit is specifically: using Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to restore the first mask image to the second mask image of the original size proportionally;

[0072] The cropping unit is specifically: reading out the alpha channel of each pixel point in the image uploaded by the user to obtain a first matrix, reading out the alpha channel of each pixel point in the second mask image to obtain a second matrix, and calling Core.min to merge the first matrix and the second matrix to obtain a third matrix; replacing the alpha channel in the image uploaded by the user with the third matrix to obtain the required image.

[0073] In this embodiment, preferably, the generating diagram module specifically includes: processing the main product diagram to obtain a first gray-background diagram with a gray background; processing the first gray-background diagram through the Depth algorithm to obtain a first depth image; processing the first gray-background diagram through the Lineart algorithm to obtain a first line drawing; inputting the first depth image and the first line drawing into Controlnet, setting the redrawing amplitude to 0.85, and setting the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively, and then performing drawing through the diffusion model to generate a product plan view; processing the generated product plan view through the depth algorithm to obtain a second depth image; processing the product gray-background diagram through the lineart algorithm to obtain a second line drawing; sending the second depth image and the second line drawing into Controlnet respectively, setting the redrawing amplitude to 1.0, setting the reference weights of the lineart module and the depth module in Controlnet to 0.45 and 0.55 respectively, and then inputting the scene description words into the diffusion model to be guided by Controlnet for drawing to generate a result diagram.

[0074] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device adopted for the method in the first embodiment of the present invention belongs to the scope of protection of the present invention.

[0075] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as detailed in the third embodiment.

[0076] Embodiment Three

[0077] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.

[0078] Since the electronic device introduced in this embodiment is the device adopted for implementing the method in the first embodiment of this application, based on the method introduced in the first embodiment of this application, those skilled in the art can understand the specific implementation manner and various variations of the electronic device in this embodiment, so the specific implementation of how this electronic device realizes the method in the embodiments of this application will not be introduced in detail here. Any device adopted by those skilled in the art for implementing the method in the embodiments of this application belongs to the scope of protection of this application.

[0079] Based on the same inventive concept, this application provides a storage medium corresponding to the first embodiment, as detailed in the fourth embodiment.

[0080] Example 4

[0081] This embodiment provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be realized.

[0082] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0083] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be realized by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0084] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0086] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.

Claims

1. A method for changing the background of a commodity, characterized in that: The steps include: Step 1: Based on the product image uploaded by the user, a cutout algorithm is used to obtain the main image of the product; Step 2: Set the application scenario, input the set prompt word format into the large language model according to the main picture of the product and the reference scenario, and the large language model generates the scenario description word; Step 3: Input the product entity and scene description words into the diffusion model to generate a result graph.

2. A method for changing the background of a commodity according to claim 1, characterized in that: The step 2 specifically includes: setting a scene library, the scene library includes multiple background pictures, setting corresponding background prompt words for each background picture, selecting a background picture, inputting the product main body picture and the background prompt words corresponding to the background picture into the large language model through the set prompt word format, and the large language model generates scene description words.

3. A method for changing the background of a commodity according to claim 1, characterized in that: The step 1 is specifically as follows: Step 11: The pixels of the product image uploaded by the user are judged. If the pixels of the product image are less than 2000×2000, the image is cut out directly through the visual intelligence open platform to obtain the required product main image; if the pixels of the image are greater than or equal to 2000×2000, proceed to step 12; Step 12, use Imgproc.resize in OpenCV and use Imgproc.INTER_LANC Z0S4 algorithm to scale the image in equal proportion to obtain a scaled image; Step 13: Perform a cutout operation on the zoomed image through the visual intelligence open platform to obtain a first mask image; Step 14: Use Imgproc.resize in OpenCV and Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the first mask image to the second mask image of the original size; Step 15: read out the transparent channel of each pixel in the picture uploaded by the user to obtain a first matrix, read out the transparent channel of each pixel in the second mask image to obtain a second matrix, call Core.min to merge the first matrix with the second matrix to obtain a third matrix; The transparent channel in the picture uploaded by the user is replaced by the third matrix to obtain the desired image.

4. A method for changing the background of a commodity according to claim 3, characterized in that: The step 3 is specifically as follows: processing the main image of the product to obtain a first gray background image with a gray background; processing the first gray background image through the Depth algorithm to obtain a first depth image; processing the first gray background image through the Lineart algorithm to obtain a first line drawing image; passing the first depth image and the first line drawing image into Controlnet, setting the redrawing amplitude to 0.85, and setting the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively, and then drawing through the diffusion model to generate a product plan view; processing the generated product plan view through the depth algorithm to obtain a second depth image; processing the gray background image of the product through the lineart algorithm to obtain a second line drawing image; sending the second depth image and the second line drawing image to Controlnet respectively, setting the redrawing amplitude to 1.0, and setting the reference weights of the lineart module and the depth module in Controlnet to 0.45 and 0.55 respectively, and then inputting the scene description words into the diffusion model to guide the drawing through Controlnet to generate a result image.

5. A device for changing the background of a commodity, characterized in that: include: The cutout module uses a cutout algorithm to obtain the main image of the product based on the product image uploaded by the user; Generate a description word module, set an application scenario, and input the set prompt word format into the large language model according to the product main image and the reference scenario, and the large language model generates a scenario description word; The graph generation module inputs the product body and scene description words into the diffusion model to generate the result graph.

6. The device for changing the background of a commodity according to claim 5, characterized in that: The description word generation module specifically includes: setting a scene library, the scene library includes multiple background pictures, setting corresponding background prompt words for each background picture, selecting a background picture, inputting the product main body picture and the background prompt words corresponding to the background picture into the large language model through the set prompt word format, and the large language model generates scene description words.

7. The device for changing the background of a commodity according to claim 5, characterized in that: The cutout module is specifically: The judgment unit judges the pixels of the product image uploaded by the user. If the pixels of the product image are less than 2000×2000, the visual intelligence open platform is directly used to perform a cutout operation to obtain the required product main image; if the pixels of the image are greater than or equal to 2000×2000, the process proceeds to step 12; The scaling unit uses Imgproc.resize in OpenCV and Imgproc.INTER_LAN CZ0S4 algorithm to scale the image in equal proportion to obtain a scaled image; An operation unit performs a cutout operation on the zoomed image through the visual intelligence open platform to obtain a first mask image; The restoration unit uses Imgproc.INTER_LANCZ0S4 algorithm in OpenCV to restore the first mask image to a second mask image of the original size in equal proportion. The cutout unit reads the transparent channel of each pixel in the picture uploaded by the user to obtain the first matrix, reads the transparent channel of each pixel in the second mask image to obtain the second matrix, and calls Core.min to merge the first matrix with the second matrix to obtain the third matrix; The transparent channel in the picture uploaded by the user is replaced by the third matrix to obtain the desired image.

8. The device for changing the background of a commodity according to claim 7, characterized in that: The image generation module is specifically as follows: processing the main image of the product to obtain a first gray background image with a gray background; processing the first gray background image through the Depth algorithm to obtain a first depth image; processing the first gray background image through the Lineart algorithm to obtain a first line drawing; passing the first depth image and the first line drawing image into Controlnet, setting the redrawing amplitude to 0.85, and setting the reference weights of the lineart module and the depth module in Controlnet to 0.5 and 0.6 respectively, and then drawing through the diffusion model to generate a product plan map; processing the generated product plan map through the depth algorithm to obtain a second depth image; processing the gray background image of the product through the lineart algorithm to obtain a second line drawing; sending the second depth image and the second line drawing image to Controlnet respectively, setting the redrawing amplitude to 1.0, and setting the reference weights of the lineart module and the depth module in Controlnet to 0.45 and 0.55 respectively, and then inputting the scene description words into the diffusion model to guide the drawing through Controlnet to generate the result image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.