Background descriptor generation method and device, equipment and medium

Through the visual function based on the large language model, combined with cutouts and user-defined scenes, variable background descriptors are generated, which solves the problem of rigid and single background prompt words in the prior art, and achieves the improvement of generation quality and optimization of language model expression effect.

CN120069063APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108227.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When generating background prompt words, it is difficult for the prior art to fully consider the complexity of the scene, the diversity of language and people's cognitive habits, resulting in the generated background prompt words being rigid and single, and it is impossible to effectively drive the language model to generate language content that is close to reality and innovative.

Method used

Using visual functions based on large language models, variable, user-defined background descriptors are generated by cutting and filling white background maps, combined with user-defined scenes and prompt words.

Benefits of technology

The quality of the generation of background prompt words is improved, making them closely related to the actual background scene, and the expression effect of the language model is optimized to generate language content that is closer to reality to meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069063A_ABST
    Figure CN120069063A_ABST
Patent Text Reader

Abstract

The invention provides a background descriptor generation method and device, equipment and a medium, and the method comprises the steps: carrying out the matting of a commodity image uploaded by a user, and obtaining a commodity main body image; setting the background of the commodity main body as a white background to obtain a white background main body picture; inputting the white-background main body graph into a large language model to obtain a main body description word; inputting a scene defined by a user and the main body descriptor into the large language model to obtain a scene descriptor; inputting the scene description word and the main body description word into a large language model to obtain a position description word; inputting the white-background main body picture into a large language model to obtain a light visual angle description word; the set background cue word is input into the large language model, the main body description word, the scene description word, the position description word and the light view angle description word are input into the large language model to obtain the background description word, and the variable cue word which can be customized by a user is generated based on the large language model, so that the actual demand is better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and particularly to a method, device, equipment and medium for generating background description words. Background Art

[0002] In the application of the StableDiffusion model, effective background prompts are the key to realizing image generation. There are certain limitations in the existing strategies for automatically generating background prompts, resulting in the disconnection between the generated language content and the actual background scene, which affects the expression effect and application effect of the language model.

[0003] The existing method generates background prompts based on fixed templates or simple context information, and it is difficult to fully consider factors such as the complexity of the scene, the diversity of language, and people's cognitive habits. Therefore, the generated background prompts may be too rigid and single, and cannot effectively drive the language model to generate language content that is close to reality and full of innovation. When using this background prompt to generate pictures, the generated pictures do not meet the actual needs. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for generating background description words, which generate variable and user-customizable prompts based on a large language model, and are more in line with actual needs.

[0005] In a first aspect, the present invention provides a method for generating background description words, including the following steps:

[0006] Step 1: Cut out the user-uploaded product picture to obtain a product main body picture;

[0007] Step 2: Set the background of the product main body to white background to obtain a white-background main body picture; input the white-background main body picture into the large language model to obtain a main body description word;

[0008] Step 3: Input the set scene prompt into the large language model, and then input the user-defined scene and the main body description word into the large language model to obtain a scene description word;

[0009] Step 4: Input the set position prompt into the large language model, and then input the scene description word and the main body description word into the large language model to obtain a position description word;

[0010] Step 5: Input the set lighting perspective prompt into the large language model, and then input the white-background main body picture into the large language model to obtain a lighting perspective description word;

[0011] Step 6: Input the set background prompt into the large language model, and then input the subject description word, scene description word, location description word, and lighting perspective description word into the large language model to obtain the background description word.

[0012] In a second aspect, the present invention provides a background description word generation device, including:

[0013] A matte extraction module that extracts the uploaded product image of the user to obtain a product main body image;

[0014] A subject description module that sets the background of the product main body to a white background to obtain a white-background main body image; inputs the white-background main body image into the large language model to obtain the subject description word;

[0015] A scene description module that inputs the set scene prompt into the large language model, and then inputs the user-defined scene and the subject description word into the large language model to obtain the scene description word;

[0016] A location description module that inputs the set location prompt into the large language model, and then inputs the scene description word and the subject description word into the large language model to obtain the location description word;

[0017] A lighting perspective description module that inputs the set lighting perspective prompt into the large language model, and then inputs the white-background main body image into the large language model to obtain the lighting perspective description word;

[0018] A generated description word module that inputs the set background prompt into the large language model, and then inputs the subject description word, scene description word, location description word, and lighting perspective description word into the large language model to obtain the background description word.

[0019] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in the first aspect.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.

[0021] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0022] The present invention adopts a vision function based on a large language model to generate variable and user-customizable prompt words. First, it can greatly improve the generation quality of background prompt words and ensure their close correlation with the actual background scene. Second, it optimizes the expression effect of the language model so that it can generate more realistic language content. Effectively generate background prompt words to adapt to the StableDiffusion model, realize more practical images, and meet user needs.

[0023] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives the specific implementation manners of the present invention. Brief Description of the Drawings

[0024] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0025] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0026] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiment

[0027] The overall idea of the technical solution in the embodiments of the present application is as follows:

[0028] Based on the large language models LLAMA and LLAMA Vision, first, through a matting algorithm, the commodity main body is extracted from the complex background to generate a corresponding grayscale commodity matte image. Then, the background of the commodity image is filled with white to form the final image. Next, the large language model LLAMA Vision extracts the commodity description and generates relevant scene descriptions in combination with the scenes set by the user. Then, LLAMA determines the placement position of the commodity according to the commodity description and the scene setting, and generates relevant content of decorative items to enhance the visual appeal. In addition, LLAMA Vision also extracts the lighting and perspective descriptions of the image, and finally summarizes all the information to generate a comprehensive and detailed prompt word.

[0029] Specifically as follows:

[0030] Step 1: Extract the commodity main body from the original commodity image and generate a corresponding grayscale commodity matte image.

[0031] Step 2: Process the background of the commodity image and fill it with a white background image.

[0032] Among them, the calculation formula for filling the white background can be expressed as:

[0033] S = S 1 *a + S 2 *(1 - a)

[0034] Where S is the filled image, S 1 is the original product image, S 2 is the white background image, and a is the grayscale image of the product mask in Step 1. This formula generates the final filling effect by weighted combination of the original image and the white background image according to the grayscale values.

[0035] Step 3: Extract the main description of the picture in Step 2 based on the large language model llama vision

[0036] Exemplarily, in an embodiment of the present application, a question is asked to llama vision, such as: describethis picture in 10words? Then, llama vision will output the corresponding result, such as: A modernribbedsofainwhitebackground.

[0037] Step 4: Generate a relevant scenario description based on the large language model LLAMA by combining Step 3 and the user-defined scenario.

[0038] Where the user-defined scenario refers to a specific situation envisioned by the user. Due to the unpredictability of the user input, the envisioned scenario may be too broad or narrow, which may cause problems in subsequent steps. Therefore, using the large language model LLAMA for summarization can help generate a more appropriate scenario to ensure the smooth progress of subsequent processes. In this way, the user's scenario envisioning can be effectively adjusted and optimized to better meet the actual needs.

[0039] Exemplarily, in an embodiment of the present application, a question is asked to llama, and the prompt is: Here isaproduct description foryou.Based on the description,you need to imagineascenario where the product can be placed,such as living room,etc. To improve the stability of the output, the following cases are given for its reference: For example: user: A modernribbed sofa.\nkeyword: indoor.\n system: livingroom\n.

[0040] Step 5: Obtain the placement location of the product based on the large language model LLAMA through the product descriptions in Step 3 and Step 4.

[0041] Exemplarily, in an embodiment of the present application, when asking questions to Llama, the complete sentence is: Based on the product description and corresponding scenario given to you, you need to imagine a location where the product can be placed, such as a desktop, floor, sofa, etc. For example: user: {"subject": "A vintage metal lantern.", "Environment": "living room"} system: floor Now let's start. user: {"subject": "A modern ribbed sofa.", "Environment": "living room"}. Output: wooden table.

[0042] Step Six: Based on the large language model Llama, combine Steps Three, Four, and Five, as well as the item information defined by the user, to generate relevant content for the decorative items.

[0043] Among them, the item information defined by the user refers to specific decorative items that the user hopes to add to the background, such as bookshelves, sofas, etc.

[0044] The role of decorative items is to enrich the content of the picture, enhance visual appeal, and highlight the main elements. At the same time, they help control the overall structure of the picture and make the visual layout more harmonious. By reasonably selecting and arranging these decorative items, the aesthetic feeling and layering of the work can be effectively improved, ensuring that the audience's attention is focused on the main elements, thus creating a vivid and interesting visual experience.

[0045] Exemplarily, in an embodiment of the present application, when asking questions to Llama. The input includes: the product main body description in Step Three. The scenario description in Step Four, the placement location description in Step Five, and the item information defined by the user, combined with specific prompt word descriptions, such as: According to the provided information, you need to provide decorative elements for the image.

[0046] Step Seven: Based on the large language model LlamaVision, extract the descriptive words of the lighting and perspective of the picture in Step Two.

[0047] Exemplarily, in an embodiment of the present application, ask questions to llama vision, such as: What perspective was this picture taken from, such as front view, side view, etc. What type of lighting is used during filming, such as natural lighting, sunlight. The output format is json: {"Viewpoint": "side view", "Lighting": "sunlight"}.

[0048] Step Eight: Use the large language model LLAMA to summarize all the above information and organize it into a comprehensive and detailed prompt.

[0049] Exemplarily, in an embodiment of the present application, ask questions to LLAMA, such as: Summarize the JSON information provided below into one sentence, no more than 75 words. for example: user: {"Subject": "Modern Tapered Holder", \n"Placed": "countertop", \n"Environment": "in kitchen",

[0050] "Viewpoint": "side view", "Object": "white wall in the background, fruit bowl with apples, kitchen utensil holders, utensils, ceramic tiles", "Lighting": "natural lighting",} system: Modern Tapered holder resting on the kitchen countertop, white wall in the background, fruit bowl with apples, kitchen utensil holders, utensils, ceramic tiles, natural lighting, side view, detailed texture, High-Resolution Image.

[0051] Example 1

[0052] As Figure 1 shown, this embodiment provides a method for generating background description words, including the following steps:

[0053] Step 1: Cut out the product image uploaded by the user to obtain the main product image;

[0054] Step 2: Set the background of the main product to a white background to obtain a white-background main image; input the white-background main image into the large language model to obtain the main description words;

[0055] Step 3: Input the set scene prompt words into the large language model, and then input the user-defined scene and the main description words into the large language model to obtain the scene description words;

[0056] Step 4: Input the set position prompt words into the large language model, and then input the scene description words and the main description words into the large language model to obtain the position description words;

[0057] Step 5: Input the set lighting perspective prompt words into the large language model, and then input the white-background main image into the large language model to obtain the lighting perspective description words;

[0058] Step 6: Input the set background prompt into the large language model, and then input the subject description word, scene description word, location description word, and lighting perspective description word into the large language model to obtain the background description word. Input this background description word into the StableDiffusion model to guide the generation of its pictures.

[0059] In this embodiment, preferably, step 1 is specifically as follows: Judge the pixels of the product image uploaded by the user. If the pixels of the product image are less than 2000×2000, directly perform a matting operation through the Visual Intelligence Open Platform to obtain the required product main image; if the pixels of the product image are greater than or equal to 2000×2000, use Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image proportionally to obtain a scaled image; perform a matting operation on the scaled image through the Visual Intelligence Open Platform to obtain a first mask image; use Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the first mask image to the original size as a second mask image; read the alpha channel of each pixel point in the product image uploaded by the user to obtain a first matrix, read the alpha channel of each pixel point in the second mask image to obtain a second matrix, and call Core.min to merge the first matrix and the second matrix to obtain a third matrix; replace the alpha channel in the product image uploaded by the user with the third matrix to obtain the required product main image.

[0060] Due to the matting size limit of the Alibaba Cloud Visual Intelligence Open Platform, when the longest side exceeds 2000 pixels, it is necessary to scale the image proportionally. Call Imgproc.resize with the Imgproc.INTER_LANCZOS4 algorithm to perform image scaling.

[0061] Call Imgproc.resize with the Imgproc.INTER_LANCZOS4 algorithm to perform image scaling. This algorithm can reduce the impact of artifacts while maintaining the edge sharpness.

[0062] Obtain the matting result according to the black and white image + original image:

[0063] a. Take out the alpha channel of the original image. Call Core.min, which will take the minimum value of the transparency channels of the mask and the original image, ensuring that the original information of the transparency channel is retained. Through this step of processing, the lines of the obtained matting can be made more perfect;

[0064] b. Remove the alpha channel of the original image and add the black and white image as the new alpha channel for layer merging;

[0065] c. Set the color value of the transparent area to black. This step is to reduce the size of the cut-out result image and save storage costs.

[0066] The core code is as follows:

[0067] / / Create a fully transparent matrix with the same size as the original image for comparison;

[0068] Mat compareAlpha = new Mat(outImg.size(), CvType.CV_8UC1, Scalar.all(0.0));

[0069] / / Used to save the comparison result;

[0070] Mat compareResult = new Mat();

[0071] / / Compare the alpha channel of the original image with alpha. The positions with a value of 0 in the resulting mask are the transparent positions in the original image;

[0072] Core.compare(outPlanes.get(3), compareAlpha, compareResult, Core.CMP_EQ);

[0073] / / Create a fully black matrix with the same size as the original image;

[0074] Mat black = new Mat(outImg.size(), outImg.type(), Scalar.all(0));

[0075] / / Copy the black color in black to the corresponding positions in outImg according to the mask;

[0076] Core.bitwise_and(black, outImg, outImg, compareResult);

[0077] In this embodiment, preferably, step 6 is specifically as follows: Input the set accessory prompt words into the large language model, and then input the subject description words, scene description words, position description words, and accessory description information into the large language model to obtain accessory description words;

[0078] Input the set background prompt words into the large language model, and then input the subject description words, scene description words, accessory description words, position description words, and lighting perspective description words into the large language model to obtain background description words.

[0079] In this embodiment, preferably, the background prompt words include the set output format.

[0080] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.

[0081] Embodiment 2

[0082] As Figure 2 shown, in this embodiment, a background description word generation apparatus is provided, including:

[0083] A matte extraction module that extracts the matte of the product image uploaded by the user to obtain a product main body image;

[0084] A main body description module that sets the background of the product main body to a white background to obtain a white background main body image; inputs the white background main body image into a large language model to obtain main body description words;

[0085] A scene description module that inputs the set scene prompt words into a large language model, and then inputs the user-defined scene and the main body description words into the large language model to obtain scene description words;

[0086] A position description module that inputs the set position prompt words into a large language model, and then inputs the scene description words and the main body description words into the large language model to obtain position description words;

[0087] A lighting perspective description module that inputs the set lighting perspective prompt words into a large language model, and then inputs the white background main body image into the large language model to obtain lighting perspective description words;

[0088] A description word generation module that inputs the set background prompt words into a large language model, and then inputs the main body description words, scene description words, position description words, and lighting perspective description words into the large language model to obtain background description words.

[0089] In this embodiment, preferably, the matte extraction module specifically operates as follows: It determines the pixels of the product image uploaded by the user. If the pixels of the product image are less than 2000×2000, it directly performs matte extraction operations through the Visual Intelligence Open Platform to obtain the required product main image. If the pixels of the product image are greater than or equal to 2000×2000, it uses the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image proportionally to obtain a scaled image. It performs matte extraction operations on the scaled image through the Visual Intelligence Open Platform to obtain a first mask image. It uses the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to restore the first mask image proportionally to the original size to obtain a second mask image. It reads the alpha channel of each pixel point in the product image uploaded by the user to obtain a first matrix, reads the alpha channel of each pixel point in the second mask image to obtain a second matrix, and calls Core.min to merge the first matrix and the second matrix to obtain a third matrix. It replaces the alpha channel in the product image uploaded by the user with the third matrix to obtain the required product main image.

[0090] In this embodiment, preferably, the descriptor generation module specifically operates as follows: It inputs the set accessory prompt words into the large language model, and then inputs the main body descriptor, scene descriptor, position descriptor, and accessory description information into the large language model to obtain accessory descriptor words.

[0091] It inputs the set background prompt words into the large language model, and then inputs the main body descriptor, scene descriptor, accessory descriptor words, position descriptor, and lighting perspective descriptor into the large language model to obtain background descriptor words.

[0092] In this embodiment, preferably, the background prompt words include the set output format.

[0093] Since the device introduced in the second embodiment of the present invention is the device used to implement the method in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0094] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment. For details, see the third embodiment.

[0095] Embodiment Three

[0096] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in Embodiment 1 can be implemented.

[0097] Since the electronic device introduced in this embodiment is the device adopted to implement the method in Embodiment 1 of the present application, based on the method introduced in Embodiment 1 of the present application, those skilled in the art can understand the specific implementation manners and various forms of variation of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in the embodiment of the present application belongs to the scope protected by the present application.

[0098] Based on the same inventive concept, the present application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.

[0099] Embodiment 4

[0100] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be implemented.

[0101] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0105] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative only and not used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of the claims of the present invention.

Claims

1. A method for generating background description words, characterized in that: The steps include: Step 1: Cut out the product image uploaded by the user to obtain the main image of the product; Step 2: Set the background of the product body to white to obtain a white-background main body image; input the white-background main body image into the large language model to obtain the main body description word; Step 3: input the set scene prompt words into the large language model, and then input the user-defined scene and the subject description words into the large language model to obtain the scene description words; Step 4: input the set location prompt word into the large language model, and then input the scene description word and the subject description word into the large language model to obtain the location description word; Step 5: input the set lighting angle prompt words into the large language model, and then input the white background main image into the large language model to obtain the lighting angle description words; Step 6: input the set background prompt words into the large language model, and then input the subject description words, scene description words, location description words and lighting perspective description words into the large language model to obtain background description words.

2. The method for generating background description words according to claim 1, characterized in that: The step 1 is specifically as follows: judging the pixels of the product image uploaded by the user, if the pixels of the product image are less than 2000×2000, directly performing a cutout operation through the visual intelligence open platform to obtain the required product main image; if the pixels of the product image are greater than or equal to 2000×2000, using Imgproc.resize in OpenCV and using the Imgproc.INTER_LANCZ0S4 algorithm to scale the product image in proportion to obtain a scaled image; performing a cutout operation on the scaled image through the visual intelligence open platform to obtain a first mask image; restoring the first mask image to a second mask image of the original size in proportion through Imgproc.resize in OpenCV and using the Imgproc.INTER_LANCZ0S4 algorithm; reading out the transparent channel of each pixel in the product image uploaded by the user to obtain a first matrix, reading out the transparent channel of each pixel in the second mask image to obtain a second matrix, calling Core.min to merge the first matrix with the second matrix to obtain a third matrix; The transparent channel in the product image uploaded by the user is replaced by the third matrix to obtain the required product main image.

3. The method for generating background description words according to claim 1, characterized in that: The step 6 specifically includes: inputting the set modifier prompt word into the large language model, and then inputting the subject description word, scene description word, location description word and modifier description information into the large language model to obtain the modifier description word; The set background prompt words are input into the large language model, and then the subject description words, scene description words, modifier description words, location description words and lighting perspective description words are input into the large language model to obtain background description words.

4. The method for generating background description words according to claim 1, characterized in that: The background prompt words include the format of setting the output.

5. A background description word generation device, characterized in that: include: The cutout module cuts out the product image uploaded by the user to obtain the main image of the product; The subject description module sets the background of the product subject to white to obtain a white-background subject image; the white-background subject image is input into the large language model to obtain subject description words; A scene description module inputs the set scene prompt words into the large language model, and then inputs the user-defined scene and the subject description words into the large language model to obtain scene description words; A location description module inputs the set location prompt word into the large language model, and then inputs the scene description word and the subject description word into the large language model to obtain the location description word; The lighting perspective description module inputs the set lighting perspective prompt words into the large language model, and then inputs the white background main image into the large language model to obtain the lighting perspective description words; A description word generation module inputs the set background prompt words into the large language model, and then inputs the subject description words, scene description words, location description words and lighting perspective description words into the large language model to obtain background description words.

6. The background description word generating device according to claim 5, characterized in that: The cutout module is specifically as follows: the pixels of the product image uploaded by the user are judged. If the pixels of the product image are less than 2000×2000, the cutout operation is directly performed through the visual intelligence open platform to obtain the required product main body image; if the pixels of the product image are greater than or equal to 2000×2000, the product image is proportionally scaled using the Imgproc.resize in OpenCV and the Imgproc.INTER_LANCZ0S4 algorithm to obtain a scaled image; the scaled image is cutout through the visual intelligence open platform to obtain a first mask image; the first mask image is proportionally restored to a second mask image of the original size through the Imgproc.resize in OpenCV and the Imgproc.INTER_LANCZ0S4 algorithm; the transparent channel of each pixel in the product image uploaded by the user is read out to obtain a first matrix, the transparent channel of each pixel in the second mask image is read out to obtain a second matrix, and Core.min is called to merge the first matrix with the second matrix to obtain a third matrix; The transparent channel in the product image uploaded by the user is replaced by the third matrix to obtain the required product main image.

7. The background description word generating device according to claim 5, characterized in that: The generating description word module specifically includes: inputting the set modifier prompt word into the large language model, and then inputting the subject description word, scene description word, location description word and modifier description information into the large language model to obtain the modifier description word; The set background prompt words are input into the large language model, and then the subject description words, scene description words, modifier description words, location description words and lighting perspective description words are input into the large language model to obtain background description words.

8. The background description word generating device according to claim 5, characterized in that: The background prompt words include the format of setting the output.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.