Character commodity graph generation method, device and equipment
By using pre-trained text product diagrams to generate models and combining image patching large models, the problem of inconsistency between product diagrams and text in the prior art is solved, and high-quality and diverse text product diagrams are generated, which is suitable for a wide range of practical applications.
Patent Information
- Application Number
- CN202510224167.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The existing product image generation scheme cannot generate product images that are highly consistent with the text content, resulting in the inability to widely promote them in actual applications.
Using the pre-trained text product image generation model method, a highly consistent text product image is generated by inputting the description text and the original product image into the model, and combining the image patching large model. The model includes a text telephony separation module, an image depth map module, a text position encoder, a text encoding module and a product image segmentation module.
It realizes the generation of product images that are highly consistent with the text, improves the quality and diversity of generated images, can be widely promoted in practical applications, and has aesthetics and practicality.
Smart Images

Figure CN120163901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method, device and equipment for generating a text product image. Background Art
[0002] At present, artificial intelligence technology, which is in a stage of rapid development, has injected innovative power into traditional industries, opened up breakthrough emerging industries, and has become an important force for technological change and progress in social industries. Artificial intelligence content automatic generation technology is obviously one of the hottest emerging technologies at present, breaking through the traditional content creation methods, improving creation efficiency and quality, and promoting innovation and development in various industries. Generating image content based on text guidance has become a popular research direction. This technology uses a personalized description text and combines artificial intelligence models to generate image content that matches the description text. The proposal and development of large models is an important progress in the field of artificial intelligence in recent years. As the performance of large models has been verified, the application ecology adapted to large models has also been gradually improved, ensuring the possibility of implementation in practical applications, making solutions based on large model technology widely respected. Therefore, using large model technology to complete text-based image generation has become a popular solution.
[0003] As the technology of generating images with text guidance matures, this technical solution has derived applications in various sub-fields, among which significant results have been achieved in product image generation. Text product images generated based on text guidance can meet user needs and give full play to the creative performance of large model technology. Most previous product image generation solutions have some shortcomings, such as the large difference between the generated image content and the text description, the destruction of the original product content, the low degree of product image fusion, and the inability to generate text content. These problems limit the usability in practical applications. Summary of the invention
[0004] In view of this, the purpose of the present invention is to propose a method, device and equipment for generating text product images, aiming to solve the problem that existing product image generation schemes cannot generate product images that are highly consistent with the text content, resulting in the inability to be widely promoted in practical applications.
[0005] To achieve the above object, the present invention provides a method for generating a text product image, which is implemented based on a pre-trained text product image generation model, and the method includes:
[0006] Inputting the description text into the text product image generation model to obtain a first processing result;
[0007] Inputting the original product image to be processed into the text product image generation model to obtain a second processing result;
[0008] Input the first processing result and the second processing result into the image inpainting large model in the text commodity image generation model to generate a text commodity image; wherein, the text commodity image generation model includes a text prompt separation module, an image depth map module, a text position encoder, a text encoding module, and a commodity image segmentation module.
[0009] Preferably, the inputting the description text into the text commodity image generation model to obtain the first processing result includes:
[0010] Perform text processing on the description text through the text prompt separation module, the image depth map module, the text position encoder, and the text encoding module in the text commodity image generation model to obtain the first processing result including an image depth map, text generation coordinates, a text contour map, and a text guidance map.
[0011] Preferably, the inputting the original commodity image to be processed into the text commodity image generation model to obtain the second processing result includes:
[0012] Input the original commodity image into the commodity image segmentation module in the text commodity image generation model to obtain the second processing result including the commodity image.
[0013] Preferably, after the inputting the original commodity image to be processed into the text commodity image generation model to obtain the second processing result, it further includes:
[0014] Based on a preset image expected size, input the second processing result combined with the image description text into the commodity position encoder to generate a commodity generation position, and input the commodity generation position into the image inpainting large model together; wherein, the image description text is obtained by processing the description text through the text prompt separation module.
[0015] Preferably, the inputting the first processing result and the second processing result into the image inpainting large model in the text commodity image generation model includes:
[0016] Input a preset scene style in the stylization enhancement module into the image inpainting large model to generate a stylized text commodity image.
[0017] Preferably, the performing text processing on the description text through the text prompt separation module, the image depth map module, the text position encoder, and the text encoding module in the text commodity image generation model to obtain the first processing result including an image depth map, text generation coordinates, a text contour map, and a text guidance map includes:
[0018] Pass the description text through the text prompter separation module to obtain an image description text and text content, where the image description text is used to describe the expected generated content, and the text content is used to describe the specific text expected in the generated content;
[0019] Pass the image description text through the image depth map module to obtain the image depth map;
[0020] Pass the image description text and the text content through the text position encoder to obtain the text generation coordinates;
[0021] Pass the text content through the text encoding module to obtain the text guidance map and the text contour map.
[0022] Preferably, the training process of the image depth map module includes:
[0023] Construct a first pre-training model through multiple convolutional layers, input a preset number of first training data pairs into the first pre-training model and train based on a preset pixel loss function to obtain the image depth map module; where the first training data pair includes a preset number of descriptions and depth maps.
[0024] Preferably, the training process of the text position encoder includes:
[0025] Construct a second pre-training model by adopting the encoder structure of U-net, input a preset number of second training data pairs into the second pre-training model and train based on a preset IOU loss function to obtain the text position encoder; where the second training data pair includes a description, the original image width and height, and the text extraction result, and the text extraction result is the text content extracted from the original image.
[0026] Preferably, the training process of the commodity position encoder includes:
[0027] Construct a third pre-training model by adopting the encoder structure of U-net, input a preset number of third training data pairs into the third pre-training model and train based on a preset IOU loss function to obtain the commodity position encoder; where the third training data pair includes a description, the original image width and height, and the commodity separation extraction result, and the commodity separation extraction result is the commodity image separated from the original image.
[0028] Preferably, the training process of the image inpainting large model includes:
[0029] Input the constructed image training data into a model constructed with a convolutional layer and a Transformer module, and train it with a preset pixel loss function and an OCR loss function to obtain the large image repair model; wherein, the image training data includes the original image, the commodity image, the text content, and the description text.
[0030] To achieve the above object, the present invention provides a device for generating a text commodity map. The device is implemented based on a pre-trained text commodity map generation model, and the device includes:
[0031] A text processing unit, configured to input the description text into the text commodity map generation model to obtain a first processing result;
[0032] An image processing unit, configured to input the original commodity image to be processed into the text commodity map generation model to obtain a second processing result;
[0033] A commodity map generation unit, configured to input the first processing result and the second processing result into the large image repair model in the text commodity map generation model to generate a text commodity map; wherein, the text commodity map generation model includes a text prompt separation module, an image depth map module, a text position encoder, a text encoding module, and a commodity map segmentation module.
[0034] To achieve the above object, the present invention also proposes a device for generating a text commodity map, including a processor, a memory, and a computer program stored in the memory. The computer program is executed by the processor to implement the steps of a method for generating a text commodity map as described in the above embodiment.
[0035] To achieve the above object, the present invention also proposes a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the steps of a method for generating a text commodity map as described in the above embodiment.
[0036] To achieve the above object, the present invention also proposes a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of a method for generating a text commodity map as described in the above embodiment are implemented.
[0037] Beneficial effects:
[0038] The above solution is based on a pre-trained text-to-product-image generation model. By extracting text information from multiple dimensions, it ensures the accuracy and diversity of the generated text. Without compromising the quality of product image generation, the text can be highly integrated into the product image, thus achieving efficient and accurate generation of text product images that contain products, background content, and text content. The generated text product images exhibit both aesthetic appeal and practicality and can be directly used for product display and extensive promotion.
[0039] The above solution is based on generating a product image. When the user does not specify the position of the product to be generated, a product position encoder is used to provide a position that fits the description text, improving the position of the product in the generated image and greatly enhancing the layout rationality of the output text product image, making the generated text product image more in line with the user's expectations and needs.
[0040] The above solution, when a more stylized text product image needs to be generated, by introducing a stylization enhancement module, can further enhance the style in a specific scenario when generating the text product image, making the generated image have stronger artistic sense and expressiveness, thus meeting the user's needs in different scenarios and enhancing the diversity and practicality of the generated image.
[0041] The above solution extracts the text content to be generated from the description text. Through a text position encoder, a text encoding module, etc., the text content is converted into various guiding information to guide content generation from multiple dimensions. It can ensure the accuracy of the generated text content and also highlight the diversity of the generated text content, making it more compatible with the entire generated text product image. In order to highly integrate the product image with the text content and the content generated from the description text, an image depth module is used to initially obtain an image depth map to guide the generation content of the image inpainting large model; the image inpainting large model combines the intermediate information converted from the original product image, text, etc., can losslessly retain the content of the product image, accurately generate the text content, and the background content can also highly match the text description. Moreover, due to the control of the accuracy of the generated text content in this solution, not only simple English characters can be generated, but also complex Chinese text content can be generated in the image. Combining with the current mature large model ecosystem, this solution can generate high-quality text product images according to text guidance and can be quickly implemented in practical applications, having practical application significance.
[0042] The above solution improves the performance and accuracy of the corresponding modules through the training of the image depth map module, text position encoder, and product position encoder, ensuring the stability and reliability of each module during the generation process, thereby improving the overall generation quality and efficiency; at the same time, the training of the image inpainting large model enables the generated text product images to have better background integration and detail expressiveness while maintaining the integrity of the product image and text content. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0044] Figure 1 It is a schematic flowchart of a method for generating a text commodity graph provided by an embodiment of the present invention.
[0045] Figure 2 It is a schematic overall flowchart of generating a text commodity graph provided by an embodiment of the present invention.
[0046] Figure 3 It is a schematic structural diagram of a device for generating a text commodity graph provided by an embodiment of the present invention.
[0047] The realization of the invention purpose, functional features and advantages will be further described with reference to the embodiments and the drawings. Specific Embodiments
[0048] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0049] The following elaborates on the content of the present invention in detail with reference to the embodiments.
[0050] Refer to Figure 1 As shown, it is a schematic flowchart of a method for generating a text commodity graph provided by an embodiment of the present invention.
[0051] In this embodiment, the method is implemented based on a pre-trained text commodity graph generation model. The method includes:
[0052] S01, input the description text into the text commodity graph generation model to obtain a first processing result;
[0053] S02. Input the original product image to be processed into the text product image generation model to obtain a second processing result;
[0054] S03. Input the first processing result and the second processing result into the image inpainting large model in the text product image generation model to generate a text product image; wherein, the text product image generation model includes a text prompt separation module, an image depth map module, a text position encoder, a text encoding module, and a product image segmentation module.
[0055] Further, inputting the description text into the text product image generation model to obtain a first processing result includes:
[0056] S11. Perform text processing on the description text through the text prompt separation module, the image depth map module, the text position encoder, and the text encoding module in the text product image generation model to obtain the first processing result including an image depth map, text generation coordinates, a text contour map, and a text guidance map.
[0057] Further, in step S11, the performing text processing on the description text through the text prompt separation module, the image depth map module, the text position encoder, and the text encoding module in the text product image generation model to obtain the first processing result including an image depth map, text generation coordinates, a text contour map, and a text guidance map includes:
[0058] S11-1. Pass the description text through the text prompt separation module to obtain an image description text and text content, wherein the image description text is used to describe the expected generated content, and the text content is used to describe the specific text expected in the generated content;
[0059] S11-2. Pass the image description text through the image depth map module to obtain the image depth map;
[0060] S11-3. Pass the image description text and the text content through the text position encoder to obtain the text generation coordinates;
[0061] S11-4. Pass the text content through the text encoding module to obtain the text guidance map and the text contour map.
[0062] Further, inputting the original product image to be processed into the text product image generation model to obtain a second processing result includes:
[0063] S12. Input the original product image into the product image segmentation module in the text product image generation model to obtain the second processing result including the product image.
[0064] Further, input the first processing result and the second processing result into the image inpainting large model in the text commodity image generation model to generate a text commodity image, including:
[0065] S13. Input the commodity image, the image depth map, the text generation coordinates, the text contour map, and the text guidance map into the image inpainting large model in the text commodity image generation model to generate a text commodity image; wherein, the text commodity image includes background content, and the generation of the background content is generated by a descriptive text guidance large model.
[0066] In this embodiment, a text commodity image generation model is constructed and trained. The text commodity image generation model includes a text prompt separation module, an image depth map module, a text position encoder, a text encoding module, a commodity segmentation module, and an image inpainting large model; further, the text commodity image generation model also includes a commodity position encoder and a style enhancement module, and these two modules can be selectively applied according to user expectations. Refer to Figure 2 As shown, the entire process of generating a text commodity image includes a text processing stage and an image processing stage. Specifically:
[0067] In the text processing stage, through the extraction of text information in multiple dimensions, the accuracy and diversity of the generated text are ensured, and it is not limited to Chinese or English. First, according to the text / descriptive text text input by the user, the text prompt separation module f split is used. This module contains a preset rule processing logic, that is, f split (text) = (text prompt , text content ). Its purpose is to be able to process the text text to obtain the image description text text prompt and the text content text content . The image description text and the text content are respectively used to describe the expected generated content and the specific text expected in the generated content. The image description text uses a lightweight image depth map module to quickly obtain an image depth map image depth that can fit the description text. The image description text and the text content pass through the text position encoder to obtain the text generation coordinates position text ; both the image depth map and the text generation coordinates are for making the generated content more integrated and coordinated. In order to accurately inject the text content into the image generation process and accurately maintain the text content, by using the text encoding module to render and draw the text content, a text guidance map image text based on different font styles is obtained, and the text contour map image contour is extracted.(The text outline map only contains the edge part of the text, while the text guidance map contains the complete text); the text guidance map ensures the accuracy of the generated text, and the text outline map avoids the singularity of the generated text.
[0068] In the image processing stage, according to the original product image image input by the user product , first use the product image segmentation module to separate the product placed in a complex background to obtain a product image with a clean background. When there is a specified product at the position of the generated image, input the product image, the image depth map, the text generation coordinates, the text outline map, and the text guidance map into the image inpainting large model to obtain the text product map.
[0069] Further, after inputting the original product image to be processed into the text product map generation model to obtain the second processing result, it further includes:
[0070] Based on the preset expected image size, input the second processing result combined with the image description text into the product position encoder to generate the product generation position, and input the product generation position into the image inpainting large model; wherein, the image description text is obtained by processing the description text through the text prompt separation module.
[0071] Further, inputting the first processing result and the second processing result into the image inpainting large model in the text product map generation model includes:
[0072] By introducing the preset scene style in the stylization enhancement module into the image inpainting large model, a stylized text product map is generated.
[0073] In another embodiment, when there is no clear specification of the product position in the generated image, according to the expected generated image size, the product image combined with the image description text passes through the product position encoder to obtain the product generation position position product ; then, combined with the output of the previous text processing stage, input the image depth map, the text generation coordinates, the text outline map, the text guidance map, the expected image size, and the product generation position into the image inpainting large model. Under the guidance of the image description text, the image inpainting large model can ensure that the product map is not damaged and has the ability to generate background content and text; further, according to the selection of whether to further enhance the style in a specific scene, the style enhancement module can be superimposed, and finally a text product map image is output output .
[0074] Further, the training process of the image depth map module includes:
[0075] The first pre-training model is composed of multiple convolutional layers. The constructed preset number of first training data pairs are input into the first pre-training model and trained based on a preset pixel loss function to obtain the image depth map module. Among them, the first training data pair includes a preset number of descriptions and depth maps.
[0076] Further, the training process of the text position encoder includes:
[0077] The second pre-training model is constructed by adopting the encoder structure of U-net. The constructed preset number of second training data pairs are input into the second pre-training model and trained based on a preset IOU loss function to obtain the text position encoder. Among them, the second training data pair includes a description, the original image width and height, and the text extraction result, and the text extraction result is the text content extracted from the original image.
[0078] Further, the training process of the commodity position encoder includes:
[0079] The third pre-training model is constructed by adopting the encoder structure of U-net. The constructed preset number of third training data pairs are input into the third pre-training model and trained based on a preset IOU loss function to obtain the commodity position encoder. Among them, the third training data pair includes a description, the original image width and height, and the commodity separation extraction result, and the commodity separation extraction result is the commodity image separated from the original image.
[0080] Further, the training process of the image inpainting large model includes:
[0081] The constructed image training data is input into a model constructed by a convolutional layer and a transformer module and trained with a preset pixel loss function and an OCR loss function to obtain the image inpainting large model. Among them, the image training data includes the original image, the commodity image, the text content, and the description text.
[0082] In this embodiment, data preparation and preprocessing are first performed: This solution includes 2 million image training data sets, and all of the images contain target commodity images. The preprocessing of the image training data mainly includes: standardizing the size of the image, separating and extracting the commodities in the image, recognizing and extracting the text in the image, and accurately labeling and pairing the image text descriptions. After preprocessing, each piece of image training data contains: the original image, the commodity image, the text content, and the description text. In the inference stage, only the description text and the original commodity image need to be input into the text commodity image generation model to generate the final text commodity image.
[0083] By combining the content of each module, the input data is transformed and finally input into the image inpainting large model. Specifically:
[0084] Text prompter separation module, which processes the input text through preset character rules using only specific logic (special text character recognition) to separately extract the image description text and the text content without the need for training.
[0085] Image depth map module, which is a pre-trained model composed of 7 convolutional layers and can convert the input text prompt to end-to-end output to obtain an image depth , and this model is trained with 300,000 pairs of description t and depth map I depth data pairs. Among them, the pixel loss Loss depth = D (t) - I depth is adopted, where D (t) represents the depth map generated by the image depth map module using the description t.
[0086] Text position encoder, which is a pre-trained model using the encoder structure of U-net (U-shaped network, an architecture based on convolutional neural network (CNN)), maps the text content and the image description text into two-dimensional features, and obtains the center point, width, and height of the text generation area through encoder regression. The training data of this model is 500,000 data pairs extracted from 2 million data of this solution, which are composed of the description text, the original image width and height, and the text extraction result (the text content extracted from the original image). Among them, the IOU loss (Intersection over Union Loss) is adopted.
[0087] Text encoding module, which adopts various font styles and selects a font using specific logic, draws the text content in a set format, and further extracts the outline of the text from the drawn result, and can obtain the text outline map and the text guidance map without training.
[0088] Commodity position encoder, which is a pre-trained model using the encoder structure of U-net, maps the commodity image and the image description text to the same feature level, and obtains the center point, width, and height of the commodity image in the generation area through encoder regression. The training data of this model is 300,000 data pairs extracted from 2 million data of this solution, which are composed of the description text, the original image width and height, and the commodity separation and extraction result. The IOU loss is also adopted.
[0089] The image inpainting large model mainly consists of a convolutional layer and a Transformer module. By constructing a model structure for image region inpainting and further training the model with 2 million pieces of data, in addition to using the same denoising loss as the original model, a pixel loss and an OCR loss (Optical Character Recognition loss) are added. Among them, the output text commodity image image output = F(text prompt , image depth , image text , iamge contour , position product , position text , size), where size is the preset size of the generated text commodity image. The pixel loss ensures the accuracy of the generated image, that is, Loss pixel = image - image output , where image is the original image; the OCR loss ensures the accuracy of the generated text, that is, Loss OCR = OCR(image) - OCR(image output ), where OCR is the text recognition function in the image.
[0090] In this solution, a text commodity image is generated through text guidance. Among them, the image depth map module, the text position encoder, and the commodity position encoder are independently pre-trained. The commodity segmentation module and the style enhancement adopt existing solutions (for example, the commodity segmentation module mainly includes a model composed of convolutions and is obtained through self-training; the style enhancement module mainly adopts the LoRA model (Low-Rank Adaptation of Large Language Models) obtained through self-training). Combining the text prompting separation module and the text encoding module, a method for generating a text commodity image based on text guidance is realized by training the image inpainting large model.
[0091] Refer to Figure 3 The following figure shows a schematic structural diagram of a device for generating a text commodity image provided by an embodiment of the present invention.
[0092] In this embodiment, the device is implemented based on a pre-trained text commodity image generation model. The device 20 includes:
[0093] A text processing unit 21, configured to input a description text into the text commodity image generation model to obtain a first processing result;
[0094] An image processing unit 22 for inputting the original product image to be processed into the text product image generation model to obtain a second processing result;
[0095] A product image generation unit 23 for inputting the first processing result and the second processing result into an image inpainting large model in the text product image generation model to generate a text product image; wherein, the text product image generation model includes a text prompting separation module, an image depth map module, a text position encoder, a text encoding module, and a product image segmentation module.
[0096] Furthermore, the product image generation unit is further configured to generate a stylized text product image by introducing a preset scene style in the stylization enhancement module into the image inpainting large model.
[0097] In another embodiment, the device further includes:
[0098] An image position generation unit for combining the second processing result with image description text based on a preset desired image size, inputting the combined result into a product position encoder to generate a product generation position, and inputting the product generation position into the image inpainting large model; wherein, the image description text is obtained by processing the description text through the text prompting separation module.
[0099] Each unit module of the device 20 can respectively execute the corresponding steps in the above method embodiments, so the unit modules will not be elaborated here. For details, please refer to the descriptions of the corresponding steps above.
[0100] An embodiment of the present invention further provides a device for generating a text product image. The device includes the device for generating a text product image as described above. Among them, the device for generating a text product image can adopt Figure 3 the structure of the embodiment, and correspondingly, it can execute Figure 1 the technical solutions of the method embodiments shown. The implementation principle and technical effects are similar. For details, please refer to the relevant records in the above embodiments and will not be elaborated here.
[0101] The device includes: devices with a photographing function such as mobile phones, digital cameras, or tablet computers, or devices with image processing functions, or devices with image display functions. The device may include components such as a memory, a processor, an input unit, a display unit, and a power supply.
[0102] Among them, the memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as an image playback function, etc.); the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide access to the memory for the processor and the input unit.
[0103] The input unit can be used to receive input digital or character or image information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, in addition to including a camera, the input unit of this embodiment can also include a touch-sensitive surface (such as a touch display screen) and other input devices.
[0104] The display unit can be used to display information input by the user or information provided to the user and various graphical user interfaces of the device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit can include a display panel. Optionally, the display panel can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor to determine the type of touch event. Subsequently, the processor provides a corresponding visual output on the display panel according to the type of touch event.
[0105] The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium can be the computer-readable storage medium included in the memory in the above embodiment; it can also exist separately and be a computer-readable storage medium not assembled into the device. At least one instruction is stored in the computer-readable storage medium, and the instruction is loaded and executed by the processor to implement Figure 1 the method for generating the text commodity map shown. The computer-readable storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0106] The embodiment of the present invention also provides a computer program product, including a computer program / instructions, and the computer program / instructions are loaded and executed by the processor to implement Figure 1 the method for generating a text commodity map shown.
[0107] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the apparatus embodiments, device embodiments, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments.
[0108] Also, in this document, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0109] The above description shows and describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments. Instead, it can be used in various other combinations, modifications, and environments, and can be changed within the scope of the inventive concept herein through the above teachings or the skills or knowledge in the relevant field. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for generating a text product image, characterized in that: The method is implemented based on a pre-trained text product image generation model, and the method includes: Inputting the description text into the text product image generation model to obtain a first processing result; Inputting the original product image to be processed into the text product image generation model to obtain a second processing result; The first processing result and the second processing result are input into the image repair model in the text product image generation model to generate a text product image; wherein the text product image generation model includes a text prompt separation module, an image depth map module, a text position encoder, a text encoding module and a product image segmentation module.
2. The method for generating a text product image according to claim 1, characterized in that: The inputting the description text into the text product image generation model to obtain a first processing result includes: The description text is processed through the text prompt separation module, image depth map module, text position encoder, and text encoding module in the text product image generation model to obtain the first processing result including the image depth map, text generation coordinates, text contour map, and text guide map.
3. The method for generating a text product image according to claim 2, characterized in that: The step of inputting the original product image to be processed into the text product image generation model to obtain a second processing result includes: The original product image is input into the product image segmentation module in the text product image generation model to obtain the second processing result including the product image.
4. The method for generating a text product image according to claim 1, characterized in that: After inputting the original product image to be processed into the text product image generation model to obtain the second processing result, the method further includes: Based on a preset expected image size, the second processing result is combined with the image description text and input into the product position encoder to generate a product generation position, and the product generation position is also input into the image repair model; wherein the image description text is obtained by processing the description text through the text prompt separation module.
5. The method for generating a text product image according to claim 1, characterized in that: The step of inputting the first processing result and the second processing result into the image repair model in the text product image generation model includes: By introducing the preset scene style in the stylized enhancement module into the image repair model, a stylized text product image is generated.
6. The method for generating a text product image according to claim 2, characterized in that: The step of processing the description text through the text prompt separation module, the image depth map module, the text position encoder, and the text encoding module in the text product image generation model to obtain the first processing result including the image depth map, the text generation coordinates, the text contour map, and the text guide map includes: The description text is passed through the text prompt separation module to obtain an image description text and text content, wherein the image description text is used to describe the desired generated content, and the text content is used to describe the specific text expected in the generated content; Passing the image description text through an image depth map module to obtain the image depth map; Pass the image description text and the text content through the text position encoder to obtain the text generation coordinates; The text content is passed through the text encoding module to obtain the text guide map and the text outline map.
7. The method for generating a text product image according to claim 1, characterized in that: The training process of the image depth map module includes: A first pre-training model is constructed by multiple convolutional layers, a preset number of first training data pairs are input into the first pre-training model and trained based on a preset pixel loss function to obtain the image depth map module; wherein the first training data pairs include a preset number of descriptions and depth maps; The training process of the text position encoder includes: A second pre-training model is constructed by adopting a U-net encoder structure, a preset number of second training data pairs are input into the second pre-training model and trained based on a preset IOU loss function to obtain the text position encoder; wherein the second training data pair includes a description, an original image width and height, and a text extraction result, and the text extraction result is the text content extracted from the original image; The training process of the image inpainting large model includes: The constructed image training data is input into a model constructed by a convolutional layer and a transformer module, and is trained in a preset pixel loss function and an OCR loss function to obtain the image repair model; wherein the image training data includes the original image, the product image, the text content and the description text.
8. The method for generating a text product image according to claim 4, characterized in that: The training process of the commodity position encoder includes: A third pre-training model is constructed by adopting the encoder structure of U-net, and a preset number of third training data pairs are input into the third pre-training model and trained based on a preset IOU loss function to obtain the product position encoder; wherein the third training data pair includes a description, an original image width and height, and a product separation extraction result, and the product separation extraction result is a product image separated from the original image.
9. A device for generating a text product image, characterized in that: The device is implemented based on a pre-trained text commodity image generation model, and the device includes: A text processing unit, used for inputting the description text into the text product image generation model to obtain a first processing result; An image processing unit, used for inputting the original product image to be processed into the text product image generation model to obtain a second processing result; The product image generation unit is used to input the first processing result and the second processing result into the image repair model in the text product image generation model to generate a text product image; wherein the text product image generation model includes a text prompt separation module, an image depth map module, a text position encoder, a text encoding module and a product image segmentation module.
10. A device for generating a text product image, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory, wherein the computer program is executed by the processor to implement the steps of a method for generating a text product image as described in any one of claims 1 to 8.