Image generation method and device

By using inquiry dialogue data and decision factor generation model, combined with the target literary graph model and quality audit model, the problem of artificial dependence and high repetition rate in product display graph generation is solved, efficient and personalized image generation is achieved, and the creative production capacity of e-commerce platforms is improved.

CN120563191APending Publication Date: 2025-08-29阿里巴巴(中国)网络技术有限公司
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510517827.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the prior art, the generation of product display pictures relies on manual production, resulting in high costs, high repetition rates, uneven quality, lack of targeting, and unable to meet the precise needs of the e-commerce field, resulting in user fatigue and weak merchant supply.

Method used

By obtaining inquiry dialogue data and product information of different groups of people, using the decision factor generation model to mine decision factors, construct a population-decision factor knowledge base, and generating personalized images based on the target literary graph model, including preprocessing, image generation prompts and image template selection, and using the image quality audit model and click-through rate prediction model to optimize the image generation process.

Benefits of technology

It realizes low repetition rate and high-quality personalized image generation, improves buyer and seller matching efficiency, shortens the design cycle, and improves user experience and merchant creative production effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563191A_ABST
    Figure CN120563191A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, and relates to the technical field of artificial intelligence. The image generation method comprises the following steps: acquiring inquiry dialogue data of a plurality of users belonging to different crowd categories and corresponding commodity information; inputting the inquiry dialogue data and the corresponding commodity information and decision factor generation prompts into a decision factor generation model to generate decision factors of different crowd types and construct a crowd-decision factor knowledge base; extracting a target decision factor from the crowd-decision factor knowledge base according to the target crowd category of the to-be-processed commodity; and inputting the commodity information of the to-be-processed commodity and the target decision factor to a pre-trained target text graph model to generate a target image. According to the inquiry dialogue data of different users, the mining decision factor and the target decision factor corresponding to the to-be-processed commodity, the image is generated by using the target text graph model, crowd-differentiated image production is realized, the repetition rate is low, and the quality is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to an image generation method and device. Background Art

[0002] At present, with the rapid development of e-commerce, product display pictures play a vital role in attracting consumers' attention and improving purchase conversion rates.

[0003] Traditionally, product display images are mostly manually produced, resulting in high labor costs and inconsistent quality across images, headlines, and points of interest. Furthermore, manual production can lead to a high rate of repetition. According to some studies, manual creative production results in a repetition rate exceeding 70%. Highly homogenized content within the same industry can lead to user fatigue and negatively impact conversion rates.

[0004] Some research uses generative AI (Artificial Intelligence Generated Content, or AIGC) to generate product display images to reduce costs. However, generative AI still relies on a library of existing content, resulting in high duplication rates, and existing models lack specificity for the e-commerce sector, resulting in low-quality display images. Summary of the Invention

[0005] Based on this, the present application provides an image generation method and device to achieve the generation of images with low repetition rate and high quality.

[0006] According to one aspect of the present application, an image generation method is proposed, comprising: obtaining inquiry conversation data of multiple users belonging to different population categories and their corresponding product information, wherein the product information includes at least a product image, a product title, and CPV information; inputting the inquiry conversation data and its corresponding product information and a pre-constructed decision factor generation prompt into a decision factor generation model to generate decision factors for different population categories; constructing a population-decision factor knowledge base based on the decision factors and their corresponding populations; extracting a target decision factor from the population-decision factor knowledge base based on the target population category of the product to be processed; and inputting the product information and the target decision factor of the product to be processed into a pre-trained target text graph model to generate a target image.

[0007] According to some embodiments, the product information and target decision factors of the product to be processed are input into a pre-trained target text graph model to generate a target image, including: based on a pre-built prompt word template, according to the CPV information and target decision factors of the product to be processed, obtaining a prompt generation instruction; inputting the prompt generation instruction into the pre-trained text graph model to generate an image generation prompt; inputting the product image of the product to be processed and the image generation prompt into the pre-trained target text graph model to obtain the target image.

[0008] According to some embodiments, the product information and target decision factors of the product to be processed are input into a pre-trained target cultural graph model to generate a target image, including: selecting a target image template from a pre-built image template library based on the product to be processed; and inputting the product image, target decision factors and target image template of the product to be processed into the pre-trained target cultural graph model to obtain the target image.

[0009] According to some embodiments, the image generation prompt includes a model information prompt and / or a raw image background prompt and / or a decision factor template position prompt; the model information prompt includes a model gender information prompt, a model age information prompt and / or a model posture information prompt.

[0010] According to some embodiments, before the product information and target decision factors of the product to be processed are input into a pre-trained target text graph model to generate a target image, it also includes: obtaining multiple input data samples; inputting each input data sample into a pre-selected first preset number of image generation models, generating a second preset number of image generation results based on each image generation model, thereby obtaining a third preset number of image generation results corresponding to each input data sample, wherein the third preset number is the product of the first preset number and the second preset number; scoring the image generation results, and statistically calculating the distribution of the image generation results based on the scoring results; determining a first scoring threshold and a second scoring threshold based on the distribution of the image generation results; selecting a high-quality sample-low-quality sample pair from the third preset number of image generation results corresponding to each input data sample based on the first scoring threshold and the second scoring threshold; constructing a training sample based on each input data sample and its corresponding high-quality sample-low-quality sample pair, thereby obtaining a training data set; and training the target text graph model based on the training data set.

[0011] According to some embodiments, after inputting the product information and target decision factors of the product to be processed into a pre-trained target text graph model to generate a target image, it also includes: inputting the target image into a pre-trained image quality review model to obtain a quality review result.

[0012] According to some embodiments, the image quality audit model is pre-trained through the following steps: obtaining a set of positive and negative image data samples, wherein the set of positive and negative image data samples includes a plurality of positive image data samples whose image quality meets preset requirements and a plurality of negative image data samples whose image quality does not meet preset requirements; constructing an initial audit model based on the Transformer encoder and MLP; training the initial audit model based on the set of positive and negative image data samples until a preset end condition is met, and outputting the training result as the image quality audit model.

[0013] According to some embodiments, after inputting the product information and target decision factors of the product to be processed into a pre-trained target text-based graph model to generate a target image, it also includes: inputting the product information and target image of the product to be processed into a pre-built precise ranking click-through rate prediction model to obtain an estimated click-through rate; and sorting and displaying the products to be processed according to the estimated click-through rate.

[0014] According to some embodiments, after sorting and displaying the products to be processed according to the estimated click-through rate, it also includes: obtaining the actual click-through rate of the products to be processed displayed based on the target image; comparing the actual click-through rate with the estimated click-through rate, and fine-tuning the target text image model based on the comparison result.

[0015] According to some embodiments, the precise ranking click-through rate prediction model includes: a first feature extraction module, used to extract product features and / or user features of the product to be processed to obtain a first feature; a second feature extraction module, used to extract the creative features of the target image to obtain a second feature; a feature fusion module, used to fuse the first feature and the second feature to obtain a fused feature; and a prediction output module, used to map the fused feature to an estimated click-through rate.

[0016] According to one aspect of the present application, an image generation device includes: a basic data acquisition module for acquiring inquiry conversation data of multiple users belonging to different population categories and their corresponding product information, wherein the product information includes at least product images, product titles and CPV information; a decision factor generation module for inputting the inquiry conversation data and its corresponding product information and pre-built decision factor generation prompts into a decision factor generation model to generate decision factors for different population categories; a knowledge base construction module for constructing a population-decision factor knowledge base based on decision factors and their corresponding populations; a decision factor extraction module for extracting target decision factors from the population-decision factor knowledge base based on the target population category of the product to be processed; and an image result generation module for inputting the product information and target decision factors of the product to be processed into a pre-trained target text graph model to generate a target image.

[0017] According to one aspect of the present application, an electronic device is proposed, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.

[0018] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.

[0019] Through the above-described embodiments provided by this application, product preference differences among users of different demographic groups are obtained based on inquiry conversation data. Decision factors for these different demographic groups are then mined using a decision factor generation model to form a demographic-decision factor knowledge base. Target decision factors are extracted based on the target demographic of the product being processed. A target text-to-image model is then used to generate text-to-image, achieving demographic-differentiated image production with low duplication and high quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application.

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without exceeding the scope of protection required by this application.

[0022] Figure 1 A flowchart of the image generation method provided in an embodiment of the present application;

[0023] Figure 2 One of the flow charts of the embodiment of the present application for inputting the commodity information and target decision factors of the commodity to be processed into a pre-trained target text graph model to generate a target image;

[0024] Figure 3 Flowchart of target image generation for the four links of the target cultural graph model provided in the embodiment of the present application;

[0025] Figure 4 The second flowchart of the embodiment of the present application provides inputting the commodity information and target decision factors of the commodity to be processed into the pre-trained target text graph model to generate the target image;

[0026] Figure 5 A flowchart of training the target text graph model provided in an embodiment of the present application;

[0027] Figure 6 A flowchart of the image quality review model training provided in an embodiment of the present application;

[0028] Figure 7 A flowchart of the method for predicting the estimated click-through rate and using the estimated click-through rate to sort and display products provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of the network architecture of the precise ranking click-through rate prediction model provided in an embodiment of the present application;

[0030] Figure 9 A block diagram of an image generation device provided in an embodiment of the present application;

[0031] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0033] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0034] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0035] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0036] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below could be referred to as the second component without departing from the teachings of the present invention. As used herein, the term "and / or" includes any one and all combinations of one or more of the associated listed items.

[0037] Currently, the production of product display images on e-commerce platforms faces multiple contradictions. From the perspective of the buyer experience, e-commerce users face a conflict between precise demand and information obscurity, with genuine buyers' needs often drowned out by marketing noise. Statistics show that manual creative production results in a duplication rate of over 70% in advertising materials. The high homogeneity of content within the same industry leads to user fatigue, and users demand a richer creative element to enhance their shopping experience. From the perspective of merchant supply, most merchants have strong offline production capabilities but lack online marketing capabilities, resulting in limited ability to express and produce products. This leads to a weak content supply and uneven quality of materials such as images, headlines, and benefit points. From the perspective of platform commercialization, most platforms currently lack industry-specific and refined creative production capabilities, remaining focused on the "person to product" matching logic and lacking the real-time dynamic matching capabilities of "person to creativity."

[0038] Based on this, the present application proposes an image generation method and device, and the specific implementation method can refer to the following embodiments.

[0039] Figure 1 This is a flow chart of the image generation method provided in the embodiment of the present application. Figure 1 As shown, the method includes steps S110 to S150.

[0040] In step S110 , inquiry dialogue data of multiple users belonging to different population categories and their corresponding product information are obtained, wherein the product information at least includes product images, product titles and CPV information.

[0041] The inquiry conversation data obtained is the inquiry conversation data of multiple users. The inquiry conversation data is specifically a complete record of communication between users (buyers) and sellers about a certain product through channels such as email, online chat, and telephone. It covers user needs, product specifications, price negotiations, etc., which can reflect the user's concerns and be used to explore the user's product preferences.

[0042] In order to increase the applicability of the constructed crowd-decision factor knowledge base, the multiple users obtained need to belong to different crowd categories.

[0043] According to an example embodiment, the population categories include category B buyers, category C buyers, cross-border buyers, and the like.

[0044] While obtaining the inquiry conversation data, it is also necessary to obtain the product information corresponding to the inquiry conversation data. The product information includes but is not limited to the product image, product title and product CPV information.

[0045] It's worth noting that CPV (Category-Property-Value) is a system for structured description and management of product information, primarily used in e-commerce. Its core goal is to achieve standardization, searchability, and intelligent application of product information through the hierarchical definition of product categories, properties, and values. CPV information includes information about a product's category, properties, and values.

[0046] In step S120 , the inquiry dialogue data and its corresponding product information and the pre-built decision factor generation prompt are input into a decision factor generation model to generate decision factors for different population categories.

[0047] The decision factor generation model is a neural network model with textual reasoning capabilities. This application does not limit the specific training steps for the decision factor generation model. A commonly used neural network model with textual reasoning capabilities can be used. Alternatively, a neural network can be trained using inquiry conversation data labeled with decision factors and their corresponding product information sample data to obtain a decision factor generation model. Alternatively, a commonly used neural network model with textual reasoning capabilities can be fine-tuned using inquiry conversation data labeled with decision factors and their corresponding product information sample data to obtain a decision factor generation model.

[0048] By utilizing the reasoning ability of the decision factor generation model, decision factor mining is performed on the inquiry dialogue data of different users and the corresponding product information.

[0049] According to an example embodiment, the decision factor generation model generates decision factors for different population categories by performing OCR text extraction on product images in product information and performing structured understanding on product titles and CPV information in the product information.

[0050] According to an example embodiment, a prompt for generating decision factors is as follows: read a buyer-seller inquiry dialogue to identify the buyer's decision factors and the merchant's acceptance capacity, and output a message that satisfies the JSON format "Merchant Acceptance Capability": {"Large quantity and good price":"","Large quantity in stock":"","Quick delivery":""Support customization":"","Support sampling":"","Drop shipping":"Cross-border preferred":"","Support invoices":"","Complete qualifications":""Source factory":""","Downstream hot selling":""").

[0051] It should be noted that JSON (JavaScript Object Notation) is a lightweight data exchange format.

[0052] Decision factors include the product features that buyers in the corresponding population category value most when making purchasing decisions, such as product functions, product marketing, product services, etc.

[0053] In step S130 , a population-decision factor knowledge base is constructed based on the decision factors and their corresponding populations.

[0054] In step S140 , target decision factors are extracted from a crowd-decision factor knowledge base according to the target crowd category of the product to be processed.

[0055] The crowd category of the product to be processed is recorded as the target crowd category. The decision factor corresponding to the target crowd category is extracted from the crowd-decision factor knowledge base as the target decision factor.

[0056] In step S150 , the product information of the product to be processed and the target decision factor are input into a pre-trained target context graph model to generate a target image.

[0057] The target text graph model is a model that can generate corresponding images based on natural language descriptions. Its core is to convert text semantics into visual expressions through cross-modal understanding and generation capabilities. This application does not limit the specific training steps of the target text graph model. A commonly used neural network model that can generate corresponding images based on natural language descriptions can be used; data that establishes an association between the generated image and the product information of the product to be processed and its target decision factors can be used to train the neural network to obtain the target text graph model; data that establishes an association between the generated image and the product information of the product to be processed and its target decision factors can be used to fine-tune a commonly used neural network model that can generate corresponding images based on natural language descriptions to obtain the target text graph model.

[0058] The product information of the product to be processed and the target decision factor of the product to be processed are used as the input of the target cultural graph model to generate the target image, and the output is a high-quality image that matches the target population category of the product to be processed.

[0059] This application uses inquiry conversation data from users in different demographic groups to identify differences in product preferences. This approach then uses a decision factor generation model to mine decision factors for these demographic groups, forming a demographic-decision factor knowledge base. Target decision factors are extracted based on the target demographic of the product being processed. Using a target text-to-image model, text-to-image generation is performed, enabling the production of demographically differentiated images with low duplication and high quality. This significantly shortens the product image design and production cycle, improving buyer-seller matching efficiency.

[0060] According to the exemplary embodiment, step S150 is primarily divided into two categories during the AI ​​image generation process: the apparel industry and industries other than the apparel industry. This application will further describe the specific execution details of step S150 under the above two industry categories in subsequent embodiments.

[0061] According to some embodiments, reference Figure 2 In step S150, the product information and target decision factors of the product to be processed are input into the pre-trained target cultural graph model to generate a target image, which can be specifically achieved through steps S210 to S230.

[0062] In step S210, based on the pre-built prompt word template, according to the CPV information of the commodity to be processed and the target decision factor, a prompt generation instruction is obtained.

[0063] According to the CPV information of the product to be processed (such as brand, color, style, decoration, etc.) and the target decision factor, a prompt generation instruction is formed in combination with the prompt word template.

[0064] According to an exemplary embodiment, the prompt word template includes the answer identity, the specific content of the thought chain, the specific requirements of the instruction, context information, precautions, etc.

[0065] Furthermore, in order to improve the quality of the generated image, before step S210 , the process further includes: pre-processing the product image of the product to be processed.

[0066] In the case where the product image is an image of a clothing model (ie, an image of clothing with a portrait), the pre-processing step includes: first performing portrait segmentation, and then performing clothing segmentation.

[0067] The content of portrait segmentation is to replace the portrait part with a white border. The content of clothing segmentation is to separate the clothing from the background.

[0068] When the product image is a flat clothing image (ie, a clothing image without a portrait), the pre-processing step includes: performing clothing segmentation.

[0069] When the product image is a flat product image (ie, a non-clothing image without a portrait), the pre-processing step includes: performing product segmentation. Specifically, the product segmentation involves separating the product from the background.

[0070] This application does not specifically limit the execution module of preprocessing, and appropriate image processing technology can be selected according to actual conditions.

[0071] In step S220 , the prompt generation instruction is input into the pre-trained Wenshengwen model to generate an image generation prompt.

[0072] The Vincent model is a model that can generate text output that meets specific requirements based on input text. This application does not limit the specific training steps of the Vincent model. A commonly used neural network model that can generate text output that meets specific requirements based on input text can be used; text sample data labeled with text output that meets specific requirements can also be used to train the neural network to obtain the Vincent model; text sample data labeled with text output that meets specific requirements can also be used to fine-tune a commonly used neural network model that can generate text output that meets specific requirements based on input text to obtain the Vincent model applicable to this step.

[0073] For image generation in the apparel industry, according to some embodiments, image generation prompts include model information prompts and / or raw image background prompts; model information prompts include model gender information prompts, model age information prompts and / or model posture information prompts.

[0074] The image generation prompts generated based on the Wenshengwen model include model posture prompts of different genders and ages and raw image background prompts.

[0075] For image generation in non-apparel industries, the Wenshengwen model can utilize a deep learning-based natural language processing model. Specifically, for non-apparel products, CPV information and target decision factors are obtained, combined with prompt word templates to form prompt generation instructions. These instructions are then fed into the deep learning-based natural language processing model to generate personalized product image prompts (i.e., image generation prompts).

[0076] Furthermore, in order to improve the quality of image generation prompts and generate personalized product image prompts, in some embodiments, a natural language processing model based on deep learning is used to further optimize the image generation prompts generated by using a pre-trained text-generated text model and combining it with a prompt word template.

[0077] In step S230 , the product image of the product to be processed and the image generation prompt are input into the pre-trained target text-based graph model to obtain the target image.

[0078] To improve efficiency, in some embodiments, reference Figure 3 For the apparel industry, the target model's model image, pre-processed product images, product images of the product to be processed, and image generation prompts can be simultaneously input into the target cultural graph model to obtain the target image. For non-apparel industries, the target image can be simultaneously input into the target cultural graph model to obtain the target image.

[0079] Furthermore, in order to explain the prompt word template in more detail, a specific embodiment is given as follows:

[0080] The prompt word template includes:

[0081] "You are a prompt engineer and an advertising and commercial photography expert. When I tell you a product [keyword], you will modify and generate a paragraph based on the content of [keyword], recorded as X.

[0082] Among them, the format of the input [keyword] is as follows: {'cate3_name': third-level category name, 'ad_color': color, 'ad_fengge': style, 'ad_xiushi': modifier}.

[0083] These all describe some characteristics of the product. Please refer to these characteristics, think and associate step by step, and generate description text X that matches these product characteristics according to the requirements below.

[0084] X will be used to present a description of the background image of a commercial photography for an advertisement. This is called a description word. X has a fixed parameter format: product photography, solo, [1], [2], [3], [4], minimalist photography, consensus background, highlighting the product, masterpiece, best quality, shot by Sony Alpha A7 III.

[0085] Among them, X's content always starts with the descriptor 'product photography, solo,' and ends with 'minimalist photography, consensus background, highlighting the product, masterpiece, best quality, shot by Sony Alpha A7 III.

[0086] [1] It should be a color and shooting style that matches the [keyword]. It can be a specific photographic device, a photographer, or an abstract synonym, such as the style of a famous movie or TV series.

[0087] [2] It should be a descriptive word about the scene or background of the advertisement that fits the theme and style of the [Keyword]. It should be a detailed and specific description of the background, rich in scene elements and details. The description of the scene should be no less than 15 words.

[0088] [3] is the main visual shot and is a commonly used descriptive technique of photographic angles.

[0089] [4] is the detailed information of the overall style. It is a collection of ideas for a product advertising picture and should be expressed in no less than 10 words that are suitable for the main picture and style.

[0090] ###

[0091] Notice:

[0092] 1. In each parameter [1][2][3][4], there should be no descriptive words related to the product referred to by the [keyword], nor should there be any physical objects similar to the [keyword]. Instead, the description should focus only on the background scene and shooting style.

[0093] 2. In each parameter, there should be no descriptions of people or content related to people, nor should there be any human body parts.

[0094] 3. The generated X content must be presented in English.

[0095] 4. Do not repeat the descriptions in [Keywords]. Try to associate related scenes based on these descriptions.

[0096] ###

[0097] If you understand the above description, here are the keywords to enter:

[0098] "['cate3_name':${itm_cate3_name],'ad_color':$(ad_color),'ad_fengge':${ad_fengge},'ad_xiushi${ad_xiushi}}"

[0099] Just return the generated X without any other content.

[0100] ”

[0101] The image generation method provided in this application can be used to generate personalized background images, maintaining the consistency of the generated image and the original image.

[0102] According to some embodiments, reference Figure 4 In step S150, the product information and target decision factors of the product to be processed are input into the pre-trained target cultural graph model to generate a target image, which can be specifically achieved through steps S410 to S420.

[0103] In step S410 , a target image template is selected from a pre-built image template library according to the product to be processed.

[0104] For non-apparel industries, target images can also be generated by combining raw product images with product templates. First, the target image template must be selected. The image template library contains pre-created industry-specific templates created by designers. The template that matches the product being processed is selected as the target image template.

[0105] The selection process can be achieved by training the model through keyword matching, or other methods can be used, which are not limited in this application.

[0106] In step S420 , the product image of the product to be processed, the target decision factor, and the target image template are input into the pre-trained target text graph model to obtain the target image.

[0107] Then, products are populated based on the target template. The target text graph model follows the target image template's image layout and selling point text, using the product images of the target product and the target decision factors to populate the composite image with different demographic preferences, achieving differentiated selling points for each demographic in image generation.

[0108] In some embodiments, reference Figure 3 In AI image production, the target cultural graph model can be divided into four stages: clothing model image generation, flattened clothing image generation, (non-clothing) product image generation, and (non-clothing) product template image generation. Clothing model image generation, flattened clothing image generation, and product image generation can be implemented using steps S210-S230, while product template image generation can be implemented using steps S410-S420.

[0109] According to some embodiments, reference Figure 5 Before step S120, steps S510 to S570 are also included.

[0110] In step S510 , a plurality of input data samples are obtained.

[0111] It should be noted that the purpose of steps S510 to S570 is to generate images by using multiple raw image models including the model to be trained, and to align the raw image model indicators in combination with the CTR (Click-Through Rate) reward model and various reinforcement learning schemes, so as to obtain a high-performance (especially high prior CTR) target raw image model without the need for complex parameter training.

[0112] On this basis, the specific content of the input data sample can be selected according to the actual situation. The input data sample can include the product information and the corresponding target decision factor in step S150, or Figure 3The following information may be included: prompt (image generation prompt), original image (product image of the product to be processed) and mask image (pre-processed product image); it may also include only product images or natural language prompt text, which is not limited in this application.

[0113] The input data samples are the basis for constructing direct preference optimization data (DPO).

[0114] In step S520, each input data sample is input into a pre-selected first preset number of image generation models, and a second preset number of image generation results are generated based on each image generation model, thereby obtaining a third preset number of image generation results corresponding to each input data sample, wherein the third preset number is the product of the first preset number and the second preset number.

[0115] During the DPO data construction process, a first preset number of image generation models are first used to generate images. The first preset number of image generation models includes a model to be trained according to the present application and multiple reference models.

[0116] Each model generates a second preset number of images as image generation results, that is, after each input data sample is input, a third preset number of different images will be generated.

[0117] The first preset number is recorded as M, the second preset number is recorded as N, and the third preset number is M×N.

[0118] In step S530 , the image generation results are scored, and the distribution of the image generation results is statistically analyzed based on the scoring results.

[0119] Score each image generation result, collect statistics on the scores of all generated images, remove abnormal samples, and then analyze their distribution.

[0120] To improve the interpretability of the steps, while scoring each image generation result, the seed and input corresponding to the image generation result can also be recorded.

[0121] In step S540 , a first scoring threshold and a second scoring threshold are determined according to the distribution of the image generation results.

[0122] Based on the distribution of scores, two thresholds are set: a first score threshold and a second score threshold.

[0123] The first scoring threshold is the threshold for selecting high-quality items. Images exceeding this threshold will be marked as high-quality samples. The second scoring threshold is the threshold for selecting low-quality items. Images below this threshold will be marked as low-quality samples.

[0124] According to an example embodiment, the first scoring threshold is a value at which the linear interpolation falls within the top 10%.

[0125] According to an example embodiment, the second scoring threshold is a value at which the last 30% of the linear interpolation values ​​lie.

[0126] In step S550 , a high-quality sample-low-quality sample pair is selected from a third preset number of image generation results corresponding to each input data sample according to the first scoring threshold and the second scoring threshold.

[0127] According to the set first scoring threshold and second scoring threshold, from the output generated by the input of each unified input data sample, a pair that meets the standards is constructed based on the high-quality samples and the low-quality samples to obtain a high-quality sample-low-quality sample pair.

[0128] In step S560, a training sample is constructed based on each input data sample and its corresponding high-quality sample-low-quality sample pair, thereby obtaining a training data set.

[0129] The constructed training dataset is DPO data.

[0130] In step S570, a target text graph model is trained based on the training data set.

[0131] Using DPO data and based on the Diff-DPO strategy, the above-mentioned pair was used to align the model to be trained and the reference model to obtain the target text graph model.

[0132] Naturally, based on the training dataset obtained in the above steps, the Diff-DPO strategy can be replaced by other common training algorithms to train the target text graph model.

[0133] According to some embodiments, after step S150 , step S160 is further included.

[0134] In step S160 , the target image is input into a pre-trained image quality audit model to obtain a quality audit result.

[0135] In order to save labor audit costs, an image quality audit model is used to judge the image quality, identify various complex problems in the generated images, and obtain quality audit results.

[0136] The quality review result may include a score for the target image, or may directly obtain a result of whether the target image passes (ie, the image quality meets the requirements) or fails (ie, the image quality does not meet the requirements) through a set score threshold.

[0137] According to some embodiments, reference Figure 6The training steps of the image quality audit model in step S160 can be specifically implemented through steps S610 to S630.

[0138] In step S610, a positive and negative image data sample set is obtained, wherein the positive and negative image data sample set includes a plurality of positive image data samples whose image quality meets the preset requirements and a plurality of negative image data samples whose image quality does not meet the preset requirements.

[0139] The image positive and negative data sample sets mentioned in the steps are high-quality and low-quality image positive and negative data sample sets constructed based on the model's understanding of the image problem. They can be directly obtained by using existing image datasets, directly annotating existing image data, or using image generation models, and this application does not impose any restrictions on this.

[0140] According to an exemplary embodiment, the plurality of positive image data samples whose image quality meets the preset requirements are image samples with normal picture logic, normal cutout, normal portrait, and normal text. The plurality of negative image data samples whose image quality does not meet the preset requirements are image samples with abnormal picture logic, abnormal cutout, deformed portrait, and abnormal text.

[0141] In step S620, an initial audit model is constructed based on the Transformer encoder and MLP.

[0142] The basic architecture of the image quality review model is constructed using the Transformer encoder and MLP, which is denoted as the initial review model.

[0143] Among them, MLP (Multilayer Perceptron) is an artificial intelligence model based on feedforward neural network.

[0144] Furthermore, to further improve the model's audit accuracy, an image segmentation module can be set before the Transformer encoder. In other words, the initial audit model is constructed using the ViT (Vision Transformer) + MLP model structure.

[0145] Among them, ViT is an image classification model based on the Transformer architecture.

[0146] In step S630, the initial audit model is trained based on the positive and negative image data sample sets until a preset end condition is met, and the training result is output as the image quality audit model.

[0147] This application does not limit the specific training algorithm of the image quality review model, and an appropriate training algorithm can be selected according to actual conditions.

[0148] Similarly, the preset end condition may be a threshold value of a certain performance indicator, and this application does not impose any specific limitation on this.

[0149] The trained image quality review model is applied online to filter low-quality AIGC images efficiently and at low cost.

[0150] In actual application, the target image is input into a pre-trained image quality review model. It is first globally encoded through the trained Transformer encoder, and then the global features are gradually fused as output through the self-attention mechanism. The category token vector in the Transformer output sequence is extracted as the input of the trained MLP. After linear mapping, the image quality review result is output to achieve efficient classification.

[0151] According to the exemplary embodiment, the recall rate of the trained image quality audit model reaches 85.5% and the accuracy rate reaches 83.9%.

[0152] According to some embodiments, reference Figure 7 After step S150, the process also includes predicting the estimated click-through rate and using the estimated click-through rate to sort and display products, which is specifically implemented through steps S710 to S720.

[0153] Step S710: Input the product information and target image of the product to be processed into a pre-built precise ranking click rate prediction model to obtain an estimated click rate.

[0154] Step S720: Sort and display the products to be processed according to the estimated click-through rates.

[0155] Specifically, when the generated target image is used in e-commerce search and recommendation scenarios, creative optimization can also be integrated into the fine ranking module, and the creative scoring directly affects the ranking results of the advertisement.

[0156] This application constructs a precise ranking click-through rate prediction model to sort advertisements. In step S710, the constructed precise ranking click-through rate prediction model is used to predict the click-through rate of the product to obtain the estimated click-through rate. In step S720, the products are sorted and displayed according to the predicted estimated click-through rate. The overall merchant creative production-advertising creative delivery is fully automated, which shortens the design and delivery cycle of product creativity, greatly improves the merchant creative production and advertising creative delivery effects, and realizes real-time dynamic matching of people and creativity at the traffic delivery end.

[0157] According to some embodiments, after step S720, steps S730 to S740 are further included.

[0158] In step S730, the actual click rate of the product to be processed displayed based on the target image is obtained.

[0159] Get the final real online click-through rate.

[0160] In step S740 , the actual click-through rate is compared with the estimated click-through rate, and the target document graph model is fine-tuned according to the comparison result.

[0161] The final online click-through rate effect is fed back to image production, the actual click-through rate is compared with the estimated click-through rate, and the target text-based image model is optimized based on the difference.

[0162] According to some embodiments, reference Figure 8 , the CTR prediction model of precise ranking includes:

[0163] The first feature extraction module is used to extract product features and / or user features of the product to be processed to obtain a first feature.

[0164] According to an example embodiment, the first feature may include user, billboard, and context features. The first feature extraction module may use an existing click rate prediction model, such as Figure 8 As shown, the first feature extraction module (i.e. Figure 8 The network architecture of the CTR network tower in the image can include multiple FCNs (Fully Convolutional Networks).

[0165] The second feature extraction module is used to extract the creative features of the target image to obtain the second feature.

[0166] According to example embodiments, the second feature may include a creative ID, a creative content feature, and the like.

[0167] The Creative ID is a unique symbol system used to identify the characteristics of creative content, such as the ID number of a pending product idea. Creative content characteristics include the target image, title, and selling points that are used to attract users and convey information.

[0168] The second feature extraction module (i.e. Figure 8 The creative network tower includes a neural network architecture that can extract text features and image features. This application does not limit its specific network architecture.

[0169] The feature fusion module is used to fuse the first feature and the second feature to obtain a fused feature.

[0170] The prediction output module is used to map the fused features into the estimated click-through rate.

[0171] The Logits (the original output value of the last layer of the neural network) of the final refined click-through rate prediction model can be mapped to the estimated click-through rate through a preset mapping relationship.

[0172] The following describes an apparatus embodiment of the present application, which can be used to perform the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, reference can be made to the method embodiment of the present application.

[0173] Figure 9 A block diagram illustrating an image generating apparatus according to an exemplary embodiment is shown.

[0174] Figure 9 The device shown can execute the aforementioned image generation method according to the embodiment of the present application.

[0175] like Figure 9 As shown, the image generating device may include:

[0176] See also Figure 9 And referring to the above description, the basic data acquisition module 910 is used to obtain inquiry dialogue data of multiple users belonging to different population categories and their corresponding product information, wherein the product information at least includes product images, product titles and CPV information.

[0177] The decision factor generation module 920 is used to input the inquiry dialogue data and its corresponding product information and pre-built decision factor generation prompts into the decision factor generation model to generate decision factors for different population categories.

[0178] The knowledge base construction module 930 is used to construct a population-decision factor knowledge base based on decision factors and their corresponding populations.

[0179] The decision factor extraction module 940 is used to extract target decision factors from the population-decision factor knowledge base according to the target population category of the product to be processed.

[0180] The image result generation module 950 is used to input the product information of the product to be processed and the target decision factor into the pre-trained target text graph model to generate a target image.

[0181] The device performs functions similar to the method provided above. For other functions, please refer to the previous description and will not be repeated here.

[0182] Figure 10 An electronic device according to an exemplary embodiment of the present application is shown. Figure 10 1000 according to this embodiment of the present application will be described. Figure 10 The electronic device 1000 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0183] like Figure 10As shown, electronic device 1000 is implemented as a general-purpose computing device. Components of electronic device 1000 may include, but are not limited to, at least one processing unit 1010, at least one storage unit 1020, a bus 1030 connecting various system components (including storage unit 1020 and processing unit 1010), a display unit 1040, and the like.

[0184] The storage unit stores program codes, which can be executed by the processing unit 1010, so that the processing unit 1010 performs the methods described in this specification according to various exemplary embodiments of the present application. For example, the processing unit 1010 can perform the aforementioned methods.

[0185] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 10201 and / or a cache memory unit 10202 , and may further include a read-only memory unit (ROM) 10203 .

[0186] The storage unit 1020 may also include a program / utility 10204 having a set (at least one) of program modules 10205, such program modules 10205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0187] Bus 1030 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0188] The electronic device 1000 can also communicate with one or more external devices 300 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1000, and / or any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1060. The network adapter 1060 can communicate with other modules of the electronic device 1000 via the bus 1030. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 1000, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0189] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software or by combining software with necessary hardware. The technical solution according to the embodiment of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present application.

[0190] The software product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0191] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0192] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0193] The computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the computer-readable medium implements the aforementioned functions.

[0194] Those skilled in the art will appreciate that the modules described above can be distributed in the device according to the description of the embodiment, or can be modified accordingly to be used in one or more devices that are different from the embodiment. The modules of the above embodiment can be combined into one module or further divided into multiple submodules.

[0195] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0196] While the exemplary embodiments of the present application have been specifically illustrated and described above, it should be understood that the present application is not limited to the detailed structures, configurations, or implementations described herein; rather, the present application is intended to encompass various modifications and equivalent configurations within the spirit and scope of the appended claims.

Claims

1. An image generation method, characterized in that: include: Obtain inquiry conversation data of multiple users belonging to different demographic categories and their corresponding product information, wherein the product information at least includes product images, product titles, product descriptions, and CPV information; Inputting the inquiry dialogue data and its corresponding product information and pre-built decision factor generation prompts into a decision factor generation model to generate decision factors for different population categories; Constructing a population-decision factor knowledge base based on the decision factors and their corresponding populations; Extracting target decision factors from the population-decision factor knowledge base according to the target population category of the product to be processed; The commodity information of the commodity to be processed and the target decision factor are input into a pre-trained target text graph model to generate a target image.

2. The method according to claim 1, characterized in that The step of inputting the commodity information of the commodity to be processed and the target decision factor into a pre-trained target text graph model to generate a target image includes: Based on a pre-built prompt word template, according to the CPV information of the commodity to be processed and the target decision factor, a prompt generation instruction is obtained; inputting the prompt generation instruction into a pre-trained Wenshengwen model to generate an image generation prompt; The product image of the product to be processed and the image generation prompt are input into a pre-trained target text-based graph model to obtain the target image.

3. The method according to claim 1, characterized in that The step of inputting the commodity information of the commodity to be processed and the target decision factor into a pre-trained target text graph model to generate a target image includes: Selecting a target image template from a pre-built image template library according to the product to be processed; The product image of the product to be processed, the target decision factor and the target image template are input into a pre-trained target text graph model to obtain a target image.

4. The method according to claim 2, characterized in that The image generation prompt includes a model information prompt and / or a raw image background prompt and / or a decision factor template position prompt; The model information prompt includes a model gender information prompt, a model age information prompt and / or a model posture information prompt.

5. The method according to claim 1, wherein Before inputting the commodity information of the commodity to be processed and the target decision factor into the pre-trained target cultural graph model to generate the target image, the method further includes: Get multiple input data samples; Inputting each of the input data samples into a preselected first preset number of image generation models, generating a second preset number of image generation results based on each of the image generation models, thereby obtaining a third preset number of image generation results corresponding to each of the input data samples, wherein the third preset number is the product of the first preset number and the second preset number; Scoring the image generation results, and calculating the distribution of the image generation results based on the scoring results; determining a first scoring threshold and a second scoring threshold according to the distribution of the image generation results; selecting, according to the first scoring threshold and the second scoring threshold, a high-quality sample-low-quality sample pair from a third preset number of image generation results corresponding to each of the input data samples; Constructing a training sample based on each of the input data samples and the corresponding high-quality sample-low-quality sample pair, thereby obtaining a training data set; Based on the training data set, the target text graph model is obtained through training.

6. The method according to claim 1, characterized in that After inputting the commodity information of the commodity to be processed and the target decision factor into a pre-trained target cultural graph model to generate a target image, the method further includes: The target image is input into a pre-trained image quality audit model to obtain a quality audit result.

7. The method according to claim 6, characterized in that The image quality audit model is pre-trained through the following steps: Acquire a set of positive and negative image data samples, wherein the set of positive and negative image data samples includes a plurality of positive image data samples whose image quality meets preset requirements and a plurality of negative image data samples whose image quality does not meet the preset requirements; Build an initial audit model based on the Transformer encoder and MLP; Based on the positive and negative image data sample sets, the initial audit model is trained until a preset end condition is met, and the training result is output as the image quality audit model.

8. The method according to claim 1, characterized in that After inputting the commodity information of the commodity to be processed and the target decision factor into a pre-trained target cultural graph model to generate a target image, the method further includes: Inputting the product information of the product to be processed and the target image into a pre-built precise ranking click rate prediction model to obtain an estimated click rate; The products to be processed are sorted and displayed according to the estimated click rates.

9. The method according to claim 8, characterized in that After sorting and displaying the products to be processed according to the estimated click rates, the method further includes: Obtaining a real click rate of the product to be processed displayed based on the target image; The actual click-through rate is compared with the estimated click-through rate, and the target document graph model is fine-tuned according to the comparison result.

10. The method according to claim 8, characterized in that The precise ranking click rate prediction model includes: A first feature extraction module is used to extract product features and / or user features of the product to be processed to obtain a first feature; A second feature extraction module is used to extract the creative feature of the target image to obtain a second feature; a feature fusion module, configured to fuse the first feature and the second feature to obtain a fused feature; The prediction output module is used to map the fusion features into an estimated click-through rate.

11. An image generating device, characterized in that: include: A basic data acquisition module is used to obtain inquiry dialogue data of multiple users belonging to different population categories and their corresponding product information, wherein the product information at least includes product images, product titles, product descriptions and CPV information; A decision factor generation module, configured to input the inquiry dialogue data and its corresponding product information and pre-built decision factor generation prompts into a decision factor generation model to generate decision factors for different population categories; A knowledge base construction module, used to construct a population-decision factor knowledge base based on the decision factors and their corresponding populations; A decision factor extraction module is used to extract target decision factors from the population-decision factor knowledge base according to the target population category of the product to be processed; The image result generation module is used to input the product information of the product to be processed and the target decision factor into the pre-trained target cultural graph model to generate a target image.

12. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Model training method and device, video generation method and device, electronic equipment and storage medium

    CN121305435A

  • Voice-driven intelligent picture book generation method and device, electronic equipment and storage medium

    CN121393444A

  • Image defect detection method, electronic equipment, storage medium and product

    CN121746395A

  • Image defect detection methods, electronic devices, storage media and products

    CN121746395B