Image generation method and apparatus, and device and storage medium
By using food description information and training data corresponding to image samples in the image generation model, high-quality food images that accurately convey the raw and mature state of the goods are generated, solving the problem of inaccurate image generation in the prior art and improving the accuracy of user purchases.
Patent Information
- Application Number
- PCT/CN2024/117717
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art is difficult to generate high-quality product images that accurately convey the raw and mature status of the product, resulting in users being misled when purchasing.
By obtaining the prompt text containing food description information, input it to the pre-trained image generation model, a food image matching the target food name and raw and cooked state information is generated. The training samples of the image generation model include food description samples and corresponding food image samples, and the food description samples contain food name and raw and cooked state information.
It realizes the generation of high-quality food images with accurate raw and mature state, improving the accuracy and user experience of image generation.
Smart Images

Figure CN2024117717_19062025_PF_FP_ABST
Abstract
Description
Image generation method, device, equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 15, 2023, with application number 202311733895.5 and application name “Image generation method, device, equipment and storage medium”, the entire contents of which are incorporated by reference into the application. Technical Field
[0002] The present disclosure relates to the field of machine learning technology, and in particular to image generation methods, devices, equipment, and storage media. Background Art
[0003] In application scenarios like local life, the platform can provide merchants with store services, allowing them to manage their online stores within the platform. For example, merchants can display a variety of different products in their online stores. By uploading product photos and adding descriptions and specifications, merchants can showcase the product's features and advantages to users. Understandably, the quality of product images plays a crucial role. Product images must accurately convey product information to enhance user awareness and interest in the product. Therefore, generating high-quality product images has become a pressing technical challenge.
[0004] Application Contents
[0005] To overcome the problems existing in the related art, the present disclosure provides an image generation method, apparatus, device and storage medium.
[0006] According to a first aspect of the embodiments of this specification, there is provided an image generation method, the method comprising:
[0007] Obtaining a prompt text containing food description information, wherein the food description information includes a target food name and target rawness or cookedness status information indicating whether the food is raw or cooked;
[0008] The prompt text is input into a pre-trained image generation model, so that the image generation model generates a food image that matches the target food name and the target rawness / cooking status information based on the input prompt text; wherein the training samples of the image generation model include: food description samples and corresponding food image samples, and the food description samples include food name samples and rawness / cooking status information samples of the food in the food image samples;
[0009] Output the food image generated by the image generation model.
[0010] According to a second aspect of the embodiments of this specification, there is provided an image generation method, the method comprising:
[0011] Obtaining the target food name input by the user and sending it to the server, wherein the server is configured to obtain target rawness or cookedness information indicating whether the food is raw or cooked, and then executing the steps of the image generation method embodiment of the first aspect;
[0012] Obtain and display the food image output by the server.
[0013] According to a third aspect of the embodiments of this specification, there is provided an image generating apparatus, including:
[0014] an acquisition module, configured to acquire a prompt text containing food description information, wherein the food description information includes a target food name and target rawness or cookedness status information indicating whether the food is raw or cooked;
[0015] An image generation module is configured to input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image that matches the target food name and the target raw / cooked state information based on the input prompt text; wherein the training samples of the image generation model include: food description samples and corresponding food image samples, wherein the food description samples include food name samples and raw / cooked state information samples of the food in the food image samples;
[0016] An image output module is used to output the food image generated by the image generation model.
[0017] According to a third aspect of the embodiments of this specification, there is provided an image generating apparatus, the apparatus comprising:
[0018] An acquisition module is configured to: acquire the target food name input by the user and send it to a server, wherein the server is configured to acquire target rawness or cookedness information indicating whether the food is raw or cooked, and then execute the steps of the image generation method embodiment of the first aspect;
[0019] The display module is used to obtain and display the food image output by the server.
[0020] According to a fourth aspect of the embodiments of this specification, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the image generation method embodiment described in the first aspect are implemented.
[0021] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image generation method embodiment described in the first aspect are implemented.
[0022] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0023] In the embodiment of this specification, the training samples for training the image generation model include: food description samples and corresponding food image samples, wherein the food description samples include food name samples and rawness / cooking status information samples of the food in the food image samples; therefore, the image generation model is capable of generating food images with accurate rawness / cooking status. During inference, the model in this embodiment obtains prompt text containing food description information, wherein the food description information includes the target food name and target rawness / cooking status information indicating whether the food is raw or cooked; the prompt text is input into the pre-trained image generation model, so that the image generation model can generate a food image that matches the target food name and target rawness / cooking status information based on the input prompt text. Therefore, this embodiment is capable of generating high-quality food images with accurate rawness / cooking status.
[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the disclosure.
[0026] FIG1 is a schematic diagram of an application scenario according to an exemplary embodiment of this specification.
[0027] FIG2A is a flowchart of an image generating method according to an exemplary embodiment of this specification.
[0028] FIG2B is a schematic diagram of a training data set of an image generation model according to an exemplary embodiment of this specification.
[0029] FIG2C is an overall flowchart of image generation according to an exemplary embodiment of this specification.
[0030] FIG2D is a diagram of an application scenario of image generation according to an exemplary embodiment of this specification.
[0031] FIG3 is a flowchart of an image generating method according to an exemplary embodiment of this specification.
[0032] FIG4 is a block diagram of an image generating apparatus according to an exemplary embodiment of this specification.
[0033] FIG5 is a hardware structure diagram of a computer device where an image generating apparatus is located according to an exemplary embodiment of this specification. DETAILED DESCRIPTION
[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.
[0035] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. As used in this specification and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0036] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0037] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0038] As shown in Figure 1, this specification shows an application scenario diagram according to an exemplary embodiment, which includes a platform, merchants and users. Among them, the platform has built a service end and provides users with a client, through which users can use the services provided by the platform; different from the client for users, the platform also provides the store with a client for merchants, through which merchants can use the merchant services provided by the platform, such as store management services, where merchants can manage their own online store information, including store name, address, contact information, etc., through the platform; it can also include product display services, where merchants can display their own products on the platform, and show users the characteristics and advantages of the products by uploading product photos, adding descriptions and specifications, etc. In actual applications, it can also include order management services or delivery services, etc., which are not described in detail in this embodiment.
[0039] The applicant of this application has found that the quality of product images plays an important role, and product images need to be able to accurately convey product information in order to enhance users’ cognition and interest in the products. The study found that in some application scenarios, whether the product information can be accurately conveyed in the product image has a certain impact on users who want to purchase the product. For example, in the catering industry, the merchant’s products are food, and some foods exist in two states: cooked food that can be eaten directly and raw food prefabricated products that require subsequent processing, such as barbecue, dumplings and other prominent categories. For products in the above two states, if the raw and cooked food state of the product picture does not match the raw and cooked food state of the product itself, it will cause greater interference to the user’s purchase and judgment, for example:
[0040] (1) In the catering product scene, some categories of prepared food products are paired with pictures of raw food preparations;
[0041] For example, a barbecue shop sells pork belly skewers, which are cooked food—that is, grilled pork belly—but the images displayed in the shop are of raw pork belly skewers. Another shop sells Tomahawk lamb chops, which are raw and uncooked, but the images displayed are of cooked food. The rawness and cookedness information in these food images doesn't match the information displayed on the shop's website, misleading consumers. One reason for this problem is that merchants lack the ability to accurately produce images of raw and cooked food.
[0042] (2) Some tools for automatically synthesizing product images will select food images from the image library and fill them into the template based on the actual synthesis needs. Since some categories in the image library (such as barbecue and other related categories) contain a large amount of raw food in the product images, the food in the barbecue category product images automatically generated by the tool is in a raw state, which is inconsistent with the actual cooked state of barbecue products.
[0043] Based on this, the embodiment of this specification provides an image generation solution that can automatically generate food images with accurate raw and cooked status.
[0044] As shown in FIG2A , FIG2A is a flowchart of an image generation method according to an exemplary embodiment of this specification, which may include the following steps:
[0045] Step 202: Obtain prompt text containing food description information.
[0046] The food description information includes the target food name and target rawness or cookedness status information indicating whether the food is raw or cooked.
[0047] Step 204: input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image that matches the target food name and the target raw / cooked state information based on the input prompt text.
[0048] The training samples of the image generation model include: food description samples and corresponding food image samples, and the food description samples include food name samples and raw / cooked state information samples of the food in the food image samples.
[0049] Step 206: Output the food image generated by the image generation model.
[0050] The image generation method of this embodiment can be run on a computer device, including but not limited to a server, a cloud server, a server cluster, a tablet computer, a personal digital assistant, a laptop computer, or a desktop computer. In some examples, the model involved in the image generation method of this embodiment can be deployed on a server, for example, it can be deployed on a server connected to the merchant's client in Figure 1. The merchant's client can receive user input and send it to the server, which runs the image generation method of this embodiment and sends the result to the client. In other examples, the model involved in the image generation method of this embodiment can also be deployed on the merchant's client, and this embodiment does not impose any restrictions on this.
[0051] The raw and cooked status in this embodiment can be used to indicate whether the food has been cooked or heated. Among them, raw food can refer to ingredients or foods that have not been cooked or cooked and are in their original state. Raw food usually refers to fresh, unprocessed food in its original state, such as uncooked vegetables, meat, seafood, etc. It can also include foods that have been simply processed, such as pickled foods, etc. For example, in the catering industry, products such as lamb and shrimp paste are hot pot ingredients, and chicken breast in retail stores, etc. Cooked food can refer to ingredients or foods that have been properly cooked or heated, such as boiled, fried, roasted, steamed, and other cooking or heating methods.
[0052] In this embodiment, the image generation model can be pre-built, and the construction process can be executed on a single computer device or on multiple devices in a computer device cluster. A training dataset can be prepared in advance, and this embodiment does not limit the method for obtaining the training dataset. In this embodiment, the training samples are designed to include: food description samples and corresponding food image samples, where the food description samples include at least a food name sample and a raw / cooked state information sample.
[0053] The specific structure of the image generation model of this embodiment can be independently constructed as needed, and the solution of this embodiment can be used to improve an existing pre-trained image generation model. Optionally, the image generation model may include a generative adversarial network (GAN) model, a flow-based generation model, or a diffusion model, etc., which is not limited in this embodiment.
[0054] The food description sample may include additional information, which can be configured as needed in practical applications and is not limited in this embodiment. When the food description sample contains more information, the image generation model can obtain richer information related to the food image sample, thereby learning rich knowledge about the relationship between the food description sample and the food image sample, thereby improving the model's reasoning accuracy. In some examples, the food description information may also include the target food category; the food description sample may also include a food category information sample.
[0055] Inputting the prompt text into a pre-trained image generation model so that the image generation model generates a food image matching the target food name and the target cooked state information based on the input prompt text may include:
[0056] The prompt text is input into a pre-trained image generation model so that the image generation model generates a food image that matches the target food name, the target food category and the target raw / cooked state information based on the input prompt text.
[0057] In this embodiment, food category information samples of the category to which the food belongs are added to the food description samples, so that the image generation model can learn knowledge of different categories, so that the generated food images match the food categories, and the quality of the generated images is improved.
[0058] In some examples, the food description sample may be obtained by:
[0059] Obtain food attribute information of multiple food products from a product database;
[0060] Obtaining a product image of each food product from the product database, and extracting image attribute information from the product image;
[0061] A task prompt text is generated based on the food attribute information and the image attribute information and then input into a pre-trained language model so that the language model outputs a food description sample based on the task prompt text; wherein, the task prompt text is used to instruct the language model to convert the food attribute information and the image attribute information into natural sentences as a food expert.
[0062] As an example, as shown in FIG2B , this specification shows a schematic diagram of a training data set for an image generation model according to an exemplary embodiment. The data can be selected from the database of a local life service platform. The local life service platform provides merchants with store management services. With the authorization of the merchant, a database can be maintained to record the goods configured by the merchant for the store. The database can store product data of various products. Alternatively, the platform can also pre-construct a knowledge graph of goods in the catering industry, obtain food attribute information of a variety of food products from the knowledge graph, and obtain product images of food products from the database, etc. The food attribute information of this embodiment can be flexibly configured according to actual needs. For example, product information including product name, raw material information, taste information, ingredient information, cooking method information, etc. can be obtained from the knowledge graph. The image attribute analysis can be used to obtain the image perspective, whether it is a white background image, raw and cooked food labels or container information, etc. from the product image.
[0063] In order to construct a model suitable for the generation of raw and cooked food, this embodiment designs a training data set. The process of constructing a training data set can be to obtain historical data stored in the platform, which can include two parts: product name and product image, where the product name belongs to text data and the product image belongs to picture data. For the product name, various product attribute information of the product can be extracted based on the pre-built knowledge graph of the catering industry; the knowledge graph is a structured knowledge representation method used to describe entities, relationships and attributes in the real world. It is a graph-based data model that constructs a knowledge graph by representing entities and relationships as nodes and edges. The knowledge graph usually consists of three main components: entities, relationships and attributes. In the scenario of this embodiment, a knowledge graph for the catering industry can be pre-built, in which entities represent specific product attributes, such as product names. Relationships represent connections or interactions between entities, such as the relationship between products. Attributes represent the characteristics or properties of entities, such as numerical values in specifications.
[0064] For product images, various image attribute information of the product images can be extracted based on image attribute understanding algorithms, etc. For example, the products here can be dishes in the catering industry, and the dishes here can be food provided to users in a restaurant. In this embodiment, the food can refer to dishes in a restaurant.
[0065] Finally, based on the language model, the product attribute information and image attribute information obtained from the knowledge graph are polished to form natural sentences.
[0066] For example, you can obtain the name of the dish and / or the identification ID of the dish in the knowledge graph, and query the attribute information corresponding to the dish from the knowledge graph. Among them, there are many kinds of attribute information of dishes. Based on the construction of the training data of this embodiment, it is designed that the attribute information is attribute information related to the product image, such as raw materials, taste, cooking method, container and other attribute information. These attribute information will be included in the product image content, thereby affecting the quality of the image generated by the image generation model. As an example, please see the following Table 1, which shows a variety of different product attribute information:
[0067] Table 1:
[0068] Image attribute information can be flexibly configured based on actual needs, and may include, for example, viewing angle information, background color information, or raw / cooked state information. Viewing angle information refers to the viewing angle of the product in the product image, and background color information refers to the background color of the product image. For example, see Table 2 below, which shows various types of image attribute information:
[0069] Table 2:
[0070] As an example, extracting image attribute information from the product image includes a combination of one or more of the following steps:
[0071] Inputting the product image into a preset perspective recognition model, obtaining perspective type information of the product in the product image identified by the perspective recognition model as image attribute information, wherein the perspective type information includes a frontal perspective, a top-down perspective, or an oblique perspective;
[0072] Inputting the commodity image into a preset rawness / readiness state recognition model, and obtaining the rawness / readiness state information of the commodity in the commodity image recognized by the rawness / readiness state recognition model as image attribute information;
[0073] Pixel information in the product image is obtained to determine the background color in the product image as image attribute information.
[0074] For example, we can use a three-way image classification method to identify the perspective type of an image. First, we label multiple product image samples from different perspectives. The product image samples are labeled into three categories: front view, oblique view, or top view. Then, we train a classification model to obtain a perspective recognition model.
[0075] For example, the raw / cooked food recognition model takes product images as input and outputs raw / cooked food category labels. This model can be trained on a pre-set neural network using product image samples labeled with raw or cooked food information.
[0076] As an example, the background color in the product image of this embodiment can be two colors, one is white and the other is non-white. There are many ways to identify a white background. For example, a color product image can be converted into a grayscale image, or an edge area can be cropped from the product image. For example, if the product image is a rectangle, the edge area of the four sides of the rectangle can be cropped, and the pixel information of the edge area is obtained to obtain the histogram of the grayscale image. The histogram of an image refers to the statistical number of occurrences of the grayscale value of each pixel in an image, thereby obtaining a graph representing the grayscale distribution. Based on the statistical information of the histogram of the edge area, it can be determined whether the product image is an image with a white background. The grayscale value of the background area is counted. If the number of pixels in the edge area with a grayscale value greater than 240 accounts for a proportion of the pixels in the product image that is greater than a threshold (the threshold can be configured as needed, such as 55%), it can be determined to be a white background image.
[0077] The language model of this embodiment may refer to a natural language processing model with a huge parameter scale and learning ability.
[0078] In this embodiment, the various attribute knowledge of some dishes is stored in a structured manner, and the text is expressed as word text, such as {container: 'stone pot', cooking method: 'boil', flavor: 'spicy'}. The structured understanding results of the attribute information of product images can also usually be presented in the form of word text, such as {perspective: 'front view', style: 'white background image', raw and cooked food: 'raw food'}. Based on this, this embodiment uses a language model to combine all the above word texts and polish them into a natural and complete sentence text to construct the prompt sentence in the training data.
[0079] As an example, the task prompt text can include the role information of "food expert", which can indicate the role the language model needs to play and activate the model's professional capabilities. The task prompt text can also include the task information of "polishing food attribute information and image attribute information and converting them into natural sentences", which can provide the language model with a clear task definition so that it can perform the task correctly. The task prompt text can specifically be the following example: "If you are a foodie, please polish the following food information into a sentence: {product name, category name, raw materials, cooking method, taste, utensils, image background color, subject perspective, raw and cooked food}".
[0080] As an example, please see Table 3 below:
[0081] Table 3:
[0082] In some examples, the image generation model can be a diffusion model, such as a stable diffusion model or a latent diffusion model. When implemented using a diffusion model, based on the working principle of the diffusion model, a noise image corresponding to the food image sample can be generated. The noise image can be a Gaussian noise image, for example. Furthermore, the noise image and the food description sample corresponding to the food image sample are input into the diffusion model. The diffusion model can perform feature extraction, calculate predicted noise based on the text features extracted from the food description sample, and further generate a predicted image based on the predicted noise.
[0083] In some examples, the diffusion model can include a text encoder and an image generation network.
[0084] The diffusion model can be obtained by extracting text features from an input food description sample through the text encoder, generating a predicted image based on the text features through the image generation network, and training with minimizing the error between the predicted image and the food image sample as the optimization goal.
[0085] The image generation network can include a U-Net neural network. The U-Net consists of an encoder and a decoder, both of which are composed of ResNet blocks. The encoder compresses the image representation (feature map) into a lower-resolution image, and the decoder decodes the lower-resolution image back into a higher-resolution image. To prevent the U-Net from losing important information during downsampling, a skip connection is usually added between the encoder's downsampling ResNet and the decoder's upsampling ResNet. This connects the encoder and decoder at the same level, thereby combining shallow and deep information to better generate images. In other examples, the image generation model can extract text features from the target rawness and cookedness status information in the prompt text alone, or it can extract features from the prompt text and then concatenate the text features of the target rawness and cookedness status information with the text features of the prompt text, and generate an image based on the concatenated features. The feature extraction of the prompt text can be the text encoder already in the diffusion model, while the target rawness and cookedness status information can be implemented by adding a text feature extraction network to the diffusion model, for example, based on a pre-trained language model, thereby improving the image generation model's ability to understand the text of rawness and cookedness.
[0086] Based on the image generation model designed in the above embodiment, when training the model, parameters that need to be updated in the model can be specified, for example, all or part of the model parameters. In some examples, considering that the image generation model of this embodiment is improved using a large open source model, which itself has a large number of parameters, to reduce the amount of computation, the model parameters updated during training can be the parameters of the text encoder and image generation network.
[0087] A loss function can also be designed for model training. In machine learning, the target in a model's training samples refers to the true value corresponding to each sample; in the image generation model of this embodiment, the target is the food image sample. During model training, the loss is calculated by comparing the difference between the model's predicted value and the target. The model parameters are then updated using a backpropagation algorithm to ensure that the model's predictions are as close to the true values as possible.
[0088] A model's loss function measures the difference, or error, between the model's predicted output and the true label. The loss function is optimized during training, helping the model adjust parameters to minimize the difference between the output and the true value. Examples of loss functions include the mean squared error (MSE) and others.
[0089] When the diffusion model is selected as the image generation model, the diffusion model calculates predicted noise based on the splicing features and further generates a predicted image based on the predicted noise. Therefore, the error can be the error between the predicted noise and the actual noise of the noise image.
[0090] In some examples, the image generation model includes a diffusion model and at least one LoRA (Low-Rank Adaptation of Large Language Models) branch network connected to the diffusion model; the image generation model is trained as follows:
[0091] A preset diffusion model is trained using a first training data set, and a LoRA branch network corresponding to each specific category in at least one specific category is added to the trained diffusion model; the training samples in the first training data set include food image samples corresponding to multiple food categories;
[0092] The following training is performed for each of the LoRA branch networks as the current LoRA branch network: the model parameters of the diffusion model and the network parameters of other LoRA branch networks other than the current LoRA branch network are frozen, and the current LoRA branch network is trained using a second training data set of a specific category corresponding to the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples contained in the training samples in the second training data set and the food image samples.
[0093] In this embodiment, the diffusion model can be trained by constructing a training data set using existing dish data. Optionally, the training data of the diffusion model can include data on all types of food, such as dish data of all categories, so as to obtain a general diffusion model. Considering that in general scenarios, it is difficult to enumerate all dish data containing all information for the diffusion model to learn, and the model may also find it difficult to pay attention to subtle knowledge, resulting in the model possibly generating some food images of slightly poor quality. For example, a food called "yellow duck call" is literally understood by most people as a food related to ducks, and the model also understands it, but the food is actually fish. There is also extremely fine-grained dish generation, such as fried egg snail noodles, braised egg snail noodles, bomb snail noodles, etc. In this scenario, it is difficult for the image generation model to grasp more details.
[0094] In order to improve the performance of the model in some sub-sectors, a LoRA branch network corresponding to a specific category is designed in this embodiment. LoRA was originally applied to the field of natural language and is used to fine-tune large language models. Since the number of parameters in large models exceeds 100 billion and the training cost is too high, LoRA adopts a method of only training low-rank matrices. When used, the parameters of the LoRA model are injected into the SD model, thereby changing the generation style of the SD model, etc. It can be expressed using the data formula W=W0+BA, where W0 is the parameter (Weights) of the initial SD model, BA is the low-rank matrix, that is, the parameter of the LoRA model, and W represents the final SD model parameter after being affected by the LoRA model. The BA matrix is the training target of the LoRA model. The LoRA model can fine-tune the Clip text encoder (text encoder of the Clip model) in the SD model as the text feature extraction layer and some network layers in Unet, such as the linear layer of CrossAttention. The whole process is a linear relationship. It can be considered that the original SD model is superimposed on the LoRA model to obtain a model with a completely new effect.
[0095] This embodiment designs a diffusion model that is first trained using a first training dataset. This first training dataset contains a large amount of data, such as training samples covering all categories. After the SD model is trained, it can be connected to one or more LoRA branch networks, where each LoRA branch network corresponds to a specific category. In actual applications, one or more specific categories can be configured as needed. In this way, under the control of each LoRA branch network, the SD model can output food images that match the specific category corresponding to each LoRA branch network.
[0096] The number of specific categories can be configured as needed. Specific specific categories can include categories where the SD model has difficulty controlling details and has poor image quality. The SD model can be tested, and the test data set contains data from multiple categories. Food images generated by the SD model for food description information of different categories are obtained. The quality of the food images generated by the SD model for different categories is judged. This can be done manually or by calculating the error between the image generated by the SD and the food image sample corresponding to the food description information. Multiple data sets can be prepared for each category, and a comprehensive judgment can be made based on the prediction results of the SD model for multiple data sets of each category, etc., to determine one or more categories with poor image generation quality of the SD model as specific categories. Of course, it can also be configured as needed; for example, it can be barbecue ingredients or seafood, etc. This is not limited in this embodiment.
[0097] For example, if n specific categories are set, for any category n among the n categories,i , set the corresponding second training data set, in which the training samples are category n i The second training dataset can be derived from the first training dataset used to train the diffusion model. The number of the second training dataset can be small, for example, dozens to hundreds of data, and can be configured as needed in actual applications. i When training the corresponding LoRA branch network, the SD model and the LoRA branch networks corresponding to other categories do not participate in the training, and the parameters can be frozen. In this way, the trained LoRA branch network can learn the association between the food description information and food images of the category, thereby controlling the diffusion model to generate food images that match the category. In this embodiment, each LoRA branch network can act on the text feature extraction layer in the SD model and the query (key) and value (value) layers in the U-Net network respectively.
[0098] After the image generation model is trained, it can be used to generate images that match the actual rawness or doneness of the food. For example, by inputting a prompt text containing "food name, rawness or doneness information" into the model, the model can generate a matching food image. Optionally, the prompt text can also contain "food name, food category, rawness or doneness information," as needed. Figure 2C shows the overall image generation flow chart according to an exemplary embodiment of this specification.
[0099] Optionally, based on the target food category in the prompt text, the LoRA branch network corresponding to the target food category can be selected in the trained model. The other LoRA branch networks are not selected and are in an inoperative state. The SD model and the LoRA branch network corresponding to the target food category generate the corresponding food image based on the text features of the input prompt text.
[0100] In some examples, inputting the prompt text into a pre-trained image generation model so that the image generation model generates a food image that matches the target food name and the target cooked state information based on the input prompt text may include:
[0101] Obtaining the target food category contained in the prompt text;
[0102] In response to determining that the target food category has a corresponding target LoRA branch network, freezing the other LoRA branch networks in the image generation model except the target LoRA branch network, and inputting the prompt text into the image generation model including the diffusion model and the target LoRA branch network, so that the image generation model including the diffusion model and the target LoRA branch network generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text;
[0103] In response to determining that the target food category does not have a corresponding target LoRA branch network, all LoRA branch networks in the image generation model are frozen, and the prompt text is input into the image generation model including the diffusion model, so that the image generation model including the diffusion model generates a food image that matches the target food name, the target food category and the target raw / cooked status information based on the input prompt text.
[0104] In this embodiment, the image generation model includes multiple LoRA branch networks corresponding to different categories. Optionally, data recording specific categories and LoRA branch networks can be generated, such as a mapping dictionary. When the model is used, the target food category in the prompt text may have a corresponding LoRA branch network, or it may not. Through the above embodiment, it is possible to query whether there is a corresponding LoRA branch network from the mapping dictionary based on the target food category in the prompt text, and control the working status of the LoRA branch network in the model.
[0105] In this embodiment, there are multiple ways to obtain the prompt text; for example, a user interface can be provided for the user to input food description information. As an example, a text input box corresponding to the food description information can be provided, and the user can enter the food description information through the text box on the user interface. Optionally, the food description information can be subdivided into food name, food category, or raw or cooked status, and corresponding text boxes can be provided for each type of subdivided information to receive user input information through each text box. The user interface can also include a prompt message prompting the user to enter the raw or cooked status information of the food, so as to prompt the user to enter the raw or cooked status information of the food.
[0106] Optionally, obtaining the prompt text containing food description information includes:
[0107] The target food name input by the user is obtained, and the target raw / cooked state information corresponding to the target food name is searched from a preset raw / cooked food database.
[0108] In this embodiment, the user may not enter the target raw or cooked food status information; a raw or cooked food database may be pre-built, which may store data of "category-dish name-raw or cooked food label", and the corresponding raw or cooked food status information may be queried based on the target food name entered by the user.
[0109] Optionally, obtaining the prompt text containing food description information may include:
[0110] In response to receiving food description information of the merchant to be configured obtained from the store configuration page and sent by the merchant client, generating a prompt text containing the food description information using the food description information;
[0111] Outputting the food image generated by the image generation model may include:
[0112] The food image generated by the image generation model is configured in the store configuration page.
[0113] In this embodiment, it can be applied to the scenario where merchants configure store merchandise. When merchants use the merchant client, the merchant client provides a store configuration page, which may include a configuration function for dishes. Users can enter food description information for the food to be configured, and this embodiment can generate prompt text. Among them, the raw and cooked status of the food can also be predicted based on the characteristics of the dishes sold in the merchant's store, for example, based on the categories of various dishes sold in the merchant's store, and / or based on the raw and cooked status of the dishes sold in the merchant's store history; for example, most of the dishes sold in the merchant's store history are cooked food, and the target raw and cooked status information can be determined to be cooked food, etc. The prompt text is input into the image generation model, and the food image generated by the image generation model is obtained and configured in the store configuration page, so that the merchant does not need to shoot and configure the food image by himself.
[0114] As shown in FIG2D , this specification shows a schematic diagram of an image generation according to an exemplary embodiment. The image generation model of this embodiment can be deployed on a user-side device or a physical device deployed on the server side; this embodiment does not limit this. The user interface shows a text input box corresponding to "Please enter the name of the dish" for the user to enter the name of the dish. The figure takes the user entering "steak" as an example. Optionally, the server can automatically obtain the raw or cooked status, or the user interface also shows a text input box corresponding to "Please enter raw or cooked food" for the user to enter the target raw or cooked status information. The user runs a Western restaurant. The figure takes the user entering "cooked food" as an example.
[0115] The user input box in the user interface captures the two types of information entered above. A prompt is generated based on the user's input and fed into the image generation model. The image generation model then generates an image based on the input data. The figure shows a model-generated image of a cooked steak, which matches the dish sold in the user's store. If the user enters a raw option, the model generates a raw steak.
[0116] FIG3 is a flowchart of another image generation method according to an exemplary embodiment of the present specification. The method may include:
[0117] Step 302: Obtain the target food name input by the user and send it to the server. The server is used to obtain target raw or cooked state information indicating whether the food is raw or cooked, and then execute the steps of the above-mentioned image generation method embodiment.
[0118] Step 304: Obtain and display the food image output by the server.
[0119] This embodiment can be applied to the client, and the client can also obtain the food images automatically generated by the server and then configure them in the merchant's store, thereby allowing the merchant to reduce the configuration of product images in the store. The specific implementation process of this embodiment can be referred to the previous embodiment and will not be repeated here.
[0120] Corresponding to the aforementioned embodiments of the image generating method, this specification also provides embodiments of an image generating apparatus and a computer device used therein.
[0121] As shown in FIG4 , FIG4 is a block diagram of an image generating device according to an exemplary embodiment of the present specification, the device including:
[0122] An acquisition module 41 is configured to acquire a prompt text containing food description information, wherein the food description information includes a target food name and target rawness or cookedness information indicating whether the food is raw or cooked;
[0123] An image generation module 42 is configured to input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image that matches the target food name and the target raw / cooked state information based on the input prompt text; wherein the training samples of the image generation model include: food description samples and corresponding food image samples, wherein the food description samples include food name samples and raw / cooked state information samples of the food in the food image samples;
[0124] The image output module 43 is used to output the food image generated by the image generation model.
[0125] In some examples, the food description information further includes a target food category; the food description sample further includes a food category information sample;
[0126] The image generation module 42 inputs the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target cooked state information based on the input prompt text, including:
[0127] The prompt text is input into a pre-trained image generation model so that the image generation model generates a food image that matches the target food name, the target food category and the target raw / cooked state information based on the input prompt text.
[0128] In some examples, the image generation model includes a diffusion model and at least one LoRA branch network connected to the diffusion model; the image generation model is trained as follows:
[0129] A preset diffusion model is trained using a first training data set, and a LoRA branch network corresponding to each specific category in at least one specific category is added to the trained diffusion model; the training samples in the first training data set include food image samples corresponding to multiple food categories;
[0130] The following training is performed for each of the LoRA branch networks as the current LoRA branch network: the model parameters of the diffusion model and the network parameters of other LoRA branch networks other than the current LoRA branch network are frozen, and the current LoRA branch network is trained using a second training data set of a specific category corresponding to the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples contained in the training samples in the second training data set and the food image samples.
[0131] In some examples, the diffusion model includes a text encoder and an image generation network;
[0132] The diffusion model is obtained by extracting text features from an input food description sample through the text encoder, generating a predicted image based on the text features through the image generation network, and training with the optimization goal of minimizing the error between the predicted image and the food image sample.
[0133] In some examples, the image generation module is specifically configured to:
[0134] Obtaining the target food category contained in the prompt text;
[0135] In response to determining that the target food category has a corresponding target LoRA branch network, freezing the other LoRA branch networks in the image generation model except the target LoRA branch network, and inputting the prompt text into the image generation model including the diffusion model and the target LoRA branch network, so that the image generation model including the diffusion model and the target LoRA branch network generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text;
[0136] In response to determining that the target food category does not have a corresponding target LoRA branch network, all LoRA branch networks in the image generation model are frozen, and the prompt text is input into the image generation model including the diffusion model, so that the image generation model including the diffusion model generates a food image that matches the target food name, the target food category and the target raw / cooked status information based on the input prompt text.
[0137] In some examples, in the image generation module 42, the food description sample is obtained by:
[0138] Obtain food attribute information of multiple food products from a product database;
[0139] Obtaining a product image of each food product from the product database, and extracting image attribute information from the product image;
[0140] A task prompt text is generated based on the food attribute information and the image attribute information and then input into a pre-trained language model so that the language model outputs a food description sample based on the task prompt text; wherein, the task prompt text is used to instruct the language model to convert the food attribute information and the image attribute information into natural sentences as a food expert.
[0141] In some examples, extracting image attribute information from the product image includes a combination of one or more of the following steps:
[0142] Inputting the product image into a preset perspective recognition model, obtaining perspective type information of the product in the product image identified by the perspective recognition model as image attribute information, wherein the perspective type information includes a frontal perspective, a top-down perspective, or an oblique perspective;
[0143] Inputting the commodity image into a preset rawness / readiness state recognition model, and obtaining the rawness / readiness state information of the commodity in the commodity image recognized by the rawness / readiness state recognition model as image attribute information;
[0144] Pixel information in the product image is obtained to determine the background color in the product image as image attribute information.
[0145] In some examples, the acquisition module is specifically used to: acquire the target food name input by the user, and query the target raw and cooked state information corresponding to the target food name from a preset raw and cooked food database.
[0146] In some examples, the acquisition module is specifically configured to: in response to receiving food description information of the merchant to be configured obtained from the store configuration page and sent by the merchant client, generate a prompt text containing the food description information using the food description information;
[0147] The image output module is specifically used to configure the food image generated by the image generation model in the store configuration page.
[0148] The implementation process of the functions and effects of each module in the above-mentioned image generation device is specifically described in the implementation process of the corresponding steps in the above-mentioned image generation method, and will not be repeated here.
[0149] The embodiments of the image generating device of this specification can be applied to computer equipment, such as servers or terminal devices. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, as shown in Figure 5, it is a hardware structure diagram of the computer device where the image generating device of this specification is located. In addition to the processor 510, memory 530, network interface 520, and non-volatile memory 540 shown in Figure 5, the computer device where the image generating device 531 in the embodiment is located, usually according to the actual function of the computer device, can also include other hardware, which will not be described in detail.
[0150] Accordingly, an embodiment of this specification further provides a computer program product, including a computer program, which implements the steps of the aforementioned image generation method embodiment when executed by a processor.
[0151] Accordingly, an embodiment of this specification also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the image generation method embodiment when executing the program.
[0152] Accordingly, an embodiment of this specification further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the embodiment of the image generating method are implemented.
[0153] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0154] The above embodiments can be applied to one or more computer devices, where the computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0155] The computer device may be any electronic product that can interact with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.
[0156] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0157] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0158] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0159] The steps of the various methods above are divided only for the purpose of clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.
[0160] The phrases "specific examples" or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0161] Other embodiments of the present invention will readily occur to those skilled in the art upon consideration of the present invention and practice of the invention. This specification is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present invention being indicated by the following claims.
[0162] It should be understood that the present description is not limited to the exact structure that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.
[0163] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A method for generating an image, the method comprising: Obtaining a prompt text containing food description information; wherein the food description information includes a target food name and target rawness or cookedness status information indicating whether the food is raw or cooked; The prompt text is input into a pre-trained image generation model, so that the image generation model generates a food image that matches the target food name and the target raw-cooked state information based on the input prompt text; wherein the training samples of the image generation model include: food description samples and corresponding food image samples, and the food description samples include food name samples and raw-cooked state information samples of the food in the food image samples; Output the food image generated by the image generation model.
2. According to the method of claim 1, the food description information further includes a target food category; the food description sample further includes a food category information sample; The step of inputting the prompt text into a pre-trained image generation model so that the image generation model generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text comprises: The prompt text is input into a pre-trained image generation model so that the image generation model generates a food image that matches the target food name, the target food category and the target raw / cooked state information based on the input prompt text.
3. According to the method of claim 2, the image generation model comprises a diffusion model and at least one LoRA branch network connected to the diffusion model; the image generation model is trained in the following manner: The preset diffusion model is trained using the first training data set, and a LoRA branch network corresponding to each specific category in at least one specific category is added to the trained diffusion model; the training samples in the first training data set include food image samples corresponding to multiple food categories; Each of the LoRA branch networks is used as the current LoRA branch network to perform the following training: freeze the model parameters of the diffusion model and the network parameters of other LoRA branch networks outside the current LoRA branch network, and use the second training data set of the specific category corresponding to the current LoRA branch network to train the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples contained in the training samples in the second training data set and the food image samples.
4. The method according to claim 3, wherein the diffusion model comprises a text encoder and an image generation network; The diffusion model is obtained by extracting text features from an input food description sample by the text encoder, and generating a predicted image based on the text features by the image generation network, with minimizing the error between the predicted image and the food image sample as the optimization goal.
5. The method according to claim 3, wherein the step of inputting the prompt text into a pre-trained image generation model so that the image generation model generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text comprises: Obtaining the target food category contained in the prompt text; In response to determining that the target food category has a corresponding target LoRA branch network, freezing other LoRA branch networks in the image generation model except the target LoRA branch network, and inputting the prompt text into the image generation model including the diffusion model and the target LoRA branch network, so that the image generation model including the diffusion model and the target LoRA branch network generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text; In response to determining that the target food category does not have a corresponding target LoRA branch network, all LoRA branch networks in the image generation model are frozen, and the prompt text is input into the image generation model including the diffusion model, so that the image generation model including the diffusion model generates a food image that matches the target food name, the target food category and the target raw / cooked status information based on the input prompt text.
6. A method for generating an image, the method comprising: Obtaining the target food name input by the user and sending it to the server, the server is used to obtain the target rawness or cookedness status information indicating whether the food is raw or cooked, and then execute the steps of any one of the methods described in claims 1 to 5; Obtain and display the food image output by the server.
7. An image generating device, comprising: An acquisition module, used to acquire a prompt text containing food description information, wherein the food description information includes a target food name and target rawness or cookedness status information indicating whether the food is raw or cooked; An image generation module is used to input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw-cooked state information based on the input prompt text; wherein the training samples of the image generation model include: food description samples and corresponding food image samples, and the food description samples include food name samples and raw-cooked state information samples of the food in the food image samples; The image output module is used to output the food image generated by the image generation model.
8. An image generating device, the device comprising: An acquisition module, used to: acquire the target food name input by the user and send it to the server, wherein the server is used to acquire the target raw or cooked state information indicating whether the food is raw or cooked, and then execute the steps of any one of the methods of claims 1 to 5; The display module is used to obtain and display the food image output by the server.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any one of the methods of claims 1 to 6.
Citation Information
Patent Citations
Dish information display method and device, storage medium and electronic device
CN116484083A
Text-to-image generation method based on fine-grained semantic reward
CN116883530A
Image generation method and device, equipment and storage medium
CN117710530A
Learning object / action pairs for recipe ingredients
US20170228364A1