Image generation method, apparatus, device, and storage medium
By constructing an image generation model, and using a pre-trained model and a diffusion model combined with a LoRA branch network, accurate images of food in both raw and cooked states are generated, solving the problem of misleading product images and improving users' awareness and interest in the products.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, the raw/cooked state of product images does not match the actual product state, leading to misleading purchases by users. Images generated by automated synthesis tools also have similar problems, lacking the ability to generate accurate raw/cooked product images.
By constructing an image generation model, using a pre-trained model and a diffusion model combined with a LoRA branch network, the training samples include food descriptions and image samples to generate accurate images of food in both raw and cooked states. The diffusion model is then used to generate images that match the food descriptions.
It generates high-quality food images that accurately represent the raw and cooked states, solving the problem of misleading product images and enhancing users' understanding and interest in the products.
Smart Images

Figure CN117710530B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of machine learning technology, and in particular to image generation methods, apparatus, devices, and storage media. Background Technology
[0002] In local services and other application scenarios, platforms can provide store services for merchants, allowing them to manage their online stores. For example, merchants can showcase various products in their online stores, uploading product photos, adding descriptions and specifications to demonstrate product features and advantages to users. Understandably, the quality of product images plays a crucial role; product images need to accurately convey product information to enhance user awareness and interest. Therefore, generating high-quality product images has become a pressing technical challenge. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides an image generation method, apparatus, device and storage medium.
[0004] According to a first aspect of the embodiments of this specification, an image generation method is provided, the method comprising:
[0005] Obtain prompt text containing food description information, wherein the food description information includes the target food name and target raw or cooked status information indicating whether the food is raw or cooked;
[0006] The prompt text is input into a pre-trained image generation model, which generates a food image that matches the target food name and the target raw / cooked status information based on the input prompt text. The training samples of the image generation model include food description samples and corresponding food image samples. The food description samples include food name samples and raw / cooked status information samples of the food in the food image samples.
[0007] Output the food image generated by the image generation model.
[0008] According to a second aspect of the embodiments of this specification, an image generation method is provided, the method comprising:
[0009] The system obtains the target food name input by the user and sends it to the server. The server then obtains the target raw or cooked state information that indicates whether the food is raw or cooked, and executes the steps of the image generation method embodiment described in the first aspect.
[0010] Obtain and display the food images output by the server.
[0011] According to a third aspect of the embodiments of this specification, an image generation apparatus is provided, comprising:
[0012] The acquisition module is used to acquire prompt text containing food description information, wherein the food description information includes the target food name and target raw or cooked status information indicating whether the food is raw or cooked.
[0013] An image generation module is used to input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw / cooked status information based on the input prompt text; wherein, the training samples of the image generation model include: food description samples and corresponding food image samples, the food description samples including food name samples and raw / cooked status information samples of the food in the food image samples;
[0014] The image output module is used to output the food image generated by the image generation model.
[0015] According to a third aspect of the embodiments of this specification, an image generation apparatus is provided, the apparatus comprising:
[0016] The acquisition module is used to: acquire the target food name input by the user and send it to the server. The server is used to acquire the target raw or cooked status information that indicates whether the food is raw or cooked, and then execute the steps of the image generation method embodiment described in the first aspect.
[0017] The display module is used to: acquire and display food images output by the server.
[0018] According to a fourth aspect of the embodiments of this specification, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the image generation method embodiments described in the first aspect above.
[0019] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the image generation method embodiments described in the first aspect above.
[0020] The technical solutions provided in the embodiments of this specification may include the following beneficial effects:
[0021] In this embodiment, the training samples for training the image generation model include: food description samples and corresponding food image samples. The food description samples include food name samples and raw / cooked status information samples from the food image samples. Therefore, the image generation model is capable of generating food images with accurate raw / cooked status. During inference, this embodiment obtains prompt text containing food description information, wherein the food description information includes the target food name and target raw / cooked status information indicating whether the food is raw or cooked. The prompt text is input into the pre-trained image generation model, so that the image generation model can generate food images matching the target food name and target raw / cooked status information based on the input prompt text. Therefore, this embodiment can generate high-quality food images with accurate raw / cooked status.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this specification and, together with the specification, serve to explain the principles of this disclosure.
[0024] Figure 1 This is a schematic diagram illustrating an application scenario according to an exemplary embodiment of this specification.
[0025] Figure 2A This is a flowchart illustrating an image generation method according to an exemplary embodiment of this specification.
[0026] Figure 2B This is a schematic diagram of the training dataset of an image generation model illustrated in this specification according to an exemplary embodiment.
[0027] Figure 2C This is an overall flowchart illustrating the generation of an image according to an exemplary embodiment of this specification.
[0028] Figure 2D This is an application scenario diagram generated from an image shown in this specification based on an exemplary embodiment.
[0029] Figure 3 This is a flowchart illustrating an image generation method according to an exemplary embodiment of this specification.
[0030] Figure 4 This is a block diagram illustrating an image generation apparatus according to an exemplary embodiment of this specification.
[0031] Figure 5 This specification is a hardware structure diagram of a computer device containing an image generation apparatus, according to an exemplary embodiment. Detailed Implementation
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0033] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0035] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.
[0036] like Figure 1The diagram shown is an application scenario illustration based on an exemplary embodiment of this specification, including a platform provider, merchants, and users. The platform provider builds a server and provides a client application to users, who can use the services offered by the platform through the client. Unlike the client application for users, the platform provider also provides a client application for merchants, who can use the merchant services offered by the platform through this client. For example, merchants can manage their online store information through the platform, including store name, address, and contact information. It may also include product display services, where merchants can showcase their products by uploading product photos, adding descriptions and specifications, and demonstrating product features and advantages to users. In practical applications, it may also include order management services or delivery services, which will not be elaborated upon in this embodiment.
[0037] The inventors of this application have discovered that the quality of product images plays a crucial role. Product images need to accurately convey product information to enhance users' awareness and interest in the product. The research found that in certain application scenarios, the accuracy of product information conveyed in product images has a significant impact on potential buyers. For example, in the food service industry, merchants sell food products, some of which exist in two states: ready-to-eat cooked food and raw pre-prepared food requiring further processing, such as barbecue and dumplings. For these two types of products, if the raw / cooked state depicted in the product image does not match the actual raw / cooked state of the product, it can significantly interfere with users' selection and judgment. For example:
[0038] (1) In the context of food and beverage products, some categories of cooked food products are accompanied by images of raw pre-prepared food products;
[0039] One barbecue restaurant sells pork belly skewers, which are actually cooked (grilled) pork belly, but the images displayed on the store page show raw pork belly skewers. Another restaurant sells tomahawk lamb chops, which are supposed to be raw (uncooked), but the images show them as cooked. These discrepancies between the food images and the actual cooking status shown on the store's website are misleading to consumers. One reason for this is that the merchants lack the ability to create accurate images of the raw and cooked food.
[0040] (2) Some automated product image synthesis tools select food images from the image library to fill the template based on actual synthesis needs. Since some categories (such as barbecue and related categories) in the image library contain a large number of raw foods in their product images, the food in the barbecue category product images automatically generated by the tool is in a raw state, which does not match the actual cooked state of barbecue category products.
[0041] Based on this, the embodiments of this specification provide an image generation scheme that can automatically generate accurate images of food in both raw and cooked states. The embodiments of this specification will now be described in detail.
[0042] like Figure 2A As shown, Figure 2A This is a flowchart illustrating an image generation method according to an exemplary embodiment, which may include the following steps:
[0043] Step 202: Obtain the prompt text containing food description information.
[0044] The food description information includes the target food name and target raw or cooked status information indicating whether the food is raw or cooked.
[0045] Step 204: Input the prompt text into a pre-trained image generation model so that the image generation model generates a food image that matches the target food name and the target raw / cooked status information based on the input prompt text.
[0046] The training samples of the image generation model include: food description samples and corresponding food image samples. The food description samples include food name samples and raw / cooked status information samples of the food in the food image samples.
[0047] Step 206: Output the food image generated by the image generation model.
[0048] The image generation method of this embodiment can run on a computer device, including but not limited to servers, cloud servers, server clusters, tablet computers, personal digital assistants, laptops, or desktop computers. In some examples, the model involved in the image generation method of this embodiment can be deployed on a server, for example, it can be deployed on a... Figure 1 In the server-side interface with the merchant's client, the merchant's client can receive user input and send it to the server. The server then runs the image generation method of this embodiment and sends the result back to the client. In other examples, the model involved in the image generation method of this embodiment can also be deployed in the merchant's client; this embodiment does not impose any restrictions on this.
[0049] In this embodiment, the terms "raw" and "cooked" can be used to indicate whether food has been cooked or heated. "Raw food" refers to ingredients or food in their raw, unprocessed state. Raw food typically refers to fresh, unprocessed food, such as uncooked vegetables, meat, and seafood. It can also include food that has undergone simple processing, such as marinated food. Examples include lamb and shrimp paste used in hot pot in the catering industry, and chicken breast sold in retail stores. "Cooked food" refers to ingredients or food that have undergone appropriate cooking or heating methods, such as boiling, stir-frying, grilling, or steaming.
[0050] In this embodiment, the image generation model can be pre-built, and the building process can be executed on a single computer device or multiple devices in a cluster of computer devices. A training dataset can be prepared in advance; this embodiment does not limit the method of obtaining the training dataset. This embodiment designs the training samples to include: food description samples and corresponding food image samples, wherein the food description samples at least include food name samples and raw / cooked status information samples.
[0051] The image generation model in this embodiment can be constructed according to its specific structure as needed, and the scheme in this embodiment can be used to improve existing pre-trained image generation models. Optionally, the image generation model may include: a generative adversarial network (GAN) model, a flow-based model, or a diffusion model, etc., and this embodiment does not limit it.
[0052] The food description samples can contain more information, and can be configured as needed in practical applications; this embodiment does not impose any limitations on this. When the food description samples contain more information, the image generation model can acquire richer information related to the food image samples, thereby learning a wealth of knowledge about the relationship between the food description samples and the food image samples, and improving the model's inference accuracy. In some examples, the food description information may also include the target food category; the food description samples may also include food category information samples.
[0053] The step of inputting the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text, may include:
[0054] The prompt text is input into a pre-trained image generation model, which generates a food image that matches the target food name, the target food category, and the target raw / cooked status information based on the input prompt text.
[0055] In this embodiment, food category information samples of the food category to which the food belongs are added to the food description samples. This allows the image generation model to learn knowledge of different categories, so that the generated food images match the food categories and improve the quality of the generated images.
[0056] In some examples, the food description sample may be obtained in the following ways:
[0057] Retrieve food attribute information for various food products from the product database;
[0058] Obtain product images of each food product from the product database, and extract image attribute information from the product images;
[0059] After generating task prompt text based on the food attribute information and the image attribute information, the text is input into a pre-trained language model so that the language model outputs food description samples based on the task prompt text; wherein, the task prompt text is used to instruct the language model to convert the food attribute information and the image attribute information into natural language in the role of a food expert.
[0060] As an example, such as Figure 2B The diagram illustrates a training dataset for an image generation model according to an exemplary embodiment of this specification. This dataset can be constructed from data selected from the databases of platforms such as local life service platforms. Local life service platforms provide store management services to merchants. With the merchant's authorization, they can maintain a database recording the products configured by the merchant for the store. This database can store product data for various products. Alternatively, the platform can pre-build a knowledge graph of products in the catering industry, obtain food attribute information for various food products from the knowledge graph, and retrieve product images of food products from the database, etc. The food attribute information in this embodiment can be flexibly configured according to actual needs. For example, product information including product name, raw material information, flavor information, ingredient information, cooking method information, etc., can be obtained from the knowledge graph. Image attribute analysis can be used to obtain image perspective, whether it is a white background image, raw / cooked food labels, or container information, etc., from the product image.
[0061] To construct a model suitable for generating both raw and cooked food, this embodiment designs a training dataset. The process of constructing the training dataset involves acquiring historical data stored on the platform, which may include two parts: product names and product images. Product names are text data, and product images are image data. For product names, various product attribute information can be extracted based on a pre-constructed knowledge graph of the catering industry. A knowledge graph is a structured knowledge representation method used to describe entities, relationships, and attributes in the real world. It is a graph-based data model that constructs a knowledge graph by representing entities and relationships as nodes and edges. A knowledge graph typically consists of three main components: entities, relationships, and attributes. In this embodiment, a knowledge graph of the catering industry can be pre-constructed. In this knowledge graph, entities represent specific product attributes, such as product names. Relationships represent connections or interactions between entities, such as relationships between products. Attributes represent the characteristics or properties of entities, such as numerical values in specifications.
[0062] For product images, various image attribute information can be extracted based on image attribute understanding algorithms. For example, the product here could be a dish in the catering industry, which refers to food provided to users in a restaurant. In this embodiment, the food can refer to the dishes served in a restaurant.
[0063] Finally, based on the language model, the product attribute information and image attribute information obtained from the knowledge graph are refined to form natural sentences.
[0064] For example, the name of a dish and / or its identifier ID in a knowledge graph can be obtained, and the corresponding attribute information can be retrieved from the knowledge graph. The attribute information of a dish can be varied. Based on the training data constructed in this embodiment, the attribute information is designed to be related to the product image, such as ingredients, flavor, cooking method, and container. This attribute information is included in the product image content and thus affects the quality of the image generated by the image generation model. As an example, please see Table 1 below, which shows various different product attribute information:
[0065] Table 1:
[0066] property For example raw material Chicken wings, kelp, tofu, fish, quail eggs, ... Cooking methods Stir-fry, steam, braise, boil, deep-fry, top with sauce, grill, ... taste Sweet and sour, Orleans, creamy, garlicky, spicy... container Platters, hot pot, stone pot, basin, earthenware pot, steamer, dry pot, pot, tin foil...
[0067] The image attribute information can be flexibly configured according to actual needs, and may include perspective information, background color information, or raw / cooked status information. Perspective information refers to the shooting angle of the product in the product image, and background color information refers to the background color of the product image. As an example, please see Table 2 below, which shows several different types of image attribute information:
[0068] Table 2:
[0069]
[0070]
[0071] As an example, extracting image attribute information from the product image includes a combination of one or more of the following steps:
[0072] The product image is input into a preset view recognition model, and the view type information of the product in the product image recognized by the view recognition model is obtained as image attribute information. The view type information includes frontal view, top view, or oblique view.
[0073] The product image is input into a preset raw / cooked state recognition model, and the raw / cooked state information of the product in the product image recognized by the raw / cooked state recognition model is obtained as image attribute information;
[0074] Pixel information in the product image is obtained to determine the background color in the product image as image attribute information.
[0075] As an example, a three-class image classification method can be used to determine the viewpoint type of an image. First, multiple product image samples with various viewpoints are labeled. The labels for the product image samples include front view, oblique view, or top view. Figure 3 Classification model is then trained to obtain the viewpoint recognition model.
[0076] As an example, the input data for the raw / cooked food status recognition model is product images, and the output is the category label for raw or cooked food. The raw / cooked food status recognition model can be obtained by training a pre-defined neural network with product image samples labeled with raw or cooked food information.
[0077] As an example, the background color in the product image of this embodiment can be two types: white and non-white. There are several ways to identify a white background. For example, a color product image can be converted to a grayscale image. Another method is to crop out the edge regions from the product image. If the product image is rectangular, the edge regions of the four sides of the rectangle can be cropped, and the pixel information of the edge regions can be obtained to generate a histogram of the grayscale image. An image histogram is a chart that statistically represents the grayscale distribution by counting the number of times each pixel's grayscale value appears in an image. Based on the statistical information of the histogram of the edge regions, it can be determined whether the product image has a white background. The grayscale values of the background regions are statistically analyzed. If the proportion of pixels with grayscale values >240 in the edge regions to the total number of pixels in the product image is greater than a threshold (the threshold can be configured as needed, such as 55%), then it can be determined to be a white background image.
[0078] In this embodiment, the language model can refer to a natural language processing model with a huge parameter scale and learning ability.
[0079] In this embodiment, the various attribute knowledge of some dishes is stored in a structured manner, and the text representation is word text, such as {container: 'stone pot', cooking method: 'boil', flavor: 'spicy'}. The structured understanding results of the attribute information of product images can also usually be presented in word text, such as {viewpoint: 'front view', style: 'white background image', raw / cooked food: 'raw food'}. Based on this, this embodiment uses a language model to combine all the above word texts, refine them into a natural and complete sentence, and construct the prompt statement in the training data.
[0080] As an example, the task prompt text could include the role information of a "food expert," thus instructing the language model on the role it needs to play and activating the model's expertise. The task prompt text could also include the task information of "refining food attribute information and image attribute information into natural language sentences," providing the language model with a clear task definition so that it can perform the task correctly. A specific example of task prompt text would be: "If you are a food connoisseur, please refine the following food information into a single sentence: {product name, category name, ingredients, cooking method, flavor, utensils, image background color, subject perspective, raw / cooked food}."
[0081] As an example, please see Table 3 below:
[0082] Table 3:
[0083]
[0084] In some examples, the image generation model can be a diffusion model, such as a stable diffusion model or a latent diffusion model. When implementing a diffusion model, based on its working principle, it can generate a noisy image corresponding to the food image sample. This noisy image could be a Gaussian noise image, etc. Further, the noisy image and the food description sample corresponding to the food image sample are input into the diffusion model. The diffusion model can perform feature extraction, calculate predicted noise based on the text features extracted from the food description sample, and then generate a predicted image based on the predicted noise.
[0085] In some examples, the diffusion model may include a text encoder and an image generation network.
[0086] The diffusion model can be trained by extracting text features from the input food description sample through the text encoder, generating a predicted image based on the text features through the image generation network, and minimizing the error between the predicted image and the food image sample as the optimization objective.
[0087] Image generation networks can include U-Net neural networks. A U-Net consists of an encoder and a decoder, both composed of ResNet blocks. The encoder compresses the image representation (feature map) into a lower-resolution image, and the decoder decodes the lower-resolution image back into a higher-resolution image. To prevent the U-Net from losing important information during downsampling, skip connections are typically added between the encoder's downsampling ResNet and the decoder's upsampling ResNet. This connects encoders and decoders at the same level, combining shallow and deep information to generate better images. In other examples, the image generation model can extract textual features from the target's raw / cooked state information in the prompt text, or it can extract features from the prompt text itself and then concatenate the textual features of the target's raw / cooked state information with the textual features of the prompt text, generating an image based on the concatenated features. The feature extraction of the prompt text can be achieved using an existing text encoder in the diffusion model, while the target's raw / cooked state information can be obtained by adding a textual feature extraction network to the diffusion model, such as a pre-trained language model, thereby improving the image generation model's ability to understand the raw / cooked state text.
[0088] Based on the image generation model designed in the above embodiments, the parameters that need to be updated in the model can be specified during training, such as all or some of the model's parameters. In some examples, the image generation model in this embodiment is an improvement on a large open-source model. Since the large model itself has a huge number of parameters, in order to reduce the computational load, the model parameters updated during training can be the parameters of the text encoder and the image generation network.
[0089] Furthermore, the loss function during model training can be designed. In the field of machine learning, in the training samples of a model, the target refers to the true value corresponding to each sample; in the image generation model of this embodiment, the target is the food image sample. During model training, the loss is calculated by comparing the difference between the model's predicted value and the target, and the backpropagation algorithm is used to update the model's parameters so that the model's prediction results are as close as possible to the true values.
[0090] The model's loss function is a function used to measure the difference or error between the model's predicted output and the true label. The loss function is optimized during training to help the model minimize the difference between the output and the true value by adjusting its parameters. Examples of loss functions include the Mean Squared Error (MSE) loss function.
[0091] When the diffusion model is selected as the image generation model, the diffusion model calculates the predicted noise based on the stitching features and further generates a predicted image based on the predicted noise. Therefore, the error can be the error between the predicted noise and the actual noise in the noise image.
[0092] In some examples, the image generation model includes a diffusion model and at least one LoRA (Low-Rank Adaptation of Large Language Models) branch network connected to the diffusion model; the image generation model is trained in the following manner:
[0093] The pre-defined diffusion model is trained using the first training dataset. A LoRA branch network corresponding to each specific category in at least one specific category is added to the trained diffusion model. The training samples in the first training dataset contain food image samples corresponding to multiple food categories.
[0094] For each of the LoRA branch networks, the following training is performed as the current LoRA branch network: the model parameters of the diffusion model and the network parameters of other LoRA branch networks besides the current LoRA branch network are frozen, and the current LoRA branch network is trained using the second training dataset of the specific category corresponding to the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples and the food image samples contained in the training samples in the second training dataset.
[0095] In this embodiment, existing food data can be used to construct a training dataset to train the diffusion model. Optionally, the training data for the diffusion model can include data on all types of food, such as food data from all categories, thus obtaining a general diffusion model. Considering that in general scenarios, it is difficult to exhaustively collect all food data containing all information for the diffusion model to learn from, and the model may also struggle to capture subtle details, potentially leading to the generation of some food images of slightly lower quality. For example, a food named "Yellow Duck Call" might be understood literally as a duck-related food, and the model would also interpret it that way, but the food is actually fish. There are also extremely fine-grained food generation scenarios, such as fried egg snail rice noodles, braised egg snail rice noodles, and bomb snail rice noodles, where the image generation model struggles to grasp more details.
[0096] To improve the model's performance in specific sub-domains, this embodiment designs a LoRA branch network corresponding to specific categories. LoRA was initially applied in the natural language processing domain to fine-tune large language models. Since large models have over a hundred billion parameters, the training cost is too high. Therefore, LoRA employs a method of training only low-rank matrices. When in use, the parameters of the LoRA model are injected into the SD model, thereby changing the SD model's generation style, etc. This can be expressed using the formula W = W0 + BA, where W0 is the initial SD model's parameters (Weights), BA is the low-rank matrix (i.e., the LoRA model's parameters), and W represents the final SD model parameters after being influenced by the LoRA model. The BA matrix is the training target of the LoRA model. The LoRA model can fine-tune the Clip text encoder (the text feature extraction layer in the SD model) and some network layers in UNet, such as the linear layer of CrossAttention. The entire process is linear; it can be considered as the original SD model superimposed with the LoRA model to obtain a completely new model.
[0097] This embodiment designs a method where a diffusion model is first trained using a first training dataset. This first training dataset contains a large amount of data, such as training samples covering all categories. After the SD model is trained, one or more LoRA branch networks can be connected, where each LoRA branch network corresponds to a specific category. In practical applications, one or more specific categories can be configured as needed. Thus, under the control of each LoRA branch network, the SD model can output food images that match the specific category corresponding to each LoRA branch network.
[0098] The number of specific categories can be configured as needed. Specific categories may include those where the SD model struggles to capture details or has poor image quality. The SD model can be tested using a test dataset containing data from multiple categories. Food images generated by the SD model for different categories of food description information are obtained. The quality of the generated food images for different categories is judged based on the SD model's performance. This judgment can be made manually or by calculating the error between the image generated by the SD model and the corresponding food image sample. Multiple datasets can be prepared for each category, and the prediction results of the SD model for each category can be comprehensively evaluated to determine one or more categories with poor image generation quality as specific categories. Of course, this can also be configured as needed; for example, it could be barbecue ingredients or seafood, etc. This embodiment does not limit this.
[0099] Taking the setting of n specific categories as an example, for any category n among the n categories i Set up the corresponding second training dataset, where the training samples are categories n. i The second training dataset consists of food image samples and food description samples, which can be derived from the first training dataset used to train the diffusion model. The second training dataset can be small, ranging from tens to hundreds of samples, and can be configured as needed in practice. Training categories n i When training the corresponding LoRA branch network, the LoRA branch networks for the SD model and other categories do not participate in the training; their parameters can be frozen. In this way, the trained LoRA branch network can learn the association between the food description information and the food image for that category, thereby controlling the diffusion model to generate food images matching that category. In this embodiment, each LoRA branch network can be applied to the text feature extraction layer in the SD model and the query (key) and value (value) layers in the U-Net network, respectively.
[0100] After training, the image generation model can be used to generate images that correspond to the actual raw / cooked state of food. For example, by inputting prompt text containing "food name and raw / cooked state information" into the model, it can generate a matching food image. Optionally, the prompt text can also be "food name, food category, and raw / cooked state information," depending on the requirements. Figure 2C The diagram shown is an overall flowchart illustrating image generation according to an exemplary embodiment of this specification.
[0101] Optionally, based on the target food category in the prompt text, the LoRA branch network corresponding to the target food category can be selected from the trained model. Other LoRA branch networks are not selected and are in an inactive state. The SD model and the LoRA branch network corresponding to the target food category generate the corresponding food image based on the text features of the input prompt text.
[0102] In some examples, inputting the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text, may include:
[0103] Obtain the target food category contained in the prompt text;
[0104] In response to determining that the target food category has a corresponding target LoRA branch network, the other LoRA branch networks in the image generation model except for the target LoRA branch network are frozen, and the prompt text is input into the image generation model containing the diffusion model and the target LoRA branch network, so that the image generation model containing the diffusion model and the target LoRA branch network generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text;
[0105] In response to determining that there is no corresponding target LoRA branch network for the target food category, all LoRA branch networks in the image generation model are frozen, and the prompt text is input into the image generation model containing the diffusion model, so that the image generation model containing the diffusion model generates a food image that matches the target food name, the target food category, and the target raw / cooked state information based on the input prompt text.
[0106] In this embodiment, the image generation model contains multiple LoRA branch networks corresponding to different categories. Optionally, data recording specific categories and LoRA branch networks can be generated, such as a mapping dictionary. When using the model, the target food category in the prompt text may or may not have a corresponding LoRA branch network. Through the above embodiment, the model can query the mapping dictionary to see if there is a corresponding LoRA branch network based on the target food category in the prompt text, and control the working state of the LoRA branch network in the model.
[0107] In this embodiment, there are multiple ways to obtain the prompt text; for example, a user interface can be provided for the user to input food description information. As an example, a text input box corresponding to the food description information can be provided, through which the user can input the food description information. Optionally, the food description information can be further subdivided into food name, food category, or raw / cooked status, and corresponding text boxes can be provided for each subdivision, so that the user input information can be received through each text box. The user interface can also include a prompt message prompting the user to input the raw / cooked status information of the food.
[0108] Optionally, obtaining the prompt text containing food description information includes:
[0109] Obtain the target food name input by the user, and query the target raw / cooked status information corresponding to the target food name from the preset raw / cooked food database.
[0110] In this embodiment, the user may also choose not to input the target raw / cooked status information; a raw / cooked food database can be pre-built, which can store data of "category-dish name-raw / cooked food label", and the corresponding raw / cooked food status information can be retrieved based on the target food name input by the user.
[0111] Optionally, the step of obtaining the prompt text containing food description information may include:
[0112] In response to receiving food description information of the food to be configured from the store configuration page sent by the merchant client, a prompt text containing the food description information is generated using the food description information;
[0113] The output of the food image generated by the image generation model may include:
[0114] Configure the food images generated by the image generation model on the store configuration page.
[0115] This embodiment can be applied to scenarios where merchants configure their store's products. When a merchant uses a merchant client, the client provides a store configuration page, which can include functions for configuring dishes. Users can input food description information for the food to be configured, and this embodiment can generate prompt text. The raw / cooked state of the food can also be predicted based on the characteristics of the dishes sold in the merchant's store, for example, based on the categories of various dishes sold in the merchant's store, and / or based on the raw / cooked state of dishes sold in the merchant's store history; for example, if most of the dishes sold in the merchant's store history are cooked food, the target raw / cooked state information can be determined as cooked food, etc. The prompt text is input to an image generation model, which generates a food image and configures it on the store configuration page, thus eliminating the need for merchants to take and configure food images themselves.
[0116] like Figure 2D The diagram illustrates an image generation method according to an exemplary embodiment of this specification. The image generation model in this embodiment can be deployed on a user-side device or a physical device deployed on a server; this embodiment does not limit this. The user interface shows a text input box for "Please enter the dish name," allowing the user to enter the dish name. The diagram uses "steak" as an example. Optionally, the server can automatically obtain the cooked / raw status, or the user interface can also show a text input box for "Please enter raw or cooked food," allowing the user to enter the target cooked / raw status information. The user operates a Western restaurant; the diagram uses "cooked food" as an example.
[0117] The user interface provides two types of information: cooked and raw. The input boxes allow the model to capture these two types of information and generate prompt text, which is then fed into the image generation model. The model generates an image based on the input data. The image shown is a cooked steak, matching the dishes sold in the user's shop. If the user inputs raw food, the model will generate a raw steak.
[0118] like Figure 3 The diagram shown is a flowchart illustrating another image generation method according to an exemplary embodiment of this specification. The method may include:
[0119] Step 302: Obtain the target food name input by the user and send it to the server. The server is used to obtain the target raw or cooked status information that indicates whether the food is raw or cooked, and then execute the steps of the aforementioned image generation method.
[0120] Step 304: Obtain and display the food image output by the server.
[0121] This embodiment can be applied to a client-side application. The client can also obtain food images automatically generated by the server and configure them in the merchant's store, thereby reducing the number of product images required for the merchant's store. The specific implementation process of this embodiment can be referred to the foregoing embodiments, and will not be repeated here.
[0122] Corresponding to the embodiments of the aforementioned image generation method, this specification also provides embodiments of an image generation apparatus and the computer equipment on which it is applied.
[0123] like Figure 4 As shown, Figure 4 This is a block diagram illustrating an image generation apparatus according to an exemplary embodiment of the present specification. The apparatus includes:
[0124] The acquisition module 41 is used to acquire a prompt text containing food description information, wherein the food description information includes the target food name and target raw or cooked status information indicating whether the food is raw or cooked.
[0125] Image generation module 42 is used to input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw / cooked status information based on the input prompt text; wherein, the training samples of the image generation model include: food description samples and corresponding food image samples, the food description samples including food name samples and raw / cooked status information samples of the food in the food image samples;
[0126] Image output module 43 is used to output the food image generated by the image generation model.
[0127] In some examples, the food description information also includes the target food category; the food description sample also includes a food category information sample;
[0128] The image generation module 42 inputs the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw / cooked status information based on the input prompt text, including:
[0129] The prompt text is input into a pre-trained image generation model, which generates a food image that matches the target food name, the target food category, and the target raw / cooked status information based on the input prompt text.
[0130] In some examples, the image generation model includes a diffusion model and at least one LoRA branch network connected to the diffusion model; the image generation model is trained in the following manner:
[0131] The pre-defined diffusion model is trained using the first training dataset. A LoRA branch network corresponding to each specific category in at least one specific category is added to the trained diffusion model. The training samples in the first training dataset contain food image samples corresponding to multiple food categories.
[0132] For each of the LoRA branch networks, the following training is performed as the current LoRA branch network: the model parameters of the diffusion model and the network parameters of other LoRA branch networks besides the current LoRA branch network are frozen, and the current LoRA branch network is trained using the second training dataset of the specific category corresponding to the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples and the food image samples contained in the training samples in the second training dataset.
[0133] In some examples, the diffusion model includes a text encoder and an image generation network;
[0134] The diffusion model is trained by the text encoder to extract text features from the input food description samples, and the image generation network to generate a predicted image based on the text features, with the optimization objective being to minimize the error between the predicted image and the food image sample.
[0135] In some examples, the image generation module is specifically used for:
[0136] Obtain the target food category contained in the prompt text;
[0137] In response to determining that the target food category has a corresponding target LoRA branch network, the other LoRA branch networks in the image generation model except for the target LoRA branch network are frozen, and the prompt text is input into the image generation model containing the diffusion model and the target LoRA branch network, so that the image generation model containing the diffusion model and the target LoRA branch network generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text;
[0138] In response to determining that there is no corresponding target LoRA branch network for the target food category, all LoRA branch networks in the image generation model are frozen, and the prompt text is input into the image generation model containing the diffusion model, so that the image generation model containing the diffusion model generates a food image that matches the target food name, the target food category, and the target raw / cooked state information based on the input prompt text.
[0139] In some examples, the food description samples in the image generation module 42 are obtained in the following manner:
[0140] Retrieve food attribute information for various food products from the product database;
[0141] Obtain product images of each food product from the product database, and extract image attribute information from the product images;
[0142] After generating task prompt text based on the food attribute information and the image attribute information, the text is input into a pre-trained language model so that the language model outputs food description samples based on the task prompt text; wherein, the task prompt text is used to instruct the language model to convert the food attribute information and the image attribute information into natural language in the role of a food expert.
[0143] In some examples, extracting image attribute information from the product image includes a combination of one or more of the following steps:
[0144] The product image is input into a preset view recognition model, and the view type information of the product in the product image recognized by the view recognition model is obtained as image attribute information. The view type information includes frontal view, top view, or oblique view.
[0145] The product image is input into a preset raw / cooked state recognition model, and the raw / cooked state information of the product in the product image recognized by the raw / cooked state recognition model is obtained as image attribute information;
[0146] Pixel information in the product image is obtained to determine the background color in the product image as image attribute information.
[0147] In some examples, the acquisition module is specifically used to: acquire the target food name input by the user, and query the target raw / cooked status information corresponding to the target food name from a preset raw / cooked food database.
[0148] In some examples, the acquisition module is specifically used to: in response to receiving food description information of the food to be configured from the store configuration page sent by the merchant client, generate a prompt text containing the food description information using the food description information;
[0149] The image output module is specifically used to: configure the food image generated by the image generation model in the store configuration page.
[0150] The specific implementation process of the functions and roles of each module in the above-mentioned image generation device can be found in the implementation process of the corresponding steps in the above-mentioned image generation method, and will not be repeated here.
[0151] The embodiments of the image generation apparatus described in this specification can be applied to computer devices, such as servers or terminal devices. The apparatus embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by its processor reading the corresponding computer program instructions from non-volatile memory into memory and executing them. From a hardware perspective, such as... Figure 5 The diagram shown is a hardware structure diagram of a computer device containing the image generation apparatus described in this specification. (Except for...) Figure 5 In addition to the processor 510, memory 530, network interface 520, and non-volatile memory 540 shown, the computer device in which the image generation device 531 is located in the embodiment may also include other hardware depending on the actual function of the computer device, which will not be described in detail here.
[0152] Accordingly, this specification also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned image generation method embodiments.
[0153] Accordingly, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the image generation method embodiments.
[0154] Accordingly, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the image generation method embodiments.
[0155] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0156] The above embodiments can be applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0157] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0158] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0159] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0160] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0161] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0162] The terms "specific example" or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with the embodiments or examples, which are included in at least one embodiment or example of this specification. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0163] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this specification are indicated by the following claims.
[0164] It should be understood that this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is limited only by the appended claims.
[0165] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
Claims
1. An image generation method, the method comprising: Obtain prompt text containing food description information; wherein, the food description information includes the target food name, the target food category, and target raw or cooked status information indicating whether the food is raw or cooked. The prompt text is input into a pre-trained image generation model, which generates a food image that matches the target food name, the target food category, and the target raw / cooked status information based on the input prompt text. The training samples of the image generation model include food description samples and corresponding food image samples. The food description samples include food name samples, food category information samples, and raw / cooked status information samples of the food in the food image samples. Output the food image generated by the image generation model; The image generation model comprises a diffusion model and at least one LoRA branch network connected to the diffusion model; the image generation model is trained in the following manner: The pre-defined diffusion model is trained using a first training dataset containing food image samples corresponding to multiple food categories. The trained diffusion model is then supplemented with a LoRA branch network corresponding to each specific category in at least one specific category. Each LoRA branch network is trained as the current LoRA branch network as follows: the model parameters of the diffusion model and the network parameters of other LoRA branch networks besides the current LoRA branch network are frozen, and the current LoRA branch network is trained using the second training dataset of the specific category corresponding to the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples and the food image samples contained in the training samples of the second training dataset.
2. The method according to claim 1, wherein the diffusion model includes a text encoder and an image generation network; The diffusion model is trained by the text encoder to extract text features from the input food description samples, and the image generation network generates a predicted image based on the text features, with the optimization objective being to minimize the error between the predicted image and the food image sample.
3. The method according to claim 1, wherein inputting the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text, comprises: Obtain the target food category contained in the prompt text; In response to determining that the target food category has a corresponding target LoRA branch network, the other LoRA branch networks in the image generation model except for the target LoRA branch network are frozen, and the prompt text is input into the image generation model containing the diffusion model and the target LoRA branch network, so that the image generation model containing the diffusion model and the target LoRA branch network generates a food image matching the target food name and the target raw / cooked state information based on the input prompt text; In response to determining that there is no corresponding target LoRA branch network for the target food category, all LoRA branch networks in the image generation model are frozen, and the prompt text is input into the image generation model containing the diffusion model, so that the image generation model containing the diffusion model generates a food image that matches the target food name, the target food category, and the target raw / cooked state information based on the input prompt text.
4. The method according to claim 1, wherein the food description sample is obtained in the following manner: Retrieve food attribute information for various food products from the product database; Obtain product images of each food product from the product database, and extract image attribute information from the product images; After generating task prompt text based on the food attribute information and the image attribute information, the text is input into a pre-trained language model, so that the language model outputs food description samples based on the task prompt text; wherein... The task prompt text is used to instruct the language model to convert the food attribute information and the image attribute information into natural language in the role of a food expert.
5. The method according to claim 4, wherein extracting image attribute information from the product image comprises a combination of one or more of the following steps: The product image is input into a preset viewpoint recognition model, and the viewpoint type information of the product in the product image recognized by the viewpoint recognition model is obtained as image attribute information, wherein... The viewpoint type information includes a frontal viewpoint, a top-down viewpoint, or an oblique viewpoint. The product image is input into a preset raw / cooked state recognition model, and the raw / cooked state information of the product in the product image recognized by the raw / cooked state recognition model is obtained as image attribute information; Pixel information in the product image is obtained to determine the background color in the product image as image attribute information.
6. The method according to claim 1, wherein obtaining the prompt text containing food description information includes: Obtain the target food name input by the user, and query the target raw / cooked status information corresponding to the target food name from the preset raw / cooked food database.
7. The method according to claim 1, wherein obtaining the prompt text containing food description information includes: In response to receiving food description information of the food to be configured from the store configuration page sent by the merchant client, a prompt text containing the food description information is generated using the food description information; The output of the food image generated by the image generation model includes: Configure the food images generated by the image generation model on the store configuration page.
8. An image generation method, the method comprising: The system obtains the target food name input by the user and sends it to the server. The server then obtains the target raw or cooked state information that indicates whether the food is raw or cooked, and executes the steps of any one of the methods described in claims 1 to 7. Obtain and display the food images output by the server.
9. An image generation apparatus, the apparatus comprising: The acquisition module is used to acquire prompt text containing food description information, wherein the food description information includes the target food name, the target food category, and the target raw or cooked status information indicating whether the food is raw or cooked. An image generation module is used to input the prompt text into a pre-trained image generation model, so that the image generation model generates a food image matching the target food name, the target food category, and the target raw / cooked status information based on the input prompt text; wherein, the training samples of the image generation model include: food description samples and corresponding food image samples, the food description samples including food name samples, food category information samples, and raw / cooked status information samples of the food in the food image samples; The image output module is used to output the food image generated by the image generation model; The image generation model comprises a diffusion model and at least one LoRA branch network connected to the diffusion model; the image generation model is trained in the following manner: The pre-defined diffusion model is trained using a first training dataset containing food image samples corresponding to multiple food categories. The trained diffusion model is then supplemented with a LoRA branch network corresponding to each specific category in at least one specific category. Each LoRA branch network is trained as the current LoRA branch network as follows: the model parameters of the diffusion model and the network parameters of other LoRA branch networks besides the current LoRA branch network are frozen, and the current LoRA branch network is trained using the second training dataset of the specific category corresponding to the current LoRA branch network, so that the current LoRA branch network learns the association relationship between the specific category samples and the food image samples contained in the training samples of the second training dataset.
10. An image generation apparatus, the apparatus comprising: The acquisition module is used to: acquire the target food name input by the user and send it to the server, wherein the server is used to acquire the target raw or cooked status information representing whether the food is raw or cooked, and then execute the steps of any one of the methods described in claims 1 to 7. The display module is used to: acquire and display food images output by the server.
11. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
12. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method of any one of claims 1 to 8.