Method and System for Generating Fine-grained Pictures of Dishes

Through the ingredients sequence layer, pre-trained encoder and unet model combined with the online adaptive decoder, the problem of low dishes image quality caused by different ingredients sequence and cooking methods is solved, and high-quality, stable and real dishes image generation is achieved.

CN120107418BActive Publication Date: 2025-07-25JIANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510592047.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-25
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

When generating dishes images, the order of ingredients is different and the cooking methods are different, the image quality is low, the diversity and reality are lacking, and it is difficult to accurately bind the food properties and cooking methods, resulting in low visual error and text matching.

Method used

The food sequence information is obtained through the ingredient sequence layer, combined with the pre-trained encoder and unet model for inverse diffusion and denoising processing, the online adaptive decoder is used to generate refined pictures of the dishes, and the adaptive weight matrix and entity cross attention mechanism are introduced to bind the food properties, and a simplified text prompt method is adopted.

Benefits of technology

It improves the detail fidelity and realism of the dish images, enhances the stability and generalization ability of the generated images, ensures that the ingredient shape matches the cooking method, reduces computing resource consumption, and improves the quality and consistency of the generated dishes images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107418B_ABST
    Figure CN120107418B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for generating refined dish pictures, belonging to the field of artificial intelligence technology, including: obtaining ingredient order information through an ingredient order layer according to the order of ingredient attributes, where the ingredient order layer is used to model the order of ingredient attributes to generate ingredient and ingredient order prompt information; performing encoding processing on multiple ingredient text prompt word information and multiple ingredient attribute binding information through a pre-trained encoder to obtain multiple embedding vectors and multiple conditional feature vectors respectively; performing inverse diffusion and denoising processing on a preset noise vector, embedding vector, and conditional feature vector through a preset unet model to generate latent image features; and performing decoding processing on the latent image features through an online adaptive decoder to generate refined dish pictures. The present invention can improve the quality of the generated dish images when facing different ingredient orders and cooking methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and specifically relates to a method and system for generating refined dish pictures. Background Art

[0002] With the rapid development of artificial intelligence and computer vision technologies, image generation technology has become an important research direction in multiple fields. In the catering industry, high-quality dish pictures have an important impact on consumer decisions. However, there are many limitations in traditional dish image acquisition methods. In cross-modal dish image generation technology, text-to-image generation is a core technology, mainly aiming to generate corresponding dish images through natural language descriptions, and breakthrough progress has been made in this field. Currently, in the task of generating dish images based on descriptions, the research mainly focuses on the application of GAN.

[0003] However, the existing technologies have the following problems: First, the image quality generated by existing GANs is limited, which is prone to mode collapse, resulting in a limited variety of generated images and lack of diversity; Second, GAN may not be able to handle the fine details of ingredients well in the food image generation task, resulting in a lack of realism in the generated images; Third, traditional diffusion models may not be able to accurately capture the color, texture, and shape of ingredients without specific optimization, resulting in dish images that do not match reality; Fourth, it is difficult for existing methods to accurately bind ingredient attributes, and different cooking methods may lead to visual errors. For example, "scrambled eggs" and "boiled eggs" in scrambled eggs with tomatoes are easily confused, resulting in the model generating incorrect images. Fifth, existing methods lack modeling of the ingredient order. Traditional methods mainly generate images based on text descriptions, but ignore the impact of the ingredient addition order on the final image. For example, adding ketchup first and adding ketchup later may affect the overall color distribution, and it is difficult for existing technologies to capture such subtle changes; Sixth, the matching degree between text and image is low. Due to structural limitations, GAN is difficult to achieve accurate text-to-image conversion, and traditional diffusion models may generate images that do not fully match the text description without optimization, resulting in a large deviation between the ingredient form, arrangement, or texture and user expectations. The above technical problems will lead to a low quality of generated dish images when the ingredient order and cooking methods are different. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method and system for generating refined dish pictures, which can solve the problem of low quality of generated dish images in the existing technology when the ingredient order and cooking methods are different.

[0005] To solve the above technical problems, this application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for generating refined pictures of dishes, the method including:

[0007] According to the ingredient attribute order, obtain ingredient order information through the ingredient order layer, where the ingredient order layer is used to perform modeling processing on the ingredient attribute order to generate ingredients and ingredient order prompt information;

[0008] According to multiple ingredient text prompt word information and multiple ingredient attribute binding information, perform encoding processing through a pre-trained encoder to respectively obtain multiple embedding vectors and multiple conditional feature vectors;

[0009] Perform inverse diffusion and denoising processing on a preset noise vector, the embedding vector, and the conditional feature vector through a preset unet model to generate latent image features;

[0010] According to the latent image features, perform decoding processing through an online adaptive decoder to generate refined pictures of dishes.

[0011] As an optional implementation manner of the first aspect of the present application, the ingredient order layer includes an ingredient composition embedding layer, a first fully connected layer, and a second fully connected layer. The ingredient composition embedding layer is used to embed the ingredient order at the end of the ingredient to generate ingredients and ingredient order. The first fully connected layer is used to map the ingredients and ingredient order to generate an embedding vector. The second fully connected layer is used to perform mapping processing on the text order embedding obtained by multiplying the embedding vector and a preset adaptive weight matrix to generate ingredient and ingredient order prompt information.

[0012] As an optional implementation manner of the first aspect of the present application, the unet model includes an entity cross-attention layer and a denoising layer. The entity cross-attention layer is used to perform inverse diffusion processing on the noise vector, the embedding vector, and the conditional feature vector to generate an intermediate latent representation. The denoising layer is used to perform denoising processing on the intermediate latent representation to generate latent image features.

[0013] As an optional implementation manner of the first aspect of the present application, the process of performing decoding processing through an online adaptive decoder to generate refined pictures of dishes includes: inputting the latent image features into the online adaptive decoder to obtain generated dish pictures; inputting the generated dish pictures into a text-graphic describer to perform text extraction to obtain text information; obtaining a spherical distance loss function based on the text information and a preset Prompt prompt word feature; updating the weights of the online adaptive decoder based on the spherical distance loss function and a preset adversarial loss function to obtain an updated online adaptive decoder; performing training fine-tuning on multiple generated dish images based on the updated online adaptive decoder to generate refined pictures of dishes.

[0014] As an alternative implementation of the first aspect of the present application, the process of obtaining the potential image features includes: binding and mapping the ingredient names and ingredient attributes in the prompt text to obtain the weighted sum of each ingredient and ingredient attribute; obtaining the weight attention degree between the generated picture and the prompt text based on the weighted sum of each ingredient and ingredient attribute; constructing the prompt text based on a preset template; obtaining the weight attention degree between the key vector of the prompt text and the word vector of the prompt text; obtaining the cross-attention score based on the weight attention degree between the generated picture and the prompt text and the weight attention degree between the key vector of the prompt text and the word vector of the prompt text; mapping the cross-attention score to an online adaptive decoder for denoising processing to obtain the potential image features.

[0015] As an alternative implementation of the first aspect of the present application, the mathematical expression of the cross-attention score is:

[0016] ;

[0017] where, represents the cross-attention score, represents the weight, represents the number of ingredients, represents the weighted sum of each ingredient and ingredient attribute, represents the activation function, represents the generated image features, represents the key vector of the prompt text, represents the word vector of the prompt text, represents the dimension, represents the matrix transpose, represents the key.

[0018] As an alternative implementation of the first aspect of the present application, the mathematical expression of the template is:

[0019] ;

[0020] where, represents the mathematical expression of the template, represents other types of text, represents multiple cooking methods of the dish, represents multiple ingredients of the dish.

[0021] In a second aspect, an embodiment of the present application provides a refined dish picture generation system, and the system includes:

[0022] An ingredient order module, configured to input the ingredient attribute order into an ingredient order layer for modeling processing to obtain ingredient order information;

[0023] A first pre-training encoder module for encoding multiple food ingredient text prompt information to generate multiple embedding vectors;

[0024] A second pre-training encoder module for encoding multiple food ingredient attribute binding information to generate multiple conditional feature vectors;

[0025] A UNet module for processing a preset noise vector, the embedding vector, and the conditional feature vector to generate latent image features;

[0026] An online adaptive decoding module for decoding the latent image features to obtain a refined dish picture;

[0027] The food ingredient order module, the first pre-training encoder module, the second pre-training encoder module, the UNet module, and the online adaptive decoding module jointly build an AdaDMOE model.

[0028] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0029] In a fourth aspect, an embodiment of the present application provides a readable storage medium with a program or instruction stored thereon. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0030] Compared with the prior art, the beneficial effects of a method for generating a refined dish picture proposed by the present invention are as follows: By introducing an online adaptive decoder, combining a multi-scale discriminator and text feature matching, the present invention improves the detail fidelity of the generated dish image, making it clearer, more realistic, and conforming to the text prompt information. By adopting an adaptive weight matrix, the influence of the food ingredient order on the finally generated image is optimized, so that different combinations of food ingredients in different orders are more visually in line with the actual cooking habits, enhancing the realism of the dish image. Through the entity relationship cross-attention mechanism, the food ingredient components and cooking methods are bound to ensure that the visual form of the food ingredients conforms to the corresponding cooking methods, avoiding incorrect matching problems such as scrambled eggs being incorrectly generated as boiled eggs. At the same time, this technology adopts a concise text prompt method based on template sentences, which significantly reduces the consumption of computing resources compared with traditional long text descriptions, improves the model processing efficiency, and ensures the accuracy and consistency of the generated dish image. In addition, this solution ignores non-critical factors such as the serving container and background elements, improving the stability and generalization ability of the generated image, enabling the model to generate dish images that meet expectations under different text prompts, and overall improving the quality and application value of the dish image generation model. Description of the Drawings

[0031] Figure 1 It is a flowchart of a method for generating refined dish pictures provided by the first embodiment of the present application;

[0032] Figure 2 It is a diagram of the AdaDMOE model provided by the first embodiment of the present application;

[0033] Figure 3 It is the internal structure diagram of a system for generating refined dish pictures provided by the second embodiment of the present application. Detailed implementation manners

[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0035] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0036] Next, in conjunction with the accompanying drawings, a method and system for generating refined dish pictures provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0037] Embodiment 1

[0038] Please refer to Figure 1 , which is a flowchart of a method for generating refined dish pictures proposed in the first embodiment of the present application. The proposed method includes steps S1 to S4.

[0039] Step S1: According to the order of ingredient attributes, obtain ingredient order information through the ingredient order layer, and the ingredient order layer is used to model the order of ingredient attributes to generate ingredient and ingredient order prompt information.

[0040] Specifically, the ingredient sequence layer includes an ingredient component embedding layer, a first fully connected layer, and a second fully connected layer. The ingredient component embedding layer is used to embed the ingredient sequence into the end of the ingredient to generate the ingredient and the ingredient sequence. The first fully connected layer is used to map the ingredient and the ingredient sequence to generate an embedding vector. The second fully connected layer is used to perform a mapping process on the text sequence embedding obtained by multiplying the embedding vector by a preset adaptive weight matrix to generate ingredient and ingredient sequence prompt information.

[0041] The present invention uses sine and cosine functions to generate the encoding for each sequence, and this method has a certain regularity among different sequences. For the sequence and the embedding dimension , the mathematical expression of the sequence encoding is:

[0042] ;

[0043] ;

[0044] Among them, represents the order of the word, represents the index of the embedding dimension, and the range of this index is to , is the channel dimension of the pre-trained CLIP encoder. represents the cosine sequence embedding, represents the sine sequence embedding.

[0045] Furthermore, the process of generating the embedding vector includes:

[0046] ;

[0047] Among them, represents the embedding vector, represents the first fully connected layer, represents the ingredient sequence and the ingredient word vector, represents the sequence encoding.

[0048] Furthermore, the process of performing a mapping process on the text sequence embedding obtained by multiplying the embedding vector by a preset adaptive weight matrix to generate ingredient and ingredient sequence prompt information includes:

[0049] ;

[0050] Among them, represents the ingredient and ingredient sequence prompt information, represents the second fully connected layer, represents the adaptive weight matrix 's initial coefficient.

[0051] Since the order of different ingredients has a greater or lesser impact on the appearance of the final generated dish image, the research of the present invention aims to control the fine-grained impact of ingredient components in different orders. The adaptive weight matrix preset during the entire training process of generating the dish image is continuously updated. The adaptive weight matrix is multiplied by the ingredient order and the ingredient word vector, and then passes through the second fully connected layer to obtain the ingredient and ingredient order prompt information .

[0052] The input of the ingredient order layer of the present invention is the ingredient component and the ingredient order embedding, and the order information embedding of the ingredient is added to the end of each ingredient. The ingredient order layer of the present invention does not require training. After the normalized text passes through the NLTK natural language tool, the ingredients and cooking methods can be obtained. The ingredients and cooking methods are matched according to the principle of proximity. For example, "stir-fried beef with celery" will become "stir-fried celery" and "stir-fried beef", binding the ingredients and the ingredient relationship, and extracting the order information.

[0053] In addition, during the process of generating the dish image, it is not only necessary to include the meaning of the ingredients, but also to represent the position and order of the ingredients in the sentence. In order to emphasize the different added ingredient orders and control the fine-grained consistency of the finally generated dish pictures, the present invention introduces an online adaptive encoder and an adaptive weight matrix. The online adaptive encoder receives the ingredient and the ingredient position order embedding, and the ingredient and the ingredient position order embedding are added to the end of the ingredient text embedding to improve the authenticity of the finally generated dish image due to different ingredient addition orders. Finally, the adaptive weight matrix is used to adapt to different prompt order information, enhancing the generalization ability.

[0054] Step S2: According to multiple ingredient text prompt word information and multiple ingredient attribute binding information, perform encoding processing through a pre-trained encoder to obtain multiple embedding vectors and multiple conditional feature vectors respectively.

[0055] Specifically, the pre-trained encoder of the present invention specifically refers to a pre-trained CLIP encoder. The CLIP encoder is a multi-modal pre-trained model that embeds images and texts into the same semantic space through contrastive learning. Its core idea is to achieve cross-modal semantic alignment through contrastive training of a large-scale image-text pair. The present invention divides the CLIP encoder into a first pre-trained CLIP encoder and a second pre-trained CLIP encoder. The first pre-trained CLIP encoder is used to encode multiple ingredient text prompt word information to obtain multiple embedding vectors. The generated embedding vectors capture the semantic features and ingredient order features of the text description. The second pre-trained CLIP encoder is used to encode multiple ingredient attribute binding information to obtain multiple conditional feature vectors. The multiple embedding vectors and multiple conditional feature vectors are used to guide the generation of the latent representation vector.

[0056] Step S3: Perform inverse diffusion and denoising processing on the preset noise vector, embedding vector, and conditional feature vector through a preset UNet model to generate latent image features.

[0057] The UNet model used in the present invention specifically includes multiple layers of entity cross-attention layers and one denoising layer. The entity cross-attention layer is used to perform inverse diffusion processing on the noise vector to obtain an intermediate latent representation , and the denoising layer is used to perform denoising processing on the intermediate latent representation to obtain latent image features .

[0058] Specifically, the specific process of obtaining latent image features includes: binding and mapping the ingredient names and ingredient attributes in to obtain the weighted sum of each ingredient and ingredient attribute; obtaining the weight attention between the generated picture and the prompt text based on the weighted sum of each ingredient and ingredient attribute; constructing the prompt text based on a preset template; obtaining the weight attention between the key vector of the prompt text and the word vector of the prompt text; obtaining the cross-attention score based on the weight attention between the generated picture and the prompt text and the weight attention between the key vector of the prompt text and the word vector of the prompt text; mapping the cross-attention score into an online adaptive decoder for denoising processing to obtain latent image features.

[0059] Specifically, the mathematical expression for the weighted sum of each ingredient and ingredient attribute is:

[0060] ;

[0061] where represents the weighted sum of each ingredient and ingredient attribute, represents the number of ingredients.

[0062] Further, the process of obtaining the weight attention between the generated picture and the prompt text includes: mapping into to obtain the weight attention between the generated picture and the prompt text, and the mathematical expression is:

[0063] ;

[0064] where represents the weight attention between the generated picture and the prompt text, represents the generated image features, represents the key, represents matrix transpose, represents the dimension, representation activation function

[0065] Furthermore, the process of obtaining the weighted attention between the key vector of the prompt text and the word vector of the prompt text includes: By aligning the ingredients and ingredient attributes with the prompt text, the noise of some background factors is ignored. To enhance the generalization of the text prompt and weaken the influence of the dish container and background elements on the generated image. Then, an alignment and cross-fusion are performed on the original input text and the generated image, and the mathematical expression is:

[0066] ;

[0067] where represents the weighted attention between the key vector of the prompt text and the word vector of the prompt text, represents the key vector of the prompt text, represents the word vector of the prompt text.

[0068] Furthermore, the process of obtaining the cross-attention score includes:

[0069] ;

[0070] where represents the cross-attention score, represents the weight, represents the number of ingredients, represents the weighted sum of each ingredient and ingredient attribute, represents the generated image feature, represents the key vector of the prompt text, represents the word vector of the prompt text, represents the dimension, represents the matrix transpose, represents the key.

[0071] Furthermore, when constructing the prompt text based on a preset template, the mathematical expression of the preset template is:

[0072] ;

[0073] where represents the mathematical expression of the template, represents other types of text, represents multiple cooking methods of the dish, represents multiple ingredients of the dish.

[0074] Step S4: According to the latent image feature, perform decoding processing through an online adaptive decoder to generate a refined picture of the dish.

[0075] The online adaptive decoder of the present invention is used to receive latent image features to generate multiple dish images, and perform training fine-tuning on each generated dish image to obtain updated generated dish images. The online adaptive decoder used in the present invention includes three parts. One is a pre-trained online adaptive decoder, and the weight parameters of the pre-trained online adaptive decoder are updated during each training fine-tuning. Another is a trainable multi-scale discriminator, and the parameters of the discriminator are not updated. The other is a text-image descriptor, which is used to perform picture extraction processing on the dish images generated by the online adaptive decoder to obtain picture extraction text.

[0076] Specifically, the process of generating refined dish pictures includes: inputting latent image features into the online adaptive decoder to obtain generated dish pictures; inputting the generated dish pictures into the text-image descriptor for text extraction to obtain text information; obtaining a spherical distance loss function based on the text information and preset Prompt prompt word features; updating the weights of the online adaptive decoder based on the spherical distance loss function and a preset adversarial loss function to obtain an updated online adaptive decoder; performing training fine-tuning on multiple generated dish images based on the updated online adaptive decoder to generate refined dish pictures.

[0077] Spherical distance loss function The mathematical expression of

[0078] ;

[0079] Among them, is used to guide the generated image to be more in line with the text prompt, represents the picture extraction text, from the generated dish pictures, is the prompt text Prompt.

[0080] Preset adversarial loss function The mathematical expression of

[0081] ;

[0082] Among them, represents the generated picture, represents the real image, represents the expected calculation, represents the result of encoding the generated picture, represents the result of encoding the real image.

[0083] The weighted spherical distance loss function of the present invention and the preset adversarial loss function are used as the total loss function (satisfying the relationship ) and utilize this total loss function to update the weights of the online adaptive decoder, obtain the updated online adaptive decoder, and then, based on the updated online adaptive decoder, perform training fine-tuning on multiple generated dish images to obtain the updated generated dish images.

[0084] In summary, the beneficial effects of step S4 of the present invention are as follows:

[0085] First of all, the present invention utilizes the total loss function to update the weights of the online adaptive decoder, which can fine-tune each generated dish picture, improving the detail fidelity of the generated dish pictures.

[0086] Secondly, the present invention also uses two conditional mechanisms to bind the attribute relationships between ingredients. The first conditional mechanism is: describe the basic information of the dish, its ingredients, and the cooking method. The second conditional mechanism is: the dish ingredients and the cooking method are bound together in a weighted manner, and the attributes of each ingredient are bound to the ingredient. For example, "stir-fry" and "egg" in scrambled eggs are an ingredient and an ingredient attribute, and eggs will not appear alone. By introducing these two conditional mechanisms, the model can accurately match the dish image with the corresponding text description, especially in terms of image-text consistency.

[0087] Then, during the entire process of generating dish images, the text prompt information is completed using a pre-trained CLIP model. However, in existing models, it may be that the dataset does not clearly represent the attributes of ingredients or the text encoder has inaccurate understanding of long texts. The output image of scrambled eggs with tomatoes may include an image of boiled eggs. It is necessary to study and solve the problem of different visual forms of ingredients under different cooking methods and strengthen the bundling of ingredient attributes after different cooking methods. The present invention obtains the cross-attention score in the cross-attention layer. The present invention uses the entity relationship cross-attention layer as a key component for image generation to process the relationships between the dish name, dish ingredients, and dish attributes. This layer fuses the attribute features between ingredients through the entity cross-attention layer, thereby generating a dish image that conforms to the text description.

[0088] Finally, in traditional dish image generation models, the input prompt text is often a long and complex text description, which requires extremely large computing resources. To generate dish images, the features required by the model are only ingredients, the cooking method of ingredients, and the input order of ingredients. And the main architecture for processing text prompts is Transformer, and the computational complexity of the self-attention mechanism of Transformer is , when generating images according to long text descriptions, the computational resource consumption increases exponentially during processing. To solve this problem, the input of the present invention is a concise prompt text based on a template sentence pattern. This method generates a structured and consistent prompt text by extracting the key elements of ingredients and their cooking methods. The present invention constructs the prompt text based on a preset template, which has higher conciseness compared to traditional long text descriptions. This template design is simple and clear, and it is easier to obtain text features during the actual processing, which can optimize the text prompt and improve the computational efficiency.

[0089] The present invention jointly constructs an AdaDMOE model based on the ingredient order layer, the first pre-trained encoder layer, the second pre-trained encoder layer, the unet layer, and the online adaptive decoding layer. The AdaDMOE model of the present invention is a generative model based on the latent space. By modeling the ingredient order of the prompt text, binding ingredient attributes, and then controlling the generation of latent space feature images through the pre-trained model, and finally fine-tuning the latent space features through the online adaptive decoder to generate refined dish pictures. The AdaDMOE model proposed by the present invention is as Figure 2 shown.

[0090] Embodiment 2

[0091] Please refer to Figure 3 , which shows a schematic structural diagram of a refined dish picture generation system proposed in the second embodiment of the present application. The system includes:

[0092] The ingredient order module 100 is used to input the ingredient attribute order into the ingredient order layer for modeling processing to obtain ingredient order information;

[0093] The first pre-trained encoder module 200 is used to encode multiple ingredient text prompt word information to generate multiple embedding vectors;

[0094] The second pre-trained encoder module 300 is used to encode multiple ingredient attribute binding information to generate multiple conditional feature vectors;

[0095] The unet module 400 is used to process the preset noise vector, embedding vector, and conditional feature vector to generate latent image features;

[0096] The online adaptive decoding module 500 is used to decode the latent image features to obtain refined dish pictures;

[0097] The ingredient order module 100, the first pre-trained encoder module 200, the second pre-trained encoder module 300, the unet module 400, and the online adaptive decoding module 500 jointly construct an AdaDMOE model.

[0098] A refined dish picture generation system in an embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make specific limitations.

[0099] A refined dish picture generation system in an embodiment of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiment of the present application does not make specific limitations.

[0100] A refined dish picture generation system provided in an embodiment of the present application can implement Figures 1 to 2 each process implemented by a refined dish picture generation method in the method embodiment. To avoid repetition, it will not be elaborated here.

[0101] The beneficial effects of a refined dish picture generation system provided by the present invention are as follows: First, by introducing an online adaptive decoding module, combining a multi-scale discriminator and a text feature matching mechanism, the detail fidelity of the generated dish image is improved, making the generated image more conform to the visual characteristics of real dishes; Second, by processing a preset noise vector, an embedding vector, and a conditional feature vector through a U-Net module to generate potential image features, the stability and generalization ability of the generated image can be improved, enabling the model to generate dish images that meet expectations under different text prompts, and overall improving the quality and application value of the dish image generation model.

[0102] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned refined dish picture generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0103] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of the method for generating refined pictures of dishes and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0104] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.

[0105] It should be noted that in the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0107] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A method for generating refined pictures of dishes, characterized in that, Including: According to the ingredient attribute order, obtain ingredient order information through the ingredient order layer, where the ingredient order layer is used to model the ingredient attribute order to generate ingredients and ingredient order prompt information; The ingredient order layer includes an ingredient composition embedding layer, a first fully connected layer, and a second fully connected layer. The ingredient composition embedding layer is used to embed the ingredient order at the end of the ingredient to generate ingredients and ingredient order. The first fully connected layer is used to map the ingredients and ingredient order to generate an embedding vector. The second fully connected layer is used to perform a mapping process on the text order embedding obtained by multiplying the embedding vector by a preset adaptive weight matrix to generate ingredients and ingredient order prompt information; According to multiple ingredient text prompt word information and multiple ingredient attribute binding information, perform encoding processing through a pre-trained encoder to obtain multiple embedding vectors and multiple conditional feature vectors respectively; Perform inverse diffusion and denoising processing on a preset noise vector, the embedding vector, and the conditional feature vector through a preset unet model to generate latent image features; The unet model includes an entity cross-attention layer and a denoising layer. The entity cross-attention layer is used to perform inverse diffusion processing on the noise vector, the embedding vector, and the conditional feature vector to generate an intermediate latent representation. The denoising layer is used to perform denoising processing on the intermediate latent representation to generate latent image features; According to the latent image features, perform decoding processing through an online adaptive decoder to generate a refined dish picture. Among them, the process of performing decoding processing through an online adaptive decoder to generate a refined dish picture includes: inputting the latent image features into the online adaptive decoder to obtain a generated dish picture; inputting the generated dish picture into a text descriptor to perform text extraction to obtain text information; obtaining a spherical distance loss function based on the text information and a preset Prompt prompt word feature; updating the weights of the online adaptive decoder based on the spherical distance loss function and a preset adversarial loss function to obtain an updated online adaptive decoder; performing training fine-tuning on multiple generated dish images based on the updated online adaptive decoder to generate a refined dish picture.

2. The method for generating refined pictures of dishes according to claim 1, wherein The process of obtaining the latent image features includes: Bind and map the ingredient names and ingredient attributes in the prompt text to obtain the weighted sum of each ingredient and ingredient attribute; Based on the weighted sum of each ingredient and ingredient attribute, obtain the weight attention between the generated picture and the prompt text; Construct a prompt text based on a preset template; Obtain the weight attention between the key vector of the prompt text and the word vector of the prompt text; Obtain a cross-attention score based on the weight attention between the generated picture and the prompt text and the weight attention between the key vector of the prompt text and the word vector of the prompt text; Map the cross-attention score into the online adaptive decoder for denoising processing to obtain latent image features.

3. A method for generating refined pictures of dishes according to claim 2, characterized in that, The mathematical expression of the cross-attention score is: ; Among them, represents the cross-attention score, represents the weight, represents the number of food ingredients, represents the weighted sum of each food ingredient and food ingredient attributes, represents the activation function, represents the generated image features, represents the key vector of the prompt text, represents the word vector of the prompt text, represents the dimension, represents the matrix transpose, represents the key.

4. A method for generating refined pictures of dishes according to claim 2, characterized in that, The mathematical expression of the template is: ; Among them, represents the mathematical expression of the template, represents other types of text, represents multiple cooking methods of the dish, represents multiple ingredients of the dish.

5. A refined picture generation system for dishes, characterized in that, Including: The ingredient order module is used to input the ingredient attribute order into the ingredient order layer for modeling processing to obtain ingredient order information; The ingredient order module includes an ingredient component embedding layer, a first fully connected layer, and a second fully connected layer. The ingredient component embedding layer is used to embed the ingredient order to the end of the ingredient to generate an ingredient and an ingredient order. The first fully connected layer is used to map the ingredient and the ingredient order to generate an embedding vector. The second fully connected layer is used to perform a mapping process on the text order embedding obtained by multiplying the embedding vector by a preset adaptive weight matrix to generate ingredient and ingredient order prompt information; The first pre-trained encoder module is used to encode multiple ingredient text prompt word information to generate multiple embedding vectors; The second pre-trained encoder module is used to encode multiple ingredient attribute binding information to generate multiple conditional feature vectors; The unet module is used to process a preset noise vector, the embedding vector, and the conditional feature vector to generate latent image features; the unet module includes an entity cross-attention layer and a denoising layer. The entity cross-attention layer is used to perform inverse diffusion processing on the noise vector, the embedding vector, and the conditional feature vector to generate an intermediate latent representation. The denoising layer is used to perform denoising processing on the intermediate latent representation to generate latent image features; The online adaptive decoding module is used to decode the latent image features to obtain refined dish pictures; wherein, the process of decoding through the online adaptive decoding module to generate refined dish pictures includes: inputting the latent image features into the online adaptive decoder to obtain generated dish pictures; inputting the generated dish pictures into a text-graphic descriptor to perform text extraction to obtain text information; obtaining a spherical distance loss function based on the text information and a preset Prompt prompt word feature; updating the weights of the online adaptive decoder based on the spherical distance loss function and a preset adversarial loss function to obtain an updated online adaptive decoder; training and fine-tuning multiple generated dish images based on the updated online adaptive decoder to generate refined dish pictures; The ingredient order module, the first pre-trained encoder module, the second pre-trained encoder module, the unet module, and the online adaptive decoding module jointly build the AdaDMOE model.

6. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for generating refined dish pictures as described in any one of claims 1-4 are implemented.

7. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of a method for generating refined dish pictures as described in any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • A method of generating food image from recipe

    CN112017255A

  • Image generation method and device, equipment and storage medium

    CN117710530A