Dish refined picture generation method and system
By introducing a food order layer and an online adaptive decoder into the dish image generation system, combining multi-scale discriminator and text feature matching, the impact of food order and cooking methods on image quality is solved, and a higher quality and realistic dish image generation is achieved.
Patent Information
- Application Number
- CN202510592047.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, when generating dishes images, the order of ingredients is different and the cooking methods are different, the generated images are of low quality and lack realism and diversity.
By introducing a food order layer and an online adaptive decoder, combining multi-scale discriminator and text feature matching, the impact of food order on images is optimized, and the ingredients and cooking methods are bound through the entity relationship cross attention mechanism.
It improves the detail fidelity and realism of the dish images, makes the generated images clearer and conforms to text prompt information, and enhances the generalization ability of the model and the quality of image generation.
Smart Images

Figure CN120107418A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a method and system for generating refined images of dishes. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technology, image generation technology has become an important research direction in many fields. In the catering industry, high-quality food images have an important impact on consumer decisions, but the traditional way of obtaining food images has many limitations. In cross-modal food image generation technology, text-generated images are a core technology. The main purpose is to generate corresponding food images through natural language descriptions, and this field has made breakthrough progress. At present, in the task of generating food images based on descriptions, research mainly focuses on the application of GAN.
[0003] However, the existing technology has the following problems: First, the quality of the images generated by the existing GAN is limited, which can easily lead to mode collapse, resulting in a limited number of generated images and a lack of diversity; Second, GAN may not be able to handle the fine details of the ingredients well in the task of generating food images, resulting in the lack of realism in the generated images; Third, the traditional diffusion model may not be able to accurately capture the color, texture and shape of the ingredients without specific optimization, resulting in the generated dish images not being consistent with reality; Fourth, it is difficult for the existing methods to accurately bind the attributes of the ingredients, and different cooking methods may lead to visual errors. For example, the "scrambled eggs" and "boiled eggs" in tomato scrambled eggs are easily confused, resulting in the model generating wrong images. Fifth, the existing methods lack modeling of the order of ingredients. Traditional methods mainly generate images based on text descriptions, but ignore the impact of the order of adding ingredients on the final image. For example, putting ketchup first or later may affect the overall color distribution, and it is difficult for existing technologies to capture such subtle changes; Sixth, the matching degree between text and image is low. Due to structural limitations, GAN is difficult to achieve accurate text-to-image conversion, and traditional diffusion models may generate images that do not fully match the text description without optimization, resulting in a large deviation between the shape, arrangement or texture of the ingredients and user expectations. The above technical problems will lead to the problem of low quality of the generated dish images when the order of ingredients is different and the cooking methods are different. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a method and system for generating refined images of dishes, which can solve the problem in the prior art that the quality of the generated dish images is low when the order of ingredients is different and the cooking methods are different.
[0005] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a method for generating a refined dish image, the method comprising: According to the order of ingredient attributes, ingredient order information is obtained through an ingredient order layer, wherein the ingredient order layer is used to model the order of ingredient attributes to generate ingredient and ingredient order prompt information; According to the information of multiple food text prompt words and the information of multiple food attribute bindings, encoding is performed through a pre-trained encoder to obtain multiple embedding vectors and multiple conditional feature vectors respectively; Performing inverse diffusion and denoising processing on a preset noise vector, the embedding vector, and the conditional feature vector through a preset unet model to generate potential image features; According to the potential image features, decoding processing is performed through an online adaptive decoder to generate a refined picture of the dish.
[0006] As an optional implementation of the first aspect of the present application, the ingredient sequence layer includes an ingredient component embedding layer, a first fully connected layer and a second fully connected layer. The ingredient component embedding layer is used to embed the ingredient sequence into the end of the ingredient to generate ingredients and ingredient sequence. The first fully connected layer is used to map the ingredients and ingredient sequence to generate an embedding vector. The second fully connected layer is used to map the text sequence embedding obtained by multiplying the embedding vector and a preset adaptive weight matrix to generate ingredient and ingredient sequence prompt information.
[0007] As an optional implementation of the first aspect of the present application, the UNET model includes an entity cross-attention layer and a denoising layer. The entity cross-attention layer is used to perform inverse diffusion processing on the noise vector, the embedding vector, and the conditional feature vector to generate an intermediate latent representation, and the denoising layer is used to perform denoising processing on the intermediate latent representation to generate latent image features.
[0008] As an optional implementation of the first aspect of the present application, the process of decoding through an online adaptive decoder to generate a refined dish picture includes: inputting the potential image features into the online adaptive decoder to obtain a generated dish picture; inputting the generated dish picture into a graphic descriptor to extract text and obtain text information; obtaining a spherical distance loss function based on the text information and a preset Prompt word feature; updating the weights of the online adaptive decoder based on the spherical distance loss function and a preset adversarial loss function to obtain an updated online adaptive decoder; and training and fine-tuning multiple generated dish images based on the updated online adaptive decoder to generate a refined dish picture.
[0009] As an optional implementation of the first aspect of the present application, the process of obtaining the potential image features includes: binding and mapping the ingredient names and ingredient attributes in the prompt text to obtain the weighted sum of each ingredient and ingredient attribute; based on the weighted sum of each ingredient and ingredient attribute, obtaining the weighted attention between the generated picture and the prompt text; constructing the prompt text based on a preset template; obtaining the weighted attention between the key vector of the prompt text and the word vector of the prompt text; obtaining a cross-attention score based on the weighted attention between the generated picture and the prompt text, and the weighted attention between the key vector of the prompt text and the word vector of the prompt text; mapping the cross-attention score to an online adaptive decoder for denoising to obtain the potential image features.
[0010] As an optional implementation of the first aspect of the present application, the mathematical expression of the cross attention score is: ; in, represents the cross attention score, represents the weight, Indicates the number of ingredients. represents the weighted sum of each ingredient and ingredient attribute, express Activation function, represents the generated image features, A key vector representing the prompt text, The word vector representing the prompt text, Represents the dimension, represents the matrix transpose, Indicates a key.
[0011] As an optional implementation of the first aspect of the present application, the mathematical expression of the template is: ; in, A mathematical expression representing the template, Represents other types of text, Multiple ways to cook representative dishes. Represents multiple ingredients of a dish.
[0012] In a second aspect, an embodiment of the present application provides a system for generating a refined dish image, the system comprising: The ingredient sequence module is used to input the ingredient attribute sequence into the ingredient sequence layer for modeling and processing to obtain ingredient sequence information; A first pre-trained encoder module is used to encode multiple food text prompt word information to generate multiple embedding vectors; A second pre-trained encoder module is used to encode the plurality of food attribute binding information to generate a plurality of conditional feature vectors; unet module, used for processing the preset noise vector, the embedding vector, and the conditional feature vector to generate potential image features; An online adaptive decoding module is used to decode the potential image features to obtain a refined picture of the dish; The ingredient sequence module, the first pre-trained encoder module, the second pre-trained encoder module, the unet module and the online adaptive decoding module jointly build the AdaDMOE model.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0015] Compared with the prior art, the method for generating a refined dish image proposed in the present invention has the beneficial effect that, by introducing an online adaptive decoder, the present invention combines a multi-scale discriminator and text feature matching to improve the detail fidelity of the generated dish image, making it clearer, more realistic, and in line with the text prompt information. An adaptive weight matrix is used to optimize the influence of the order of ingredients on the final generated image, so that the combination of ingredients in different orders is more visually in line with actual cooking habits, and the realism of the dish image is enhanced. Through the entity relationship cross-attention mechanism, the ingredients and cooking methods are bound to ensure that the visual form of the ingredients conforms to the corresponding cooking methods, avoiding the problem of mismatching such as scrambled eggs mistakenly generating boiled eggs. At the same time, the present technology adopts a concise text prompt method based on template sentences, which greatly reduces the consumption of computing resources, improves the model processing efficiency, and ensures the accuracy and consistency of the generated dish image compared to the traditional lengthy text description. In addition, the scheme ignores non-critical factors such as dish containers and background elements, improves the stability and generalization ability of the generated image, and enables the model to generate expected dish images under different text prompts, which improves the quality and application value of the dish image generation model as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of a method for generating a refined dish image provided by the first embodiment of the present application; Figure 2 It is an AdaDMOE model diagram provided by the first embodiment of the present application; Figure 3 This is an internal structure diagram of a dish refinement image generation system provided in the second embodiment of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0019] In the following, in combination with the accompanying drawings, a method and system for generating refined images of dishes provided in an embodiment of the present application are described in detail through specific embodiments and their application scenarios.
[0020] Example 1 See also Figure 1 , which is a flow chart of a method for generating refined dish images proposed in the first embodiment of the present application. The proposed method includes steps S1 to S4.
[0021] Step S1: According to the order of ingredient attributes, ingredient order information is obtained through an ingredient order layer. The ingredient order layer is used to model the order of ingredient attributes to generate ingredient and ingredient order prompt information.
[0022] Specifically, the ingredient sequence layer includes an ingredient embedding layer, a first fully connected layer and a second fully connected layer. The ingredient embedding layer is used to embed the ingredient sequence into the end of the ingredient to generate ingredients and ingredient sequence. The first fully connected layer is used to map the ingredients and ingredient sequence to generate an embedding vector. The second fully connected layer is used to map the text sequence embedding obtained by multiplying the embedding vector and a preset adaptive weight matrix to generate ingredient and ingredient sequence prompt information.
[0023] The present invention uses sine and cosine functions to generate the code for each sequence, and this method has a certain regularity between different sequences. and embedding dimension , the mathematical expression of sequential coding is: ; ; in, Indicates the order of words, Represents the index of the embedding dimension, the range of the index is arrive , is the channel dimension of the pre-trained CLIP encoder. represents the cosine order embedding, represents a sinusoidal sequential embedding.
[0024] Furthermore, the process of generating an embedding vector includes: ; in, represents the embedding vector, represents the first fully connected layer, represents the order of ingredients and the word vector of ingredients, Indicates sequential encoding.
[0025] Further, the text sequence embedding obtained by multiplying the embedding vector and the preset adaptive weight matrix is mapped to generate the ingredients and the ingredient sequence prompt information, including: ; in, Indicates the ingredients and the order of ingredients. represents the second fully connected layer, Represents the adaptive weight matrix The initial coefficient of .
[0026] Since different orders of ingredients have a greater or lesser impact on the appearance of the final dish image, this paper studies the effect of different orders of ingredients on fine-grainedness. The preset adaptive weight matrix is continuously updated during the entire training process of generating dish images. By multiplying the ingredients order and the ingredients word vector, and then passing through the second fully connected layer Get tips on ingredients and their order .
[0027] The input of the ingredient sequence layer of the present invention is the ingredient composition and ingredient sequence embedding, and the ingredient sequence information embedding is added at the end of each ingredient. The ingredient sequence layer of the present invention does not require training. After the normalized text passes through the NLTK natural language tool, the ingredients and cooking methods can be obtained. The ingredients and cooking methods are matched according to the proximity principle. For example, fried celery with beef will become fried celery and fried beef. The relationship between ingredients is bound, and the sequence information is extracted.
[0028] In addition, in the process of generating dish images, it is necessary not only to include the meaning of the ingredients, but also to indicate the position and order of the ingredients in the sentence. In order to emphasize the different order of added ingredients and control the fine-grained consistency of the final generated dish image, the present invention introduces an online adaptive encoder and an adaptive weight matrix. The online adaptive encoder receives the position order embedding of the ingredients and the ingredients. The position order embedding of the ingredients and the ingredients is added to the end of the ingredient text embedding to improve the authenticity of the final generated dish image due to the different order of adding ingredients. Finally, the adaptive weight matrix is to adapt to different prompt order information, which enhances the generalization ability.
[0029] Step S2: According to the multiple food text prompt word information and the multiple food attribute binding information, encoding processing is performed through a pre-trained encoder to obtain multiple embedding vectors and multiple conditional feature vectors respectively.
[0030] Specifically, the pre-trained encoder of the present invention specifically refers to a pre-trained CLIP encoder, which is a multimodal pre-trained model that embeds images and texts into the same semantic space through contrastive learning. The core idea is to achieve cross-modal semantic alignment through contrastive training of large-scale image-text pairs. The present invention divides the CLIP encoder into a first pre-trained CLIP encoder and a second pre-trained CLIP encoder, and the first pre-trained CLIP encoder is used to encode multiple food text prompt word information to obtain multiple embedding vectors, and the generated embedding vectors capture the semantic features of the text description and the order features of the food. The second pre-trained CLIP encoder is used to encode multiple food attribute binding information to obtain multiple conditional feature vectors. Multiple embedding vectors and multiple conditional feature vectors are used to guide the generation of potential representation vectors.
[0031] Step S3: Perform inverse diffusion and denoising processing on the preset noise vector, embedding vector, and conditional feature vector through a preset unet model to generate latent image features.
[0032] The unet model used in the present invention specifically includes multiple entity cross attention layers and a denoising layer. The entity cross attention layer is used to Perform inverse diffusion to obtain the intermediate latent representation , the denoising layer is used to transform the intermediate potential representation Perform denoising to obtain potential image features .
[0033] Specifically, the specific process of obtaining potential image features includes: binding and mapping prompt text The ingredient names and ingredient attributes in the image are obtained to obtain the weighted sum of each ingredient and ingredient attribute; based on the weighted sum of each ingredient and ingredient attribute, the weighted attention between the generated image and the prompt text is obtained; the prompt text is constructed based on a preset template; the weighted attention between the key vector of the prompt text and the word vector of the prompt text is obtained; the cross-attention score is obtained based on the weighted attention between the generated image and the prompt text, and the weighted attention between the key vector of the prompt text and the word vector of the prompt text; the cross-attention score is mapped to the online adaptive decoder for denoising to obtain the potential image features.
[0034] Specifically, the mathematical expression of the weighted sum of each ingredient and ingredient attribute is: ; in, represents the weighted sum of each ingredient and ingredient attribute, Indicates the number of ingredients.
[0035] Furthermore, the process of obtaining the weighted attention between the generated image and the prompt text includes: Map to To obtain the weighted attention between the generated image and the prompt text, the mathematical expression is: ; in, Represents the weighted attention between the generated image and the prompt text, represents the generated image features, Indicates the key, represents the matrix transpose, Represents the dimension, express Activation function.
[0036] Furthermore, the process of obtaining the weighted attention between the key vector of the prompt text and the word vector of the prompt text includes: aligning the prompt text with the ingredients and their attributes, ignoring the noise of some background factors. In order to enhance the generalization of the text prompt, the influence of the container for the dish and the background elements on the generated image is weakened. Then align and cross-fuse the original input text with the generated image, and the mathematical expression is: ; in, Represents the weighted attention between the key vector of the prompt text and the word vector of the prompt text, A key vector representing the prompt text, Word vector representing the prompt text. Further, get the cross attention score The process includes: ; in, represents the cross attention score, represents the weight, Indicates the number of ingredients. represents the weighted sum of each ingredient and ingredient attribute, represents the generated image features, A key vector representing the prompt text, The word vector representing the prompt text, Represents the dimension, represents the matrix transpose, Indicates a key.
[0037] Furthermore, when constructing the prompt text based on the preset template, the mathematical expression of the preset template is: ; in, A mathematical expression representing the template, Represents other types of text, Multiple ways to cook representative dishes. Represents multiple ingredients of a dish.
[0038] Step S4: Based on the potential image features, decoding is performed through an online adaptive decoder to generate a refined picture of the dish.
[0039] The online adaptive decoder of the present invention is used to receive potential image features to generate multiple dish images, and train and fine-tune each generated dish image to obtain an updated generated dish image. The online adaptive decoder used in the present invention includes three parts: one is a pre-trained online adaptive decoder, and the weight parameters of the pre-trained online adaptive decoder are updated during each training and fine-tuning process; one is a trainable multi-scale discriminator, and the parameters of the discriminator are not updated; and the other is a picture-text descriptor, which is used to perform picture extraction processing on the dish image generated by the online adaptive decoder to obtain picture extracted text.
[0040] Specifically, the process of generating a refined dish image includes: inputting the latent image features into the online adaptive decoder to obtain the generated dish image; inputting the generated dish image into the image-text descriptor to extract text and obtain text information; obtaining the spherical distance loss function based on the text information and the preset Prompt prompt word features; updating the weights of the online adaptive decoder based on the spherical distance loss function and the preset adversarial loss function to obtain an updated online adaptive decoder; training and fine-tuning multiple generated dish images based on the updated online adaptive decoder to generate refined dish images.
[0041] Spherical distance loss function The mathematical expression is: ; in, Used to guide the generated image to be more consistent with the text prompt, Indicates the text extracted from the image, which comes from the generated dish image. The prompt text is Prompt.
[0042] Preset adversarial loss function The mathematical expression is: ; in, Indicates the generated image. represents the real image, represents the expected calculation, Represents the result of encoding the generated image. Represents the result of encoding the real image.
[0043] Weighted spherical distance loss function of the present invention And the preset adversarial loss function As the total loss function (Satisfying the relation ), and use the total loss function The weight of the online adaptive decoder is updated to obtain an updated online adaptive decoder, and then based on the updated online adaptive decoder, a plurality of generated dish images are trained and fine-tuned to obtain updated generated dish images.
[0044] In summary, the beneficial effects of step S4 of the present invention can be divided into the following points: First, the present invention uses the total loss function By updating the weights of the online adaptive decoder, each generated dish image can be fine-tuned, improving the detail fidelity of the generated dish images.
[0045] Secondly, the present invention also uses two conditional mechanisms to bind the attribute relationship between ingredients. The first conditional mechanism is: describing the basic information of the dish and its ingredients and cooking methods. The second conditional mechanism is: the dish ingredients and cooking methods are bound together in a weighted manner, and the attributes of each ingredient are bound to the ingredient ingredients. For example, "fried" and "eggs" in scrambled eggs are an ingredient and an ingredient attribute, and eggs will not appear alone. By introducing these two conditional mechanisms, the model can accurately match the dish image with the corresponding text description, especially in terms of image-text consistency.
[0046] Then, in the whole process of generating dish images, the text prompt information is completed using the pre-trained CLIP model. However, in the existing model, it may be that there is no clear attribute of the ingredients in the data set or the text encoder does not understand the long text accurately. The output image of scrambled eggs with tomatoes will include the image of boiled eggs. It is necessary to study and solve the problem of different visual forms of ingredients under different cooking methods, and strengthen the bundling of ingredients and ingredient attributes after different cooking methods. The present invention obtains the cross-attention score in the cross-attention layer. The present invention uses the entity relationship cross-attention layer as the key component of image generation to process the relationship between the dish name, dish ingredients and dish attributes. This layer fuses the attribute features between ingredients and ingredients through the entity cross-attention layer, thereby generating a dish image that conforms to the text description.
[0047] Finally, in traditional dish image generation models, the input prompt text is often a lengthy and complex text description, which requires huge computing resources. To generate dish images, the model only needs the ingredients, cooking methods, and input order of ingredients. The main architecture for processing text prompts is Transformer, and the computational complexity of Transformer's self-attention mechanism is , to generate images based on long text descriptions, the consumption of computing resources during processing increases exponentially. In order to solve this problem, the input of the present invention is a concise prompt text based on a template sentence. The method generates a structured and consistent prompt text by extracting the key elements of ingredients and their methods. The present invention constructs prompt text based on a preset template, which is more concise than traditional long text descriptions. This template design is concise and clear, and it is easier to obtain text features in the actual processing process, which can optimize text prompts and improve computing efficiency.
[0048] The present invention builds an AdaDMOE model based on an ingredient sequence layer, a first pre-trained encoder layer, a second pre-trained encoder layer, a unet layer, and an online adaptive decoding layer. The AdaDMOE model of the present invention is a generative model based on a latent space. By modeling the prompt text ingredient sequence and binding the ingredient attributes, the latent space feature image is generated through a pre-trained model, and finally the latent space feature is fine-tuned through an online adaptive decoder to generate a refined dish image. The AdaDMOE model proposed by the present invention is as follows Figure 2 shown.
[0049] Example 2 See also Figure 3 , which is a schematic diagram of the structure of a system for generating refined images of dishes proposed in the second embodiment of the present application, and the system includes: The ingredient sequence module 100 is used to input the ingredient attribute sequence into the ingredient sequence layer for modeling and processing to obtain ingredient sequence information; A first pre-trained encoder module 200 is used to encode a plurality of food text prompt word information to generate a plurality of embedding vectors; A second pre-trained encoder module 300 is used to encode the plurality of food attribute binding information to generate a plurality of conditional feature vectors; unet module 400, used to process the preset noise vector, embedding vector, and conditional feature vector to generate potential image features; An online adaptive decoding module 500 is used to decode the potential image features to obtain a refined picture of the dish; The ingredient sequence module 100, the first pre-trained encoder module 200, the second pre-trained encoder module 300, the unet module 400 and the online adaptive decoding module 500 jointly build the AdaDMOE model.
[0050] A dish refined image generation system in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0051] A dish refined image generation system in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0052] The system for generating refined pictures of dishes provided in the embodiment of the present application can achieve Figure 1 to Figure 2 In order to avoid repetition, the various processes of implementing a method for generating a refined dish image in the method embodiment will not be described here.
[0053] The beneficial effects of a dish refined image generation system provided by the present invention are: first, by introducing an online adaptive decoding module, combined with a multi-scale discriminator and a text feature matching mechanism, the detail fidelity of the generated dish image is improved, so that the generated image is more in line with the visual characteristics of the real dish; second, the preset noise vector, embedding vector, and conditional feature vector are processed by the unet module to generate latent image features, which can improve the stability and generalization ability of the generated image, so that the model can generate expected dish images under different text prompts, and the quality and application value of the dish image generation model are improved as a whole.
[0054] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned embodiment of the method for generating a refined picture of a dish is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0055] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned embodiment of a method for generating a refined dish image is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0056] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0057] It should be noted that, in the present invention, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises one..." does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be noted that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0058] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0059] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for generating a refined dish image, characterized in that: include: According to the order of ingredient attributes, ingredient order information is obtained through an ingredient order layer, wherein the ingredient order layer is used to model the order of ingredient attributes to generate ingredient and ingredient order prompt information; According to the information of multiple food text prompt words and the information of multiple food attribute bindings, encoding is performed through a pre-trained encoder to obtain multiple embedding vectors and multiple conditional feature vectors respectively; Performing inverse diffusion and denoising processing on a preset noise vector, the embedding vector, and the conditional feature vector through a preset unet model to generate potential image features; According to the potential image features, decoding processing is performed through an online adaptive decoder to generate a refined picture of the dish.
2. A method for generating a refined dish image according to claim 1, characterized in that: The ingredient sequence layer includes an ingredient embedding layer, a first fully connected layer, and a second fully connected layer. The ingredient embedding layer is used to embed the ingredient sequence into the end of the ingredient to generate ingredients and ingredient sequence. The first fully connected layer is used to map the ingredients and ingredient sequence to generate an embedding vector. The second fully connected layer is used to map the text sequence embedding obtained by multiplying the embedding vector and a preset adaptive weight matrix to generate ingredient and ingredient sequence prompt information.
3. The method for generating a refined dish image according to claim 1, characterized in that: The UNET model includes an entity cross-attention layer and a denoising layer. The entity cross-attention layer is used to perform inverse diffusion processing on the noise vector, the embedding vector, and the conditional feature vector to generate an intermediate latent representation, and the denoising layer is used to perform denoising processing on the intermediate latent representation to generate latent image features.
4. The method for generating a refined dish image according to claim 1, characterized in that: The process of decoding through an online adaptive decoder to generate a refined picture of the dish includes: Inputting the potential image features into the online adaptive decoder to obtain a generated dish image; Input the generated dish picture into the image-text descriptor to extract text and obtain text information; Obtaining a spherical distance loss function based on the text information and preset Prompt word features; Based on the spherical distance loss function and the preset adversarial loss function, the weight of the online adaptive decoder is updated to obtain an updated online adaptive decoder; Based on the updated online adaptive decoder, multiple generated dish images are trained and fine-tuned to generate refined dish images.
5. The method for generating a refined dish image according to claim 3, characterized in that: The process of obtaining the potential image features includes: Bind and map the ingredient name and ingredient attributes in the prompt text to obtain the weighted sum of each ingredient and ingredient attribute; Based on the weighted sum of each ingredient and ingredient attribute, the weighted attention between the generated image and the prompt text is obtained; Build prompt text based on preset templates; Get the weighted attention between the key vector of the prompt text and the word vector of the prompt text; Obtaining a cross attention score based on the weighted attention between the generated image and the prompt text, and the weighted attention between the key vector of the prompt text and the word vector of the prompt text; The cross-attention scores are mapped to an online adaptive decoder for denoising to obtain latent image features.
6. A method for generating a refined dish image according to claim 5, characterized in that: The mathematical expression of the cross attention score is: ; in, represents the cross attention score, represents the weight, Indicates the number of ingredients. represents the weighted sum of each ingredient and ingredient attribute, express Activation function, represents the generated image features, A key vector representing the prompt text, The word vector representing the prompt text, Represents the dimension, represents the matrix transpose, Indicates a key.
7. The method for generating a refined dish image according to claim 5, characterized in that: The mathematical expression of the template is: ; in, A mathematical expression representing the template, Represents other types of text, Multiple ways to cook representative dishes. Represents multiple ingredients of a dish.
8. A system for generating refined pictures of dishes, characterized in that: include: The ingredient sequence module is used to input the ingredient attribute sequence into the ingredient sequence layer for modeling and processing to obtain ingredient sequence information; A first pre-trained encoder module is used to encode multiple food text prompt word information to generate multiple embedding vectors; A second pre-trained encoder module is used to encode the plurality of food attribute binding information to generate a plurality of conditional feature vectors; unet module, used for processing the preset noise vector, the embedding vector, and the conditional feature vector to generate potential image features; An online adaptive decoding module is used to decode the potential image features to obtain a refined picture of the dish; The ingredient sequence module, the first pre-trained encoder module, the second pre-trained encoder module, the unet module and the online adaptive decoding module jointly build the AdaDMOE model.
9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of a method for generating a refined dish image as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a method for generating a refined dish image as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
A method of generating food image from recipe
CN112017255A
Automatic ingredient production system and method for prefabricated dishes
CN117496255A
Image generation method and device, equipment and storage medium
CN117710530A
Recipe-to-food controllable generation method and device based on pre-trained image-text matching model
CN118052895A
Diffusion model food generation method and system based on collaborative relationship perception
CN119169135A