Display method and device of dish information, storage medium and electronic device
By using a pre-trained generative model to automatically generate dish images that match the input ingredients, the problem of time-consuming dish information acquisition in existing technologies is solved, and the matching accuracy and efficiency are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
- Filing Date
- 2023-03-31
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for displaying dish information often result in long processing times due to the low matching degree between the dishes and the input ingredients.
Using a pre-trained generative model, the ingredient information of the reference ingredients is used as the control condition for image generation. This automatically generates dish images that match the input ingredients and displays them on the target terminal screen.
It improved the matching accuracy and efficiency of obtaining menu information, reduced the acquisition time, increased the richness of information display, and enhanced users' interest in creating new dishes.
Smart Images

Figure CN116484083B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the Internet field, and more specifically, to a method and apparatus for displaying menu information, a storage medium, and an electronic device. Background Technology
[0002] Currently, users can use relevant apps to search for dishes when cooking, and then refer to the searched dishes and related recipes to assist in cooking. When searching for recipes, users usually enter one or more ingredients they already have. The app's backend then recommends dishes and related recipes to the user based on recipe ratings, the relevance of the ingredients in the recipe to the entered ingredients, and the user's browsing history.
[0003] However, the way the above-mentioned dish information is displayed recommends dishes that already exist and contain the ingredient. In addition to the currently entered ingredient, it may also include other types of related ingredients. If the user does not have other related ingredients, cooking cannot be completed. Therefore, users need to spend a long time browsing the displayed dish list to find the dish they need.
[0004] It is evident that the methods for displaying dish information in related technologies suffer from a technical problem: the low matching degree between the displayed dishes and the input ingredients leads to a long time consumption in obtaining dish information. Summary of the Invention
[0005] This application provides a method and apparatus for displaying dish information, a storage medium, and an electronic device to at least solve the technical problem in the related art that the time required to obtain dish information is long due to the low matching degree between the displayed dishes and the input ingredients.
[0006] According to one aspect of the embodiments of this application, a method for displaying dish information is provided, comprising: obtaining a dish generation request from a target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients; inputting the ingredient information of the set of reference ingredients into a target generation model to obtain a target dish image, wherein the target generation model is pre-trained and used to automatically generate a dish image matching the ingredients indicated by the input ingredient information as a control condition for image generation; and responding to the dish generation request, controlling the display of the target dish image on the screen of the target terminal.
[0007] In an exemplary embodiment, after obtaining the dish generation request sent by the target terminal, the method further includes: extracting the ingredient image of the first reference ingredient from the set of reference ingredients and the ingredient identifier of the second reference ingredient from the dish generation request, wherein the ingredient information of the first reference ingredient is the ingredient image of the first reference ingredient, and the ingredient information of the second reference ingredient is the ingredient identifier of the second reference ingredient.
[0008] In an exemplary embodiment, the step of inputting the ingredient information of the set of reference ingredients into the target generation model to obtain a target dish image includes: inputting the ingredient information of the set of reference ingredients and the kitchenware information of the target kitchenware into the target generation model to obtain the target dish image. The target generation model is further configured to use the input ingredient information and the input kitchenware information as control conditions for image generation, and automatically generate a dish image that matches the ingredients indicated by the input ingredient information and the kitchenware indicated by the input kitchenware information.
[0009] In an exemplary embodiment, the step of inputting the ingredient information of the set of reference ingredients and the kitchenware information of the target kitchenware into the target generation model to obtain the target dish image includes: inputting the ingredient information of the set of reference ingredients, the kitchenware information of the target kitchenware, and the indication information of the target cooking method into the target generation model to obtain the target dish image. The target generation model is further configured to use the input ingredient information, the input kitchenware information, and the input cooking method indication information as control conditions for image generation, and automatically generate dish images that match the ingredients indicated by the input ingredient information, the kitchenware indicated by the input kitchenware information, and the cooking method indicated by the input indication information.
[0010] In an exemplary embodiment, after obtaining the dish generation request sent by the target terminal, the method further includes: if a target text instruction is extracted from the dish generation request and a kitchenware identifier is parsed from the target text instruction, the parsed kitchenware identifier is determined as the kitchenware information of the target kitchenware; if a target text instruction is extracted from the dish generation request and a kitchenware identifier is not parsed from the target text instruction, or if no text instruction is extracted from the dish generation request, the preset kitchenware information matching the ingredient category of the set of reference ingredients is determined as the kitchenware information of the target kitchenware.
[0011] In an exemplary embodiment, the step of inputting the ingredient information of the set of reference ingredients into a target generation model to obtain a target dish image includes: inputting the ingredient information of each reference ingredient in the set of reference ingredients as a control condition into a target diffusion model, so as to perform the following processing operations through the target diffusion model, wherein the target generation model is the target diffusion model: converting each input information as a control condition of the target diffusion model into an input feature corresponding to each input information, wherein the control condition of the target diffusion model includes the ingredient information of the set of reference ingredients; performing feature fusion with the input features corresponding to each input information to obtain a target fusion feature, wherein the target fusion feature is used to characterize the control condition of the target diffusion model; inputting the image features of the initial noisy image and the target fusion feature as the weight matrix of the attention layer of the denoising network in the target diffusion model into the attention layer to obtain the target dish features output by the denoising network, wherein the target dish features are image features obtained after denoising the image features of the initial noisy image; and decoding the target dish features to obtain the target dish image.
[0012] In one exemplary embodiment, the method further includes: acquiring training dish images and ingredient information of a set of training ingredients corresponding to the training dish images; using the ingredient information of each training ingredient in the set of training ingredients as a control condition for an initial generation model; and using the training dish images to train the initial generation model to obtain the target generation model.
[0013] According to another aspect of the embodiments of this application, a display device for dish information is also provided, comprising: a first acquisition unit, configured to acquire a dish generation request from a target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients; an input unit, configured to input the ingredient information of the set of reference ingredients into a target generation model to obtain a target dish image, wherein the target generation model is pre-trained and configured to use the input ingredient information as a control condition for image generation to automatically generate a dish image matching the ingredients indicated by the input ingredient information; and a control unit, configured to respond to the dish generation request and control the display of the target dish image on the screen of the target terminal.
[0014] In one exemplary embodiment, the apparatus further includes: an extraction unit, configured to, after obtaining the dish generation request sent by the target terminal, extract from the dish generation request an ingredient image of a first reference ingredient in the set of reference ingredients and an ingredient identifier of a second reference ingredient in the set of reference ingredients, wherein the ingredient information of the first reference ingredient is the ingredient image of the first reference ingredient, and the ingredient information of the second reference ingredient is the ingredient identifier of the second reference ingredient.
[0015] In an exemplary embodiment, the input unit includes: a first input module, configured to input the ingredient information of the set of reference ingredients and the kitchenware information of the target kitchenware into the target generation model to obtain the target dish image, wherein the target generation model is further configured to use the input ingredient information and the input kitchenware information as control conditions for image generation, and automatically generate a dish image that matches the ingredients indicated by the input ingredient information and the kitchenware indicated by the input kitchenware information.
[0016] In an exemplary embodiment, the first input module includes an input submodule, configured to input the ingredient information of the set of reference ingredients, the kitchenware information of the target kitchenware, and the indication information of the target cooking method into the target generation model to obtain the target dish image. The target generation model is further configured to use the input ingredient information, the input kitchenware information, and the input cooking method indication information as control conditions for image generation, and automatically generate dish images that match the ingredients indicated by the input ingredient information, the kitchenware indicated by the input kitchenware information, and the cooking method indicated by the input indication information.
[0017] In one exemplary embodiment, the apparatus further includes: a first determining unit, configured to, after acquiring the dish generation request sent by the target terminal, determine the parsed kitchen utensil identifier as the kitchen utensil information of the target kitchen utensil if a target text instruction is extracted from the dish generation request and a kitchen utensil identifier is parsed from the target text instruction; and a second determining unit, configured to, if a target text instruction is extracted from the dish generation request and a kitchen utensil identifier is not parsed from the target text instruction, or if no text instruction is extracted from the dish generation request, determine the preset kitchen utensil information matching the ingredient category of the set of reference ingredients as the kitchen utensil information of the target kitchen utensil.
[0018] In an exemplary embodiment, the input unit includes: a second input module, configured to input the ingredient information of each reference ingredient in the set of reference ingredients as a control condition into a target diffusion model, so as to perform the following processing operations through the target diffusion model, wherein the target generation model is the target diffusion model: converting each input information as a control condition of the target diffusion model into an input feature corresponding to each input information, wherein the control condition of the target diffusion model includes the ingredient information of the set of reference ingredients; performing feature fusion with the input features corresponding to each input information to obtain a target fusion feature, wherein the target fusion feature is used to characterize the control condition of the target diffusion model; inputting the image features of the initial noisy image and the target fusion feature as a weight matrix of the attention layer of the denoising network in the target diffusion model into the attention layer to obtain the target dish feature output by the denoising network, wherein the target dish feature is an image feature obtained after denoising the image features of the initial noisy image; and decoding the target dish feature to obtain the target dish image.
[0019] In one exemplary embodiment, the apparatus further includes: a second acquisition unit, configured to acquire training dish images and ingredient information of a set of training ingredients corresponding to the training dish images; and a training unit, configured to use the ingredient information of each training ingredient in the set of training ingredients as a control condition for an initial generation model, and use the training dish images to train the initial generation model to obtain the target generation model.
[0020] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-mentioned method for displaying dish information when it is run.
[0021] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for displaying menu information through the computer program.
[0022] In this embodiment, a method is adopted to automatically generate dish images by using the ingredient information of reference ingredients as the image generation control condition of the generation model. The method obtains a dish generation request from the target terminal, whereby the dish generation request requests the use of a set of reference ingredients to generate a matching dish. The ingredient information of a set of reference ingredients is input into the target generation model to obtain a target dish image. The target generation model is pre-trained and used to automatically generate a dish image matching the ingredients indicated by the input ingredient information, using the input ingredient information as a control condition for image generation. In response to the dish generation request, the target dish image is displayed on the screen of the target terminal. After the terminal receives a dish generation request, the ingredient information of the reference ingredients is used as the control condition (equivalent to the constraint condition for dish generation) for the image generation model. Thus, the image generation model can automatically generate a dish image that matches the ingredients indicated by the input ingredient information. Since the dish image is generated directly rather than matched from existing dish images, the matching degree between the displayed dish image and the input ingredient information can be improved, thereby reducing the technical effect of obtaining dish information. This solves the technical problem in related technologies where the dish information display method has a long time consumption for obtaining dish information due to the low matching degree between the displayed dish and the input ingredients. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the hardware environment for an optional method of displaying menu information according to an embodiment of this application;
[0026] Figure 2 This is a flowchart illustrating an optional method for displaying menu information according to an embodiment of this application;
[0027] Figure 3 This is a flowchart illustrating another optional method for displaying menu information according to an embodiment of this application;
[0028] Figure 4 This is a flowchart illustrating another optional method for displaying menu information according to an embodiment of this application;
[0029] Figure 5 This is a structural block diagram of an optional dish information display device according to an embodiment of this application;
[0030] Figure 6 This is a structural block diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] According to one aspect of the embodiments of this application, a method for displaying menu information is provided. This method for displaying menu information is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned method for displaying menu information can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to terminal device or clients installed on terminal device. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0034] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0035] The method for displaying menu information in this embodiment can be executed by server 104, terminal device 102, or both server 104 and terminal device 102. Alternatively, the method for displaying menu information in this embodiment can be executed by a client installed on the terminal device 102.
[0036] Taking the method of displaying dish information in this embodiment, executed by server 104, as an example, Figure 2 This is a flowchart illustrating an optional method for displaying menu information according to an embodiment of this application, as shown below. Figure 2 As shown, the process of this method may include the following steps:
[0037] Step S202: Obtain the dish generation request from the target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients.
[0038] The method for displaying dish information in this embodiment can be applied to scenarios where matching dish images are generated based on the ingredient information of a set of reference ingredients. Here, the ingredient information of the reference ingredients can be in multiple modalities such as images and text (identifiers). The types of reference ingredients can be any positive integer greater than or equal to 1. When using two or more types of reference ingredients, multiple information modalities can be arbitrarily combined, that is, multimodal input. For example, they can all be in image form, all in text form, or both image form and text form. The generated dish image can be a dish image containing reference ingredients, for example, it can be a dish image containing only reference ingredients.
[0039] In real life, users may encounter scenarios where, when they want to cook, they can use relevant apps to search for dishes and refer to the search results and related recipes to assist in cooking. When searching for recipes, users typically enter the names or images of one or more ingredients they currently have. The app's backend then recommends dishes and related recipes to the user based on recipe ratings, the relevance of the ingredients in the recipe to the entered ingredients, and the user's browsing history.
[0040] However, the way the above-mentioned dish information is displayed recommends dishes that already exist and contain the ingredient. In addition to the currently entered ingredient, it may also include other types of related ingredients. If the user does not have other related ingredients, cooking cannot be completed. Therefore, users need to spend a long time browsing the displayed dish list to find the dish they need.
[0041] To address at least part of the aforementioned problems, in this embodiment, a pre-trained generative model for generating dishes is used. This model can use the ingredient information of the input reference ingredients as a control condition for image generation, directly generating dish images that match the input reference ingredients. The generated dish images can display existing dishes or dishes that have not yet been tried (i.e., dishes not tried in the existing recipe system). Since matching dish images are generated directly based on the ingredient information of the input reference ingredients and displayed directly on the target terminal's screen, the time for obtaining dish information can be reduced, improving the efficiency of information acquisition. Furthermore, generating images of dishes that have not yet been tried can enhance the richness of information display and increase user interest in creating new dishes. Here, the target terminal can be a smart home device within a smart kitchen ecosystem (e.g., a smart refrigerator) or a user-related smart device (e.g., a smartphone, a smart tablet, etc.). Additionally, the target terminal has a screen capable of displaying images.
[0042] In this embodiment, when a user needs to view a picture of a dish to be prepared, the target terminal can first be activated through voice input, manual touch, or other activation operations. After activating the target terminal, it can send a dish generation request to the target server via user trigger, requesting the generation of a matching dish image using a set of reference ingredients. The target server can then obtain and parse the dish generation request instruction from the target terminal to determine the ingredient information of the set of reference ingredients.
[0043] The ingredient information for a set of reference ingredients can be in a multimodal form, such as images or labels. Here, the ingredient information for each reference ingredient in a set of reference ingredients can be an image of that reference ingredient or an ingredient identifier (e.g., ingredient name, ingredient number, etc.). The ingredient information for a set of reference ingredients can be carried in the dish generation request, or it can be determined based on the ingredient indication information (e.g., the ingredient identifier mentioned above) carried in the dish generation request. For example, the ingredient information for a set of reference ingredients may be a set of images of the reference ingredients. The dish generation request carries the ingredient name of a certain reference ingredient, and the target server searches for an image that matches the ingredient name of the reference ingredient from the pre-set image database to obtain the image of that reference ingredient.
[0044] Optionally, the ingredient information of a set of reference ingredients can be a set of ingredient images, and these ingredient images are carried in the dish generation request. For the target terminal, the ingredient information of a set of reference ingredients can be entered through the dish image generation interface (which can be the display interface within a specified application running on the target terminal). Users can directly select ingredient images of certain reference ingredients from their local storage in the dish image generation interface, or they can use the image acquisition component on the interface to photograph or scan the reference ingredients to obtain ingredient images, or they can directly enter the ingredient identifier of the reference ingredients, allowing the target terminal to automatically obtain the matching ingredient image (either locally or from the server). For example, in response to the ingredient image of a third reference ingredient entered in the dish image generation interface, the target terminal can obtain the ingredient image of the third reference ingredient. For example, in response to the ingredient identifier of the fourth reference ingredient entered in the dish image generation interface, the preset ingredient image of the fourth reference ingredient can be obtained (either saved locally on the terminal or obtained from the server). If there are multiple preset ingredient images of the fourth reference ingredient, the multiple preset ingredient images of the fourth reference ingredient can be displayed in the dish image generation interface, and in response to the received selection operation performed on the ingredient image among the multiple preset ingredient images, the ingredient image of the fourth reference ingredient can be determined as the selected ingredient image.
[0045] Optionally, the ingredient information of a set of reference ingredients can be a set of ingredient images of reference ingredients, and at least some of the ingredient images of the reference ingredients in the set of reference ingredients are obtained by the target server from the database in the manner described above based on the ingredient identifiers (e.g., ingredient names) of at least some of the reference ingredients carried in the dish generation request.
[0046] Optionally, the ingredient information of a set of reference ingredients may include ingredient images of at least some of the reference ingredients and ingredient identifiers of the other reference ingredients besides at least some of them, or, a set of ingredient identifiers for the reference ingredients. In the above scenario, the ingredient information of a set of reference ingredients may be included in the dish generation request.
[0047] In addition, the dish image can also be generated based on the ingredient information of a set of reference ingredients and at least one of the following: utensil information indicating the cooking utensils used to cook the set of reference ingredients (e.g., wok, casserole, pressure cooker, etc.), and cooking parameters indicating the cooking method of the set of reference ingredients (e.g., pan-frying, stir-frying, steaming, braising, etc.). The utensil information can be an image of the cooking utensils or a utensil identifier. The identifier can be in text form (e.g., iron pot) or in other forms such as numbers or symbols (e.g., 1, * representing an iron pot). Optionally, the utensil information can be carried in the dish generation request or obtained in a similar way to obtaining the ingredient information described above; this embodiment does not limit this.
[0048] Step S204: Input a set of reference ingredient information into the target generation model to obtain the target dish image. The target generation model is a pre-trained model used to automatically generate dish images that match the ingredients indicated by the input ingredient information, taking the input ingredient information as a control condition for image generation.
[0049] After receiving a dish generation request, the target server can obtain the ingredient information of each reference ingredient in a set of reference ingredients using any method of obtaining ingredient information, and input the obtained ingredient information of each reference ingredient into a pre-trained generation model to obtain a target dish image. The aforementioned pre-trained generation model is the target generation model, and the target dish image is a dish image that matches the ingredients indicated by the input ingredient information of the reference ingredients. It may only contain dish images that match the ingredients indicated by the aforementioned set of reference ingredients.
[0050] Here, the target generation model can be a deep learning model with multimodal input, which uses the input ingredient information as a control condition for image generation to automatically generate dish images that match the ingredients indicated by the input ingredient information. It can use existing recipe data as a reference to simulate dish images that match the input reference ingredients. Here, the dish images are automatically generated using the input ingredient information as a control condition for image generation, rather than being matched from existing dish images. Therefore, it can represent dish images that have not appeared in the existing recipe data. Here, the existing recipe data can include dish images of existing dishes, as well as other dish reference information, such as ingredient names, cooking methods, and cooking utensils. The training and use of the target generation model can be on the same device or on different devices, which is not limited in this embodiment.
[0051] It should be noted that the target generation model can include one or more control conditions, which can be used to constrain the information contained in the food images generated by the target generation model. Here, the control conditions can be a conditional mechanism (i.e., conditional control), so that food images that meet the control conditions can be automatically generated.
[0052] For example, a multimodal input target generation model can be trained, taking photos of leftover food from the refrigerator as input and generating photos of the final product. For instance, if the input photos include carrots and eggs, the model can generate photos of the dishes that might be made from them.
[0053] Step S206: In response to the dish generation request, control the display of the target dish image on the target terminal's screen.
[0054] In response to a dish generation request, the target terminal's screen can be controlled to display a target dish image. This display can occur after the target server receives the target dish image and sends a dish generation response to the target terminal. The response may include the target dish image. Upon receiving the response, the target terminal can parse it to obtain the target dish image and display it on its screen. This display can be executed directly after the target terminal receives the response, or it can be triggered by a detected display action. For example, the target terminal's interface may display a "View Dish Image" button to indicate that the image of the dish to be created has been generated and can be viewed. The user can click "View Dish Image" or similar triggering action to control the display of the target dish image on the target terminal's screen. Other display timings are also possible, but this embodiment does not limit this.
[0055] Before displaying the target dish image on the target terminal's screen, the target terminal's screen may display a loading indicator or animation such as "Generating" to indicate that the target dish image is being generated. It may also display existing information such as recipes and dish images containing the aforementioned set of reference ingredients for reference. It may also display a brief introduction to the aforementioned set of reference ingredients, such as the effects and contraindications of the reference ingredients. This embodiment does not limit this. After the target dish image is displayed on the target terminal's screen, the user can decide whether to try cooking based on the image. This not only provides a certain entertainment function but also has certain practical value. After seeing the generated image, the user may actually try it and may even develop a new dish. The target dish image can also be further processed by adjusting colors, decorating, etc. This embodiment does not limit this.
[0056] It's important to note that the server and target terminal can interact via WebSocket. WebSocket is a protocol that enables full-duplex communication over a single TCP (Transmission Control Protocol) connection, simplifying data exchange between the client and server and allowing the server to proactively push data to the client (e.g., the target application). In the WebSocket API (Application Programming Interface), the client and server only need to complete a single handshake to establish a persistent connection and perform bidirectional data transmission.
[0057] Through steps S202 to S206 above, a dish generation request from the target terminal is obtained. This request requests the generation of a matching dish using a set of reference ingredients. The ingredient information of the reference ingredients is input into the target generation model to obtain a target dish image. The target generation model is pre-trained and uses the input ingredient information as a control condition for image generation, automatically generating a dish image that matches the ingredients indicated by the input ingredient information. In response to the dish generation request, the target dish image is displayed on the target terminal's screen. This solves the technical problem in related technologies where the display method for dish information suffers from a low matching degree between the displayed dish and the input ingredients, resulting in long processing times for dish information acquisition, thus shortening the time required to acquire dish information.
[0058] In one exemplary embodiment, after obtaining the dish generation request sent by the target terminal, the above method further includes:
[0059] S11, extract the ingredient image of the first reference ingredient from a set of reference ingredients and the ingredient identifier of the second reference ingredient from a set of reference ingredients from the dish generation request. The ingredient information of the first reference ingredient is the ingredient image of the first reference ingredient, and the ingredient information of the second reference ingredient is the ingredient identifier of the second reference ingredient.
[0060] In this embodiment, after receiving a dish generation request, information related to a set of reference ingredients can be extracted from the request. Similar to the previous embodiments, the dish generation request may carry ingredient images of at least some of the reference ingredients, such as an image of the first reference ingredient, which the target server can extract from the request. The request may also carry ingredient identifiers of at least some of the reference ingredients, such as an identifier of the second reference ingredient, which the target server can extract from the request. The types and quantities of the first and second reference ingredients can be flexibly set by the user as needed. This embodiment does not impose any limitations on this.
[0061] In this embodiment, the ingredient information of a set of reference ingredients may include: an ingredient image of a first reference ingredient, and may also include an ingredient identifier of a second reference ingredient, or an ingredient image obtained from the ingredient identifier of the second reference ingredient. Here, what can be extracted from the dish generation request may also be other content representing cooking parameter information, such as the aforementioned cookware identifier or image, or, for example, cooking method indication information.
[0062] For example, if the quantity of the first reference ingredient and the second reference ingredient are both set to 1, and the dish generation request includes the ingredient name of the tomato and the image of the egg, after receiving the dish generation request sent by the target terminal, the ingredient identifier ("tomato") and the ingredient image (egg image) can be extracted from the dish generation request. Among them, the egg belongs to the first reference ingredient and the tomato belongs to the second reference ingredient.
[0063] This embodiment extracts reference ingredient images and / or ingredient identifiers from the dish generation request, which increases the flexibility of ingredient information configuration and improves the user experience.
[0064] In one exemplary embodiment, the ingredient information of a set of reference ingredients is input into the target generation model to obtain a target dish image, including:
[0065] S21, input the ingredient information of a set of reference ingredients and the kitchenware information of the target kitchenware into the target generation model to obtain the target dish image. The target generation model is also used to use the input ingredient information and the input kitchenware information as control conditions for image generation, and automatically generate dish images that match the ingredients indicated by the input ingredient information and the kitchenware indicated by the input kitchenware information.
[0066] In this embodiment, dishes cooked with the same ingredients using different cooking utensils can be different. For example, stir-fried beef with tomatoes and braised beef brisket with tomatoes can be cooked in a wok and a casserole or pressure cooker, respectively. In order to improve the rationality of the generated dish images, the ingredient information of each reference ingredient and the cooking utensils of the target cooking utensils can be input into the target generation model to obtain the target dish image.
[0067] Here, the ingredient information for each reference ingredient can be similar to that in the aforementioned embodiments, and will not be repeated here. The kitchenware information for the target kitchenware can be multimodal, such as images, icons, etc. For example, the kitchenware information for the target kitchenware can be a default image or icon information, or it can be carried in the dish request information. It can be an image directly obtained from the dish request information, or an icon directly obtained from the dish request information, or an image indicated by an icon obtained from the dish request information (from local storage or obtained from the server side). In addition, it can also be a preset cooking tool that matches a set of input reference ingredients. The matching relationship can be generated based on big data, determined based on user habits, or preset through other methods. The target kitchenware can include, but is not limited to, rice cookers, gas stoves, steamers, pressure cookers, etc. In this embodiment, there are no limitations on the type of target kitchenware, the type of kitchenware information for the target kitchenware, or the method of determining the target kitchenware.
[0068] Optionally, the target generation model can be a pre-trained multimodal input model that automatically generates dish images matching the ingredients and utensils indicated by the input ingredient information and the input utensil information, respectively, as control conditions for image generation. In other words, the control conditions of the target generation model include the ingredient information of each reference ingredient in a set of input reference ingredients, as well as the utensil information of the input target utensil.
[0069] For example, a recipe generation request may include information about the kitchen appliances (e.g., a picture of a gas stove), the ingredient name (tomato), and the ingredient (egg). By inputting the ingredient identifier ("tomato"), the ingredient picture (egg picture), and the kitchen appliance information (e.g., gas stove picture) into the target generation model, a recipe picture (tomato and egg cooked on a gas stove) can be obtained that matches the ingredient and appliance indicated by the input ingredient information and the input kitchen appliance information.
[0070] In this embodiment, the input ingredient information and the input kitchenware information are used as control conditions to input into the generation model to obtain the dish images output by the generation model, which can improve the matching degree of dish display and improve the user experience.
[0071] In one exemplary embodiment, ingredient information of a set of reference ingredients and kitchenware information of the target kitchenware are input into the target generation model to obtain a target dish image, including:
[0072] S31, input a set of reference ingredient information, target kitchenware information, and target cooking method instruction information into the target generation model to obtain a target dish image. The target generation model is also used to use the input ingredient information, input kitchenware information, and input cooking method instruction information as control conditions for image generation, and automatically generate dish images that match the ingredients indicated by the input ingredient information, the kitchenware indicated by the input kitchenware information, and the cooking method indicated by the input instruction information.
[0073] In this embodiment, the same kitchen utensil may be used for more than one cooking method. For example, an iron wok can be used for stir-frying, braising, or pan-frying. In order to improve the rationality of the generated dish images and obtain dish images that are closer to the user's expectations, the ingredient information of each reference ingredient, the kitchen utensil information of the target kitchen utensil, and the instruction information of the target cooking method can be input into the target generation model to obtain dish images that match the ingredients indicated by the input ingredient information, the kitchen utensil indicated by the input kitchen utensil information, and the cooking method indicated by the input instruction information.
[0074] Here, the indication information for the target cooking method can be the aforementioned cooking parameters, which can be carried in the dish request information, determined based on the target kitchenware information, or determined based on a set of reference ingredients. The indication information for the target cooking method can include, but is not limited to, various types such as frying, stir-frying, steaming, and braising. The ingredient information for each reference ingredient and the kitchenware information for the target kitchenware can be similar to those in the previous embodiments, and will not be repeated here.
[0075] Optionally, the target generation model can be a pre-trained multimodal input model that automatically generates dish images matching the ingredients and utensils indicated by the input ingredient information, the input kitchen utensil information, and the input cooking method instructions, respectively, as control conditions for image generation. In other words, the control conditions of the target generation model include the ingredient information of each reference ingredient in a set of input reference ingredients, the kitchen utensil information of the input target, and the instruction information of the target cooking method.
[0076] For example, a recipe generation request includes the cooking method instruction "stir-fry," the kitchenware information (image of a gas stove), the ingredient name "tomato," and the ingredient "egg." By inputting the ingredient identifier ("tomato"), the ingredient image (egg image), the kitchenware information (gas stove image), and the cooking method instruction "stir-fry" into the target generation model, a recipe image (stir-fried tomatoes and eggs cooked on a gas stove) can be obtained that matches the ingredient indicated by the input ingredient information, the kitchenware indicated by the input kitchenware information, and the cooking method indicated by the input instruction.
[0077] In this embodiment, the input ingredient information, input kitchen utensil information, and input cooking method instructions are used as control conditions and input into the generation model to obtain the dish images output by the generation model. This can improve the matching degree between the generated dish images and the user's expectations, thereby improving the user experience.
[0078] In one exemplary embodiment, after obtaining the dish generation request sent by the target terminal, the above method further includes:
[0079] S41, if the target text instruction is extracted from the dish generation request and the kitchenware identifier is parsed from the target text instruction, the parsed kitchenware identifier is determined as the kitchenware information of the target kitchenware;
[0080] S42, if the target text instruction is extracted from the dish generation request but the kitchenware identifier is not parsed from the target text instruction, or if the text instruction is not extracted from the dish generation request, the preset kitchenware information that matches the ingredient category of a set of reference ingredients is determined as the kitchenware information of the target kitchenware.
[0081] In this embodiment, the dish generation request sent by the target terminal may carry target text instructions. After obtaining the dish generation request sent by the target terminal, the target text instructions can be extracted from the dish generation request and parsed to determine the type of information carried in the target text instructions, such as kitchenware identification, ingredient identification, cooking method instructions, etc.
[0082] In an optional embodiment, if the target text instruction is extracted from the dish generation request and the kitchenware identifier is parsed from the target text instruction, the parsed kitchenware identifier can be identified as the kitchenware information of the target kitchenware. For example, if the target text instruction in the dish request information is "use a wok to stir-fry a dish", then the kitchenware identifier "wok" can be parsed from the target text instruction, and the kitchenware identifier "wok" can be identified as the kitchenware information of the target kitchenware.
[0083] In another optional embodiment, if the target text instruction is extracted from the dish generation request, but the kitchen utensil identifier is not parsed from the target text instruction (e.g., the target text instruction extracted from the dish generation request is "use tomatoes to make a dish"), or if no text instruction is extracted from the dish generation request, a preset kitchen utensil identifier can be matched with the ingredient category of a set of input reference ingredients, and the matched kitchen utensil identifier can be determined as the kitchen utensil information of the target kitchen utensil. Here, the matching relationship of the target kitchen utensil matched according to the ingredient category can be based on the user's cooking habits, or it can be based on big data defaults, or it can be a matching relationship determined by other methods. This embodiment does not limit this.
[0084] This embodiment improves the flexibility of kitchenware selection and enhances the user experience by parsing kitchenware information from recipe generation requests or matching preset kitchenware information based on input reference ingredients.
[0085] In one exemplary embodiment, the ingredient information of a set of reference ingredients is input into the target generation model to obtain a target dish image, including:
[0086] S51, the ingredient information of each reference ingredient in a set of reference ingredient information is input as a control condition into the target diffusion model, so as to perform the following processing operations through the target diffusion model, wherein the target generation model is the target diffusion model:
[0087] Each input information that serves as the control condition for the target diffusion model is converted into an input feature corresponding to each input information. The control condition for the target diffusion model includes a set of ingredient information of reference ingredients.
[0088] The input features corresponding to each input information are fused to obtain the target fusion features, which are used to characterize the control conditions of the target diffusion model.
[0089] The image features of the initial noisy image and the target fusion features are used as the weight matrix of the attention layer of the denoising network in the target diffusion model and input into the attention layer to obtain the target dish features output by the denoising network. The target dish features are the image features obtained after denoising the image features of the initial noisy image.
[0090] The characteristics of the target dish are decoded to obtain an image of the target dish.
[0091] The target generation model can be a target diffusion model, such as a conditional diffusion model. Here, the diffusion model is a deep learning model for image generation, specifically a generative model based on an encoder-decoder architecture. The training process of the diffusion model is divided into a diffusion phase (noise addition) and a de-diffusion phase (denoising). In the diffusion phase, noise is continuously added to the original data (e.g., the original image of a dish) to transform the data from its original distribution to the desired distribution; for example, Gaussian noise is continuously added to transform the original data distribution into a normal distribution. In the de-diffusion phase, a denoising neural network is used to restore the data from its normal distribution back to its original distribution. The application of the diffusion model can be the aforementioned de-diffusion phase. Optionally, conditional mechanisms can be introduced into the diffusion operation, using cross-attention to achieve multimodal training, thereby realizing the conditional image generation task.
[0092] In this embodiment, the ingredient information of each reference ingredient in a set of reference ingredient information can be used as a control condition input into the target diffusion model to guide the reverse generation (denoising) process. Here, the control condition may also include the kitchenware information of the target kitchenware, the instruction information of the cooking method, etc. The control condition can be multimodal, such as text, images, etc. In this embodiment, the content and form of the control condition are not limited.
[0093] Based on multimodal control conditions, each input information (e.g., ingredient information of each reference ingredient, kitchenware information of the target cooking method, instruction information of the target cooking method, etc.) that serves as the control conditions of the target diffusion model can be converted into input features corresponding to each input information. By fusing the input features corresponding to each input information, the target fusion features can be obtained. Here, the role of the target fusion features is to characterize the control conditions of the target diffusion model. Feature fusion can be implemented based on concatenation or based on a specific pre-trained neural network. In this embodiment, the method of feature fusion is not limited.
[0094] The image features of the initial noisy image and the target fusion features are used as the weight matrix of the attention layer of the denoising network in the target diffusion model. This inputs the target dish features output by the denoising network. Here, the target dish features are the image features obtained after denoising the image features of the initial noisy image. The initial noisy image can be a pure Gaussian noise image with independent features, which can be obtained by continuously adding noise based on existing recipe data during model training. Decoding the obtained target dish features yields the target dish image. Here, the target dish image can be a dish image that matches the control conditions and contains only the input reference ingredients.
[0095] This embodiment improves the efficiency and relevance of food image generation by using a conditional diffusion model, thereby enhancing the user experience.
[0096] In one exemplary embodiment, the above method further includes:
[0097] S61, Obtain the training dish image and the ingredient information of a set of training ingredients corresponding to the training dish image;
[0098] S62, the ingredient information of each training ingredient in a set of training ingredients is used as a control condition for the initial generation model. The initial generation model is trained using training dish images to obtain the target generation model.
[0099] The target generation model can be trained with reference to existing recipe data. In this embodiment, training dish images and a set of training ingredient information corresponding to the training dish images are obtained. The ingredient information of each training ingredient in the set of training ingredients is used as a control condition for the initial generation model. The initial generation model is trained using the training dish images to obtain the target generation model.
[0100] Here, the training dish images can be existing dish images, which can be obtained from the network or stored locally. The ingredient information of a set of training ingredients corresponding to the training dish images can be obtained by recognizing the training dish images, input by the user, or obtained through other methods. The ingredient information of a set of training ingredients can be in the form of images, in the form of tags, or in a multimodal form that combines both images and tags. The initial generation model can be an untrained generation model that can be used to generate dish images. The training and use of the generation model can be on the same device or on different devices. This embodiment does not limit this. Optionally, the control conditions of the initial generation model can also include cooking parameter information such as the information of the target kitchen utensils and the indication information of the cooking method. This information can be contained in the existing recipe data or inferred from the obtained ingredient information.
[0101] Optionally, in this embodiment, kitchenware information can be inferred from the acquired ingredient information. Here, there can be one or more inferred kitchenware identifiers. When there is only one inferred kitchenware identifier, it can be directly used as a control condition for the initial generation model. When there are multiple inferred kitchenware identifiers, the frequently used kitchenware identifiers can be prioritized as input information for the initial generation model based on user habits, or other priorities can be determined.
[0102] For example, the training process for a target generation model may include the following steps:
[0103] Step 1: Collect recipe data. Recipe data includes: final image of the dish, ingredient names, ingredient images, and preparation method.
[0104] Step 2: Work backward from the preparation method to deduce the kitchen utensils that can be used;
[0105] Step 3, according to Figure 3 The example shows the training of a target generation model.
[0106] In this embodiment, based on reference dish images, the initial generation model is trained using corresponding ingredient information and kitchen utensil identifiers to obtain a trained generation model, which can improve the reliability of model training and enhance the rationality of dish image generation.
[0107] The method for displaying dish information in this application embodiment will be explained below with reference to optional examples. This optional example provides a scheme for generating dishes based on user ingredient images and instructions, which can generate matching dish images according to the input ingredient information. Combined with... Figure 4 As shown, the process of displaying dish information in this optional example may include the following steps:
[0108] Step 1: The user inputs the dish generation instruction "Stir-fry a dish using a gas stove" and the ingredient information of tomatoes and eggs into the pre-trained diffusion model. The diffusion model is pre-trained with reference to existing recipes and can use any of the four ingredient input formats (picture of tomatoes and name of eggs, picture of tomatoes and picture of eggs, name of tomatoes and name of eggs, and name of tomatoes and picture of eggs) as input.
[0109] Step 2: Obtain the final image of the dish from the Diffusion model output.
[0110] This optional example demonstrates how to construct a multimodal input method for ingredient information and use a diffusion model to generate dish images. This allows for the generation of images of dishes not previously found in recipes. It can provide possible dish images when trying new ingredient combinations, generate a final dish image for reference when ingredients are limited and existing dishes cannot be made, and can be used for preliminary screening during dish development. It is also applicable to other scenarios.
[0111] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0113] According to another aspect of the embodiments of this application, a display device for displaying dish information for implementing the above-described method for displaying dish information is also provided. Figure 5This is a structural block diagram of an optional dish information display device according to an embodiment of this application, such as... Figure 5 As shown, the device may include:
[0114] The first acquisition unit 502 is used to acquire the dish generation request of the target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients;
[0115] The input unit 504 is connected to the first acquisition unit 502 and is used to input the ingredient information of a set of reference ingredients into the target generation model to obtain the target dish image. The target generation model is pre-trained and is used to use the input ingredient information as a control condition for image generation to automatically generate a dish image that matches the ingredients indicated by the input ingredient information.
[0116] The control unit 506, connected to the input unit 504, is used to control the display of the target dish image on the screen of the target terminal in response to the dish generation request.
[0117] It should be noted that the first acquisition unit 502 in this embodiment can be used to execute the above step S202, the input unit 504 in this embodiment can be used to execute the above step S204, and the control unit 506 in this embodiment can be used to execute the above step S206.
[0118] Through the above modules, a dish generation request from the target terminal is obtained. This request requests the generation of a matching dish using a set of reference ingredients. The ingredient information of the reference ingredients is input into the target generation model to obtain a target dish image. The target generation model is pre-trained and uses the input ingredient information as a control condition for image generation, automatically generating a dish image that matches the ingredients indicated by the input ingredient information. In response to the dish generation request, the target dish image is displayed on the target terminal's screen. This solves the technical problem in related technologies where the dish information display method is time-consuming due to low matching degree between the displayed dish and the input ingredients, thus shortening the dish information acquisition time.
[0119] In one exemplary embodiment, the above-described apparatus further includes:
[0120] The extraction unit is used to extract, after obtaining the dish generation request sent by the target terminal, the ingredient image of the first reference ingredient from a set of reference ingredients and the ingredient identifier of the second reference ingredient from a set of reference ingredients, wherein the ingredient information of the first reference ingredient is the ingredient image of the first reference ingredient and the ingredient information of the second reference ingredient is the ingredient identifier of the second reference ingredient.
[0121] In one exemplary embodiment, the input unit includes:
[0122] The first input module is used to input the ingredient information of a set of reference ingredients and the kitchenware information of the target kitchenware into the target generation model to obtain the target dish image. The target generation model is also used to use the input ingredient information and the input kitchenware information as control conditions for image generation, and automatically generate dish images that match the ingredients indicated by the input ingredient information and the kitchenware indicated by the input kitchenware information.
[0123] In one exemplary embodiment, the first input module includes:
[0124] The input submodule is used to input a set of reference ingredient information, target kitchenware information, and target cooking method instructions into the target generation model to obtain target dish images. The target generation model is also used to use the input ingredient information, input kitchenware information, and input cooking method instructions as control conditions for image generation, and automatically generate dish images that match the ingredients indicated by the input ingredient information, the kitchenware indicated by the input kitchenware information, and the cooking method indicated by the input instructions.
[0125] In one exemplary embodiment, the above-described apparatus further includes:
[0126] The first determining unit is used to determine the parsed kitchenware identifier as the kitchenware information of the target kitchenware after obtaining the dish generation request sent by the target terminal, after extracting the target text instruction from the dish generation request and parsing the kitchenware identifier from the target text instruction.
[0127] The second determining unit is used to determine the preset kitchenware information that matches the ingredient category of a set of reference ingredients as the kitchenware information of the target kitchenware when the target text instruction is extracted from the dish generation request and the kitchenware identifier is not parsed from the target text instruction, or when the text instruction is not extracted from the dish generation request.
[0128] In one exemplary embodiment, the input unit includes:
[0129] The second input module is used to input the ingredient information of each reference ingredient from a set of reference ingredient information as a control condition into the target diffusion model, so as to perform the following processing operations through the target diffusion model, wherein the target generation model is the target diffusion model:
[0130] Each input information that serves as the control condition for the target diffusion model is converted into an input feature corresponding to each input information. The control condition for the target diffusion model includes a set of ingredient information of reference ingredients.
[0131] The input features corresponding to each input information are fused to obtain the target fusion features, which are used to characterize the control conditions of the target diffusion model.
[0132] The image features of the initial noisy image and the target fusion features are used as the weight matrix of the attention layer of the denoising network in the target diffusion model and input into the attention layer to obtain the target dish features output by the denoising network. The target dish features are the image features obtained after denoising the image features of the initial noisy image.
[0133] The characteristics of the target dish are decoded to obtain an image of the target dish.
[0134] In one exemplary embodiment, the above-described apparatus further includes:
[0135] The second acquisition unit is used to acquire training dish images and a set of training ingredient information corresponding to the training dish images;
[0136] The training unit is used to take the ingredient information of each training ingredient in a set of training ingredients as a control condition for the initial generation model, and use the training dish images to train the initial generation model to obtain the target generation model.
[0137] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules, as part of a device, can operate in environments such as... Figure 1 The hardware environment shown can be implemented through software or hardware, and the hardware environment includes the network environment.
[0138] According to another aspect of the embodiments of this application, a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to execute program code for displaying dish information according to any of the above-described methods in the embodiments of this application.
[0139] Optionally, in this embodiment, the storage medium may be located on at least one of the network devices in the network shown in the above embodiment.
[0140] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps:
[0141] S1, Obtain the dish generation request from the target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients;
[0142] S2, input the ingredient information of a set of reference ingredients into the target generation model to obtain the target dish image. The target generation model is a pre-trained model that uses the input ingredient information as a control condition for image generation and automatically generates a dish image that matches the ingredients indicated by the input ingredient information.
[0143] S3, in response to the dish generation request, controls the display of the target dish image on the target terminal's screen.
[0144] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated in this embodiment.
[0145] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.
[0146] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described method for displaying menu information is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0147] Figure 6 This is a structural block diagram of an optional electronic device according to an embodiment of this application, such as... Figure 6 As shown, it includes a processor 602, a communication interface 604, a memory 606, and a communication bus 608. The processor 602, communication interface 604, and memory 606 communicate with each other via the communication bus 608.
[0148] Memory 606 is used to store computer programs;
[0149] When processor 602 executes a computer program stored in memory 606, it performs the following steps:
[0150] S1, Obtain the dish generation request from the target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients;
[0151] S2, input the ingredient information of a set of reference ingredients into the target generation model to obtain the target dish image. The target generation model is a pre-trained model that uses the input ingredient information as a control condition for image generation and automatically generates a dish image that matches the ingredients indicated by the input ingredient information.
[0152] S3, in response to the dish generation request, controls the display of the target dish image on the target terminal's screen.
[0153] Optionally, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic device and other devices.
[0154] The memory may include RAM, or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0155] As an example, the memory 606 described above may include, but is not limited to, the first acquisition unit 502, the input unit 504, and the control unit 506 in the dish information display device. Furthermore, it may include, but is not limited to, other module units in the dish information display device, which will not be elaborated upon in this example.
[0156] The processors mentioned above can be general-purpose processors, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; they can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0157] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.
[0158] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. The device that implements the above method for displaying menu information can be a terminal device, such as a smartphone (e.g., an Android phone, an iOS phone), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal devices. Figure 6This does not limit the structure of the aforementioned electronic device. For example, the electronic device may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.
[0159] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, ROM, RAM, disk or optical disk, etc.
[0160] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0161] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0162] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0163] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0164] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the solution provided in this embodiment, depending on actual needs.
[0165] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or at least two units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0166] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for displaying menu information, characterized in that, include: Obtain the dish generation request from the target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients; The ingredient information of the set of reference ingredients is input into the target generation model to obtain the target dish image. The target generation model is a pre-trained model that uses the input ingredient information as a control condition for image generation to automatically generate a dish image that matches the ingredients indicated by the input ingredient information. In response to the dish generation request, control the display of the target dish image on the screen of the target terminal; The step of inputting the ingredient information of the set of reference ingredients into the target generation model to obtain the target dish image includes: inputting the ingredient information of each reference ingredient in the set of reference ingredients as a control condition into the target diffusion model, so as to perform the following processing operations through the target diffusion model, wherein the target generation model is the target diffusion model: converting each input information as a control condition of the target diffusion model into an input feature corresponding to each input information, wherein the control condition of the target diffusion model includes the ingredient information of the set of reference ingredients; performing feature fusion with the input features corresponding to each input information to obtain target fusion features, wherein the target fusion features are used to characterize the control condition of the target diffusion model; inputting the image features of the initial noisy image and the target fusion features as the weight matrix of the attention layer of the denoising network in the target diffusion model into the attention layer to obtain the target dish features output by the denoising network, wherein the target dish features are the image features obtained after denoising the image features of the initial noisy image; and decoding the target dish features to obtain the target dish image; The ingredient information of the set of reference ingredients is in a multimodal form, and the ingredient information of each reference ingredient in the set of reference ingredients includes: an ingredient image of each reference ingredient and an ingredient identifier of each reference ingredient.
2. The method according to claim 1, characterized in that, After obtaining the dish generation request sent by the target terminal, the method further includes: Extract the ingredient image of the first reference ingredient from the set of reference ingredients and the ingredient identifier of the second reference ingredient from the set of reference ingredients from the dish generation request. The ingredient information of the first reference ingredient is the ingredient image of the first reference ingredient, and the ingredient information of the second reference ingredient is the ingredient identifier of the second reference ingredient.
3. The method according to claim 1, characterized in that, The step of inputting the ingredient information of the set of reference ingredients into the target generation model to obtain the target dish image includes: The ingredient information of the set of reference ingredients and the kitchenware information of the target kitchenware are input into the target generation model to obtain the target dish image. The target generation model is also used to automatically generate dish images that match the ingredients indicated by the input ingredient information and the kitchenware indicated by the input kitchenware information, respectively, as control conditions for image generation.
4. The method according to claim 3, characterized in that, The step of inputting the ingredient information of the set of reference ingredients and the kitchenware information of the target kitchenware into the target generation model to obtain the target dish image includes: The target generation model inputs the ingredient information of the set of reference ingredients, the kitchenware information of the target kitchenware, and the instruction information of the target cooking method to obtain the target dish image. The target generation model is also used to automatically generate dish images that match the ingredients indicated by the input ingredient information, the kitchenware indicated by the input kitchenware information, and the cooking method indicated by the input instruction information, respectively, as control conditions for image generation.
5. The method according to claim 3, characterized in that, After obtaining the dish generation request sent by the target terminal, the method further includes: If the target text instruction is extracted from the dish generation request and the kitchenware identifier is parsed from the target text instruction, the parsed kitchenware identifier is determined as the kitchenware information of the target kitchenware; If the target text instruction is extracted from the dish generation request but no kitchenware identifier is parsed from the target text instruction, or if no text instruction is extracted from the dish generation request, the preset kitchenware information matching the ingredient category of the set of reference ingredients is determined as the kitchenware information of the target kitchenware.
6. The method according to claim 1, characterized in that, The method further includes: Obtain images of training dishes and information about a set of training ingredients corresponding to the images of training dishes; The ingredient information of each training ingredient in the set of training ingredients is used as a control condition for the initial generation model. The initial generation model is trained using the training dish images to obtain the target generation model.
7. A device for displaying menu information, characterized in that, include: The first acquisition unit is used to acquire the dish generation request of the target terminal, wherein the dish generation request is used to request the generation of a matching dish using a set of reference ingredients; The input unit is used to input the ingredient information of the set of reference ingredients into the target generation model to obtain the target dish image. The target generation model is pre-trained and is used to automatically generate a dish image that matches the ingredients indicated by the input ingredient information by taking the input ingredient information as a control condition for image generation. A control unit is configured to, in response to the dish generation request, control the display of the target dish image on the screen of the target terminal; The input unit includes a second input module, configured to input the ingredient information of each reference ingredient in the set of reference ingredients as a control condition into the target diffusion model, so as to perform the following processing operations through the target diffusion model, wherein the target generation model is the target diffusion model: converting each input information as a control condition of the target diffusion model into an input feature corresponding to each input information, wherein the control condition of the target diffusion model includes the ingredient information of the set of reference ingredients; performing feature fusion with the input features corresponding to each input information to obtain a target fusion feature, wherein the target fusion feature is used to characterize the control condition of the target diffusion model; inputting the image features of the initial noisy image and the target fusion feature as the weight matrix of the attention layer of the denoising network in the target diffusion model into the attention layer to obtain the target dish feature output by the denoising network, wherein the target dish feature is the image feature obtained after denoising the image features of the initial noisy image; and decoding the target dish feature to obtain the target dish image; The ingredient information of the set of reference ingredients is in a multimodal form, and the ingredient information of each reference ingredient in the set of reference ingredients includes: an ingredient image of each reference ingredient and an ingredient identifier of each reference ingredient.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 6.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 6 through the computer program.