Recipe recommendation method and intelligent refrigerator

By integrating the camera and processor in the smart refrigerator, and using the recipe recommendation model of the food identification sub-model and the big language sub-model, the problem that the smart refrigerator cannot recommend recipes when offline is solved, the offline function of recipe recommendations is realized and user privacy is protected.

CN120218231APending Publication Date: 2025-06-27SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214389.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Smart refrigerators cannot effectively recommend recipes when offline or off-network, and there is a risk of leaking user privacy data.

Method used

By integrating cameras, monitors, memory and processors in a smart refrigerator, using recipe recommendation models of food identification sub-models, projection layers and large language sub-models, obtain and process food images and user prompt words, generate recommended recipes, and display them on the monitor.

Benefits of technology

It realizes the recommendation of recipes when the smart refrigerator is offline, protects user privacy, and avoids the risk of data leakage caused by real-time connection with the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218231A_ABST
    Figure CN120218231A_ABST
Patent Text Reader

Abstract

The invention discloses a recipe recommendation method and an intelligent refrigerator, and the method comprises the steps: obtaining a first food material image stored in the intelligent refrigerator, and obtaining a first target prompt word corresponding to each first food material; the first food material image and the first target prompt word are input into a recipe recommendation model to obtain a recommended recipe, the recipe recommendation model comprises a food material recognition sub-model, a projection layer and a big language sub-model, the food material recognition sub-model is used for coding the first food material image to obtain a corresponding food material code, and the projection layer is used for projecting the food material code to the big language sub-model; the projection layer is used for converting the food material code into a feature vector, and the large language sub-model is used for determining the recommended recipe based on the feature vector and the first target prompt word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of electronic technology, including but not limited to a recipe recommendation method and an intelligent refrigerator. Background Art

[0002] Currently, intelligent refrigerators not only have the function of storing food, but also can recommend balanced diet recipes for users. The related technology for recommending recipes requires connecting to a server, and there is a risk of leaking users' private data when connecting to the server to implement recipe recommendation.

[0003] How to implement recipe recommendation when the intelligent refrigerator is not connected to the server, that is, offline or disconnected from the network, has become a technical problem to be solved urgently. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a recipe recommendation method and an intelligent refrigerator.

[0005] The technical solution of the embodiments of the present application is implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a recipe recommendation method, including:

[0007] Obtain a first food ingredient image stored in the intelligent refrigerator, and obtain a first target prompt word corresponding to each first food ingredient;

[0008] Input the first food ingredient image and the first target prompt word into a recipe recommendation model to obtain a recommended recipe, where the recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is used to encode the first food ingredient image to obtain a corresponding food ingredient code, the projection layer is used to convert the food ingredient code into a feature vector, and the large language sub-model is used to determine the recommended recipe based on the feature vector and the first target prompt word.

[0009] In a second aspect, the embodiments of the present application provide an intelligent refrigerator, including: a camera, a display, a memory, and a processor. The memory stores a computer program that can run on the processor. The camera is used to capture a first food ingredient image stored in the intelligent refrigerator; the processor is used to obtain the first food ingredient image, and obtain a first target prompt word corresponding to each first food ingredient; input the first food ingredient image and the first target prompt word into a recipe recommendation model to obtain a recommended recipe, where the recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is used to encode the first food ingredient image to obtain a corresponding food ingredient code, the projection layer is used to convert the food ingredient code into a feature vector, and the large language sub-model is used to determine the recommended recipe based on the feature vector and the first target prompt word; the display is used to display the recommended recipe.

[0010] In a third aspect, the embodiments of the present application provide a recipe recommendation device, including:

[0011] The first acquisition module is configured to acquire the first food ingredient images stored in the smart refrigerator and obtain the first target prompt words corresponding to each first food ingredient.

[0012] The recipe recommendation module is configured to input the first food ingredient images and the first target prompt words into a recipe recommendation model to obtain recommended recipes. The recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is configured to encode the first food ingredient images to obtain corresponding food ingredient encodings. The projection layer is configured to convert the food ingredient encodings into feature vectors. The large language sub-model is configured to determine the recommended recipes based on the feature vectors and the first target prompt words.

[0013] In a fourth aspect, an embodiment of the present application provides a storage medium storing executable instructions, which when executed by a processor, are configured to acquire the first food ingredient images stored in the smart refrigerator and obtain the first target prompt words corresponding to each first food ingredient; input the first food ingredient images and the first target prompt words into a recipe recommendation model to obtain recommended recipes. The recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is configured to encode the first food ingredient images to obtain corresponding food ingredient encodings. The projection layer is configured to convert the food ingredient encodings into feature vectors. The large language sub-model is configured to determine the recommended recipes based on the feature vectors and the first target prompt words.

[0014] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program or instructions, which when executed by a processor, are configured to acquire the first food ingredient images stored in the smart refrigerator and obtain the first target prompt words corresponding to each first food ingredient; input the first food ingredient images and the first target prompt words into a recipe recommendation model to obtain recommended recipes. The recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is configured to encode the first food ingredient images to obtain corresponding food ingredient encodings. The projection layer is configured to convert the food ingredient encodings into feature vectors. The large language sub-model is configured to determine the recommended recipes based on the feature vectors and the first target prompt words. Description of the Drawings

[0015] Figure 1A It is a schematic flowchart of the implementation of a recipe recommendation method provided by an embodiment of the present application;

[0016] Figure 1B It is a schematic structural diagram of a recipe recommendation model provided by an embodiment of the present application;

[0017] Figure 2A It is a schematic flowchart of the implementation of training a recipe recommendation model provided by an embodiment of the present application;

[0018] Figure 2BSchematic diagram of training a recipe recommendation model using a first data set provided by an embodiment of the present application;

[0019] Figure 2C Schematic diagram of training a recipe recommendation model using a second data set provided by an embodiment of the present application;

[0020] Figure 3 Schematic diagram of an implementation process for prompting supplementary ingredients provided by an embodiment of the present application;

[0021] Figure 4A Schematic diagram of the structure of an intelligent refrigerator provided by an embodiment of the present application;

[0022] Figure 4B Schematic diagram of an implementation process of an offline recipe recommendation method provided by an embodiment of the present application;

[0023] Figure 5 Schematic diagram of the composition structure of a recipe recommendation device provided by an embodiment of the present application;

[0024] Figure 6 Schematic diagram of a hardware entity of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the embodiments of the application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0026] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0027] In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0029] An embodiment of the present application provides a recipe recommendation method, as Figure 1A shown, the method includes:

[0030] Step S110: Obtain the first food images stored in the smart refrigerator and obtain the first target prompt words corresponding to each first food item.

[0031] Here, a smart refrigerator refers to a type of refrigerator that can perform intelligent control of the refrigerator and intelligent management of food. At least one camera can be set in the smart refrigerator to capture the food ingredients stored in the smart refrigerator, obtain at least one first food image, and send the first food image to the processor of the smart refrigerator, so that the processor can determine the food ingredients stored in the smart refrigerator by identifying the at least one first food image obtained. Among them, the camera can be set to take pictures at regular intervals or trigger the shooting manually to adapt to different usage scenarios. For example, 12 high-definition built-in cameras can be set at different positions in the smart refrigerator to capture the food ingredients stored in the refrigerator compartment in real time to obtain the first food images.

[0032] In the implementation process, features can be extracted from each recognized first food image, and these features can include color, shape, texture, etc. In some embodiments, the extracted features can be matched with a predefined food ingredient prompt word library. The prompt word library can contain the accurate names, aliases, shelf life information (such as "fresh", "about to expire", etc.) of various food ingredients, as well as possible cooking suggestions (such as "suitable for grilling", "recommended to refrigerate", etc.). According to the matching results, the first target prompt words corresponding to each food ingredient can be generated. The prompt words should be concise and clear, and can accurately reflect the status and characteristics of the food ingredients.

[0033] In some embodiments, the display set in the smart refrigerator can also be used to communicate with the user using the smart refrigerator to obtain the user's usage requirements for each recognized food ingredient. For example, the user can put forward requirements such as "want to eat foods that are helpful for skin care". As a result, the first target prompt words corresponding to each first food ingredient can be obtained based on the user's usage requirements and the recognized food ingredients. Among them, the display can be a touch screen display, which can use the touch screen to communicate with the user, and a voice communication device can also be set to communicate with the user using voice prompts to obtain the user's usage requirements for the food ingredients.

[0034] Step S120: Input the first food image and the first target prompt words into a recipe recommendation model to obtain a recommended recipe. Among them, the recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is used to encode the first food image to obtain a corresponding food ingredient code, the projection layer is used to convert the food ingredient code into a feature vector, and the large language sub-model is used to determine the recommended recipe based on the feature vector and the first target prompt words.

[0035] Here, the recipe recommendation model set in the smart refrigerator includes a food ingredient recognition sub-model, a projection layer, and a large language model. Among them, the food ingredient recognition sub-model is used to encode the input first food ingredient image to obtain the corresponding food ingredient code. The projection layer is used to convert the input food ingredient code into a feature vector that can be recognized by the large language sub-model. The large language sub-model can output a recommended recipe based on the input feature vector and the first target prompt word, and use the display set in the smart refrigerator to display the recommended recipe for the user.

[0036] Figure 1B It is a schematic structural diagram of a recipe recommendation model provided by an embodiment of the present application. As Figure 1B shown, the recipe recommendation model includes a dedicated food ingredient recognition network (Image Encoder) 11, a projection layer (Projection Layer) 12, and a pre-trained large language model (Large Language Models, LLM) 13. Among them,

[0037] The dedicated food ingredient recognition network 11, that is, the food ingredient recognition sub-model, can be used as an image encoder. Using the dedicated food ingredient recognition network 11 as the image encoder can effectively improve the food ingredient recognition accuracy of the multi-modal large model and avoid the hallucination phenomenon.

[0038] The projection layer 12 is used to convert the image code output by the image encoder 11 into elements (tokens) that can be accepted and understood by the large language model, that is, to output a feature vector. The projection layer is usually a linear layer, and its main function is to project the input data from a high-dimensional space to a low-dimensional space. In this process, the data is transformed through matrix multiplication, but does not pass through a non-linear activation function. Therefore, the output of the projection layer is a linear combination of the input data. For example, a multi-layer perceptron (MLP) can be used as the projection layer. The MLP is a feedforward artificial neural network model composed of an input layer, a hidden layer, and an output layer, where the hidden layer can be one or more layers. The MLP as the projection layer can be used for feature extraction. Through the non-linear transformation and fully connected structure of the MLP, the original input data can be converted into a more expressive feature vector.

[0039] The pre-trained large language model 13 can output recommended recipes based on the input feature vectors and the first target prompt words. For example, the large language model can be models such as LLaMA2 and Baichuan. Among them, LLaMA2 uses a training dataset of 2 trillion tokens, giving it powerful language understanding and generation capabilities. It can be applied to various natural language processing tasks such as question answering systems, text generation, and machine translation. The Baichuan model has more than 100 billion parameters and can be applied to various scenarios such as intelligent question answering, text generation, machine translation, and recommendation systems. Moreover, Baichuan has obvious advantages in Chinese tasks.

[0040] In the embodiments of the present application, first, the first food ingredient image stored in the smart refrigerator is obtained, and the first target prompt words corresponding to each first food ingredient are obtained. Then, the first food ingredient image and the first target prompt words are input into the recipe recommendation model to obtain the recommended recipe. In this way, since the ingredient recognition sub-model is used in the recipe recommendation model, the accuracy of ingredient recognition can be effectively improved. And since the recipe recommendation model is set in the intelligent control system of the smart refrigerator, it can be realized that when the smart refrigerator is not connected to the network, that is, in an offline state, by taking the first food ingredient image and obtaining the first target prompt words, and then outputting the recommended recipe offline.

[0041] In some embodiments, the embodiments of the present application further provide a method for obtaining the first target prompt words, which can be implemented through the following steps:

[0042] Step S130: Use the ingredient recognition sub-model to recognize the first food ingredient image to obtain the first food ingredient recognition result;

[0043] Here, the ingredient recognition sub-model can be an ingredient recognition sub-model that has been trained. This ingredient recognition model can accurately recognize a variety of ingredients, including vegetables, fruits, meats, seafood, etc.

[0044] In some embodiments, before inputting the image into the ingredient recognition sub-model, some preprocessing steps can be performed to improve the recognition accuracy. These steps may include resizing the image, cropping, normalizing, denoising, etc.

[0045] The preprocessed first food ingredient image is input into the ingredient recognition sub-model. The ingredient recognition sub-model will automatically extract the features in the image and classify the ingredients based on these features. One or more possible ingredient labels, as well as the confidence level of each label (that is, the probability that the model believes the image is the corresponding ingredient), will be output. Select the label with the highest confidence level as the first food ingredient recognition result.

[0046] Step S140: In response to a preset prompt word input instruction, obtain the first preset prompt word;

[0047] Here, the preset prompt word input instruction can be triggered by the user. When the user determines the first preset prompt word, the user can trigger this preset prompt word input instruction. For example, the user can use the display to input the first preset prompt word of "long-term fitness" and trigger the preset prompt word input instruction, so that the intelligent control system of the smart refrigerator can obtain the first preset prompt word of "long-term fitness".

[0048] Step S150: Combine the first food ingredient recognition result and the first preset prompt word to obtain the first target prompt word.

[0049] In the implementation process, the first food ingredient recognition result and the first preset prompt word can be synthesized to obtain the first target prompt word. For example: The first food ingredient recognition result is "cucumber, tomato, milk, pork", and the first preset prompt word is "I am a fitness person", then the first target prompt word can be obtained as "I am a fitness person. There are ingredients such as cucumber, tomato, milk, and pork in the refrigerator. Please help me recommend a healthy recipe".

[0050] In some embodiments, before obtaining the first preset prompt word in response to the preset prompt word input instruction, the method further includes:

[0051] Display a prompt word template;

[0052] Here, the prompt word template is a tool that can help users make personalized settings according to their specific needs and preferences. The prompt word template usually contains a series of questions or prompts to guide users to input relevant information, so as to generate prompt words that meet their requirements.

[0053] In the implementation process, the prompt word template can be displayed using the display set on the smart refrigerator.

[0054] Obtain the preset prompt word input instruction, where the preset prompt word input instruction is an input instruction triggered by the user and used to determine the first preset prompt word based on the prompt word template.

[0055] In the implementation process, through the displayed prompt word template, users can be guided to fill in the information of each part one by one. For example, interaction can be carried out through online forms, chatbots, etc. After the user fills in, collect the information they provide, and organize and classify it. Ensure that the user's needs and preferences are accurately understood.

[0056] For example, the prompt word template includes "May I ask your age", "May I ask your gender", "What is your occupation", "What are your requirements for the recommended recipes", etc. The user can set their first preset prompt word as "I am a female and I need a fat-reducing meal" according to the prompt word template.

[0057] During the implementation process, when the user determines that the first preset prompt word set can meet the prompt word requirements for the recommended recipe, a preset prompt word input instruction can be triggered to input the first preset prompt word into the recipe recommendation model set in the smart refrigerator.

[0058] In the embodiments of the present application, first, the first food ingredient image is recognized by the food ingredient recognition sub-model to obtain the first food ingredient recognition result; then, in response to the preset prompt word input instruction, the first preset prompt word is obtained; finally, the first food ingredient recognition result and the first preset prompt word are combined to obtain the first target prompt word. In this way, the obtained first target prompt word includes the first preset prompt word input by the user and the first food ingredient recognition result, and can more accurately express the user's needs and the existing food ingredients in text form, providing richer input for the large language model, so that the output of the large language model is more effective and accurate and can meet the user's needs for the recommended recipe.

[0059] In some embodiments, the embodiments of the present application provide a method for training a recipe recommendation model, as Figure 2A shown, which can be implemented through the following steps:

[0060] Step S210, obtain a trained food ingredient recognition sub-model, a projection layer with parameters initialized to follow a Gaussian distribution, and a trained large language sub-model;

[0061] During the implementation process, a model related to food ingredient recognition can be obtained first, and a food ingredient picture data set is collected and labeled. The data set is used to train the model, and after the training is completed, the model weights and configuration files are saved to obtain a trained food ingredient recognition sub-model.

[0062] The weights of the projection layer are initialized using a Gaussian distribution (normal distribution) to obtain a projection layer with parameters initialized to follow a Gaussian distribution.

[0063] During the implementation process, a large language model to be trained can be obtained from an existing pre-trained model library, and a text data set is collected and labeled. The data set is used to train the model, and after the training is completed, the model weights and configuration files are saved to obtain a trained large language sub-model.

[0064] Step S220, adjust the parameters of the projection layer while fixing the parameters of the food ingredient recognition sub-model and the parameters of the large language sub-model.

[0065] During the implementation process, when training the recipe recommendation model, the parameters of the ingredient recognition sub-model and the large language sub-model can be fixed, and the projection layer connecting the ingredient recognition sub-model and the large language sub-model can be trained to optimize the recipe recommendation effect. Here, the projection layer is a connection module between the ingredient recognition sub-model and the large language sub-model. In order to enable all feature vectors input into the large language sub-model to be understood by the large language model and make correct content outputs, the projection layer plays a crucial role, so the projection layer is trained.

[0066] In the embodiment of the present application, first, obtain the trained ingredient recognition sub-model, the projection layer with parameter initialization following a Gaussian distribution, and the trained large language sub-model; then, while fixing the parameters of the ingredient recognition sub-model and the large language sub-model, adjust the parameters of the projection layer. In this way, since the projection layer is a connection module between the ingredient recognition sub-model and the large language sub-model, in order to enable all feature vectors input into the large language sub-model to be understood by the large language model and make correct content outputs, the projection layer plays a crucial role. Therefore, adjusting the parameters of the projection layer can achieve the effect of optimizing the recipe recommendation, so that the trained recipe recommendation model can output recommended recipes more accurately and efficiently.

[0067] In some embodiments, the above step S220 "while fixing the parameters of the ingredient recognition sub-model and the large language sub-model, adjust the parameters of the projection layer" can be implemented through the following steps:

[0068] Step 221: Use the first data set to adjust the parameters of the projection layer while fixing the parameters of the ingredient recognition sub-model and the large language sub-model, and obtain the recipe recommendation model that has completed the first-stage training;

[0069] Here, the first data set can be the pre-training data pairs generated by existing multi-modal models. The characteristics of this first data set are large scale and low quality. During the implementation process, a pre-trained multi-modal model can be selected as the basis first. This model should have the ability to process images, texts, etc., and has shown good performance in the recipe recommendation task. Then, existing image and text data sets can be used to match images and texts through methods such as keyword search and image description generation to form image-text pairs. The multi-modal model itself can also be used to generate image-text pairs. For example, given an image, the model can be used to generate a related text description; conversely, given a text description, the model can be used to generate a matching image. For example, given an image of the ingredients inside a refrigerator, the model can be used to generate a related recipe recommendation.

[0070] Figure 2BThe training schematic diagram of training the recipe recommendation model using the first dataset provided by the embodiments of the present application is as follows Figure 2B As shown, during the training process, the output feature vector of the projection layer 12 and the recipe recommendation text instruction can be input into the large language model 13 using the first dataset to train the recipe recommendation model. For example, the recipe recommendation text instruction can be "I am a fitness person, please recommend healthy recipes for me."

[0071] As Figure 2B shown, during the training process, the parameter freezing module is the dedicated ingredient recognition network 11, that is, the ingredient recognition sub-model, and the pre-trained large language module 13, that is, the large language sub-model. The parameter update module is the projection layer 12.

[0072] Step 222: Use the second dataset to train the recipe recommendation model that has completed the first-stage training to obtain a recipe recommendation model that has completed the second-stage training;

[0073] Among them, the quantity and quality of the second dataset are higher than the data quality of the second dataset.

[0074] In some embodiments, the first dataset is obtained using a multi-modal large language model; the second dataset is obtained by screening the data in the first dataset.

[0075] Here, the second dataset can be thousands of pairs of high-quality labeled datasets screened from the first dataset. During the implementation process, since the training data pairs in the first dataset may contain noise or inaccurate information, data cleaning and filtering are performed. It can be achieved through manual inspection, using rules or models for automatic filtering, etc., to obtain the second dataset that has completed screening to ensure the quality and accuracy of the data pairs.

[0076] During the implementation process, using the second dataset to train the recipe recommendation model that has completed the first-stage training is crucial for improving the performance of the model.

[0077] In the embodiments of the present application, first, using the first dataset, while fixing the parameters of the ingredient recognition sub-model and the parameters of the large language sub-model, train the projection layer to obtain a recipe recommendation model that has completed the first-stage training; then use the second dataset to train the recipe recommendation model that has completed the first-stage training to obtain a recipe recommendation model that has completed the second-stage training. In this way, through two-stage training, it is possible to realize the pre-training of the recipe recommendation model using the first dataset and the fine-tuning of the recipe recommendation model using the second dataset. By reasonably selecting the dataset, the training effect of the recipe recommendation model can be effectively improved, and the generalization ability and performance of the recipe recommendation model can be improved.

[0078] In some embodiments, the above step S221, "using the first data set, while fixing the parameters of the ingredient recognition sub-model and the parameters of the large language sub-model, adjusting the parameters of the projection layer to obtain a recipe recommendation model that has completed the first stage of training", can be implemented through the following steps:

[0079] Step 2211: Obtain a second preset prompt word, where the second preset prompt word is determined by the user based on a prompt word template;

[0080] Here, the second preset prompt word can be, for example, Figure 2B the "recipe recommendation text instruction" as shown. In the implementation process, the user can view the prompt word template based on the display screen of the smart refrigerator, and then determine the second preset prompt word. For example, the second preset prompt word can be "I am a fitness person, please recommend healthy recipes for me".

[0081] Step 2212: Using the first data set and the second preset prompt word, while fixing the parameters of the ingredient recognition sub-model and the parameters of the large language sub-model, adjust the parameters of the projection layer to obtain the recipe recommendation model that has completed the first stage of training.

[0082] In the implementation process, as Figure 2B shown, in the process of training the recipe recommendation model using the first data set, the "recipe recommendation text instruction" and the feature vector output by the projection layer 12 can be simultaneously input into the trained large language model 13 to train the recipe recommendation model, and the parameters of the projection layer are adjusted during the training process.

[0083] In some embodiments, a fixed number of iterations can be preset in advance, and the training is completed when the number of training times meets the fixed iteration times; a loss function can also be set. In consecutive multiple iterations, if the value of the loss function changes very little or tends to be stable, this may indicate that the model has converged and the training can be stopped.

[0084] In the embodiments of the present application, first, a second preset prompt word is obtained, and then, using the first data set and the second preset prompt word, while fixing the parameters of the ingredient recognition sub-model and the parameters of the large language sub-model, the parameters of the projection layer are adjusted to obtain a recipe recommendation model that has completed the first stage of training. In this way, in the first training stage, the model uses the first data set including a large amount of data and the second preset prompt word set in advance for learning, and these data cover a wide range of fields and language structures. This enables the model to master the underlying laws of the language, such as lexical semantics, syntactic structures, etc., thereby improving its generalization ability when facing new data.

[0085] In some embodiments, the above step S222, "using the second data set to train the recipe recommendation model that has completed the first-stage training to obtain a recipe recommendation model that has completed the second-stage training", can be implemented through the following steps:

[0086] Step 2221: Obtain a second target prompt word based on the second data set, where the second target prompt word is obtained by combining the second ingredient recognition result and the second preset prompt word. The second ingredient recognition result is obtained by using the ingredient recognition sub-model to recognize the second ingredient images in the second data set, and the second preset prompt word is determined by the user based on the prompt word template;

[0087] Here, the second target prompt word is obtained by combining the second ingredient recognition result and the second preset prompt word. Figure 2C This is a training schematic diagram of using the second data set to train the recipe recommendation model provided by the embodiments of the present application. As Figure 2C shown, during the training process, the second data set can be used to input the output feature vector of the projection layer 12 and the second target prompt word into the large language model 13 to train the recipe recommendation model. For example, the second ingredient recognition result can be "There are cucumbers, tomatoes, milk, and pork in the refrigerator", and the second target prompt word can be "I am a fitness person". Then, the second target prompt word obtained by combining the second ingredient recognition result and the second preset prompt word can be "I am a fitness person. There are ingredients such as cucumbers, tomatoes, milk, and pork in the refrigerator. Please help me recommend healthy recipes".

[0088] Step 2222: Use the second data set and the second target prompt word to train the recipe recommendation model that has completed the first-stage training to obtain a recipe recommendation model that has completed the second-stage training.

[0089] During the implementation process, as Figure 2C shown, the second data set can be used to input the output feature vector of the projection layer 12 and the second target prompt word into the large language model 13 to train the recipe recommendation model, that is, to fine-tune the projection layer 12 to obtain a recipe recommendation model that has completed the second-stage training.

[0090] In the embodiments of the present application, first, a second target prompt word is obtained based on the second data set, and then the recipe recommendation model that has completed the first-stage training is trained using the second data set and the second target prompt word to obtain a recipe recommendation model that has completed the second-stage training. In this way, since the second target prompt word is obtained by combining the second ingredient recognition result and the second preset prompt word, compared with the second preset prompt word used in the first-stage training, the expressed text is richer and can more accurately express the conditions and requirements of the recommendation, enabling fine-tuning of the recipe recommendation model. In the fine-tuning stage (the second stage), the selected second data set is used for training, which enables the model to more precisely adapt to a specific scenario or task. Through fine-tuning, the performance of the recipe recommendation model in the recipe recommendation task can be significantly improved.

[0091] In some embodiments, the step "input the first ingredient image and the first target prompt word into the recipe recommendation model to obtain a recommended recipe" in step S120 above can be implemented through the following process: when the smart refrigerator is not connected to the network, input the first ingredient image and the first target prompt word into the recipe recommendation model stored in the smart refrigerator to obtain the recommended recipe.

[0092] Here, when the smart refrigerator is not connected to the network, it means that the smart refrigerator cannot access Internet resources in real time, such as online databases, remote servers, or cloud computing services. In this case, the smart refrigerator can provide recipe recommendation services based on locally stored data and functions. The recipe recommendation model is pre-trained and stored in the local storage of the smart refrigerator, and may contain a large amount of recipe information and related technologies such as image recognition and natural language processing. This recipe recommendation model can generate a recipe recommendation that matches the input based on the input ingredient image and target prompt word.

[0093] The first ingredient image can be captured by a camera built into the smart refrigerator or input into the smart refrigerator by other means (such as manual upload).

[0094] The first target prompt word is input by the user and can be used to describe the user's requirements for recipe recommendations, the type or taste of the dish they want to make, etc.

[0095] During the implementation process, the first ingredient image and the first target prompt word can be input into the recipe recommendation model stored locally. The recipe recommendation model can search and match in the locally stored recipe database based on the input ingredients and target prompt words, generate one or more recommended recipes that match the input, and display them to the user through the user interface of the display installed in the smart refrigerator.

[0096] In the embodiments of the present application, in the case where the smart refrigerator is not connected to the network, by using the recipe recommendation model stored locally, the recipe recommendation service can still be provided for users. This recommendation method can protect the user's privacy data and effectively avoid the risk of leaking privacy data.

[0097] In some embodiments, the embodiments of the present application further provide a method for prompting to supplement ingredients, as Figure 3 shown, which can be implemented through the following steps:

[0098] Step S310: Determine that there is a target ingredient in the recommended recipe that is not stored in the smart refrigerator;

[0099] In some embodiments, due to the recommended requirements for users, there may be at least one target ingredient in the output recommended recipe that is not stored in the smart refrigerator. For example, when the user's recommended requirement is "low-fat meal", it is recognized that lettuce or broccoli in the low-fat meal is not stored in the smart refrigerator.

[0100] Step S320: Output a prompt message to prompt the user to supplement the target ingredient.

[0101] During the implementation process, the prompt message can be output based on the target ingredient to prompt the user to supplement the target ingredient. For example, the prompt message "Lettuce is also needed to make a low-fat meal" can be output to prompt the user to supplement lettuce so as to be able to complete the production of the recommended low-fat meal.

[0102] In the embodiments of the present application, first, it is determined that there is a target ingredient in the recommended recipe that is not stored in the smart refrigerator; then, a prompt message is output to prompt the user to supplement the target ingredient. In this way, the diversity of the recommended recipes can be improved, that is, the recommended recipes can include target ingredients that are not stored in the smart refrigerator. And further enable the user to complete the production of the recommended recipe by supplementing the target ingredient according to the output prompt message.

[0103] Figure 4A The following is a schematic structural diagram of a smart refrigerator provided by an embodiment of the present application, as Figure 4A shown. The smart refrigerator at least includes a camera 41, a display 42, a memory 43, and a processor 44, wherein,

[0104] The camera 41 is used to capture a first food ingredient image stored in the smart refrigerator.

[0105] A processor 44 is configured to obtain the first food ingredient image and obtain first target prompt words corresponding to each first food ingredient; input the first food ingredient image and the first target prompt words into a recipe recommendation model to obtain a recommended recipe, where the recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is configured to encode the first food ingredient image to obtain a corresponding food ingredient code, the projection layer is configured to convert the food ingredient code into a feature vector, and the large language sub-model is configured to determine the recommended recipe based on the feature vector and the first target prompt words.

[0106] A display 42 is configured to display the recommended recipe.

[0107] A memory 43 stores a computer program that can run on the processor 44.

[0108] In the embodiment of the present application, for the intelligent refrigerator provided, the camera first obtains the first food ingredient image in the intelligent refrigerator. Then, the processor uses the recipe recommendation model stored in the memory to obtain the first food ingredient image and obtain first target prompt words corresponding to each first food ingredient. Then, the first food ingredient image and the first target prompt words are input into the recipe recommendation model to obtain a recommended recipe, and the display is used to display the recommended recipe. In this way, since the food ingredient recognition sub-model is used in the recipe recommendation model, the accuracy of food ingredient recognition can be effectively improved. And since the recipe recommendation model is set in the intelligent control system of the intelligent refrigerator, it can be realized that when the intelligent refrigerator is not connected to the network, that is, in an offline state, by taking the first food ingredient image and obtaining the first target prompt words, and then outputting the recommended recipe offline.

[0109] In the related art, the working mode followed by the intelligent refrigerator recipe recommendation system is as follows: The food ingredients existing in the refrigerator are identified through a built-in camera and a food ingredient recognition model, the types of food ingredients recognized by the recognition model are sent to the server side, and a recipe recommendation request is sent to the server. The server searches for recipe recommendation results through a certain retrieval method and then returns the results to the intelligent refrigerator terminal. Finally, the intelligent refrigerator provides options for food ingredient recognition to the user through a display device. Since the intelligent refrigerator needs to interact with the server during the process of outputting the recommended recipe, there is a risk of leaking user privacy. And this solution can output the recommended recipe offline, thereby protecting user privacy.

[0110] An embodiment of the present application provides an offline recipe recommendation method, which is applied to an intelligent refrigerator, as Figure 4B shown, and can be implemented through the following steps:

[0111] Step S410, construct a data set;

[0112] Here, this step can be used to collect the dataset required for the dedicated network training and perform data cleaning. During the construction of this dataset, in order to save the cost of manual annotation, the existing Multimodal Large Language Models (LLMs) can be utilized to generate the original image and recipe dataset, and then thousands of pairs of low-quality datasets and thousands of pairs of high-quality datasets can be obtained through manual screening and correction, which are respectively used for the model training in the first stage and the second stage.

[0113] Step S420: Construct a knowledge-based multimodal large model and train the knowledge-based multimodal large model;

[0114] Figure 1B It is a schematic structural diagram of a multimodal large model provided by an embodiment of the present application. As Figure 1B shown, the multimodal large model includes a dedicated ingredient recognition network (Image Encoder) 11, a projection layer (Projection Layer) 12, and a pre-trained large language model 13, where

[0115] This multimodal large model (recipe recommendation model) can be used to receive text and image data and output the recommended recipes in text form.

[0116] The dedicated ingredient recognition network (Image Encoder) 11 can be used as an image encoder. Using the dedicated ingredient recognition network 11 as the image encoder can effectively improve the ingredient recognition accuracy of the multimodal large model and avoid hallucination phenomena. The projection layer 12 is used to convert the image encoding output by the image encoder 11 into elements (tokens) that can be accepted and understood by the large language model. Here, since the input to the pre-trained large language model 13 includes text and images, and the LLM is a language model that can receive text as input, the image content needs to be converted or encoded into a form that the LLM can understand before being input into the LLM, which is the above-mentioned token.

[0117] When initializing this multimodal large model, the weights of the pre-trained dedicated ingredient recognition network 11 can be adopted. The parameters of the projection layer 12 are initialized to follow the standard Gaussian distribution, and the initialization of the large language model 13 adopts the weights of the pre-trained large model, such as LLaMA2.

[0118] During the process of model training, freeze the parameters of the dedicated ingredient recognition network 11 and the parameters of the large language model 13, and train the projection layer 12. Here, the projection layer 12 is an interface module between the image encoder 11 and the large language model 13 in the multi-modal large model. In order to enable all tokens input to the large language model 13 to be understood by the large language model 13 and produce correct content output, the projection layer 12 plays a crucial role, so train this projection layer 12.

[0119] During implementation, the training of the multi-modal large model can be completed in two stages as follows:

[0120] Stage 1 training:

[0121] Dataset: Use an existing multi-modal model to generate pre-training data pairs (large scale, low quality).

[0122] Network structure: Freeze the image encoder 11 and the large language model 13, and fine-tune the projection layer 12.

[0123] Model training: The completed trained model is used as the pre-training model for stage 2.

[0124] Figure 1B It can be used as a schematic diagram of stage 1 training, as Figure 1B shown. During the training process, input the output tokens of the projection layer 12 and the recipe recommendation text instruction into the large language model 13 to train the multi-modal large model. For example, the recipe recommendation text instruction can be "I am a fitness person, please help me recommend healthy recipes".

[0125] Stage 2 training:

[0126] Dataset: Screen thousands of pairs of high-quality labeled datasets.

[0127] Network structure: Freeze the image encoder 11 and the large language model 13, and fine-tune the projection layer 12.

[0128] Model training: Use the model that has completed stage 1 training as the pre-training model for stage 2.

[0129] Figure 2C This is a schematic diagram of stage 2 training provided by an embodiment of the present application, as Figure 2C shown. During the training process, input the output tokens of the projection layer 12 and the text prompt words output based on the image encoder 11 into the large language model 13. Among them, the text prompt words are text prompt words obtained based on the ingredient recognition result and the prompt word template. That is, the text prompt words can be synthetic text prompt words. For example, the text prompt words can be "I am a fitness person, and there are cucumbers, tomatoes, milk, and pork in the refrigerator. Please help me recommend healthy recipes".

[0130] Step S430: Use a multimodal large model for recipe recommendation.

[0131] After the model training is completed, the multimodal large model can be deployed as the core component of the intelligent refrigerator recipe recommendation system in the intelligent refrigerator system.

[0132] During the inference process, the model receives the pictures and prompt words transmitted by the built-in camera of the refrigerator. The prompt words are personalized through interaction with the prompt word template and combined with the prediction results of the dedicated ingredient recognition model to form the final text prompt words.

[0133] The above-mentioned trained dedicated knowledge multimodal large model can receive pictures and text prompt words and output recipe recommendation results for users to refer to.

[0134] In the embodiment of the present application, a new type of intelligent refrigerator recipe recommendation system is provided. This system can, under offline conditions, recommend recipes that meet the personalized needs of different people based on the existing ingredients through the built-in camera of the refrigerator. The main components of this new type of intelligent refrigerator recipe recommendation system include a knowledge-based multimodal large language model, a built-in camera of the refrigerator, and a display device. Users only need to simply operate (such as selecting a prompt word template through a click operation) to set personalized preferences, and they can automatically and quickly obtain recipes and cooking methods that meet their personalized needs. In this way, the intelligent refrigerator does not need to be connected to the server and can work normally under offline and disconnected network conditions to achieve recipe recommendation, realizing fast and real-time recommendation, meeting the personalized needs of users, and protecting the privacy data security of users.

[0135] Based on the foregoing embodiments, the embodiment of the present application provides a recipe recommendation device. The device includes each module included, each module includes each sub-module, and each sub-module includes units, which can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; during the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0136] Figure 5 It is a schematic diagram of the composition structure of the recipe recommendation device provided by the embodiment of the present application. As Figure 5 shown, the device 500 includes:

[0137] The first acquisition module 510 is configured to acquire a first food ingredient image stored in the smart refrigerator and obtain a first target prompt word corresponding to each first food ingredient;

[0138] The food ingredient recommendation module 520 is configured to input the first food ingredient image and the first target prompt word into a recipe recommendation model to obtain a recommended recipe, where the recipe recommendation model includes a food ingredient recognition sub-model, a projection layer, and a large language sub-model. The food ingredient recognition sub-model is configured to encode the first food ingredient image to obtain a corresponding food ingredient code, and the projection layer is configured to convert the food ingredient code into a feature vector. The large language sub-model is configured to determine the recommended recipe based on the feature vector and the first target prompt word.

[0139] In some embodiments, the food ingredient recommendation device further includes an identification module, a second acquisition module, and a combination module. The identification module is configured to use the food ingredient recognition sub-model to identify the first food ingredient image to obtain a first food ingredient recognition result. The second acquisition module is configured to acquire a first preset prompt word in response to a preset prompt word input instruction. The combination module is configured to combine the first food ingredient recognition result and the first preset prompt word to obtain the first target prompt word.

[0140] In some embodiments, the following modules can be used to train the recipe recommendation model: a third acquisition module and a training module. The third acquisition module is configured to acquire a trained food ingredient recognition sub-model, a projection layer with parameters initialized to follow a Gaussian distribution, and a trained large language sub-model. The training module is configured to adjust the parameters of the projection layer while fixing the parameters of the food ingredient recognition sub-model and the large language sub-model.

[0141] In some embodiments, the training module includes a first training sub-module and a second training sub-module. The first training sub-module is configured to use a first data set to adjust the parameters of the projection layer while fixing the parameters of the food ingredient recognition sub-model and the large language sub-model to obtain a recipe recommendation model that has completed the first stage of training. The second training sub-module is configured to use a second data set to train the recipe recommendation model that has completed the first stage of training to obtain a recipe recommendation model that has completed the second stage of training. The quality of the second data set is higher than the quality of the second data set.

[0142] In some embodiments, the first data set is obtained using a multi-modal large language model; the second data set is obtained by screening the data in the first data set.

[0143] In some embodiments, the first training sub-module includes a first acquisition unit and a first training unit. Among them, the first acquisition unit is configured to acquire a second preset prompt word, where the second preset prompt word is determined by the user based on a prompt word template; the first training unit is configured to use the first data set and the second preset prompt word to adjust the parameters of the projection layer while fixing the parameters of the ingredient recognition sub-model and the parameters of the large language sub-model, so as to obtain the recipe recommendation model that has completed the first stage of training.

[0144] In some embodiments, the second training sub-module includes a second acquisition unit and a second training unit. Among them, the second acquisition unit is configured to acquire a second target prompt word based on the second data set, where the second target prompt word is obtained by combining the second ingredient recognition result and the second preset prompt word. The second ingredient recognition result is obtained by using the ingredient recognition sub-model to recognize the second ingredient image in the second data set, and the second preset prompt word is determined by the user based on a prompt word template; the second training unit is configured to use the second data set and the second target prompt word to train the recipe recommendation model that has completed the first stage of training, so as to obtain the recipe recommendation model that has completed the second stage of training.

[0145] In some embodiments, the ingredient recommendation module 520 is further configured to, when the smart refrigerator is not connected to the network, input the first ingredient image and the first target prompt word into the recipe recommendation model stored in the smart refrigerator to obtain the recommended recipe.

[0146] In some embodiments, the recipe recommendation device further includes a determination module and an output module. Among them, the determination module is configured to determine that there is a target ingredient in the recommended recipe that is not stored in the smart refrigerator; the output module is configured to output a prompt message to prompt the user to supplement the target ingredient.

[0147] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0148] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0149] Correspondingly, the embodiments of the present application provide a storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the recipe recommendation method provided in the above embodiments.

[0150] Correspondingly, the embodiments of the present application provide an electronic device. Figure 6 It is a schematic diagram of a hardware entity of the electronic device provided in the embodiments of the present application. As Figure 6 shown, the hardware entity of the device 600 includes: a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the program, it implements the steps in the recipe recommendation method provided in the above embodiments.

[0151] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or already processed by the processor 602 and each module in the electronic device 600 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (RAM).

[0152] It should be pointed out here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0153] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above processes does not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0154] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0155] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0156] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, each functional unit in the embodiments of the present application can be all integrated in a processing unit, or each unit can be separately a unit, or two or more units can be integrated in one unit; the above integrated unit can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0158] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: various media that can store program codes such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0159] Alternatively, if the above integrated units of the present application are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present application, in essence or the parts that contribute to the related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc.) to execute all or part of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: various media that can store program codes such as removable storage devices, ROM, magnetic disks, or optical discs.

[0160] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0161] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0162] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0163] As described above, it is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A recipe recommendation method, applied to a smart refrigerator, comprising: Acquire first food images stored in the smart refrigerator, and obtain first target prompt words corresponding to each first food; The first ingredient image and the first target prompt word are input into a recipe recommendation model to obtain a recommended recipe, wherein the recipe recommendation model includes an ingredient recognition sub-model, a projection layer and a large language sub-model, the ingredient recognition sub-model is used to encode the first ingredient image to obtain a corresponding ingredient code, the projection layer is used to convert the ingredient code into a feature vector, and the large language sub-model is used to determine the recommended recipe based on the feature vector and the first target prompt word.

2. The method of claim 1, further comprising: Using the food recognition sub-model to recognize the first food image, to obtain a first food recognition result; In response to a preset prompt word input instruction, obtaining a first preset prompt word; The first food identification result and the first preset prompt word are combined to obtain the first target prompt word.

3. The method of claim 1, wherein training the recipe recommendation model comprises: Obtain the trained food recognition sub-model, the projection layer whose parameters are initialized to obey the Gaussian distribution, and the trained large language sub-model; When the parameters of the food recognition sub-model and the parameters of the large language sub-model are fixed, the parameters of the projection layer are adjusted.

4. The method of claim 3, wherein adjusting the parameters of the projection layer while fixing the parameters of the food recognition sub-model and the parameters of the large language sub-model comprises: Using the first data set, while fixing the parameters of the food recognition sub-model and the parameters of the large language sub-model, adjusting the parameters of the projection layer to obtain a recipe recommendation model that completes the first stage of training; Using the second data set, training the recipe recommendation model that has completed the first stage of training, to obtain a recipe recommendation model that has completed the second stage of training; The quantity quality of the second data set is higher than the data quality of the second data set.

5. The method as claimed in claim 4, wherein the first data set is obtained by using a multimodal large language model; and the second data set is obtained by screening data in the first data set.

6. The method of claim 4, wherein the first data set is used to adjust the parameters of the projection layer while fixing the parameters of the ingredient recognition sub-model and the parameters of the large language sub-model to obtain a recipe recommendation model that completes the first stage of training, comprising: Acquire a second preset prompt word, wherein the second preset prompt word is determined by the user based on a prompt word template; By using the first data set and the second preset prompt word, and fixing the parameters of the ingredient recognition sub-model and the large language sub-model, the parameters of the projection layer are adjusted to obtain the recipe recommendation model that has completed the first stage of training.

7. The method of claim 6, wherein the step of training the recipe recommendation model that has completed the first-stage training using the second data set to obtain the recipe recommendation model that has completed the second-stage training comprises: acquiring a second target prompt word based on the second data set, wherein the second target prompt word is obtained by combining a second food recognition result and a second preset prompt word, the second food recognition result is obtained by using the food recognition sub-model to recognize a second food image in the second data set, and the second preset prompt word is determined by a user based on a prompt word template; The recipe recommendation model that has completed the first stage of training is trained using the second data set and the second target prompt word to obtain a recipe recommendation model that has completed the second stage of training.

8. The method according to any one of claims 1 to 7, wherein inputting the first ingredient image and the first target prompt word into a recipe recommendation model to obtain a recommended recipe comprises: When the smart refrigerator is not connected to the Internet, the first ingredient image and the first target prompt word are input into a recipe recommendation model stored in the smart refrigerator to obtain the recommended recipe.

9. The method according to any one of claims 1 to 7, further comprising: Determining that there is a target ingredient in the recommended recipe that is not stored in the smart refrigerator; Output prompt information to prompt the user to replenish the target food.

10. A smart refrigerator, comprising a camera, a display, a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the camera is used to capture an image of a first food stored in the smart refrigerator; The processor is used to acquire the first food image and obtain the first target prompt word corresponding to each first food; The first ingredient image and the first target prompt word are input into a recipe recommendation model to obtain a recommended recipe, wherein the recipe recommendation model includes an ingredient recognition sub-model, a projection layer and a large language sub-model, the ingredient recognition sub-model is used to encode the first ingredient image to obtain a corresponding ingredient code, the projection layer is used to convert the ingredient code into a feature vector, and the large language sub-model is used to determine the recommended recipe based on the feature vector and the first target prompt word; the display is used to display the recommended recipe.