Recipe recommendation system based on transfer learning fine-tuning multi-modal large model

Through a multimodal large model fine-tuned based on transfer learning, the deficiencies in image recognition accuracy, personalized recommendations, and multimodal data processing in the dietary health analysis system were resolved, achieving high-precision, personalized dietary health management and improving the intelligence and adaptability of the system.

CN120809083APending Publication Date: 2025-10-17SHANGHAI UNIV OF MEDICINE & HEALTH SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510926173.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing dietary health analysis system suffers from insufficient image recognition accuracy, lack of personalized analysis, insufficient multimodal data processing capabilities, and insufficient transfer learning fine-tuning, resulting in low recognition accuracy and recommendations that do not meet user needs.

Method used

A multimodal large model based on transfer learning fine-tuning is used to generate personalized nutritional recipes through data collection, storage, analysis and recipe recommendation modules, combined with image feature extraction and multimodal large models. LoRA is used to fine-tune the model to improve its adaptability.

Benefits of technology

It achieves high-precision dietary health analysis, can accurately identify dishes in complex backgrounds, provide personalized nutritional recommendations, improve the intelligence and adaptability of the system, and meet the health management needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809083A_ABST
    Figure CN120809083A_ABST
Patent Text Reader

Abstract

The invention discloses a recipe recommendation system based on a transfer learning fine-tuning multi-modal large model, and relates to the technical field of diet health analysis. In the system, a data acquisition module is used for acquiring meal multi-modal data of a user at the current moment; the meal multi-modal data of the user at the current moment are respectively sent to the data storage module and the analysis module; an image feature extraction algorithm is deployed in the data storage module, and the data storage module is used for classifying and storing the meal multi-modal data by using the image feature extraction algorithm; a multi-modal large model is deployed in the analysis module, and the analysis module is used for determining the nutrition intake condition of the user by using the multi-modal large model, the meal multi-modal data of the user at the current moment and the meal multi-modal data of the user at each historical moment in a preset historical period before the current moment; and the recipe recommendation module is used for generating a nutrition recipe based on the nutrition intake condition and dynamically presenting the nutrition recipe. According to the invention, personalized recommendation of user recipes is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of dietary health analysis, in particular to a recipe recommendation system based on fine-tuning of a multi-modal large model through transfer learning. BACKGROUND

[0002] With the improvement of people's living standards and the increasing awareness of health, dietary health has gradually evolved into a key issue of concern in modern society. Reasonable diet not only provides adequate nutrition, but also effectively prevents and improves various chronic diseases such as obesity, diabetes, hypertension, etc. Traditional dietary health management mainly relies on manual recording and nutritionist analysis, which has problems such as difficulty in data collection, low analysis accuracy, poor user participation, etc., and cannot achieve personalized and real-time health management. Therefore, how to provide convenient, efficient and accurate dietary health analysis services through intelligent means has become a technical problem to be solved by the current scientific research and industry.

[0003] With the increasing demand for health management, in-depth analysis and scientific management of dietary health have become a core issue to be tackled in modern society. However, the current dietary health analysis system still has certain limitations in the technical aspect, as follows:

[0004] 1. Single image recognition model precision is insufficient: Current dietary health analysis systems mostly rely on image recognition models to identify and analyze user's dining pictures. Due to the large difference in appearance characteristics of different dishes, a single image recognition model often has difficulty in accurately distinguishing a variety of dishes, especially in complex backgrounds (such as when multiple dishes, tableware or interference are placed on the dining table), resulting in low recognition accuracy and classification errors. There are still great challenges for dish recognition in diversified and complex backgrounds. 2. Lack of personalized dietary analysis: Current dietary health analysis systems mostly use traditional nutrition assessment models, lacking analysis and response to user's personalized needs. Each user has different dietary habits, health conditions and preferences, while current systems mostly rely on fixed rules or standardized models for dietary recommendations. This approach not only fails to effectively reflect the user's personalized needs, but also greatly increases the risk of recommending diets that do not meet the actual situation, affecting the accuracy of the system and the user's reliance. 3. Insufficient multi-modal data processing capability: Dietary health analysis systems usually rely on image, text, sound and other multi-modal data for comprehensive analysis. However, current systems mostly use independent processing methods when dealing with these multi-modal data, failing to fully utilize the complementarity between modalities. For example, image data can provide visual information of dishes, while text data (such as food names, food material information) can provide more detailed nutrient composition analysis, but current systems fail to effectively integrate these information, resulting in insufficient depth and accuracy of data analysis. Therefore, how to improve the comprehensive performance of the system under the synergistic effect of multi-modal data is an important problem in current technology. 4. Insufficient transfer learning fine-tuning, poor model self-adaptation ability: Many current dietary health analysis systems rely on models trained from scratch, which often require a large amount of labeled data and computing resources, and the training time is relatively long. Although deep learning technology has made significant progress in recent years, current systems mostly fail to fully utilize the advantages of transfer learning. However, current dietary health analysis systems still have deficiencies in the application of transfer learning, failing to effectively utilize deep learning models trained on large-scale datasets for fine-tuning, resulting in insufficient intelligence and self-adaptation ability of the system, and failing to meet the individual needs of different users. SUMMARY

[0005] The purpose of the present application is to provide a recipe recommendation system based on transfer learning fine-tuning of multi-modal large models to solve the problem of failing to meet user's personalized recipe recommendation.

[0006] To achieve the above-mentioned purpose, the present application provides the following solutions:

[0007] The application provides a recipe recommendation system based on a multi-modal large model fine-tuned by transfer learning, comprising a data acquisition module, a data storage module, an analysis module and a recipe recommendation module; the data acquisition module, the data storage module and the recipe recommendation module are connected with the analysis module, and the data acquisition module is connected with the data storage module;

[0008] The data acquisition module is used for acquiring multi-modal meal data of the current moment of the user, and sending the multi-modal meal data of the current moment of the user to the data storage module and the analysis module respectively; the multi-modal meal data comprises meal pictures and meal text data;

[0009] The image feature extraction algorithm is deployed in the data storage module, and the data storage module is used for classifying and storing the multi-modal meal data by using the image feature extraction algorithm;

[0010] The multi-modal large model is deployed in the analysis module, and the analysis module is used for determining the nutritional intake of the user by using the multi-modal large model, the multi-modal meal data of the current moment of the user and the multi-modal meal data of each historical moment in a preset historical period before the current moment;

[0011] The recipe recommendation module is used for generating a nutritional recipe based on the nutritional intake and dynamically presenting the nutritional recipe.

[0012] Optionally, the recipe recommendation system based on the multi-modal large model fine-tuned by transfer learning further comprises an associated information library calling module; the associated information library calling module is connected with the data storage module and the analysis module respectively;

[0013] The associated information library calling module is used for calling the multi-modal meal data of each historical moment in a preset historical period before the current moment of the user from the data storage module, and sending the multi-modal meal data of each historical moment to the analysis module.

[0014] Optionally, the data acquisition module comprises an image acquisition unit and a text acquisition unit;

[0015] The image acquisition unit is used for acquiring meal pictures of the current moment of the user;

[0016] The text acquisition unit is used for acquiring meal text data of the current moment of the user.

[0017] Optionally, the data acquisition module further comprises a data transmission unit; the data transmission unit is connected with the image acquisition unit, the text acquisition unit, the data storage module and the analysis module respectively;

[0018] The data transmission unit is configured to transmit the meal multi-modal data of the current time of the user to the data storage module and the analysis module, respectively.

[0019] Optionally, the data storage module comprises a classification unit and a cloud storage, and the classification unit is connected with the cloud storage and the association information library calling module, respectively.

[0020] The image feature extraction algorithm is deployed in the classification unit, and the classification unit is configured to classify the meal multi-modal data by using the image feature extraction algorithm.

[0021] The cloud storage is configured to store the classified meal multi-modal data.

[0022] Optionally, the analysis module comprises an image feature extraction unit, a classification and recognition unit and a nutrition calculation unit connected in sequence, the image feature extraction unit is connected with the data acquisition module and the data storage module, respectively, and the nutrition calculation unit is connected with the recipe recommendation module.

[0023] The multi-modal large model is deployed in the image feature extraction unit, and the image feature extraction unit is configured to extract features of the meal pictures of the current time of the user and the meal pictures of each historical time in a preset historical period before the current time by using the multi-modal large model, to obtain feature vectors of the meal pictures.

[0024] The classification and recognition unit is configured to classify and recognize the types and ingredients of the meals based on the feature vectors of the meal pictures, to obtain recognition results.

[0025] The nutrition calculation unit is configured to calculate the nutrition intake of the user based on the recognition results, the multi-modal large model, the meal text data of the current time of the user and the meal text data of each historical time in a preset historical period before the current time.

[0026] Optionally, the image feature extraction unit comprises an image preprocessing sub-unit and a feature extraction sub-unit, the image preprocessing sub-unit is connected with the data acquisition module, the data storage module and the feature extraction sub-unit, respectively, and the feature extraction sub-unit is connected with the classification and recognition unit.

[0027] The image preprocessing sub-unit is configured to preprocess the meal pictures of the current time of the user and the meal pictures of each historical time in a preset historical period before the current time, to obtain corresponding preprocessed meal pictures, and the preprocessing comprises denoising, cropping, size adjustment and color optimization.

[0028] The multi-modal large model is deployed in the feature extraction subunit, and the feature extraction subunit is configured to extract features of each pre-processed meal picture by using the multi-modal large model to obtain a feature vector corresponding to each meal picture.

[0029] Optionally, the multi-modal large model is an open-source multi-modal large model trained by using LoRA for transfer learning.

[0030] Optionally, the recipe recommendation module comprises a nutritional recipe generation unit and a display unit, and the nutritional recipe generation unit is connected with the analysis module and the display unit, respectively.

[0031] The nutritional recipe generation unit is configured to generate a nutritional recipe based on the nutritional intake condition.

[0032] The display unit is configured to dynamically present the nutritional recipe.

[0033] Optionally, the recipe recommendation module further comprises a recipe adjustment unit, and the recipe adjustment unit is connected with the nutritional recipe generation unit and the display unit, respectively.

[0034] The recipe adjustment unit is configured to collect dietary preferences, allergy information and regional dietary habits of a user, and adjust the nutritional recipe generated by the nutritional recipe generation unit by using the dietary preferences, the allergy information and the regional dietary habits to obtain an adjusted nutritional recipe.

[0035] According to the specific embodiments provided in the present application, the following technical effects are disclosed:

[0036] The application discloses a recipe recommendation system based on a multi-modal large model fine-tuned by transfer learning, a data acquisition module is used to acquire multi-modal meal data of a user at a current time, and the multi-modal meal data of the user at the current time is sent to a data storage module and an analysis module respectively; an image feature extraction algorithm is deployed in the data storage module, and the data storage module is used to classify and store the multi-modal meal data by using the image feature extraction algorithm; a multi-modal large model is deployed in the analysis module, and the analysis module is used to determine a nutrition intake condition of the user by using the multi-modal large model and the multi-modal meal data of the user at the current time and multi-modal meal data of each historical time in a preset historical period before the current time; and a recipe recommendation module is used to generate a nutrition recipe based on the nutrition intake condition, and dynamically present the nutrition recipe. The system is systematically improved from aspects of multi-modal data fusion, personalized nutrition analysis and transfer learning fine-tuning. The multi-modal large model obtained by the knowledge transfer of a large-scale pre-training model and the innovative fine-tuning strategy realizes high precision, diversification and dynamic self-adaptation of diet health analysis, thereby meeting the personalized health management needs of the user, and solving the limitations in recognition accuracy, personalized recommendation, multi-modal fusion and model self-adaptation. Not only the accuracy of image recognition can be improved, but also more intelligent health diet evaluation and personalized recommendation can be realized. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0038] Figure 1 The structure schematic diagram of the recipe recommendation system based on the multi-modal large model fine-tuned by transfer learning is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0039] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0040] The purpose of the present application is to provide a recipe recommendation system based on a multi-modal large model fine-tuned by transfer learning, which aims to realize personalized recipe recommendation for users.

[0041] In order to make the above objectives, features and advantages of the present application more apparent, more comprehensible, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] In an exemplary embodiment, as shown in Figure 1 a recipe recommendation system based on transfer learning fine-tuning of a multi-modal large model is provided, comprising: a data acquisition module, a data storage module, an analysis module and a recipe recommendation module; the data acquisition module, the data storage module and the recipe recommendation module are connected with the analysis module, and the data acquisition module is connected with the data storage module.

[0043] The data acquisition module is used for acquiring the current meal multi-modal data of the user, and sending the current meal multi-modal data of the user to the data storage module and the analysis module respectively; the meal multi-modal data includes meal picture and meal text data.

[0044] The data storage module is deployed with an image feature extraction algorithm, and is used for classifying and storing the meal multi-modal data by using the image feature extraction algorithm.

[0045] The analysis module is deployed with a multi-modal large model, and is used for determining the nutritional intake of the user by using the multi-modal large model and the current meal multi-modal data of the user and the meal multi-modal data of each historical moment in a preset historical period before the current moment.

[0046] The recipe recommendation module is used for generating a nutritional recipe based on the nutritional intake, and dynamically presenting the nutritional recipe.

[0047] As an optional implementation, the recipe recommendation system based on transfer learning fine-tuning of a multi-modal large model further comprises: an associated information library calling module; the associated information library calling module is connected with the data storage module and the analysis module respectively.

[0048] The associated information library calling module is used for calling the meal multi-modal data of each historical moment in a preset historical period before the current moment of the user from the data storage module, and sending the meal multi-modal data of each historical moment to the analysis module.

[0049] Specifically, the associated information library calling module automatically retrieves the meal picture data of the user in the last week from the cloud database, organizes the data by using time sequence, and supports data filtering according to time, category or personalized tags marked by the user, so as to efficiently call target data. The associated information library calling module also has data consistency verification and repeated checking functions to ensure that the called data is complete and error-free. The historical data and the currently uploaded meal picture data are input into the multi-modal large model for comprehensive analysis, so as to facilitate the model to comprehensively evaluate the recent nutritional intake trend and structure of the user.

[0050] In an exemplary embodiment, the functions of the association information base calling module specifically include:

[0051] Automatically retrieve the user's meal picture data in the cloud database in the past week.

[0052] Organize data in time series using efficient retrieval algorithms, support multi-dimensional filtering of time, category or label.

[0053] With data verification and repeated inspection, ensure data consistency and availability.

[0054] As an optional implementation, the data acquisition module includes: image acquisition unit and text acquisition unit.

[0055] The image acquisition unit is used to acquire the user's current meal picture.

[0056] The text acquisition unit is used to acquire the user's current meal text data.

[0057] As an optional implementation, the data acquisition module further includes: data transmission unit; the data transmission unit is connected with the image acquisition unit, the text acquisition unit, the data storage module and the analysis module respectively.

[0058] The data transmission unit is used to send the user's current meal multi-modal data to the data storage module and the analysis module respectively.

[0059] Specifically, the image acquisition unit uses the smartphone camera to take high-resolution pictures of the user's table meal, and is equipped with automatic focusing, color correction and brightness optimization functions to ensure picture clarity and authenticity. The data transmission unit uploads the collected meal pictures to the cloud server through wireless network (such as Wi-Fi or cellular data) encryption. In actual application, the user can choose immediate upload or delayed upload mode, and support batch transmission of multiple pictures to improve the convenience of use.

[0060] In an exemplary embodiment, the functions of the data storage module specifically include:

[0061] The function of classifying and retrieving the user's uploaded pictures based on timestamp.

[0062] Use distributed storage technology to support efficient management and retrieval of large-scale image data.

[0063] Use image feature extraction algorithm to archive meal pictures according to dish category or nutritional component features.

[0064] Integrate data encryption and redundancy backup strategy to ensure data security and stability.

[0065] As an optional implementation, the data storage module comprises a classification unit and a cloud storage; the classification unit is connected with the cloud storage and the association information library calling module respectively.

[0066] The image feature extraction algorithm is deployed in the classification unit, and the classification unit is configured to classify the meal multi-modal data by using the image feature extraction algorithm.

[0067] The cloud storage is configured to store the classified meal multi-modal data.

[0068] Specifically, the data storage module indexes the pictures based on timestamps through the cloud database to realize fast retrieval and classified storage. In order to improve the data management efficiency, the module is built-in with an image feature extraction algorithm, which can intelligently classify and archive according to the dish category or nutritional ingredient characteristics of the meal. In order to protect data security and privacy, the data storage module adopts multi-layer encryption, access permission control strategy, and supports regular data backup and recovery function, so as to realize efficient, stable and safe storage of large-scale image data.

[0069] As an optional implementation, the analysis module comprises an image feature extraction unit, a classification and recognition unit and a nutrition calculation unit connected in sequence; the image feature extraction unit is connected with the data acquisition module and the data storage module respectively, and the nutrition calculation unit is connected with the recipe recommendation module.

[0070] A multi-modal large model is deployed in the image feature extraction unit, and the image feature extraction unit is configured to extract features of the meal pictures at the current moment and the meal pictures at each historical moment in a preset historical period before the current moment by using the multi-modal large model, to obtain feature vectors of the meal pictures.

[0071] The classification and recognition unit is configured to classify and recognize the types and ingredients of the meals based on the feature vectors of the meal pictures, to obtain the recognition result.

[0072] The nutrition calculation unit is configured to calculate the nutrition intake of the user based on the recognition result, the multi-modal large model, the meal text data at the current moment of the user and the meal text data at each historical moment in a preset historical period before the current moment.

[0073] As an optional implementation, the image feature extraction unit comprises an image preprocessing subunit and a feature extraction subunit; the image preprocessing subunit is connected with the data acquisition module, the data storage module and the feature extraction subunit respectively, and the feature extraction subunit is connected with the classification and recognition unit.

[0074] The image preprocessing subunit is configured to preprocess the meal picture of the current moment of the user and the meal pictures of each historical moment in a preset historical period before the current moment to obtain corresponding preprocessed meal pictures. The preprocessing includes denoising, cropping, size adjustment, and color optimization.

[0075] The feature extraction subunit is configured to deploy a multi-modal large model, and the feature extraction subunit is configured to extract features of each preprocessed meal picture by using the multi-modal large model to obtain a feature vector corresponding to each meal picture.

[0076] As an optional implementation, the multi-modal large model is an open-source multi-modal large model that is fine-tuned by using low-rank adaptation (LoRA) for transfer learning.

[0077] Specifically, the training and deployment process of the multi-modal large model includes:

[0078] Step 101: Obtain the Food-101 dataset, check the integrity of the Food-101 dataset, and divide the training set and the test set. At the same time, perform size adjustment, normalization, and data enhancement processing on the pictures in the Food-101 dataset. Then, construct a data loader to support batch loading and data enhancement, and ensure efficient training. Finally, encode the class labels into a format that can be recognized by the model, and prepare to input the model for fine-tuning.

[0079] Step 102: Install the corresponding deep learning framework and dependent libraries. Load the pre-trained model weights and fine-tuned data, set up an efficient data loading pipeline, support distributed training and batch processing. Optimize the running environment, improve the efficiency and stability of model fine-tuning through mixed precision training, memory management, and caching strategies.

[0080] Step 103: Load the pre-trained model and freeze most of its parameters to reduce training resource consumption. Then define the LoRA module, perform trainable adjustment on part of the weights by inserting a low-rank decomposition layer, set the optimizer and learning rate strategy. Then, load the data into the model, perform forward propagation and back propagation, and only update the parameters of the LoRA layer. Finally, save the fine-tuned LoRA weights to facilitate lightweight loading and application of the model in the inference stage.

[0081] Step 104: Load the preprocessed data and the fine-tuned pre-trained model into memory and ensure that the hardware device is running normally. Then enter the training loop, sequentially perform the processes of data batch loading, forward propagation loss calculation, and backward propagation parameter updating, and only optimize the specified fine-tuning layer (such as the LoRA layer). At the same time, record the training indicators such as loss and accuracy in real time, and periodically evaluate the model performance on the validation set. After training is completed, save the fine-tuned model weights to prepare for subsequent inference or deployment.

[0082] Step 105: When evaluating the model on the validation set, set the fine-tuned model to evaluation mode to freeze parameters, and load validation data batch by batch for forward propagation. Calculate the performance indicators of the model on the validation set, such as loss value, accuracy, F1 score or AUC, etc., to measure the generalization ability of the model. Record and analyze the evaluation results to find potential optimization directions or verify whether the model has reached the expected performance.

[0083] In a further embodiment, the LoRA fine-tuning method comprises:

[0084] Freeze most of the pre-trained model parameters, and only introduce low-rank matrices A and B in the QKV projection matrix and the feedforward neural network weight matrix for fine-tuning, to reduce the computational and storage overhead.

[0085] Further, in an exemplary embodiment, the multi-modal large model is designed based on the Transformer architecture; the Transformer model is sequentially composed of an input layer, an encoding layer, and a decoding layer. The function of the input layer is to convert the text question into word units and generate the corresponding word sequence. The encoding layer contains multi-head self-attention mechanism, feedforward neural network, and residual connection and layer normalization, which are connected in turn. The role of multi-head self-attention mechanism is to explore the internal relationship of word sequence, focusing on capturing key information in the text. The feedforward neural network then performs nonlinear mapping on the output of the attention mechanism, thereby enhancing the feature expression ability of the model and processing complex patterns. And residual connection and layer normalization are used to optimize the training process of query vector, key vector and value vector, to extract context information. The decoding layer generates an answer to the input question by combining the context information and key features. Based on the above architecture design, the multi-modal large model can efficiently process data of multiple modalities and make personalized recipe recommendations combined with the user's personalized information, so it is suitable for health management and diet recommendation systems. The following is the specific process of the multi-modal large model in actual operation.

[0086] Step 201: The user uploads personal health information to establish a personalized health profile. Health information includes basic physical indicators (such as weight, height, blood pressure, etc.) and possible health risk assessment data. These information will be used for model initialization and personalized analysis.

[0087] Step 202: The user takes a picture of today's meal and uploads the picture data. The system pre-processes the uploaded picture (including size adjustment, image enhancement) and extracts features, so as to extract features related to food types and intake, providing a basis for subsequent analysis.

[0088] Step 203: The user makes a simple text description of the meal for the day and uploads the text data. The text data includes dish name, ingredient composition, ingredient list, dish origin, dietary preference, food name, portion size, and other supplementary descriptions (such as cooking method). The system converts the text data into structured input that can be recognized by the multi-modal large model, and performs multi-modal alignment processing with the image features.

[0089] Step 204: The multi-modal large model analyzes and makes decisions based on the multi-modal data uploaded by the user, health records, and historical data. The multi-modal large model uses deep learning methods to evaluate dietary patterns and nutritional structure, and gives optimization suggestions based on user health information.

[0090] Step 205: The multi-modal large model makes analysis and dietary health recommendations and feeds back to the user. The system generates personalized dietary recommendations to help users balance nutrient intake and improve health status. At the same time, it supports visual analysis reports to help users understand the scientific basis of the recommendations.

[0091] Through the above steps, the multi-modal large model organically integrates visual features (such as the features of the meal picture) and text features (such as the user's text description of the meal) to accurately identify the type of meal and its main nutritional components. Finally, the multi-modal large model outputs meal classification labels, nutritional component estimates, and calorie information to provide personalized dietary recommendations for users

[0092] Specifically, the analysis module first preprocesses the input meal picture (such as denoising, cropping, size adjustment, and color optimization), then extracts the deep feature information of the meal through the LoRA fine-tuned multi-modal large model, and applies classification and recognition algorithms to accurately classify the meal and identify its nutritional components. The nutrition calculation unit generates a nutrition report including macro and micronutrient intake levels for the user based on the identified components and the nutrition database. This nutrition report can reveal the user's nutritional intake structure over a period of time, marking missing and excessive nutrients, and providing data support for healthy diet management.

[0093] As an optional implementation, the recipe recommendation module includes: a nutritional recipe generation unit and a display unit; the nutritional recipe generation unit is connected with the analysis module and the display unit respectively.

[0094] The nutritional recipe generation unit is used to generate nutritional recipes based on the nutritional intake situation.

[0095] The display unit is used to dynamically present the nutritional recipes.

[0096] As an optional implementation, the recipe recommendation module further includes: a recipe adjustment unit; the recipe adjustment unit is connected with the nutritional recipe generation unit and the display unit respectively.

[0097] The recipe adjustment unit is configured to collect dietary preferences, allergy information and regional dietary habits of the user, and adjust the nutritional recipe generated by the nutritional recipe generation unit according to the dietary preferences, allergy information and regional dietary habits, to obtain an adjusted nutritional recipe.

[0098] Specifically, after the analysis module completes the nutrition assessment, the recipe recommendation module automatically matches the food materials with the recipe database according to the nutrition intake report of the user. The recipe recommendation module intelligently filters suitable dishes and corresponding cooking steps in combination with the dietary preferences, allergy information and regional dietary habits of the user, and recommends nutritionally balanced recipes for breakfast, lunch and dinner, respectively. The recommended content includes the dish name, food material list, nutritional composition and calorie information for each meal. The user can view and download the recommended recipes and corresponding shopping lists in an intuitive graphic and text manner through the mini-program client, to realize convenient healthy diet planning.

[0099] In further embodiments, the functions of the recipe recommendation module specifically include:

[0100] According to the nutrition intake report, the food materials are intelligently matched with the recipe database, and nutritionally balanced recipes for breakfast, lunch and dinner are formulated.

[0101] The dietary preferences, allergy information and regional habits of the user are comprehensively considered to ensure that the recommended recipes are both nutritionally rich and meet the user's taste.

[0102] Detailed cooking steps, nutritional composition and calorie information are provided, and are presented in a graphic and text manner through the mini-program client.

[0103] Support is provided for generating a shopping list for subsequent purchase and cooking by the user.

[0104] Through the above embodiments, the present application uses a multi-modal large model based on transfer learning fine-tuning (LoRA method) to comprehensively analyze the meal pictures uploaded by the user and the historical data, thereby realizing accurate nutrition intake assessment and providing personalized and nutritionally balanced recipe recommendations. Compared with traditional methods, the present application only needs to fine-tune a small amount of parameters, which not only maintains the extensive knowledge of the pre-trained large model, but also significantly reduces the calculation and resource costs. At the same time, the present application adopts efficient data management, safe storage, flexible calling and personalized recommendation strategies, so that the system has excellent performance and scalability in actual application.

[0105] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0106] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, it should be understood that the application encompasses all possible combinations of the technical features unless such a combination is not technically possible.

[0107] The principles and implementation manners of the present application are described herein by using specific examples, and the above embodiments are only used to help understand the system of the present application and its core idea; meanwhile, according to the idea of the present application, a person skilled in the art will make changes in specific implementation manners and application scopes. In conclusion, the content of the present application should not be understood as a limitation.

Claims

1. A recipe recommendation system based on transfer learning and fine-tuning a multimodal large model, characterized by: The recipe recommendation system based on transfer learning fine-tuning multimodal large model includes: a data acquisition module, a data storage module, an analysis module and a recipe recommendation module; the data acquisition module, the data storage module and the recipe recommendation module are all connected to the analysis module, and the data acquisition module is connected to the data storage module; The data acquisition module is used to collect the user's current meal multimodal data and send the user's current meal multimodal data to the data storage module and the analysis module respectively; the meal multimodal data includes: meal pictures and meal text data; An image feature extraction algorithm is deployed in the data storage module, and the data storage module is used to classify and store meal multimodal data using the image feature extraction algorithm; The analysis module is configured to utilize a multimodal large model and the multimodal data of the user's current meal and the multimodal data of each historical moment in a preset historical period before the current moment to determine the user's nutritional intake; The recipe recommendation module is used to generate nutritional recipes based on the nutritional intake situation and dynamically present the nutritional recipes.

2. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 1 is characterized in that: The recipe recommendation system based on transfer learning fine-tuning multimodal large model further includes: an associated information library calling module; the associated information library calling module is connected to the data storage module and the analysis module respectively; The associated information library calling module is used to call the meal multimodal data of each historical moment in a preset historical period before the current moment of the user from the data storage module, and send the meal multimodal data of each historical moment to the analysis module.

3. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 2 is characterized in that: The data acquisition module includes: an image acquisition unit and a text acquisition unit; The image acquisition unit is used to acquire the user's meal picture at the current moment; The text collection unit is used to collect the user's meal text data at the current moment.

4. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 3 is characterized in that: The data acquisition module further includes: a data transmission unit; the data transmission unit is connected to the image acquisition unit, the text acquisition unit, the data storage module and the analysis module respectively; The data transmission unit is used to send the user's current meal multimodal data to the data storage module and the analysis module respectively.

5. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 4 is characterized in that: The data storage module includes: a classification unit and a cloud storage; the classification unit is connected to the cloud storage and the associated information library calling module respectively; The image feature extraction algorithm is deployed in the classification unit, and the classification unit is used to classify the meal multimodal data using the image feature extraction algorithm; The cloud storage is used to store the classified meal multimodal data.

6. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 1 is characterized in that: The analysis module includes: an image feature extraction unit, a classification and recognition unit, and a nutrition calculation unit connected in sequence; the image feature extraction unit is connected to the data acquisition module and the data storage module respectively, and the nutrition calculation unit is connected to the recipe recommendation module; The multimodal large model is deployed in the image feature extraction unit, and the image feature extraction unit is used to use the multimodal large model to extract features from the user's meal pictures at the current moment and meal pictures at each historical moment in a preset historical period before the current moment, to obtain a feature vector for each meal picture; The classification and recognition unit is used to classify and recognize the type and ingredients of each meal based on the feature vector of each meal image to obtain a recognition result; The nutrition calculation unit is used to calculate the user's nutritional intake based on the recognition result, the multimodal large model, the user's current meal text data, and the meal text data of each historical moment in a preset historical period before the current moment.

7. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 6 is characterized in that: The image feature extraction unit includes: an image preprocessing subunit and a feature extraction subunit; the image preprocessing subunit is connected to the data acquisition module, the data storage module and the feature extraction subunit respectively, and the feature extraction subunit is connected to the classification and recognition unit; The image preprocessing subunit is used to preprocess the user's meal picture at the current moment and meal pictures at each historical moment in a preset historical period before the current moment to obtain corresponding preprocessed meal pictures; the preprocessing includes: denoising, cropping, resizing and color optimization; The multimodal large model is deployed in the feature extraction subunit, and the feature extraction subunit is used to use the multimodal large model to extract features from each preprocessed meal image to obtain a feature vector corresponding to each meal image.

8. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 7 is characterized in that: The multimodal large model is an open source multimodal large model that has been fine-tuned and trained using LoRA for transfer learning.

9. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 1 is characterized in that: The recipe recommendation module includes: a nutrition recipe generation unit and a display unit; the nutrition recipe generation unit is connected to the analysis module and the display unit respectively; The nutritional recipe generating unit is used to generate a nutritional recipe based on the nutritional intake situation; The display unit is used to dynamically present nutritional recipes.

10. The recipe recommendation system based on transfer learning fine-tuning multimodal large model according to claim 9 is characterized in that: The recipe recommendation module further includes: a recipe adjustment unit; the recipe adjustment unit is connected to the nutrition recipe generation unit and the display unit respectively; The recipe adjustment unit is used to collect the user's dietary preferences, allergy information and regional dietary habits, and use the dietary preferences, allergy information and regional dietary habits to adjust the nutritional recipe generated by the nutritional recipe generation unit to obtain an adjusted nutritional recipe.