Intelligent food nutrient content evaluation method and system based on multi-modal large model, terminal and storage medium

Through multimodal large models and deep learning algorithms, combined with adaptive feature alignment and cross-modal attention mechanisms, the model parameters are dynamically updated, which solves the problem of insufficient adaptability to new food types and achieves high-precision and personalized food nutrition assessment.

CN120613031AActive Publication Date: 2025-09-09THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511114731.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-09
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing AI-based intelligent food nutrition assessment methods have deficiencies in adaptability to new food types and the depth of multimodal data fusion, resulting in poor accuracy and comprehensiveness of assessment results, making it difficult to meet the needs of large-scale, real-time, and personalized nutrition assessment.

Method used

A large multimodal model combined with an adaptive feature alignment algorithm and a cross-modal attention mechanism is used for feature extraction, alignment, and fusion to build an initial deep learning model. The model parameters are updated through incremental learning and memory enhancement mechanisms to dynamically learn knowledge about new food types and output nutritional content assessment results.

Benefits of technology

It significantly improves the accuracy and robustness of food nutritional assessment, can adapt to the emergence of new food types while maintaining high-precision assessment of old food types, provide personalized nutritional recommendations, and meet real-time and large-scale data processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613031A_ABST
    Figure CN120613031A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal data processing, and discloses a multi-modal large model-based food nutritional ingredient content intelligent evaluation method and system, a terminal and a storage medium, and the method comprises the following steps: obtaining and preprocessing multi-modal data to obtain target multi-modal data; performing feature extraction, feature alignment and feature fusion on the target multi-modal data by adopting a pre-trained multi-modal large model to obtain multi-modal fusion features; when a new food type appears, updating parameters of the intermediate deep learning model through an incremental learning and memory enhancement mechanism to obtain a target deep learning model; and inputting the multi-modal fusion features into the target deep learning model, and outputting a nutritional ingredient content evaluation result. Through the multi-modal large model and the continuous learning algorithm, high-precision dynamic evaluation of food nutritional ingredients is achieved, the method can adapt to appearance of new food types, and high-precision evaluation of old food types can be kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing, and in particular to a method, system, terminal and computer-readable storage medium for intelligently evaluating the nutritional content of food based on a multimodal large model. Background Art

[0002] As global attention to healthy eating continues to grow, food nutrition assessment has become an important part of public health. Traditional food nutrition assessment methods rely on manual records, laboratory testing, or preset food databases. These methods are not only time-consuming and labor-intensive, but also difficult to adapt to the rapid changes in food types and individualized needs. In recent years, with the development of artificial intelligence technology, especially the rise of multimodal large models, intelligent assessment systems based on images, text, and voice have gradually become a research hotspot. Existing technologies have already produced a variety of food nutrition intelligent assessment systems that are used to intelligently assess the content of the main nutrients in food and automatically generate personalized plans based on the customer's dietary needs, integrating nutrition assessment, planning, storage, and sharing functions.

[0003] However, the existing intelligent food nutrition assessment methods based on artificial intelligence technology still have shortcomings in adaptability to new food types and the depth of multimodal data fusion, resulting in poor accuracy and comprehensiveness of food nutrition assessment results, and it is difficult to meet the needs of large-scale, real-time and personalized nutrition assessment.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, system, terminal and computer-readable storage medium for intelligent evaluation of food nutrient content based on a multimodal large model, aiming to solve the problem that existing intelligent food nutrition evaluation methods still have deficiencies in adaptability to new food types and depth of multimodal data fusion, resulting in poor accuracy and comprehensiveness of evaluation results.

[0006] To achieve the above-mentioned object of the invention, the present invention provides a method for intelligently evaluating the nutritional content of food based on a multimodal large model. The method comprises: Acquiring multimodal data and preprocessing the multimodal data to obtain target multimodal data; Using a pre-trained multimodal large model and combining it with an adaptive feature alignment algorithm and a cross-modal attention mechanism, feature extraction, feature alignment, and feature fusion are performed on the target multimodal data to obtain multimodal fusion features; Building an initial deep learning model using a deep learning algorithm, and training the initial deep learning model to obtain an intermediate deep learning model; When a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned new knowledge corresponding to the new food type; The multimodal fusion features are input into the target deep learning model, and the target deep learning model outputs a nutrient content assessment result.

[0007] Optionally, the acquiring multimodal data and preprocessing the multimodal data to obtain target multimodal data specifically includes: Acquiring multimodal data of the food to be evaluated collected by a smart device, wherein the multimodal data includes images, voice, and text; Performing denoising, enhancement and segmentation processing on the image to obtain a target image; Recognizing the speech using speech recognition technology to obtain recognized text; Cleaning, word segmentation, and format alignment are performed on the text and the recognized text to obtain a target text; The target multimodal data includes the target image and the target text.

[0008] Optionally, the pre-trained multimodal large model is used in combination with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fusion features, specifically including: Using an image encoder in a pre-trained multimodal large model to extract features from the target image to obtain image depth features; Using a text encoder in a pre-trained multimodal large model to extract features from the target text to obtain text semantic features; An adaptive feature alignment algorithm is used to map the features of different modalities to a unified feature space to obtain aligned features of different modalities, where the features of different modalities include image depth features and text semantic features: ; in, Indicates the aligned The characteristics of the modality, represents the feature alignment function, express The characteristics of the modality, Represents an image, Represents text; The cross-modal attention mechanism is used to calculate the importance weight coefficients corresponding to different modalities, and the aligned features of different modalities are weightedly fused based on the importance weight coefficients to obtain the multimodal fusion features of the food to be evaluated: ; in, represents multimodal fusion features, express Importance weight coefficient corresponding to the mode.

[0009] Optionally, the adopting a deep learning algorithm to construct an initial deep learning model, and training the initial deep learning model to obtain an intermediate deep learning model specifically includes: Constructing an initial deep learning model using a deep learning algorithm, wherein the initial deep learning model includes an input layer, a feature processing layer, and an output layer; Acquire historical multimodal data and historical labels corresponding to the historical multimodal data, and perform preprocessing, feature extraction, feature alignment, and feature fusion on the historical multimodal data to obtain historical multimodal fusion features; Training the initial deep learning model using the historical multimodal fusion features and the historical labels to obtain an intermediate deep learning model; The historical multimodal data includes old multimodal data corresponding to old food types.

[0010] Optionally, when a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned new knowledge corresponding to the new food type, specifically including: When a new food type appears, new multimodal data corresponding to the new food type and a new label corresponding to the new multimodal data are obtained, and the new multimodal data are preprocessed, feature extracted, aligned, and fused to obtain new multimodal fusion features; Based on the new multimodal fusion features and the new labels, the original parameters of the intermediate deep learning model are updated through incremental learning technology to obtain new parameters of the intermediate deep learning model, so as to continuously learn new knowledge corresponding to the new food type: ; in, represents the new parameters of the intermediate deep learning model, represents the original parameters of the intermediate deep learning model, represents the learning rate, represents the gradient of the loss function with respect to the model parameters, represents the loss function, represents the new multimodal fusion feature, Indicates a new tag; Based on the new parameters and the original parameters of the intermediate deep learning model, the new parameters of the intermediate deep learning model are updated by memory enhancement technology to obtain a target deep learning model, so as to periodically review the old knowledge corresponding to the old food type: ; in, represents the parameters of the target deep learning model, represents the regularization coefficient, represents the gradient of the regularization term with respect to the model parameters, represents the regularization term.

[0011] Optionally, inputting the multimodal fusion features into the target deep learning model, and the target deep learning model outputting a nutrient content assessment result, specifically includes: Inputting the multimodal fusion features of the food to be evaluated into the input layer of the target deep learning model, and the input layer transmits the multimodal fusion features to the feature processing layer of the target deep learning model; The feature processing layer processes the multimodal fusion features and outputs the processing results to the output layer of the target deep learning model; The output layer maps the processing result into the nutritional content of the food and outputs the nutritional content evaluation result of the food to be evaluated; Obtaining user health data, and using a collaborative filtering algorithm or a content-based recommendation algorithm to obtain user-customized nutritional advice and health risk warnings based on the user health data and the nutrient content assessment results; The nutrient content assessment results, the nutritional recommendations, and the health risk warnings are visualized through interactive charts and graphical interfaces.

[0012] Optionally, the method for intelligently evaluating the nutritional content of food based on a multimodal large model further includes: Receive user feedback and use the user feedback as a reference factor to update the parameters of the intermediate deep learning model through the incremental learning and memory enhancement mechanism.

[0013] To achieve the above-mentioned object of the invention, the present invention further provides a food nutrient content intelligent assessment system based on a multimodal large model, the food nutrient content intelligent assessment system based on a multimodal large model comprising: Data acquisition and preprocessing module: used to acquire multimodal data and preprocess the multimodal data to obtain target multimodal data; Multimodal large model fusion module: used to use a pre-trained multimodal large model and combine it with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fusion features; Deep learning model construction module: used to construct an initial deep learning model using a deep learning algorithm, and train the initial deep learning model to obtain an intermediate deep learning model; Continuous learning module: used to update the parameters of the intermediate deep learning model through incremental learning and memory enhancement mechanism when a new food type appears, so as to obtain the target deep learning model that has learned the new knowledge corresponding to the new food type; Nutrient content assessment module: used to input the multimodal fusion features into the target deep learning model, and the target deep learning model outputs the nutrient content assessment result.

[0014] To achieve the above-mentioned purpose of the invention, the present invention also provides a terminal, which includes: a memory, a processor, and an intelligent evaluation program for the nutritional content of food based on a multimodal large model stored in the memory and runnable on the processor. When the intelligent evaluation program for the nutritional content of food based on a multimodal large model is executed by the processor, the steps of the intelligent evaluation method for the nutritional content of food based on a multimodal large model as described above are implemented.

[0015] To achieve the above-mentioned purpose of the invention, the present invention also provides a computer-readable storage medium, which stores a program for intelligently evaluating the nutrient content of food based on a multimodal large model. When the program for intelligently evaluating the nutrient content of food based on a multimodal large model is executed by a processor, the steps of the method for intelligently evaluating the nutrient content of food based on a multimodal large model as described above are implemented.

[0016] In the present invention, multimodal data is acquired and preprocessed to obtain target multimodal data; a pre-trained multimodal large model is used in combination with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment and feature fusion on the target multimodal data to obtain multimodal fusion features; a deep learning algorithm is used to construct an initial deep learning model, and the initial deep learning model is trained to obtain an intermediate deep learning model; when a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned the new knowledge corresponding to the new food type; the multimodal fusion features are input into the target deep learning model, and the target deep learning model outputs a nutrient content assessment result. The present invention deeply integrates multiple modal data such as images, text, and voice through a multimodal large model, significantly improving the precision and robustness of food nutrition assessment and the accuracy of food nutrition assessment results. Through the continuous learning algorithm, it can dynamically learn new knowledge about new food types without forgetting old knowledge. It can not only adapt to the emergence of new food types, but also maintain high-precision assessment of old food types, which solves the limitations of existing methods when facing new food types and improves the comprehensiveness of food nutrition assessment results. Based on user health data and food nutrition assessment results, it provides users with customized nutrition recommendations to meet the personalized health needs of different users. It is real-time and convenient, and can efficiently and quickly process large-scale data and complete food nutrition assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of a preferred embodiment of the method for intelligently evaluating the nutritional content of food based on a multimodal large model of the present invention; Figure 2 This is a structural diagram of a preferred embodiment of the intelligent evaluation system for food nutrient content based on a multimodal large model of the present invention; Figure 3 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0019] With growing global attention to healthy eating, food nutrition assessment has become a crucial component of public health. Traditional food nutrition assessment methods rely on manual record-keeping, laboratory testing, or pre-defined food databases. These methods are not only time-consuming and labor-intensive, but also struggle to adapt to the rapidly changing food types and individualized needs. In recent years, with the advancement of artificial intelligence (AI), particularly the rise of large multimodal models, intelligent assessment systems based on images, text, and voice have become a research hotspot. A variety of intelligent food nutrition assessment systems have emerged. For example, SmartPlate is an AI-powered meal planning platform that intelligently assesses the content of key nutrients in foods and automatically generates personalized plans based on the dietary needs of the customer, integrating nutritional assessment, planning, storage, and sharing. However, existing AI-based intelligent food nutrition assessment methods still lack adaptability to new food types and the depth of multimodal data integration. This results in poor accuracy and comprehensiveness in food nutrition assessment results, and makes it difficult to meet the needs of large-scale, real-time, and personalized nutrition assessment.

[0020] The application of artificial intelligence (AI) technology has made significant progress in nutritional assessment. For example, using image recognition technology, systems can automatically identify food images and estimate their nutritional content, simplifying dietary assessment methods. Furthermore, some systems not only provide personalized nutritional recommendations but also incorporate smart shopping lists and cooking guides, further enhancing the user experience. While these applications demonstrate the enormous potential of AI in nutrition and health, existing AI-based nutritional assessment methods still have limitations in terms of multimodal data fusion and continuous learning capabilities.

[0021] In order to solve the above technical problems, the present invention provides an intelligent evaluation method for the nutritional content of food based on a multimodal large model, which obtains multimodal data and preprocesses the multimodal data to obtain target multimodal data; uses a pre-trained multimodal large model and combines it with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment and feature fusion on the target multimodal data to obtain multimodal fusion features; uses a deep learning algorithm to construct an initial deep learning model, and trains the initial deep learning model to obtain an intermediate deep learning model; when a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned the new knowledge corresponding to the new food type; inputs the multimodal fusion features into the target deep learning model, and the target deep learning model outputs the nutritional content evaluation result. The present invention deeply integrates multiple modal data such as images, text, and voice through a multimodal large model, significantly improving the precision and robustness of food nutrition assessment and the accuracy of food nutrition assessment results. Through the continuous learning algorithm, it can dynamically learn new knowledge about new food types without forgetting old knowledge. It can not only adapt to the emergence of new food types, but also maintain high-precision assessment of old food types, which solves the limitations of existing methods when facing new food types and improves the comprehensiveness of food nutrition assessment results. Based on user health data and food nutrition assessment results, it provides users with customized nutrition recommendations to meet the personalized health needs of different users. It is real-time and convenient, and can efficiently and quickly process large-scale data and complete food nutrition assessment.

[0022] The application content will be further explained below through description of embodiments in conjunction with the accompanying drawings.

[0023] The preferred embodiment of the intelligent evaluation method of food nutrient content based on the multimodal large model of the present invention is as follows: Figure 1 As shown, specifically including: S1. Acquire multimodal data and preprocess the multimodal data to obtain target multimodal data.

[0024] In one implementation of this embodiment, acquiring multimodal data and preprocessing the multimodal data to obtain target multimodal data specifically includes: Acquiring multimodal data of the food to be evaluated collected by a smart device, wherein the multimodal data includes images, voice, and text; Performing denoising, enhancement and segmentation processing on the image to obtain a target image; Recognizing the speech using speech recognition technology to obtain recognized text; Cleaning, word segmentation, and format alignment are performed on the text and the recognized text to obtain a target text; The target multimodal data includes the target image and the target text.

[0025] Specifically, multimodal data is first collected. A user uses a smart device (such as a smartphone, tablet, or smartwatch) to capture an image of the food to be evaluated and input a text or voice description related to the food to be evaluated. The image, text description (i.e., text), and voice description (i.e., voice) are used as the multimodal data of the food to be evaluated. The collected multimodal data is then preprocessed. This process includes: denoising, enhancing, and segmenting the image to extract key regions to obtain a target image. Image preprocessing steps such as image resizing, pixel value normalization, and channel order conversion may also be added to meet the input requirements of the image encoder. Speech recognition technology is used to convert speech into text (i.e., recognized text) to reduce model storage space. The text (including the original captured text and the recognized text obtained through speech recognition) is cleaned (removing useless characters and standardizing the encoding), segmented, and formatted (adding special tags, truncation, or padding) to obtain the target text. By preprocessing the multimodal data, the present invention obtains target multimodal data that meets the input requirements of the multimodal large model.

[0026] S2. Using a pre-trained multimodal large model and combining it with an adaptive feature alignment algorithm and a cross-modal attention mechanism, feature extraction, feature alignment, and feature fusion are performed on the target multimodal data to obtain multimodal fusion features.

[0027] In one implementation of this embodiment, the pre-trained multimodal large model is used in combination with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fusion features, specifically including: Using an image encoder in a pre-trained multimodal large model to extract features from the target image to obtain image depth features; Using a text encoder in a pre-trained multimodal large model to extract features from the target text to obtain text semantic features; An adaptive feature alignment algorithm is used to map the features of different modalities to a unified feature space to obtain aligned features of different modalities, where the features of different modalities include image depth features and text semantic features: ; in, Indicates the aligned The characteristics of the mode ( Aligned image depth features, Semantic features of the aligned text), represents the feature alignment function, express The characteristics of the mode ( represents the image depth feature, represents the semantic features of the text), Represents an image, Represents text; it should be noted that the features of the image modality are image depth features, and the features of the text modality are text semantic features; The cross-modal attention mechanism is used to calculate the importance weight coefficients corresponding to different modalities, and the aligned features of different modalities are weightedly fused based on the importance weight coefficients to obtain the multimodal fusion features of the food to be evaluated: ; in, represents multimodal fusion features, express Importance weight coefficient corresponding to the mode, Represents the importance weight coefficient corresponding to the aligned image depth feature (i.e., the importance weight coefficient corresponding to the image modality), Represents the importance weight coefficient corresponding to the aligned text semantic features (i.e., the importance weight coefficient corresponding to the text modality).

[0028] Specifically, the pre-trained multimodal large model of the present invention includes an image encoder and a text encoder. For example, the multimodal large model CLIP (Contrastive Language–Image Pre-training) is used. CLIP includes an image encoder and a text encoder, wherein the image encoder is similar to ViT (Vision Transformer) and the text encoder is similar to RoBERTa (Robustly Optimized Bidirectional Encoder Representations from Transformers Approach). The present invention adopts a pre-trained multimodal large model (such as CLIP) and combines it with an adaptive feature alignment algorithm and a cross-modal attention mechanism to extract features of different modalities from different modal data, map the features of different modalities to a unified feature space, and fuse the aligned features of different modalities. Before feature fusion, the weights of the features of different modalities are dynamically adjusted through the attention mechanism.

[0029] S3. Use a deep learning algorithm to build an initial deep learning model, and train the initial deep learning model to obtain an intermediate deep learning model.

[0030] In one implementation of this embodiment, the use of a deep learning algorithm to construct an initial deep learning model and training the initial deep learning model to obtain an intermediate deep learning model specifically includes: Constructing an initial deep learning model using a deep learning algorithm, wherein the initial deep learning model includes an input layer, a feature processing layer, and an output layer; Acquire historical multimodal data and historical labels corresponding to the historical multimodal data, and perform preprocessing, feature extraction, feature alignment, and feature fusion on the historical multimodal data to obtain historical multimodal fusion features; Training the initial deep learning model using the historical multimodal fusion features and the historical labels to obtain an intermediate deep learning model; The historical multimodal data includes old multimodal data corresponding to old food types.

[0031] Specifically, the present invention adopts a deep learning algorithm (such as the Transformer architecture) to predict the nutritional content of food. First, it is necessary to build an initial deep learning model. The core components include an input layer (receiving multimodal fusion features), a feature processing layer (such as a Transformer encoder, processing multimodal fusion features) and an output layer (such as a fully connected prediction head, mapping the output of the feature processing layer to the prediction target); then initialize the model, define the input / output dimensions and initialization parameters; then obtain historical multimodal data of old food types (i.e., known food types / existing food types) and their corresponding historical labels, use the historical multimodal data and the historical labels as training data sets to train the initial deep learning model, update the parameters through forward propagation, loss calculation, and back propagation, and save the trained model parameters as an intermediate deep learning model.

[0032] S4. When a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned the new knowledge corresponding to the new food type.

[0033] In one implementation of this embodiment, when a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned new knowledge corresponding to the new food type, specifically including: When a new food type appears, new multimodal data corresponding to the new food type and a new label corresponding to the new multimodal data are obtained, and the new multimodal data are preprocessed, feature extracted, aligned, and fused to obtain new multimodal fusion features; Based on the new multimodal fusion features and the new labels, the original parameters of the intermediate deep learning model are updated through incremental learning technology to obtain new parameters of the intermediate deep learning model, so as to continuously learn new knowledge corresponding to the new food type: ; in, represents the new parameters of the intermediate deep learning model, represents the original parameters of the intermediate deep learning model, represents the learning rate, represents the gradient of the loss function with respect to the model parameters, represents the loss function, represents the new multimodal fusion feature, Indicates a new tag; Based on the new parameters and the original parameters of the intermediate deep learning model, the new parameters of the intermediate deep learning model are updated by memory enhancement technology to obtain a target deep learning model, so as to periodically review the old knowledge corresponding to the old food type: ; in, represents the parameters of the target deep learning model, represents the regularization coefficient, represents the gradient of the regularization term with respect to the model parameters, represents the regularization term.

[0034] Specifically, the present invention utilizes a continuous learning algorithm (i.e., incremental learning and memory enhancement mechanisms) to dynamically learn the characteristics and nutritional information of new food types without forgetting existing knowledge. Based on this mechanism, the present invention continuously learns new knowledge through incremental learning techniques and regularly reviews old knowledge through regularization and memory replay techniques (i.e., memory enhancement technologies), preventing the model from forgetting old knowledge as it learns new ones.

[0035] S5. Input the multimodal fusion features into the target deep learning model, and the target deep learning model outputs the nutrient content assessment result.

[0036] In one implementation of this embodiment, the multimodal fusion feature is input into the target deep learning model, and the target deep learning model outputs a nutrient content assessment result, specifically including: Inputting the multimodal fusion features of the food to be evaluated into the input layer of the target deep learning model, and the input layer transmits the multimodal fusion features to the feature processing layer of the target deep learning model; The feature processing layer processes the multimodal fusion features and outputs the processing results to the output layer of the target deep learning model; The output layer maps the processing result into the nutritional content of the food and outputs the nutritional content evaluation result of the food to be evaluated; Obtaining user health data, and using a collaborative filtering algorithm or a content-based recommendation algorithm to obtain user-customized nutritional advice and health risk warnings based on the user health data and the nutrient content assessment results; The nutrient content assessment results, the nutritional recommendations, and the health risk warnings are visualized through interactive charts and graphical interfaces.

[0037] Specifically, the present invention calculates the nutritional content of food based on the fused multimodal features (i.e., multimodal fusion features) and the updated model parameters (i.e., the parameters of the target deep learning model), and provides personalized nutritional recommendations based on the user's health data. The specific process includes: using the target deep learning model to process the multimodal fusion features to predict the nutritional content of the food (i.e., obtain a nutritional content assessment result), and the assessment result includes the content of major nutrients such as calories, protein, fat, and carbohydrates; combining with a preset nutritional database, querying the detailed nutritional information of the food to be assessed in real time, and comparing the query results with the assessment results to ensure the accuracy and comprehensiveness of the assessment results; based on the user's health data (such as age, gender, health goals, dietary preferences, etc.), using collaborative filtering algorithms or content-based recommendation algorithms, providing users with customized nutritional recommendations (including personalized health guidance suggestions and dietary plans) to help users prevent potential health risks; through interactive charts and graphical interfaces, the assessment results and recommendations are intuitively presented to users, including the proportion of nutrients and their comparison with recommended intakes. It should be noted that the architecture of the target deep learning model is the same as that of the initial deep learning model, both of which include input layer, feature processing layer and output layer.

[0038] In one implementation of this embodiment, the method for intelligently assessing the nutritional content of food based on a multimodal large model further includes: Receive user feedback and use the user feedback as a reference factor to update the parameters of the intermediate deep learning model through the incremental learning and memory enhancement mechanism.

[0039] Specifically, obtain user satisfaction feedback on the evaluation results (such as ratings and comments), and use the feedback as input to the continuous learning algorithm to further optimize the model performance.

[0040] In another implementation of this embodiment, the method for intelligently assessing the nutritional content of food based on a multimodal large model further includes: Design an interactive interface through which users can input food information in various ways such as text, voice or image, and input satisfaction feedback on the evaluation results; users receive evaluation results and personalized suggestions through the interactive interface; the layout and functions of the interactive interface are automatically adjusted according to the user's usage habits and preferences to provide a more personalized user experience; the interactive interface supports input and output in multiple languages ​​to meet the needs of different user groups; in addition to providing nutritional assessments and suggestions, the interactive interface can also display nutritional knowledge, health tips and other content to help users improve their health awareness.

[0041] The present invention achieves dynamic evaluation and optimization of food nutritional components by fusing multimodal data such as images, text, voice and user feedback, combined with a continuous learning algorithm; it can automatically adapt to the emergence of new food types while maintaining high-precision evaluation of known food types; it has the characteristics of multimodal data fusion, continuous learning ability and personalized service, providing users with convenient and accurate food nutritional assessment services.

[0042] In addition, based on the above-mentioned method for intelligently evaluating the nutritional content of food based on a multimodal large model, the present invention also provides an intelligent evaluation system for the nutritional content of food based on a multimodal large model, wherein a preferred embodiment of the intelligent evaluation system for the nutritional content of food based on a multimodal large model is as follows: Figure 2 As shown, specifically including: Data acquisition and preprocessing module 01: used to acquire multimodal data and preprocess the multimodal data to obtain target multimodal data; Multimodal large model fusion module 02: uses a pre-trained multimodal large model and combines it with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fusion features; Deep learning model construction module 03: used to construct an initial deep learning model using a deep learning algorithm, and train the initial deep learning model to obtain an intermediate deep learning model; Continuous Learning Module 04: used to update the parameters of the intermediate deep learning model through incremental learning and memory enhancement mechanism when a new food type appears, to obtain a target deep learning model that has learned the new knowledge corresponding to the new food type; Nutrient content assessment module 05: used to input the multimodal fusion features into the target deep learning model, and the target deep learning model outputs the nutrient content assessment result.

[0043] In addition, based on the above-mentioned method and system for intelligent evaluation of food nutrient content based on multimodal large model, the present invention also provides a terminal, wherein a preferred embodiment of the terminal is as follows: Figure 3 As shown, it specifically includes a processor 10, a memory 20 and a display 30. Figure 3 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0044] In some embodiments, the memory 20 can be an internal storage unit of the terminal, such as the terminal's hard drive or memory. In other embodiments, the memory 20 can also be an external storage device of the terminal, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, or a Flash Card equipped on the terminal. Furthermore, the memory 20 can also include both the terminal's internal storage unit and an external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as program code for the terminal. The memory 20 can also be used to temporarily store data that has been output or is about to be output. In one embodiment, the memory 20 stores a program 40 for intelligently assessing the nutritional content of food based on a multimodal macromodel. The program 40 for intelligently assessing the nutritional content of food based on a multimodal macromodel can be executed by the processor 10, thereby implementing the steps of the method for intelligently assessing the nutritional content of food based on a multimodal macromodel in this application.

[0045] In some embodiments, the processor 10 can be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing a food nutrient content intelligent assessment program 40 based on a multimodal large model.

[0046] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0047] In one embodiment, when the processor 10 executes the food nutrient content intelligent assessment program 40 based on the multimodal large model in the memory 20, the steps of the food nutrient content intelligent assessment method based on the multimodal large model as described above are implemented.

[0048] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program for intelligently evaluating the nutrient content of food based on a multimodal large model. When the program for intelligently evaluating the nutrient content of food based on a multimodal large model is executed by a processor, the steps of the method for intelligently evaluating the nutrient content of food based on a multimodal large model as described above are implemented.

[0049] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0050] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When executed, the program can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0051] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for intelligently evaluating the nutritional content of food based on a multimodal large model, characterized in that: The method for intelligently evaluating the nutritional content of food based on a multimodal large model includes: Acquiring multimodal data and preprocessing the multimodal data to obtain target multimodal data; Using a pre-trained multimodal large model and combining it with an adaptive feature alignment algorithm and a cross-modal attention mechanism, feature extraction, feature alignment, and feature fusion are performed on the target multimodal data to obtain multimodal fusion features; Building an initial deep learning model using a deep learning algorithm, and training the initial deep learning model to obtain an intermediate deep learning model; When a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned new knowledge corresponding to the new food type; The multimodal fusion features are input into the target deep learning model, and the target deep learning model outputs a nutrient content assessment result.

2. The method for intelligently evaluating the nutritional content of food based on a multimodal large model according to claim 1, characterized in that: The acquiring multimodal data and preprocessing the multimodal data to obtain target multimodal data specifically includes: Acquiring multimodal data of the food to be evaluated collected by a smart device, wherein the multimodal data includes images, voice, and text; Performing denoising, enhancement and segmentation processing on the image to obtain a target image; Recognizing the speech using speech recognition technology to obtain recognized text; Cleaning, word segmentation, and format alignment are performed on the text and the recognized text to obtain a target text; The target multimodal data includes the target image and the target text.

3. The method for intelligently evaluating the nutritional content of food based on a multimodal large model according to claim 2, characterized in that: The pre-trained multimodal large model is combined with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fusion features, specifically including: Using an image encoder in a pre-trained multimodal large model to extract features from the target image to obtain image depth features; Using a text encoder in a pre-trained multimodal large model to extract features from the target text to obtain text semantic features; An adaptive feature alignment algorithm is used to map the features of different modalities to a unified feature space to obtain aligned features of different modalities, where the features of different modalities include image depth features and text semantic features: ; in, Indicates the aligned The characteristics of the modality, represents the feature alignment function, express The characteristics of the modality, Represents an image, Represents text; The cross-modal attention mechanism is used to calculate the importance weight coefficients corresponding to different modalities, and the aligned features of different modalities are weightedly fused based on the importance weight coefficients to obtain the multimodal fusion features of the food to be evaluated: ; in, represents multimodal fusion features, express Importance weight coefficient corresponding to the mode.

4. The method for intelligently evaluating the nutritional content of food based on a multimodal large model according to claim 3, wherein: The method of constructing an initial deep learning model using a deep learning algorithm and training the initial deep learning model to obtain an intermediate deep learning model specifically includes: Constructing an initial deep learning model using a deep learning algorithm, wherein the initial deep learning model includes an input layer, a feature processing layer, and an output layer; Acquire historical multimodal data and historical labels corresponding to the historical multimodal data, and perform preprocessing, feature extraction, feature alignment, and feature fusion on the historical multimodal data to obtain historical multimodal fusion features; Training the initial deep learning model using the historical multimodal fusion features and the historical labels to obtain an intermediate deep learning model; The historical multimodal data includes old multimodal data corresponding to old food types.

5. The method for intelligently evaluating the nutritional content of food based on a multimodal large model according to claim 4, characterized in that: When a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned new knowledge corresponding to the new food type, specifically including: When a new food type appears, new multimodal data corresponding to the new food type and a new label corresponding to the new multimodal data are obtained, and the new multimodal data are preprocessed, feature extracted, aligned, and fused to obtain new multimodal fusion features; Based on the new multimodal fusion features and the new labels, the original parameters of the intermediate deep learning model are updated through incremental learning technology to obtain new parameters of the intermediate deep learning model, so as to continuously learn new knowledge corresponding to the new food type: ; in, represents the new parameters of the intermediate deep learning model, represents the original parameters of the intermediate deep learning model, represents the learning rate, represents the gradient of the loss function with respect to the model parameters, represents the loss function, represents the new multimodal fusion feature, Indicates a new tag; Based on the new parameters and the original parameters of the intermediate deep learning model, the new parameters of the intermediate deep learning model are updated by memory enhancement technology to obtain a target deep learning model, so as to periodically review the old knowledge corresponding to the old food type: ; in, represents the parameters of the target deep learning model, represents the regularization coefficient, represents the gradient of the regularization term with respect to the model parameters, represents the regularization term.

6. The method for intelligently evaluating the nutritional content of food based on a multimodal large model according to claim 4, characterized in that: Inputting the multimodal fusion features into the target deep learning model, and the target deep learning model outputting a nutrient content assessment result, specifically includes: Inputting the multimodal fusion features of the food to be evaluated into the input layer of the target deep learning model, and the input layer transmits the multimodal fusion features to the feature processing layer of the target deep learning model; The feature processing layer processes the multimodal fusion features and outputs the processing results to the output layer of the target deep learning model; The output layer maps the processing result into the nutritional content of the food and outputs the nutritional content evaluation result of the food to be evaluated; Obtaining user health data, and using a collaborative filtering algorithm or a content-based recommendation algorithm to obtain user-customized nutritional advice and health risk warnings based on the user health data and the nutrient content assessment results; The nutrient content assessment results, the nutritional recommendations, and the health risk warnings are visualized through interactive charts and graphical interfaces.

7. The method for intelligently evaluating the nutritional content of food based on a multimodal large model according to claim 1, characterized in that: The method for intelligently evaluating the nutritional content of food based on the multimodal large model also includes: Receive user feedback and use the user feedback as a reference factor to update the parameters of the intermediate deep learning model through the incremental learning and memory enhancement mechanism.

8. An intelligent food nutrient content assessment system based on a multimodal large model, characterized by: The food nutrient content intelligent assessment system based on the multimodal large model includes: Data acquisition and preprocessing module: used to acquire multimodal data and preprocess the multimodal data to obtain target multimodal data; Multimodal large model fusion module: used to use a pre-trained multimodal large model and combine it with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fusion features; Deep learning model construction module: used to construct an initial deep learning model using a deep learning algorithm, and train the initial deep learning model to obtain an intermediate deep learning model; Continuous learning module: used to update the parameters of the intermediate deep learning model through incremental learning and memory enhancement mechanism when a new food type appears, so as to obtain the target deep learning model that has learned the new knowledge corresponding to the new food type; Nutrient content assessment module: used to input the multimodal fusion features into the target deep learning model, and the target deep learning model outputs the nutrient content assessment result.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an intelligent evaluation program for the nutritional content of food based on a multimodal large model, which is stored in the memory and can be run on the processor. When the intelligent evaluation program for the nutritional content of food based on a multimodal large model is executed by the processor, the steps of the intelligent evaluation method for the nutritional content of food based on a multimodal large model are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program for intelligently evaluating the nutrient content of food based on a multimodal large model. When the program for intelligently evaluating the nutrient content of food based on a multimodal large model is executed by a processor, the steps of the method for intelligently evaluating the nutrient content of food based on a multimodal large model as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-modal knowledge generation method and device based on feedback enhancement

    CN117035074A

  • Multimodal model knowledge updating method based on knowledge representation and dynamic prompt

    CN118627610A

  • Multimodal neural biological signal processing method and apparatus, and server and storage medium

    WO2024108483A1