A food nutrition ingredient content intelligent evaluation method, system, terminal and storage medium based on a multi-modal large model

By combining multimodal large models and deep learning algorithms with adaptive feature alignment and cross-modal attention mechanisms, the model parameters are dynamically updated, solving the problem of insufficient adaptability to new food types and achieving high-precision, personalized food nutrition assessment to meet large-scale and real-time requirements.

CN120613031BActive Publication Date: 2025-11-21THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114731.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing AI-based intelligent food nutrition assessment methods have shortcomings in adaptability to new food types and depth of multimodal data fusion, resulting in poor accuracy and comprehensiveness of assessment results, making it difficult to meet the needs of large-scale, real-time, and personalized nutrition assessment.

Method used

A multimodal large model is used in combination with an adaptive feature alignment algorithm and a cross-modal attention mechanism to extract, align and fuse features, construct an initial deep learning model, and update the model parameters through incremental learning and memory enhancement mechanisms to dynamically learn knowledge of new food types and output nutritional content assessment results.

Benefits of technology

It significantly improves the accuracy and robustness of food nutrition assessment, can adapt to the emergence of new food types, while maintaining high-precision assessment of old food types, provides personalized nutrition advice, and meets the needs of real-time and large-scale data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613031B_ABST
    Figure CN120613031B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of multi-modal data processing, and discloses a food nutrition ingredient content intelligent evaluation method, system, terminal and storage medium based on a multi-modal large model, the method comprising the following steps: obtaining multi-modal data and performing pretreatment to obtain target multi-modal data; adopting a pre-trained multi-modal large model to perform feature extraction, feature alignment and feature fusion on the target multi-modal data to obtain multi-modal fusion features; when a new food type appears, updating parameters of an intermediate deep learning model through an incremental learning and memory enhancement mechanism to obtain a target deep learning model; and inputting the multi-modal fusion features into the target deep learning model to output a nutrition ingredient content evaluation result. Through the multi-modal large model and the continuous learning algorithm, high-precision dynamic evaluation of food nutrition ingredients is realized, the emergence of new food types can be adapted to, and high-precision evaluation of old food types can be maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data processing technology, and in particular to a method, system, terminal, and computer-readable storage medium for intelligent assessment of food nutrient content based on a multimodal large model. Background Technology

[0002] With increasing global focus on healthy eating, food nutrition assessment has become a crucial component of public health. Traditional methods rely on manual recording, laboratory testing, or pre-set food databases, which are not only time-consuming and labor-intensive but also ill-suited to the rapid changes in food types and individualized needs. In recent years, with the development of artificial intelligence technology, especially the rise of multimodal large models, intelligent assessment systems based on images, text, and voice have gradually become a research hotspot. Existing technologies have yielded various intelligent food nutrition assessment systems for intelligently assessing the content of key nutrients in food and automatically generating personalized plans based on client dietary needs, integrating nutrition assessment, planning, storage, and sharing functions.

[0003] However, existing intelligent food nutrition assessment methods based on artificial intelligence technology still have shortcomings in terms of adaptability to new food types and depth of multimodal data fusion, resulting in poor accuracy and comprehensiveness of food nutrition assessment results, and making it difficult to meet the needs of large-scale, real-time, and personalized nutrition assessment.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide a method, system, terminal, and computer-readable storage medium for intelligent assessment of food nutrient content based on a multimodal large model. This aims to address the shortcomings of existing intelligent food nutrient assessment methods in terms of adaptability to new food types and depth of multimodal data fusion, which result in poor accuracy and comprehensiveness of assessment results.

[0006] To achieve the above-mentioned objectives, this invention provides an intelligent assessment method for the content of food nutrients based on a multimodal large model. The method includes:

[0007] Acquire multimodal data and preprocess the multimodal data to obtain target multimodal data;

[0008] A pre-trained multimodal large model is used in conjunction with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fused features;

[0009] An initial deep learning model is constructed by using a deep learning algorithm, and the initial deep learning model is trained to obtain an intermediate deep learning model;

[0010] When a new food type appears, the parameters of the intermediate deep learning model are updated by an incremental learning and memory enhancement mechanism to obtain a target deep learning model that has learned new knowledge corresponding to the new food type;

[0011] The multi-modal fusion features are input into the target deep learning model, and the target deep learning model outputs a nutritional ingredient content evaluation result.

[0012] Optionally, the multi-modal data is obtained and preprocessed to obtain target multi-modal data, specifically including:

[0013] Multi-modal data of a food to be evaluated collected by an intelligent device is obtained, wherein the multi-modal data includes images, voices, and texts;

[0014] The images are denoised, enhanced, and segmented to obtain target images;

[0015] The voices are recognized by using a voice recognition technology to obtain recognized texts;

[0016] The texts and the recognized texts are cleaned, segmented, and format-aligned to obtain target texts;

[0017] The target multi-modal data includes the target images and the target texts.

[0018] Optionally, the target multi-modal data is feature-extracted, feature-aligned, and feature-fused by using a pre-trained multi-modal large model in combination with a self-adaptive feature alignment algorithm and a cross-modal attention mechanism to obtain multi-modal fusion features, specifically including:

[0019] The target images are feature-extracted by using an image encoder in the pre-trained multi-modal large model to obtain image deep features;

[0020] The target texts are feature-extracted by using a text encoder in the pre-trained multi-modal large model to obtain text semantic features;

[0021] Different modal features are respectively mapped to a unified feature space by using a self-adaptive feature alignment algorithm to obtain aligned different modal features, wherein the different modal features include image deep features and text semantic features:

[0022] ;

[0023] wherein, indicates the aligned a feature of a modality, representing a feature alignment function, representing a feature of a modality, representing an image, representing text;

[0024] An importance weight coefficient corresponding to different modalities is calculated by using a cross-modal attention mechanism, and features of different modalities after alignment are weighted and fused based on the importance weight coefficient to obtain a multi-modal fusion feature of the food to be evaluated:

[0025] ;

[0026] wherein, representing a multi-modal fusion feature, representing an importance weight coefficient corresponding to a modality.

[0027] Optionally, the initial deep learning model is constructed by using a deep learning algorithm, and the initial deep learning model is trained to obtain an intermediate deep learning model, and specifically includes:

[0028] An initial deep learning model is constructed by using a deep learning algorithm, wherein the initial deep learning model includes an input layer, a feature processing layer and an output layer;

[0029] Historical multi-modal data and historical labels corresponding to the historical multi-modal data are obtained, and the historical multi-modal data are preprocessed, feature extracted, feature aligned and feature fused to obtain historical multi-modal fusion features;

[0030] The initial deep learning model is trained by using the historical multi-modal fusion features and the historical labels to obtain an intermediate deep learning model;

[0031] Wherein, the historical multi-modal data include old multi-modal data corresponding to old food types.

[0032] Optionally, when a new food type appears, the parameters of the intermediate deep learning model are updated by using an incremental learning and memory enhancement mechanism to obtain a target deep learning model that has learned new knowledge corresponding to the new food type, and specifically includes:

[0033] When a new food type appears, new multi-modal data corresponding to the new food type and new labels corresponding to the new multi-modal data are obtained, and the new multi-modal data are preprocessed, feature extracted, feature aligned and feature fused to obtain new multi-modal fusion features;

[0034] updating the original parameters of the intermediate deep learning model by an incremental learning technique based on the new multi-modal fusion features and the new labels to obtain new parameters of the intermediate deep learning model for continuously learning new knowledge corresponding to the new food type:

[0035] ;

[0036] wherein, denotes the new parameters of the intermediate deep learning model, denotes the original parameters of the intermediate deep learning model, denotes a learning rate, denotes a gradient of a loss function to model parameters, denotes the loss function, denotes the new multi-modal fusion features, denotes the new labels;

[0037] updating the new parameters of the intermediate deep learning model by a memory enhancement technique based on the new parameters and the original parameters of the intermediate deep learning model to obtain a target deep learning model for regularly reviewing old knowledge corresponding to the old food type:

[0038] ;

[0039] wherein, denotes parameters of the target deep learning model, denotes a regularization coefficient, denotes a gradient of a regularization term to model parameters, denotes the regularization term.

[0040] Optionally, the inputting the multi-modal fusion features into the target deep learning model, the target deep learning model outputs a nutritional ingredient content evaluation result, specifically comprising:

[0041] inputting the multi-modal fusion features of the food to be evaluated into an input layer of the target deep learning model, the input layer transmitting the multi-modal fusion features to a feature processing layer of the target deep learning model;

[0042] the feature processing layer processes the multi-modal fusion features and outputs a processing result to an output layer of the target deep learning model;

[0043] the output layer maps the processing result into a food nutritional ingredient content and outputs a nutritional ingredient content evaluation result of the food to be evaluated;

[0044] Obtaining user health data, and according to the user health data and the nutrition ingredient content evaluation result, using a collaborative filtering algorithm or a content-based recommendation algorithm to obtain user customized nutrition suggestions and health risk prompts;

[0045] Visualizing the nutrition ingredient content evaluation result, the nutrition suggestions and the health risk prompts through an interactive chart and a graphical interface.

[0046] Optionally, the food nutrition ingredient content intelligent evaluation method based on a multi-modal large model further comprises:

[0047] Receiving user feedback, and in the process of updating the parameters of the intermediate deep learning model through an incremental learning and memory enhancement mechanism, taking the user feedback as a reference factor to update the parameters of the intermediate deep learning model.

[0048] To achieve the above-mentioned purposes, the present application further provides a food nutrition ingredient content intelligent evaluation system based on a multi-modal large model, which comprises:

[0049] A data acquisition and preprocessing module for acquiring multi-modal data and preprocessing the multi-modal data to obtain target multi-modal data;

[0050] A multi-modal large model fusion module for using a pre-trained multi-modal large model and combining a self-adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment and feature fusion on the target multi-modal data to obtain multi-modal fusion features;

[0051] A deep learning model construction module for constructing an initial deep learning model using a deep learning algorithm and training the initial deep learning model to obtain an intermediate deep learning model;

[0052] A continuous learning module for updating the parameters of the intermediate deep learning model through an incremental learning and memory enhancement mechanism when a new food type appears to obtain a target deep learning model that has learned new knowledge corresponding to the new food type;

[0053] A nutrition ingredient content evaluation module for inputting the multi-modal fusion features into the target deep learning model, and the target deep learning model outputs a nutrition ingredient content evaluation result.

[0054] To achieve the above-mentioned purposes of the application, the application further provides a terminal, comprising a memory, a processor, and a food nutrition content intelligent evaluation program based on a multi-modal large model stored on the memory and executable on the processor, wherein the food nutrition content intelligent evaluation program based on the multi-modal large model implements the steps of the food nutrition content intelligent evaluation method based on the multi-modal large model when executed by the processor.

[0055] To achieve the above-mentioned purposes of the application, the application further provides a computer readable storage medium storing a food nutrition content intelligent evaluation program based on a multi-modal large model, wherein the food nutrition content intelligent evaluation program based on the multi-modal large model implements the steps of the food nutrition content intelligent evaluation method based on the multi-modal large model when executed by a processor.

[0056] In the application, multi-modal data is acquired and preprocessed to obtain target multi-modal data; a pre-trained multi-modal large model is used in combination with a self-adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multi-modal data to obtain multi-modal fusion features; a deep learning algorithm is used to construct an initial deep learning model, and the initial deep learning model is trained to obtain an intermediate deep learning model; when a new food type appears, the parameters of the intermediate deep learning model are updated through an incremental learning and memory enhancement mechanism to obtain a target deep learning model that has learned new knowledge corresponding to the new food type; the multi-modal fusion features are input into the target deep learning model, and the target deep learning model outputs a nutrition content evaluation result. The application significantly improves the precision and robustness of food nutrition evaluation and improves the accuracy of the food nutrition evaluation result by deeply fusing multi-modal data such as images, texts, and voices through a multi-modal large model. The application can dynamically learn new knowledge of new food types without forgetting old knowledge through a continuous learning algorithm, which not only adapts to the appearance of new food types but also maintains high-precision evaluation of old food types, solves the limitations of existing methods when facing new food types, and improves the comprehensiveness of the food nutrition evaluation result. Based on user health data and food nutrition evaluation results, the application provides customized nutrition suggestions to meet the individual health needs of different users. The application has real-time and convenient features and can efficiently and quickly process large-scale data and complete food nutrition evaluation. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 is a flowchart of a preferred embodiment of the food nutrition content intelligent evaluation method based on a multi-modal large model of the application;

[0058] Figure 2is a structural diagram of a preferred embodiment of a food nutrition content intelligent evaluation system based on a multi-modal large model of the present application;

[0059] Figure 3 is a structural diagram of a preferred embodiment of a terminal of the present application. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of the present application clearer and more explicit, the present application will be further described in detail below with reference to the drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0061] With the increasing global attention to healthy eating, food nutrition evaluation has become an important part of public health. Traditional food nutrition evaluation methods rely on manual recording, laboratory testing or pre-set food databases. These methods not only consume time and effort, but also are difficult to adapt to the rapid changes in food types and individual needs. In recent years, with the development of artificial intelligence technology, especially the rise of multi-modal large models, intelligent evaluation systems based on images, texts and voices have gradually become a research hotspot. Existing technologies have appeared various food nutrition intelligent evaluation systems, for example, SmartPlate is an AI (Artificial Intelligence) driven meal plan platform, which is used to intelligently evaluate the content of main nutritional components of food, and automatically generate personalized solutions according to the dietary needs of customers, integrating nutrition evaluation, plan development, storage and sharing functions. However, the existing food nutrition intelligent evaluation methods based on artificial intelligence technology still have deficiencies in the adaptability of new food types and the depth of multi-modal data fusion, resulting in poor accuracy and comprehensiveness of food nutrition evaluation results, and difficulty in meeting the large-scale, real-time and personalized nutrition evaluation needs.

[0062] In the field of nutrition evaluation, the application of artificial intelligence technology has made significant progress. For example, through image recognition technology, the system can automatically identify food images and estimate their nutritional components, thereby simplifying the dietary evaluation method. In addition, some systems not only provide personalized nutrition recommendations, but also combine intelligent shopping lists and cooking guides, further improving the user experience. Although these applications demonstrate the great potential of artificial intelligence in the field of nutrition and health, existing artificial intelligence-based nutrition evaluation methods still have limitations in multi-modal data fusion and continuous learning capabilities.

[0063] To solve the above technical problems, the application provides a food nutrition ingredient content intelligent evaluation method based on a multi-modal large model, multi-modal data is acquired and preprocessed to obtain target multi-modal data; a pre-trained multi-modal large model is used in combination with a self-adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment and feature fusion on the target multi-modal data to obtain multi-modal fusion features; a deep learning algorithm is used to construct an initial deep learning model, and the initial deep learning model is trained to obtain an intermediate deep learning model; when a new food type appears, the parameters of the intermediate deep learning model are updated through an incremental learning and memory enhancement mechanism to obtain a target deep learning model that has learned new knowledge corresponding to the new food type; the multi-modal fusion features are input into the target deep learning model, and the target deep learning model outputs a nutrition ingredient content evaluation result. The application significantly improves the precision and robustness of food nutrition evaluation by deeply fusing various modal data such as images, texts and voices through a multi-modal large model, and improves the accuracy of the food nutrition evaluation result; through a continuous learning algorithm, new knowledge of a new food type can be dynamically learned without forgetting old knowledge, which not only adapts to the appearance of new food types, but also maintains high-precision evaluation of old food types, solves the limitations of existing methods when facing new food types, and improves the comprehensiveness of the food nutrition evaluation result; based on user health data and food nutrition evaluation results, customized nutrition suggestions are provided for users to meet the personalized health needs of different users; the application has real-time and convenient performance, can efficiently and quickly process large-scale data and complete food nutrition evaluation.

[0064] The application content will be further described through the description of the embodiments in combination with the drawings.

[0065] The preferred embodiment of the food nutrition ingredient content intelligent evaluation method based on a multi-modal large model of the application is shown in Figure 1 , and specifically includes:

[0066] S1, acquiring multi-modal data and preprocessing the multi-modal data to obtain target multi-modal data.

[0067] In one implementation manner of the embodiment, the acquiring multi-modal data and preprocessing the multi-modal data to obtain target multi-modal data specifically includes:

[0068] Acquiring multi-modal data of a food to be evaluated collected by an intelligent device, wherein the multi-modal data includes images, voices and texts;

[0069] Performing denoising, enhancement and segmentation processing on the images to obtain target images;

[0070] recognize the voice by using a voice recognition technology to obtain recognized text;

[0071] perform cleaning, word segmentation and format alignment processing on the text and the recognized text to obtain target text;

[0072] The target multi-modal data comprises the target image and the target text.

[0073] Specifically, first, multi-modal data is collected, and a user captures an image of a food to be evaluated, inputs a text description or a voice description related to the food to be evaluated through a smart device (such as a smart phone, a tablet computer or a smart watch), and takes the image, the text description (i.e., text) and the voice description (i.e., voice) as multi-modal data of the food to be evaluated. Then, the collected multi-modal data is preprocessed, and the specific process comprises: performing denoising, enhancement and segmentation processing on the image, extracting a key region, obtaining a target image, and also adding image preprocessing steps such as image size adjustment, pixel value normalization and channel order conversion to adapt to the requirements of an image encoder for input; recognizing the voice into text (i.e., recognized text) by using a voice recognition technology to reduce the model storage space; performing cleaning (removing useless symbols, standardizing coding), word segmentation and format alignment (adding special markers, truncation or padding) processing on the text (including the originally collected text and the recognized text) to obtain target text. The present application obtains target multi-modal data meeting the input requirements of a multi-modal large model by preprocessing multi-modal data.

[0074] S2, using a pre-trained multi-modal large model and combining a self-adaptive feature alignment algorithm and a cross-modal attention mechanism, performing feature extraction, feature alignment and feature fusion on the target multi-modal data to obtain multi-modal fusion features.

[0075] In one implementation manner of the embodiment, the using a pre-trained multi-modal large model and combining a self-adaptive feature alignment algorithm and a cross-modal attention mechanism, performing feature extraction, feature alignment and feature fusion on the target multi-modal data to obtain multi-modal fusion features specifically comprises:

[0076] using an image encoder in the pre-trained multi-modal large model to perform feature extraction on the target image to obtain image deep features;

[0077] using a text encoder in the pre-trained multi-modal large model to perform feature extraction on the target text to obtain text semantic features;

[0078] using a self-adaptive feature alignment algorithm to respectively map features of different modalities to a unified feature space to obtain features of different modalities after alignment, wherein the features of different modalities comprise the image deep features and the text semantic features:

[0079] ;

[0080] wherein, denote the aligned features of the modalities (image depth features and text semantic features), the aligned image depth features, the aligned text semantic features), denote a feature alignment function, denote features of the modalities (image depth features and text semantic features), denote the image depth features, denote the text semantic features), denote an image, denote a text; it is to be noted that the features of the image modality are the image depth features, and the features of the text modality are the text semantic features;

[0081] The importance weight coefficients corresponding to different modalities are calculated by using the cross-modal attention mechanism, and the aligned features of different modalities are weighted and fused based on the importance weight coefficients to obtain the multi-modal fusion features of the food to be evaluated:

[0082] ;

[0083] wherein, denote the multi-modal fusion features, denote the importance weight coefficients corresponding to different modalities, denote the importance weight coefficients corresponding to the aligned image depth features (i.e. the importance weight coefficients corresponding to the image modality), denote the importance weight coefficients corresponding to the aligned text semantic features (i.e. the importance weight coefficients corresponding to the text modality).

[0084] Specifically, the pre-trained multi-modal large model of the present application includes an image encoder and a text encoder, such as using a multi-modal large model CLIP (Contrastive Language-Image Pre-training), which includes an image encoder similar to ViT (Vision Transformer) and a text encoder similar to RoBERTa (Robustly Optimized Bidirectional Encoder Representations from Transformers Approach). The present application uses a pre-trained multi-modal large model (such as CLIP), and combines a self-adaptive feature alignment algorithm and a cross-modal attention mechanism to extract features of different modalities from different modal data, map the features of different modalities to a unified feature space, and fuse the aligned features of different modalities, wherein the weights of the features of different modalities are dynamically adjusted by the attention mechanism before feature fusion.

[0085] S3, constructing an initial deep learning model using a deep learning algorithm and training the initial deep learning model to obtain an intermediate deep learning model.

[0086] In one implementation of the present embodiment, constructing an initial deep learning model using a deep learning algorithm and training the initial deep learning model to obtain an intermediate deep learning model specifically includes:

[0087] constructing an initial deep learning model using a deep learning algorithm, wherein the initial deep learning model includes an input layer, a feature processing layer, and an output layer;

[0088] obtaining historical multi-modal data and historical labels corresponding to the historical multi-modal data, and pre-processing, feature extraction, feature alignment, and feature fusion of the historical multi-modal data to obtain historical multi-modal fusion features;

[0089] training the initial deep learning model using the historical multi-modal fusion features and the historical labels to obtain an intermediate deep learning model;

[0090] wherein the historical multi-modal data includes old multi-modal data corresponding to old food types.

[0091] Specifically, the present application adopts a deep learning algorithm (such as a Transformer architecture) to predict the nutritional ingredient content of food. First, an initial deep learning model needs to be constructed, the core components of which include an input layer (receiving multi-modal fusion features), a feature processing layer (such as a Transformer encoder, processing multi-modal fusion features), and an output layer (such as a fully connected prediction head, mapping the output of the feature processing layer to the prediction target). Then, the model is initialized, the input / output dimensions and the initialization parameters are defined. Next, the historical multi-modal data of old food types (i.e. known food types / existing food types) and their corresponding historical labels are obtained. The historical multi-modal data and the historical labels are used as a training data set to train the initial deep learning model. The parameters are updated through forward propagation, loss calculation, and back propagation, and the trained model parameters are saved as an intermediate deep learning model.

[0092] S4, when a new food type appears, the parameters of the intermediate deep learning model are updated through an incremental learning and memory enhancement mechanism to obtain a target deep learning model that has learned new knowledge corresponding to the new food type.

[0093] In one implementation of the present embodiment, when a new food type appears, the parameters of the intermediate deep learning model are updated through an incremental learning and memory enhancement mechanism to obtain a target deep learning model that has learned new knowledge corresponding to the new food type, specifically including:

[0094] When a new food type appears, the new multi-modal data corresponding to the new food type and the new labels corresponding to the new multi-modal data are obtained, and the new multi-modal data is preprocessed, feature extracted, feature aligned, and feature fused to obtain new multi-modal fusion features;

[0095] Based on the new multi-modal fusion features and the new labels, the original parameters of the intermediate deep learning model are updated through an incremental learning technique to obtain new parameters of the intermediate deep learning model to continuously learn new knowledge corresponding to the new food type:

[0096] ;

[0097] wherein, denotes the new parameters of the intermediate deep learning model, denotes the original parameters of the intermediate deep learning model, denotes the learning rate, denotes the gradient of the loss function with respect to the model parameters, denotes the loss function, denotes the new multi-modal fusion features, denotes the new labels;

[0098] updating new parameters of the intermediate deep learning model by a memory enhancement technique based on the new parameters and the original parameters of the intermediate deep learning model to obtain a target deep learning model for regularly reviewing old knowledge corresponding to the old food type:

[0099] ;

[0100] wherein, denotes parameters of the target deep learning model, denotes a regularization coefficient, denotes a gradient of the regularization term on the model parameters, denotes the regularization term.

[0101] Specifically, the present application can dynamically learn the characteristics and nutritional information of new food types without forgetting existing knowledge through a continuous learning algorithm (i.e., incremental learning and memory enhancement mechanism). Based on the incremental learning and memory enhancement mechanism, the present application continuously learns new knowledge through incremental learning technology, and regularly reviews old knowledge through regularization and memory replay technology (i.e., memory enhancement technology) to prevent the model from forgetting old knowledge when learning new knowledge.

[0102] S5, inputting the multi-modal fusion feature into the target deep learning model, and outputting a nutritional ingredient content evaluation result by the target deep learning model.

[0103] In one implementation of the embodiment, the inputting the multi-modal fusion feature into the target deep learning model and outputting a nutritional ingredient content evaluation result by the target deep learning model specifically includes:

[0104] inputting the multi-modal fusion feature of the food to be evaluated into an input layer of the target deep learning model, and transmitting the multi-modal fusion feature to a feature processing layer of the target deep learning model by the input layer;

[0105] processing the multi-modal fusion feature by the feature processing layer, and outputting a processing result to an output layer of the target deep learning model;

[0106] mapping the processing result into a food nutritional ingredient content by the output layer, and outputting a nutritional ingredient content evaluation result of the food to be evaluated;

[0107] obtaining user health data, and obtaining user customized nutritional suggestions and health risk prompts by using a collaborative filtering algorithm or a content-based recommendation algorithm according to the user health data and the nutritional ingredient content evaluation result;

[0108] visualizing the nutritional ingredient content evaluation result, the nutritional suggestions and the health risk prompts through an interactive chart and a graphical interface.

[0109] Specifically, according to the fused multi-modal features (i.e., multi-modal fusion features) and the updated model parameters (i.e., parameters of the target deep learning model), the nutritional ingredient content of the food is calculated, and personalized nutrition recommendations are provided according to the user's health data. The specific process includes: using the target deep learning model to process the multi-modal fusion features to predict the nutritional ingredient content of the food (i.e., to obtain the nutritional ingredient content evaluation result), and the evaluation result includes the content of main nutritional ingredients such as calories, protein, fat, and carbohydrates; combining a pre-set nutrition database, real-time query of detailed nutrition information of the food to be evaluated, comparison of the query result with the evaluation result to ensure the accuracy and comprehensiveness of the evaluation result; according to the user's health data (such as age, gender, health goals, dietary preferences, etc.), using a collaborative filtering algorithm or a content-based recommendation algorithm to provide customized nutrition recommendations (including personalized health guidance recommendations and meal plans) for the user, helping the user to prevent potential health risks; through interactive charts and graphical interfaces, the evaluation result and the recommendations are intuitively displayed to the user, including the proportion of nutritional ingredients and their comparison with recommended intake, etc. It should be noted that the architecture of the target deep learning model is the same as that of the initial deep learning model, and both include an input layer, a feature processing layer, and an output layer.

[0110] In one implementation of the embodiment, the multi-modal large model-based intelligent evaluation method of the nutritional ingredient content of food further comprises:

[0111] Receiving user feedback and updating the parameters of the intermediate deep learning model by taking the user feedback as a reference factor in the process of updating the parameters of the intermediate deep learning model through incremental learning and memory enhancement mechanism.

[0112] Specifically, the user's satisfaction feedback (such as rating, comment) on the evaluation result is obtained, and the feedback is taken as the input of the continuous learning algorithm to further optimize the model performance.

[0113] In another implementation of the embodiment, the multi-modal large model-based intelligent evaluation method of the nutritional ingredient content of food further comprises:

[0114] An interactive interface is designed, through which the user inputs food information in various ways such as text, voice, or image, and inputs satisfaction feedback on the evaluation result through the interactive interface; the user receives the evaluation result and personalized recommendations through the interactive interface; according to the user's usage habits and preferences, the layout and functions of the interactive interface are automatically adjusted to provide more personalized user experience; the interactive interface supports input and output in multiple languages to meet the needs of different user groups; in addition to providing nutrition evaluation and recommendations, the interactive interface can also display nutrition knowledge, health tips, and other content to help users improve their health awareness.

[0115] The application realizes dynamic evaluation and optimization of food nutritional components by fusing multi-modal data such as images, texts, voices and user feedback, combining a continuous learning algorithm; can automatically adapt to the appearance of new food types while maintaining high-precision evaluation of known food types; has the characteristics of multi-modal data fusion, continuous learning ability and personalized service, and provides convenient and accurate food nutrition evaluation services for users.

[0116] In addition, based on the above-mentioned multi-modal large model-based food nutritional component content intelligent evaluation method, the application also provides a multi-modal large model-based food nutritional component content intelligent evaluation system, wherein the preferred embodiment of the multi-modal large model-based food nutritional component content intelligent evaluation system is as shown in Figure 2 , and specifically includes:

[0117] The data acquisition and preprocessing module 01 is used for acquiring multi-modal data and preprocessing the multi-modal data to obtain target multi-modal data;

[0118] The multi-modal large model fusion module 02 is used for adopting a pre-trained multi-modal large model and combining a self-adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment and feature fusion on the target multi-modal data to obtain multi-modal fusion features;

[0119] The deep learning model construction module 03 is used for constructing an initial deep learning model by using a deep learning algorithm and training the initial deep learning model to obtain an intermediate deep learning model;

[0120] The continuous learning module 04 is used for updating parameters of the intermediate deep learning model by an incremental learning and memory enhancement mechanism when a new food type appears to obtain a target deep learning model that has learned new knowledge corresponding to the new food type;

[0121] The nutritional component content evaluation module 05 is used for inputting the multi-modal fusion features into the target deep learning model, and the target deep learning model outputs a nutritional component content evaluation result.

[0122] In addition, based on the above-mentioned multi-modal large model-based food nutritional component content intelligent evaluation method and system, the application also correspondingly provides a terminal, wherein the preferred embodiment of the terminal is as shown in Figure 3 , and specifically includes a processor 10, a memory 20 and a display 30. Figure 3 Only part of the components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0123] The memory 20 can be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 can also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, and the like. Further, the memory 20 can include both an internal storage unit and an external storage device of the terminal. The memory 20 is used to store application software and various data installed on the terminal, such as program codes of the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 20 stores the food nutrition content intelligent evaluation program based on a multi-modal large model 40, which can be executed by the processor 10 to implement the steps of the food nutrition content intelligent evaluation method based on a multi-modal large model in the present application.

[0124] The processor 10 can be a Central Processing Unit (CPU), a microprocessor, or other data processing chip in some embodiments, which is used to run program codes or process data stored in the memory 20, such as the food nutrition content intelligent evaluation program based on a multi-modal large model 40.

[0125] The display 30 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, and the like in some embodiments. The display 30 is used to display information of the terminal and to display a visualized user interface.

[0126] In an embodiment, the processor 10 implements the steps of the food nutrition content intelligent evaluation method based on a multi-modal large model as described above when executing the food nutrition content intelligent evaluation program based on a multi-modal large model 40 in the memory 20.

[0127] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a food nutrition content intelligent evaluation program based on a multi-modal large model, which implements the steps of the food nutrition content intelligent evaluation method based on a multi-modal large model as described above when executed by a processor.

[0128] It should be noted that, in the present document, the terms "comprises / comprising" or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more limitations, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0129] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer readable storage medium readable by a computer. When the program is executed, it can include the processes of the above-mentioned method embodiments. The computer readable storage medium can be a memory, a disk, an optical disk, etc.

[0130] It should be understood that the application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all these improvements and changes shall fall within the protection scope of the appended claims of the present application.

Claims

1. A method for intelligent assessment of food nutrient content based on a multimodal large model, characterized in that, The intelligent assessment method for food nutrient content based on a multimodal large model includes: Acquire multimodal data and preprocess the multimodal data to obtain target multimodal data; A pre-trained multimodal large model is used in conjunction with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data to obtain multimodal fused features; An initial deep learning model is constructed using a deep learning algorithm, and the initial deep learning model is trained to obtain an intermediate deep learning model. When a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain the target deep learning model that has learned the new knowledge corresponding to the new food type. The multimodal fusion features are input into the target deep learning model, and the target deep learning model outputs the nutritional content assessment results. When a new food type appears, the parameters of the intermediate deep learning model are updated through incremental learning and memory enhancement mechanisms to obtain a target deep learning model that has learned the new knowledge corresponding to the new food type. Specifically, this includes: When a new food type appears, new multimodal data corresponding to the new food type and new labels corresponding to the new multimodal data are obtained, and the new multimodal data is preprocessed, feature extracted, feature aligned and feature fused to obtain new multimodal fused features. Based on the new multimodal fusion features and the new labels, the original parameters of the intermediate deep learning model are updated using incremental learning techniques to obtain new parameters for the intermediate deep learning model, so as to continuously learn new knowledge corresponding to the new food type. in, This represents the new parameters of the intermediate deep learning model. This represents the original parameters of the intermediate deep learning model. Indicates the learning rate. This represents the gradient of the loss function with respect to the model parameters. Represents the loss function. This indicates new multimodal fusion features. Indicates a new tag; Based on the new parameters and the original parameters of the intermediate deep learning model, the new parameters of the intermediate deep learning model are updated using memory enhancement techniques to obtain the target deep learning model, which periodically reviews old knowledge corresponding to old food types. in, These represent the parameters of the target deep learning model. Represents the regularization coefficient. This represents the gradient of the regularization term with respect to the model parameters. This represents the regularization term.

2. The intelligent assessment method for food nutrient content based on a multimodal large model according to claim 1, characterized in that, The process of acquiring multimodal data and preprocessing the multimodal data to obtain target multimodal data specifically includes: Acquire multimodal data of the food to be evaluated collected by a smart device, wherein the multimodal data includes images, voice and text; The image is then subjected to denoising, enhancement, and segmentation processes to obtain the target image; The speech is identified using speech recognition technology to obtain the identified text; The text and the identified text are cleaned, segmented, and formatted to obtain the target text. The target multimodal data includes the target image and the target text.

3. The intelligent assessment method for food nutrient content based on a multimodal large model according to claim 2, characterized in that, The method employs a pre-trained multimodal large model combined with an adaptive feature alignment algorithm and a cross-modal attention mechanism to perform feature extraction, feature alignment, and feature fusion on the target multimodal data, obtaining multimodal fused features, specifically including: The target image is used to extract features by employing an image encoder in a pre-trained multimodal large model to obtain image depth features; The target text is used to extract features by employing a text encoder in a pre-trained multimodal large model to obtain text semantic features; An adaptive feature alignment algorithm is used to map features from different modalities to a unified feature space, resulting in aligned features for each modality. These features include image depth features and text semantic features. in, The features representing the aligned i-mode are... Represents the feature alignment function. Features representing the i-modality: image represents an image, and text represents text. A cross-modal attention mechanism is used to calculate the importance weight coefficients corresponding to different modalities. Based on these importance weight coefficients, the aligned features of different modalities are weighted and fused to obtain the multimodal fusion features of the food to be evaluated. in, Indicates multimodal fusion features, This represents the importance weight coefficient corresponding to mode i.

4. The intelligent assessment method for food nutrient content based on a multimodal large model according to claim 3, characterized in that, The process of constructing an initial deep learning model using a deep learning algorithm and training the initial deep learning model to obtain an intermediate deep learning model specifically includes: An initial deep learning model is constructed using a deep learning algorithm, wherein the initial deep learning model includes an input layer, a feature processing layer, and an output layer; Historical multimodal data and corresponding historical labels are obtained, and the historical multimodal data is preprocessed, feature extracted, feature aligned and fused to obtain historical multimodal fused features. The initial deep learning model is trained using the historical multimodal fusion features and the historical labels to obtain an intermediate deep learning model; The historical multimodal data includes old multimodal data corresponding to old food types.

5. The intelligent assessment method for food nutrient content based on a multimodal large model according to claim 4, characterized in that, The step of inputting the multimodal fusion features into the target deep learning model, and the target deep learning model outputting the nutritional content assessment result, specifically includes: The multimodal fusion features of the food to be evaluated are input into the input layer of the target deep learning model, and the input layer transmits the multimodal fusion features to the feature processing layer of the target deep learning model. The feature processing layer processes the multimodal fusion features and outputs the processing results to the output layer of the target deep learning model; The output layer maps the processing results into the nutritional content of food and outputs the nutritional content assessment results of the food to be evaluated. Acquire user health data, and based on the user health data and the nutritional component content assessment results, use collaborative filtering algorithms or content-based recommendation algorithms to obtain user-customized nutritional advice and health risk warnings; The nutrient content assessment results, nutritional recommendations, and health risk warnings are visualized through interactive charts and graphical interfaces.

6. The intelligent assessment method for food nutrient content based on a multimodal large model according to claim 1, characterized in that, The intelligent assessment method for food nutrient content based on a multimodal large model also includes: The system receives user feedback and uses it as a reference factor to update the parameters of the intermediate deep learning model during the process of updating the parameters of the intermediate deep learning model through incremental learning and memory enhancement mechanisms.

7. A smart food nutrient content assessment system based on a multimodal large model, wherein the smart food nutrient content assessment system based on a multimodal large model is applied to the smart food nutrient content assessment method based on a multimodal large model as described in any one of claims 1-6, characterized in that, The intelligent food nutrient content assessment system based on a multimodal large model includes: Data acquisition and preprocessing module: used to acquire multimodal data and preprocess the multimodal data to obtain target multimodal data; Multimodal large model fusion module: Used to perform feature extraction, feature alignment and feature fusion on the target multimodal data by using a pre-trained multimodal large model and combining an adaptive feature alignment algorithm and a cross-modal attention mechanism to obtain multimodal fused features; Deep learning model building module: used to build an initial deep learning model using deep learning algorithms, and to train the initial deep learning model to obtain an intermediate deep learning model; Continuous learning module: When a new food type appears, it updates the parameters of the intermediate deep learning model through incremental learning and memory enhancement mechanisms to obtain the target deep learning model that has learned the new knowledge corresponding to the new food type. Nutritional content assessment module: used to input the multimodal fusion features into the target deep learning model, and the target deep learning model outputs the nutritional content assessment results.

8. A terminal, characterized in that, The terminal includes: a memory, a processor, and a smart food nutrient content assessment program based on a multimodal large model stored in the memory and executable on the processor. When the smart food nutrient content assessment program based on a multimodal large model is executed by the processor, it implements the steps of the smart food nutrient content assessment method based on a multimodal large model as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a food nutrient content intelligent assessment program based on a multimodal large model. When the food nutrient content intelligent assessment program based on a multimodal large model is executed by a processor, it implements the steps of the food nutrient content intelligent assessment method based on a multimodal large model as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-modal knowledge generation method and device based on feedback enhancement

    CN117035074A

  • Multimodal model knowledge updating method based on knowledge representation and dynamic prompt

    CN118627610A