A method and system for healthy diet planning based on multimodal fusion

By acquiring food and user information through multimodal fusion technology and generating personalized dietary plans using large models, the problem of existing dietary planning systems being unable to meet personalized needs is solved, thus achieving precise health dietary management.

CN120636699BActive Publication Date: 2026-01-30HEALTH HOPE (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511014664.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2026-01-30
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing dietary planning systems cannot fully consider individual differences in age, gender, physical condition, eating habits, and nutritional needs, and cannot provide personalized dietary plans that match each person's health condition and lifestyle, making it difficult to meet the special dietary needs of different groups.

Method used

A health diet planning method based on multimodal fusion is adopted. By acquiring multimodal information about food (images, text, and voice), combined with the user's personal information and physical health status, a personalized diet plan is generated using large model reasoning, and the plan is optimized through a dynamically updated nutrition knowledge base.

Benefits of technology

It provides precise recipe recommendations, ingredient pairing suggestions, and dietary restrictions to help users achieve reasonable diet and health management, and meet the special nutritional needs and dietary restrictions of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636699B_ABST
    Figure CN120636699B_ABST
Patent Text Reader

Abstract

This application discloses a method and system for healthy dietary planning based on multimodal fusion. The method includes: acquiring multimodal information about food; extracting and fusing features from the multimodal information using multimodal fusion technology to obtain food identification results; constructing a user profile of the target user based on the user information of the target user, the user information including at least the target user's dietary needs and physical health status information; combining the food identification results and the user profile, using large-scale model inference to generate a dietary planning scheme suitable for the target user; optimizing the dietary planning scheme according to a dynamically updated nutrition knowledge base, and outputting the optimized dietary planning scheme to the target user. This application's solution can take into account the special nutritional needs and dietary restrictions of different users, providing accurate recipe recommendations, ingredient pairing suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of health management and artificial intelligence technology, and in particular to a health diet planning method and system based on multimodal fusion. Background Technology

[0002] Food is the material foundation of people's lives, and good eating habits can prevent various chronic diseases (such as obesity and diabetes). Food recommendations have a wide range of practical applications, such as recommending recipes based on nutritional balance and recommending medicinal diets for certain diseases. Existing dietary planning systems often use general dietary recommendations without fully considering individual differences in age, gender, physical condition, eating habits, and nutritional needs. They cannot provide personalized dietary plans that suit each person's health condition and lifestyle, and are unable to meet the special needs of different groups, such as the special dietary requirements of diabetic patients, hypertensive patients, athletes, and the elderly. Summary of the Invention

[0003] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a healthy diet planning method and system based on multimodal fusion to solve the technical problems such as the lack of personalization in existing diet planning.

[0004] In a first aspect, the present invention provides a healthy diet planning method based on multimodal fusion, the method comprising:

[0005] The multimodal information of food is acquired, and feature extraction and fusion of the multimodal information are performed based on multimodal fusion technology to obtain food recognition results;

[0006] A user profile of the target user is constructed based on the target user's user information, which includes at least the target user's dietary needs and physical health information.

[0007] Combining the food identification results with the user profile, a dietary planning scheme suitable for the target user is generated using large model inference;

[0008] The dietary plan is optimized based on a dynamically updated nutrition knowledge base, and the optimized dietary plan is output to the target user.

[0009] This application solution, based on the powerful learning and reasoning capabilities of a large model, combines users' personal information (such as age, gender, body indicators, health status, dietary habits, etc.) with food recognition results to tailor personalized healthy diet plans for users. It can also take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, food pairing suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0010] In one technical solution of the above-mentioned health diet planning method based on multimodal fusion, the multimodal information includes at least a food image, text description, and voice information of the food to be identified. The multimodal fusion technology is used to extract and fuse features from the multimodal information to obtain the food identification result, including:

[0011] A convolutional neural network is used to extract visual features from the food image;

[0012] Semantic information in the text description is obtained by parsing using a natural language processing model;

[0013] The speech information is identified using speech recognition technology to obtain the corresponding speech recognition results;

[0014] The visual features, semantic information, and speech recognition results are cross-modal aligned and fused to obtain the food recognition information.

[0015] This application utilizes multimodal fusion technology to combine information from multiple modalities such as images, text, and voice to more accurately and comprehensively identify food, thereby improving the accuracy and reliability of food identification.

[0016] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, the step of extracting and fusing features from the multimodal information based on multimodal fusion technology to obtain food recognition results further includes:

[0017] Obtain fused features based on the visual features, the semantic information, and the speech recognition results;

[0018] Based on the fusion features, match food description data from a preset food database;

[0019] The food identification results are corrected based on the food description data.

[0020] This application utilizes food description data from a pre-set food database to correct the food identification results, thereby further improving the accuracy of the food identification results.

[0021] In one technical solution of the above-mentioned multimodal fusion-based healthy diet planning method, the step of constructing the user profile of the target user based on the target user's user information includes:

[0022] Obtain the target user's dietary needs and physical health information;

[0023] Determine the appropriate dietary range for the target user based on the aforementioned health status information;

[0024] A user profile of the target user is generated based on the stated dietary needs, dietary range, and physical health information.

[0025] In one technical solution of the above-mentioned multimodal fusion-based healthy diet planning method, the step of combining the food identification results and the user profile to generate a diet planning scheme suitable for the target user using large model inference includes:

[0026] Based on the food recognition results and the user profile, multiple candidate food combinations are generated for the target user.

[0027] The nutritional balance index and health benefit score of each candidate food combination are calculated by combining the food identification results and the nutritional synergy between foods.

[0028] A dietary plan suitable for the target user is generated based on the nutritional balance index and the health benefit score.

[0029] In one technical solution of the above-mentioned multimodal fusion-based healthy diet planning method, generating multiple candidate food combinations for the target user based on the food identification results and the user profile includes:

[0030] Analyze the nutritional needs of the target users based on the user profile;

[0031] Based on the nutritional requirements and the food identification results, multiple candidate food combinations are generated that are adapted to the health goals of the target user.

[0032] In one of the above-mentioned technical solutions of the health diet planning method based on multimodal fusion, the health goal includes at least one of weight loss, muscle gain, and blood sugar control.

[0033] In one technical solution of the above-mentioned multimodal fusion-based healthy diet planning method, the dynamically updated nutrition knowledge base is implemented in the following way:

[0034] Real-time access to the latest nutritional research data, food databases, or user feedback;

[0035] The parameters and rule base of the large model are updated using incremental learning techniques.

[0036] This application proposes a nutrition knowledge base that is configured to have dynamic learning and updating capabilities. It can integrate the latest food nutrition research results, ingredient data, health cases, and other information in real time, and continuously optimize the rule base of the large model (such as the nutrition analysis rule base and dietary planning strategies). This enables the system to adapt to constantly changing scientific knowledge and user needs, and maintain the scientific and timely nature of dietary planning.

[0037] In one technical solution of the above-mentioned multimodal fusion-based healthy diet planning method, the method further includes:

[0038] The multimodal information is preprocessed, and the preprocessing includes at least data cleaning and data standardization.

[0039] The proposed solution, by preprocessing multimodal information, can remove invalid, redundant, or erroneous data from the multimodal information, thereby improving the accuracy of the information.

[0040] In a second aspect, the present invention provides a health diet planning system based on multimodal fusion, the system being used to implement the method as described in any one of the first aspects, the system comprising:

[0041] The food recognition module is used to acquire multimodal information about food, and to extract and fuse features from the multimodal information based on multimodal fusion technology to obtain food recognition results.

[0042] The user profile building module is used to build a user profile of the target user based on the user information of the target user, wherein the user information includes at least the target user’s dietary needs and physical health information.

[0043] The dietary planning module is used to combine the food identification results and the user profile, and use large model inference to generate a dietary planning scheme suitable for the target user.

[0044] The plan optimization module is used to optimize the dietary plan based on a dynamically updated nutrition knowledge base and output the optimized dietary plan to the target user.

[0045] In a third aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the computer program is executed by the processor, implements the multimodal fusion-based healthy diet planning method as described in any of the first aspects.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed, it implements the multimodal fusion-based healthy diet planning method as described in any one of the first aspects.

[0047] The above-described technical solutions of the present invention have at least one or more of the following beneficial effects:

[0048] In implementing the technical solution of this invention, based on the powerful learning and reasoning capabilities of the large model, combined with the user's personal information (such as age, gender, body indicators, health status, eating habits, etc.) and food recognition results, a personalized healthy diet plan is tailored for the user. It can also take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, food pairing suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0049] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0050] The disclosure of this invention will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. Furthermore, similar numbers in the drawings are used to denote similar components, wherein:

[0051] Figure 1 This is a flowchart of the healthy diet planning method based on multimodal fusion provided in Embodiment 1 of this application;

[0052] Figure 2 This is a schematic diagram of the structure of the health diet planning system based on multimodal fusion provided in Embodiment 2 of this application.

[0053] Figure 3 This is a schematic diagram of the structure of the computer device provided in Embodiment 4 of this application. Detailed Implementation

[0054] Some embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0055] As described in the background section, existing dietary planning systems often use general dietary recommendations without fully considering individual differences in age, gender, physical condition, eating habits, and nutritional needs. They cannot provide personalized dietary plans tailored to each person's health condition and lifestyle, and struggle to meet the specific needs of different groups, such as the special dietary requirements of diabetic patients, hypertensive patients, athletes, and the elderly.

[0056] Based on this, this application proposes a health diet planning method and system based on multimodal fusion. This method is based on the powerful learning and reasoning capabilities of large models, combined with users' personal information (such as age, gender, body indicators, health status, eating habits, etc.) and food recognition results, to tailor a personalized health diet plan for users. It can also take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, food pairing suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0057] Example 1

[0058] Figure 1 This is a flowchart of a multimodal fusion-based healthy diet planning method provided in Embodiment 1 of this application. This method is applicable to various scenarios, such as home diet management (e.g., uploading food images via mobile terminals to obtain healthy diet planning schemes), restaurant ordering assistance (identifying menu items and recommending combinations that meet the user's health needs), health management institutions (providing standardized or customized diet plans for special groups), and school canteens. It provides users with corresponding food recognition and diet planning services in different scenarios, meeting their healthy eating needs in various life situations, and has broad application prospects and social value.

[0059] Reference Figure 1 As shown, the method includes the following steps:

[0060] S100: Obtain multimodal information of food, extract and fuse features of the multimodal information based on multimodal fusion technology, and obtain food recognition results.

[0061] Traditional food identification methods may rely solely on single-modal information, such as image recognition or text descriptions, making it difficult to accurately identify food types, ingredients, cooking methods, and other information. For example, the same food prepared in different ways can appear very different, and image recognition alone may lead to misidentification; while relying solely on the food's name fails to capture its actual appearance features and details, resulting in inaccurate identification. Therefore, this application utilizes multimodal information about food to achieve more accurate and comprehensive food identification, improving the accuracy and reliability of food identification.

[0062] In this embodiment, the method of acquiring multimodal information is not limited. Without departing from the inventive concept of this application, the method of acquiring information for each modality can be set according to the actual application scenario. For example, for the application scenario of family diet management, food images can be acquired using the camera of a smart device, text descriptions can be text input by the target user, and voice information can be descriptions of food uploaded by the target user, etc., which will not be elaborated here.

[0063] S200: Construct a user profile of the target user based on the target user's user information, wherein the user information includes at least the target user's dietary needs and physical health information.

[0064] It should be noted that the dietary requirements in this application include, but are not limited to, the target user's dietary preferences, recommended foods suitable for the target user's consumption based on the target user's physical health condition, and restricted foods for the target user's consumption. Physical health information includes, but is not limited to, the target user's step length, heart rate, sleep quality, blood pressure, daily diet, exercise time, and sleep patterns.

[0065] It should be noted that the user information in this application embodiment includes not only the target user's dietary needs and physical health information, but also information such as age, gender, and eating habits, which will not be listed here.

[0066] By constructing user profiles based on user information such as the target users' dietary needs and physical health status, and then using these user profiles to generate dietary plans suitable for the target users, it is possible to better take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, food pairing suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0067] S300: Combining the food recognition results and the user profile, a dietary planning scheme suitable for the target user is generated using large model inference.

[0068] This application utilizes the powerful learning and reasoning capabilities of large models, combined with the user profiles of target users and food recognition results, to tailor personalized healthy diet plans for target users.

[0069] It should be noted that the large models in the embodiments of this application include, but are not limited to, multimodal large models. These models are capable of processing and understanding multiple modal information (such as text, images, video, and audio) and can achieve more complex tasks through cross-modal interaction. These models typically consist of multiple parts, including modal encoders, input projectors, language model backbones, output projectors, and modality generators. Their training process mainly includes two stages: multimodal pre-training and multimodal instruction fine-tuning.

[0070] Multimodal pre-training refers to machine learning methods that integrate data from multiple modalities (such as text, images, and speech) for pre-training, aiming to improve the model's generalization ability and performance in cross-modal tasks. For example, combining text and images can lead to a more comprehensive understanding of the scene, thereby improving the accuracy of downstream tasks (such as visual question answering and speech recognition). Visual instruction tuning (VIT) is a method that extends instruction tuning techniques to multimodal scenarios, aiming to enable large language models to perform complex tasks by combining visual information. For example, answering questions or performing classification based on image content. This process adjusts the model weights using a specific dataset (containing image and instruction pairs) to achieve synergy between vision and language.

[0071] S400: Optimize the dietary plan based on a dynamically updated nutrition knowledge base, and output the optimized dietary plan to the target user.

[0072] It should be noted that the proposed solution configures the nutrition knowledge base to have dynamic learning and updating capabilities, enabling it to integrate the latest food nutrition research results, ingredient data, health cases, and other information in real time. This allows for continuous optimization of the rule base of the large model (such as the nutrition analysis rule base and dietary planning strategies), enabling the system to adapt to constantly changing scientific knowledge and user needs, and maintaining the scientific rigor and timeliness of dietary planning.

[0073] In some specific embodiments, the multimodal information includes at least a food image, text description, and voice information of the food to be identified. The step of extracting and fusing features from the multimodal information based on multimodal fusion technology to obtain the food identification result includes:

[0074] A convolutional neural network is used to extract visual features from the food image;

[0075] Semantic information in the text description is obtained by parsing using a natural language processing model;

[0076] The speech information is identified and analyzed using speech recognition technology to obtain the corresponding speech recognition results;

[0077] The visual features, semantic information, and speech recognition results are cross-modal aligned and fused to obtain the food recognition information.

[0078] In some specific embodiments, visual feature extraction from food images is achieved using a convolutional neural network (CNN), including but not limited to the following steps:

[0079] 1. Input preprocessing

[0080] The acquired food images are standardized, including but not limited to size normalization (e.g., adjusting to 224×224 pixels), illumination compensation (e.g., histogram equalization), and noise filtering (e.g., Gaussian blurring), to improve the robustness of subsequent feature extraction.

[0081] Multi-layer convolution feature extraction

[0082] Image features are extracted step-by-step using convolutional layers of deep convolutional neural networks (such as pre-trained models like ResNet50 and EfficientNet). For example, for shallow features, basic visual features such as edges, colors, and textures (e.g., the color and surface texture of food) can be captured using low-level convolutional kernels (e.g., 3×3 or 5×5). For deep features, semantic-level features (e.g., the hierarchical structure and geometric shape of food) are extracted using high-level convolutional kernels combined with pooling operations (e.g., max pooling). Preferably, the convolutional layers can use the ReLU activation function and batch normalization (BatchNorm) to accelerate training convergence.

[0083] Feature mapping and dimensionality reduction

[0084] Global average pooling (GAP) is used to convert the feature map output from the last convolutional layer into a fixed-dimensional feature vector, preserving the statistical properties of spatial information. Optionally, principal component analysis (PCA) or fully connected layers are used for further dimensionality reduction to reduce computational redundancy.

[0085] In some specific embodiments, the semantic information in the text description can be realized through a natural language processing (NLP) model, including but not limited to the following steps:

[0086] 1. Text preprocessing and precoding

[0087] The input text description is standardized, including but not limited to word segmentation, stop word filtering, and stemming. The processed text is then converted into word embeddings, such as by using a pre-trained word vector model (e.g., Word2Vec, GloVe) or by directly using the sub-word embedding layer of the Transformer model (e.g., BERT's Tokenizer).

[0088] Contextual semantic modeling

[0089] Deep semantic features are extracted using Transformer-based pre-trained language models (such as BERT, RoBERTa, or domain-adapted FoodBERT), including: capturing the contextual associations of keywords in the text through a self-attention mechanism, and outputting a context-related vector representation of each word.

[0090] Key information structure extraction

[0091] Structured semantic information can be extracted using any of the following methods: named entity recognition, relation extraction, or sentiment analysis.

[0092] In some specific embodiments, the recognition of voice information can be achieved through voice recognition technology, including but not limited to the following steps:

[0093] Speech signal preprocessing

[0094] The input speech signal is denoised (e.g., environmental noise suppression based on spectral subtraction), framed, and windowed to improve the robustness of speech feature extraction. Acoustic features are extracted from the preprocessed speech signal to generate a temporal feature sequence.

[0095] Speech recognition model decoding

[0096] An end-to-end automatic speech recognition model (such as Conformer or Whisper) is used for phoneme-to-text conversion, including: capturing local and global temporal features of speech through convolutional layers and temporal attention mechanisms (such as Conformer Block); optimizing the decoding results by combining pre-trained language models (such as N-gram or neural language models) to resolve homonym ambiguity. Finally, the text transcription result with timestamps is output.

[0097] Semantic post-processing and error correction

[0098] Error correction is performed based on relevant domain knowledge bases (such as food databases), and colloquial expressions are standardized to improve the accuracy of speech recognition.

[0099] In some specific embodiments, when performing cross-modal alignment and fusion of the visual features, semantic information, and speech recognition results, the visual features, semantic information, and speech recognition results can first be mapped to a unified semantic space. For example, a shared projection matrix can be used to reduce the dimensionality of each modality feature, and cross-modal contrastive learning can constrain matching modal features to be spatially adjacent, while non-matching features are far away. Then, the confidence score of each modality (such as image classification probability, speech recognition likelihood) is calculated, and the fusion weight is calculated based on the confidence score of each modality. Then, a multimodal Transformer layer is used to realize fine-grained feature interaction. For example, for visual-text interaction, key regions can be located by calculating the attention weights of image regions and text keywords; for speech-image correction, when the information of speech description and image detection conflicts, the conflict detection module is triggered to reweight the evidence, etc. Finally, the fused multimodal features are input into the multi-layer perceptual layer of a large model for inference and prediction to obtain food recognition information.

[0100] In some specific embodiments, the step of extracting and fusing features from the multimodal information based on multimodal fusion technology to obtain food recognition results further includes:

[0101] Obtain fused features based on the visual features, the semantic information, and the speech recognition results;

[0102] Based on the fusion features, match food description data from a preset food database;

[0103] The food identification results are corrected based on the food description data.

[0104] Specifically, after obtaining the fused features of visual features, semantic information, and speech recognition results, the comprehensive similarity between these fused features and candidate foods in a pre-defined food database can be calculated. A clustering algorithm (such as k-means) is then used to quickly filter out the top-ranked (e.g., Top-50) candidate foods. Multi-dimensional matching of these candidate foods is then performed, including but not limited to component consistency verification and nutritional rule filtering, outputting the food with the highest matching degree and its corresponding food description data. Finally, the food recognition results are corrected based on this food description data to further improve the accuracy of the food recognition results.

[0105] In some specific embodiments, constructing the user profile of the target user based on the target user's user information includes:

[0106] Obtain the target user's dietary needs and physical health information;

[0107] Determine the appropriate dietary range for the target user based on the aforementioned health status information;

[0108] A user profile of the target user is generated based on the stated dietary needs, dietary range, and physical health information.

[0109] It's important to note that by creating user profiles based on the target users' dietary needs and health conditions, a balance can be achieved between individual dietary preferences and nutritional needs from the user's perspective. This considers both the user's food preferences and their overall physical condition, recommending foods they enjoy while meeting their nutritional requirements. Furthermore, it takes into account the complex and variable health information of different users, such as whether they have diabetes or food allergies.

[0110] In some specific embodiments, the step of combining the food identification results and the user profile to generate a suitable dietary planning scheme for the target user using large model inference includes:

[0111] Based on the food recognition results and the user profile, multiple candidate food combinations are generated for the target user.

[0112] The nutritional balance index and health benefit score of each candidate food combination are calculated by combining the food identification results and the nutritional synergy between foods.

[0113] A dietary plan suitable for the target user is generated based on the nutritional balance index and the health benefit score.

[0114] Specifically, when calculating the nutritional balance index, core nutritional data (such as macronutrients, micronutrients, functional components, and health risk indicators) of each food can be obtained based on food identification results. Combining this with the synergistic effects between foods, a pre-established multidimensional scoring system is used to calculate the nutritional balance index and health benefit score for each candidate food combination. Preferably, the scoring parameters can be dynamically adjusted based on the user's health status information. For example, for diabetic patients, the carbohydrate weight can be reduced and the dietary fiber score increased; for fitness enthusiasts, the proportion of protein and BCAA amino acid scores can be increased.

[0115] It should be noted that when generating dietary plans for target users, this application not only considers the natural attributes of food itself, such as its nutritional components, calories, appearance, smell, and taste, but also the synergistic effects of nutrients between foods. This allows for a more accurate and profound understanding of food types and other semantic levels, thereby more effectively completing the dietary planning work.

[0116] In some specific embodiments, generating multiple candidate food combinations for the target user based on the food recognition results and the user profile includes:

[0117] Analyze the nutritional needs of the target users based on the user profile;

[0118] Based on the nutritional requirements and the food identification results, multiple candidate food combinations are generated that are adapted to the health goals of the target user.

[0119] Specifically, a deep analysis of user profiles is conducted to examine the nutritional needs of target users, including but not limited to basal metabolic parameters, disease management needs, nutritional gap detection, and personal preference constraints. Finally, based on the food identification results and the target users' nutritional needs and health goals, multiple suitable candidate food combinations are generated for each user.

[0120] In some specific embodiments, health goals include at least one of weight loss, muscle gain, and blood sugar control.

[0121] In some specific embodiments, the dynamically updated nutrition knowledge base is implemented in the following ways:

[0122] Real-time access to the latest nutritional research data, food databases, or user feedback;

[0123] The parameters and rule base of the large model are updated using incremental learning techniques.

[0124] It should be noted that this application does not limit the specific access methods for the latest nutritional research data, food databases, or user feedback. Without violating the inventive concept of this application, the method can be selected according to actual needs. For example, it can establish a connection with an authoritative nutritional database through an API interface to regularly capture the latest research results, or use natural language processing (NLP) to parse academic papers, automatically extract key conclusions, and convert them into calculable nutritional rules.

[0125] In some specific embodiments, the method further includes:

[0126] The multimodal information is preprocessed, and the preprocessing includes at least data cleaning and data standardization.

[0127] As an illustrative and not limiting explanation, the cleaning of food images in this application includes, but is not limited to, removing low-quality images, performing background segmentation and food region extraction, or outlier filtering (e.g., automatically removing interfering regions using a semantic segmentation model when non-food objects are identified). The cleaning of text descriptions includes, but is not limited to, error correction and redundant information removal. The cleaning of speech information includes, but is not limited to, environmental noise suppression and spoken language normalization.

[0128] In some specific embodiments, data standardization processing for multimodal information includes, but is not limited to, feature dimension unification processing and semantic alignment preprocessing. For example, feature dimension unification processing for food images includes scaling images of different resolutions to a preset pixel size or normalizing the RGB channels of the image. Feature dimension unification processing for text descriptions includes using a word segmenter to uniformly truncate / paste to a preset word length. Feature dimension unification processing for speech information includes unifying MFCC features to a preset dimension / frame and cyclically padding with zeros when the duration is insufficient.

[0129] Example 2

[0130] Corresponding to Embodiment 1 above, this application also provides a health diet planning system based on multimodal fusion. This system is used to implement the health diet planning method based on multimodal fusion provided in any of Embodiment 1. In this embodiment, content that is the same as or similar to that in Embodiment 1 above can be referred to the above description and will not be repeated hereafter. (Refer to...) Figure 2As shown, the system includes:

[0131] The food recognition module 10 is used to acquire multimodal information about food, and to extract and fuse features from the multimodal information based on multimodal fusion technology to obtain food recognition results;

[0132] User profile building module 20 is used to build a user profile of the target user based on the user information of the target user, wherein the user information includes at least the target user’s dietary needs and physical health information.

[0133] The dietary planning module 30 is used to combine the food identification results and the user profile to generate a dietary planning scheme suitable for the target user using large model reasoning.

[0134] The scheme optimization module 40 is used to optimize the dietary planning scheme based on a dynamically updated nutrition knowledge base, and output the optimized dietary planning scheme to the target user.

[0135] In some specific embodiments, the food recognition module 10 is specifically used for:

[0136] Visual features in the food image are extracted using a convolutional neural network; semantic information in the text description is obtained by parsing using a natural language processing model; speech recognition technology is used to identify and parse the speech information to obtain the corresponding speech recognition result; the visual features, the semantic information, and the speech recognition result are cross-modal aligned and fused to obtain the food recognition information.

[0137] In some specific embodiments, the food recognition module 10 is further used for:

[0138] Obtain fusion features based on the visual features, the semantic information, and the speech recognition results; match food description data from a preset food database according to the fusion features; and correct the food recognition results according to the food description data.

[0139] In some specific embodiments, the user profile building module 20 is specifically used for:

[0140] Obtain the dietary needs and physical health information of the target user; determine the dietary range suitable for the target user based on the physical health information; generate a user profile of the target user based on the dietary needs, the dietary range, and the physical health information.

[0141] In some specific embodiments, the diet planning module 30 is specifically used for:

[0142] Based on the food identification results and the user profile, multiple candidate food combinations are generated for the target user; the nutritional balance index and health benefit score of each candidate food combination are calculated by combining the food identification results and the nutritional synergy between foods; and a dietary plan suitable for the target user is generated based on the nutritional balance index and the health benefit score.

[0143] In some specific embodiments, the diet planning module 30 is further used for:

[0144] The nutritional needs of the target user are analyzed based on the user profile; and multiple candidate food combinations that match the health goals of the target user are generated by combining the nutritional needs with the food identification results.

[0145] In some specific embodiments, the health goal includes at least one of weight loss, muscle gain, and blood sugar control.

[0146] In some specific embodiments, the scheme optimization module 40 is specifically used for:

[0147] It accesses the latest nutritional research data, food databases, or user feedback in real time; and uses incremental learning technology to update the parameters and rule base of the large model.

[0148] In some specific embodiments, the food recognition module 10 is further used for:

[0149] The multimodal information is preprocessed, and the preprocessing includes at least data cleaning and data standardization.

[0150] Example 3

[0151] Corresponding to Embodiment 1 above, this application also provides a computer device, including: a processor and a memory, wherein the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it executes the health diet planning method based on multimodal fusion provided in any of the above embodiments.

[0152] in, Figure 3 An exemplary computer device 1500 is shown, which may specifically include a processor 1510, a video display adapter 1511, a disk drive 1512, an input / output interface 1513, a network interface 1514, and a memory 1520. The processor 1510, video display adapter 1511, disk drive 1512, input / output interface 1513, network interface 1514, and memory 1520 can communicate with each other via a communication bus 1530.

[0153] The processor 1510 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided by the present invention.

[0154] The memory 1520 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1520 can store the operating system 1521 for controlling the operation of the electronic device, and the basic input / output system (BIOS) for controlling the low-level operations of the electronic device. Additionally, it can store a web browser 1523, a data storage management system 1524, and a device identification information processing system 1525, etc. The aforementioned device identification information processing system 1525 can be the application program that specifically implements the aforementioned steps in this embodiment of the invention. In summary, when implementing the technical solution provided by this invention through software or firmware, the relevant program code is stored in the memory 1520 and is called and executed by the processor 1510.

[0155] Input / output interface 1513 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0156] Network interface 1514 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0157] The bus includes a pathway for transmitting information between various components of the device, such as processor 1510, video display adapter 1511, disk drive 1512, input / output interface 1513, network interface 1514, and memory 1520.

[0158] In addition, the electronic device can also obtain information on specific claim conditions from the virtual resource object claim condition information database for condition judgment, and so on.

[0159] It should be noted that although the above-described device only shows the processor 1510, video display adapter 1511, disk drive 1512, input / output interface 1513, network interface 1514, memory 1520, bus, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the present invention, and not necessarily all the components shown in the figures.

[0160] Example 4

[0161] Corresponding to Embodiment 1 above, this application also provides a computer-readable storage medium. In this embodiment, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description and will not be repeated hereafter.

[0162] The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the multimodal fusion-based healthy diet planning method as described above.

[0163] In some implementations of this application, when the computer program is executed by the processor, it can also implement the steps corresponding to the method described in Embodiment 1. Please refer to the detailed description in Embodiment 1, which will not be repeated here.

[0164] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0165] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0166] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0167] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0168] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0169] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A health diet planning method based on multi-modal fusion, characterized in that, The method comprises: acquiring multi-modal information of food, extracting and fusing features of the multi-modal information based on multi-modal fusion technology, and obtaining food recognition results; constructing a user portrait of a target user based on user information of the target user, the user information at least including dietary needs and physical health status information of the target user; combining the food recognition results and the user portrait, and generating a meal planning scheme suitable for the target user by using a large model inference; optimizing the meal planning scheme based on a dynamically updated nutrition knowledge base, and outputting the optimized meal planning scheme to the target user; wherein the multi-modal information at least includes food images, text descriptions, and voice information of food to be identified, the feature extraction and fusion of the multi-modal information based on the multi-modal fusion technology to obtain food recognition results comprising: extracting visual features in the food images using a convolutional neural network; analyzing and obtaining semantic information in the text description through a natural language processing model; identifying and analyzing the voice information using voice recognition technology to obtain corresponding voice recognition results; aligning and fusing the visual features, the semantic information, and the voice recognition results across modalities to obtain the food recognition results. 2.The health diet planning method based on multi-modal fusion according to claim 1, characterized in that, The feature extraction and fusion of the multi-modal information based on the multi-modal fusion technology to obtain food recognition results further comprises: obtaining fusion features based on the visual features, the semantic information, and the voice recognition results; matching food description data in a pre-set food database according to the fusion features; correcting the food recognition results according to the food description data. 3.The health diet planning method based on multi-modal fusion according to claim 1, characterized in that, The construction of the user portrait of the target user based on the user information of the target user comprises: acquiring dietary needs and physical health status information of the target user; determining a dietary range suitable for the target user according to the physical health status information; generating a user portrait of the target user according to the dietary needs, the dietary range, and the physical health status information. 4.The health diet planning method based on multi-modal fusion according to claim 1, wherein, The combination of the food recognition results and the user portrait to generate a meal planning scheme suitable for the target user by using a large model inference comprises: generating multiple candidate food material combinations for the target user based on the food recognition results and the user portrait; calculating a nutritional balance index and a health benefit score of each candidate food material combination in combination with the food recognition results and the nutritional synergistic effect among foods; generating a meal planning scheme suitable for the target user according to the nutritional balance index and the health benefit score. 5.The health diet planning method based on multi-modal fusion according to claim 4, characterized in that, The generation of multiple candidate food material combinations for the target user based on the food recognition results and the user portrait comprises: analyzing the nutritional needs of the target user according to the user portrait; combining the nutritional needs and the food recognition results to generate multiple candidate food material combinations adapted to the health goals of the target user. 6.The health diet planning method based on multi-modal fusion according to claim 5, characterized in that, The health goals include at least one of weight loss, muscle gain, and blood sugar control.

7. The health diet planning method based on multi-modal fusion according to any one of claims 1 to 6, characterized in that, The dynamically updated nutrition knowledge base is achieved by the following means: accessing the latest nutrition research data, food database or user feedback in real time; updating the parameters and rule base of the large model using incremental learning techniques. 8.The health diet planning method based on multi-modal fusion according to any one of claims 1 to 6, characterized in that, The method further comprises: preprocessing the multi-modal information, including at least data cleaning and data standardization. 9.A health diet planning system based on multi-modal fusion, characterized in that, The system is used to implement the method of any one of claims 1-8, and the system comprises: a food recognition module for obtaining multi-modal information of food, extracting and fusing features of the multi-modal information based on multi-modal fusion technology, and obtaining food recognition results; a user portrait construction module for constructing a user portrait of a target user based on user information of the target user, the user information including at least dietary needs and physical health status information of the target user; a meal planning module for combining the food recognition results and the user portrait, and generating a meal planning scheme suitable for the target user using a large model inference; a scheme optimization module for optimizing the meal planning scheme according to a dynamically updated nutrition knowledge base, and outputting the optimized meal planning scheme to the target user; wherein the multi-modal information includes at least food images, text descriptions and voice information of the food to be recognized, and the feature extraction and fusion of the multi-modal information based on multi-modal fusion technology to obtain food recognition results comprises: extracting visual features in the food images using a convolutional neural network; analyzing and obtaining semantic information in the text description through a natural language processing model; recognizing and analyzing the voice information using voice recognition technology to obtain corresponding voice recognition results; aligning and fusing the visual features, semantic information and voice recognition results across modalities to obtain the food recognition results.

Citation Information

Patent Citations

  • Intelligent nutritional diet recommendation method based on multi-source heterogeneous data fusion

    CN112614564A

  • Method and apparatus for identifying food using multimodal model

    EP4557237A1