Healthy diet planning method and system based on multi-modal fusion

Through multimodal fusion technology and large models, combined with user information and food recognition results, personalized meal planning plans are generated, which solves the problem that existing meal planning systems cannot meet individual differences and realizes accurate healthy meal management.

CN120636699AActive Publication Date: 2025-09-12HEALTH HOPE (BEIJING) TECH CO LTD

Patent Information

Application Number
CN202511014664.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-09-12
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

The existing dietary planning system cannot fully consider individual differences in age, gender, physical condition, eating habits, nutritional needs, etc., and cannot provide each person with a personalized dietary plan that suits their own health status and lifestyle, and it is difficult to meet the special needs of different groups of people.

Method used

A healthy diet planning method based on multimodal fusion is adopted. By obtaining multimodal information of food (such as images, text, and voice), combined with the user's personal information and health status, a large model is used to generate personalized diet planning plans, and the plans are optimized through a dynamically updated nutrition knowledge base to provide accurate recipe recommendations and dietary advice.

Benefits of technology

It has achieved personalized healthy meal planning tailored for users, taking into account the special nutritional needs and dietary restrictions of different users, providing accurate recipe recommendations and ingredient pairing suggestions to help users achieve reasonable diet and health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636699A_ABST
    Figure CN120636699A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a healthy diet planning method and system based on multi-modal fusion, and the method comprises the steps: obtaining the multi-modal information of food, carrying out the feature extraction and fusion of the multi-modal information based on a multi-modal fusion technology, and obtaining a food recognition result; constructing a user portrait of a target user based on user information of the target user, wherein the user information at least comprises diet requirements and body health condition information of the target user; in combination with the food recognition result and the user portrait, generating a diet planning scheme suitable for the target user by utilizing large model reasoning; and optimizing the dietary planning scheme according to the dynamically updated nutrition knowledge base, and outputting the optimized dietary planning scheme to the target user. According to the scheme, special nutritional requirements and diet restrictions of different users can be considered, accurate recipe recommendation, food material matching suggestions, diet taboo and the like are provided, and the users are helped to realize reasonable diet and health management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of health management and artificial intelligence technology, and in particular to a healthy diet planning method and system based on multimodal fusion. Background Art

[0002] Food is the material basis of people's lives, and good eating habits can prevent various chronic diseases (such as obesity, diabetes, etc.). Food recommendations have a very wide range of practical applications, such as recommending recipes based on nutritional combinations, recommending medicinal foods for certain diseases, etc. Existing dietary planning systems often use general dietary recommendations, which do not fully take into account individual differences in age, gender, physical condition, eating habits, nutritional needs, etc. It is impossible to provide each person with a personalized dietary plan that suits their own health status and lifestyle, and it is difficult to meet the special needs of different groups of people, such as diabetics, hypertensive patients, athletes, the elderly, etc. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a healthy diet planning method and system based on multimodal fusion to address technical problems such as the lack of personalization in diet planning in the prior art.

[0004] In a first aspect, the present invention provides a healthy diet planning method based on multimodal fusion, the method comprising: Acquiring multimodal information of food, performing feature extraction and fusion on the multimodal information based on multimodal fusion technology, and obtaining food recognition results; Building a user profile of the target user based on the user information of the target user, wherein the user information at least includes the dietary needs and physical health information of the target user; Combining the food recognition results and the user profile, using large model reasoning to generate a meal plan suitable for the target user; The dietary planning scheme is optimized according to the dynamically updated nutritional knowledge base, and the optimized dietary planning scheme is output to the target user.

[0005] This application solution, based on the powerful learning and reasoning capabilities of the large model, combines the user's personal information (such as age, gender, physical indicators, health status, eating habits, etc.) and food recognition results to tailor personalized healthy meal plans for users. It can also take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, ingredient combination suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0006] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, the multimodal information includes at least a food image, a text description, and voice information of the food to be identified. The multimodal fusion technology is used to extract and fuse the multimodal information to obtain the food identification result, which includes: extracting visual features from the food image using a convolutional neural network; Obtaining semantic information from the text description through natural language processing model analysis; Recognize the voice information using voice recognition technology to obtain corresponding voice recognition results; The visual features, the semantic information, and the speech recognition results are cross-modally aligned and fused to obtain the food recognition information.

[0007] This application scheme uses multimodal fusion technology to combine information from multiple modalities such as images, text, and voice to more accurately and comprehensively identify food, thereby improving the accuracy and reliability of food identification.

[0008] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, the step of extracting and fusing features of the multimodal information based on the multimodal fusion technology to obtain food recognition results further includes: Acquire a fusion feature based on the visual feature, the semantic information, and the speech recognition result; matching food description data in a preset food database according to the fusion feature; The food recognition result is corrected according to the food description data.

[0009] The present application scheme utilizes the food description data in the preset food database to correct the food recognition results, thereby further improving the accuracy of the food recognition results.

[0010] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, constructing a user profile of the target user based on the user information of the target user includes: Obtaining dietary needs and health status information of the target user; Determining a diet range suitable for the target user based on the health status information; A user profile of the target user is generated based on the dietary needs, the dietary range, and the physical health status information.

[0011] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, combining the food recognition results and the user profile and using large model reasoning to generate a diet plan suitable for the target user includes: generating a plurality of candidate food combinations for the target user based on the food recognition result and the user portrait; Calculating the nutritional balance index and health benefit score of each candidate food combination based on the food identification results and the nutritional synergy between foods; A dietary planning program suitable for the target user is generated based on the nutritional balance index and the health benefit score.

[0012] In a technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, generating multiple candidate ingredient combinations for the target user based on the food recognition result and the user portrait includes: Analyzing the nutritional needs of the target user based on the user portrait; Combining the nutritional requirements and the food identification results, a plurality of candidate food combinations that are compatible with the health goals of the target user are generated.

[0013] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, the health goal includes at least one of weight loss, muscle gain, and blood sugar control.

[0014] In one technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, the dynamically updated nutrition knowledge base is implemented in the following manner: Real-time access to the latest nutrition research data, food ingredient databases or user feedback; Incremental learning technology is used to update the parameters and rule base of the large model.

[0015] This application plan configures the nutrition knowledge base to have the ability to dynamically learn and update, and can integrate the latest food nutrition research results, ingredient data, health cases and other information in real time, and continuously optimize the rule base of the large model (such as the nutrition analysis rule base and dietary planning strategies, etc.), so that the system can adapt to the ever-changing scientific knowledge and user needs, and maintain the scientific nature and timeliness of dietary planning.

[0016] In a technical solution of the above-mentioned healthy diet planning method based on multimodal fusion, the method further includes: The multimodal information is preprocessed, where the preprocessing includes at least data cleaning and data standardization.

[0017] The present application solution can remove invalid, redundant or erroneous data from the multimodal information by preprocessing the multimodal information, thereby improving the accuracy of the information.

[0018] In a second aspect, the present invention provides a healthy diet planning system based on multimodal fusion, the system being configured to implement the method described in any one of the first aspects, the system comprising: A food recognition module is used to obtain multimodal information of food, extract and fuse features of the multimodal information based on multimodal fusion technology, and obtain food recognition results; A user portrait building module, configured to build a user portrait of a target user based on the user information of the target user, wherein the user information at least includes the dietary needs and physical health information of the target user; A meal planning module is used to combine the food recognition results and the user profile and use large model reasoning to generate a meal planning plan suitable for the target user; A plan optimization module is used to optimize the meal planning plan according to the dynamically updated nutrition knowledge base and output the optimized meal planning plan to the target user.

[0019] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the healthy diet planning method based on multimodal fusion as described in any one of the first aspects is implemented.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, characterized in that when the computer program is executed, it implements the healthy diet planning method based on multimodal fusion as described in any one of the first aspects.

[0021] The above one or more technical solutions of the present invention have at least one or more of the following beneficial effects: In implementing the technical solution of the present invention, based on the powerful learning and reasoning capabilities of the large model, combined with the user's personal information (such as age, gender, physical indicators, health status, eating habits, etc.) and food recognition results, a personalized healthy meal plan is tailored for the user. It can also take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, ingredient matching suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0022] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The disclosure of the present invention will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Furthermore, similar numbers in the drawings represent similar components, wherein: Figure 1 This is a flow chart of a healthy diet planning method based on multimodal fusion provided in Example 1 of the present application; Figure 2 This is a structural diagram of the healthy diet planning system based on multimodal fusion provided in Example 2 of the present application.

[0024] Figure 3 This is a structural diagram of the computer device provided in Example 4 of the present application. DETAILED DESCRIPTION

[0025] Some embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0026] As mentioned in the background, existing meal planning systems often use generic dietary recommendations that fail to fully account for individual differences in age, gender, physical condition, eating habits, and nutritional needs. These systems are unable to provide personalized meal plans tailored to each individual's health and lifestyle, and struggle to meet the specific dietary needs of different populations, such as those with diabetes, hypertension, athletes, and the elderly.

[0027] Based on this, this application proposes a healthy meal planning method and system based on multimodal fusion. This method is based on the powerful learning and reasoning capabilities of large models, combined with the user's personal information (such as age, gender, physical indicators, health status, eating habits, etc.) and food recognition results, to tailor personalized healthy meal plans for users. It can also take into account the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, ingredient matching suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0028] Example 1 Figure 1This is a flowchart of a healthy diet planning method based on multimodal fusion, provided in Example 1 of this application. This method is applicable to a variety of application scenarios, such as family diet management (e.g., taking food images through a mobile terminal and uploading them to obtain healthy diet planning solutions), restaurant ordering assistance (identifying menu items and recommending combinations that meet the user's health needs), health management agencies (providing standardized or customized diet plans for special groups of users), and school cafeterias. In different scenarios, it provides users with corresponding food identification and diet planning services to meet their healthy diet needs in various life scenarios, and has broad application prospects and social value.

[0029] Reference Figure 1 As shown, the method includes the following steps: S100: Acquire multimodal information of food, perform feature extraction and fusion on the multimodal information based on multimodal fusion technology, and obtain food recognition results.

[0030] Traditional food recognition methods may rely solely on a single modality, such as image recognition or text descriptions, making it difficult to accurately identify food types, ingredients, cooking methods, and other information. For example, the appearance of the same food cooked in different ways may vary significantly, and image recognition alone may lead to misjudgment. Relying solely on the text of the food name cannot capture the actual appearance characteristics and details of the food, resulting in inaccurate recognition. To this end, this application uses multimodal information about food to more accurately and comprehensively identify food, thereby improving the accuracy and reliability of food recognition.

[0031] In the embodiments of this application, the method for acquiring multimodal information is not limited. Without violating the inventive concept of this application, the method for acquiring information of each modality can be set according to the actual application scenario. For example, for the application scenario of family diet management, the camera of the smart device can be used to capture food images, etc. The text description can be text entered by the target user, and the voice information can be a description of the food uploaded by the target user, etc., which will not be detailed here.

[0032] S200: Constructing a user profile of the target user based on user information of the target user, where the user information at least includes dietary needs and physical health information of the target user.

[0033] It should be noted that the dietary requirements in this application include, but are not limited to, the target user's dietary preferences, as well as recommended ingredients suitable for the target user and restricted ingredients based on the target user's health status. Health status information includes, but is not limited to, the target user's step length, heart rate, sleep quality, blood pressure, daily diet, exercise time, sleep status, and other information.

[0034] It should be noted that the user information in the embodiment of the present application not only includes the target user's dietary needs and physical health information, but also includes age, gender, eating habits and other information, which will not be listed here one by one.

[0035] By building a user profile based on user information such as the target user's dietary needs and physical health status, and then combining this user profile to generate a meal planning plan suitable for the target user, we can better consider the special nutritional needs and dietary restrictions of different users, and provide accurate recipe recommendations, ingredient pairing suggestions, dietary taboos, etc., to help users achieve reasonable diet and health management.

[0036] S300: Combining the food recognition result and the user portrait, a meal planning plan suitable for the target user is generated using large model reasoning.

[0037] This application plan utilizes the powerful learning and reasoning capabilities of the large model, combined with the user portrait of the target user and the food recognition results, to tailor a personalized healthy meal plan for the target user.

[0038] It should be noted that the large models in the embodiments of the present application include but are not limited to multimodal large models (Multimodal Large Models), which can process and understand multiple modal information (such as text, images, video and audio) and can achieve more complex tasks through cross-modal interaction. These models are usually composed of multiple parts, including modal encoders, input projectors, language model backbones, output projectors and modal generators. Their training process mainly includes two stages: multimodal pre-training and multimodal instruction fine-tuning.

[0039] Among them, multimodal pre-training refers to a machine learning method that integrates multimodal data (such as text, images, and speech) for pre-training, aiming to improve the model's generalization ability and performance in cross-modal tasks. For example, combining text and images can provide a more comprehensive understanding of the scene, thereby improving the accuracy of downstream tasks (such as visual question answering and speech recognition). Multimodal instruction tuning refers to a method that extends instruction fine-tuning technology to multimodal scenarios, aiming to enable large language models to perform complex tasks in combination with visual information. For example, answer questions or perform classification based on image content. This process adjusts the model weights through a specific dataset (containing image and instruction pairs) to achieve synergy between vision and language.

[0040] S400: Optimizing the dietary planning scheme according to the dynamically updated nutritional knowledge base, and outputting the optimized dietary planning scheme to the target user.

[0041] It should be noted that this application scheme configures the nutrition knowledge base to have the ability to dynamically learn and update, so that it can integrate the latest food nutrition research results, ingredient data, health cases and other information in real time, and continuously optimize the rule base of the large model (such as the nutrition analysis rule base and dietary planning strategy, etc.), so that the system can adapt to the ever-changing scientific knowledge and user needs, and maintain the scientific nature and timeliness of dietary planning.

[0042] In some specific embodiments, the multimodal information includes at least a food image, a text description, and voice information of the food to be identified, and the step of extracting and fusing features of the multimodal information based on the multimodal fusion technology to obtain a food identification result includes: extracting visual features from the food image using a convolutional neural network; Obtaining semantic information from the text description through natural language processing model analysis; Using speech recognition technology to identify and analyze the speech information to obtain corresponding speech recognition results; The visual features, the semantic information, and the speech recognition results are cross-modally aligned and fused to obtain the food recognition information.

[0043] In some specific embodiments, visual feature extraction from food images is implemented using a convolutional neural network (CNN), including but not limited to the following steps: 1. Input preprocessing The collected food images are standardized, including but not limited to size normalization (such as adjusting to 224×224 pixels), illumination compensation (such as histogram equalization), and noise filtering (such as Gaussian blur), to improve the robustness of subsequent feature extraction.

[0044] Multi-layer convolution feature extraction Image features are extracted layer by layer using the convolutional layers of a deep convolutional neural network (such as pre-trained models like ResNet50 and EfficientNet). For shallow features, for example, low-level convolution kernels (such as 3×3 or 5×5) can be used to capture basic visual features such as edges, color, and texture (e.g., the color and surface texture of food). For deeper features, semantic features (e.g., the hierarchical structure and geometric shape of food) can be extracted using high-level convolution kernels combined with pooling operations (such as max pooling). Preferably, the convolution layers use the ReLU activation function and batch normalization (BatchNorm) to accelerate training convergence.

[0045] Feature Mapping and Dimensionality Reduction Global average pooling (GAP) is used to convert the feature map output by the last convolution layer into a fixed-dimensional feature vector, preserving the statistical properties of the spatial information. Optionally, principal component analysis (PCA) or a fully connected layer is used to further reduce the dimensionality and reduce computational redundancy.

[0046] In some specific embodiments, the semantic information in the text description can be realized through a natural language processing (NLP) model, including but not limited to the following steps: 1. Text preprocessing and precoding The input text description is standardized, including but not limited to word segmentation, stop word filtering, and stemming. The processed text is converted into a word embedding vector (Word Embedding), such as using a pre-trained word vector model (such as Word2Vec, GloVe) or directly using the subword embedding layer of the Transformer model (such as BERT's Tokenizer).

[0047] Contextual semantic modeling Use a Transformer-based pre-trained language model (such as BERT, RoBERTa, or the domain-adapted FoodBERT) to extract deep semantic features, including capturing the contextual associations of keywords in the text through the self-attention mechanism and outputting a context-dependent vector representation of each token.

[0048] Structured extraction of key information Extract structured semantic information through any of named entity recognition, relation extraction, or sentiment analysis.

[0049] In some specific embodiments, the recognition of voice information can be achieved through voice recognition technology, including but not limited to the following steps: Speech signal preprocessing The input speech signal is subjected to noise reduction (e.g., spectral subtraction-based ambient noise suppression), framing, and windowing to improve the robustness of speech feature extraction. The acoustic features of the preprocessed speech signal are extracted to generate a time series feature sequence.

[0050] Speech recognition model decoding End-to-end automatic speech recognition models (such as Conformer and Whisper) are used for phoneme-to-text conversion. This involves capturing local and global temporal features of speech through convolutional layers and temporal attention mechanisms (such as the Conformer Block). Pre-trained language models (such as N-grams or neural language models) are then combined to optimize decoding results and resolve homophone ambiguity. Finally, the resulting text transcript is output with a timestamp.

[0051] Semantic post-processing and error correction Improve the accuracy of speech recognition by correcting errors based on relevant domain knowledge bases (such as food databases) and normalizing spoken expressions.

[0052] In some specific embodiments, when cross-modally aligning and fusing the visual features, semantic information, and speech recognition results, the visual features, semantic information, and speech recognition results can first be mapped into a unified semantic space. For example, a shared projection matrix can be used to reduce the dimensionality of each modal feature. Cross-modal contrastive learning can be used to constrain matching modal features to be spatially adjacent and non-matching features to be distant. Confidence scores (e.g., image classification probability, speech recognition likelihood) are then calculated for each modality, and fusion weights are calculated based on the confidence scores. A multimodal Transformer layer is then used to implement fine-grained feature interaction. For example, for visual-text interaction, key regions can be located by calculating attention weights between image regions and text keywords. For speech-image deflection correction, when conflicting information between the speech description and image detection is detected, the conflict detection module is triggered to reweight the evidence. Finally, the fused multimodal features are input into the multi-layer perception layer of the larger model for inference and prediction, yielding food identification information.

[0053] In some specific embodiments, extracting and fusing features of the multimodal information based on the multimodal fusion technology to obtain food recognition results further includes: Acquire a fusion feature based on the visual feature, the semantic information, and the speech recognition result; matching food description data in a preset food database according to the fusion feature; The food recognition result is corrected according to the food description data.

[0054] Specifically, after obtaining a fusion feature of visual features, semantic information, and speech recognition results, the system calculates the comprehensive similarity between this fusion feature and candidate foods in a preset food database. Using a clustering algorithm (such as k-means), the system quickly screens the top-ranked (e.g., top 50) candidate foods. The system then performs a multi-dimensional matching of these candidate foods, including but not limited to ingredient consistency verification and nutritional rule filtering, outputting the highest-matching food and its corresponding food description. Finally, the food recognition results are corrected based on this food description to further improve their accuracy.

[0055] In some specific embodiments, constructing a user profile of the target user based on the user information of the target user includes: Obtaining dietary needs and health status information of the target user; Determining a diet range suitable for the target user based on the health status information; A user profile of the target user is generated based on the dietary needs, the dietary range, and the physical health status information.

[0056] It's important to note that by building a user profile for each target user based on their dietary needs and health status, it's possible to achieve an organic balance between individual dietary preferences and nutritional needs from the perspective of the eater. This takes into account both the eater's food preferences and their physical condition, enabling recommendations for their favorite foods while still meeting their nutritional needs. Furthermore, it can factor in the complex and ever-changing physical information of each eater, such as diabetes and food allergies.

[0057] In some specific embodiments, combining the food recognition result and the user profile to generate a meal plan suitable for the target user using a large model inference includes: generating a plurality of candidate food combinations for the target user based on the food recognition result and the user portrait; Calculating the nutritional balance index and health benefit score of each candidate food combination based on the food identification results and the nutritional synergy between foods; A dietary planning program suitable for the target user is generated based on the nutritional balance index and the health benefit score.

[0058] Specifically, when calculating the nutritional balance index, the core nutritional data of each food (such as macronutrients, micronutrients, functional ingredients, and health risk indicators) can be obtained based on the food identification results. Combined with the nutritional synergy between foods, a pre-established multi-dimensional scoring system is used to calculate the nutritional balance index and health benefit score of each candidate ingredient combination. Preferably, the scoring parameters can be dynamically adjusted based on the user's health status information. For example, for diabetic patients, the carbohydrate weight can be reduced and the dietary fiber score can be increased. For fitness enthusiasts, the protein and BCAA amino acid score ratios can be increased.

[0059] It should be noted that when generating a meal planning plan for the target user, this application not only takes into account the natural properties of the food itself, such as the nutritional composition, calories, appearance, smell, taste, etc. of the food, but also takes into account the nutritional synergy between foods, so as to more accurately obtain a deep understanding of food types and other semantic levels, and further effectively complete the meal planning work.

[0060] In some specific embodiments, generating a plurality of candidate food combinations for the target user based on the food recognition result and the user portrait includes: Analyzing the nutritional needs of the target user based on the user portrait; Combining the nutritional requirements and the food identification results, a plurality of candidate food combinations that are compatible with the health goals of the target user are generated.

[0061] Specifically, the user profile is deeply analyzed to assess the target user's nutritional needs, including but not limited to basic metabolic parameters, disease management needs, nutritional gap detection, and personal preference constraints. Finally, the food recognition results are combined with the target user's nutritional needs and health goals to generate multiple candidate food combinations suitable for the target user.

[0062] In some specific embodiments, the health goal includes at least one of weight loss, muscle gain, and blood sugar control.

[0063] In some specific embodiments, the dynamically updated nutrition knowledge base is implemented in the following manner: Real-time access to the latest nutrition research data, food ingredient databases or user feedback; Incremental learning technology is used to update the parameters and rule base of the large model.

[0064] It should be noted that this application does not limit the specific access method for the latest nutritional research data, food database or user feedback. Without violating the inventive concept of this application, it can be selected according to actual needs, such as establishing a connection with an authoritative nutritional database through an API interface, regularly capturing the latest research results, or using natural language processing (NLP) to parse academic papers, automatically extract key conclusions, and convert them into computable nutritional rules.

[0065] In some specific embodiments, the method further comprises: The multimodal information is preprocessed, where the preprocessing includes at least data cleaning and data standardization.

[0066] As an illustrative, non-limiting illustration, the cleaning of food images in this application includes, but is not limited to, removing low-quality images, performing background segmentation and food region extraction, or filtering outliers (e.g., automatically removing interfering areas when non-food objects are identified using a semantic segmentation model). Cleaning of text descriptions includes, but is not limited to, error correction and redundant information removal. Cleaning of speech information includes, but is not limited to, ambient noise suppression and spoken language normalization.

[0067] In some specific embodiments, data standardization processing of multimodal information includes, but is not limited to, unified feature dimension processing, semantic alignment preprocessing, and the like. For example, unified feature dimension processing of food images includes uniformly scaling images of different resolutions to preset pixels or performing mean normalization on the RGB channels of the image. Unified feature dimension processing of text descriptions includes uniform truncation / padding to a preset token length using a word segmenter. Unified feature dimension processing of speech information includes unifying MFCC features into preset dimensions / frames and cyclically padding zeros when the duration is insufficient.

[0068] Example 2 Corresponding to the above-mentioned embodiment 1, the present application also provides a healthy meal planning system based on multimodal fusion, which is used to implement the healthy meal planning method based on multimodal fusion provided in any one of the embodiments 1. In this embodiment, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 As shown, the system includes: The food recognition module 10 is used to obtain multimodal information of food, extract and fuse features of the multimodal information based on multimodal fusion technology, and obtain food recognition results; A user portrait building module 20 is configured to build a user portrait of a target user based on the user information of the target user, wherein the user information at least includes the dietary needs and physical health information of the target user; A meal planning module 30 is configured to combine the food recognition results and the user profile and generate a meal planning solution suitable for the target user using a large model inference; Program optimization module 40, for optimizing the meal planning program according to the dynamically updated nutrition knowledge base, and outputting the optimized meal planning program to the target user. In some specific embodiments, the food identification module 10 is specifically used to: A convolutional neural network is used to extract visual features from the food image; semantic information in the text description is obtained through analysis using a natural language processing model; the voice information is recognized and analyzed using speech recognition technology to obtain corresponding speech recognition results; the visual features, the semantic information, and the speech recognition results are cross-modally aligned and fused to obtain the food identification information.

[0069] In some specific embodiments, the food identification module 10 is further configured to: Acquire a fusion feature based on the visual feature, the semantic information and the speech recognition result; match food description data in a preset food database according to the fusion feature; and correct the food recognition result according to the food description data.

[0070] In some specific embodiments, the user portrait building module 20 is specifically used to: Obtain the target user's dietary needs and physical health information; determine a dietary range suitable for the target user based on the physical health information; and generate a user profile of the target user based on the dietary needs, the dietary range, and the physical health information.

[0071] In some specific embodiments, the meal planning module 30 is specifically configured to: Based on the food recognition results and the user portrait, multiple candidate ingredient combinations are generated for the target user; the nutritional balance index and health benefit score of each candidate ingredient combination are calculated based on the food recognition results and the nutritional synergy between foods; and a meal planning plan suitable for the target user is generated based on the nutritional balance index and the health benefit score.

[0072] In some specific embodiments, the meal planning module 30 is further configured to: The nutritional needs of the target user are analyzed according to the user portrait; and a plurality of candidate food combinations that are compatible with the health goals of the target user are generated by combining the nutritional needs and the food identification results.

[0073] In some specific embodiments, the health goal includes at least one of weight loss, muscle gain, and blood sugar control.

[0074] In some specific embodiments, the solution optimization module 40 is specifically used to: Real-time access to the latest nutrition research data, food material database or user feedback; using incremental learning technology to update the parameters and rule base of the large model.

[0075] In some specific embodiments, the food identification module 10 is further configured to: The multimodal information is preprocessed, where the preprocessing includes at least data cleaning and data standardization.

[0076] Example 3 Corresponding to the above-mentioned embodiment 1, the present application also provides a computer device, including: a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the healthy diet planning method based on multimodal fusion provided in any one of the above-mentioned embodiments is executed.

[0077] in, Figure 3The computer device 1500 is shown as an example and may include a processor 1510, a video display adapter 1511, a disk drive 1512, an input / output interface 1513, a network interface 1514, and a memory 1520. The processor 1510, the video display adapter 1511, the disk drive 1512, the input / output interface 1513, the network interface 1514, and the memory 1520 may be communicatively connected via a communication bus 1530.

[0078] The processor 1510 may be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and may be used to execute relevant programs to implement the technical solutions provided by the present invention.

[0079] The memory 1520 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1520 can store an operating system 1521 for controlling the operation of the electronic device and a basic input and output system (BIOS) for controlling the low-level operations of the electronic device. In addition, a web browser 1523, a data storage management system 1524, and a device identification information processing system 1525 can also be stored. The above-mentioned device identification information processing system 1525 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present invention. In short, when the technical solution provided by the present invention is implemented by software or firmware, the relevant program code is stored in the memory 1520 and is called and executed by the processor 1510.

[0080] The input / output interface 1513 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component within the device (not shown) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc. Output devices may include a display, speaker, vibrator, indicator light, etc.

[0081] The network interface 1514 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0082] The bus comprises a pathway that transmits information between various components of the device (eg, processor 1510 , video display adapter 1511 , disk drive 1512 , input / output interface 1513 , network interface 1514 , and memory 1520 ).

[0083] In addition, the electronic device can also obtain information on specific collection conditions from the virtual resource object collection condition information database for use in condition judgment, etc.

[0084] It should be noted that although the above device only shows a processor 1510, a video display adapter 1511, a disk drive 1512, an input / output interface 1513, a network interface 1514, a memory 1520, a bus, etc., in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include components necessary to implement the solution of the present invention, and does not necessarily include all the components shown in the figure.

[0085] Example 4 Corresponding to the above-mentioned embodiment 1, the embodiment of the present application further provides a computer-readable storage medium, wherein, in this embodiment, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated later.

[0086] The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by the processor, the healthy diet planning method based on multimodal fusion as described above is implemented.

[0087] In some implementations, in the embodiments of the present application, when the computer program is executed by the processor, it can also implement steps corresponding to the method described in Example 1. Please refer to the detailed description in Example 1 and will not be repeated here.

[0088] From the above description of the embodiments, it is clear that those skilled in the art will clearly understand that the present invention can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the methods described in various embodiments of the present invention, or portions thereof.

[0089] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0090] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0091] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0092] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0093] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A healthy diet planning method based on multimodal fusion, characterized in that: The method comprises: Acquiring multimodal information of food, performing feature extraction and fusion on the multimodal information based on multimodal fusion technology, and obtaining food recognition results; Building a user profile of the target user based on the user information of the target user, wherein the user information at least includes the dietary needs and physical health information of the target user; Combining the food recognition results and the user profile, using large model reasoning to generate a meal plan suitable for the target user; The dietary planning scheme is optimized according to the dynamically updated nutritional knowledge base, and the optimized dietary planning scheme is output to the target user.

2. The healthy diet planning method based on multimodal fusion according to claim 1 is characterized in that: The multimodal information includes at least a food image, a text description, and voice information of the food to be identified. The step of extracting and fusing features of the multimodal information based on the multimodal fusion technology to obtain a food identification result includes: extracting visual features from the food image using a convolutional neural network; Obtaining semantic information from the text description through natural language processing model analysis; Using speech recognition technology to identify and analyze the speech information to obtain corresponding speech recognition results; The visual features, the semantic information, and the speech recognition results are cross-modally aligned and fused to obtain food recognition information.

3. The healthy diet planning method based on multimodal fusion according to claim 2, characterized in that: The extracting and fusing features of the multimodal information based on the multimodal fusion technology to obtain the food recognition result further includes: Acquire a fusion feature based on the visual feature, the semantic information, and the speech recognition result; matching food description data in a preset food database according to the fusion feature; The food recognition result is corrected according to the food description data.

4. The healthy diet planning method based on multimodal fusion according to claim 1, characterized in that: The constructing a user profile of the target user based on the user information of the target user includes: Obtaining dietary needs and health status information of the target user; Determining a diet range suitable for the target user based on the health status information; A user profile of the target user is generated based on the dietary needs, the dietary range, and the physical health status information.

5. The healthy diet planning method based on multimodal fusion according to claim 1, characterized in that: Combining the food recognition result and the user portrait and using a large model to infer and generate a meal plan suitable for the target user includes: generating a plurality of candidate food combinations for the target user based on the food recognition result and the user portrait; Calculating the nutritional balance index and health benefit score of each candidate food combination based on the food identification results and the nutritional synergy between foods; A dietary planning program suitable for the target user is generated based on the nutritional balance index and the health benefit score.

6. The healthy diet planning method based on multimodal fusion according to claim 5, characterized in that: Generating a plurality of candidate food combinations for the target user based on the food recognition result and the user portrait includes: Analyzing the nutritional needs of the target user based on the user portrait; Combining the nutritional requirements and the food identification results, a plurality of candidate food combinations that are compatible with the health goals of the target user are generated.

7. The healthy diet planning method based on multimodal fusion according to claim 6, characterized in that: The health goal includes at least one of weight loss, muscle gain, and blood sugar control.

8. The healthy diet planning method based on multimodal fusion according to any one of claims 1 to 7, characterized in that: The dynamically updated nutrition knowledge base is implemented in the following ways: Real-time access to the latest nutrition research data, food ingredient databases or user feedback; Incremental learning technology is used to update the parameters and rule base of the large model.

9. The healthy diet planning method based on multimodal fusion according to any one of claims 1 to 7, characterized in that: The method further comprises: The multimodal information is preprocessed, where the preprocessing includes at least data cleaning and data standardization.

10. A healthy diet planning system based on multimodal fusion, characterized in that: The system is used to implement the method according to any one of claims 1 to 9, and the system includes: A food recognition module is used to obtain multimodal information of food, extract and fuse features of the multimodal information based on multimodal fusion technology, and obtain food recognition results; A user portrait building module, configured to build a user portrait of a target user based on the user information of the target user, wherein the user information at least includes the dietary needs and physical health information of the target user; A meal planning module is used to combine the food recognition results and the user profile and use large model reasoning to generate a meal planning plan suitable for the target user; A plan optimization module is used to optimize the meal planning plan according to the dynamically updated nutrition knowledge base and output the optimized meal planning plan to the target user.

Citation Information

Patent Citations

  • Food recommendation method and system based on multi-modal information association analysis

    CN111159539A

  • Intelligent nutritional diet recommendation method based on multi-source heterogeneous data fusion

    CN112614564A

  • Intelligent food identification and diet analysis system based on multi-modal interaction

    CN120126124A

  • Method and apparatus for identifying food using multimodal model

    EP4557237A1

  • Apparatus for providing information about a food product

    GB201705975D0

Cited By

  • Diabetes diet management method, system and equipment based on multi-modal data fusion

    CN122091101A