Food nutrient content prediction method and system based on cross-modal attention mechanism
By adopting a cross-modal attention mechanism in food nutrition assessment, extracting and fusing multimodal features of food images, the problem of insufficient accuracy and efficiency of nutrition assessment in the prior art is solved, and rapid and accurate nutritional component prediction is achieved.
Patent Information
- Application Number
- CN202210027210.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-11
AI Technical Summary
The existing food nutrition assessment methods are insufficient in terms of accuracy and efficiency, especially the need for end-to-end, fast and effective nutrition assessment is not met.
The food nutritional component content prediction method is adopted with a cross-modal attention mechanism, and the food image is extracted and attention mapped by multimodal feature and fused RGB and deep modal features to predict the content of each nutrient component in the food.
It achieves rapid and accurate prediction of the content of nutrients such as quality, calories, fat, etc. in food, improves the accuracy and speed of nutritional evaluation, and meets the needs of efficient and fast nutritional evaluation.
Smart Images

Figure CN114529790B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food nutrition assessment, and in particular to a method and system for predicting food nutrient content based on a cross-modal attention mechanism. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Regarding the evaluation of food nutritional content, professional food nutritional assessors are slightly lacking in accuracy in estimating food nutritional content, and professionals cannot meet people's daily food nutritional assessment needs. Therefore, in order to alleviate the shortage of nutritional assessment professionals and improve the accuracy and efficiency of nutritional assessment, a large number of dietary assessment applications and food nutritional estimation systems based on smartphone images have emerged, such as FatSceret, bitesnap, and keepoa. Although they can roughly calculate the calorie content, they require manual weighing and manual input of portion size and other information. This process is cumbersome, time-consuming, and prone to errors. In addition, the identification, overlap, and portion size estimation of food types will affect the accuracy of nutritional estimation, so these applications and systems are still lacking in accuracy and efficiency.
[0004] At present, some end-to-end food nutrition estimation methods estimate the nutritional content based on a single RGB image, or simply add RGB and depth images to predict the calorie, mass, carbohydrate, protein and fat content. Although the feasibility of nutrition assessment based on food images has been proven, the image features used for nutrition prediction have not been fully explored, and no more robust prediction results have been obtained. Furthermore, these methods require a series of complex operations, and there are problems such as cumbersome steps and slow speed. In terms of nutrition assessment accuracy and speed, existing methods cannot meet the requirements of end-to-end, fast and effective nutrition assessment. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes a method and system for predicting the nutritional content of food based on a cross-modal attention mechanism, which extracts multimodal features of food images and simultaneously predicts the nutritional content by fusing attention multimodal features, thereby improving prediction accuracy and speed.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for predicting food nutrient content based on a cross-modal attention mechanism, comprising:
[0008] Label the nutritional ingredients and content of food image samples to train the prediction model;
[0009] Perform multimodal feature extraction on the food image to be tested, perform attention mapping on the feature map of each modality to obtain a weight map, perform outer product on the feature map of any modality and the weight map of other modalities, and add the outer product result to the feature map of the modality to perform feature fusion;
[0010] According to the fused feature map, the prediction model is used to obtain the content of each nutrient in the food image to be tested.
[0011] As an optional implementation, the process of extracting multimodal features from the food image to be tested includes using a ResNet-101 network as a backbone network to extract four levels of RGB modal features and deep modal features.
[0012] As an optional implementation, the process of extracting multimodal features from the food image to be tested includes reducing the dimension of the extracted multimodal features through a 1×1 convolution layer, and obtaining a feature map of each modal feature through two residual convolution units.
[0013] As an optional implementation, the process of performing attention mapping on the feature map of each modality includes obtaining a weight map by subjecting the feature map to global average pooling, 1×1 convolution reorganization and activation function.
[0014] As an optional implementation, the fused features under each modal branch are added after passing through the convolution layer to obtain a final fused feature map, and the fused feature map is passed through a chain residual pool to capture contextual information.
[0015] As an optional embodiment, the nutritional components include calories, mass, fat, protein and carbohydrates.
[0016] As an alternative embodiment, the mean absolute error and the percentage of the mean absolute error are used to evaluate the accuracy of the nutrient content prediction.
[0017] In a second aspect, the present invention provides a food nutrient content prediction system based on a cross-modal attention mechanism, comprising:
[0018] A training module configured to annotate a sample set of food images with nutritional ingredients and contents, thereby training a prediction model;
[0019] The feature processing module is configured to perform multimodal feature extraction on the food image to be tested, perform attention mapping on the feature map of each modality to obtain a weight map, perform outer product on the feature map of any modality and the weight map of other modalities, and add the outer product result to the feature map of the modality to perform feature fusion;
[0020] The prediction module is configured to obtain the content of each nutrient component in the food image to be tested by using a prediction model based on the fused feature map.
[0021] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.
[0022] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] The present invention provides a method and system for predicting the nutrient content of food with a cross-modal attention mechanism, which can quickly and accurately predict the content of nutrients such as mass, calories, and fat in food, and perform predictions in an end-to-end manner. The nutrient content can be efficiently predicted in terms of both accuracy and speed, meeting the demand for efficient and fast nutrition assessment methods, significantly improving the detection speed and accuracy of food nutrition assessment methods, and effectively improving the stability and feasibility of food nutrition assessment applications in daily life.
[0025] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0027] Figure 1 A flow chart of a method for predicting food nutrient content using a cross-modal attention mechanism provided in Example 1 of the present invention;
[0028] Figure 2 A diagram of the cross-modal attention feature fusion structure provided in Example 1 of the present invention;
[0029] Figure 3 The attention structure diagram provided by Embodiment 1 of the present invention;
[0030] Figure 4 This is a flowchart of a single iteration in the training process provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0031] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0032] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0033] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0035] Example 1
[0036] This embodiment provides an end-to-end cross-modal attention mechanism for predicting food nutrient content, which aims to use deep learning methods to combine multimodal information such as RGB and depth in food images to perform food recognition, classification, and calorie and macronutrient content prediction. Figure 1 As shown, specifically including;
[0037] Label the nutritional ingredients and content of food image samples to train the prediction model;
[0038] Perform multimodal feature extraction on the food image to be tested, perform attention mapping on the feature map of each modality to obtain a weight map, perform outer product on the feature map of any modality and the weight map of other modalities, and add the outer product result to the feature map of the modality to perform feature fusion;
[0039] According to the fused feature map, the prediction model is used to obtain the content of each nutrient in the food image to be tested.
[0040] In this embodiment, food images are collected in different environmental scenes so that the collected food images can include different types of interference that exist in a normal dining environment for users, so as to represent a real restaurant environment;
[0041] Normalize the size of the food images to 600×400, add and weigh the food items one by one to the plate in increments to record the quality of each type of food, and record the name and quality of the food items in turn;
[0042] Refer to the food nutrition database for the calories, protein, carbohydrates, fat and other nutritional content per gram of each food to mark the real value, which is convenient for the subsequent generation of model training targets;
[0043] After annotating the food images with nutritional ingredients and nutritional content, the food dataset is divided into training sets and test sets to train the prediction model.
[0044] In this embodiment, the prediction model includes an image feature extraction network, an attentive multi-modal feature fusion network (AMFFNet) and a feature refinement network (RefineNet). The ResNet-101 network is used to extract the multimodal features of food images, and the extracted four-level RGB features and depth features are fused through the AMFFNet network. The fused features are further refined through RefineNet to obtain a feature map rich in detail information and semantic information.
[0045] In this embodiment, the image feature extraction network uses the ResNet-101 network as the backbone network, and uses the extracted four-level RGB and depth features as the input of the AMFFNet network, that is, the outputs of the four residual blocks conv2, conv3, conv4, and conv5 are used as the input of the AMFFNet network, and then the output of the AMFF network is further fused through the RefineNet network.
[0046] In this embodiment, in order to fully integrate multi-scale features, RefineNet with residual learning is used to perform multi-level feature fusion. The RefineNet network includes RefineNet-1, RefineNet-2, RefineNet-3, and RefineNet-4. Except for RefineNet-4, other RefineNet modules receive the fused features from AMFF and the refined features of the previous module; RefineNet-1 outputs the final fused features, which contain more detailed information and semantic information for better nutrition estimation in the future.
[0047] In order to obtain more powerful features, in addition to the ordinary feature extraction process, multimodal feature fusion is very necessary. Most of the existing multimodal fusion methods follow a later fusion method and fail to fully fuse. The cross-modal application of attention mechanisms for feature maps of different modalities is complementary in visual tasks. When a strongly desired target appears at a certain position in a certain modality, the judgment of the position is improved not only in the expected modality, but also in another modality. For the processing of multimodal data, the cross-modal attention mechanism has shown its effectiveness. In order to more fully exploit RGB features and deep features, this embodiment proposes an efficient multimodal feature fusion network.
[0048] This embodiment adopts a cross-modal attention mechanism, which models the importance of each channel and enhances or suppresses different channels for different tasks, effectively capturing the information between channels while ensuring that the dimension remains unchanged.
[0049] The attention multimodal feature fusion network adopts the AMFFNet network, such as Figure 2 As shown in the figure, the extracted multimodal features are first passed through a 1×1 convolution layer to reduce the dimension and reduce the amount of calculation to facilitate subsequent training; each modal feature is then passed through two residual convolution units (RCU) to obtain a feature map, and the feature map is subjected to attention mapping by the channel attention module (CAM).
[0050] In the RGB mode, the feature map of the deep modality branch is subjected to attention mapping to obtain a weight map, that is, the feature map of the deep modality is subjected to global average pooling, 1×1 convolution reorganization, and activation function to obtain a weight map, and the weight map is multiplied with the feature map of the RGB branch (Element-wise multiplication), and the multiplication result is added to the feature map of the RGB branch (Element-wise addition) to obtain the feature map after cross-modal attention mechanism fusion;
[0051] After the deep branch performs the same processing, the output results of the two branches are added after passing through the convolution layer to obtain a fused feature map; finally, the fused feature map is passed through the chained residual pooling to capture a wide range of contextual information.
[0052] like Figure 3 As shown, the feature map obtained after the residual convolution unit RCU is F = {F 1 ,F 2 ,……F C}, F∈RH ×W×C ; In the attention module CAM, the feature map is firstly subjected to global average pooling to obtain the output O∈R H×W×C , H and W represent the width and height of the feature map respectively, and C represents the number of channels;
[0053]
[0054] Where k is the kth channel in the output O, O k It is the output after global average pooling.
[0055] Then, a 1×1 convolutional layer is used to learn the relationship between each channel to obtain the output of the same channel as the output O; the sigmoid activation function is used for the output result, which not only increases the nonlinearity of the network, but also limits the weight of each feature channel to [0,1].
[0056] Assume that the output after the activation function is ω∈R H×W×C :
[0057]
[0058] Among them, σ represents the sigmoid function, Represents a 1×1 convolution operation with C channels.
[0059] Perform the outer product of the feature map and the weight map under the relative mode, and then add the outer product output to the original feature map of the corresponding branch to obtain the final output Y∈R H×W×C It is expressed as follows:
[0060]
[0061]
[0062] in, represents the outer product, represents the addition of corresponding elements, F RGB represents the feature map corresponding to the RGB branch output by the RCU module, F depth Represents the feature map corresponding to the depth branch output by the RCU module, ω RGB and ω depth They respectively represent the weights of the corresponding channels obtained through each branch of the attention module.
[0063] Through the above operations, the feature map F is transformed into a new feature map Y, which contains more information useful for prediction than F.
[0064] In this embodiment, it is assumed that the current feature map is Y∈R H×W×C, and obtain the predicted results of nutrient content through multi-scale prediction, global average pooling, and fully connected layers. Considering that the results output by nutrient estimation and end-to-end methods are numerical, this embodiment uses mean absolute error MAE and the percentage value of mean absolute error to measure the accuracy of nutrient prediction.
[0065] Nutrients include calories, mass, and macronutrients. Macronutrients include carbohydrates, fats, and proteins. Calories are in kilocalories, and mass and macronutrients are in grams. The mean absolute error percentage represents the percentage of the mean absolute error to the average of all true values. The lower the mean absolute error percentage value for each evaluation indicator, the higher the accuracy of the nutritional content assessment.
[0066] It is expressed as follows;
[0067]
[0068] in, is the predicted value of i for a given test image, y i is the true value of image i.
[0069] In this embodiment, the total loss function adopts a multi-loss function. The total loss function integrates mass loss, calorie loss and macronutrient loss, and uses the mean absolute error as the regression loss of mass, calories and macronutrients. The optimization algorithm is used to further optimize the problem of excessive swing amplitude of the loss function during the update, thereby accelerating the convergence speed of the network.
[0070] The total loss function is composed of five regression losses. The quality loss, calorie loss and macronutrient loss (carbohydrate, fat and protein) constitute the total loss function. The total loss of multiple tasks is defined as shown in formula (6):
[0071]
[0072] Among them, y m ,y cal , is the true value of mass, calories and macronutrients. M includes the total content of carbohydrates, fats and proteins. are the network's predictions for mass, calories, and macronutrients.
[0073] By calculating the loss function, continuously performing gradient back propagation to update the prediction model parameters, repeatedly iterating and evaluating, the optimal prediction model is obtained. The training process of a single iteration network is as follows: Figure 4 shown.
[0074] This embodiment provides a method and system for predicting the nutritional content of food with a cross-modal attention mechanism, and provides an end-to-end network structure. First, RGB and depth images are used as network inputs, and the ResNet network is used as the backbone network to extract multimodal features of food images; the extracted features of the four levels are fused based on attention multimodal feature fusion and feature refinement, so that the feature map has richer detail information and semantic information; finally, after multi-scale prediction, global average pooling, and a fully connected layer, the final calorie, mass, fat, protein, and carbohydrate content are output. To meet the pursuit of healthy and nutritious diets, by understanding and tracking the nutritional content of the food eaten, users can make more scientific and reasonable dietary choices, thereby solving many problems such as the shortage of nutrition professionals and the low accuracy of nutrition assessments.
[0075] Example 2
[0076] This embodiment provides a food nutrient content prediction system based on a cross-modal attention mechanism, including:
[0077] A training module configured to annotate a sample set of food images with nutritional ingredients and contents, thereby training a prediction model;
[0078] The feature processing module is configured to perform multimodal feature extraction on the food image to be tested, perform attention mapping on the feature map of each modality to obtain a weight map, perform outer product on the feature map of any modality and the weight map of other modalities, and add the outer product result to the feature map of the modality to perform feature fusion;
[0079] The prediction module is configured to obtain the content of each nutrient component in the food image to be tested by using a prediction model based on the fused feature map.
[0080] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.
[0081] In further embodiments, there is also provided:
[0082] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, it will not be described in detail here.
[0083] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0084] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0085] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method described in Example 1 is completed.
[0086] The method in Example 1 can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it is not described in detail here.
[0087] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0088] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A method for predicting food nutrient content based on a cross-modal attention mechanism, characterized in that: include: The food image sample set is labeled with nutrients and contents to train the prediction model; the prediction model includes an image feature extraction network, an attention multimodal feature fusion network and a feature refinement network. The ResNet-101 network is used to extract the multimodal features of the food image, and the extracted four-level RGB features and deep features are fused through the AMFFNet network. The fused features are further refined through RefineNet to obtain a feature map rich in detail information and semantic information; Perform multimodal feature extraction on the food image to be tested, perform attention mapping on the feature map of each modality to obtain a weight map, perform outer product on the feature map of any modality and the weight map of other modalities, and add the outer product result to the feature map of the modality to perform feature fusion; add the fused features under each modality branch after passing through the convolution layer to obtain the final fused feature map, and capture the context information of the fused feature map through the chain residual pool; According to the fused feature map, the prediction model is used to obtain the content of each nutrient in the food image to be tested.
2. The method for predicting food nutrient content based on a cross-modal attention mechanism as claimed in claim 1, characterized in that: The process of multimodal feature extraction for the food images to be tested includes using the ResNet-101 network as the backbone network to extract four levels of RGB modal features and deep modal features.
3. The method for predicting food nutrient content based on a cross-modal attention mechanism as claimed in claim 1, characterized in that: The process of extracting multimodal features from the food image to be tested includes reducing the dimension of the extracted multimodal features through a 1×1 convolution layer, and then obtaining a feature map of each modal feature through two residual convolution units.
4. The method for predicting food nutrient content based on a cross-modal attention mechanism as claimed in claim 1, characterized in that: The process of attention mapping the feature map of each modality includes obtaining a weight map after global average pooling, 1×1 convolution reorganization and activation function of the feature map.
5. The method for predicting food nutrient content based on a cross-modal attention mechanism as claimed in claim 1, characterized in that: The fused features under each modal branch are added after passing through the convolution layer to obtain the final fused feature map, and the fused feature map is passed through the chain residual pool to capture the context information.
6. The method for predicting food nutrient content based on a cross-modal attention mechanism as claimed in claim 1, characterized in that: The nutritional components include calories, mass, fat, protein and carbohydrates.
7. The method for predicting food nutrient content based on a cross-modal attention mechanism as claimed in claim 1, characterized in that: The accuracy of nutrient content predictions was assessed using mean absolute error and percentage of mean absolute error.
8. A food nutrient content prediction system based on a cross-modal attention mechanism, characterized in that: include: A training module is configured to annotate the nutritional components and contents of a food image sample set to train a prediction model; the prediction model includes an image feature extraction network, an attention multimodal feature fusion network, and a feature refinement network. The ResNet-101 network is used to extract multimodal features of food images, and the extracted four-level RGB features and deep features are fused through an AMFFNet network. The fused features are further refined through RefineNet to obtain a feature map rich in detail information and semantic information; The feature processing module is configured to perform multimodal feature extraction on the food image to be tested, perform attention mapping on the feature map of each modality to obtain a weight map, perform outer product on the feature map of any modality and the weight map of other modalities, and add the outer product result to the feature map of the modality to perform feature fusion; add the fused features under each modality branch after passing through the convolution layer to obtain the final fused feature map, and capture context information on the fused feature map through a chain residual pool; The prediction module is configured to obtain the content of each nutrient component in the food image to be tested by using a prediction model based on the fused feature map.
9. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 7 is completed.
10. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for image-text cross-mode sentiment classification based on compact bilinear fusion
CN107066583A
Cross-modal Hash method and system based on multi-modal attention mechanism
CN113095415A