Food nutrition detection model, method and equipment

Through a multimodal fusion food nutrition detection model, the visual, depth and text description features are integrated and dimensionality reduction features are integrated, and the problem of insufficient correlation between characteristics and nutrition in the prior art is solved, which significantly improves the accuracy and reliability of food nutrition detection.

CN120107955APending Publication Date: 2025-06-06INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510149072.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing vision-based food nutrition assessment methods rely on accurate food category identification and comprehensive nutrition databases, and the extracted features are mainly used for image classification, lack of information directly related to nutritional components, affecting the accuracy and reliability of the assessment.

Method used

A food nutrition detection model is proposed, including feature extraction module, enhancement and fusion module, feature integration module and nutrition regression head. By integrating visual features, depth features and text description features, dimensionality reduction features are integrated and reduced to improve the correlation between features and nutrition.

Benefits of technology

It significantly improves the accuracy of food nutrition testing, can better understand the portion size and composition of food, enhances the nutrition-related semantic information in the characteristics, and improves the performance and reliability of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107955A_ABST
    Figure CN120107955A_ABST
Patent Text Reader

Abstract

The invention provides a food nutrition detection model. The food nutrition detection model comprises a feature extraction module, an enhancement and fusion module, a feature integration module and a nutrition regression head. The feature extraction module is used for performing feature extraction according to the food image to obtain a visual feature map and performing feature extraction according to the depth image of the food image to obtain a depth feature map. And the enhancement and fusion module is used for obtaining a fusion feature map according to the food image, the visual feature map and the depth feature map. And the feature integration module is used for carrying out integration and dimension reduction on the fused feature map to obtain a feature vector after dimension reduction. And the nutrition regression head is used for predicting the nutritional ingredient value of the food in the food image according to the feature vector after dimension reduction. According to the invention, the accuracy of food nutrition detection is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and deep learning, and in particular to a food nutrition detection model, method and device. Background Art

[0002] The statements in this section are merely intended to provide background information related to the present invention to help understanding the present invention. Such background information does not necessarily constitute prior art.

[0003] With the development of deep learning, vision-based food nutrition assessment has become a cutting-edge research field. It can directly predict the nutritional content of food by analyzing food images, making it possible to achieve fast, automated and accurate assessment. Previous food nutrition assessment methods were implemented through a multi-stage evaluation process, which required identifying the specific type of food and then querying the relevant nutrition database to estimate the nutritional value of the food. However, this method relies on accurate food category recognition algorithms and requires the construction of a comprehensive nutrition database, which limits its further development. With the development of computer vision technology, researchers have begun to explore end-to-end food nutrition assessment methods, directly extracting features from food images and using them to predict nutritional content. This method avoids dependence on food category recognition, but places higher requirements on the nutritional relevance of the extracted features. In current vision-based food nutrition assessment research, deep learning models designed for image classification are usually used directly to extract image features and use them to predict the nutritional content of food. The problem with this method is that the features extracted by these models are mainly optimized for identifying images of different categories. They focus more on the discriminative features of images rather than information directly related to nutritional content. Therefore, while such methods can automate the process of image-to-nutrient content prediction, the features they rely on may not be strongly correlated with the actual nutritional value of the food, which may affect the accuracy and reliability of the assessment. Summary of the invention

[0004] To solve the above problems, the present invention provides a food nutrition detection model, method and equipment.

[0005] According to a first aspect, the present invention provides a food nutrition detection model, comprising: a feature extraction module, an enhancement and fusion module, a feature integration module and a nutrition regression head; the feature extraction module is used to perform feature extraction based on a food image to obtain a visual feature map, and to perform feature extraction based on a depth image of the food image to obtain a depth feature map; the enhancement and fusion module is used to extract a text description feature map of the food image based on the food image, and to obtain a fused feature map based on the text description feature map of the food image, the visual feature map and the depth feature map; the feature integration module is used to integrate and reduce the dimension of the fused feature map to obtain a feature vector after dimensionality reduction; the nutrition regression head is used to predict the nutritional value of the food in the food image based on the feature vector after dimensionality reduction.

[0006] Preferably, the enhancement and fusion module includes a first feature pyramid network, a second feature pyramid network, a first feature fusion module, a second feature fusion module and a visual-language contrast pre-training model; wherein the first feature pyramid network is used to refine the visual feature map to obtain a first refined feature map; the second feature pyramid network is used to refine the depth feature map to obtain a second refined feature map; the first feature fusion module is used to fuse the first refined feature map with the second refined feature map to obtain a first fused feature map; the visual-language contrast pre-training model is used to obtain a text description feature map of the food image based on the food image; the second feature fusion module is used to fuse the first fused feature map with the text description feature map to obtain the fused feature map.

[0007] Preferably, the visual-language contrast pre-training model includes a visual encoder; the visual encoder is used to obtain a contrast language-image pre-training feature map based on the food image; wherein the contrast language-image pre-training feature map includes a text description feature map of the food image.

[0008] Preferably, the feature integration module includes a position encoding module, a self-attention mechanism and a global average pooling module; the position encoding module is used to obtain the position information of the fused feature map based on the fused feature map; the self-attention mechanism is used to mine and integrate the multi-scale information of the fused feature map based on the position information to obtain the integrated feature map; the global average pooling module is used to convert the integrated feature map into the feature vector after dimensionality reduction.

[0009] Preferably, the self-attention mechanism includes a spatial self-attention mechanism and a hierarchical self-attention mechanism; the spatial self-attention mechanism is used to mine the multi-scale information between different spatial positions of the fused feature map according to the position information; the hierarchical self-attention mechanism is used to integrate the multi-scale information between different spatial positions of the fused feature map to obtain the integrated feature map.

[0010] Preferably, residual connections are introduced into the spatial self-attention mechanism and the hierarchical self-attention mechanism respectively.

[0011] Preferably, the feature integration module further includes a multi-layer perceptron, and the multi-layer perceptron is used to process the integrated feature map before the global average pooling module.

[0012] Preferably, the nutritional component values ​​predicted by the nutritional regression include calories, fat content, carbohydrate content and protein content.

[0013] According to the second aspect, the present invention provides a food nutrition detection method, which includes: inputting a food image and a depth image of the food image into the nutrition detection model described in any one of the first aspects; the nutrition detection model outputs the nutritional value of the food in the food image.

[0014] According to the third aspect, the present invention provides a food nutrition detection device, comprising: a processor; a storage device for storing one or more programs, when the one or more programs are executed by the processor, the food nutrition detection device implements the steps of the method described in the second aspect.

[0015] The depth feature-based enhancement and fusion module in the present invention can fuse visual information and spatial information of food from depth images, and enhance the information in food images. The feature integration module can integrate and refine the extracted features, improve the correlation between features and nutrition to improve the evaluation performance, and enable a better understanding of the portion and composition of food, thereby significantly improving the accuracy of food nutrition detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a structural schematic diagram of a nutrition detection model according to an embodiment of the present invention; Figure 2 is an extraction schematic diagram of a feature extraction module according to an embodiment of the present invention; Figure 3 is a schematic diagram of the structure and workflow of the enhancement and fusion module of the nutrition detection model according to one embodiment of the present invention; Figure 4 It is a schematic diagram of the structure and workflow of a feature integration module of a nutrition detection model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments herein are only used for illustration and are not intended to limit the present invention. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to those of ordinary skill in the art that these specific details are not necessarily used to implement the present invention. In other examples, in order to avoid confusing the present invention, known procedures, materials or methods are not specifically described.

[0018] Existing technologies can meet the needs of food nutrition assessment to a certain extent. However, multi-stage food nutrition assessment technology relies on accurate food category recognition algorithms, which are limited by the limited number of foods and nutrition databases and cannot handle foods outside of existing categories. The nutritional relevance of food visual features extracted by end-to-end food nutrition assessment technology is low, which affects further improvement of accuracy.

[0019] When the inventors were conducting research on food nutrition assessment models based on deep learning, they found that it was difficult to accurately assess food nutrition from food images in the prior art because traditional computer vision and deep learning methods found it difficult to learn food nutrition-related knowledge from end-to-end training. This may be because the visual features of the food itself are complex, and end-to-end training only establishes a connection between the visual features and nutrition of food images, while information such as food composition and portion size is relatively scarce, resulting in poor performance. After research on computer vision, deep learning, and food nutrition assessment-related technologies, the inventors found that this defect can be solved by introducing multimodal information such as deep images and image description text into the deep learning method for fusion, increasing the amount of available information, and improving the deep learning method's understanding of food nutrition, thereby improving the performance of the assessment.

[0020] 1. Model Structure like Figure 1 As shown, according to an embodiment of the present invention, a food nutrition detection model based on multimodal fusion is provided. The nutrition detection model includes: a feature extraction module 104, an enhancement and fusion module 101, a feature integration module 102 and a nutrition regression head 103. Among them, the feature extraction module 104 is used to extract features from food images to obtain visual feature maps, and to extract features from depth images of food images to obtain depth feature maps. Figure 2An embodiment of the feature extraction module 104 is shown, which includes a first backbone network 1041 and a second backbone network 1042. The first backbone network 1041 is used to extract features from food images to obtain a visual feature map, and the second backbone network 1042 is used to extract features from the depth image of the food image to obtain a depth feature map. In some embodiments, the food image may be an RGB image. The depth image may be a depth map simultaneously acquired by a depth camera when acquiring the food image, or may be a depth map generated based on the food image by a depth estimation basic model based on a monocular image. The first backbone network 1041 and the second backbone network 1042 may be backbone network structures commonly used in various computer vision and deep learning fields, such as ResNet, VGG, etc. In some embodiments, the visual feature map and the depth feature map may be single-scale or multi-scale, depending on the specific structure of the backbone network. Figure 1 The enhancement and fusion module 101 shown is used to extract a text description feature map of the food image based on the food image, and fuse the text description feature map, the visual feature map and the depth feature map of the food image to obtain a fused feature map. In some embodiments, a visual-language contrast pre-training model can be used to extract a text description feature map of the food image based on the food image. The visual-language contrast pre-training model will be introduced in detail later. The enhancement and fusion module 101 can use a multimodal fusion method to fuse feature maps from depth images corresponding to food images at different scales, so as to achieve the purpose of multimodal fusion and enhancement of nutrition-related semantic information in features. Among them, the multimodal fusion method can be any method in deep learning that can be used to fuse multiple modal features. Figure 1 The feature integration module 102 shown is used to integrate and reduce the dimensionality of the fused feature map to obtain a reduced-dimensional feature vector. The feature integration module integrates and reduces the dimensionality of the high-dimensional complex feature map output by the enhancement and fusion module for subsequent regression tasks. Figure 1 The nutritional regression head 103 shown is used to predict the nutritional value of food based on the feature vector after dimensionality reduction.

[0021] Figure 3 FIG. 1 is a schematic diagram of the structure and workflow of the enhancement and fusion module 101 according to an embodiment of the present invention. Figure 3As shown, the enhancement and fusion module 101 includes a first feature pyramid network 301, a second feature pyramid network 302, a vision-language comparison pre-trained model (Vision-Language Pre-trained Models) 303, a first feature fusion module 304 and a second feature fusion module 305. The visual feature map is input into the first feature pyramid network 301 for refinement processing to obtain a first refined feature map. The depth feature map is input into the second feature pyramid network 302 for refinement processing to obtain a second refined feature map. Then, the first refined feature map and the second refined feature map are fused by the first feature fusion module 304 to obtain a first fused feature map. The first feature pyramid network 301 and the second feature pyramid network 302 both use feature pyramid networks (FPN for short), which is a commonly used structure in deep learning target detection models, and is used to construct multi-scale feature representations to detect targets of different sizes. The core idea of ​​FPN is to fuse feature information of different scales to obtain a feature pyramid containing rich information. FPN solves the shortcomings of target detection in dealing with multi-scale changes, so that the model can simultaneously use high-level semantic information and bottom-level contour information to effectively detect targets of different sizes. In addition, in order to make up for the lack of semantics, enhance detail retention and meet multi-scale requirements, according to an embodiment of the present invention, it is also necessary to refine the feature map. In FPN, refinement refers to the process of refining and improving the features layer by layer by fusing top-down semantic information with bottom-level high-resolution features, so that the features of each scale retain both local spatial details and stronger high-level semantics. Refinement has the characteristics of downward propagation of high-level semantics (i.e., supplementing categories and contextual information from deep semantic features of the network), lateral connection fusion (i.e., adding or convolutionally fusing with low-level features of the same resolution to retain details and resolution) and multi-scale consistency (i.e., features at all levels have rich semantics at different resolutions, improving sensitivity to small targets and details). The first fused feature map not only includes the original image information, but also introduces depth information from the depth image, which can help the model understand spatial information such as the amount of food. The food image is input into the trained visual-language contrast pre-training model 303 and resized to obtain a contrastive language-image pre-training feature map (CLIP) and a CLIP feature map (i.e., a text description feature map of the food image). The CLIP feature map contains text description information of the image, which can help the model better understand the food ingredients, raw materials, cooking methods, and other information in the image, thereby achieving more accurate food nutrition assessment.Among them, the visual-language contrast pre-training model 303 is a multimodal pre-training model that can simultaneously process and understand image and language data, so as to be applied to a variety of visual, language and cross-modal tasks. By pre-training on large-scale visual and language data, the model can capture the complex relationship between image and text, so that the model can show excellent performance in tasks such as image description, visual question answering, cross-modal retrieval, etc., and is particularly suitable for enhancing semantic analysis in food nutrition assessment. The visual-language contrast pre-training model 303 can maximize the similarity of matching image-text pairs and minimize the similarity of unmatched image-text pairs through contrastive learning methods to learn the matching relationship between image and text. Through different mechanisms, such as self-attention, joint attention and visual-semantic embedding cross-modal contrastive learning, deep fusion between image and text is achieved. Large-scale unlabeled data can be used for pre-training, and the cross-modal understanding ability of the model can be improved through tasks such as image-text matching, masked language modeling and masked visual modeling. The model can predict the most relevant text fragments or images through natural language instructions without directly optimizing specific tasks, showing wide potential in a variety of application scenarios. like. Figure 3 As shown, the second feature fusion module 305 is finally used to perform feature fusion on the CLIP feature map and the feature map fused for the first time, and a fused feature map 306 is obtained. Therefore, it can be seen that the enhancement and fusion module 101 fuses the feature maps from the depth image corresponding to the food image at different scales, and further fuses the feature maps containing the image text description information extracted by the visual-language contrast pre-training model at different scales, so as to achieve the purpose of multimodal fusion and enhancement of nutrition-related semantic information in the features. Through this fusion, the fused feature map obtained not only contains information from the color image and the depth map, but also integrates rich semantic information from the text description, which greatly improves the correlation between the features and food nutrition.

[0022] Figure 4 FIG. 1 is a schematic diagram of the structure and workflow of the feature integration module 102 according to an embodiment of the present invention. Figure 4 As shown, the feature integration module 102 includes a position encoding module 401, a self-attention mechanism 402, a multi-layer perceptron 403, and a global average pooling module 404. The above modules of the feature integration module 102 are described in detail below.

[0023] First, the position encoding module 401 performs position encoding on the input fusion feature map to enhance its position information, and uses the self-attention mechanism 402 to mine and integrate the multi-scale information of the fusion feature map according to the position information to obtain the integrated feature map. In some embodiments, the self-attention mechanism 402 includes a spatial self-attention mechanism 405 and a hierarchical self-attention mechanism 406. The spatial self-attention mechanism 405 is used to capture and enhance the interaction between different spatial positions of the fusion feature map, and to mine the multi-scale information between different spatial positions of the fusion feature map. In some embodiments, residual connections are introduced in the spatial self-attention mechanism to increase information flow and model stability. Residual connection is a technology used to build deep neural networks in deep learning. The core idea of ​​residual connection is to allow signals in the network to bypass one or more layers and pass directly, which can solve the degradation problem in deep neural network training, that is, as the number of network layers increases, the performance of the network decreases instead. In a standard deep learning model, each additional layer should theoretically be able to learn more complex features, but in reality, as the number of layers increases, the gradient may become very small (gradient vanishing problem) or very large (gradient explosion problem) during back propagation, which makes the network difficult to train. Residual connections solve this problem by introducing a path that directly connects the input and output, so that the network can maintain signal transmission by simply learning the identity mapping, thereby avoiding the problem of gradient vanishing or exploding. Residual connections 407 are introduced into the spatial self-attention mechanism, that is, the input information in the spatial self-attention mechanism is added to the output of the spatial self-attention mechanism as the input of the hierarchical self-attention mechanism. In some embodiments, before the input of the spatial self-attention mechanism, layer normalization processing can also be performed, that is, a normalization technique used to stabilize the neural network training process does not depend on the size of the batch, but normalizes all features in a single sample. Then, the hierarchical self-attention mechanism 406 is used to continue processing to integrate and understand information between different scales to obtain an integrated feature map. In some embodiments, the information integrity is maintained by residual connection 407 in the hierarchical self-attention mechanism, and its principle is the same as that of the spatial self-attention mechanism, which will not be described in detail here. In some embodiments, layer normalization processing can also be performed before inputting the hierarchical self-attention mechanism. Then, in order to improve the feature expression and introduce nonlinear transformation, a multi-layer perceptron 403 is used to process the integrated feature map. Figure 4 As shown, in some embodiments, information integrity is maintained in the multi-layer perceptron through residual connection 407, and its principle is the same as the spatial self-attention mechanism, which will not be described in detail here. Figure 4As shown, in some embodiments, layer normalization processing can also be performed before inputting into the multi-layer perceptron. Finally, the feature map is converted into a feature vector with reduced dimension using the global average pooling module 404. Therefore, as mentioned above, the feature integration module 102 uses a variety of self-attention techniques to make the model pay attention to the feature information of multi-modality, multi-scale and different spatial positions, and integrates and reduces the high-dimensional complex feature map output by the enhancement and fusion module for subsequent regression tasks.

[0024] In some embodiments, the feature vector output by the feature integration module is input into the nutrition regression head 103 to obtain detailed nutritional value. In one embodiment, the nutritional value includes calories, fat content, carbohydrate content, and protein content. The nutrition regression head 103 is a specific network structure part in the deep learning model for regressing from food images to nutritional estimation, which is responsible for regressing various nutritional components from the feature vectors generated in the previous steps. Specifically, the nutrition regression head 103 is a deep learning network structure for predicting food nutritional components (such as calories, fat content, carbohydrate content, protein content, etc.). The nutrition regression head 103 can be a variety of deep learning network structures. For example, in order to improve the accuracy of food nutritional assessment, a multi-task learning mechanism (i.e., a multi-task learning architecture) can be introduced. Multi-task learning refers to learning multiple related tasks simultaneously in a unified model, which can help the model generalize better and improve the performance of each task. According to one embodiment of the present invention, in addition to the main nutritional assessment task, the nutrition regression head 103 also introduces an auxiliary task: weight estimation. By simultaneously predicting the weight of the food, the model can be helped to estimate the nutritional components more accurately, because the nutritional components are usually related to the weight of the food. According to one embodiment of the present invention, the nutrition regression head 103 is a multi-task learning network architecture, including a Dropout layer, a multilayer perceptron (MLP) with shared parameters, and 5 MLPs with the same structure, which are responsible for the regression tasks of specific nutrients. First, the purpose of introducing the Dropout layer in the nutrition regression head 103 is to reduce overfitting and improve the robustness of the model by randomly discarding a part of the neurons. Then, the MLP with shared parameters is used to extract common features that are beneficial to all tasks, while each independent MLP performs precise regression for specific nutrients, thereby achieving accurate prediction of nutrients. In addition, through hard parameter sharing, the model can promote each other with auxiliary tasks such as weight estimation while learning to predict nutrients, further improving the generalization ability of the model. It is a well-known technology to use the nutrition regression head to predict nutrients, which will not be described in detail here.

[0025] 2. Model Training 1. Dataset The model proposed in the embodiment of the present invention is trained on the vision-based food nutrition assessment dataset Nutrition5k, which covers more than 250 food categories and has detailed annotations of multiple nutrients (i.e., real labels). Nutrition5k consists of food videos shot from multiple angles and color-depth image pairs captured by a depth camera directly above the food. In order to ensure data integrity and prevent information leakage, the ratio of the training set to the test set of the dataset is 5:1.

[0026] 2. Training process The image pairs consisting of the color food images and their corresponding depth images in the training set are input into the nutrition detection model of the embodiment of the present invention for layer-by-layer calculation to predict the nutritional value of the food in the food image. The predicted result and the true label are calculated through the loss function to obtain the loss value. The gradient of the loss function with respect to each parameter is calculated by the chain rule. Starting from the output layer, propagate backward layer by layer to calculate the gradient of each layer. Use the calculated gradient to update the network parameters. Repeat the above steps until the preset number of iterations is reached.

[0027] 3. Loss Function In some embodiments, the Lp loss function is used to perform model training. The loss function is based on the classic Lp norm and is defined as follows: in, Indicates the preset number of iterations; Indicates The output of the model after iterations; Indicates the The true label corresponding to the iteration; 𝑝 represents the preset parameter. By introducing the parameter 𝑝, the sensitivity of the loss function to the error is adjusted, so that the model can focus more on the key deviations in the prediction. The value range of the parameter 𝑝 is (0,1], which allows the model to flexibly balance the relationship between linear error and nonlinear error when dealing with errors. When 𝑝 is close to 0, the model tends to focus on large errors; when 𝑝 is close to 1, all errors are treated more equally.

[0028] The Lp loss function first calculates the absolute error between the model output and the target value, then performs a custom nonlinear transformation on the error, and finally sums the processed error through an exponential function to achieve a comprehensive evaluation of the error. The nonlinear transformation is accomplished by scaling the logarithm of the error and applying an exponential function, which helps to avoid over-sensitivity to extreme values ​​while maintaining the sensitivity of the loss function. In order to ensure numerical stability and avoid potential numerical problems, clipping operations are introduced in the loss calculation process to ensure that the loss value is within a reasonable range.

[0029] In addition, the training of the model of the embodiment of the present invention adopts the adaptive moment estimation method (Adaptive MomentEstimation, Adam) as the optimization algorithm, the initial learning rate is set to 0.0001, and the weight decay is also set to 0.0001. The learning rate adjustment strategy adopts the cosine annealing method to gradually reduce the learning rate to achieve better training effect. The whole training process includes 200 training cycles, of which the first 15 cycles are used as warm-up periods. During the warm-up period, the learning rate gradually increases from 0 to the set initial value. The batch size of the training is set to 16. Lp is used as the loss function, where the value of 𝑝 is set to 0.8 to adjust the sensitivity of the loss. In addition, considering that the feature integration layer integrates different dimensions of the feature map, in order to avoid overly complex design leading to model degradation, bitwise addition can be selected as the fusion method of the semantically enhanced multi-scale multimodal fusion module. Among them, bitwise addition refers to adding pixels or feature vectors at the same position of multiple feature maps according to corresponding positions. This is a well-known technology and will not be repeated here.

[0030] For data augmentation strategies, synchronous random horizontal flipping and synchronous random vertical flipping are used to ensure that the input color image and depth image remain consistent when flipped to reduce the noise introduced by asynchronous flipping. In addition, multi-resolution training is introduced, that is, each training batch randomly selects a resolution size from six preset sizes (256×352, 288×384, 320×448, 352×480, 384×512, 480×640). This method not only enhances the adaptability of the model to inputs of different sizes, but also improves its overall robustness.

[0031] 3. Experimental Results and Analysis 1. Evaluation indicators In the task of visual-based food nutrition assessment, the percentage mean absolute error (PMAE) is often used as the evaluation indicator for each value predicted by the method. It is defined as the mean absolute error between the predicted value and the true value divided by the mean of the true value, and the formula is: in, represents the predicted value, Represents the true value. PMAE provides an intuitive way to measure the level of error. The smaller the value, the more accurate the evaluation result. In order to comprehensively evaluate the predictive ability of the method for all nutrients, the mean PMAE of all nutrients is used to measure the accuracy of the visual-based nutritional assessment method.

[0032] Table 1 2. Backbone network selection In order to select a backbone network that is more suitable for the vision-based nutrition assessment task, all networks adopt a dual-branch structure to process color images and depth maps, and only use the output of the last classification layer of the network as the feature representation, and realize the fusion of multimodal data through feature splicing. The experimental results are shown in Table 1. According to the experimental results, the network based on visual Transformer shows better performance in the nutrition assessment task. This reflects that in the vision-based nutrition assessment task, the long-range feature dependency in the image is more critical than the local texture feature. Swin Transformer is a classic computer vision model based on Transformer. Swin Transformer adopts a hierarchical structure to adapt to image processing of different scales and introduces a window self-attention mechanism to effectively reduce the computational complexity. The design of the model allows it to process images through four stages, and each stage further refines the image representation through Patch Merging and Swin Transformer Block, thereby optimizing the information conversion from pixel to patch (image block) level. This structure not only retains the hierarchical nature of CNN, but also enhances the model's ability to handle global dependencies through window-based self-attention and moving window self-attention mechanisms. It can gradually process and integrate local and global information in the image, which is crucial for understanding the complex visual features of food and associating them with nutrition. Therefore, Swin Transformer is selected as the backbone network.

[0033] 3. Results Analysis Table 2 The performance of various methods on RGB images and depth images of the Nutrition5k dataset is compared, as shown in Table 2. It can be seen that the model based on the visual Transformer (i.e., OpenSeeD-Transformer) is particularly good at capturing long-distance dependencies in images, and it performs particularly well in predicting the fat content and weight of food, reaching 21.4% and 8.87% accuracy, respectively. IMIR-Net uses a convolutional neural network to achieve an average nutritional prediction performance that exceeds that of the visual Transformer. It also uses CLIP to semantically enhance the features extracted by the backbone network, using additional food raw material text as input to CLIP. The embodiment of the present invention achieves the best average nutritional prediction performance, reaching 18.2%, while the PMAE values ​​of energy, fat, carbohydrates, and protein reach 12.8%, 21.7%, 19.8%, and 18.4%, respectively.

[0034] Accurate prediction of food weight is crucial for nutritional assessment. Even without using an additional segmentation network, the embodiment of the present invention achieved a competitive result of 9.04% in food weight prediction. The PMAE with all indicators combined still achieved the best evaluation performance, proving the effectiveness of the embodiment of the present invention.

[0035] 4. Ablation experiment The model proposed in the embodiment of the present invention is mainly composed of a semantically enhanced multi-scale and multi-modal enhancement and fusion module and a feature integration module. In order to prove the effectiveness of each part and explore the role of more sophisticated modules, a total of four groups of ablation experiments are designed, as shown in Table 3. For methods that do not use CLIP semantic enhancement, the features are directly integrated after multi-modal multi-scale fusion, and the multi-scale features from CLIP are no longer fused. For methods that do not use multi-scale information, the image and depth backbone networks and the CLIP pre-trained ResNet101 only use the last layer of features for fusion and enhancement, and no longer use multi-scale features and two feature pyramid networks. For methods that do not use depth information, only the features of the image backbone network and CLIP will be used. For methods that do not use the feature integration layer, the feature map will directly perform global average pooling to obtain the feature vector, and self-attention will no longer be calculated.

[0036] Table 3 In the ablation experiment, the model proposed in the embodiment of the present invention still achieved the best performance in the average PMAE and overall PMAE of nutritional evaluation, which were 18.17% and 16.34%, respectively. After removing the CLIP semantic enhancement, the model's prediction of energy content was slightly improved (from 12.83% to 12.77%), while the protein content prediction decreased more significantly (from 18.40% to 18.96%), resulting in a slight increase in the average value of its nutritional content prediction error (from 18.17% to 18.33%). This phenomenon may indicate that CLIP-based semantic enhancement plays a key role in understanding complex semantic information, such as food types and ingredients, and its absence causes the model to have a lower accuracy in predicting specific nutrients. In addition, when multi-scale information is not used, the model's performance in predicting fat content is slightly improved (from 21.68% to 21.14%), while the prediction of other nutrients is not significantly affected. This may indicate that multi-scale information is less important for capturing the details of fat content, or it may be caused by the confrontation between multiple tasks. In general, multi-scale information does not improve the task much, which may be due to the fact that the multi-resolution training used in the training played a certain role as a substitute. Without depth information, the prediction performance of all nutrients decreased significantly, especially the prediction of energy content increased from 12.83 to 14.68, which is significantly higher than other methods. This result strongly suggests that depth information is essential for a comprehensive understanding of the three-dimensional structure of food, and its absence causes the model to be unable to accurately capture the overall nutritional value of food. The estimation of weight also increased accordingly (from 9.04% to 10.48%), which may be because the lack of depth information causes the model to overestimate the volume and weight of food. Finally, when the feature integration layer is not used, a decrease in the performance of carbohydrate prediction is observed (from 19.77% to 20.57%), indicating that the effective fusion of different nutritional information by the feature integration layer is the key to improving prediction accuracy. Although this configuration has little effect on the prediction of other nutrients, its slight decrease in weight prediction performance (from 9.04% to 9.44%) may reflect the importance of feature integration in maintaining the consistency of the overall model estimate. These ablation results reveal the important contributions of CLIP semantic enhancement, multi-scale information, depth information, and feature integration layers to improving the prediction accuracy of food nutritional composition and weight estimation. The removal of each component affects the performance of the model in a unique way, emphasizing the necessity of comprehensively utilizing these features to achieve excellent prediction results.

[0037] According to one embodiment of the present invention, a food nutrition detection method based on multimodal fusion is provided, including: inputting the food image and its depth image into the above-mentioned trained nutrition detection model respectively, and the nutrition detection model outputs the nutritional value of the food. A food nutrition assessment method based on multimodal fusion proposed in the present invention can improve the accuracy of the model for assessing nutritional components from food images. In some implementations, before feature extraction of the image, the food images in the data set can be preprocessed, including operations such as cropping, scaling and normalization of the image, to adapt to the subsequent detection model.

[0038] The embodiment of the present invention proposes a food nutrition assessment method based on multimodal fusion, which enhances the nutrition-related semantic information in the features by introducing multimodal information such as depth images and image description texts, and uses self-attention technology to integrate multimodal, multi-scale and feature information of different spatial positions, thereby significantly improving the accuracy of the food nutrition assessment model.

[0039] According to one embodiment of the present invention, there is provided a food nutrition detection device, comprising: a processor and a storage device, wherein the storage device is used to store one or more programs, and when the one or more programs are executed by the processor, the food nutrition detection device implements the steps of the above-mentioned food nutrition detection method.

[0040] It should be noted that although the above describes the various steps in a specific order, it does not mean that the various steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0041] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0042] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a protruding structure in a groove on which instructions are stored, and any suitable combination thereof.

[0043] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A food nutrition detection model, comprising: Feature extraction module, enhancement and fusion module, feature integration module and nutrition regression head; The feature extraction module is used to extract features from the food image to obtain a visual feature map, and to extract features from the depth image of the food image to obtain a depth feature map; The enhancement and fusion module is used to extract a text description feature map of the food image according to the food image, and obtain a fusion feature map according to the text description feature map of the food image, the visual feature map and the depth feature map; The feature integration module is used to integrate and reduce the dimension of the fused feature map to obtain a feature vector after dimension reduction; The nutrition regression head is used to predict the nutritional value of the food in the food image according to the feature vector after dimension reduction.

2. The detection model according to claim 1, wherein: The enhancement and fusion module includes a first feature pyramid network, a second feature pyramid network, a first feature fusion module, a second feature fusion module and a visual-language contrast pre-training model; The first feature pyramid network is used to refine the visual feature map to obtain a first refined feature map; The second feature pyramid network is used to refine the depth feature map to obtain a second refined feature map; The first feature fusion module is used to fuse the first refined feature map with the second refined feature map to obtain a first fused feature map; The visual-language contrast pre-training model is used to obtain a text description feature map of the food image according to the food image; The second feature fusion module is used to fuse the first fused feature map with the text description feature map to obtain the fused feature map.

3. The detection model according to claim 2, wherein: The visual-language contrast pre-trained model includes a visual encoder; The visual encoder is used to obtain a comparative language-image pre-trained feature map according to the food image; The comparative language-image pre-trained feature map includes a text description feature map of the food image.

4. The detection model according to claim 1, wherein: The feature integration module includes a position encoding module, a self-attention mechanism and a global average pooling module; The position encoding module is used to obtain the position information of the fused feature map according to the fused feature map; The self-attention mechanism is used to mine and integrate the multi-scale information of the fused feature map according to the position information to obtain an integrated feature map; The global average pooling module is used to convert the integrated feature map into the feature vector after dimensionality reduction.

5. The detection model according to claim 6, wherein: The self-attention mechanism includes a spatial self-attention mechanism and a hierarchical self-attention mechanism; The spatial self-attention mechanism is used to mine multi-scale information between different spatial positions of the fused feature map according to the position information; The hierarchical self-attention mechanism is used to integrate multi-scale information between different spatial positions of the fused feature map to obtain the integrated feature map.

6. The detection model according to claim 5, wherein: Residual connections are introduced into the spatial self-attention mechanism and the hierarchical self-attention mechanism respectively.

7. The detection model according to claim 5, wherein: The feature integration module also includes a multi-layer perceptron, which is used to process the integrated feature map before the global average pooling module.

8. The detection model according to any one of claims 1 to 7, wherein: The nutritional component values ​​predicted by the nutritional regression head include calories, fat content, carbohydrate content and protein content.

9. A method for detecting food nutrition, wherein: include: Inputting a food image and a depth image of the food image into a nutrition detection model as described in any one of claims 1 to 8; The nutrition detection model outputs the nutrition value of the food in the food image.

10. A food nutrition testing device, characterized in that: include: processor; A storage device for storing one or more programs, when the one or more programs are executed by the processor, enables the food nutrition detection device to implement the steps of the method as claimed in claim 9.

Citation Information

Cited By

  • Bread food detection system and method

    CN120707495A

  • CLIP-based diet recognition and nutrition analysis method

    CN120766884A

  • A method for estimating food volume or weight in a container based on artificial intelligence deep learning and computer vision

    TWI941254B