Intelligent food identification and nutrition analysis method based on deep learning

Through deep learning combined with feature fusion methods of depth estimation and three-dimensional point cloud generation, the problem that traditional two-dimensional image processing methods cannot extract three-dimensional structure information of food is solved, and accurate prediction of food weight and calories is achieved, and intelligent nutrition management and health monitoring is supported.

CN120279548AActive Publication Date: 2025-07-08ZHEJIANG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510748959.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional two-dimensional image processing methods cannot effectively extract and predict physical properties such as volume, weight and calories in food images, and lack three-dimensional structural information.

Method used

Deep learning method is used to combine depth estimation and three-dimensional point cloud generation, and two-dimensional image and three-dimensional point cloud information are combined through feature fusion technology. Features are extracted using Mask R-CNN, MiDaS depth estimation model and PointNet model, and food weight and calories are predicted through a multi-layer perceptron.

Benefits of technology

Improves the accuracy of predicting physical properties such as food volume, weight and calories, and enables automated estimation of nutritional content in food, helps to develop a scientific diet plan and provide personalized nutrition advice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279548A_ABST
    Figure CN120279548A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent food identification and nutrition analysis method based on deep learning, and belongs to the technical field of image analysis. The method comprises the following steps: 1, collecting input data, and preprocessing the data; 2, determining a position and a mask area of a target in the scene, and generating a dense depth map; 3, constructing a food three-dimensional point cloud, denoising the three-dimensional point cloud, and removing abnormal points; 4, obtaining a three-dimensional feature, extracting a two-dimensional feature from the image, and carrying out weight fusion; 5, splicing the obtained features, and reinforcing the fused feature information; 6, predicting the weight and calorie of the food by regression; and 7, optimizing model prediction and verifying a prediction effect. According to the method, feature fusion in a 2D + 3D modal mode is adopted, data enhancement is performed on an original image, 2D features are extracted, a 3D point cloud is generated at the same time, and multi-modal feature fusion is performed to predict the calorie and the weight of the food, so that consumers can be helped to accurately estimate indexes such as the calorie and the weight of the food and formulate scientific diet methods and suggestions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image analysis, and particularly relates to an intelligent food recognition and nutrition analysis method based on deep learning. Background Art

[0002] In the field of food image analysis, traditional methods mainly focus on aspects such as classification, detection, segmentation, and recognition of two-dimensional images. With the development of deep learning technology, especially the widespread application of convolutional neural networks, the accuracy and efficiency of image analysis have been significantly improved. However, food images not only contain rich visual information such as color, shape, and texture, but their physical properties such as volume, weight, and calories often require more complex modeling methods. To effectively extract and predict physical properties such as volume, weight, and calories from food images, traditional two-dimensional image processing methods may not provide sufficient spatial information. For this reason, in recent years, multi-modal learning methods based on depth estimation, three-dimensional point cloud generation, and feature fusion have become a research hotspot.

[0003] Depth estimation technology generates a depth map by inferring the distance between each pixel and the camera from a 2D image. The depth map provides a dimension related to the spatial geometry of the objects in the image. These depth information can be generated by depth estimation algorithms. A 3D point cloud is a data structure composed of multiple points with spatial coordinates, which can effectively describe the three-dimensional shape of an object. Through the depth map, each pixel in the two-dimensional image can be mapped to three-dimensional space to obtain a set of point cloud data, and further extract the geometric features of the object. For two-dimensional images, using a convolutional neural network is the most common feature extraction method. A CNN can learn local features in the image such as edges, textures, colors, etc. as well as global features. However, relying solely on the features of two-dimensional images cannot fully reflect the three-dimensional structure of food. Therefore, it is necessary to combine other information such as depth maps and point clouds to improve the prediction ability of the model. Traditional image feature extraction methods are not sufficient to process the three-dimensional information of food images. With the gradual popularization of 3D data, researchers have proposed various point cloud processing networks, such as PointNet and PointNet++. These networks can process the spatial positions of each point in the point cloud data and effectively extract three-dimensional features of food objects, such as volume, shape, surface roughness, etc. In practical applications, multi-modal learning methods that combine 2D and 3D features have received increasing attention. By fusing features from two-dimensional images and three-dimensional point clouds, it is possible to better capture the appearance and form of food, especially foods with complex shapes. For example, by adding an attention mechanism or multi-scale feature learning method to the model, the network can automatically select the most relevant 2D and 3D features for fusion in different tasks.

[0004] For food images, feature fusion techniques can combine the visual features of two-dimensional images with three-dimensional spatial information. Through feature fusion, the understanding of the object's appearance and geometric shape by the model can be enhanced, thereby improving the prediction ability of physical attributes such as food volume, weight, and calories. After completing feature extraction and fusion, the next task is to predict physical attributes such as volume, weight, and calories. Usually, a regression model such as a multi-layer perceptron (MLP) will input the fused features into a neural network and train them through a regression loss function such as mean squared error (MSE), and finally output the predicted values of physical attributes.

[0005] This method has broad application potential in multiple fields. It can help consumers estimate the calories, nutritional components, volume, and weight of food, so as to formulate a scientific diet plan. It can automatically estimate the quantity, weight, etc. of food during the production process and improve production efficiency. Through the automatic detection and analysis of food by intelligent devices, it can help doctors and dietitians provide personalized diet advice for patients. With the continuous progress of deep learning and computer vision technologies, multi-modal learning methods that combine depth estimation, 3D point cloud generation, feature extraction, and regression prediction will become the mainstream of research. These methods will be increasingly applied in fields such as food image analysis, intelligent nutrition management, and health monitoring, promoting the development of intelligent catering, food safety, and health management. Summary of the Invention

[0006] Aiming at the above problems existing in the prior art, the purpose of the present invention is to provide an intelligent food recognition and nutrition analysis method based on deep learning.

[0007] The present invention provides the following technical solutions: An intelligent food recognition and nutrition analysis method based on deep learning, including the following steps: S1. Collect scene images as the original input data and preprocess the collected image data; S2. Determine the position and mask area of the target in the scene and generate a dense depth map of the scene; S3. Construct a 3D point cloud of the food, denoise the 3D point cloud, and remove abnormal points; S4. Obtain 3D features representing the food spatial structure, and use different CNN models to extract 2D features from the image and perform weight fusion to obtain fused 2D features; S5. Concatenate the obtained 3D features and 2D fused features, and use a self-attention mechanism to strengthen the fused feature information; S6. Based on the fused feature information, use a multi-layer perceptron model to regress and predict the food weight and calories; S7. Optimize the model prediction and verify the prediction effect.

[0008] Further, the specific process of S1 is as follows: Collect the scene image through the RGB camera as the input data, ensuring that the image quality is clear, the illumination is uniform, and the background is simple, and store the obtained image in the image database. Use the image processing tool to perform Gaussian filtering and noise reduction on the captured image; uniformly adjust the image size to the model input size and perform normalization processing to make the pixel values within a unified range.

[0009] Further, the specific process of S2 is as follows: S2.1. Input the preprocessed image into the Mask R-CNN model to identify the food category in the image and determine the specific location and precise code area of the target; S2.2. Adopt the MiDaS depth estimation model to predict the relative depth value of each pixel in the image through the deep convolutional neural network and generate a dense depth map of the scene.

[0010] Further, the specific process of S3 is as follows: Based on the mask determined in S2 and the generated depth map, use the camera internal parameters in the intelligent device to calculate the three-dimensional spatial coordinates of each pixel of the target, construct the food three-dimensional point cloud, and perform denoising processing on the constructed three-dimensional point cloud. Use wavelet transform to remove abnormal points and improve the quality of the point cloud data.

[0011] Further, the specific process of S4 is as follows: S4.1. Adopt the PointNet model to extract features from the denoised three-dimensional point cloud to obtain three-dimensional features that can accurately represent the food spatial structure; S4.2. Augment the original image data to improve the diversity of the data and the generalization ability of the model; S4.3. Respectively use different CNN models to extract two-dimensional features from the augmented image, and perform weight fusion on the extracted features of different models to obtain the fused two-dimensional features.

[0012] Further, the specific process of S5 is as follows: Use the fused features obtained in S4 as the input, combine with the existing food database, construct the training data samples, use the real food weight and calories as the labels, and input the constructed training data samples into the multi-layer perceptron MLP to regressively predict the food weight and calories.

[0013] By adopting the above technologies, compared with the prior art, the beneficial effects of the present invention are as follows: The present invention adopts a feature fusion method in two modalities of 2D + 3D to perform data enhancement on the original image, extracts 2D features using different deep learning models, and simultaneously generates 3D point clouds by depth estimation and masked images. Finally, multi-modal feature fusion of 2D and 3D is performed to predict the calorie and weight of food. Through this method, consumers can be helped to accurately estimate indicators such as the calorie and weight of food, so as to formulate scientific diet methods and suggestions. Description of the Drawings

[0014] Figure 1 is the algorithm flowchart of the present invention; Figure 2 is the flowchart of the mobile phone online food recognition and analysis of the present invention; Figure 3 is the flowchart of the online and offline system of the present invention. Detailed Description of the Invention

[0015] This case provides an intelligent food recognition and nutrition analysis method based on deep learning, and the algorithm process is as Figure 1 shown, Figure 2 is the flowchart of the mobile phone online food recognition and analysis, Figure 3 is the flowchart of the overall online and offline system. The detailed implementation method includes the following steps: Step 1. Obtain the original image data. Collect the scene image through an RGB camera as the input data, ensure that the image quality is clear, the illumination is uniform, and the background is simple, and store the obtained image in the image database. A total of 513 images of 11 different angles are collected.

[0016] Step 2. Use the image processing tool to perform Gaussian filtering denoising on the captured image; uniformly adjust the image size to the model input size and perform normalization processing to make the pixel values within a unified range.

[0017] Step 3. Object detection and instance segmentation. Input the preprocessed image into the Mask R-CNN model to identify the food category in the image and determine the specific position and precise mask area of the target. Equation (1) is: ; (1) where is the classification loss, used to identify the target category; is the bounding box regression loss, determining the specific position of the target; is the mask prediction loss, used to determine the precise contour of the target.

[0018] Step 4. Adopt the MiDaS depth estimation model to predict the relative depth value of each pixel in the image through a deep convolutional neural network to generate a dense depth map of the scene. Equation (2) is: ; (2) where , represents the depth value predicted by the network at the i,j-th pixel, , represents the true depth value at the i,j-th pixel.

[0019] Step 5. Based on the mask in Step 3, the depth map obtained in Step 4 and the corresponding depth value Z, then determine the ray direction of each pixel in the camera coordinate system with the help of the internal parameters of the RGB camera as shown in Equation (3). Finally, multiply the direction vector by the corresponding depth value to stretch the two-dimensional pixel points into the actual three-dimensional space, and then obtain the set of spatial points on the food surface; aggregating these points constitutes the original three-dimensional point cloud of the food. Equation (3) is: ; ; ; (3) where x,y represent the pixel coordinates in the original image, C x、 C y modify to represent the horizontal and vertical coordinates of the image center in pixel coordinates, f x、 f y respectively represent the equivalent focal lengths of the camera in the x and y directions, X, Y are the image plane coordinates after removing the principal point offset and unifying the focal length differences in the horizontal and vertical directions, and Z is the depth value.

[0020] Step 6. Denoise the three-dimensional point cloud generated in Step 4, use wavelet transform to remove outliers and improve the quality of the point cloud data. The processed point cloud is represented as Equation (4): ; (4) where is the original three-dimensional point cloud, W is the wavelet transform function, T is the threshold function, is the inverse wavelet transform function.

[0021] To evaluate the quality of the generated three-dimensional point cloud, select the error from the actual volume label as the evaluation index as shown in Equation (5), and compare it with the traditional three-dimensional point cloud generation methods stereo vision and the deep learning method U-Net. As shown in Table 1, the results show that the three-dimensional point cloud calculated by Mask-RCNN and depth estimation Midas has the best effect; ; (5) where is the predicted volume, is the true volume; Table 1 Quality assessment of point clouds generated by different methods 。

[0022] Step 7. Use the PointNet model to extract features from the denoised 3D point cloud to obtain 3D features that can accurately represent the spatial structure of the food. The experimental parameters are set as shown in Table 2. The Point aggregation feature is given by Equation (6): ; (6) where represents the local feature extraction function for the point in the point cloud, represents the feature transformation function; Table 2 PointNet model training parameter settings 。

[0023] Step 8. Perform data augmentation such as rotation, translation, and scaling on the original image to increase data diversity and improve the generalization ability of the model.

[0024] Step 9. Since the image dataset is large, three different CNN models, namely EfficientNet-B0, ResNet18, and MobileNetV2, are used to extract 2D features from the augmented images. The experimental parameters are set as shown in Table 3. The features extracted by different models in Step 8 are weighted and fused to obtain the fused 2D features. The feature fusion is given by Equation (7): ; (7) where , , represent the feature fusion weights of each model, satisfying: ; Table 3 CNN model training parameter settings 。

[0025] To compare the impacts of different models on the image dataset, a control experiment was conducted. The image dataset was divided into a training set and a validation set at a ratio of 8:2. The last fully connected layer of the model was changed to a linear layer for regression. The evaluation metrics were RMSE, MAE, and MAPE, and their calculation formulas are shown in (8)-(10). The experimental results are shown in Tables 2 and 3. Table 4 shows the comparison of weight predictions of different models, and Table 5 shows the comparison of calorie predictions of different models. In actual prediction, first, the EfficientNet-B0, MobileNetV2, and ResNet18 networks are used to extract multi-level visual features. Then, at the end of each model, global average pooling is used to compress the spatial features into a fixed-length vector, and then a fully connected layer is connected to output a continuous calorie value. The results show that the model fused by average weighting achieved more accurate predictions in terms of weight and volume: ;(8) ;(9) ;(10) where 、 、n represent the true value, the predicted value, and the total number of samples; Table 4 Comparison of weight predictions of different models on the validation set ; Table 5 Comparison of calorie predictions of different models on the validation set 。

[0026] Step 10. Concatenate the three-dimensional features extracted in Step 7 and the two-dimensional fusion features extracted in Step 9, and use the self-attention mechanism to strengthen the feature information after fusion. The attention mechanism is shown in formula (11): ;(11) where Q, K, and V represent the query vector, the key vector, and the value vector respectively, is the vector dimension.

[0027] Step 11. Use the fusion features in Step 10 as the input, combine with the existing food database, construct training data samples, and use the true food weight and calories as labels. Input the constructed training data samples into a multi-layer perceptron (MLP) to regressively predict the food weight and calories. The prediction formula (12) is as follows: ;(12) Step 12. Compare the predicted weight and calories of the model with the true labels. When the prediction error exceeds the set threshold (such as ±10%), these samples are recollected and added to the training data to improve the model accuracy.

[0028] Step 13. To analyze the role of the grouping method in the food prediction task, ablation experiments were conducted and compared for different modality methods respectively. As shown in Table 6, the results indicate that the same method has a better prediction effect on food weight than on calorie prediction, and the multi-modal method is overall better than the single-modal method; Table 6 Comparison of multi-modal and single-modal methods 。

[0029] Step 14. Deploy the trained model weights to the smartphone. Specifically: export the model file in the general format ONNX; then use the official conversion tool ONNX-TFLite to compress it into a.ptlite file and enable quantization during conversion to significantly reduce the volume and accelerate inference; then copy this lightweight model file into the assets directory of the Android project and introduce the corresponding inference library in the Gradle configuration; finally, instantiate the lightweight inference engine in the code to load the model file, so as to accurately deploy it to the smartphone. The above steps are the offline stage. When the user takes a new image, the convolutional network model in the offline stage is relatively large. To reduce the computational burden, a single CNN model is used to extract two-dimensional features. The captured image will be automatically input into the model deployed on the mobile phone through the preprocessing steps 2 to 11 for prediction, and the weight and calorie will be automatically estimated and displayed in real time, as Figure 2 。The stage where the user takes a new picture and uploads it for prediction is called the online stage, as Figure 3 。

Claims

1. An intelligent food recognition and nutrition analysis method based on deep learning, characterized in that, It includes the following steps: S1. Collect the scene image as the original input data and preprocess the collected image data; S2. Determine the position and mask area of the target in the scene and generate the dense depth map of the scene; S3. Construct the 3D point cloud of the food, denoise the 3D point cloud, and remove the abnormal points; S4. Obtain the 3D features representing the food spatial structure, and use different CNN models to extract 2D features from the image and perform weight fusion to obtain the fused 2D features; S5. Concatenate the obtained 3D features and 2D fused features, and adopt the self-attention mechanism to strengthen the fused feature information; S6. Based on the fused feature information, regressively predict the food weight and calories; S7. Optimize the model prediction and verify the prediction effect.

2. The intelligent food recognition and nutrition analysis method based on deep learning according to claim 1, wherein, The specific process of S1 is as follows: Collect the scene image through the RGB camera as the input data, store the obtained image in the image database, and use the image processing tool to perform Gaussian filtering denoising on the captured image; Uniformly adjust the image size to the model input size and perform normalization processing to make the pixel values within a unified range.

3. An intelligent food recognition and nutrition analysis method based on deep learning according to claim 1, characterized in that, The specific process of S2 is as follows: S2.

1. Input the preprocessed image into the Mask R-CNN model to identify the food category in the image and determine the specific position and precise code area of the target; S2.

2. Adopt the MiDaS depth estimation model to predict the relative depth value of each pixel in the image through the depth convolutional neural network and generate the dense depth map of the scene.

4. The intelligent food recognition and nutrition analysis method based on deep learning according to claim 3, characterized in that, The specific process of S3 is as follows: Based on the mask determined in S2 and the generated depth map, use the camera internal parameters in the intelligent device to calculate the 3D spatial coordinates of each pixel of the target, construct the 3D point cloud of the food, and perform denoising processing on the constructed 3D point cloud. Adopt wavelet transform to remove the abnormal points and improve the quality of the point cloud data.

5. The intelligent food recognition and nutrition analysis method based on deep learning according to claim 4, characterized in that, The specific process of S4 is as follows: S4.

1. Adopt the PointNet model to extract features from the denoised 3D point cloud to obtain the 3D features that can accurately represent the food spatial structure; S4.

2. Augment the original image data to improve the data diversity and the generalization ability of the model; S4.

3. Respectively use different CNN models to extract 2D features from the augmented image, and perform weight fusion on the extracted features of different models to obtain the fused 2D features.

6. The intelligent food recognition and nutrition analysis method based on deep learning according to claim 5, characterized in that, The specific process of S5 is as follows: Take the fused features obtained in S4 as the input, combine with the existing food database, construct the training data samples, use the real food weight and calories as the labels, and input the constructed training data samples into the multi-layer perceptron MLP to regressively predict the food weight and calories.

Citation Information

Patent Citations

  • Intelligent nutrition management method and system based on image processing

    CN116884572A

  • Method for analyzing lithology of tunnel face by fusing image and three-dimensional point cloud

    CN118898698A

  • Scene object classification method and system based on multi-view depth image

    CN119169288A

  • Imaging system for object recognition and assessment

    US20160150213A1