A method for calculating food calories based on binocular vision

CN121747097BActive Publication Date: 2026-09-18JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511814156.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-09-18
Estimated Expiration
2045-12-04

AI Technical Summary

Technical Problem

本发明待解决的技术问题是如何克服现有食物热量计算技术中存在的计算误差大、适应性差的问题

Benefits of technology

本发明通过引入边缘-视差协同网络,同步提取边缘与视差特征,利用边缘特征指导立体匹配,并以匹配结果反向优化边缘,形成正向促进循环,显著提升了弱纹理食物的视差计算精度。同时,根据食物形态自适应选择等效柱体法或点云重建法进行体积计算,有效应对规则与不规则形状食物的三维重建需求。结合轻量级分割与分类网络,精准识别食物细分品类及烹饪方式,并通过具备反馈校正与数据聚合学习机制的动态数据库,实现热量计算的个性化与持续优化。整体方案系统性地解决了弱纹理、不规则食物的分割与体积计算难题,最终实现热量误差≤10%,为家庭饮食管理、健康监测等场景提供了高效、可靠的技术支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747097B_ABST
    Figure CN121747097B_ABST
Patent Text Reader

Abstract

This invention discloses a food calorie calculation method based on binocular vision, belonging to the fields of computer vision and nutritional analysis technology. The method first uses a calibrated binocular camera to acquire food images and performs image preprocessing. Then, a semi-global block matching algorithm combined with weighted least squares filtering is used to calculate and optimize the disparity map, and an edge-disparity collaborative network is introduced to improve the matching accuracy of foods with weak textures. Subsequently, a U-Net network is used to segment the food region, and after edge optimization, an accurate contour is output; a lightweight classification network is used to identify the specific food category. Furthermore, based on the food shape, an equivalent cylinder method or point cloud reconstruction method is adaptively selected for 3D volume reconstruction, and the volume is converted to mass using a density database. Finally, a calorie database containing cooking method information is queried to obtain accurate calorie values. This invention significantly outperforms existing technologies, providing an efficient and reliable solution for family diet management and health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of computer vision, deep learning, and nutritional analysis, specifically to a method for calculating food calories based on binocular vision. Background Technology

[0002] With increasing health awareness, the demand for accurate calculation of daily food calorie intake is growing. Existing methods for calculating food calories are mainly divided into three categories: Manual recording method: Users manually input the food name and estimated weight through the APP, which relies on subjective judgment. The error is usually >30%, and it cannot distinguish sub-categories (such as white rice and brown rice). Monocular vision method: Estimating food volume based on a single 2D image, it relies on reference objects such as a standard plate and cannot handle occlusion or irregular shapes, such as stacked noodles or broccoli, with a volume error >20%; General binocular vision method: This method uses conventional binocular parallax to calculate depth, but it is not optimized for food scenes. Foods with weak textures (such as tofu and porridge) are prone to parallax holes, and the blurred boundary between the plate and the food leads to segmentation errors, resulting in a calorie error of >15%.

[0003] Furthermore, existing technologies do not take into account the impact of cooking methods on calories (e.g., fried foods have more than 50% more calories than boiled foods), and the database is static, which cannot adapt to the calorie differences of foods from different origins and varieties, thus limiting its practicality. Summary of the Invention

[0004] [Technical Issues] The technical problem to be solved by this invention is how to overcome the problems of large calculation errors and poor adaptability in existing food calorie calculation technologies.

[0005] [Technical Solution] To address the above issues, this invention provides a food calorie calculation method based on binocular vision, achieving the following objectives: resolving the segmentation and volume calculation accuracy problems of foods with weak textures and irregular shapes; accurately identifying food subcategories and cooking methods to improve classification accuracy; constructing a dynamic database with feedback correction and data aggregation learning mechanisms to achieve personalized and iterative optimization of calorie calculation; specifically, the database allows for manual correction of automatically calculated food categories or qualities, and updates exclusive density or calorie preference values ​​based on correction feedback data to achieve personalized adaptation.

[0006] In a first aspect, the present invention provides a method for calculating food calories based on binocular vision, comprising the following steps: S1. Image Acquisition and Preprocessing: The binocular images of the food in the plate are acquired using a calibrated binocular camera. Distortion correction, epipolar alignment, and image quality optimization are performed on the binocular images. The initial disparity map is calculated using a semi-global block matching algorithm, and the initial disparity map is optimized using weighted least squares filtering to obtain the optimized disparity map. S2, Food Region Segmentation: Input the image processed in S1 into a lightweight semantic segmentation network to obtain classification masks for food, plates, and background; perform edge optimization and post-processing on the classification masks to output accurate food region contours; S3. Food Classification and Recognition: Based on the outline of the food region, crop out the food sub-image from the original image and input the food sub-image into the lightweight classification network to obtain the specific category of the food; S4. Volume and Mass Calculation: Based on the optimized disparity map and the food region outline, calculate the true depth and physical size of the food, and then select the equivalent cylinder method or point cloud reconstruction method to obtain the volume of the food through three-dimensional reconstruction according to the food shape; query the pre-established density database according to the specific category of the food to obtain the corresponding density, and convert the volume of the food into the mass of the food. S5. Calorie Calculation: Based on the specific category and quality of the food, query the food calorie database that records cooking method information to obtain the final food calorie value.

[0007] Optionally, the calculation of the disparity map in step S1 further includes: using an edge-disparity collaborative network to calculate the disparity map, simultaneously extracting edge features and disparity features from the left and right images, using the edge features to guide the stereo matching process, and using the matching results to optimize the edge extraction in reverse. The edge-parallax collaborative network includes: A weighted twin backbone network is used to extract 1 / 4 resolution feature maps from the left and right images; The edge detection branch, consisting of at least two convolutional layers and a sigmoid activation function, is used to generate a 1 / 4 resolution edge feature map from the left image feature map. The fusion module is used to fuse the feature maps of the left and right images with the edge feature maps, and construct the initial cost volume through a one-dimensional correlation layer; The context pyramid module uses multi-scale convolutional kernels to extract prior information about the scene. The parallax encoder, consisting of 16 residual blocks and dilated convolutions with a dilation rate of 2, is used to extract deep parallax features. The residual pyramid decoder is used to perform multi-level upsampling and residual correction on deep disparity features, and output a full-resolution disparity map and edge map. The twin backbone network is a ResNet-50 model. The context pyramid module uses 1×1, 3×3, 5×5 and 7×7 convolutional kernels for multi-scale feature extraction. The residual pyramid decoder performs 4 levels of upsampling.

[0008] Optionally, the lightweight semantic segmentation network in step S2 is a U-Net structure.

[0009] Optionally, in step S2, edge optimization and post-processing specifically involve: using Canny edge detection to refine the edges of the food region mask, and using morphological closing operations to fill the mask region.

[0010] Optionally, the lightweight classification network in step S3 is the EfficientNet-B0 model.

[0011] Optionally, in step S4, the calculation of the true depth of the food is specifically achieved using the following formula:

[0012] in, For true depth, For camera focal length, The binocular baseline distance. For disparity values, The minimum disparity value offset.

[0013] Optionally, in step S4, the volume of the food is obtained through three-dimensional reconstruction using the equivalent cylinder method. Specifically, the projected area of ​​the food is combined with the height calculated from the depth information, and the actual volume of the food is approximated by the cylinder volume.

[0014] Optionally, in step S4, the volume of the food is obtained by three-dimensional reconstruction using a point cloud reconstruction method. Specifically, the optimized disparity map is reprojected into a three-dimensional point cloud; food point clouds are extracted from the three-dimensional point cloud based on the food region contour; Poisson surface reconstruction is performed on the food point cloud to generate a three-dimensional mesh model; and the volume of the three-dimensional mesh model is calculated as the volume of the food.

[0015] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described binocular vision-based food calorie calculation method.

[0016] Thirdly, the present invention provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to implement the steps of the above-described binocular vision-based food calorie calculation method when executing the computer program.

[0017] [Beneficial Effects] This invention introduces an edge-parallax collaborative network to simultaneously extract edge and parallax features. Edge features guide stereo matching, and the matching results are used to optimize the edges, creating a positive feedback loop that significantly improves the parallax calculation accuracy for foods with weak textures. Simultaneously, it adaptively selects either the equivalent cylinder method or point cloud reconstruction method for volume calculation based on the food's shape, effectively addressing the 3D reconstruction needs of both regular and irregular shaped foods. Combined with a lightweight segmentation and classification network, it accurately identifies food subcategories and cooking methods. Furthermore, through a dynamic database with feedback correction and data aggregation learning mechanisms, it achieves personalized and continuous optimization of calorie calculation. The overall solution systematically solves the challenges of segmentation and volume calculation for foods with weak textures and irregular shapes, ultimately achieving a calorie error of ≤10%, providing efficient and reliable technical support for scenarios such as family diet management and health monitoring. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is the overall flowchart provided by the present invention.

[0020] Figure 2 This is a comparison chart of the disparity map optimization effect provided by the present invention.

[0021] Figure 3 This is a flowchart of the edge-parallax collaborative network processing provided by the present invention. Detailed Implementation

[0022] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Example 1: This embodiment uses a common family diet management scenario as an example to demonstrate how the method of the present invention calculates the calorie value of food. Before the method of the present invention is executed, the binocular cameras need to be calibrated: using the Zhang Zhengyou calibration method, a 12×9 checkerboard with a grid side length of 20mm is used as the calibration board, and 20-30 sets of binocular images at different angles and distances are collected. By calling the stereoCalibrate function in the OpenCV library, the intrinsic parameter matrices of the left and right cameras (the focal lengths fx and fy of the left and right cameras, and the principal point coordinates cx and cy), distortion coefficients (radial distortion k1-k6, tangential distortion p1-p2), and extrinsic parameter matrices (the rotation matrix R between the two cameras and the translation vector T) are calculated and obtained. Then, the method of the present invention is used to calculate the calorie value of food. The specific steps are as follows, see below. Figure 1 : S1. Image Acquisition and Preprocessing: Using a calibrated binocular camera, binocular images (left image imgL, right image imgR) of food on a plate are acquired. The OpenCV `undistort` function is called to correct distortion in imgL and imgR using pre-stored calibration parameters, eliminating monocular image distortion. The `stereoRectify` function is used to calculate the correction mapping matrix, projecting the left and right images onto the same epipolar plane (ensuring that corresponding points are searched only in the horizontal direction during subsequent disparity matching, reducing computational complexity). The `remap` function generates the corrected binocular images `rectL` and `rectR`. Finally, OpenCV's `drawEpipolarLines` function is used to verify whether the epipolar lines are horizontally aligned. This step ensures that corresponding pixels are searched only in the horizontal direction during subsequent stereo matching, significantly reducing computational complexity.

[0024] To optimize the image quality of rectL and rectR, the image encoding format is first unified to the common JPG or PNG. Then, the image is scaled to a fixed size using a bilinear interpolation algorithm (the width and height of the image must be divisible by 64 and not exceed 1024 pixels). Next, a Gaussian filter with a 3x3 convolution kernel is used for smoothing and denoising, and then a 3x3 median filter is used to remove salt-and-pepper noise.

[0025] A disparity map is a horizontal distance mapping between corresponding pixels in the left and right eye images. A larger pixel value indicates that the point is closer to the camera (depth is inversely proportional to disparity). A semi-global block matching (SGBM) algorithm is used to calculate rectL and rectR. SGBM finds corresponding pixels in the left and right eye images through "block matching + global constraints." The block matching process is as follows: for each pixel in the left image, within a specified range on the same horizontal line in the right image (after epipolar alignment), the optimal matching point is found using "pixel block similarity" (such as the sum of absolute differences (SAD) and sum of squared differences (SSD)) to obtain the initial disparity. The global constraint process is as follows: the initial disparity is optimized through path cost aggregation (such as paths in 8 directions) to avoid disparity discontinuities caused by local optima (such as discontinuities in food edges). The SGBM algorithm can balance accuracy and real-time performance to obtain the initial disparity map. Subsequently, the initial disparity map was optimized using weighted least squares (WLS) filtering. The edge information of rectL in the left image was used to guide the smoothing process, which smoothed the noise in the flat area while preserving the sharpness of the food edges, and finally obtained a high-precision optimized disparity map disp.

[0026] See Figure 2 To better address the challenge of matching stereoscopic foods with weak textures, such as rice, an edge-disparity collaborative network can be used for disparity map calculation. This network can simultaneously extract edge and disparity features from the left and right images, explicitly guiding the stereo matching process using edge features, and simultaneously using the matching results as feedback to optimize edge extraction. This collaborative mechanism can effectively fill disparity holes in weakly textured regions, generating a more accurate disparity map (disp).

[0027] For the specific process of edge-parallax collaborative networks, please refer to [link / reference]. Figure 3 : (1) The preprocessed left and right eye RGB images (their width and height must be divisible by 64) are input into a Siamese backbone network with shared weights. In this embodiment, the backbone network is preferably a ResNet-50 model (with the last fully connected layer removed) to extract the left and right image feature maps with 1 / 4 resolution and 512 channels; (2) The feature map of the left image is processed through the edge detection branch. This branch consists of two convolutional layers and a sigmoid activation function, which ultimately generates an edge feature map with a resolution of 1 / 4. (3) The left image feature map and right image feature map obtained in step (1) are fused with the edge feature map obtained in step (2), and the initial cost volume (size is 1 / 4 resolution, depth is dmax / 4) is constructed through a one-dimensional correlation layer. (4) The left image feature map, the initial cost body and the edge feature map are spliced ​​together and fused through a 1×1 convolutional layer to form a 512-channel mixed feature, which is used to prepare for subsequent deep feature extraction. (5) The hybrid features are sequentially passed through the context pyramid module (using four types of convolution kernels, 1×1, 3×3, 5×5, and 7×7, to perform parallel convolution and fusion to extract multi-scale scene prior information) and the disparity encoder (containing 16 residual blocks and using dilated convolution with a dilation rate of 2), and finally output deep disparity features with 1 / 8 resolution and 1024 channels. (6) Finally, the deep disparity features are upsampled at four levels (D3→D2→D1→D0 are generated sequentially) and the residual is corrected through the residual pyramid decoder to gradually restore the full resolution. The final full resolution disparity map (D0) and the optimized full resolution edge map are output synchronously.

[0028] In this process, edge features are explicitly introduced into the stereo matching process to guide the construction and optimization of the cost volume, realizing the guiding role of edges in matching. Simultaneously, information generated during disparity decoding is fed back to the edge extraction path, achieving dynamic optimization of edge features and enabling reverse optimization of edges through matching results. This forms a closed-loop mechanism where edges and disparity mutually promote and synergistically optimize each other, which is the key to improving the matching accuracy of weakly textured regions in this invention.

[0029] Before performing food region segmentation, we first need to prepare the dataset required by the segmentation model: 1. Data collection: Take 1000+ sets of binocular images of food in different scenes (covering Chinese / Western food, different plate colors, and complex backgrounds such as tablecloths / tableware). Each set includes: the corrected left image (as segmentation input) and the labeled image (using LabelMe or VGG Image Annotator to manually label 3 types of regions: background (0), plate (1), and food (2). The labeling should be accurate to the edges of the food, such as vegetable leaves and rice grains).

[0030] 2. Data Augmentation: To avoid model overfitting, online augmentation is performed on the training set: Geometric transformations: random rotation (±10°), horizontal flip (synchronous flipping required for binocular images), scaling (0.8-1.2x); Pixel transformation: Randomly adjust brightness (±20%), contrast (±15%), and saturation (±15%) to simulate different lighting environments.

[0031] Train a lightweight semantic segmentation network using the prepared dataset.

[0032] In this embodiment, the preferred training parameter configuration for the lightweight U-Net model is as follows: The training framework is PyTorch or Tensorflow, with a batch size of 8 and 20-30 training epochs. The loss function is a weighted average of cross-entropy loss and Dice loss, preferably with weights of 0.5 and 0.5 respectively in this embodiment. The Dice loss function optimizes segmentation of small food regions (such as nuts) and effectively avoids class imbalance. The AdamW optimizer is used with an initial learning rate of 1e-4, employing a dynamic learning rate adjustment strategy, decaying to a factor of 0.5 every 5 epochs. The segmentation accuracy on the validation set must reach above 90%.

[0033] S2, Food Region Segmentation: The corrected left image rectL is input into a pre-trained lightweight semantic segmentation network. This network adopts a U-Net structure, which is particularly suitable for deployment in embedded devices (such as mobile phones and edge computing modules). It can quickly and accurately output the probability of each pixel belonging to the background, plate, or food, and generate the corresponding three-class mask image.

[0034] The initial food mask output by the model is refined to improve contour accuracy. Canny edge detection (with high and low thresholds set to 200 and 100 respectively) is used to extract precise edges of the food region. These edges are then fused and aligned with the contours of the initial mask, removing rough pixels from the mask edges and making them more closely resemble the true boundaries of the food (such as the edges of broccoli leaves). Subsequently, a morphological closing operation (dilation followed by erosion) using a 5x5 rectangular kernel is used to fill small holes in the food region caused by model misjudgments (such as gaps inside bread). Small noise areas (such as isolated points with an area <50 pixels) are filtered out using an area threshold, ultimately retaining the most important food areas and avoiding interference from background clutter.

[0035] After processing, the final output includes the masks for three types of regions: mask_bg (background), mask_plate (plate), and mask_food (food), as well as the outline coordinates of the food region: contour_food.

[0036] S3. Food Classification and Recognition: Based on the contour_food output in step S2, use the cv2.boundingRect function to obtain the minimum bounding rectangle (x, y, w, h) of the food region. Then, crop the food sub-image food_img = rectL[y:y+h, x:x+w] from rectL. Scale this sub-image to a standard size (e.g., 224x224 pixels) to fit the input requirements of the classification network.

[0037] Before inputting the food subimage into the lightweight classification network, we also need to prepare the dataset required by the classification model: 1. Build a dataset: cover common food categories (such as rice, noodles, beef, broccoli, apples, etc., 50-100 categories), and collect 200+ sub-images for each category (different angles, cooking methods).

[0038] 2. Data Augmentation: Similar to the data augmentation process before semantic segmentation, random cropping (preserving the main food subject) and Gaussian blur (simulating slight defocusing) are added to improve the robustness of the model.

[0039] Train a lightweight classification network using the prepared dataset.

[0040] In this implementation, the preferred training parameter configuration for the lightweight classification model is as follows: The loss function is the cross-entropy loss function, using the Adam optimizer. The initial learning is set to 5e-5, the batch size is 16, the training epochs are 15~20, and the Top-1 accuracy on the validation set needs to reach above 85%. For easily confused categories (such as rice vs. glutinous rice), the sample size can be increased separately.

[0041] This embodiment uses the EfficientNet-B0 model, which has been pre-trained with weights trained on the ImageNet dataset. 80% of the convolutional layers are frozen, retaining only the ability to extract general features. Transfer learning fine-tuning was performed using the large-scale food dataset prepared above. To adapt to the application scenario of this embodiment, the output dimension of the last fully connected layer of the model is changed to be consistent with the number of food categories, and the model finally uses the Softmax activation function.

[0042] The processed food sub-image is then fed into a pre-trained lightweight classification network for food category identification. This lightweight classification model reduces the amount of training data required, speeds up inference, and can output the specific category of food with high accuracy. For cases with multiple foods on a plate, this step can be repeated for each food item.

[0043] S4. Volume and Mass Calculation: First, calculate the true depth of the food, using the following method:

[0044] in, The actual depth of the object from the camera (unit: mm). The focal length of the left camera (obtained from the intrinsic parameter matrix calibrated by the binocular cameras, unit: pixels). The binocular baseline (the distance between the centers of the two lenses of the binocular camera, calculated from the translation vector T of the extrinsic parameters of the binocular camera) is the distance between the centers of the two lenses of the binocular camera. (Unit: mm) This represents the disparity value (in pixels) of a pixel in the disparity map. This is the minimum disparity value offset (usually 0).

[0045] The specific implementation process is as follows: extract the disparity value d_pixel of all pixels within the mask_food region (obtained from the optimized disparity map disp in step S1); for each pixel, calculate the true depth using the above formula. Filter outliers (such as...) (Values ​​>2000mm or <100mm are considered outside the effective measurement range); calculate the average depth of the food area. To avoid the influence of individual pixel noise, the distance from the food subject to the camera is considered.

[0046] Count the total number of pixels in mask_food (That is, the number of pixels with a value of 2 in the statistical mask), combined with the physical width of the camera sensor and the image pixel width, calculate the physical size of a single pixel, and then obtain the true projected area of ​​the food. (Unit: mm²)

[0047] The specific process is as follows: using the calibration parameters of the binocular cameras, the pixel spacing of the left camera... ,in The width of the camera sensor can be obtained from the camera's specification sheet; for example, a 1 / 2.3-inch sensor corresponds to a width of approximately 6.17mm. The width of the image, such as an image with a resolution of 640x480. That's 640 pixels; It indicates how many millimeters a single pixel (1 pixel) corresponds to in the real world.

[0048] To explain the physical meaning of dx in more detail, this embodiment provides an example: Total width of sensor The image pixel width is ,but In other words, under the current camera settings, one pixel in an image actually corresponds to a length of approximately 0.00964 millimeters in the real world.

[0049] Therefore, the physical area of ​​a 1×1 pixel in the real world is dx × dx (unit: (i.e., the area of ​​a single pixel). Then calculate the 2D true area. The unit is That is, the actual projected area of ​​the food. The total number of pixels representing food The product of this and the physical area of ​​a single pixel.

[0050] Then, balancing accuracy and complexity, different methods are manually selected based on the food's shape to calculate its 3D volume: Regularly shaped foods such as apples and rice balls, due to their relatively regular and compact overall shape, approximate a cylinder. Therefore, the equivalent cylinder method is chosen to calculate the 3D volume of the food. We assume the food is a "cylinder," with the base representing the actual 2D area. The height H is the thickness of the food (calculated using binocular parallax: the depth difference between the front and back edges of the food area, i.e., ...). , The depth at which the food is near the camera. (for the depth away from the camera). The volume calculation formula at this time is: (unit: ).

[0051] For irregularly shaped foods such as noodles and vegetables, due to their highly irregular shapes and significantly uneven surfaces, point cloud reconstruction is used to calculate their 3D volume. Using OpenCV's `cv2.reprojectImageTo3D` function, combined with calibration parameters, the optimized disparity map `disp` is converted into corresponding 3D point cloud data, where each point contains its (X, Y, Z) coordinates (in mm) in the real-world coordinate system. Using the precise food mask `mask_food` obtained in step S2 as an index, a point cloud set `point_cloud_food` belonging only to this irregularly shaped food is selected from the entire scene's 3D point cloud. The `create_from_point_cloud_poisson` method from the Open3D library is called to perform Poisson surface reconstruction on the extracted food point cloud `point_cloud_food`. This algorithm can robustly generate a closed, continuous triangular mesh surface model from discrete point clouds. The volume of this 3D mesh model is directly calculated using the `open3d.geometry.TriangleMesh.compute_volume` method from the Open3D library, in mm. .

[0052] Calculate food mass using the mass formula: ,in Density of food, unit: . The density value is obtained by querying a pre-established density database based on the food category. This food density database was built by querying the "Food Engineering Data Handbook" and records the density values ​​of common foods, such as rice (approximately 0.8). Beef approximately 1.03g Apples, approximately 0.9g .

[0053] S5. Heat Calculation: Based on the food classification results (food categories) and their respective calculated masses, a food calorie database containing cooking method information is queried. This database is a multi-dimensional structured database, with core fields including at least food category, cooking method, density, and calorie content per 100g edible portion, as shown in Table 1: Table 1. Food Calorie Database Fields

[0054] This database is stored using MySQL or Excel.

[0055] For example, "white rice (steamed / boiled)" contains 116 kcal per 100g; "broccoli (stir-fried)" contains 45 kcal per 100g. Finally, calculate and obtain the total calorie value of the food in the dish: ,in, Indicates the quality of cooked white rice (steamed or boiled); This indicates the quality of the stir-fried broccoli.

[0056] In the method of this invention, the database allows for manual correction of automatically calculated food categories or qualities, and updates the exclusive density or calorie preference values ​​based on the correction feedback data to achieve personalized adaptation; at the same time, through the aggregation and analysis of a large amount of data, the food density, classification model and calorie mapping relationship in the database are continuously optimized to improve the overall accuracy and adaptability of the database.

[0057] In summary, this embodiment fully and clearly demonstrates the entire process of the present invention, from image acquisition and preprocessing, food region segmentation and recognition, food volume and mass calculation to final calorie calculation and output. By introducing key technologies such as edge-parallax collaborative network optimization for weak texture matching, U-Net combined with post-processing for accurate segmentation, and adaptive volume reconstruction based on food shape, the invention systematically solves the problem of insufficient segmentation and volume calculation accuracy faced by existing technologies when processing weakly textured and irregularly shaped foods, providing a reliable technical tool for scenarios such as family diet management and health monitoring.

[0058] Example 2 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described binocular vision-based food calorie calculation method.

[0059] Example 3 This embodiment provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the above-described binocular vision-based food calorie calculation method.

[0060] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for calculating food calories based on binocular vision, characterized in that, Includes the following steps: S1. Image Acquisition and Preprocessing: The binocular images of the food in the plate are acquired using a calibrated binocular camera. Distortion correction, epipolar alignment, and image quality optimization are performed on the binocular images. The initial disparity map is calculated using a semi-global block matching algorithm, and the initial disparity map is optimized using weighted least squares filtering to obtain the optimized disparity map. S2, Food Region Segmentation: Input the image processed in S1 into a lightweight semantic segmentation network to obtain classification masks for food, plates, and background; The classification mask is subjected to edge optimization and post-processing to output an accurate food region outline; S3. Food Classification and Recognition: Based on the outline of the food region, crop out the food sub-image from the original image and input the food sub-image into the lightweight classification network to obtain the specific category of the food; S4. Volume and Mass Calculation: Based on the optimized disparity map and the food region outline, calculate the true depth and physical size of the food, and then select the equivalent cylinder method or point cloud reconstruction method to obtain the volume of the food through three-dimensional reconstruction according to the food shape; query the pre-established density database according to the specific category of the food to obtain the corresponding density, and convert the volume of the food into the mass of the food. S5. Calorie Calculation: Based on the specific category and quality of the food, query the food calorie database that records cooking method information to obtain the final food calorie value; The disparity map calculation in step S1 also includes: using an edge-disparity collaborative network to calculate the disparity map, simultaneously extracting edge features and disparity features from the left and right images, using edge features to guide the stereo matching process, and using the matching results to optimize edge extraction in reverse. The edge-parallax collaborative network includes: A weighted twin backbone network is used to extract 1 / 4 resolution feature maps from the left and right images; The edge detection branch, consisting of at least two convolutional layers and a sigmoid activation function, is used to generate a 1 / 4 resolution edge feature map from the left image feature map. The fusion module is used to fuse the feature maps of the left and right images with the edge feature maps, and construct the initial cost volume through a one-dimensional correlation layer; The context pyramid module uses multi-scale convolutional kernels to extract prior information about the scene. The parallax encoder, consisting of 16 residual blocks and dilated convolutions with a dilation rate of 2, is used to extract deep parallax features. The residual pyramid decoder is used to perform multi-level upsampling and residual correction on deep disparity features, and output a full-resolution disparity map and edge map. The twin backbone network is a ResNet-50 model. The context pyramid module uses 1×1, 3×3, 5×5 and 7×7 convolutional kernels for multi-scale feature extraction. The residual pyramid decoder performs 4 levels of upsampling.

2. The method according to claim 1, characterized in that, The lightweight semantic segmentation network in step S2 is a U-Net structure.

3. The method according to claim 2, characterized in that, In step S2, edge optimization and post-processing specifically involve: using Canny edge detection to refine the edges of the food region mask, and using morphological closing operations to fill the mask region.

4. The method according to claim 1, characterized in that, The lightweight classification network in step S3 is the EfficientNet-B0 model.

5. The method according to claim 1, characterized in that, In step S4, the calculation of the true depth of the food is specifically achieved through the following formula: in, For true depth, For camera focal length, The binocular baseline distance. For disparity values, The minimum disparity value offset.

6. The method according to claim 1, characterized in that, In step S4, the volume of the food is obtained through three-dimensional reconstruction using the equivalent cylinder method. Specifically, the projected area of ​​the food is combined with the height calculated from the depth information, and the actual volume of the food is approximated by the cylinder volume.

7. The method according to claim 1, characterized in that, In step S4, the volume of the food is obtained through three-dimensional reconstruction using the point cloud reconstruction method. Specifically, the optimized disparity map is reprojected into a three-dimensional point cloud; food point clouds are extracted from the three-dimensional point cloud based on the food region contour; Poisson surface reconstruction is performed on the food point cloud to generate a three-dimensional mesh model; and the volume of the three-dimensional mesh model is calculated as the volume of the food.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Binocular stereoscopic vision type food identification method

    CN108535252A

  • Depth estimation method and system for asynchronous binocular camera

    CN113822925A