Food material volume detection method and cooking apparatus

By employing a multi-angle camera layout and hybrid depth estimation in the steam oven, the problems of occlusion and error accumulation in monocular vision within the steam oven cavity are solved, achieving high-precision food volume detection and improving the stability of cooking control and user experience.

CN120707617BActive Publication Date: 2025-11-18HANGZHOU ROBAM APPLIANCES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511179324.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-18
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing steam ovens rely on monocular vision to detect the volume of food, which is subject to occlusion, reflection and steam interference, resulting in inaccurate depth estimation, affecting the stability of cooking control and energy consumption, and causing poor user experience. Furthermore, high-precision solutions are not suitable for enclosed and narrow cavities.

Method used

By employing a multi-angle camera layout and hybrid depth estimation, and through tilted camera installation, hydrophobic anti-fog coating, miniature heating element, temperature sensor and ring infrared fill light, combined with calibration board and lightweight semantic segmentation model, multi-view image fusion and 3D reconstruction are achieved, eliminating occlusion and error accumulation.

Benefits of technology

It improves the accuracy and robustness of food volume detection, is suitable for low-power edge devices, provides high-precision sensing support and energy consumption optimization, and meets the needs of intelligent cooking systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707617B_ABST
    Figure CN120707617B_ABST
Patent Text Reader

Abstract

The application provides a food material volume detection method and a cooking device. The method is applied to the cooking device and includes: acquiring image information of the cooking device collected by a multi-angle camera; acquiring intrinsic parameters and extrinsic parameters of the camera based on a preset calibration board; sequentially performing fusion semantic segmentation, feature alignment and hybrid depth estimation on multi-view image information to obtain a final depth map corresponding to the multi-view image information; and sequentially performing depth map conversion, point cloud registration, surface reconstruction and voxel statistics on the final depth map based on the intrinsic parameters and the extrinsic parameters to obtain a food material volume corresponding to the image information. In this way, the problems of occlusion and error accumulation of monocular vision can be solved through inclined camera layout and hybrid depth estimation, and high precision, high robustness and high integration are achieved, which is suitable for low-power edge devices and provides key sensing support and energy consumption optimization basis for an intelligent cooking system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for detecting the volume of food ingredients and a cooking device. Background Technology

[0002] Currently, most steam ovens on the market rely on single images to identify food areas and convert area to volume using preset functions, or use depth estimation from a single image to calculate food volume. However, these methods ignore blind spots in single-image imaging, leading to errors in volume estimation. Furthermore, monocular vision within the steam oven cavity is susceptible to obstruction, reflection, and steam interference, further reducing the accuracy of depth estimation. This inaccuracy in volume perception directly affects the stability of cooking control, making precise control of hot air or steam impossible, thus impacting the cooking results. Additionally, due to the lack of real-time volume information, the equipment often requires manual adjustment of heating time, resulting in energy waste. Users also struggle to obtain a personalized and precise intelligent cooking experience, mostly relying on experience and rules. While some current improvements, such as laser ranging or structured light depth cameras, can improve perception accuracy, their high cost and large size make them unsuitable for the confined space of a closed steam oven cavity. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method for detecting the volume of food ingredients and a cooking device, so as to solve the problems of occlusion and error accumulation in monocular vision by using a tilted camera layout and mixed depth estimation.

[0004] In a first aspect, embodiments of the present invention provide a method for detecting the volume of food ingredients, applied to cooking equipment. The method includes: acquiring image information of the cooking equipment captured by a multi-angle camera; acquiring intrinsic and extrinsic parameters of the camera based on a preset calibration plate; sequentially performing fusion semantic segmentation, feature alignment, and hybrid depth estimation on the multi-view image information to obtain a final depth map corresponding to the multi-view image information; and sequentially performing depth map conversion, point cloud registration, surface reconstruction, and voxel statistics on the final depth map based on the intrinsic and extrinsic parameters to obtain the volume of food ingredients corresponding to the image information.

[0005] In an optional embodiment of this application, the camera is a wide-angle camera. Multiple wide-angle cameras are symmetrically installed on both sides of the top of the cavity of the cooking device, and the combined field of view of the multiple wide-angle cameras covers the bottom area of ​​the cavity of the cooking device. The lens surface of the camera is provided with a hydrophobic anti-fog coating. The camera module integrates a micro heating element, which, in conjunction with a temperature sensor, adjusts the output power in real time to keep the lens temperature of the camera within a preset range higher than the dew point temperature. The temperature sensor is fixed inside the camera module cavity to sense the heat and humidity state of the environment around the lens of the camera.

[0006] In optional embodiments of this application, the above method further includes: using a time-division exposure mode to acquire RGB images and infrared images respectively, a combination of a ring infrared fill light and a visible light LED light source.

[0007] In an optional embodiment of this application, the calibration plate is embedded in the bottom of the cavity of the cooking device; the method further includes: popping out the calibration plate when the cooking device is first started or during periodic self-test.

[0008] In an optional embodiment of this application, the steps of obtaining the intrinsic and extrinsic parameters of the camera based on a preset calibration board include: acquiring multiple sets of dual-view images of the calibration board at different spatial positions, determining the intrinsic parameters of the camera based on the multiple sets of dual-view images; determining the radial distortion coefficient of the camera based on the intrinsic parameters; acquiring images containing the calibration board from multiple viewpoints, performing corner detection, feature point matching, and preliminary alignment and verification sequentially based on the images containing the calibration board to obtain image point pair data; determining the extrinsic parameters of the camera based on the intrinsic parameters and the image point pair data, and optimizing the reprojection error.

[0009] In optional embodiments of this application, the above method further includes: acquiring a static background image of the cooking device in a cavity state; performing inter-frame difference analysis on the current frame image and the static background image based on the image information to obtain a difference image, and extracting edge line segments from the difference image; constructing a geometric mask for the tray region based on the straight line fitting result of the edge line segments; and removing the image information of the tray region from the current frame image based on the geometric mask of the tray region.

[0010] In optional embodiments of this application, the steps of sequentially performing semantic segmentation, feature alignment, and hybrid depth estimation on multi-view image information to obtain the final depth map corresponding to the multi-view image information include: preprocessing the multi-view image information based on a lightweight semantic segmentation model to extract the semantic mask corresponding to the food region as the food mask; determining feature points within the food region based on the food mask and aligning the feature points of the multi-view image information; determining a multi-view depth map based on the aligned feature points of the multi-view image information; inputting the multi-view image information into a preset monocular depth estimation model and outputting a monocular depth map; and fusing the multi-view depth map and the monocular depth map by a weighted average to obtain the final depth map.

[0011] In an optional embodiment of this application, the steps of performing depth map conversion, point cloud registration, surface reconstruction, and voxel statistics on the final depth map based on intrinsic and extrinsic parameters to obtain the food volume corresponding to the image information include: performing depth map conversion on the final depth map to obtain three-dimensional point cloud data; performing point cloud registration on the three-dimensional point cloud data based on intrinsic and extrinsic parameters to obtain a three-dimensional food point cloud set; performing convex hull modeling and surface reconstruction on the three-dimensional food point cloud set to obtain a closed surface; dividing the space surrounding the closed surface into a three-dimensional voxel mesh; and performing statistics on all voxel units falling inside the convex hull to obtain the food volume corresponding to the image information.

[0012] In optional embodiments of this application, the above method further includes: identifying the types of food ingredients in the image information based on a lightweight classification network; determining a volume compensation coefficient based on the types of food ingredients; and adjusting the volume of the food ingredients based on the volume compensation coefficient.

[0013] Secondly, embodiments of the present invention also provide a cooking device, comprising: an image information acquisition module for acquiring image information of the cooking device captured by a multi-angle camera; a camera parameter calculation module for acquiring intrinsic and extrinsic parameters of the camera based on a preset calibration plate; an image information processing module for sequentially performing fusion semantic segmentation, feature alignment, and hybrid depth estimation on the multi-view image information to obtain a final depth map corresponding to the multi-view image information; and a food volume calculation module for sequentially performing depth map conversion, point cloud registration, surface reconstruction, and voxel statistics on the final depth map based on the intrinsic and extrinsic parameters to obtain the food volume corresponding to the image information.

[0014] The embodiments of the present invention bring the following beneficial effects:

[0015] This invention provides a method for detecting food volume and a cooking device. The method involves acquiring image information from a multi-angle camera; obtaining intrinsic and extrinsic parameters of the camera based on a preset calibration board; sequentially performing semantic segmentation, feature alignment, and hybrid depth estimation on the multi-view image information to obtain the final depth map corresponding to the multi-view image information; and then sequentially performing depth map transformation, point cloud registration, surface reconstruction, and voxel statistics on the final depth map based on the intrinsic and extrinsic parameters to obtain the food volume corresponding to the image information. This method can solve the occlusion and error accumulation problems of monocular vision through tilted camera layout and hybrid depth estimation, combining high precision, high robustness, and high integration. It is suitable for low-power edge devices and provides crucial perception support and energy consumption optimization basis for intelligent cooking systems.

[0016] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.

[0017] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 A flowchart of a method for detecting the volume of food provided in an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of the structure of a multi-view imaging system in an integrated device provided by an embodiment of the present invention;

[0021] Figure 3 A schematic diagram of a calibration plate provided in an embodiment of the present invention;

[0022] Figure 4 A flowchart of another method for detecting the volume of food provided in an embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram of the structure of a cooking device provided in an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Currently, most steam ovens on the market rely on single images to identify food areas and convert area to volume using preset functions, or use depth estimation from a single image to calculate food volume. However, these methods ignore blind spots in single-image imaging, leading to errors in volume estimation. Furthermore, monocular vision within the steam oven cavity is susceptible to obstruction, reflection, and steam interference, further reducing the accuracy of depth estimation. This inaccuracy in volume perception directly affects the stability of cooking control, making precise control of hot air or steam impossible, thus impacting the cooking results. Additionally, due to the lack of real-time volume information, the equipment often requires manual adjustment of heating time, resulting in energy waste. Users also struggle to obtain a personalized and precise intelligent cooking experience, mostly relying on experience and rules. While some current improvements, such as laser ranging or structured light depth cameras, can improve perception accuracy, their high cost and large size make them unsuitable for the confined space of a closed steam oven cavity.

[0026] Based on this, the present invention provides a method and cooking device for detecting the volume of food ingredients. Specifically, it provides a method for detecting the volume of food ingredients in a steaming and baking device based on the fusion of multi-view vision and depth estimation. By arranging dual-view cameras and combining camera calibration and image registration technology, it achieves accurate reconstruction of the three-dimensional structure and volume calculation of food ingredients in the complex environment of the steaming and baking cavity, thereby significantly improving the accuracy and robustness of volume detection.

[0027] To facilitate understanding of this embodiment, a method for detecting the volume of food ingredients disclosed in this embodiment of the invention will first be described in detail.

[0028] Example 1:

[0029] This invention provides a method for detecting the volume of food ingredients, applied to cooking equipment. See [link to relevant documentation]. Figure 1 The flowchart shown illustrates a method for detecting the volume of food ingredients, which includes the following steps:

[0030] Step S102: Acquire image information of the cooking equipment from the multi-angle camera.

[0031] Taking an integrated steam and bake cooking appliance as an example, see [link / reference]. Figure 2 The diagram shows a multi-view imaging system in an integrated device. To improve the 3D reconstruction accuracy and image quality of the integrated steam oven, this embodiment has systematically optimized the internal configuration of the integrated machine. Through collaborative optimization, this invention achieves multi-view, highly robust, and easily deployable visual perception capabilities while retaining the integrated and closed structure of the device, providing strong hardware support for high-precision 3D modeling.

[0032] Regarding camera installation, in some embodiments, the camera is a wide-angle camera. Multiple wide-angle cameras are symmetrically installed on both sides of the top of the cooking device cavity, and the combined field of view of the multiple wide-angle cameras covers the bottom area of ​​the cooking device cavity. The lens surface of the camera is provided with a hydrophobic anti-fog coating. The camera module integrates a micro heating element, which, in conjunction with a temperature sensor, adjusts the output power in real time to keep the lens temperature of the camera within a preset range higher than the dew point temperature. The temperature sensor is fixed inside the camera module cavity to sense the thermal and humidity state of the environment around the camera lens.

[0033] In this embodiment, at least two wide-angle cameras (field of view ≥ 120°) are symmetrically installed on both sides of the top of the steam oven cavity. The installation position, attitude angle (pitch angle, yaw angle) and tilt direction of the cameras have been optimized so that their combined field of view can seamlessly cover the bottom area of ​​the cavity, realizing full-view observation of the ingredients during the cooking process and effectively eliminating the image blind spots or occlusion problems existing in the traditional single-view layout.

[0034] By simultaneously acquiring image information from multiple cameras, the system can obtain the texture and geometric features of the food surface from different perspectives, improving subsequent multi-view processing. Figure 3 The system offers high integrity and accuracy in reconstruction, making it particularly suitable for food scenes with complex structures, multiple occlusions, or dynamic changes. Furthermore, this multi-view configuration enhances the system's robustness across different food placement positions, ensuring imaging stability and consistent reconstruction results.

[0035] This embodiment also incorporates anti-fog and anti-interference design: To address the adverse effects of steam on the imaging quality of the camera lens, this invention introduces a hydrophobic anti-fog coating on the surface of the camera lens and integrates a micro heating element inside the module. This heating element, in conjunction with a temperature sensor, can adjust its output power in real time, ensuring that the lens temperature is always maintained approximately 3-5°C above the dew point temperature within the cavity.

[0036] Dew point temperature refers to the critical temperature at which water vapor begins to condense into liquid water under the current humidity and air pressure conditions of the cavity. By controlling the lens temperature above this critical value, water vapor condensation on the lens surface can be effectively prevented, thus preventing lens fogging and ensuring clear and stable image acquisition.

[0037] This system uses an integrated digital temperature and humidity sensor, which features high accuracy (temperature error ±0.3℃, humidity error ±2%RH) and fast response. The sensor is fixed inside the module cavity, close to the lens area, and can sense the thermal and humidity status of the environment around the lens in real time, providing accurate data for dew point calculation and heating control.

[0038] This embodiment can use an embedded algorithm to calculate the collected temperature and relative humidity in real time, and estimate the dew point temperature based on the Magnus formula: ;in, Here, a = 17.62℃, b = 243.12℃, T is the temperature in Celsius, RH is the relative humidity, and γ is a function based on T and RH. The above embedded algorithm can run on the microprocessor of the all-in-one machine with extremely low resource consumption, supports high-frequency updates (e.g., once per second), ensuring the heating system can quickly respond to environmental changes and prevent condensation on the lens surface.

[0039] This embodiment may also be configured with a supplementary lighting system. In some embodiments, a combination of a ring-shaped infrared supplementary light and a visible light LED light source is used to acquire RGB images and infrared images respectively using a time-division exposure mode.

[0040] The system in this embodiment can be equipped with a combination of a ring-shaped infrared fill light (wavelength 850nm) and a visible light LED (light-emitting diode) light source, and adopts a time-division exposure mode to acquire RGB (red-green-blue) images and infrared images separately. The introduction of infrared fill light is mainly based on the following considerations:

[0041] 1. Reduced ambient light interference: Infrared light is in the non-visible spectrum range and is not affected by changes in visible light in the environment (such as sunlight and indoor lighting), so it can maintain stable imaging in low-illuminance or high-contrast environments.

[0042] 2. Suppressing reflective interference: Infrared light has a relatively low reflectivity to highly reflective surfaces such as metals and liquids, which helps to reduce specular highlights and glare, and improve image clarity.

[0043] 3. Enhanced surface detail acquisition: Infrared bands have higher penetration capabilities for surface textures or structural details of some materials (such as fruits, vegetables, and cooked foods), which can compensate for the lack of information in low-contrast areas of RGB images and obtain more complete information about the surface of food.

[0044] 4. Expanded applicable environment: Infrared supplementary lighting supports normal operation in complex kitchen environments such as night, obstruction, oil fumes, and steam, significantly enhancing the system's robustness and adaptability to different scenarios.

[0045] Step S104: Obtain the intrinsic and extrinsic parameters of the camera based on the preset calibration board.

[0046] The cooking device in this embodiment may be pre-installed with a calibration plate. In some embodiments, the calibration plate is embedded in the bottom of the cavity of the cooking device; the calibration plate can be ejected when the cooking device is first started or during periodic self-tests.

[0047] See Figure 3The diagram shows a calibration plate. To achieve accurate multi-camera calibration, a retractable checkerboard calibration plate (10mm spacing between black and white squares) can be pre-embedded in the bottom of the cavity. During the initial startup of the device or periodic self-test, the calibration plate will automatically pop out and provide calibration functionality, ensuring accurate intrinsic and extrinsic parameters of the dual cameras and providing reliable data support for subsequent depth estimation and 3D reconstruction.

[0048] In some embodiments, multiple sets of dual-view images of the calibration board at different spatial positions can be acquired, and the intrinsic parameters of the camera can be determined based on the multiple sets of dual-view images; the radial distortion coefficient of the camera can be determined based on the intrinsic parameters; images containing the calibration board can be acquired from multiple perspectives, and corner detection, feature point matching, and preliminary alignment and verification can be performed sequentially based on the images containing the calibration board to obtain image point pair data; the extrinsic parameters of the camera can be determined based on the intrinsic parameters and the image point pair data, and reprojection error optimization can be performed.

[0049] Camera calibration aims to obtain the camera's intrinsic and extrinsic parameters to support subsequent image correction, 3D reconstruction, and spatial transformations. Intrinsic parameters describe the internal characteristics of the camera's imaging system and are used to project from the camera coordinate system onto the image plane. Common parameters include: focal length f. x ,f y (Unit: pixel), principal point position c x ,c y (Image center) Distortion coefficients (radial distortion k1, k2, tangential distortion p1, p2, etc.).

[0050] Regarding intrinsic parameter calibration, to accurately obtain the imaging parameters of the cameras, 20 sets of dual-view images of the calibration board at different spatial positions were first acquired, and the Zhang Zhengyou calibration method was used to estimate the intrinsic parameters of each camera. This method can calculate the intrinsic parameter matrix K of each camera, in the following form: ; where f x and f y c represents the focal length of the camera in the horizontal and vertical directions. x and c y These are the coordinates of the principal points in the image. This calibration method allows for the acquisition of accurate imaging parameters, laying the foundation for subsequent depth estimation and 3D reconstruction.

[0051] Regarding distortion parameter estimation, in practical applications, camera lenses, especially wide-angle lenses, often introduce significant radial distortion, which manifests as stretching or compression at the image edges. This distortion mainly originates from the difference in light refraction paths between areas near and far from the lens's optical axis, and is a typical example of nonlinear distortion.

[0052] To further improve the geometric accuracy of the images, this system, based on intrinsic parameter calibration, simultaneously estimates the radial distortion coefficients k1 and k2 of the camera to correct image distortion errors caused by lens nonlinearity during imaging. The geometric model of radial distortion is typically expressed as: ;in, These are ideal image coordinates without distortion; is the image coordinates including distortion; r is the radial distance to the image center. ; and This is the radial distortion coefficient that needs to be estimated.

[0053] During the calibration process, image corner points can be extracted by collecting image samples from the calibration plate, and the distortion model parameters can be fitted using the least squares method to accurately correct imaging distortion.

[0054] Regarding extrinsic parameter calibration, this embodiment requires extrinsic parameter calibration to establish the spatial coordinate transformation relationship between multiple cameras, i.e., calculating the relative pose parameters between each camera, including the rotation matrix. Translation vector T. This process unifies the images from each camera to the same three-dimensional coordinate system, achieving spatial information fusion and reconstruction. The specific steps are as follows:

[0055] 1. Image Registration: First, images containing the retractable checkerboard calibration board are acquired from multiple viewpoints, and the following registration process is executed sequentially: Corner Detection: The corner points (i.e., calibration points) of the checkerboard are detected in each image as stable and reliable image features; Feature Point Matching: Corresponding corner points in images from different cameras are precisely matched to establish pixel coordinate pairs of the same spatial point in multiple images; Preliminary Alignment and Verification: The matched points are back-projected into the normalized image coordinate system using intrinsic parameters, and inconsistent matches are eliminated or optimized to ensure registration accuracy. This step provides high-quality image point pair data for subsequent essential matrix calculation.

[0056] 2. Calculation of the essential matrix: Based on the known intrinsic parameters of the camera, calculate the essential matrix E according to the point-to-point relationship. The mathematical relationship between the essential matrix and the extrinsic parameters is as follows: ;in, It is the antisymmetric matrix of the translation vector T, used to represent the cross product relationship between vectors.

[0057] 3. Essential Matrix Decomposition and Solution Space Selection: For the essential matrix... Singular value decomposition (SVD) can yield multiple candidate solutions. and Combinations (up to four possibilities). To determine a unique and physically reasonable solution, further steps are required: Triangulation: Reconstructing 3D points by inverse triangulation of the matching points; Depth consistency judgment: Selecting solutions that place the 3D points in front of the two cameras (i.e., with positive depth); Geometric consistency verification: Further verifying the structural continuity and rationality of the reconstructed point cloud to determine the final valid solution.

[0058] 4. Reprojection Error Optimization: To improve the accuracy of extrinsic parameter estimation, linear least squares or nonlinear optimization (such as Levenberg-Marquardt) is used to minimize the reprojection error for all matched image points. That is: ;in, For real image points, Represents the projection function. Let i be the number of reconstructed 3D points.

[0059] The final output is the precise extrinsic parameters between the two cameras: the rotation matrix R and the translation vector T, which provide basic support for geometric registration and spatial point reconstruction of multi-view images.

[0060] This embodiment can also provide an anti-interference image preprocessing method. In some embodiments, a static background image of the cooking device in a cavity state can be acquired; a difference image is obtained by performing inter-frame difference between the current frame image and the static background image based on the image information, and edge line segments are extracted from the difference image; a geometric mask of the tray area is constructed based on the straight line fitting result of the edge line segments; and the image information of the tray area in the current frame image is removed based on the geometric mask of the tray area.

[0061] The anti-interference image preprocessing method in this embodiment mainly includes the following steps:

[0062] 1. Dynamic Background Differentiation: First, a static background image is acquired in the cavity without any food. This serves as a reference baseline. During actual detection, the current frame image is used as a reference baseline. Inter-frame difference analysis with the background image yields the dynamically changing region: .

[0063] 2. Tray Area Localization and Exclusion: To avoid interference from tray edges on food contour extraction, a tray boundary detection algorithm based on Hough transform is further introduced: the Canny algorithm is applied to the difference image for edge extraction; Hough linear transform is used to detect edge segments in the scene; and a geometric mask for the tray area is constructed based on the linear fitting results. This is used for subsequent masking and exclusion processing.

[0064] The pseudocode implementation example is as follows:

[0065] edges=cv2.Canny(image,50,150);

[0066] lines=cv2.HoughLinesP(edges,rho=1,theta=np.pi / 180,threshold=50,minLineLength=100,maxLineGap=10);

[0067] mask_tray=draw_lines_mask(lines) # Generate a tray area mask based on line segments.

[0068] This method can accurately extract and remove image information from the tray area, retaining the pure food area, thereby improving the accuracy of subsequent volume estimation and 3D modeling.

[0069] The method provided in this embodiment of the invention improves the segmentation robustness under steam / reflective conditions by combining cross-modal anti-interference processing with infrared images and Hough transform geometric analysis.

[0070] Step S106: Sequentially perform fusion semantic segmentation, feature alignment and hybrid depth estimation on the multi-view image information to obtain the final depth map corresponding to the multi-view image information.

[0071] Regarding multi-view image registration and depth estimation, to achieve high-precision reconstruction of the three-dimensional structure of food ingredients, this embodiment proposes a multi-view image processing method that integrates semantic segmentation, feature alignment, and hybrid depth estimation. This method comprehensively utilizes binocular parallax information and monocular depth estimation results, and still exhibits good robustness and reconstruction accuracy even in complex imaging environments such as steam interference and surface reflection.

[0072] Step S108: Based on the intrinsic and extrinsic parameters, perform depth map conversion, point cloud registration, surface reconstruction, and voxel statistics on the final depth map in sequence to obtain the food volume corresponding to the image information.

[0073] Regarding 3D reconstruction and volume calculation, in order to achieve accurate reconstruction of the true 3D structure of food and quantitative volume assessment, this embodiment uses a combination of depth map conversion, point cloud registration, surface reconstruction and voxel statistics to effectively improve the reconstruction accuracy and the stability of volume estimation.

[0074] This invention provides a method for detecting food volume. The method involves acquiring image information from a cooking device using multi-angle cameras; obtaining the intrinsic and extrinsic parameters of the cameras based on a preset calibration board; sequentially performing semantic segmentation, feature alignment, and hybrid depth estimation on the multi-view image information to obtain the final depth map corresponding to the multi-view image information; and then sequentially performing depth map transformation, point cloud registration, surface reconstruction, and voxel statistics on the final depth map based on the intrinsic and extrinsic parameters to obtain the food volume corresponding to the image information. This method can solve the occlusion and error accumulation problems of monocular vision through tilted camera layout and hybrid depth estimation, combining high precision, high robustness, and high integration. It is suitable for low-power edge devices and provides crucial perception support and energy consumption optimization basis for intelligent cooking systems.

[0075] Example 2:

[0076] This embodiment provides another method for detecting food volume, which is implemented based on the above embodiment. It focuses on describing the specific implementation methods of multi-view image registration and depth estimation, and 3D reconstruction and volume calculation. See also... Figure 4 The flowchart shown illustrates another method for detecting the volume of food ingredients, which includes the following steps:

[0077] Step S402: Acquire image information of the cooking equipment from the multi-angle camera.

[0078] Step S404: Obtain the intrinsic and extrinsic parameters of the camera based on the preset calibration board.

[0079] Step S406: Sequentially perform fusion semantic segmentation, feature alignment and hybrid depth estimation on the multi-view image information to obtain the final depth map corresponding to the multi-view image information.

[0080] In some embodiments, preprocessing of multi-view image information can be performed based on a lightweight semantic segmentation model to extract the semantic mask corresponding to the food region as the food mask; feature points within the food region are determined based on the food mask, and the feature points of the multi-view image information are aligned; a multi-view depth map is determined based on the aligned feature points of the multi-view image information; the multi-view image information is input into a preset monocular depth estimation model to output a monocular depth map; the multi-view depth map and the monocular depth map are fused by a weighted average to obtain the final depth map.

[0081] Regarding the hybrid depth estimation strategy, this embodiment proposes a hybrid strategy that fuses multi-view and monocular depth estimation to further enhance depth perception capabilities in complex environments. This strategy mainly includes the following three parts:

[0082] 1. Multi-view disparity calculation: Calculate the disparity map based on aligned feature point pairs. An initial depth map is generated using triangulation. This depth estimation has strong geometric constraints and is suitable for areas with clear images.

[0083] 2. Monocular Depth Enhancement: An improved Depth Anything monocular depth estimation model (fine-tuned on a typical steaming / baking scene image set) is introduced to predict the monocular depth map using the current image as input. This model can provide supplementary depth information in certain image regions, such as low-texture areas or high-reflectivity areas.

[0084] 3. Depth map fusion:

[0085] The multi-view and monocular depth maps are combined using a weighted average method to form the final depth result. The specific formula is as follows: .

[0086] Among them, the fusion weight parameters This parameter controls the weighted fusion ratio of multi-view and monocular depth estimation results. It is dynamically adjusted based on the sharpness index of the current frame image: the sharper the image, the higher the reliability of the binocular disparity calculation, and correspondingly, the weighting ratio is reduced. The value is increased to enhance the weight of multi-view depth results; conversely, under adverse imaging conditions such as high vapor concentration or blurred images, the value is appropriately increased. This reduces the impact of multi-view depth estimation and increases the reliance on monocular depth estimation results, thereby improving the overall system's adaptability and robustness in complex environments.

[0087] This multi-source depth fusion strategy fully leverages the respective advantages of accurate multi-view depth estimation and wide applicability of monocular estimation, significantly improving the stability of 3D reconstruction and the accuracy of volume estimation under steaming and baking environments.

[0088] Step S408: Based on the intrinsic and extrinsic parameters, perform depth map conversion, point cloud registration, surface reconstruction, and voxel statistics on the final depth map in sequence to obtain the food volume corresponding to the image information.

[0089] In some embodiments, the final depth map can be transformed to obtain three-dimensional point cloud data; point cloud registration can be performed on the three-dimensional point cloud data based on intrinsic and extrinsic parameters to obtain a three-dimensional food point cloud set; surface reconstruction can be performed based on convex hull modeling of the three-dimensional food point cloud set to obtain a closed surface; the space surrounding the closed surface can be divided into a three-dimensional voxel mesh; and all voxel units falling inside the convex hull can be counted to obtain the food volume corresponding to the image information.

[0090] For ease of explanation, the following discussion will use a binocular system with the fewest cameras.

[0091] Regarding point cloud generation and registration, the depth maps corresponding to the dual-view images are first back-projected into 3D point cloud data. and This involves combining image pixel coordinates with depth values ​​and calculating the three-dimensional position of each pixel in the camera coordinate system based on camera intrinsic parameters (including focal length and principal point position).

[0092] Subsequently, based on the aforementioned extrinsic parameter calibration results (i.e., the relative pose information between the two cameras (rotation matrix R, translation vector T)), the point cloud was... Transform all points in the coordinate system to a point cloud In the coordinate system, the transformation formula is as follows: This operation enabled spatial alignment of point cloud data from two camera perspectives, laying the foundation for subsequent fine registration.

[0093] To further improve the accuracy of point cloud fusion, the Iterative Closest Point (ICP) algorithm is introduced to perform fine registration of the two sets of point clouds. ICP is a point cloud alignment algorithm, and its core idea is as follows:

[0094] Initial alignment: Align the source point cloud (e.g.) Preliminary transformation based on existing extrinsic parameters; nearest neighbor matching: for each point in the source point cloud, in the target point cloud... Find the nearest neighbor point to form a pair of corresponding points; minimize the error: solve for the new optimal rigid transformation (rotation and translation) by minimizing the sum of squared Euclidean distances between the two pairs of corresponding points; iterative update: repeat the matching and optimization until the convergence condition is met (such as the error change is below the threshold).

[0095] The final rigid transformation is used to precisely align the two sets of point clouds, and then merge them to generate a complete set of 3D food point clouds. The fusion result provides a geometric basis for subsequent 3D reconstruction and volume calculation, and is particularly suitable for high-precision modeling tasks under multi-view imaging conditions.

[0096] In this embodiment, the following volume estimation strategy is proposed for the concave areas that may exist in the complex structure of food ingredients:

[0097] 1. Convex Hull Surface Reconstruction: The AlphaShape algorithm is used to reconstruct the fused point cloud. Perform convex hull modeling to obtain a closed surface. This algorithm is better suited to natural concave and convex shapes than the traditional convex hull algorithm, effectively avoiding volume underestimation caused by surface leakage.

[0098] 2. Voxel integration method: Using the enclosing convex hull... The space is divided into a three-dimensional voxel mesh with a side length of 1 mm. All voxel elements falling inside the convex hull are counted, and the effective number of voxels is assumed to be... Then the volume V of the food can be expressed as: .

[0099] This method has advantages such as high spatial resolution and strong noise tolerance, and is suitable for volume estimation of various irregular food shapes.

[0100] Step S410: Identify the types of ingredients in the image information based on a lightweight classification network; determine the volume compensation coefficient based on the types of ingredients; and adjust the volume of the ingredients based on the volume compensation coefficient.

[0101] In this embodiment, the calculated food volume can be further optimized, namely, type-aware dynamic compensation.

[0102] In actual heating processes, different types of food often undergo physical deformation phenomena after being heated, such as shrinkage or expansion, which can lead to a systematic deviation between the visually estimated volume and the actual volume. To improve the accuracy of 3D reconstruction results, this invention proposes a type-aware dynamic compensation strategy: First, a lightweight ResNeXt classification network (or even lighter models such as ResNet or MobileNext, which can further improve detection speed) is used to identify the types of food; then, based on the type of food, a targeted volume compensation coefficient is introduced to dynamically adjust the estimation results.

[0103] Table 1

[0104]

[0105] The typical thermal behavior of food ingredients is shown in Table 1. Therefore, the compensation coefficient is automatically adjusted according to the classification results: if the food ingredient is identified as shrinking, the compensation coefficient is adjusted down proportionally based on the estimated volume in three dimensions; if the food ingredient is identified as expanding, the compensation coefficient is adjusted up proportionally; for food ingredients with high thermal stability (such as nuts), no compensation is performed by default.

[0106] The method provided in this embodiment of the invention eliminates measurement deviations caused by deformation during cooking by using a dynamic volume compensation mechanism based on a compensation coefficient library categorized by ingredient type.

[0107] To evaluate the accuracy and performance of this invention in practical applications, the system underwent comprehensive verification and optimization from two dimensions: volume measurement error calibration and inference real-time performance.

[0108] To verify the accuracy of the volume estimation, this embodiment also includes a standard block calibration experiment. A series of standard volume blocks (volume range 50–1000 cm³) were selected. 3The sample was placed in the steam oven cavity for testing, and the estimated value was recorded each time. its true value Calculate the mean relative error (MRE) as the evaluation metric:

[0109] Experimental results show that, with the support of the multi-view vision and multi-source depth fusion method proposed in this invention, the average relative error of the system is controlled below 3.5%, which is significantly better than the traditional monocular depth estimation scheme (MRE is usually higher than 8%), verifying the high accuracy and robustness of the proposed method in complex environments.

[0110] To meet the high response speed requirements of intelligent cooking scenarios, the system provided in this embodiment is deployed on an embedded platform and its operational efficiency is optimized from multiple levels. First, TensorRT (a high-performance deep learning inference engine) is used to accelerate the core deep learning model (including semantic segmentation and depth estimation networks) at the tensor level, and mixed-precision computing (FP16) is enabled, significantly improving inference speed. Simultaneously, the image processing and data transmission processes are streamlined and restructured to avoid redundant operations and resource waste. After overall optimization, the system achieves an end-to-end processing time of less than 2.5 seconds from image acquisition, preprocessing, segmentation, feature registration, 3D reconstruction to volume output, effectively meeting the stringent real-time requirements of embedded edge computing environments.

[0111] Furthermore, this embodiment also provides the ability to remotely detect and return results via WIFI, which can further improve the accuracy of food identification.

[0112] In summary, the embodiments of the present invention aim to propose a food volume detection method that is more accurate and robust, suitable for embedded devices, and is specifically designed for constrained hardware environments such as smart steam ovens.

[0113] To address the common single-view blind spot problem in current equipment, as well as the technical bottleneck of unstable depth estimation accuracy under complex imaging conditions such as steam, water mist, and reflection, this invention constructs a multi-camera-based stereo vision acquisition system, combining a lightweight semantic segmentation model with an improved hybrid depth estimation algorithm to achieve accurate reconstruction of the three-dimensional structure of food ingredients.

[0114] The algorithm provided in this invention can run end-to-end on an embedded platform. With the help of inference acceleration frameworks such as TensorRT, the overall processing time is controlled within 2.5 seconds, meeting the real-time requirements of intelligent cooking. The system also integrates infrared imaging and Hough geometric analysis technology to enhance segmentation stability in high-humidity and high-reflection environments; combines camera calibration parameters with point cloud optimization reconstruction to improve 3D reconstruction accuracy; and effectively corrects volume estimation errors caused by thermal expansion and contraction through food type recognition and compensation coefficient correction mechanisms. The overall solution combines high precision, high robustness, and high integration, making it suitable for low-power edge devices and providing crucial sensing support and energy consumption optimization for intelligent cooking systems.

[0115] Example 3:

[0116] Corresponding to the above method embodiments, this invention provides a cooking device, see [link to relevant documentation]. Figure 5 The diagram shows the structure of a cooking device, which includes:

[0117] Image information acquisition module 51 is used to acquire image information of cooking equipment captured by multi-angle cameras;

[0118] The camera parameter calculation module 52 is used to obtain the camera's intrinsic and extrinsic parameters based on a preset calibration board.

[0119] The image information processing module 53 is used to sequentially perform fusion semantic segmentation, feature alignment and hybrid depth estimation on the multi-view image information to obtain the final depth map corresponding to the multi-view image information.

[0120] The food volume calculation module 54 is used to perform depth map conversion, point cloud registration, surface reconstruction and voxel statistics on the final depth map based on intrinsic and extrinsic parameters to obtain the food volume corresponding to the image information.

[0121] This invention provides a cooking device that acquires image information from multiple angle cameras; obtains intrinsic and extrinsic parameters of the cameras based on a preset calibration board; sequentially performs semantic segmentation, feature alignment, and hybrid depth estimation on the multi-view image information to obtain the final depth map corresponding to the multi-view image information; based on the intrinsic and extrinsic parameters, sequentially performs depth map transformation, point cloud registration, surface reconstruction, and voxel statistics on the final depth map to obtain the food volume corresponding to the image information. This method can solve the occlusion and error accumulation problems of monocular vision through tilted camera layout and hybrid depth estimation, combining high precision, high robustness, and high integration. It is suitable for low-power edge devices and provides key perception support and energy consumption optimization basis for intelligent cooking systems.

[0122] The aforementioned cameras are wide-angle cameras. Multiple wide-angle cameras are symmetrically installed on both sides of the top of the cooking device's cavity. The combined field of view of the multiple wide-angle cameras covers the bottom area of ​​the cooking device's cavity. The lens surface of the camera is coated with a hydrophobic anti-fog coating. The camera module integrates a micro heating element, which, in conjunction with a temperature sensor, adjusts the output power in real time to maintain the camera lens temperature within a preset range above the dew point temperature. The temperature sensor is fixed inside the camera module cavity to sense the heat and humidity of the environment around the camera lens.

[0123] The aforementioned cooking equipment also includes: a supplementary lighting system configuration module, which is used to acquire RGB images and infrared images respectively using a time-sharing exposure mode with a combination of ring infrared supplementary light and visible light LED light source.

[0124] The aforementioned calibration plate is embedded in the bottom of the cavity of the cooking equipment; the aforementioned cooking equipment also includes: a calibration plate pre-installation module, used to pop out the calibration plate when the cooking equipment is first started or during periodic self-test.

[0125] The aforementioned camera parameter calculation module is used to acquire multiple sets of dual-view images of the calibration board at different spatial positions, determine the intrinsic parameters of the camera based on the multiple sets of dual-view images, determine the radial distortion coefficient of the camera based on the intrinsic parameters, acquire images containing the calibration board from multiple perspectives, and perform corner detection, feature point matching, and preliminary alignment and verification based on the images containing the calibration board to obtain image point pair data; determine the extrinsic parameters of the camera based on the intrinsic parameters and image point pair data, and optimize the reprojection error.

[0126] The aforementioned cooking equipment also includes: an anti-interference image preprocessing module, used to acquire a static background image of the cooking equipment in a cavity state; perform inter-frame difference analysis on the current frame image and the static background image based on the image information to obtain a difference image, and extract edge line segments from the difference image; construct a geometric mask for the tray area based on the straight line fitting results of the edge line segments; and remove the image information of the tray area in the current frame image based on the geometric mask of the tray area.

[0127] The aforementioned image information processing module is used to preprocess multi-view image information based on a lightweight semantic segmentation model, extract the semantic mask corresponding to the food region as the food mask; determine the feature points within the food region based on the food mask, and align the feature points of the multi-view image information; determine a multi-view depth map based on the aligned feature points of the multi-view image information; input the multi-view image information into a preset monocular depth estimation model, and output a monocular depth map; and fuse the multi-view depth map and the monocular depth map through a weighted average to obtain the final depth map.

[0128] The aforementioned food volume calculation module is used to perform depth map conversion on the final depth map to obtain 3D point cloud data; perform point cloud registration on the 3D point cloud data based on intrinsic and extrinsic parameters to obtain a 3D food point cloud set; perform surface reconstruction by convex hull modeling based on the 3D food point cloud set to obtain a closed surface; divide the space surrounding the closed surface into a 3D voxel mesh; and statistically analyze all voxel units falling inside the convex hull to obtain the food volume corresponding to the image information.

[0129] The aforementioned cooking equipment also includes: a type-aware dynamic compensation module, used to identify the type of ingredients in image information based on a lightweight classification network; determine the volume compensation coefficient based on the type of ingredients; and adjust the volume of ingredients based on the volume compensation coefficient.

[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the cooking equipment described above can be referred to the corresponding process in the aforementioned embodiments of the food volume detection method, and will not be repeated here.

[0131] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0132] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0133] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0134] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A food material volume detection method applied to a cooking device, characterized in that, The method comprises: acquiring image information of a cooking device collected by a multi-angle camera; obtaining intrinsic parameters and extrinsic parameters of the camera based on a preset calibration board; performing fusion semantic segmentation, feature alignment and hybrid depth estimation on the multi-view image information in sequence to obtain a final depth map corresponding to the multi-view image information; wherein the hybrid depth estimation is a hybrid strategy of multi-view and monocular depth estimation fusion, and the hybrid depth estimation includes multi-view disparity calculation, monocular depth enhancement and depth map fusion; based on the intrinsic parameters and the extrinsic parameters, performing depth map conversion, point cloud registration, surface reconstruction and voxel statistics on the final depth map in sequence to obtain a food material volume corresponding to the image information; The method further comprises: collecting a static background image of the cooking device in a cavity state; obtaining a difference image by performing inter-frame difference based on the current frame image of the image information and the static background image, and extracting edge line segments of the difference image; constructing a geometric mask of the tray area based on the straight line fitting result of the edge line segments; and removing the image information of the tray area of the current frame image based on the geometric mask of the tray area.

2. The method of claim 1, wherein, The camera is a wide-angle camera, and a plurality of wide-angle cameras are symmetrically installed on both sides of the top of the cavity of the cooking device, and the combined field of view of the plurality of wide-angle cameras covers the bottom area of the cavity of the cooking device; The lens surface of the camera is provided with a hydrophobic anti-fog coating; A micro-heating element is integrated inside the camera module, and the micro-heating element adjusts the output power in real time in combination with a temperature sensor to maintain the lens temperature of the camera above a preset range of dew point temperature; The temperature sensor is fixed inside the camera module cavity and senses the thermal and humid state of the environment around the camera lens.

3. The method of claim 1, wherein, The method further comprises: The annular infrared fill light and the visible light LED combined light source adopt a time-sharing exposure mode to respectively collect RGB images and infrared images.

4. The method of claim 1, wherein, The calibration board is pre-embedded in the bottom of the cavity of the cooking device; the method further comprises: When the cooking device is started for the first time or periodically self-checked, the calibration board is popped up.

5. The method of claim 1, wherein, The step of obtaining the intrinsic parameters and the extrinsic parameters of the camera based on the preset calibration board comprises: acquiring a plurality of sets of double-view images of the calibration board at different spatial positions, and determining the intrinsic parameters of the camera based on the plurality of sets of double-view images; determining the radial distortion coefficient of the camera based on the intrinsic parameters; acquiring images containing the calibration board under multiple views, performing corner point detection, feature point matching and preliminary alignment and verification in sequence based on the images containing the calibration board to obtain image point pair data; determining the extrinsic parameters of the camera based on the intrinsic parameters and the image point pair data, and performing re-projection error optimization.

6. The method of claim 1, wherein, The step of performing fusion semantic segmentation, feature alignment and hybrid depth estimation on the multi-view image information in sequence to obtain a final depth map corresponding to the multi-view image information comprises: The image information of multiple perspectives is preprocessed based on a lightweight semantic segmentation model to extract a semantic mask corresponding to a food material region as a food material mask; feature points in the food material region are determined based on the food material mask, and the feature points of the image information of multiple perspectives are aligned; A multi-view depth map is determined based on the aligned feature points of the image information of multiple perspectives; the image information of multiple perspectives is input into a preset monocular depth estimation model to output a monocular depth map; and the multi-view depth map and the monocular depth map are fused by weighted averaging to obtain a final depth map.

7. The method of claim 1, wherein, Based on the intrinsic parameters and the extrinsic parameters, the final depth map is sequentially subjected to depth map conversion, point cloud registration, surface reconstruction and voxel statistics to obtain a food material volume corresponding to the image information, including: The final depth map is subjected to depth map conversion to obtain three-dimensional point cloud data; The three-dimensional point cloud data is subjected to point cloud registration based on the intrinsic parameters and the extrinsic parameters to obtain a three-dimensional food material point cloud set; Surface reconstruction is performed based on the three-dimensional food material point cloud set to obtain a closed surface; The space surrounding the closed surface is divided into a three-dimensional voxel grid; and all voxel units falling inside the convex hull are counted to obtain the food material volume corresponding to the image information.

8. The method of claim 1, wherein, The method further includes: A classification network is used to identify the food material category of the image information; A volume compensation coefficient is determined based on the food material category, and the food material volume is adjusted based on the volume compensation coefficient.

9. A cooking apparatus, characterized by, The cooking device includes: An image information acquisition module is configured to acquire image information of a cooking device collected by a multi-angle camera; A camera parameter calculation module is configured to acquire intrinsic parameters and extrinsic parameters of the camera based on a preset calibration board; An image information processing module is configured to sequentially perform fusion semantic segmentation, feature alignment and hybrid depth estimation on the image information of multiple perspectives to obtain a final depth map corresponding to the image information of multiple perspectives; the hybrid depth estimation is a hybrid strategy of multi-view and monocular depth estimation fusion, and includes multi-view disparity calculation, monocular depth enhancement and depth map fusion; A food material volume calculation module is configured to sequentially perform depth map conversion, point cloud registration, surface reconstruction and voxel statistics on the final depth map based on the intrinsic parameters and the extrinsic parameters to obtain a food material volume corresponding to the image information. The cooking device further includes an anti-interference image preprocessing module configured to acquire a static background image of the cooking device in a cavity state; perform inter-frame difference based on a current frame image of the image information and the static background image to obtain a difference image, extract edge line segments of the difference image, construct a geometric mask of a tray region based on a straight line fitting result of the edge line segments, and remove image information of a tray region of the current frame image based on the geometric mask of the tray region.

Citation Information

Patent Citations

  • Food material size calculation method in oven, size identification device and oven

    CN110623555A

  • 3D information sensing method and system based on RGB and infrared images

    CN114782541A

  • Goods volume calculation method and system based on panoramic video

    CN119068042A