Method for volumetric estimation using a stereo camera

The method uses a static stereo camera to transform depth estimation into height maps for precise volume calculation of static objects, overcoming limitations of existing methods by enabling efficient and complex-free volume estimation.

WO2025229505A1PCT designated stage Publication Date: 2025-11-06STEREOLABS SAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/054411
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-04-29
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing methods for object volume estimation using stereo cameras require the object or the camera to move, and often approximate objects to simple shapes, limiting their applicability.

Method used

A method using a static stereo camera to compute depth estimation as a point cloud, transformed into a height map, and calculate volume by comparing height maps with and without the object, allowing detection and volume calculation without moving the object or camera.

Benefits of technology

Enables efficient, accurate volume estimation of objects of any shape with reduced power usage and system complexity, facilitating multiple measurements with a single camera calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025054411_06112025_PF_FP_ABST
    Figure IB2025054411_06112025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods related to volumetric estimation using a stereo camera are disclosed herein. The stereo camera, with two or more imagers, may belong to a stereo imaging system. The camera may be static. The system may determine a first height map corresponding to a reference surface, a second height map corresponding to a physical object on the reference surface, and a difference map including height differences between the first height map and the second height map. The height maps may be derived from point clouds. Based on the difference map, the system may detect the presence of the physical object. The volume of the physical object may be determined, and the object may be identified based on its volume.
Need to check novelty before this filing date? Find Prior Art

Description

Method for Volumetric Estimation Using a Stereo CameraCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Prov. Pat. App. No. 63 / 640,871, filed April 30, 2024.BACKGROUND

[0002] The present description relates to methods of image-based volume estimation using a multi-camera system, and in particular static object volume estimation using stereo camera vision. Applications include logistics, urban planning, and surveillance, where precise space optimization, inventory management, and monitoring are crucial.

[0003] Existing techniques to measure the volume of an object either require the object to be moving (e.g., on a conveyor) and having a projection device (line or points), or require the capture device (e.g., camera) to move around the object. For example, a volume estimation can be based on multiple views of a stereo camera to reconstruct the object. Furthermore, existing methods generally approximate the object to a cube or a similar simple shape. This limits the applicability of the approaches to more general object volume estimation.SUMMARY

[0004] This disclosure relates to measuring the volume of an object using a stereo camera to compute a depth estimation (e.g., a point cloud) which is then transformed into a height map. The height map can be used to compute the volume of the object by determining the difference between the height map with the object present and the height map with the object not visible. To detect the object, a contour detection method can be done based on a height map difference image. The volume measurement may be computationally efficient and applicable to objects of any shape.

[0005] The methods disclosed herein can involve determining a reference height map. The geometric characteristics of a reference surface can be determined in the form of a reference height map. The process can involve a calibration during a first time. The reference surface canbe any surface, plane or not, on which the object will be placed. For example, the reference surface can be a ground plane in a warehouse, a conveyor belt surface, or a loading surface. During this first time, a stereo camera can detect the reference plane without the object.

[0006] The data obtained during the calibration step can then be utilized to determine a three- dimensional (3D) point cloud and the projection thereof in a grid map (e.g., height map). A 3D point cloud can be captured using a stereo camera. The 3D point cloud may then be projected on a grid map. The grid map may be used to determine the volume of the object. The object may be identified (e.g., as one of a class of objects) based on its volume.

[0007] In specific embodiments, the stereo camera is static so that a single calibration can be used for all volume estimations in the reference plane. Alternatively, if the camera is not static, the calibration can be repeated or adjusted as needed when the camera is moved.

[0008] Embodiments disclosed herein allow detection and volume calculation of an object without needing either the object or the measuring device to be moving. Additionally, many object measurements may be made based on a single camera calibration. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance compared to the related art.

[0009] In specific embodiments of the invention, a method for detecting physical objects with a stereo imaging system is provided. The method comprises: determining a first height map at a first time using a pair of imagers of the stereo imaging system and determining a second height map at a second time using the pair of imagers of the stereo imaging system. The first time and the second time are different. The method also comprises determining a difference map, the difference map including a difference between a set of heights of the first height map and a set of heights of the second height map. The method also comprises detecting a presence of a physical object based on the difference map.

[0010] In specific embodiments of the invention, a stereo imaging system for detecting physical objects is provided. The system comprises: a pair of imagers; one or more processors; and one or more non-transitory computer-readable media. The non-transitory computer readable media stores instructions that, when executed by the one or more processors, cause the stereo imaging system to conduct a method. The method comprises: determining a first height map ata first time using the pair of imagers of the stereo imaging system and determining a second height map at a second time using the pair of imagers of the stereo imaging system. The first time and the second time are different. The method also comprises determining a difference map, the difference map including a difference between a set of heights of the first height map and a set of heights of the second height map. The method also comprises detecting a presence of a physical object based on the difference map.

[0011] In specific embodiments of the invention, a stereo imaging system for detecting physical objects is provided. The system comprises: a pair of imagers; a means for determining a first height map at a first time using the pair of imagers of the stereo imaging system; and a means for determining a second height map at a second time using the pair of imagers of the stereo imaging system. The first time and the second time are different. The system also comprises a means for determining a difference map, the difference map including a difference between a set of heights of the first height map and a set of heights of the second height map. The system also comprises a means for detecting a presence of a physical object based on the difference map.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings illustrate various embodiments of systems, methods, and various other aspects of the disclosure. A person with ordinary skills in the art will appreciate that the illustrated element boundaries (e.g., boxes, groups of boxes, or other shapes) in the figures represent one example of the boundaries. It may be that in some examples one element may be designed as multiple elements or that multiple elements may be designed as one element. In some examples, an element shown as an internal component of one element may be implemented as an external component in another, and vice versa. Furthermore, elements may not be drawn to scale. Non-limiting and non-exhaustive descriptions are described with reference to the following drawings. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating principles.

[0013] Figure 1 provides an example of a camera directly above a reference surface at two different times in accordance with specific embodiments of the inventions disclosed herein.

[0014] Figure 2 provides an example of a camera with an axonometric view of a reference surface at two different times in accordance with specific embodiments of the inventions disclosed herein.

[0015] Figure 3 provides an example of a point cloud projected onto a grid map in accordance with specific embodiments of the inventions disclosed herein.

[0016] Figure 4 provides an example of a reference height map, an object height map, and a difference height map in accordance with specific embodiments of the inventions disclosed herein.

[0017] Figure 5 provides a top view of an example of a difference height map in accordance with specific embodiments of the inventions disclosed herein.

[0018] Figure 6 provides a side view of an example of a difference height map with a height threshold in accordance with specific embodiments of the inventions disclosed herein.

[0019] Figure 7 provides an axonometric view of a difference height map with the areas of groups of pixels in accordance with specific embodiments of the inventions disclosed herein.

[0020] Figure 8 provides an example of a difference height map using a bounding box to estimate the volume of an object in accordance with specific embodiments of the inventions disclosed herein.

[0021] Figure 9 provides an example of a classification of an object based on its volume in accordance with specific embodiments of the inventions disclosed herein.

[0022] Figure 10 provides a flowchart of a method for volumetric estimation in accordance with specific embodiments of the inventions disclosed herein.

[0023] Figure 11 provides an example of a system with a camera, processor, and computer- readable media in accordance with specific embodiments of the inventions disclosed herein.

[0024] Figure 12 provides an example of a system for performing volumetric estimation in accordance with specific embodiments of the inventions disclosed herein.

[0025] Figure 13 provides a method for detecting the presence of a physical object in accordance with specific embodiments of the inventions disclosed herein.

[0026] Figure 14 provides a method for determining the volume of a second object in accordance with specific embodiments of the inventions disclosed herein.

[0027] Figure 15 provides a method for recalibrating a pair of imagers in accordance with specific embodiments of the inventions disclosed herein.DETAILED DESCRIPTION

[0028] Reference will now be made in detail to implementations and embodiments of various aspects and variations of systems and methods described herein. Although several exemplary variations of the systems and methods are described herein, other variations of the systems and methods may include aspects of the systems and methods described herein combined in any suitable manner having combinations of all or some of the aspects described.

[0029] Different systems and methods for volumetric estimation using a stereo camera in accordance with the summary above are described in detail in this disclosure. The methods and systems disclosed in this section are nonlimiting embodiments of the invention, are provided for explanatory purposes only, and should not be used to constrict the full scope of the invention. It is to be understood that the disclosed embodiments may or may not overlap with each other. Thus, part of one embodiment, or specific embodiments thereof, may or may not fall within the ambit of another, or specific embodiments thereof, and vice versa. Different embodiments from different aspects may be combined or practiced separately. Many different combinations and sub-combinations of the representative embodiments shown within the broad framework of this invention, that may be apparent to those skilled in the art but not explicitly shown or described, should not be construed as precluded.

[0030] Existing techniques to measure the volume of an object either require the object to be moving (e.g., on a conveyor) and a projection device (using lines or points), or require the capture device (e.g., camera) to move around the object. For instance, N. Bandi et al., Image Based Volume Estimation Using Stereo Vision, 2020 IEEE 18th International SISY, September 1, 2020, describes a volume estimation based on multiple views of a stereo camera to reconstruct the object. Furthermore, existing methods generally approximate the object to a cube or a similar simple shape. This limits the applicability of the approaches to more general object volume estimation.

[0031] A disclosed method for measuring the volume of an object uses a stereo camera (e.g., a ZED stereo camera) to compute a depth estimation (e.g., a point cloud) which is thentransformed into a height map. The height map can be used to compute the volume of the object by determining the difference between the height map with the object present and the height map with the object not visible. To detect the object, a contour detection method can be done based on a height map difference image.

[0032] The methods disclosed herein can involve determining a reference height map. The geometric characteristics of a reference surface can be determined in the form of a reference height map as explained in more detail below. The process can involve a calibration during a first time. The reference surface can be any surface, plane or not, on which the object will be placed. For example, the reference surface can be a ground plane in a warehouse, a conveyor belt surface, or a loading surface. During this first time, a stereo camera can detect the reference plane without the object.

[0033] In specific embodiments, the stereo camera is preferably static so that a single calibration can be used for all volume estimations in the reference plane. Alternatively, if the camera is not static, the calibration can be repeated or adjusted as needed.

[0034] The data obtained during the calibration step can then be utilized to determine a three- dimensional (3D) point cloud and the projection thereof in a grid map (e.g., height map). A 3D point cloud can be captured using a stereo camera. The 3D point cloud may then be projected on a grid map. The grid map may be used to determine the volume of the object. The object may be identified (e.g., as one of a class of objects) based on its volume.

[0035] Embodiments disclosed herein allow detection, volume calculation, and identification of an object without needing either the object or the measuring device to be moving.Additionally, many object measurements may be made based on a single camera calibration. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0036] Fig. 1 illustrates an example of camera 102 directly above reference surface 103 at two different times. Calibration 101 may refer to a first time while measurement 151 may refer to a second time. Camera 102 may be a stereo camera containing at least two imagers. Calibration 101 may calibrate camera 102 by determining a reference height map based on camera view 104. Between the times for calibration 101 and measurement 151, object 155 may be placedonto reference surface 103 (or removed from reference surface 103, if measurement 151 occurs before calibration 101). Measurement 151 may determine an object height map based on camera view 154. Calibration 101 may occur before measurement 151 during an initialization of the stereo camera system. In specific embodiments, calibration 101 may occur after measurement 151.

[0037] In alternative embodiments, the reference height map can be determined by inputting the position of camera 102 (e.g., attitude, orientation and altitude) with respect to reference surface 103, inputting the shape of reference surface 103 (e.g. plane) and computing a reference height map from those values. In other words, the reference height map may be computed using inputs (e.g., user inputs) about the environment (reference surface 103, camera 102, etc.) rather than be computed using measurements performed by camera 102.

[0038] The geometric characteristics of reference surface 103 can be determined in the form of a reference height map. Reference surface 103 can be any surface, plane or not, on which object 155 (or another object) may be placed. For example, reference surface 103 can be a ground plane in a warehouse, a conveyor belt surface, or a loading surface. In applications in which the object is placed on a moving surface (e.g., a conveyor belt), a single shot estimation may be used for the height map with the (moving) object. During calibration 101, camera 102 can detect the reference plane without the object. If no object has been placed in the reference plane then the "measurement" height map may be identical to the reference height map based on camera view 104.

[0039] In specific embodiments, camera 102 is static so that a single calibration 101 can be used for all volume estimations of objects on reference surface 103. Alternatively, if camera 102 is not static, calibration 101 may be repeated or adjusted as needed such that the relationship between camera 102 and reference surface 103 is the same for both calibration 101 and measurement 151.

[0040] The data obtained during calibration 101 can be utilized to determine a 3D point cloud and the projection thereof in a grid map (e.g., height map). The height map corresponding to calibration 101 may be referred to as a reference height map. The data obtained during measurement 151 can be utilized to determine a 3D point cloud and the projection thereof in agrid map (e.g., height map). The height map corresponding to measurement 151 may be referred to as an object height map. The calculated difference between the reference height map and the object height map may be referred to as a difference height map. The difference height map may be used to determine the volume of object 155. In specific embodiments, object 155 may be identified (e.g., as one of a class of objects) based on its volume. In specific embodiments, object 155 may also be identified based on other characteristics such as color and specific dimensions.

[0041] Fig. 2 illustrates an example of camera 202 with an axonometric view of reference surface 203 at two different times. The system of Fig. 2 may be similar to the system of Fig. 1, except that camera 202 has a different position relative to reference surface 203 than camera 102 has with reference surface 103. Calibration 201 may refer to a first time while measurement 251 may refer to a second time. The first time may be before or after the second time. Camera 202 may be a stereo camera containing at least two imagers. Calibration 201 may calibrate camera 202 by determining a reference height map based on camera view 204 (without object 255). Between calibration 201 and measurement 251, object 255 may be placed onto (or removed from) reference surface 203. Measurement 251 may determine an object height map (with object 255) based on camera view 254.

[0042] The data obtained during calibration 201 can be utilized to determine a 3D point cloud and the projection thereof in a grid map (e.g., height map). The height map corresponding to calibration 201 may be referred to as a reference height map. The data obtained during measurement 251 can be utilized to determine a 3D point cloud and the projection thereof in a grid map (e.g., height map). The height map corresponding to measurement 251 may be referred to as an object height map. The difference between the reference height map and the object height map may be referred to as a difference height map. The difference height map may be used to determine the volume of object 255. In specific embodiments, object 255 may be identified (e.g., as one of a class of objects) based on its volume.

[0043] Although two examples of camera position are shown (camera directly above the reference frame in Fig. 1 and camera with an axonometric view in Fig. 2), the camera may be placed in any position relative to the reference surface.

[0044] Embodiments disclosed herein allow detection, volume calculation, and identification of an object without needing either the object or the measuring device to be moving. That is, although camera 202 may be placed in a variety of different locations, camera 202 once placed may refrain from moving. Additionally, many object measurements may be made based on a single camera calibration. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0045] Fig. 3 illustrates an example point cloud 301 projected onto a grid map with grid cell 302. Point cloud 301 contains eight voxels for the sake of illustration but may contain any number of voxels. Point cloud 301 may be 3D and may be captured using a stereo camera.

[0046] Point cloud 301 may be made from data obtained during a calibration step (such as calibration 101 or calibration 201) or from a measuring step (such as measurement 151 or measurement 251). The data may include points 303 with confidence values 304. The data can be utilized to determine the projection (e.g., conversion) of point cloud 301 to a grid map.

[0047] The conversion of the point cloud can be done by dividing the environment of point cloud 301 into cells (e.g., voxels such as voxel 306) of a known dimension (for example 1cm x lcm), and then each point 303 of point cloud 301 can be projected into a specific grid cell 302 to form the grid map. Points 303 of a voxel, along with the point confidence values 304, may be combined (e.g., averaged) to form estimated centroid 305 of grid cell 302. Estimated centroid 305 may be associated with centroid confidence value 307. The grid map of grid cells 302 can then be converted into a height map.

[0048] Fig. 4 illustrates reference height map 401, object height map 431, and difference height map 461. Each height map may be converted from a grid map and / or a point cloud. Each height map shows a cross section or side view of pixels along the x-direction.

[0049] Reference height map 401 may be obtained from a reference surface (e.g., plane) without an object. Reference height map 401 can provide a height value in the z-direction for a set of pixels (e.g., columns) that are located in the x-y plane of the reference surface (although only a single set of pixels along the x-direction are shown). In alternative embodiments, reference height map 401 can be determined by inputting the position of the camera (attitude, orientation and altitude) with respect to the reference surface, inputting the shape of thereference surface (e.g. plane) and computing a reference height map from those values. In other words, reference height map 401 may be computed using inputs (e.g., user inputs) about the environment (reference surface, camera, etc.) rather than be computed using measurements performed by the camera.

[0050] Object height map 431 may be obtained from the reference surface with the object placed on the reference surface. The object may be placed in the zone of interest of the stereo cameras (e.g., in camera view, on the reference surface). Object height map 431 can provide a height value in the z-direction for a set of pixels that are located in the x-direction and y- direction of the reference surface (although only a set of pixels along the x-direction are shown). Object height map 431 can have the same x and y dimensions as reference height map 401. In specific embodiments, the reference surface may be a moving surface (e.g., a conveyor belt). In this case, a single shot estimation may be used to determine object height map 431 with the (moving) object.

[0051] Difference height map 461 may be the difference between reference height map 401 and object height map 431. Difference height map 461 can have the same x and y dimensions as reference height map 401 and object height map 431 and may be comprised of values that are equal to the difference between the values (e.g., z-values) of object height map 431 and reference height map 401. When making difference height map 461, it is possible to derive the shape of the object by assuming that all the cells of the grid map that define the object are above the reference plane.

[0052] Embodiments disclosed herein allow detection, volume calculation, and identification of an object without needing either the object or the measuring device to be moving.Additionally, many difference height maps may be made based on a single reference height map. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0053] Fig. 5 illustrates difference height map 561, which may be derived from a grid map. It is possible to derive the shape of an object used to make difference height map 561 by assuming that all the cells of the grid map that define the object are above the reference surface. Portion 501 of difference height map 561 defines the shape of the object (e.g., in the x-y plane).Portion 500 relates to the reference surface used for difference height map 561. Using standard contour extraction, it is possible to extract the object itself and compute the volume thereof.

[0054] In specific embodiments, methods disclosed herein continue with determining a volume of the object. Each pixel of difference height map 561 defines a specific cell of the corresponding grid map. Therefore, since the size of the cell is known (e.g., set when making the grid map), the volume of the object can be calculated with the sum of each pixel height of portion 501 multiplied by the size of the cell. The volume would therefore be a sum of values where the values are the products of the heights, the cell size x-dimensions, and the cell size y- dimensions. For example, Volume of object - Sum (Height(cm) * cell size-x(cm) * cell size-y (cm)).

[0055] In specific embodiments of the invention, where the object is of a known shape, the methods can conduct an operation on an approximate shape of the object. For example, if it were known that the object was approximately rectangular, the outline of the shape could be computed from a rectangle extraction on the difference height map to give the length and width of the rectangle. The average height of each cell in the rectangle could then be taken as a height value of the rectangle and all three values could be multiplied together to obtain the volume of the object. For example, Volume of object - Length * Width * Height. If the known object were nonrectangular (e.g., approximately spherical, pyramidal, etc.) then the appropriate equations may be used to determine those volumes.

[0056] In specific embodiments of the invention, the methods can also utilize a process of shape extraction conducted on the difference map. The process can involve determining a list of two-dimensional (2D) shapes where each 2D shape bounds a blob (e.g., group) of pixels from the difference map and the list of 2D shapes approximately defines the content of the difference map. The blob of pixels may be above a threshold height. The 2D shapes can be various shapes such as circles, squares, rectangles, etc. In specific embodiments, inclusion within a blob of pixels can be predicated on the height value of a given pixel in the difference map exceeding a threshold.

[0057] Difference height map 561 may be used to detect an object type. For example, calculations or measurements using difference height map 561 may detect the presence of a pallet. The calculations or measurements may also determine which type of pallet the pallet is. For example, if a detected length of a pallet is 104 cm, then the system may determine that the pallet is a EURO pallet (e.g., which are around 100 cm long). If a detected length of a pallet is 84 cm, then the system may determine that the pallet is a CHEP pallet (e.g., which are around 80 cm long).

[0058] Fig. 6 illustrates an example of difference map 600 with threshold 601. Threshold 601 indicates a minimum height for a pixel to be included within a blob of pixels determined to correspond to an object. For example, set of pixels 602 and set of pixels 604 are both below threshold 601. Accordingly, set of pixels 602 and set of pixels 604 are not recognized as corresponding to the object (e.g., may be ignored). Set of pixels 603, however, is above threshold 601. Accordingly, set of pixels 603 is recognized as corresponding to the object. In specific embodiments, threshold 601 reduces noise and increases accuracy of object measurements, such as object shape, height, and volume.

[0059] Fig. 7 illustrates an example of an axonometric view of difference height map 700 showing the heights and locations of several pixels that may correspond to an object.Difference height map 700 may show only pixel heights that satisfy (e.g., are equal to or above) a threshold. Difference height map 700 includes two sets of pixels: pixels 701 corresponding to an object, and pixels 705 that may correspond to noise. In specific embodiments, difference height map 700 may be used to determine the volume of the object. There are many ways in which difference height map 700 may be used to determine the volume of the object.

[0060] In specific embodiments, difference height map 700 may be derived from a grid map. The shape of the object used to make difference height map 700 may be derived by assuming that all the cells of the grid map that define the object are above the reference surface. Using standard contour extraction, it is possible to extract the object itself and compute the volume thereof.

[0061] Pixels 701 may correspond to area 702 in the x-y plane. Pixels 705 may correspond to area 706 in the x-y plane. In specific embodiments, detecting the presence of the object maybe based on a grouping of pixels in difference height map 700 having a minimum area. For example, area 706 may be too small, such that the system does not detect an object corresponding to pixels 705. In specific embodiments, the system may ignore pixels 705 because the corresponding area 706 is below an area threshold even if the heights of pixels 705 are above a height threshold. The system may detect an object corresponding to pixels 701 based on area 702 satisfying (e.g., meeting, exceeding) an area threshold. The thresholds may act to reduce noise and improve accuracy of the system.

[0062] In specific embodiments of the invention, one or more 2D shapes may be mapped to groupings of pixels, such as pixels 701 and pixels 705. The presence of a physical object can be determined when the areas of the shapes (e.g., in the list of 2D shapes) have an area that exceeds a minimum area. For example, pixels 705 may be associated with a square. The square may have an area (e.g., similar to area 706), which may be below a threshold. The system may ignore pixels 705 based on the area of the square being below the threshold.Pixels 701 may be associated with a rectangle. The rectangle may have an area (e.g., similar to area 702), which may be above the threshold. The system may detect an object associated with pixels 705 based on the area of the rectangle being above the threshold.

[0063] In specific embodiments, each pixel 701 of difference height map 700 defines a specific cell of the corresponding grid map. Therefore, since the size of the cell is known (e.g., set when making the grid map), the volume of the object can be calculated with the sum of the height of each pixel 701 multiplied by the size of each cell. The volume would therefore be a sum of values where the values are the products of the heights of pixels 701, the cell size x-dimensions, and the cell size y-dimensions.

[0064] In specific embodiments of the invention, where the object is of a known shape, the methods can conduct an operation on an approximate shape of the object. For example, if it were known that the object was approximately rectangular, the outline of the shape could be computed from a rectangle extraction on difference height map 700 to give the length and width of the rectangle. In the example of Fig. 7, the object is a rectangle 18 pixels by 8 pixels. The average height of each pixel 701 in the rectangle could then be taken as the height value ofthe rectangle and all three values (height, length, width) could be multiplied together to obtain the volume of the object.

[0065] In specific embodiments of the invention, the methods can also utilize a process of shape extraction conducted on difference height map 700. The process can involve determining a list of 2D shapes where each 2D shape bounds a blob of pixels from difference height map 700 and the list of 2D shapes approximately defines the content of difference height map 700. The 2D shapes can be various shapes such as circles, squares, rectangles, etc. Although the example of Fig. 7 would likely be approximated as a single rectangle, other examples of height maps may include irregular objects that are approximated as a conglomeration of multiple different shapes. That is, the object may be approximated as an equilateral triangle attached to an ellipse, or a few isosceles triangles connected to a rectangle and a circle, etc. The average height of each pixel 701 associated with the 2D shape may be taken as the height value of the shape and could be used together with a calculation of the surface area of the shape to obtain the volume of the object attributable to the given shape. The process of averaging heights of pixels within a shape and multiplying that average height by the surface area of that shape could be conducted for all the shapes attributed to (e.g., that make up) the object to derive the volume of the object.

[0066] Fig. 8 illustrates an example of difference height map 800 using bounding box 802 to estimate the volume of object 803. Bounding box 802 may be oriented along the axis of difference height map 800 or may be skewed. In specific embodiments, bounding box 802 may include x, y, and z coordinates. In specific embodiments, bounding box 802 may only define the polygon area. Bounding box 802 may be based on a detection of the contour or area of the blob of pixels corresponding to object 803. Although illustrated as a rectangle in the example of Fig. 8, bounding box 802 may be any shape (e.g., from a list of 2D shapes) that fits the pixels of object 803 within a tolerance. For example, another object may be more suited to be contained by a circle bounding box than a rectangle bounding box.

[0067] Bounding box 802 operates on the set of pixels corresponding to object 803. In specific embodiments, the set of pixels may be associated with 2D shapes from a list of 2D shapes and may fit in bounding box 802 in accordance with a desired tolerance. The average heights of thepixels in bounding box 802 may be used in combination with knowledge of the size of bounding box 802 in a product computation to compute an estimate of the volume of object 803, which is subsumed within bounding box 802. Using a bounding box to determine the volume of the object is agnostic of the object itself. The system may not be required to have seen the object before.

[0068] A reference height map, without object 803, may be determined for the region of interest. Object 803 has a specific size and may be placed on a reference surface (e.g., the floor). An object height map (with object 803) may be measured and compared to the reference map (without the object). By making a comparison of both height maps (e.g., the object height map minus the reference height map), a difference height map may be formed and bounding box 802 that defines object 803 may be detected. A detection of contour and / or area may determine bounding box 802. In specific embodiments, bounding box 802 may be different shapes. For example, when determining the volume of the object, the bounding box may be a polygon area. As another example, a bounding rectangle may be used to give an X, Y, and Z measurement.

[0069] X, Y, and Z measurements may refer to width of the bounding rectangle, height or length of the bounding rectangle, and height of the object, respectively. The X measurement may refer to the width of bounding box 802. For example, X may be calculated as the width of bounding box 802 in pixels multiplied by the size of a pixel (e.g., height map resolution). The Y measurement may refer to the height of bounding box 802. For example, Y may be calculated as the height of bounding box 802 in pixels multiplied by the size of a pixel (e.g., height map resolution in the Y-axis). The Z measurement may refer to the average height inside bounding box 802. For example, Z may be calculated using the calculated difference height map. The volume of object 803 may be calculated as Volume - X * Y * Z. Instead of searching for a specific object (e.g., through artificial intelligence (Al) classification) which may only work for already-known objects, using bounding boxes may be completely agnostic of the object to be measured. There is no need for the stereo imaging system to have seen the specific object before; the object simply needs to have a "height." Additionally, only one camera position is needed to create the object height map and thus determine the volume of the object. Inapplications in which the object is placed on a moving surface (e.g., a conveyor belt), a single shot estimation may be used for the height map with the (moving) object.

[0070] The volume of object 803 may be calculated as Volume - Sumfsubvolume of a single cell) for the all cells in the object height map (e.g., within bounding box 802). This calculation method may refrain from extracting (e.g., may not use) X, Y, and Z dimensions of object 803. This volume estimation method may be more precise than the X*Y*Z method for non-cubic objects. Volume estimation may be based on multiple pixels (e.g., subcells) with their own height.

[0071] Fig. 9 illustrates a classification of object 901. In specific embodiments, the process of determining a volume of a physical object could be used to classify the physical object as belonging to a specific class because it matches the expected volume of that class of objects (i.e., to identify a physical object as being a particular class of physical objects).

[0072] In the example of Fig. 9, object 901 has a measured volume 902 of value D. The value D may correspond to any volume, range of volume, or specific volume with a tolerance.Measured volume 902 may be compared to a list of reference volumes 903. According to the list of reference volumes 903, a value of D as a volume (e.g., measured volume 902) may correspond to object classification 904 type iv. Accordingly, object 901 may be classified as type iv. Any object with a measured volume 902 of value D may be categorized as object classification 904 type iv. As another example, any object with a measured volume of value A may be categorized as object classification 904 type i.

[0073] Fig. 10 illustrates flowchart 1000 of a method for volumetric estimation according to embodiments of the inventions disclosed herein. Steps of flowchart 1000 may be omitted, duplicated, rearranged, or otherwise deviate from the method as depicted. Flowchart 1000 may be performed by a system comprising a pair of images, one or more processors, and one or more computer readable media.

[0074] At step 1002, a reference height map may be determined. The reference height map may act to calibrate the pair of imagers to a reference frame. The reference frame can include ground plane in a warehouse, a conveyor belt surface, a loading surface, or other reference surface.

[0075] At step 1004, an object may be placed in the reference frame (e.g., onto the reference surface). The object may be placed in the reference frame manually, by a conveyor belt, by another machine, or by some other means.

[0076] At step 1006, a height map with the object may be determined. That is, a second height map may be determined where this height map is different than the reference height map (from step 1002) due to the presence of the object.

[0077] At step 1008, a difference height map may be determined. The difference height map may be the result of the height map with the object (from step 1006) minus the reference height map (from step 1002).

[0078] At step 1010, a volume of the physical object may be determined based on the difference map (from step 1008). The volume may be based on a bounding box fitting a portion of the difference map and an average height of the portion of the difference map. The volume may be based on contour extraction. The difference map may be made up of pixels of a known size and the volume may be based on the sum of the height of each pixel multiplied by the size of the pixel. The volume of the physical object may be based on the outline of the object and the shape (e.g., known shape, extracted shape) of the object. The physical object may be broken down into a conglomeration of multiple shapes and the volume of the object may be based on the surface areas and average heights of these shapes.

[0079] At step 1012, the object may be removed from the reference frame. The object may be removed from the reference frame manually, by a conveyor belt, by a machine, or by some other means. In specific embodiments, the object may be removed before the difference map or the volume of the object are determined. In other words, step 1012 may occur before (or in between) step 1008 and step 1010. In specific embodiments, step 1012 may occur before step 1006 is fully realized. That is, the system may have taken measurements to determine the height map with the object but may not have determined the height map itself when the object is removed.

[0080] At step 1014, whether the camera (imagers) has been relocated may be determined. Relocation of the camera may include any movement including translational or rotational movement. Relocation of the camera may also refer to a relocation or change of the referencesurface. For example, if the reference surface were a table and the table was adjusted (e.g., raised, tilted, etc.), then the camera may be considered relocated, as the relative position of the camera to the reference surface has changed. The camera may also be considered relocated by user input. That is, a user may request a recalibration of the camera with the reference surface. If the camera has been relocated, the method may continue to step 1002 to determine another reference height map (e.g., a third height map). In other words, the camera may be recalibrated. If the camera has not been relocated, the method may continue to step 1004. In other words, the system may repeat the process of determining the volume of an object with another (e.g., second) object including determining another (e.g., third) height map.

[0081] Embodiments disclosed herein allow detection, volume calculation, and identification of an object without needing either the object or the measuring device to be moving.Additionally, many object measurements may be made based on a single camera calibration (e.g., reference height map) with methods for recalibrating (e.g., creating new reference height maps) as needed or desired. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0082] Fig. 11 illustrates an example of system 1100 capable of performing the methods described herein. System 1100 includes camera 1101 with imager 1102 and imager 1103. In specific embodiments, system 1100 includes processor 1104 and non-transitory computer- readable media 1105. Non-transitory computer-readable media 1105 may store instructions that, when executed by processor 1104, cause system 1100 to conduct a stereo imaging method for detecting physical objects.

[0083] The stereo imaging method for detecting physical objects may include determining a first height map at a first time using imager 1102 and imager 1103, determining a second height map at a second time using imager 1102 and imager 1103 (where the first time and second time are different), determining a difference map (the difference map including a difference between a set of heights of the first height map and a set of heights of the second height map), and detecting a presence of a physical object based on the difference map.

[0084] In specific embodiments, system 1100 includes camera 1101 with imager 1102 and imager 1103 and various means for performing a method for detecting physical objects.System 1100 may include a means for determining a first height map at a first time using imager 1102 and imager 1103 of system 1100, a means for determining a second height map at a second time using imager 1102 and imager 1103 of system 1100 (where the first time and the second time are different), a means for determining a difference map (the difference map including a difference between a set of heights of the first height map and a set of heights of the second height map), and a means for detecting a presence of a physical object based on the difference map.

[0085] In specific embodiments, system 1100 may also include a combination of: means for determining a volume of the physical object based on height values of the difference map, wherein the height values are based on the difference between the set of heights of the first height map and the set of heights of the second height map; means for determining a volume of the physical object based on a bounding box fitting a portion of the difference map and an average height of the portion of the difference map; means for moving the physical object into a reference frame before the second time and after the first time; means for determining a third height map at a third time using the pair of imagers of the stereo imaging system; means for determining a second difference map, the second difference map including a difference between the set of heights of the first height map and a set of heights of the third height map; means for determining a volume of a second physical object based on the second difference map; means for detecting a movement of the pair of imagers; means for determining, in response to detecting the movement, a fourth height map at a fourth time using the pair of imagers of the stereo imaging system; and means for replacing the first height map with the fourth height map for future use in determining additional difference maps.

[0086] The above means may include any kind of processor or computing system, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), processors, microcontrollers, central processing units (CPUs), microprocessors, embedded processors, digital signal processors (DSPs), media processors, servers, etc. Means for moving the physical object into a reference frame may additionally include manual placement, conveyor belt, robotic arm, slide, drop, lift, push, pull, magnet, or any other means by human, machine, or other.

[0087] Non-transitory computer-readable media 1105 may store instructions capable of causing one or more processors to execute the means functions above. Instructions for determining a first height map may include taking a photo with each imager and extracting height information from the photos. Instructions for determining a second height map may include taking a photo with each imager and extracting height information from the photos. Instructions for determining a difference map may include instructions for calculating differences in height information of the first height map and the second height map and storing the differences. Instructions for detecting the presence of a physical object may include comparing height information of the difference map to an area threshold and to a height threshold.

[0088] In specific embodiments, instructions for determining a volume of the physical object based on height values of the difference map may include reading height values of the difference map and multiplying the height values by cell dimensions. Instructions for determining a volume of the physical object based on a bounding box fitting a portion of the difference map and an average height of the portion of the difference map may include assigning a bounding box to the difference map and calculating the average height of the portion of the difference map. Instruction for moving the physical object into a reference frame may include turning a conveyor belt on and turning a conveyor belt off. Instructions for determining a third height map may include taking a photo with each imager and extracting height information from the photos. Instructions for determining a second difference map, the second difference map including a difference between the set of heights of the first height map and a set of heights of the third height map, may include instructions for calculating differences in height information of the first height map and the third height map and storing the differences. Instructions for determining a volume of a second physical object may include comparing heights of pixels of the difference map to an area threshold and a height threshold. Instructions for detecting a movement of the pair of imagers may include taking a photo of the reference surface at a first time and taking a photo of the reference surface at a second time or may include getting data from an accelerometer; instructions for determining, in response to detecting the movement, a fourth height map may include taking a photo with each imager andextracting height information from the photos. Instructions for replacing the first height map with the fourth height map for future use in determining additional difference maps may include erasing the first height map, writing the fourth height map into memory, and tagging the fourth height map as the reference height map.

[0089] The imagers may include image sensors. The image sensors may detect and convey information used to form an image. The image sensors may convert light waves into current bursts. The imagers may include charge-coupled devices (CCDs), active-pixel sensors (CMOS), vacuum tubes, flat-panel detectors, or any kind of light image sensor.

[0090] Fig. 12 illustrates system 1200 that may perform volume computation 1223 using camera 1201 in accordance with the inventions disclosed herein. In specific embodiments of the invention, system 1200 may be used to detect the presence of a physical object or to determine the volume of a physical object. System 1200 may include aspects of system 1100. System 1200 may include camera 1201 (including imager 1202 and imager 1203), reference point cloud 1204, reference height map 1205 (including set of heights 1206), object 1213, object point cloud 1214, object height map 1215 (including set of heights 1216), difference height map 1220, shape extraction 1221, and volume computation 1223.

[0091] System 1200 may include two branches. Branch 1207 may refer to reference point cloud 1204 and reference height map 1205, each referring to a reference surface without an object. In other words, branch 1207 may refer to calibrating camera 1201. Branch 1217 may refer to object point cloud 1214 and object height map 1215, each referring to a reference surface with an object. In other words, branch 1207 may refer to measuring object 1213. Branch 1207 may be implemented before or after branch 1217.

[0092] Camera 1201 may use imager 1202 and imager 1203 to create reference point cloud 1204. Reference point cloud 1204 may be used to derive reference height map 1205.Reference height map 1205 may be a grid map or may be derived from a grid map. Reference height map 1205 may include set of heights 1206. Set of heights 1206 may refer to the heights of various pixels of the reference surface. Reference height map 1205 may be used to calibrate camera 1201 to the reference surface and the angle at which camera 1201 views the reference surface.

[0093] Camera 1201 may use imager 1202 and imager 1203 to create object point cloud 1214. Before creating object point cloud 1214, object 1213 may be placed on the reference surface. The placement of object 1213 may be unique to branch 1217 and may not occur in branch 1207. Object point cloud 1214 may be used to derive object height map 1215. Object height map 1215 may be a grid map or may be derived from a grid map. Object height map 1215 may include set of heights 1216. Set of heights 1216 may refer to the heights of various pixels of the reference surface and of the object on the reference surface. Object height map 1215 may be used to detect object 1213.

[0094] Difference height map 1220 may be calculated using reference height map 1205 and object height map 1215. Difference height map 1220 may correspond to set of heights 1206 subtracted from set of heights 1216 and may refer to the height of object 1213 without the height (or other data) of the reference surface. Difference height map 1220 may be used to detect object 1213. Difference height map 1220 may also be used to measure the shape, surface area, and / or volume of object 1213.

[0095] The shape of object 1213 may be extracted from difference height map 1220. Shape extraction 1221 may organize difference height map 1220 into one or more shapes (e.g., from a list of 2D shapes, from an expected 3D shape, etc.). Volume computation 1223 of object 1213 may be derived from shape extraction 1221. Volume computation 1223 of object 1213 may be done without prior knowledge of object 1213 by system 1200.

[0096] Embodiments disclosed herein allow detection, shape extraction 1221, volume computation 1223, and identification of object 1213 without needing either object 1213 or camera 1201 to be moving. Additionally, many difference height maps 1220 may be made based on a single reference height map 1205. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0097] Fig. 13 illustrates method 1300 of detecting the presence of a physical object in accordance with the inventions disclosed herein. Steps, or portions of steps, of method 1300 may be omitted, duplicated, rearranged, or otherwise deviate from the form shown. Method 1300 may be implemented using a stereo imaging system including a pair of imagers. The pairof imagers may be any of a variety of 3D sensors. Method 1300 may be used to detect a variety of objects or surfaces. For example, method 1300 may detect different types of pallets (e.g., EURO, CHEP, and INDUS) where the only difference is the dimension in the y-axis.

[0098] At step 1302, a first height map may be determined at a first time. The first height map may be determined using a pair of imagers of the stereo imaging system. In specific embodiments, the first time is at a calibration time and the physical object is not present in the reference plane during the first time. The reference plane may refer to a surface on which an object may be placed and may correspond to a reference frame of the imagers.

[0099] In specific embodiments, at step 1304, a physical object may be moved into the reference frame. The physical object may be moved into the reference frame after the first time but before a second time. The physical object may be moved manually, by a machine (conveyer belt, etc.), or by another means.

[0100] At step 1306, a second height map may be determined at the second time. The second height map may be determined using the pair of imagers of the stereo imaging system. The first time and second time may each refer to different times. In specific embodiments, the physical object is present in the reference plane at the second time. The second time may be before or after the first time as long as the object is present in the reference frame at the second time and is not present in the reference frame at the first time.

[0101] At step 1308, a difference map may be determined. The difference map may include a difference between a set of heights of the first height map (determined at step 1302) and a set of heights of the second height map (determined at step 1306).

[0102] At step 1310, the presence of the physical object may be detected based on the difference map. The physical object may be the physical object from step 1304. In specific embodiments, detecting the presence of the physical object may ignore values of the difference map that are below a threshold value (e.g., a height threshold, an area threshold). In specific embodiments, detecting the presence of the physical object may be based on the difference map having a minimum area. In specific embodiments, detecting the presence of the physical object may be conducted without the stereo imaging system having prior knowledge of the physical object.

[0103] Method 1300 may include many ways to determine the volume of the physical object, including step 1312 and step 1314. In specific embodiments, step 1312 may be performed, step 1314 may be performed, steps 1312 and 1314 may be performed in combination, or neither may be performed.

[0104] In specific embodiments, at step 1312, a volume of the physical object may be determined based on a bounding box fitting a portion of the difference map and on an average height of the portion of the difference map.

[0105] In specific embodiments, at step 1314, a volume of the physical object may be determined based on height values of the difference map. The height values of the difference map may be based on the difference between the set of heights of the first height map and the set of heights of the second height map.

[0106] Embodiments disclosed herein allow detection, volume calculation, and identification of an object without needing either the object or the measuring device to be moving.Additionally, many object measurements may be made based on a single camera calibration. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0107] Fig. 14 illustrates method 1400 of determining the volume of a second object. Steps, or portions of steps, of method 1400 may be omitted, duplicated, rearranged, or otherwise deviate from the form shown. Method 1400 may be implemented using a stereo imaging system. Method 1400 may be a continuation of method 1300.

[0108] At step 1401, steps of method 1300 may be performed.

[0109] In specific embodiments, at step 1402, a third height map may be determined at a third time. The third height map may be determined using the pair of imagers of the stereo imaging system. In specific embodiments, the third height map may be associated with a second physical object. The third height map may be determined before or after the first height map and / or the second height map.

[0110] At step 1404, a second difference map may be determined. The second difference map may include a difference between the set of heights of the first height map (from step 1302) and a set of height of the third height map (from step 1402).

[0111] In specific embodiments, at step 1406, a volume of the second physical object may be determined based on the second difference map (from step 1404).

[0112] Fig. 15 illustrates method 1500 of recalibrating the pair of imagers. Steps, or portions of steps, of method 1500 may be omitted, duplicated, rearranged, or otherwise deviate from the form shown. Method 1500 may be implemented using a stereo imaging system. Method 1500 may be a continuation of method 1300 or method 1400.

[0113] At step 1501, steps of method 1300 may be performed. In specific embodiments, steps of method 1400 may also be performed.

[0114] At step 1502, a movement of the pair of imagers may be detected. That is, the system may detect that the pair of imagers have moved from a first position relative to a reference surface to a second position relative to the reference surface, or the reference surface may have been moved, or the pair of imagers may have been moved to a different reference surface. The first position may be associated with the first height map.

[0115] At step 1504, a third height map may be determined at a third time. The third height map may be determined in response to detecting the movement of the pair of imagers (e.g., at step 1502). The third height map may be determined using the pair of imagers of the stereo imaging system and may be associated with the second position.

[0116] At step 1506, the first height map may be replaced with the third height map for future use in determining additional difference maps. In other words, step 1308 of method 1300 may be performed using the third height map from step 1504 (and its corresponding set of heights) rather than the first height map from step 1302 (and its corresponding set of heights). In specific embodiments, the first height map may be discarded (e.g., overwritten, erased, etc.).

[0117] Embodiments disclosed herein allow detection, volume calculation, and identification of an object without needing either the object or the measuring device to be moving.Additionally, many object measurements may be made based on a single camera calibration. Accordingly, the embodiments may allow for increased efficiency, decreased power usage, decreased system complexity, and improved performance.

[0118] At least one processor in accordance with this disclosure can include at least one non- transitory computer-readable media. The at least one processor could comprise at least onecomputational node in a network of computational nodes. The media could include cache memories on the processor. The media can also include shared memories that are not associated with a unique computational node. The media could be a shared memory, could be a shared random-access memory, and could be, for example, a double data rate (DDR) dynamic random-access memory (DRAM). The shared memory can be accessed by multiple channels. The non-transitory computer-readable media can store data required for the execution of any of the methods disclosed herein, the instruction data disclosed herein, and / or the operand data disclosed herein. The computer-readable media can also store instructions which, when executed by the system, cause the system to execute the methods disclosed herein. The concept of executing instructions is used herein to describe the operation of a device conducting any logic or data movement operation, even if the "instructions" are specified entirely in hardware (e.g., an AND gate executes an "and" instruction). The term is not meant to impute the ability to be programmable to a device.

[0119] While the specification has been described in detail with respect to specific embodiments of the invention, it will be appreciated that those skilled in the art, upon attaining an understanding of the foregoing, may readily conceive of alterations to, variations of, and equivalents to these embodiments. Any of the method steps discussed above can be conducted by a processor operating with a computer-readable non-transitory medium storing instructions for those method steps. The computer-readable medium may be memory within a personal user device or a network accessible memory. Although examples in the disclosure were generally directed to detecting the presence of an object, the same approaches could be utilized to detect the removal of an object. Similarly, although examples in the disclosure were generally directed to planar and level reference surfaces, the same approaches could be utilized on curved, irregular, tilted, and other such surfaces. These and other modifications and variations to the present invention may be practiced by those skilled in the art, without departing from the scope of the present invention, which is more particularly set forth in the appended claims.

Claims

WHAT IS CLAIMED IS:

1. A method for detecting physical objects (155; 255; 803; 1213) with a stereo imaging system (1100; 1200), comprising: determining a first height map (401; 1205) at a first time using a pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200); determining a second height map (431; 1215) at a second time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200), wherein the first time and the second time are different; determining a difference map (461; 561; 600; 700; 1220), the difference map (461; 561; 600; 700; 1220) including a difference between a set of heights (1206) of the first height map (401; 1205) and a set of heights (1216) of the second height map (431; 1215); and detecting a presence of a physical object (155; 255; 803; 1213) based on the difference map (461; 561; 600; 700; 1220).

2. The method of claim 1, wherein: the first time is at a calibration time (101; 201) and the physical object (155; 255; 803; 1213) is not present in a reference plane during the first time; and the physical object (155; 255; 803; 1213) is present in the reference plane at the second time (151; 251).

3. The method of claim 1, wherein: the detecting of the presence of the physical object (155; 255; 803; 1213) ignores values (602; 604) of the difference map (461; 561; 600; 700; 1220) that are below a threshold value (601).

4. The method of claim 1, wherein: the detecting of the presence of the physical object (155; 255; 803; 1213) is based on the difference map (461; 561; 600; 700; 1220) having a minimum area (702).

5. The method of claim 1, further comprising: determining a volume of the physical object (155; 255; 803; 1213) based on height values of the difference map (461; 561; 600; 700; 1220), wherein the height values are based on the difference between the set of heights (1206) of the first height map (401; 1205) and the set of heights (1216) of the second height map (431; 1215).

6. The method of claim 1, further comprising: determining a volume of the physical object (155; 255; 803; 1213) based on a bounding box (802) fitting a portion of the difference map (461; 561; 600; 700; 1220) and an average height of the portion of the difference map (461; 561; 600; 700; 1220).

7. The method of claim 1, further comprising: moving the physical object (155; 255; 803; 1213) into a reference frame before the second time and after the first time.

8. The method of claim 1, further comprising: determining a third height map (431; 1215) at a third time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200); determining a second difference map (461; 561; 600; 700; 1220), the second difference map (461; 561; 600; 700; 1220) including a difference between the set of heights (1206) of the first height map (401; 1205) and a set of heights of the third height map (431; 1215); anddetermining a volume of a second physical object (155; 255; 803; 1213) based on the second difference map (461; 561; 600; 700; 1220).

9. The method of claim 1, further comprising: detecting a movement of the pair of imagers (102; 202; 1102; 1103; 1202; 1203); determining, in response to detecting the movement, a third height map (401; 1205) at a third time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200); and replacing the first height map (401; 1205) with the third height map (401; 1205) for future use in determining additional difference maps.

10. The method of claim 1, wherein: the detecting of the presence of the physical object (155, 255, 803) is conducted without a prior knowledge of the physical object (155; 255; 803; 1213) by the stereo imaging system (1100; 1200).

11. A stereo imaging system (1100; 1200) for detecting physical objects (155; 255; 803;1213) comprising: a pair of imagers (102; 202; 1102; 1103; 1202; 1203); one or more processors (1104); and one or more non-transitory computer-readable media (1105) storing instructions that, when executed by the one or more processors (1104), cause the stereo imaging system (1100; 1200) to conduct a method comprising: determining a first height map (401; 1205) at a first time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200);determining a second height map (431; 1215) at a second time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200), wherein the first time and the second time are different; determining a difference map (461; 561; 600; 700; 1220), the difference map (461; 561; 600; 700; 1220) including a difference between a set of heights (1206) of the first height map (401; 1205) and a set of heights (1216) of the second height map (431; 1215); and detecting a presence of a physical object (155; 255; 803; 1213) based on the difference map (461; 561; 600; 700; 1220).

12. The stereo imaging system (1100; 1200) of claim 11, wherein: the first time is at a calibration time (101; 201) and the physical object (155; 255; 803; 1213) is not present in a reference plane during the first time; and the physical object (155; 255; 803; 1213) is present in the reference plane at the second time (151; 251).

13. The stereo imaging system (1100; 1200) of claim 11, wherein: the detecting of the presence of the physical object (155; 255; 803; 1213) ignores values (602; 604) of the difference map (461; 561; 600; 700; 1220) that are below a threshold value (601).

14. The stereo imaging system (1100; 1200) of claim 11, wherein: the detecting of the presence of the physical object (155; 255; 803; 1213) is based on the difference map (461; 561; 600; 700; 1220) having a minimum area (702).

15. The stereo imaging system (1100; 1200) of claim 11, the method further comprising: determining a volume of the physical object (155; 255; 803; 1213) based on height values of the difference map (461; 561; 600; 700; 1220), wherein the height values arebased on the difference between the set of heights (1206) of the first height map (401; 1205) and the set of heights (1216) of the second height map (431; 1215).

16. The stereo imaging system (1100; 1200) of claim 11, the method further comprising: determining a volume of the physical object (155; 255; 803; 1213) based on a bounding box (802) fitting a portion of the difference map (461; 561; 600; 700; 1220) and an average height of the portion of the difference map (461; 561; 600; 700; 1220).

17. The stereo imaging system (1100; 1200) of claim 11, the method further comprising: moving the physical object (155; 255; 803; 1213) into a reference frame before the second time and after the first time.

18. The stereo imaging system (1100; 1200) of claim 11, the method further comprising: determining a third height map (431; 1215) at a third time using the pair of imagers(102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200); determining a second difference map (461; 561; 600; 700; 1220), the second difference map (461; 561; 600; 700; 1220) including the difference between the set of heights (1206) of the first height map (401; 1205) and a set of heights of the third height map (431; 1215); and determining a volume of a second physical object (155; 255; 803; 1213) based on the second difference map (461; 561; 600; 700; 1220).

19. The stereo imaging system (1100; 1200) of claim 11, the method further comprising: detecting a movement of the pair of imagers (102; 202; 1102; 1103; 1202; 1203); determining, in response to detecting the movement, a third height map (401; 1205) at a third time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200); andreplacing the first height map (401; 1205) with the third height map (401; 1205) for future use in determining additional difference maps.

20. A stereo imaging system (1100; 1200) for detecting physical objects (155; 255; 803; 1213) comprising: a pair of imagers (102; 202; 1102; 1103; 1202; 1203); a means for determining a first height map (401; 1205) at a first time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100; 1200); a means for determining a second height map (431; 1215) at a second time using the pair of imagers (102; 202; 1102; 1103; 1202; 1203) of the stereo imaging system (1100;1200), wherein the first time and the second time are different; a means for determining a difference map (461; 561; 600; 700; 1220), the difference map (461; 561; 600; 700; 1220) including a difference between a set of heights (1206) of the first height map (401; 1205) and a set of heights (1216) of the second height map (431; 1215); and a means for detecting a presence of a physical object (155; 255; 803; 1213) based on the difference map (461; 561; 600; 700; 1220).

Citation Information

Patent Citations

  • Object volume measuring method and device

    CN110349205A

  • Method and Apparatus for Determining Volume of Object

    US20190139251A1