Z-plane identification and box size determination using 3D time-of-flight imaging

By identifying the Z-plane and filtering TOF data, the point cloud is transformed to determine the box size, solving the sensor alignment problem and enabling fast and accurate box size measurement under different lighting conditions.

CN116547559BActive Publication Date: 2026-03-10ANALOG DEVICES INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-21
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing 3D time-of-flight imaging systems have difficulty reliably aligning sensors with the environment in artificial settings, resulting in time-consuming and inconvenient box size measurements. Noise is a significant factor, especially under varying lighting conditions, and existing solutions are fragile and expensive.

Method used

By identifying the Z-plane in the environment, reducing the sensor's degrees of freedom, using a TOF sensor and processor to identify basis vectors, filtering TOF data, transforming the point cloud to determine the top and edges of the box, calculating the box size, and reducing the impact of noise under different lighting conditions.

Benefits of technology

It enables rapid and accurate measurement of box dimensions in various environments, simplifies the sensor calibration process, and improves the robustness and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116547559B_ABST
    Figure CN116547559B_ABST
Patent Text Reader

Abstract

A sensor system is provided that obtains and processes time-of-flight data (TOF) obtained in any orientation. The TOF sensor obtains distance data that describes various surfaces. A processor identifies a horizontal Z-plane in the environment and transforms the data to align with the Z-plane. In some embodiments, the environment includes a box, and the processor identifies the bottom and top of the box in the transformed data. The processor can further determine the dimensions of the box, e.g., the height between the top and bottom of the box, and the length and width of the top of the box.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 081742, filed September 22, 2020, entitled “Box Size Determination Using Three-Dimensional Time-of-Flight Imaging”, and U.S. Provisional Patent Application No. 63 / 081,775, filed September 22, 2020, entitled “World Z-Plane Recognition in Time-of-Flight Imaging”, the entire contents of which are incorporated herein by reference. Background Technology

[0003] Man-made environments are typically assigned a preferred orientation corresponding to the local direction of the Earth's gravitational field. Simply put, "up" and "down" define the natural engineered orientations of indoor environments (such as rooms) and outdoor environments (such as streets). Floors, walls, and ceilings are strongly constrained by the local gravitational direction. In particular, man-made environments are often composed of a horizontal Z-plane (e.g., tabletops, chairs, floors, sidewalks).

[0004] When a person walks through a man-made environment with a handheld 3D time-of-flight (TOF) imaging system, the sensor's angular orientation relative to the natural "up" and "down" directions is often unknown. Humans cannot reliably align the imaging system with the environment, and having the user align and realign the sensor to match the environment can be a time-consuming and frustrating process.

[0005] One potential application of TOF imaging systems is determining the dimensions of boxes. Measuring the volume of physical objects is a fundamental problem in various industrial and consumer markets, such as the packaging, transportation, and storage of goods. In typical packaging and transportation environments, humans use measuring tapes to measure the dimensions of boxes, which is a time-consuming process. Existing technological solutions are often fragile, expensive, and / or only usable in certain environments. For example, some dimensional determination solutions rely on a fixed reference frame, such as deriving the volume of a box placed on a specified surface from an image taken by a camera at a fixed position relative to that surface. Attached Figure Description

[0006] To provide a more complete understanding of this disclosure and its features and advantages, reference is made to the following description in conjunction with the accompanying drawings, wherein like reference numerals denote like parts, wherein:

[0007] Figure 1 This is a block diagram of a TOF sensor system according to some embodiments of the present disclosure.

[0008] Figure 2 The light direction of a pixel of a TOF sensor according to some embodiments of the present disclosure is shown.

[0009] Figure 3This is a flowchart illustrating a process for identifying the Z-plane in TOF data obtained in any reference frame, according to some embodiments of the present disclosure.

[0010] Figure 4 This is a flowchart illustrating a process for identifying basis vectors based on TOF data according to some embodiments of the present disclosure.

[0011] Figure 5 This is a flowchart illustrating a process for identifying the Z-plane in a point cloud based on a transformation of the identified basis vectors, according to some embodiments of the present disclosure.

[0012] Figure 6 This is a flowchart illustrating a process for determining and outputting box dimensions based on TOF data according to some embodiments of the present disclosure.

[0013] Figure 7 This is a flowchart illustrating a process for identifying the top and bottom of a box according to some embodiments of the present disclosure.

[0014] Figure 8 This is a flowchart illustrating a process for calculating the length and width of the top of a box according to some embodiments of the present disclosure.

[0015] Figure 9 This is an example image showing a box placed on a desktop according to some embodiments of the present disclosure.

[0016] Figure 10 Examples of distance data obtained by a TOF sensor according to some embodiments of this disclosure are shown.

[0017] Figure 11 An example point cloud calculated from distance data according to some embodiments of this disclosure is shown.

[0018] Figure 12A and 12B Example angular coordinates of the surface normals of points in a point cloud according to some embodiments of the present disclosure are shown.

[0019] Figure 13 This is an example histogram classifying the angular coordinates of surface normals according to some embodiments of this disclosure.

[0020] Figure 14 Examples of point clouds converted to a reference coordinate system of identified basis vectors according to some embodiments of this disclosure are shown.

[0021] Figure 15 This is an example height map obtained from a transformed point cloud according to some embodiments of the present disclosure.

[0022] Figure 16This is an example Z-profile of a height map having peak values ​​indicating various horizontal surfaces, according to some embodiments of this disclosure.

[0023] Figure 17 Four example Z-plane slices identified from height maps according to some embodiments of this disclosure are shown.

[0024] Figures 18A-18B Two sets of connection components for two different Z-plane slices are shown according to some embodiments of the present disclosure.

[0025] Figures 19A-19B The tops of two candidate boxes identified in the connected components are shown according to some embodiments of this disclosure.

[0026] Figure 20 A set of points corresponding to the connection component identified as the top of the box is shown according to some embodiments of this disclosure.

[0027] Figure 21 This is an example outline of the top of a box projected along the x-axis and y-axis according to some embodiments of this disclosure.

[0028] Figure 22 Some embodiments based on this disclosure are shown. Figure 21 The outline of the box is rotated to align with the top of the box.

[0029] Figure 23A and 23B The top width and length outline of an example box according to some embodiments of this disclosure are shown.

[0030] Figure 24 The image shows the edges of a box that are overlaid on an image obtained by a TOF sensor, according to some embodiments of the present disclosure.

[0031] Figure 25 The illustration shows identified box edges and determined box dimensions superimposed on an image obtained by a camera, according to some embodiments of the present disclosure. Detailed Implementation

[0032] Overview

[0033] Each of the systems, methods, and apparatuses disclosed herein has several innovative aspects, none of which alone is responsible for all the desired properties disclosed herein. Details of one or more implementations of the subject matter described herein are set forth in the following description and figures.

[0034] Reliable identification of the Z-plane in an environment (e.g., floor, street, tabletop) is useful in many 2D and 3D image processing applications. In particular, determining the roll and pitch angles of a time-of-flight sensor relative to the Z-plane, as well as the sensor's height relative to the Z-plane within its environment, is useful. As used herein, the Z-plane is a plane in a real-world environment that is parallel to the ground in a given environment. The Z-plane includes the ground or floor, and curved surfaces parallel to the ground or floor. In many cases, the Z-plane is perpendicular to the direction of gravity. In some cases, such as on hills or other sloping surfaces, the Z-plane (e.g., the ground, a table placed on the ground) may be slightly tilted relative to the direction of gravity.

[0035] The fundamental Z-plane is the lowest Z-plane in an image that captures the environment. For example, in an image of an environment that includes a box on a table placed on the floor, the top of the box, the tabletop, and the floor are all Z-planes, while the floor is the fundamental Z-plane. If another image includes the box and the tabletop but not the floor, then the tabletop is the fundamental Z-plane of that image.

[0036] This paper describes methods and systems for identifying the Z-plane in an environment and, in some cases, the fundamental Z-plane in the environment. The method includes extracting parameters of roll and pitch rotation angles relative to the Z-plane. In some embodiments, the method also extracts parameters of the sensor's height relative to a reference Z-plane from a single input TOF depth frame. Once these two rotation and translation parameters are extracted, the number of prior-unknown external camera calibration parameters is reduced from six (3 translations + 3 rotation angles) to three (2 translations + 1 rotation angle). When the number of unknown sensor degrees of freedom is reduced in this way, time-of-flight applications become easier and faster for the processing system. Furthermore, aligning the coordinate system axes with the Z-plane simplifies the use of time-of-flight images in a variety of applications, such as box sizing, object sizing, box packing, or obstacle detection.

[0037] This paper also describes methods and systems for measuring the dimensions of a box. One method involves receiving a 3D point cloud obtained from time-of-flight data and identifying a box within the point cloud. Specifically, the method includes identifying the top of the box in the point cloud and then identifying the surface on which the box rests, such as a tabletop or floor. The method then includes calculating the height of the box as the distance between the top of the box and the surface on which the box rests, and identifying the edge of the top of the box. The method then includes calculating the width and length profiles of the edges and determining the width and length of the box based on the width and length profiles. Quantitative height, width, and length values, measured, for example in centimeters, can be reported to a user, for example, on the display of a TOF measurement device. In some examples, the device also generates a visualization of the identified box superimposed on an image of the box, allowing the user to qualitatively confirm the calculated dimensions.

[0038] Existing box-size determination solutions are often highly susceptible to sunlight, which introduces significant noise into TOF or image data. Previous box-size systems were only suitable for indoor use or under specific lighting conditions. In some embodiments described herein, TOF measurement data is filtered to reduce the impact of visual noise, enabling TOF sensor systems to be used under a wide range of ambient lighting conditions, including indoors and outdoors. In one example, the measurement data is filtered in a first stage to identify the Z-plane in the observation environment. Because the Z-plane is relatively large, aggressive filters (e.g., large filter windows) can be used. As mentioned above, after identifying the box using the Z-plane, the box edges are identified. Since higher accuracy is required at this stage, finer filters (e.g., smaller filter windows) can be used to filter the measurement data to find the box edges.

[0039] One embodiment provides a method for identifying a Z-plane. The method includes: receiving distance data describing distances between a sensor capturing the distance data and a plurality of surfaces in the sensor's environment, wherein at least one of the surfaces is a Z-plane; generating a point cloud based on the distance data, the point cloud being in a reference frame of the sensor; identifying a basis vector representing a peak direction across the point cloud; transforming the point cloud into the reference frame containing the basis vector; and identifying a Z-plane in the transformed point cloud.

[0040] Another embodiment provides an imaging system including a Time-of-Flight (TOF) depth sensor and a processor. The TOF depth sensor acquires distance data describing the distances between the TOF depth sensor and multiple surfaces in its environment. The processor receives the distance data from the TOF depth sensor; generates a point cloud based on the distance data, the point cloud being in a reference frame of the TOF depth sensor; identifies a basis vector representing a peak direction across the point cloud; transforms the point cloud into a reference frame containing the basis vector; and identifies a Z-plane in the transformed point cloud.

[0041] Another embodiment provides a method for determining the dimensions of a physical box. The method includes: receiving distance data describing distances between a sensor and a plurality of surfaces in the sensor's environment, at least a portion of the surfaces corresponding to the box to be measured; converting the distance data into a reference frame of one of the surfaces in the sensor's environment; selecting from the plurality of surfaces in the sensor's environment a first surface corresponding to the top of the box and a second surface corresponding to a surface on which the box is placed; calculating a height between the first surface and the second surface; and calculating a length and a width based on the selected first surface corresponding to the top of the box.

[0042] Another embodiment provides an imaging system including a Time-of-Flight (TOF) depth sensor and a processor. The TOF depth sensor acquires distance data describing the distances between the TOF depth sensor and a plurality of surfaces in the environment in which the TOF depth sensor is located. The processor receives the distance data from the TOF depth sensor; converts the distance data into a reference frame of one of the surfaces in the sensor's environment; selects a first surface corresponding to the top of a box and a second surface corresponding to a surface on which the box is placed; calculates the height between the first surface and the second surface; and calculates the length and width based on the selected first surface corresponding to the top of the box.

[0043] As those skilled in the art will understand, aspects of this disclosure, particularly those described herein regarding TOF image-based Z-plane recognition and box size determination, can be embodied in various ways (e.g., as a method, system, computer program product, or computer-readable storage medium). Therefore, aspects of this disclosure can take the form of hardware embodiments, software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, which may generally be referred to herein as “circuit,” “module,” or “system.” The functionality described in this disclosure can be implemented as an algorithm executed by one or more hardware processing units (e.g., one or fewer microprocessors) of one or more computers. In various embodiments, different steps and portions of each method described herein can be executed by different processing units. Furthermore, aspects of this disclosure can take the form of a computer program product embodied in one or more computer-readable media, preferably non-transitory, on which computer-readable program code is embodied (e.g., stored). In various embodiments, for example, such a computer program can be downloaded (updated) to existing devices and systems (e.g., to existing sensing system devices and / or their controllers, etc.) or stored during the manufacture of these devices and systems.

[0044] The following detailed description presents various descriptions of specific embodiments. However, the innovations described herein can be embodied in many different ways, for example, as defined and covered by the claims and / or selected examples. In the following description, reference is made to the accompanying drawings, wherein similar reference numerals may indicate the same or functionally similar elements. It should be understood that the elements shown in the drawings are not necessarily drawn to scale. Furthermore, it will be understood that some embodiments may include more elements and / or a subset of the elements shown in the figures than are shown in the figures. In addition, some embodiments may combine any suitable combination of features from two or more figures.

[0045] The following disclosure describes various illustrative embodiments and examples for implementing the features and functions of this disclosure. While specific components, arrangements, and / or features are described below in conjunction with various exemplary embodiments, these are merely examples for simplifying this disclosure and are not intended to be limiting. It should be understood, of course, that in the development of any actual embodiment, many implementation-specific decisions must be made to achieve the developer's specific objectives, including compliance with system, business, and / or legal constraints, which may vary from implementation to implementation. Furthermore, it should be appreciated that while such development work may be complex and time-consuming, it will be a routine task for those skilled in the art who benefit from this disclosure.

[0046] In this specification, reference may be made to the spatial relationships between the various components and the spatial orientation of various aspects of the components shown in the accompanying drawings. However, as those skilled in the art will recognize upon fully reading this disclosure, the devices, components, parts, apparatuses, etc., described herein can be positioned in any desired orientation. Therefore, the use of terms such as “above,” “below,” “upper,” “lower,” “top,” “bottom,” or other similar terms to describe the spatial relationships between various components or to describe the spatial orientation of various aspects of these components should be understood as describing the relative relationships between these components or the spatial orientation of various aspects of these components, since the components described herein can be oriented in any desired orientation. When used to describe the dimensional range or other characteristics (e.g., time, pressure, temperature, length, width, etc.) of elements, operations, and / or conditions, the phrase “between X and Y” indicates a range including both X and Y.

[0047] Other features and advantages of this disclosure will be apparent from the following description and claims.

[0048] Overview of TOF Systems

[0049] Figure 1 This is a block diagram of an example sensor system 100 according to some embodiments of the present disclosure. The sensor system 100 includes a Time-of-Flight (TOF) sensor 110, a processor 120, a camera 130, a display device 140, and a memory 150. In alternative configurations, the TOF sensor system may include... Figure 1 The components shown are different, fewer, and / or additional components. Furthermore, in combination... Figure 1The functions described by one or more components shown may be distributed among the components in a different manner than described. In some embodiments, some or all of components 110-150 may be integrated into a single unit, such as a handheld unit with TOF sensor 110, a processor 120 for processing TOF data, local memory 150, and a display device 140 for displaying the output of processor 120 to a user. In some embodiments, some components may be located in different devices; for example, the handheld TOF sensor 110 may transmit TOF data to an external processing system (e.g., a computer or tablet) that stores and processes the TOF data and provides it to one or more displays for the user. The different devices may communicate via wireless or wired connections.

[0050] The TOF sensor 110 collects distance data that describes the distance between the TOF sensor 110 and various surfaces in its environment. The TOF sensor 110 may include a light source, such as a laser, and an image sensor for capturing light reflected from the surfaces. In some embodiments, the TOF sensor 110 emits light pulses and captures multiple image frames at different times to determine the amount of time it takes for the light pulses to travel to the surface and return to the image sensor. In other embodiments, the TOF sensor 110 detects a phase shift in the captured light, and the phase shift indicates the distance between the TOF sensor 110 and the various surfaces. In some embodiments, the TOF sensor 110 may generate and capture light of multiple different frequencies. If the TOF sensor 110 emits and captures light of multiple frequencies, this can help resolve ambiguous distances and help the TOF sensor 110 operate over a wider range of distances. For example, if for a first frequency, a first observed phase may correspond to a surface 0.5 meters, 1.5 meters, or 2.5 meters away, and for a second frequency, a second observed phase may correspond to a surface 0.75 meters, 1.5 meters, or 2.25 meters away, by combining these two observations, the TOF sensor 110 can determine that the surface is 1.5 meters away. Using multiple frequencies can also improve robustness to noise caused by specific frequencies of ambient light, regardless of whether phase shift or pulse return time is used to measure distance. In alternative embodiments, different types of sensors may be used instead of and / or in addition to the TOF sensor 110 to obtain distance data.

[0051] Processor 120 receives distance data from TOF sensor 110 and processes the distance data to identify various features in the environment of the TOF sensor, as described in detail herein, for example, regarding Figure 3-8In some embodiments, the distance data includes the observed distances to various surfaces measured by the TOF sensor 110 using, for example, the phase shift or pulse return time methods described above. In some embodiments, if the TOF sensor 110 measures a phase shift, the distance data received by the processor 120 from the TOF sensor 110 is phase shift data, and the processor 120 calculates the distances to the surfaces based on the phase shift data.

[0052] Camera 130 can capture image frames of the environment. Camera 130 can be a visible light camera that captures images of the environment within its visible range. In other embodiments, camera 130 is an infrared (IR) camera that captures the IR intensity of surfaces in the environment of the sensor system. The fields of view of camera 130 and TOF sensor 110 partially or completely overlap; for example, the field of view of camera 130 may be slightly larger than that of TOF sensor 110. Camera 130 can transmit the captured images to processor 120. In some embodiments, two processors or processing units may be included, for example, a first processing unit for performing the Z-plane recognition and box size determination algorithms described herein, and a second graphics processing unit for receiving images from camera 130 and generating a display based on the images and data from the first processing unit. In some embodiments, image data from camera 130 may be used to determine the level of sunlight in the environment of TOF sensor 110. In alternative embodiments, sensor system 100 may include a separate light sensor for detecting sunlight or other ambient light conditions in the environment of TOF sensor 110.

[0053] Display device 140 provides visual output to the user of sensor system 100. For example, display device 140 may display the box size and / or box volume calculated by processor 120 based on distance data from TOF sensor 110. In some embodiments, display device 140 displays an image acquired by camera 130 and overlays a visual image indicating one or more features identified in the field of view of camera 130 and TOF sensor 110 based on the distance data. For example, processor 120 may instruct display device 140 to display the outline of the box on an image of the box acquired by camera 130. The user can use the display to determine whether sensor system 100 has correctly identified the box and its edges. Sensor system 100 may include additional or alternative input and / or output devices, such as buttons, speakers, touchscreens, etc.

[0054] Memory 150 stores data for sensor system 100. For example, memory 150 stores processing instructions from processor 120 for identifying features in the environment of TOF sensor 110, such as instructions for identifying one or more Z-planes and / or for calculating the box size of an observed box. Memory 150 may temporarily store data and images acquired by camera 130 and / or TOF sensor 110 and accessed by processor 120. Memory 150 may further store image data accessed by display device 140 to generate an output display.

[0055] Figure 2 The diagram illustrates the ray direction of pixels in a TOF sensor 110 according to some embodiments of the present disclosure. Distance data obtained by the TOF sensor 110 can be arranged as a group of pixels, such as pixels 210a and 210b, within an image frame (e.g., image frame 220). Each pixel 210 has an associated ray direction 215, which points outward from the TOF sensor 110. The ray direction 215 is projected toward the image frame 220. Although in Figure 2 The image shows 25 rays and pixels, but it should be understood that the TOF sensor 110 can have more pixels. Although in Figure 2 In the example shown, image frame 220 has a square shape, but in other embodiments, image frame 220 may have other shapes. In some examples, certain pixels, such as those near the edges of image frame 220, may not be considered valid (e.g., not reliable enough) and are removed from the distance data.

[0056] For example, a first pixel 210a has a ray direction 215a extending straight from the TOF sensor 110; pixel 210a is located at the center of image frame 220. A second pixel 210b at a corner of image frame 220 is associated with ray direction 215b, which extends from the TOF sensor 110 at an angle of 30° from the center of image frame 220, for example, in the x and y directions, where image frame 220 is the xy plane in the reference frame of TOF sensor 110. TOF sensor 110 returns distance data (e.g., distance, one or more phase shifts) to the surface along the ray of each valid pixel. In one example, the first pixel 210a may have a measurement distance of 1 meter, representing the distance to a specific point on the box, and the second pixel 210b may have a measurement distance of 2 meters, representing the distance to a specific point on the wall behind the box.

[0057] Example process for identifying the Z-plane

[0058] Figure 3This is a flowchart illustrating a process 300 for identifying the Z-plane in TOF data obtained in an arbitrary reference frame, according to some embodiments of the present disclosure. A TOF sensor 110 captures distance data of an environment 310, including various surfaces in the environment. It can be assumed that at least one of these surfaces is the Z-plane. The TOF sensor 110 transmits the distance data to a processor 120. In some examples, a camera 130 captures an image of the environment, for example, transmitting the image to the processor 120 simultaneously with the TOF sensor 110 capturing the distance data.

[0059] Figure 9 and Figure 10 Two example visual representations of the inputs from camera 130 and TOF sensor 110 are shown. Figure 9 This is an example image illustrating a box placed on a desktop according to some embodiments of the present disclosure. In this example, Figure 9 An IR intensity image is shown, which illustrates a box 910 placed on a table 920, with a chair 930 to the left of the table and a floor 940. The IR intensity image can be used during visualization (e.g., to visualize the location of the extracted world Z-plane, or for visualization of other applications, such as box dimensions).

[0060] Figure 10 Examples of distance data obtained by a TOF sensor according to some embodiments of this disclosure are shown. Figure 10 The field of view of the TOF sensor 110 shown corresponds to Figure 9 The field of view of the IR intensity image is shown. Distance data is represented by shading, where different shadings represent different distances from the TOF sensor 110 to various surfaces in the environment, for example, box 910 is closer to the TOF sensor 110 than floor 940. (See also: Regarding...) Figure 2 The distance data described includes having Figure 2 The various pixels in the ray direction 215 shown.

[0061] In some embodiments, the processor 120 filters the received distance data 320. Ambient light in the environment of the TOF sensor 110 can introduce noise into the distance data. To reduce the impact of noise, a filter (e.g., an integral filter) can be applied to the distance data before performing further analysis. Filtering noise in this manner may be particularly useful if the TOF sensor 110 captures data in an outdoor environment due to noise caused by sunlight. To filter the distance data, the processor 120 can calculate an average pixel value for each pixel based on the pixel values ​​in the region surrounding the pixel. For example, the filtered pixel value for a given pixel could be the average of 11x11 or 21x21 squared pixels centered on the given pixel. In some embodiments, the processor 120 performs filtering on phase measurement data received from the TOF sensor 110, for example, the processor 120 first filters multiple phase measurements (as described above, for different frequencies), and then processes the filtered phase data to determine the distance measurement for each pixel. Alternatively, the processor 120 can filter the distance measurements, for example, if a pulse-return method is used to obtain the distance data.

[0062] In some embodiments, the filtering step may be omitted, for example, if the TOF sensor 110 is intended for use in an environment with a relatively low noise level, such as if the TOF sensor 110 is designed only for indoor use. In some cases, the processor 120 may perform filtering in response to determining that a threshold level of sunlight is present in the environment of the TOF sensor 110. Furthermore, in some embodiments, the processor 120 may perform adaptive filtering based on the type or level of ambient light in the environment of the TOF sensor 110, for example, using a larger filter window when brighter sunlight is detected, using a larger filter window when a larger frequency distribution is detected in the ambient light, or using a larger filter window when a specific frequency known to interact with the TOF sensor 110 is detected in the ambient light.

[0063] Processor 120 generates a 330-point cloud based on distance data and pixel ray direction 215. For example, for each individual pixel, processor 120 multiplies the ray direction 215 of that pixel by the measured distance from that pixel to the surface, e.g. Figure 10 The measured distance is shown. The processor 120 can retrieve the ray direction 215 from the memory 150.

[0064] The point cloud in the reference frame of the TOF sensor 110 is also referred to as the ego frame. For example, if a user holds the TOF sensor 110 at a small angle relative to the ground in the environment, the Z direction in the reference frame of the TOF sensor is not aligned with the Z plane in the environment (e.g., if the Z plane is perpendicular to the direction of gravity, the Z direction in the reference frame is at an angle relative to the direction of gravity).

[0065] Figure 11 Example point clouds calculated from distance data according to some embodiments of this disclosure are shown. Because the TOF sensor 110 is at an arbitrary angle relative to the Z-plane of the environment when capturing distance data, the point clouds generated in the reference frame of the TOF sensor 110 are difficult for humans to interpret. Such point clouds are also challenging for computer algorithms used in various applications that utilize TOF data.

[0066] Processor 120 identifies 340 basis vectors for a reference frame (also called a "world" reference frame) of surfaces in the environment. The first basis vector corresponds to a direction perpendicular to the Z-plane in the environment observed by TOF sensor 110. The second and third basis vectors are each orthogonal to the first basis vector. The basis vectors define the "world" coordinate system, i.e., a coordinate system with the Z-plane horizontal.

[0067] Figure 4 This is a flowchart illustrating an example process for identifying basis vectors based on Time-of-Flight (TOF) distance data according to some embodiments of this disclosure. Based on a 3D point cloud in an ego frame (e.g., ... Figure 11 (As shown in the point cloud), processor 120 calculates the surface normal vectors (also called surface normals) of the points in the point cloud. To calculate the surface normal for a given point, processor 120 can fit a plane to a set of points in a region surrounding a single point, and then processor 120 calculates the surface normal of the fitted plane. For planes (e.g., floors, cuboid surfaces, walls), the surface normals associated with the point cloud are fairly uniform and have some noise variations. As described above, filter 320 can reduce the noise variations in the surface normals. Surface normals can be represented by polar angles and azimuth angles in a polar coordinate system. In other embodiments, a Cartesian coordinate system can be used to represent the surface normal vectors.

[0068] Figure 12A and 12B Some embodiments according to this disclosure are shown. Figure 11 The example angular coordinates of the surface normals of points in the point cloud shown. Specifically, Figure 12A The polar angle of each calculated surface normal is shown, and Figure 12B The azimuth angles of each calculated surface normal are shown. As these figures illustrate, the Z-planes corresponding to the tops of box 910, table 920, and chair 930 have consistent surface normals on their surfaces, with some variation due to noise in the distance data. Furthermore, since each of these objects is flat along the Z-plane, they each have similar surface normals (represented by similar shading on these Z-planes in each of the two images). Conversely, the front of box 910 has orthogonal surface normals relative to the Z-plane, such as... Figure 12B The darker shadow is shown on the front of the middle box 910.

[0069] After calculating the surface normals, processor 120 extracts one or more basis vectors based on the calculated surface normals. To extract the first basis vector, processor 120 can load the coordinates of the surface normals 420, for example, processor 120 loads the polar angle and azimuth angle of each calculated surface normal. The result of this loading is a two-dimensional distribution. This distribution can be visualized using a histogram (e.g., for...). Figure 13 The two-dimensional histogram (which classifies the surface normals by their angular coordinates) visually represents the classification. Figure 13 The two-dimensional histogram shown has a strong peak 1310 (represented by dark shading) corresponding to the Z-plane direction vector in the ego frame of the TOF sensor 110. The processor 120 identifies peak coordinates, such as peak azimuth and peak polar angle, in the two-dimensional distribution of the 430 loading coordinates. The processor 120 defines a first basis vector 440 as a direction vector corresponding to the peak directions (e.g., peak azimuth and peak polar angle) of the surface normals on the point cloud, which is the surface normal of the Z-plane in the environment of the TOF sensor 110.

[0070] After selecting the first basis vector, processor 120 selects the second and third basis vectors. The second and third basis vectors are orthogonal to the first basis vector (i.e., orthogonal to the surface perpendicular to the Z-plane). The second and third basis vectors are also orthogonal to each other. The first, second, and third basis vectors define the world reference frame.

[0071] In some embodiments, processor 120 calculates the projection of the pointing direction of the TOF sensor (e.g., the ray direction 215b extending straight from the TOF sensor 110) onto the Z-plane (e.g., a plane orthogonal to the first basis vector), and processor 120 selects this projection as the second basis vector. Processor 120 selects a vector orthogonal to the first and second basis vectors as the third basis vector; processor 120 may compute the third basis vector as the cross product of the first and second basis vectors. In other embodiments, the second and third basis vectors may be selected in other ways.

[0072] return Figure 3 After identifying the basis vectors, processor 120 transforms the point cloud 350 into a reference frame of basis vectors. For example, each point in the untransformed point cloud can be defined as a vector (e.g., the product of a ray direction and a measured distance, where the ray direction is a vector in the reference frame of the TOF sensor 110, as described above). In the transformed point cloud, each point can be defined as a linear combination of basis vectors. Specifically, processor 120 can define each point as a sum of basis vectors, each basis vector multiplied by a scalar, where processor 120 determines each scalar by calculating the inner product of the points (in vector representation) in the untransformed point cloud with the basis vectors. Figure 14Examples of point clouds transformed into a reference coordinate system of identified basis vectors according to some embodiments of this disclosure are shown. The transformed point cloud is more easily interpreted by humans. Figure 11 The point cloud shown is illustrated. Box 910, table 920, chair 930, and floor 940 are visible in the transformed point cloud. Furthermore, for processor 120, the transformed point cloud is easier to use for further computations than the untransformed point cloud in the ego frame (e.g., identifying the Z-plane, identifying boxes, and determining box dimensions, as described further below). For example, aligning the point cloud with a reference frame for the Z-plane simplifies the steps of identifying and isolating the Z-plane.

[0073] The processor 120 then identifies the Z-plane in the point cloud with a 360-degree transformation. Figure 5 This is a flowchart illustrating an example process for identifying the Z-plane in a point cloud based on a transformation of the identified basis vectors, according to some embodiments of the present disclosure. First, processor 120 generates a height map from the transformed point cloud. For example, processor 120 distributes the points of the transformed point cloud into square "chimneys" and then selects a representative height for each chimney. Each chimney can be a shape of the same size, for example, 2 square meters (0.75 cm). Other sizes or shapes can be used to construct the height map. The representative height can be, for example, the top point of the chimney (maximum height), the average height, the middle height, or another height selected or calculated from the height of points falling within the chimney. Reducing a three-dimensional point cloud to a two-dimensional height map simplifies data processing and improves computational speed.

[0074] Figure 15 These are example visual representations of height maps obtained from transformed point clouds according to some embodiments of the present disclosure. Figure 15 The shadows in the diagram represent the height of each chimney. For example, the brighter shadow at the top of box 910 indicates a greater height than the darker shadow of table 920.

[0075] Processor 120 then generates a contour representation of the 520 height map. For example, processor 120 integrates the height maps in the x and y directions to obtain the Z contour of the height map. This contour represents the probability density of height in the height map. Peaks within the contour correspond to various planes, i.e., surfaces orthogonal to the first basis vector.

[0076] Figure 16 This is an example Z-profile of a height map having peak values ​​indicating various horizontal surfaces, according to some embodiments of this disclosure. Figure 16This includes four peaks: 1610, 1620, 1630, and 1640. In this example, peak 1610, at the lowest height, corresponds to floor 940. The next peak, 1620, corresponds to chair 930. The next peak, 1630, which is also the highest peak (indicating that most points within the heightmap fall on this peak), corresponds to table 920. Finally, the last peak, 1640, corresponds to the top of box 910.

[0077] Processor 120 identifies peaks in the contour representation of the heightmap as Z-planes. For example, processor 120 identifies a Z-plane for each portion of the Z-contour where the probability density drops above a given threshold (e.g., 0.01). In some embodiments, processor 120 applies one or more other rules or heuristics to the Z-contour to identify Z-planes, for example, by removing spurious noise peaks while preserving genuine weak signals, such as the bottom peaks shown on page 9. For example, processor 120 may identify peaks with at least a threshold number of associated points, or peaks whose heights fall within a given range of each other.

[0078] Processor 120 can select a specific height point (e.g., the highest point or the center point) within a given peak as the Z-plane height of that peak. In some embodiments, processor 120 selects the lowest Z-value peak as the height of the base Z-plane. Processor 120 can set the height of the base Z-plane to zero and determine the heights of other Z-planes based on the heights of other Z-planes relative to the base Z-plane. For example, processor 120 sets the height of floor peak 1610 to 0, and processor 120 determines that chair peak 1620 has a height of 0.569 meters, table peak 1630 has a height of 0.815 meters, and box peak 1640 has a height of 0.871 meters.

[0079] After identifying the Z-plane and its associated height, processor 120 associates individual points in the point cloud with the identified Z-plane. For example, if a point's height falls within the height range of the identified Z-plane, processor 120 can associate a specific point in the point cloud with the identified Z-plane. For instance, if a specific point's height is within twice the peak FWHM (full width at half the maximum value), the point is associated with the Z-plane corresponding to that peak. In other examples, additional ranges around the Z-plane height are used to associate points in the point cloud with the Z-plane.

[0080] Figure 17 Four example Z-slices from a point cloud identified from a height map according to some embodiments of the present disclosure are shown. Each set of points associated with a particular Z-plane by the processor 120 may be referred to as a Z-plane slice, or simply a Z-slice. Each Z-slice is represented with a different coloring, wherein different colored Z-slices correspond to Z-planes at different heights.

[0081] The transformed point cloud and the identified Z-plane can be used for various further processing of the point cloud data. In some examples, processor 120 can continue to locate boxes in the environment of TOF sensor 110 and determine the size and / or volume of the boxes, as described below. In other examples, processor 120 can perform other types of recognition or analysis on other types of objects in the environment of TOF sensor 110.

[0082] In some embodiments, sensor system 100 displays the output of the Z-plane recognition process to a user. For example, processor 120 may associate the recognized Z-plane with various pixels in an image acquired by camera 130 and generate a display with a visual indication of the recognized Z-plane. For example, the Z-plane may be contour- or color-coded in a display output by display device 140. Display device 140 may alternatively or additionally output the height of the determined Z-plane.

[0083] Example process for determining box size

[0084] Figure 6 This is a flowchart illustrating a process 600 for determining and outputting box dimensions based on TOF data according to some embodiments of the present disclosure. A sensor system, such as sensor system 100, receives 610 distance data describing various surfaces in the environment of sensor system 100. For example, as per... Figure 3 As described in steps 310-330, the TOF sensor 110 captures distance data, the processor 120 optionally filters the distance data, and the processor 120 generates a point cloud based on the distance data.

[0085] For the box size determination process, processor 120 may assume that at least a portion of the surface in the environment surrounding TOF sensor 110 corresponds to the box to be measured. Several additional assumptions may be made about the box being measured. Such assumptions can improve the speed and accuracy of the box size determination process, particularly in applications where rapid detection and measurement of boxes is important, such as when the box size is calculated and provided to the user in real time or near real time as the user points TOF sensor 110 at the box. These assumptions may include: the angle between adjacent surfaces of the box is reasonably close to 90° (e.g., between 85° and 95°, or in some other range); the box is within a specific distance of TOF sensor 110 (e.g., within 3 meters or 5 meters); each box size is within a specific range (e.g., at least 3 cm, or at least 10 cm; no more than 2 meters, or no more than 3 meters); the box is closed; the top surface of the box is visible to TOF sensor 110; and the box is placed on a flat, horizontal surface (i.e., the Z-plane) that is also visible to TOF sensor 110.

[0086] It should be understood that in some embodiments, one or more of these assumptions may be relaxed or removed. For example, the range between the sensor and the box, as well as the minimum and maximum box sizes, are merely exemplary, and in other embodiments, different ranges and sizes may be used. In some embodiments, the range may vary depending on the intended use of the sensor system 100 or the target user. For example, if the sensor system 100 is used to measure boxes loaded into a mobile truck (e.g., wardrobe black boxes and boxed furniture), a larger distance range and a larger box size may be used. In some embodiments, the user may be able to input the distance range and / or the maximum and minimum box sizes.

[0087] Processor 120 transforms distance data (e.g., a point cloud calculated based on distance data from TOF sensor 110) 620 into a reference frame of the surface in the environment of the TOF sensor. For example, as per [reference to...] Figure 3 As described in steps 340 and 350, processor 120 identifies the basis vectors of a reference frame of a surface in the environment of the TOF sensor 110, and processor 120 converts distance data (e.g., point cloud) into a reference frame of basis vectors. Regarding Figure 4 The process of converting point clouds into reference frames for basis vectors is described in more detail. Processor 120 can further identify the Z-plane in the transformed distance data, such as regarding... Figure 3 Step 360 and about Figure 5 A more detailed description.

[0088] Such as about Figure 5 and Figure 17 As indicated, each Z-plane can be represented as a Z-slice of heightmap data. Processor 120 selects 630 the surface corresponding to the top of the box and the surface corresponding to the bottom of the box based on the heightmap data. For example, processor 120 identifies one Z-slice as containing the top of the box and another Z-slice as the surface on which the box is placed (e.g., a floor or table) corresponding to the bottom of the box.

[0089] Figure 7 This is a flowchart illustrating an example process for identifying the top and bottom of a box according to some embodiments of the present disclosure. Processor 120 generates a 710 height map based on distance data. For example, processor 120 can generate a height map such as... Figure 5 The height map described in step 510. Processor 120 then identifies 720Z slices in the height map. For example, processor 120 generates a contour representation of the height map, identifies the Z-plane as a peak in the profile representation of the height map, and associates individual points in the distance data (e.g., points in the point cloud) with the Z-plane, as per the context... Figure 5 As described in steps 520-540. As mentioned above, each set of points associated with a particular Z-plane is called a Z-slice.

[0090] Processor 120 identifies connected components within at least some Z-slices. Each of the connected components can be the top of a candidate box. To identify connected components, processor 120 finds clusters of nearby or connected points within the Z-slice. Processor 120 can identify connected components by finding a set of pixels within the Z-slice that can be reached by moving along that Z-slice, for example, pixels that are within a threshold distance of each other. For example, processor 120 can select a specific pixel in the Z-slice and recursively add neighboring pixels that are also in the Z-slice to the connected component. Each connected component has its own height along the Z-axis in the reference frame of the basis vectors; this height corresponds to the height of the Z-slice.

[0091] Figures 18A-18B Two sets of connection components for two different Z-plane slices are shown according to some embodiments of this disclosure. In particular, Figure 18A The connectivity components within an 81.5 cm Z-section are shown, while Figure 18B The connectivity components in the 87.1 cm Z-slice are shown. Each corresponding connectivity component is assigned a different color. If a box is present in the distance data obtained by the TOF sensor 110, one of the connectivity components is expected to correspond to the top of the box.

[0092] In some embodiments, before identifying the connected components in the Z-slices, processor 120 may apply one or more rules to remove one or more Z-slices from consideration as the top of the box. For example, processor 120 may eliminate the lowest or fundamental Z-slice as potentially containing the top of the box, since it is assumed that the top of the box is above the lowest surface. Processor 120 may also remove Z-slices that are not sufficiently close to some other lower Z-slice (i.e., the potential surface for placing the box) within the heightmap (in both the x and y directions). For example, for Figure 17 As shown in the Z-plane slice diagram, processor 120 can remove the slice with Z = 56.9 cm from consideration because it is not sufficiently close to the only Z-slice below it, namely the slice with Z = 0 cm. In contrast, the 56.9 cm Z-slice is close enough to the Z-slice with Z = 81.5 cm that the Z-slice with Z = 81.5 cm cannot be excluded as potentially containing the top of the box, where the 56.9 cm Z-slice is the surface that holds the bottom of the box.

[0093] After identifying the connected components representing the top of the candidate box, processor 120 selects one of the 740 connected components as the box top. Processor 120 can apply various rules to the connected components to identify the box top. For example, processor 120 can remove very small connected components (e.g., those with a width and / or length below the minimum box size threshold mentioned above). Processor 120 can remove highly elongated or non-compact connected components (e.g., those with a large perimeter compared to the square root of their area). Processor 120 can remove connected components whose bottom (i.e., the surface on which the box sits) cannot be derived from the height map, for example, because there are no other connected components in the x and y directions or the Z slice is close enough to the connected components in the height map.

[0094] When applying these rules Figure 18A and 18B Following the connection components shown, Figure 19A and 19B The two connected components shown in the image are retained as the top of the candidate box. Figure 19A and 19B The tops of the two candidate boxes identified in the connected component are shown. Specifically, Figure 19A The connection component 1910a in the slice at Z = 81.5 cm is shown, and Figure 19B The connecting component 1910b in the slice with Z = 87.1 cm is shown. To select one of the remaining connecting components as the top of the box, processor 120 applies an additional rule that considers the shape of the convex hull polygon surrounding the connecting component, for example, how well the convex hull polygon matches the desired rectangular shape. Convex hull polygons 1920a and 1920b are... Figure 19A and 19B Each connected component in the graph is drawn. Figure 19A The convex hull polygon 1920a in the figure deviates strongly from the rectangular shape, while Figure 19B The convex hull polygon 1920b in the image is very close to a rectangle; for example, the convex hull polygon 1920b deviates from the expected rectangular shape by less than a threshold deviation. Therefore, in this example, the processor 120 selects the rectangular connected component in the 87.1cm Z-slice as the top of the box.

[0095] While several example rules for identifying box tops have been discussed above, in different embodiments, processor 120 may apply additional, fewer, or different rules to the connection component to identify box tops. In some embodiments, if multiple candidates pass each of the rules described above, processor 120 may use one or more additional rules to select among possible box tops. For example, processor 120 may select the candidate box top closest to TOF sensor 110.

[0096] After identifying the top of the box, processor 120 identifies the surface on which the box rests, corresponding to the bottom of the box. For example, processor 120 selects a Z-slice in the height map that has a lower height than the top of the box and is closest to the top of the box in the lateral direction; for example, the Z-slice closest to the identified top of the box in the x and y directions of the height map. Figure 17 In the example height diagram shown, the top of the identified box is located on a slice with Z = 81.5 cm. This also corresponds to the height of the bottom of the box.

[0097] return Figure 6 After identifying the top and bottom of the box, processor 120 calculates the box height from the top to the bottom. The box height is the difference between the heights of the Z-slices at the top and bottom of the box, for example, 87.1cm – 81.5cm = 5.6cm.

[0098] Processor 120 further calculates the length and width of the box based on the selected box top. For example, after identifying the box top, processor 120 calculates the length and width of the box top. Because the distance data of the box top is acquired at an angle and may be noisy, processor 120 may filter the box top data, rotate the box top to align it with the x and y axes, and calculate the horizontal and vertical contours of the edges to determine the length and width. For example, the trailing edge of the box (i.e., the edge of the box farthest from the TOF sensor 110) may be blurry, which may make it difficult for processor 120 to identify the trailing edge without performing additional data processing.

[0099] Figure 8 This is a flowchart illustrating an example process for calculating the length and width of the top of a box according to some embodiments of this disclosure. In some embodiments, processor 120 filters the box top data, for example, at least the distance data corresponding to the point on the identified box top in the distance data. In some embodiments, processor 120 filters all the distance data. (See also: Regarding...) Figure 3 As described, ambient light in the environment of the TOF sensor 110 can introduce noise into the distance data. To reduce the impact of noise, a filter, such as an integral filter, can be applied to the distance data before calculating the length and width of the box. Filtering noise in this way may be particularly useful if the TOF sensor 110 captures data in an outdoor environment due to noise caused by sunlight.

[0100] To filter the distance data, processor 120 can calculate an average pixel value for each pixel in the distance data, based on the pixel values ​​in the region surrounding the pixel. The processor can use a filter different from the filter described with respect to step 320. Specifically, processor 120 can use a smaller filter window than the filter used to identify the Z-plane. For example, the filtered pixel value for a given pixel could be the average of 5x5 or 7x7 square pixels centered on the given pixel. (See regarding...) Figure 3 As described, processor 120 can perform filtering on phase measurement data received from TOF sensor 110. For example, processor 120 first filters multiple phase measurements (for different frequencies, as described above), and then processes the filtered phase data to determine the distance measurement for each pixel. Alternatively, processor 120 can filter the distance measurements, for example, if a pulse return method is used to obtain the distance data.

[0101] In some embodiments, the filtering step may be omitted, for example, if the TOF sensor 110 is intended for use in an environment with a relatively low noise level, such as if the TOF sensor 110 is designed only for indoor use. In some cases, the processor 120 may perform filtering in response to determining that a threshold level of sunlight is present in the environment of the TOF sensor 110. Furthermore, in some embodiments, the processor 120 may perform adaptive filtering based on the type or level of ambient light in the environment of the TOF sensor 110, for example, using a larger filter window when brighter sunlight is detected, using a larger filter window when a larger frequency distribution is detected in the ambient light, or using a larger filter window when a specific frequency known to interact with the TOF sensor 110 is detected in the ambient light.

[0102] In step 740, processor 120 extracts a subset of points corresponding to the top of the box within the distance data transformed by 820, for example, the connection component selected as the top of the box. If filtering 810 is performed, processor 820 can calculate a second point cloud based on the filtered data (following...). Figure 3 The process described in step 330 of the process transforms the second point cloud (following the steps regarding...). Figure 3 Steps 340-350 and about Figure 4 The processor 120 performs the described processing and extracts points corresponding to the second point cloud and connected components. The processor 120 can use the same basis vectors selected during the Z-plane recognition phase to transform the point cloud based on the newly filtered data.

[0103] Figure 20A set of points corresponding to the connected component identified as the top of the box is shown according to some embodiments of this disclosure. This set of extracted points is also referred to as the box-top subcloud. To simplify the processing and understanding of the box-top, processor 120 can determine the rotation angle of the extracted box-top subcluster and rotate the box-top subcluster by 830 degrees of that rotation angle, such that the edge of the box-top is aligned with the x-axis and y-axis. For example, processor 120 projects the points in the box-top subcloud onto the x-axis and y-axis as a function of the rotation angle of the subcloud about its center. Processor 120 then calculates the sum of the x and y projections as a function of the rotation angle to generate a contour. Figure 21 This is an example outline of a box-top subcloud projected along the x and y axes according to some embodiments of this disclosure. Processor 120 identifies an azimuth rotation angle that minimizes the sum of the box-top projections. More specifically, the selected rotation angle minimizes the sum of the projections of the box-top edges onto a set of axes of a previously determined reference frame (e.g., a reference frame for the basis vectors). Processor 120 rotates the box-top subcloud through the identified azimuth angle such that the box-top subcloud is axis-aligned. Figure 22 Some embodiments based on this disclosure are shown. Figure 21 The axis of rotation of the outline in the middle is aligned with the top of the box.

[0104] After rotating the top sub-cluster of the boxes, processor 120 calculates the width and length profiles of the top of the 840 boxes. For many TOF sensors, while the leading edge of the box closest to the sensor is sharp and easily identifiable by both humans and computers, the trailing edge of the box located downstream of the TOF sensor 110 can be blurred. This results in the trailing edge being more obscured and difficult to identify from distance data. Processor 120 generates the width and length profiles of the top of the boxes by projecting points from the rotated top sub-cloud of boxes onto the horizontal and vertical axes. Figure 23A and 23B The width and length outlines of the top portion of an example box according to some embodiments of the present disclosure are shown respectively.

[0105] Processor 120 identifies leading and trailing edges in a contour with a width and length of 850. For example, processor 120 applies one or more rules to identify edges from the contour. Processor 120 may fit a line to the interior of each of the contours and define a leading edge as the location where the contour equals a set percentage of the linear fit value (e.g., 40% of the linear fit value). Processor 120 may define trailing edges by the same or different percentages of the contour equal to the linear fit value, for example, a specific value in the range of 25% to 85%. Processor 120 may further apply one or more rules to determine the trailing edge fraction. For example, the trailing edge percentage threshold may vary with the height of the box. Alternatively, the percentage threshold may be different for the shorter and longer of the top edges of two boxes. Figure 23A and 23B An example of trailing and leading edges based on width and length contour recognition is shown.

[0106] Processor 120 calculates the width and length of the top of box 860 based on the determined leading and trailing edges. Specifically, the width is the distance between the leading and trailing edges in the width projection, and the length is the distance between the leading and trailing edges in the length projection. Figure 23A and 23B These represent the width and length between the trailing and leading edges, respectively.

[0107] return Figure 6 The sensor system 100 (e.g., processor 120 and display device 140) can display the determined box dimensions to a user. For example, processor 120 can generate a display for output on display device 140, which includes a visual representation of the box as well as its height, width, and length. For example, the display can show identified edges and / or dimensions projected onto an image captured by camera 130, or an image created based on distance data from TOF sensor 110. In one embodiment, processor 120 generates an image of a three-dimensional box defined by identified leading and trailing edges, identified top and bottom surfaces of the box, and / or calculated height, width, and length. Processor 120 can also calculate the volume (length x width x height) and output the volume on display device 140.

[0108] Processor 120 can project an image of the 3D box onto the 2D image plane of camera 130 to generate an overlay image, such as an overlay image of the box's outline. The calculated width, length, and height dimensions can also be reported in a graphical display along the edges or in separate areas. A user can view the graphical display on display device 140 to qualitatively confirm that sensor system 100 has correctly identified the box and correctly identified the edges and surfaces.

[0109] Figure 24The image shows the edge of a box superimposed on an image obtained based on data from a TOF sensor, according to some embodiments of the present disclosure. Figure 24 The image in the image can be generated by the processor 120 based on a point cloud generated from distance data from the TOF sensor 110. Figure 24 It also includes the outline of the box edges superimposed on the point cloud image.

[0110] Figure 25 The illustration shows identified box edges and determined frame dimensions superimposed on an image obtained by a camera, according to some embodiments of this disclosure. In this example, the image may be an IR image obtained by an IR camera. Figure 24 It also includes the outline of the box's edge superimposed on the IR image, and displays the calculated width, length, and height of the box in the upper left corner of the display. Figure 24 The union and intersection (IoU) scores are also displayed. In some embodiments, the processor 120 calculates an IoU score that measures the overlap between the top of the box and a predefined circle 2510 appearing at the center of the field of view of the TOF sensor 110. A larger IoU score is associated with a higher accuracy in sizing, and the user can adjust the view of the TOF sensor 110 to obtain a higher IoU score. In some embodiments, the sensor system 100 may set a lower limit for IoU for reporting box dimensions; for example, if the IoU score is greater than 0.40 or another threshold, the processor 120 displays the box dimensions, and if the IoU score is below the threshold, a request to move the TOF sensor 110 is displayed to the user. Ensuring that the user orients the TOF sensor 110 relative to the box with a sufficiently high IoU can reduce errors in the box dimensions reported by the sensor system 100.

[0111] In some embodiments, the sensor system 100 may additionally or alternatively report an intensity indicator indicating the intensity measured at a specific pixel or across a set of pixels in the distance data collected by the TOF sensor 110 and / or the intensity measured at the corresponding pixel or across a set of pixels collected by the camera 130. In some cases, if the intensity measured in the region of interest in image frame 220 is too low, the processor 120 may have difficulty finding the Z-plane, determining the top size of the box, or performing other processing of the TOF distance data. The processor 120 may analyze the intensity of at least a portion of the sensor system's field of view and report the intensity to the user. Based on the reported intensity, the user can determine whether to adjust the environment, for example, by changing lighting conditions, by changing the angle of the TOF sensor 110 relative to the box or other region of interest, by moving the box to a different location (e.g., on a different Z-plane, into another room), etc., to increase the intensity. In some embodiments, if the processor 120 determines that the intensity is too low (e.g., the intensity is below a given threshold and / or the processor 120 has difficulty finding the Z-plane or the box, e.g., none of the identified connected components satisfy the rules for identifying the top of the box), the processor 120 may output instructions to the user to change the environment, sensor position, or box position to increase the intensity.

[0112] For example, if camera 130 is an IR camera, processor 120 can determine the IR intensity of at least a portion of the camera's field of view, for example, at or near the center of the image frame of camera 130. If camera 130 is a visible light camera, processor 120 can determine the intensity or brightness of visible light at or near the center of the image frame. The intensity measurement can be correlated with the reflectivity of a material in a given area, for example, the reflectivity of the box material. Since users typically point the TOF sensor 110 at the box, the Z-plane, or other regions of interest, and users can be encouraged to include the top of the box in the center of the image frame via IoU (as described above), the center of the image frame of camera 130 typically corresponds to the top of the box, other parts of the box, the Z-plane, or other regions of interest.

[0113] As a specific example, camera 130 captures an image frame having a region corresponding to image frame 220. Processor 120 can identify the intensity near the center of the image frame captured by camera 130, for example, the intensity at the position corresponding to pixel 215a in the center of image frame 220 of TOF sensor 110, or the average intensity of a group of pixels including the center of the image frame. For example, processor 120 can determine the intensity of the region corresponding to the center of the image frame. Figure 25 The average intensity of a group of pixels corresponding to circle 2510 shown in the diagram.

[0114] Select Example

[0115] The following paragraphs provide various examples of the embodiments disclosed herein.

[0116] Example 1 provides a method for identifying a Z-plane, the method comprising: receiving distance data describing distances between a sensor capturing the distance data and a plurality of surfaces in the sensor’s environment, wherein at least one of the surfaces is a Z-plane; generating a point cloud based on the distance data, the point cloud being in a reference frame of the sensor; identifying a basis vector representing a peak direction across the point cloud; transforming the point cloud into the reference frame of the basis vector; and identifying a Z-plane in the transformed point cloud.

[0117] Example 2 provides the method of Example 1, wherein the sensor is a TOF sensor that includes a light source and an image sensor.

[0118] Example 3 provides the method of Example 1, wherein the distance data is arranged in multiple pixels within an image frame of the sensor.

[0119] Example 4 provides the method of Example 3, wherein a single pixel has a distance to one of a plurality of surfaces in the sensor environment, and the single pixel has an associated ray direction describing the direction from the sensor to the surface.

[0120] Example 5 provides the method of Example 4, wherein generating a point cloud involves multiplying the ray direction of a single pixel by the distance to one of a plurality of surfaces of the single pixel.

[0121] Example 6 provides the method of Example 1, wherein the distance data is arranged as a plurality of pixels, and the method further includes filtering the distance data by calculating an average pixel value for a single pixel based on pixel values ​​in the region surrounding the single pixel.

[0122] Example 7 provides the method of Example 1, wherein identifying basis vectors includes calculating surface normals to points in the point cloud; and extracting basis vectors based on the calculated surface normals, the basis vectors representing the peak direction of the surface normals on the point cloud.

[0123] Example 8 provides the method of Example 7, where calculating the surface normal for points in a point cloud includes calculating the angular coordinates of the surface normal for those points in the point cloud.

[0124] Example 9 provides the method of Example 8, wherein extracting the basis vectors includes loading the angular coordinates of the surface normal; identifying the peak angle of each angular coordinate; and identifying the basis vectors based on the identified peak angles.

[0125] Example 10 provides the method of Example 7, wherein calculating the surface normal of a single point in a point cloud includes fitting a plane to a set of points in the region surrounding that single point.

[0126] Example 11 provides the method of Example 1, wherein the basis vector is a first basis vector, and the method further includes selecting a second basis vector orthogonal to the first basis vector and a third basis vector orthogonal to the first basis vector and the second basis vector, wherein the reference frame of the basis vector is a reference frame of the first basis vector, the second basis vector and the third basis vector.

[0127] Example 12 provides the method of Example 11, wherein the second basis vector is selected as the projection of the sensor’s pointing direction onto the Z-plane, and the third basis vector is set to be equal to the cross product of the first basis vector and the second basis vector.

[0128] Example 13 provides the method of Example 1, wherein identifying the Z-plane in a transformed point cloud includes generating a height map of the transformed point cloud; generating a contour representation of the height map, the contour representation having peak values ​​corresponding to each of a plurality of Z-planes; and identifying the Z-plane in the contour representation.

[0129] Example 14 provides the method of Example 13, wherein the identified Z-plane is a basic Z-plane, and the method further includes setting the height of the basic Z-plane to zero.

[0130] Example 15 provides the method of Example 13, further comprising associating points in the transformed point cloud with the identified Z-plane based on determining that the height of the points is within a height range associated with the identified Z-plane.

[0131] Example 16 provides an imaging system comprising: a Time-of-Flight (TOF) depth sensor for acquiring distance data describing distances between the TOF depth sensor and a plurality of surfaces in an environment in which the TOF depth sensor is located; and a processor for receiving the distance data from the TOF depth sensor; generating a point cloud based on the distance data, the point cloud being in a reference frame of the TOF depth sensor; identifying a basis vector representing a peak direction across the point cloud; transforming the point cloud into a reference frame containing the basis vector; and identifying a Z-plane in the transformed point cloud.

[0132] Example 17 provides the system of Example 16, wherein the TOF depth sensor includes a light source for illuminating the environment of the TOF depth sensor and an image sensor for sensing reflected light.

[0133] Example 18 provides the system of Example 16, wherein the TOF depth sensor has an image frame and the distance data is arranged in a plurality of pixels within the image frame.

[0134] Example 19 provides the system of Example 18, wherein a single pixel has a distance to one of a plurality of surfaces in the environment of a TOF depth sensor, and the single pixel has an associated ray direction describing the direction from the TOF depth sensor to the surface.

[0135] Example 20 provides the system of Example 19, wherein, in order to generate the point cloud, the processor multiplies the ray direction of the individual pixel by the distance to one of the plurality of surfaces of the individual pixel.

[0136] Example 21 provides the system of Example 16, which also includes a camera for capturing images of the environment of the TOF depth sensor.

[0137] Example 22 provides the system of Example 21, and also includes a display screen on which the processor displays images captured by the camera and visual indications of the identified Z-plane.

[0138] Example 23 provides the system of Example 16, and further includes a light sensor for detecting sunlight in the environment of the TOF depth sensor, wherein the processor applies a filter to the distance data in response to detecting sunlight at least at a threshold level.

[0139] Example 24 provides a method for determining the size of a physical box, the method comprising: receiving distance data describing distances between a sensor and a plurality of surfaces in the sensor’s environment, at least a portion of the surfaces corresponding to the box to be measured; converting the distance data into a reference frame of one of the surfaces in the sensor’s environment; selecting from the plurality of surfaces in the sensor’s environment a first surface corresponding to the top of the box and a second surface corresponding to a surface on which the box is placed; calculating a height between the first surface and the second surface; and calculating a length and a width based on the selected first surface corresponding to the top of the box.

[0140] Example 25 provides the method of Example 24, where the distance data is a point cloud in a reference frame of the sensor.

[0141] Example 26 provides the method of Example 25, wherein converting distance data into a reference frame of a surface in a sensor environment includes identifying a basis vector representing a peak direction across the point cloud; and transforming the point cloud into the reference frame of the basis vector.

[0142] Example 27 provides the method of Example 26, wherein identifying the basis vector includes calculating the angular coordinates of the surface normals of points in the point cloud; and extracting the basis vectors based on the calculated angular coordinates of the surface normals, the basis vectors representing the peak direction of the surface normals on the point cloud.

[0143] Example 28 provides the method of Example 24, wherein the sensor is a TOF sensor that includes a light source and an image sensor.

[0144] Example 29 provides the method of Example 24, wherein one of the surfaces used as a reference frame for transforming the distance data is a Z-plane.

[0145] Example 30 provides the method of Example 24, wherein selecting a first surface includes: identifying a plurality of connected components within the distance data of the transformation, each connected component having a respective height along the Z-axis in a reference frame of one of the surfaces; and selecting one of the plurality of connected components as the first surface by applying a set of rules to the plurality of connected components.

[0146] Example 31 provides the method of Example 30, wherein identifying the plurality of connected components includes identifying a plurality of Z slices of distance data of the transformation, each of the plurality of Z slices having a corresponding height along the Z-axis; and identifying at least one connected component of a height map pixel within each of the plurality of Z slices.

[0147] Example 32 provides the method of Example 31, wherein identifying the plurality of Z slices includes generating a height map of the distance data; generating a contour representation of the height map, the contour representation having peak values ​​corresponding to each Z slice; and identifying the plurality of Z slices from the contour representation.

[0148] Example 33 provides the method of Example 31, wherein selecting a second surface corresponding to the surface on which the box is placed includes selecting a Z-slice from a plurality of Z-slices within the lateral range of the selected first surface.

[0149] Example 34 provides the method of Example 30, wherein the set of rules applied to the plurality of connected components includes removing connected components whose width or length is less than a threshold minimum width or length; removing connected components from another connected component by at least a threshold distance; and removing connected components having a closed convex hull polygon that deviates from the desired rectangular shape by at least a threshold deviation.

[0150] Example 35 provides the method of Example 24, wherein calculating the length and width based on a selected first surface includes: extracting a subset of distance data of a transformation corresponding to the selected first surface; calculating a length profile and a width profile of the subset; identifying a first leading edge and a first trailing edge of the box within the width profile; identifying a second leading edge and a second trailing edge of the box within the length profile; and calculating the box width between the first leading edge and the second leading edge, and calculating the box length between the second leading edge and the second trailing edge.

[0151] Example 36 provides the method of Example 24, further comprising: determining a rotation angle of an extracted subset corresponding to the first selected surface, the determined angle being selected to minimize the sum of projections of the edges of the first selected surface on a set of axes of a reference frame of one of the surfaces in the environment of the sensor; and rotating the extracted subset corresponding to the first selected surface by the determined angle.

[0152] Example 37 provides the method of Example 24, wherein the transformed distance data comprises a plurality of pixels, and calculating the length and width based on a selected first surface corresponding to the top of the box comprises, for at least pixels in the selected first surface, filtering the pixel by calculating an average pixel value for each individual pixel based on pixel values ​​in the region surrounding the individual pixel; and calculating the length and width based on the filtered pixels in the selected first surface.

[0153] Example 38 provides the method of Example 24, and further includes generating a visual representation of the box, the visual representation indicating the height, width and length of the box.

[0154] Example 39 provides the method of Example 24, further comprising calculating an IoU score based on the overlap between the first surface corresponding to the top of the box and a circle in the field of view of the sensor; and generating a display including the calculated IoU score.

[0155] Example 40 provides the method of Example 24, further comprising receiving camera data from a camera having a camera field of view that at least partially overlaps with the field of view of the sensor; determining the intensity of at least a portion of the camera field of view based on the camera data; and generating a display including the determined intensity.

[0156] Example 41 provides an imaging system comprising: a Time-of-Flight (TOF) depth sensor for obtaining distance data describing distances between the TOF depth sensor and a plurality of surfaces in the environment of the TOF depth sensor; and a processor for receiving the distance data from the TOF depth sensor; converting the distance data into a reference frame of one of the surfaces in the environment of the sensor; selecting a first surface corresponding to the top of a box and a second surface corresponding to a surface on which the box is placed; calculating a height between the first surface and the second surface; and calculating a length and a width based on the selected first surface corresponding to the top of the box.

[0157] Example 42 provides the system of Example 41, wherein the TOF depth sensor includes a light source for illuminating the environment of the depth sensor and an image sensor for sensing reflected light.

[0158] Example 43 provides the system of Example 41, wherein the TOF sensor has an image frame and the distance data is arranged in a plurality of pixels within the image frame.

[0159] Example 44 provides the system of Example 43, wherein a single pixel has a distance from one of a plurality of surfaces in the environment of the TOF depth sensor, and the single pixel has an associated ray direction describing the direction from the sensor to the TOF depth surface.

[0160] Example 45 provides the system of Example 41, and also includes a camera for capturing images of the environment of the TOF depth sensor.

[0161] Example 46 provides the system of Example 45, and further includes a display screen on which the processor displays an image captured by the camera and calculated width, length and height.

[0162] Example 47 provides the system of Example 45, and also includes a display screen on which the processor displays an overlay depiction of an image captured by the camera and a selected first surface.

[0163] Example 48 provides the system of Example 47, wherein the processor further displays a plurality of box edges beneath the selected first surface on the display screen.

[0164] Other implementation notes, changes, and applications

[0165] It should be understood that not all objectives or advantages can be achieved according to any particular embodiment described herein. Therefore, for example, those skilled in the art will recognize that certain embodiments may be configured to operate in a manner that achieves or optimizes one or more advantages taught herein, without necessarily achieving other objectives or advantages taught or suggested herein.

[0166] In one example embodiment, any number of the circuits shown in the figures can be implemented on a board of an associated electronic device. This board can be a general-purpose circuit board that can house various components of the internal electronic system of the electronic device and further provide connectors for other peripheral devices. More specifically, the board can provide electrical connections through which other components of the system can communicate electrically. Any suitable processor (including digital signal processors, microprocessors, supporting chipsets, etc.), computer-readable non-transitory storage elements, etc., can be appropriately coupled to the board based on specific configuration requirements, processing requirements, computer design, etc. Other components, such as external memory, additional sensors, audio / video display controllers, and peripheral devices, can be connected to the board via cables as plug-in cards or integrated into the board itself. In various embodiments, the functionality described herein can be implemented in emulation form as software or firmware running within one or more configurable (e.g., programmable) elements arranged in a structure supporting these functions. The emulated software or firmware can be provided on a non-transitory computer-readable storage medium including instructions that allow the processor to perform these functions.

[0167] It must also be noted that all specifications, dimensions, and relationships (e.g., number of processors, logical operations, etc.) outlined herein are for illustrative and educational purposes only. Such information may be significantly altered without departing from the spirit of this disclosure or the scope of the appended claims. This specification applies only to a non-limiting example and should therefore be interpreted as a non-limiting instance. In the foregoing description, exemplary embodiments have been described with reference to specific arrangements of components. Various modifications and changes may be made to such embodiments without departing from the scope of the appended claims. Therefore, the specification and drawings should be considered illustrative rather than restrictive.

[0168] Note that the interactions can be described using two, three, four, or more components in the numerous examples provided herein. However, this is done merely for clarity and illustration. It should be understood that the system can be combined in any suitable manner. Based on similar design alternatives, any components, modules, and elements shown in the figures can be combined in a wide variety of possible configurations, all of which are clearly within the broad scope of this specification.

[0169] Please note that in this specification, references to various features (e.g., elements, structures, modules, components, steps, operations, characteristics, etc.) included in “one embodiment,” “exemplary embodiment,” “embodiment,” and “another embodiment,” etc., are intended to mean that any such feature is included in one or more embodiments of this disclosure, but may or may not be combined in the same embodiment.

[0170] Those skilled in the art can identify many other changes, substitutions, variations, alterations, and modifications, and this disclosure includes all such changes, substitutions, variations, alterations, and modifications that fall within the scope of the appended claims. Note that all optional features of the systems and methods described above can also be implemented relative to the methods or systems described herein, and the details in the examples can be used anywhere in one or more embodiments.

[0171] In order to assist the United States Patent and Trademark Office (USPTO) and the readers of any patent in this application in interpreting the appended claims, the applicant wishes to draw the attention of the applicant that: (a) the applicant does not intend to invoke section 112 of title 35(f) of the United States Code, which exists as of the date of this agreement, unless the word “component” or “step” is specifically used in a particular claim; and (b) the applicant does not intend to limit this disclosure in any way not reflected in the appended claims by any statement in the specification.

Claims

1. A method for identifying a Z-plane, the method comprising: receiving distance data, the distance data describing distances between a sensor capturing the distance data and a plurality of surfaces in an environment of the sensor, wherein at least one of the plurality of surfaces is a Z-plane, and wherein a Z-plane is a plane in a real-world environment that is parallel to the ground in a particular environment; generating a point cloud based on the distance data, the point cloud being in a reference frame of the sensor; identifying a basis vector representing a peak direction across the point cloud, including binning coordinates of surface normals of points in the point cloud, wherein a peak direction is a peak in a histogram of binned coordinates of surface normals of points in the point cloud; transforming the point cloud to a reference frame of the basis vector; and identifying a Z-plane in the transformed point cloud.

2. The method of claim 1, wherein the sensor is a time-of-flight (TOF) sensor including a light source and an image sensor.

3. The method of claim 1, wherein the distance data is arranged in a plurality of pixels within an image frame of the sensor.

4. The method of claim 3, wherein a single pixel includes a distance to one of the plurality of surfaces in an environment of the sensor, and the single pixel has an associated ray direction describing a direction from the sensor to the surface.

5. The method of claim 4, wherein generating the point cloud includes defining the single pixel as a vector that is a product of the ray direction for the single pixel and a distance to the one of the plurality of surfaces for the single pixel.

6. The method of claim 1, wherein the distance data is arranged as a plurality of pixels, the method further comprising: filtering the distance data by computing an average pixel value for a single pixel based on pixel values in a region around the single pixel.

7. The method of claim 1, wherein identifying the basis vector further comprises: computing the surface normal of points in the point cloud; and extracting the basis vector based on the computed surface normals, the basis vector representing a peak direction of surface normals over the point cloud.

8. The method of claim 7, wherein computing the surface normal of points in the point cloud includes computing angular coordinates of the surface normal of points in the point cloud.

9. The method of claim 8, wherein extracting the basis vector includes: binning the angular coordinates of the surface normal; identifying a peak angle for each angular coordinate; and identifying the basis vector based on the identified peak angles.

10. The method of claim 7, wherein computing the surface normal of a single point in the point cloud includes fitting a plane to a set of points in a region around the single point.

11. The method of claim 1, wherein the basis vector is a first basis vector, the method further comprising: ​ ​ ​ selecting a second basis vector orthogonal to the first basis vector and a third basis vector orthogonal to the first basis vector and the second basis vector, wherein the reference frame of the basis vectors is a reference frame of the first basis vector, the second basis vector, and the third basis vector.

12. The method of claim 11, wherein the second basis vector is selected to be a projection of a pointing direction of the sensor into a Z-plane, and the third basis vector is set equal to a cross product of the first basis vector and the second basis vector.

13. The method of claim 1, wherein identifying a Z-plane in the transformed point cloud comprises: generating a height map of the transformed point cloud; generating a contour representation of the height map, the contour representation having a peak value corresponding to each of a plurality of Z-planes; and identifying a Z-plane in the contour representation.

14. The method of claim 13, wherein the identified Z-plane is a base Z-plane, the method further comprising setting a height of the base Z-plane to zero.

15. The method of claim 13, further comprising associating a point in the transformed point cloud with the identified Z-plane based on determining that a height of the point is within a height range associated with the identified Z-plane.

16. An imaging system comprising: a time-of-flight (TOF) depth sensor to obtain distance data describing distances between the TOF depth sensor and a plurality of surfaces in an environment of the TOF depth sensor, wherein at least one of the plurality of surfaces is a Z-plane, and wherein a Z-plane is a plane in a real-world environment that is parallel to a ground plane in a particular environment; and a processor to: receive the distance data from the TOF depth sensor; generate a point cloud based on the distance data, the point cloud being in a reference frame of the TOF depth sensor; identify a basis vector representing a peak direction across the point cloud, including binning coordinates of surface normals of points in the point cloud, wherein a peak direction is a peak in a histogram of binned coordinates of surface normals of points in the point cloud; transform the point cloud to a reference frame of the basis vector; and identify the Z-plane in the transformed point cloud.

17. The system of claim 16, wherein the TOF depth sensor comprises a light source to illuminate an environment of the TOF depth sensor and an image sensor to sense reflected light.

18. The system of claim 16, wherein the TOF depth sensor has an image frame, and the distance data is arranged in a plurality of pixels within the image frame.

19. The system of claim 18, wherein a single pixel includes a distance to one of the plurality of surfaces in an environment of the TOF depth sensor, and the single pixel has an associated ray direction describing a direction from the TOF depth sensor to the surface.

20. The system of claim 19, wherein to generate the point cloud, the processor defines the single pixel as a vector that is a product of the ray direction for the single pixel and a distance to the one of the plurality of surfaces for the single pixel.

21. The system of claim 16, further comprising a camera to capture an image of an environment of the TOF depth sensor.

22. The system of claim 21, further comprising a display screen, the processor to display on the display screen the image captured by the camera and a visual indication of the identified Z-plane.

23. The system of claim 16, further comprising a light sensor to detect sunlight in an environment of the TOF depth sensor, wherein the processor is to apply a filter to the distance data in response to detecting at least a threshold level of sunlight.

Citation Information

Patent Citations

  • Method of Processing 3D Sensor Data to Provide Terrain Segmentation

    US20150198735A1

  • Method and system for classification of an object in a point cloud data set

    US20190370614A1

  • Control apparatus, control method, program, and moving body

    US20200231176A1