Automatic detection of obstacles by sensors with moving dimensions
By detecting and comparing obstacle depth with a threshold, image distortion caused by obstacles is suppressed, solving the problem of inaccurate dimension annotation caused by obstacles in the field of view of depth sensors, and achieving more accurate object size measurement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZEBRA TECHNOLOGIES CORP
- Filing Date
- 2024-11-19
- Publication Date
- 2026-07-31
AI Technical Summary
Obstacles in or near the depth sensor's field of view can reduce the accuracy of object size capture, leading to inaccurate dimensioning.
By detecting the region of interest in the 3D image captured by the sensor, the depth of the obstacle is determined and compared with a threshold to determine whether to suppress the dimensioning process and avoid image distortion caused by the obstacle.
It improves the accuracy of object dimensioning, avoids dimension measurement errors caused by obstacles, and ensures the precision of dimensioning.
Smart Images

Figure CN122497975A_ABST
Abstract
Description
Background Technology
[0001] Depth sensors, such as time-of-flight (ToF) sensors, can be deployed in mobile devices, such as handheld computers, to capture point clouds of objects, such as boxes or other packaging, from which the object's dimensions can be derived. However, obstacles in or near the depth sensor's field of view can degrade the quality of the capture. Attached Figure Description
[0002] The accompanying drawings (in which the same reference numerals denote the same or functionally similar elements throughout the different views) together with the following detailed description are incorporated into and form part of the specification, and serve to further illustrate embodiments including the concepts of the claimed invention, and to explain the various principles and advantages of those embodiments.
[0003] Figure 1 This is a schematic diagram of a computing device used for dimensioning objects.
[0004] Figure 2 It is shown by Figure 1 A schematic diagram of dimension artifacts introduced by nearby obstacles at the device.
[0005] Figure 3 This is a flowchart of a method for automatic detection of obstacles using sensors with moving dimensions.
[0006] Figure 4 It is shown Figure 3 The diagram illustrates the example execution of the methods in boxes 305 and 310.
[0007] Figure 5 It is shown Figure 3 A schematic diagram illustrating the example execution of the method in box 350.
[0008] Figure 6 It is shown Figure 3 Another example of the execution of the method in box 350 is illustrated.
[0009] Those skilled in the art will understand that the elements in the accompanying drawings are shown for simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some elements in the drawings may be exaggerated relative to other elements to aid in understanding embodiments of the invention.
[0010] The apparatus and method configurations have been indicated in appropriate places in the accompanying drawings by conventional symbols, which show only those specific details relevant to understanding embodiments of the invention, so as not to obscure this disclosure with details that would be obvious to those skilled in the art who benefit from the description herein. Detailed Implementation
[0011] The examples disclosed herein relate to a method comprising: capturing a three-dimensional image corresponding to an object via a sensor of a computing device; detecting a region of interest in the three-dimensional image, the region of interest corresponding to an obstacle; determining a depth from the sensor to the region of interest; comparing the determined depth with a threshold; and determining whether to pass the three-dimensional image to a dimensioning module based on the comparison of the depth of the obstacle with the threshold.
[0012] Additional examples disclosed herein relate to a computing device including: a sensor; and a processor configured to: capture a three-dimensional image corresponding to an object via the sensor; detect a region of interest in the three-dimensional image, the region of interest corresponding to an obstacle; determine a depth from the sensor to the region of interest; compare the determined depth with a threshold; and determine whether to pass the three-dimensional image to a dimensioning module based on the comparison of the depth of the obstacle with the threshold.
[0013] Figure 1 A computing device 100 is illustrated, configured to capture sensor data of a target object 104 within the field of view (FOV) of one or more sensors of the device 100. In the illustrated example, the computing device 100 is a mobile computing device, such as a tablet, smartphone, etc. The computing device 100 can be manipulated by its operator to place the target object 104 within the FOV(s) of the sensors(s), thereby capturing sensor data for subsequent processing as described below. In other examples, the computing device 100 may be implemented as a fixed computing device, for example, installed near an area where the target object 104 is placed and / or transported (e.g., a temporary storage area, conveyor belt, storage container, etc.).
[0014] In this example, object 104 is a package (e.g., a carton, a pallet, etc.), although various other objects can also be processed as described herein. The sensor data captured by device 100 may include three-dimensional (3D) images (e.g., from which the device can generate point clouds). In some examples, the sensor data captured by device 100 may also include two-dimensional (2D) images. To capture 3D images, computing device 100 may be configured to capture multiple depth measurements, each corresponding to a pixel of a depth sensor. In some examples, each pixel may also include an intensity measurement.
[0015] Then, for example, based on the calibration parameters of the depth sensor, the depth measurement values and sensor pixel coordinates can be converted into multiple points forming a point cloud. Each point is defined by three-dimensional coordinates according to a predetermined coordinate system. Thus, the point cloud defines the three-dimensional positions of corresponding points on the target object 104 and any other object within the FOV of the depth sensor, such as the support surface 108 of the support object 104.
[0016] To capture a two-dimensional image, computing device 100 can be configured to capture a two-dimensional pixel array via an image sensor. Each pixel in the array can be defined by a color value and / or a brightness (e.g., intensity) value. For example, the image can include a color image (e.g., an RGB image). As mentioned above, a three-dimensional image can also include an intensity value for each pixel, so in some examples, the two-dimensional image can also be derived from the same dataset as the three-dimensional image.
[0017] Device 100 (or, in some examples, another computing device, such as a server configured to acquire sensor data from device 100) can be configured to determine dimensions from the point cloud mentioned above, such as the width “W”, depth “D”, and height “H” of the target object 104. The dimensions determined from the point cloud can be used in various downstream processes, such as optimizing the loading arrangement for storage containers, pricing of transportation services based on package size, etc.
[0018] Figure 1 Some internal components of device 100 are also shown. For example, device 100 includes a processor 112 (e.g., a central processing unit (CPU), graphics processing unit (GPU), and / or other suitable control circuitry, microcontroller, etc.). Processor 112 is interconnected with non-transitory computer-readable storage media (such as memory 116). Memory 116 includes a combination of volatile memory (e.g., random access memory, i.e., RAM) and non-volatile memory (e.g., read-only memory, i.e., ROM, electrically erasable programmable read-only memory, i.e., EEPROM, flash memory). Memory 116 may store computer-readable instructions, which, by being executed by processor 112, configure processor 112 to perform various functions in conjunction with certain other components of device 100. Device 100 may also include a communication interface 120, enabling device 100 to exchange data with other computing devices, for example, via a local area network and / or wide area network, short-range communication links, etc.
[0019] Device 100 may also include one or more input and output devices, such as display 124 with an integrated touchscreen. In other examples, input / output devices may include any suitable combination of microphone, speaker, keyboard, data capture trigger, etc.
[0020] Device 100 further includes a depth sensor 128, which can be controlled by processor 112 to capture 3D images and generate point cloud data therefrom, as described above. Depth sensor 128 may include a time-of-flight (ToF) sensor, a stereo camera assembly, a LiDAR sensor, etc. Depth sensor 128 may be mounted on the housing of device 100, for example, on the back of the housing (opposite to display 124, e.g.) Figure 1As shown), and has an optical axis substantially perpendicular to the display 124. In some embodiments, the device 100 may also include an image sensor 132. The image sensor 132 may include a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) camera and may be controlled to capture two-dimensional (2D) images as described above. Although sensors 128 and 132 are described herein as physically separate sensors, in other examples, sensors 128 and 130 may be combined together. Thus, sensors 128 and 132 may also be referred to as sensor components, which may be separate depth sensors and image sensors, or a single combined sensor. For example, some depth sensors (such as ToF sensors) can capture both three-dimensional and two-dimensional images by capturing intensity data and depth measurements.
[0021] When implemented as a ToF sensor, depth sensor 128 may include an emitter (e.g., an infrared laser emitter) configured to illuminate a scene and a sensor (e.g., an infrared-sensitive image sensor) configured to capture reflected light from such illumination. Depth sensor 128 may include a controller configured to determine a depth measurement for each captured reflection based on the time difference between the illumination pulse and the reflection. The depth measurement for a given pixel indicates the distance between depth sensor 128 itself and the spatial point of origin of the reflection. Thus, each depth measurement can represent a point in the resulting point cloud. Depth sensor 128 and / or processor 112 may be configured to convert the depth measurements into points in a three-dimensional coordinate system to generate a point cloud from which the dimensions of object 104 can be determined.
[0022] For example, determining the size of object 104 can include detecting the support surface 108 and the upper surface 136 of object 104 in the point cloud. The height H of object 104 can be determined as the vertical distance between the upper surface 136 and the support surface 108. The width W and depth D can be determined as the size of the upper surface 136. In other examples, the size of object 104 can be determined by detecting the intersections between the planes forming the surface of object 104, by calculating the smallest hexahedron that can contain object 104, etc. Furthermore, using algorithms for irregularly shaped objects, the size of irregularly shaped objects can be labeled from the point cloud and depth images.
[0023] The detection and dimensioning of object 104 based on a point cloud captured by depth sensor 128 can be affected by various factors. Obstacles (such as the operator's fingers, clothing, etc.) that are sufficiently close to the field of view (FOV) of sensor 128 can negatively impact dimensioning accuracy, even if they do not obstruct the object 104's line of sight relative to sensor 128. For example, an obstacle sufficiently close to sensor 128 may cause the edges of object 104 to be distorted in the generated point cloud, which could lead to inaccurate dimensions of object 104. Therefore, device 100 implements certain actions to suppress dimensioning of object 104 based on a 3D image that may contain neighboring obstacles relative to sensor 128.
[0024] Memory 116 stores computer-readable instructions that are executed by processor 112. Specifically, memory 116 stores a preprocessing application 140, which, when executed by processor 112, configures processor 112 to process 3D images captured via depth sensor 128 to determine whether these images may contain nearby obstacles (e.g., obstacles close enough to sensor 128 to deform other objects in the image, such as object 104). In this example, memory 116 also stores a dimensioning application 144 (also referred to as a dimensioning module), which, when executed by processor 112, configures processor 112 to process point cloud data captured via depth sensor assembly 128 to determine the dimensions of objects (such as object 104) in the image. Figure 1 (Width, depth, and height shown).
[0025] For illustrative purposes, Application 140 and Application 144 are shown as distinct applications, but in other examples, the functionality of Application 140 and Application 144 may be integrated into a single application. In further examples, either or both of Application 140 and Application 144 may be implemented by one or more specially designed hardware and firmware components, such as FPGAs, ASICs, etc.
[0026] refer to Figure 2An example scenario is shown where a partial obstacle near sensor 128 causes partial distortion of the captured 3D image. A rear surface 200 of device 100 opposite display 124 is shown. Sensors 128 and 132 are shown disposed on the rear surface 200, with sensor 128 partially obstructed, for example, by the finger 204 of the operator of device 100. For example, when the operator holds device 100, finger 204 may rest against the rear surface 200. As previously described, sensor 128 includes a transmitter 208 and a receiver 212. When finger 204 partially obstructs transmitter 208, the high-intensity reflection of light emitted by transmitter 208 may affect receiver 212 due to the small distance between transmitter 208 and finger 204. Figure 2 The lower part of the image shows the viewfinder presented on the display 124. Although the object 104 itself is not obscured by the finger 204, some parts of the object 104 are curved in this example captured by the sensor 128, and the dimensions obtained from a 3D image containing such deformation may be inaccurate.
[0027] Go to Figure 3 A method 300 for automatic detection of sensor obstacles for moving dimensioning is illustrated. The method 300 is described below in conjunction with the execution of device 100, for example, dimensioning object 104. It will be understood from the following discussion that method 300 can also be executed by various other computing devices, including those functionally similar to those used in the method. Figure 1 The mentioned sensor 128 is a depth sensor or is connected to a depth sensor.
[0028] At box 305, device 100 is configured to capture a 3D image, for example, by capturing multiple depth measurements via sensor 128 and generating a point cloud therefrom. In some examples, at box 310, device 100 may also be configured to capture a 2D image via sensor 132 substantially simultaneously with box 305. Capturing a 2D image at box 310 is optional and may be omitted in other embodiments. As discussed below, when a nearby obstacle is detected, device 100 may optionally employ sensors from different sensors (such as...) Figure 2 As shown, two-dimensional images (captured at different locations on device 100) are used to select which notification to send to the operator of device 100. 3D images (and, if applicable, 2D images) captured at boxes 305 and 310 can be, for example, one of a sequence of captured images at an appropriate frame rate. Each captured image can be displayed on display 124, for example, to achieve... Figure 2 The electronic viewfinder function is shown.
[0029] At box 315, device 100 is configured to detect one or more regions of interest (ROIs) in a 3D image via execution of application 140. For example, the ROI detected at box 315 could correspond to object 104 and supporting surface 108. When there are such Figure 2 When there is an obstacle such as finger 204 as shown, the ROI detected at box 315 may also include the ROI corresponding to the candidate obstacle.
[0030] For example, at box 315, processor 112 can be configured to perform one or more segmentation operations to detect object 104 in the 3D image from box 305. Such segmentation operations include machine learning-based segmentation models, such as You Only Look Once (YOLO). Various other models can be used for segmentation, such as region-based convolutional neural networks (R-CNN), fast CNNs, and plane fitting algorithms such as RANdom SAmple Consensus (RANSAC). At box 310, various other segmentation operations can also be employed, including thresholding operations, edge detection operations, region growing operations, etc.
[0031] Go to Figure 4 The diagram illustrates an example 3D image 400 and an example 2D image captured by a depth sensor 128 at box 305. For example, the depth sensor 128 could be a ToF sensor configured to capture depth measurements and intensity values for each of a plurality of pixels. The 3D image 400 is shown in the form of a point cloud generated from the aforementioned depth measurements. As shown in the 3D image 400, object 104 is as described above. Figure 2 The deformation is as discussed. At box 315, processor 112 can detect two ROIs 404-1 and 404-2 (collectively referred to as ROI 404, or commonly as ROI 404) from image 400. Figure 4 As shown, ROI 404 corresponds to a region in image 400 with similar depth and / or intensity (e.g., the difference in depth and / or intensity within ROI 404 is less than the difference in depth and / or intensity between ROI 404 and adjacent portions of image 400). Therefore, ROI 404 typically corresponds to a physical object; in this example, ROI 404-1 corresponds to finger 204, while ROI 404-2 corresponds to object 104.
[0032] Processor 112 can also detect Regions of Interest (ROIs) in image 402, for example, by applying 2D segmentation techniques to intensity values captured by sensor 128, by mapping the location of ROI 404 detected from the 3D image to pixel coordinates of sensor 128, or a combination thereof. In this example, processor 112 detects a first ROI 408-1 in image 402 corresponding to finger 204 (and thus representing the same physical object as ROI 404-1), and a second ROI 408-2 corresponding to object 104 (and thus representing the same physical object as ROI 404-2).
[0033] Processor 112 can also be configured to determine certain attributes of each ROI 404 and ROI 408 detected at box 315. For example, such as Figure 4 As shown, processor 112 can determine the size of each ROI 404 (e.g., the pixel area corresponding to the pixel array of sensor 128) from 3D image 400, and determine the depth indicating the average distance between sensor 128 and the pixels representing ROI 404. From image 402, processor 112 can determine the intensity indicating the average intensity or luminance of the pixels representing ROI 408. Therefore, processor 112 can at least determine the depth of each physical object in the FOV of sensor 128, and can also determine the size and intensity of each physical object observed by sensor 128. In other examples, intensity measurements can be obtained from a separate sensor, such as sensor 132.
[0034] At boxes 325 and 330, device 100 can be configured to evaluate each ROI from box 315 to determine whether the ROI likely represents a nearby obstacle, such as... Figure 2 The finger 204 is shown. At box 325, processor 112 can be configured to determine whether the depth of the considered ROI 404 is below a threshold. In other words, processor 112 is configured to determine whether each ROI 404 corresponds to an object that is close enough to sensor 128 to potentially deform the object 104 and reduce the accuracy of the dimensions determined for the object 104.
[0035] The determination at box 325 may include comparing the depth determined for each ROI 404 with a depth threshold and determining whether the depth of each ROI 404 is below the threshold. The threshold may be, for example, a minimum depth setting for sensor 128 (e.g., manufacturer's guideline), indicating a distance at which sensor 128 might produce inaccurate results. For example, in this example, the depth threshold may be 20 cm, although various other depth thresholds may be applied to other sensors 128. Figure 4As shown, ROI 404-2 has a depth exceeding the threshold, while ROI 404-1 has a depth below the threshold. Therefore, the determination result for ROI 404-1 at box 325 is positive, while the determination result for ROI 404-2 at box 325 is negative.
[0036] In some examples, the determination at box 325 may include comparing the intensity of each ROI 408 to an intensity threshold, rather than comparing the depth of the corresponding ROI 404 to a depth threshold, or as a supplement to comparing the depth of the corresponding ROI 404 to a depth threshold. As mentioned above, nearby obstacles tend to cause high-intensity reflections at the receiver 212 of sensor 128, so an ROI 408 with high intensity can represent an obstacle close enough to sensor 128 to distort other objects in image 400. For example, refer to Figure 4 The intensities of ROI 408-1 and ROI 408-2 are 100 and 56, respectively. As will be apparent, a wide variety of intensity ranges can be achieved depending on sensor 128 and the image format generated by sensor 128. Figure 4 The values shown are purely illustrative. For example, for a threshold of 80, the determination of ROI 408-1 at box 325 is positive, while the determination of ROI 408-2 at box 325 is negative.
[0037] In some examples, depth-based and intensity-based comparisons can be used at box 325, and the determination of a given ROI 404 and its corresponding ROI 408 at box 325 can be affirmative if either or both of the two comparisons are positive (e.g., if ROI 404 is close enough to sensor 128 and / or the corresponding ROI 408 is bright enough).
[0038] When the determination result for all ROIs 404 in image 400 at box 325 is negative, device 100 can proceed to the dimensioning stage, as further described below. When the determination result for at least one ROI 404 at box 325 is positive, such as... Figure 4 As in the example, device 100 continues to box 330. At box 330, processor 112 can be configured to determine whether the size of ROI 404, whose determination result at box 325 is positive, exceeds a threshold. In other words, ROI 404 exceeding the depth threshold at box 325 does not need to be processed via box 330. In other examples, the size check at box 330 can be omitted.
[0039] The threshold can be predetermined, for example, stored in memory 140 as a component of application 140, and the threshold can be selected to ignore visual artifacts, such as small specular reflections, dust on the sensor 128 window, etc. For example, the threshold applied at box 325 can be ten pixels, and the determination results for ROI 404-1 and ROI 404-2 at box 325 are positive.
[0040] When the determination results for all ROIs 404 at frame 330 are negative, device 100 can proceed to the dimensioning stage. Specifically, processor 112 can proceed to frame 335, passing image 400 to dimensioning application 144, which can then determine and output the dimensions of object 104. As will be apparent to those skilled in the art, dimensioning application 144 can be configured (e.g., based on ROI 404) to identify object 104 in image 400, identify surface 136 and support surface 108, and determine the object's depth, width, and height. These dimensions can then be displayed on display 124, transmitted via communication interface 120 to another computing device, etc.
[0041] However, a positive determination of at least one ROI 404 at box 330 indicates that image 400 contains an ROI 404 that may represent a nearby obstacle. Therefore, when the determination of at least one ROI 404 at 330 is positive, device 100 proceeds to box 340. At box 340, device 100 may suppress dimensioning of the image from box 305, for example, by discarding image 400 without passing image 400 or any portion thereof to dimensioning application 144. In some examples, such as if application 144 relies on temporal filtering to generate dimensions from multiple frames, or if application 144 otherwise requires notification of frame discarding, processor 112 may generate an inter-application notification from application 140 to application 144 indicating that the frame is discarded.
[0042] It will now be apparent that, based on the determinations at boxes 325 and 330, device 100 selects a processing action for image 400 between suppressing dimensioning of image 400 (if image 400 may contain nearby obstacles) and transmitting image 400 for dimensioning. The selection of processing actions as described above enables device 100 to avoid generating inaccurate dimensions for object 104 when object 104 may be deformed due to nearby obstacles such as finger 204.
[0043] Furthermore, device 100 can be configured to generate a notification in certain circumstances when an image from frame 305 is discarded at frame 340. Specifically, at frame 345, processor 112 can be configured (e.g., via executing application 140) to determine whether size annotation of a predetermined limit number of images at frame 340 has been suppressed. This limit can be defined as a threshold, for example, held in memory 116. For example, the limit could be five frames (although various other limits can also be used). When the determination at frame 345 is negative, device 100 does not need to generate a notification because the nearby obstacle may only be present briefly. However, when the determination at frame 345 is positive, indicating that the nearby obstacle has been present for a threshold number of consecutive frames, device 100 can proceed to frame 350 to generate a notification, for example, on display 124 and / or via another output device.
[0044] For example, such as Figure 5 As shown, device 100 can display notification 500 on display 124, notifying the operator of device 100 that sensor 128 is obstructed and therefore dimension markings are suppressed.
[0045] In some examples, the 2D image captured at frame 310 via sensor 132 can be selected by processor 112 at frame 350. For example, an electronic viewfinder implemented by device 100 can use the 2D image(s) captured at frame 310 instead of the 3D image from frame 305. In some cases, due to the different physical locations of sensor 128 and sensor 132 (e.g., ... Figure 2 As shown), obstacles to sensor 128 may not obstruct sensor 132, such as Figure 2 The case of finger 204 shown. Therefore, if the electronic viewfinder is based on the image from sensor 132, the operator of device 100 may not notice from the viewfinder on display 124 that sensor 128 is obstructed. Therefore, processor 112 can be configured, for example, to determine whether ROI 404-1 (or any other ROI 404 at box 330 whose determination result is positive) is within the FOV of sensor 132.
[0046] Processor 112 can, for example, transform the position of ROI404-1 in the coordinate system of sensor 128 to the position in the coordinate system of sensor 132 based on a transformation defined by calibration data stored in memory 116. The calibration data can define the physical positions of sensors 128 and 132 relative to each other. The calibration data may also include other sensor parameters, such as focal length, field of view size, etc. The calibration data may include, for example, the extrinsic parameter matrix and / or intrinsic parameter matrix of each of sensors 128 and 132.
[0047] Processor 112 can then determine whether ROI 404-1 is within the FOV of sensor 132 and thus appears in the 2D image from frame 310. When ROI 404-1 is expected to be within the FOV of sensor 132, the notification at frame 350 can be omitted because the obstacle will be visible on display 124. However, when ROI 404-1 is outside the FOV of sensor 132, device 100 can generate a notification at frame 350, as described above.
[0048] Figure 6 The illustration shows an example viewfinder control action taken by device 100 on display 124, based in part on the above-described determination. For example, when ROI 404-1 is at least partially within the FOV of sensor 132, the 2D image 600 presented on display 124 (e.g., as an electronic viewfinder) includes a region 604 at least partially corresponding to ROI 404-1. Therefore, device 100 can omit the generation of a notification because the obstacle is visible on display 124, and it can be expected that the operator of device 100 will remove the obstacle. On the other hand, when ROI 404-1 is outside the FOV of sensor 132, the image 608 from frame 310 may not show any obstacle, so device 100 can generate notification 612 at frame 350.
[0049] Specific embodiments have been described in the foregoing specification. However, those skilled in the art will understand that various modifications and changes can be made without departing from the scope of the invention as set forth in the appended claims. Therefore, the specification and drawings are to be considered illustrative rather than restrictive, and all such modifications are intended to be included within the scope of this teaching.
[0050] These benefits, advantages, solutions to problems, and any one or more elements that may make any benefit, advantage, or solution occur or become more prominent are not to be construed as key, essential, or necessary features or elements of any or all claims. The invention is defined solely by the appended claims, including any amendments made during the pending period of this application and all equivalents of these claims in the patent announcement.
[0051] Furthermore, in this document, relational terms such as first and second, top and bottom, etc., may be used individually to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms “comprises,” “comprising,” “has,” “having,” “includes,” “including,” “contains,” “containing,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes, has, includes, or contains a list of elements includes not only those elements but may also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Elements beginning with "comprises," "has," "includes," or "contains" do not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes, has, includes, or contains that element, unless otherwise expressly stated herein. The term "a / an" is defined as one or more unless otherwise expressly stated herein. The terms "substantially," "essentially," "approximately," "about," or any other version of these terms are defined as being as close as understood by one of ordinary skill in the art, and in one non-limiting embodiment, these terms are defined as within 10%, in another within 5%, in yet another within 1%, and in yet another within 0.5%. The term "coupled" as used herein is defined as connected, although not necessarily directly connected or mechanically connected. A device or structure that is “configured” in a certain way is configured at least in that way, but may also be configured in ways not listed.
[0052] Certain expressions may be used in this document to list combinations of elements. Examples of such expressions include: “at least one of A, B, and C”; “one or more of A, B, and C”; “at least one of A, B, or C”; “one or more of A, B, or C”. Unless otherwise expressly stated, the above expressions cover any combination of A and / or B and / or C.
[0053] It will be understood that some embodiments may include one or more dedicated processors (or processing devices), such as microprocessors, digital signal processors, custom processors, and field-programmable gate arrays (FPGAs), and uniquely stored program instructions (including both software and firmware) that control one or more processors to implement some, most, or all of the functions of the methods and / or apparatuses described herein, in conjunction with some non-processor circuitry. Alternatively, some or all of the functions may be implemented by a state machine without stored program instructions, or in one or more application-specific integrated circuits (ASICs), wherein each function or some combination of certain functions is implemented as custom logic. Of course, a combination of these two approaches may also be used.
[0054] Furthermore, embodiments can be implemented as computer-readable storage media having computer-readable code stored thereon for programming a computer (e.g., including a processor) to perform the methods described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, hard disks, CD-ROMs, optical storage devices, magnetic storage devices, ROMs (read-only memories), PROMs (programmable read-only memories), EPROMs (erasable programmable read-only memories), EEPROMs (electrically erasable programmable read-only memories), and flash memory. Moreover, it is anticipated that those skilled in the art, while making potentially significant efforts driven by, for example, available time, current technology, and economic considerations, and numerous design choices, will be able to readily generate such software instructions and programs, as well as ICs, with minimal experimentation when guided by the concepts and principles disclosed herein.
[0055] This abstract is provided to allow the reader to quickly determine the nature of the disclosure. This abstract is submitted with the understanding that it is not intended to interpret or limit the scope or meaning of the claims. Furthermore, in the above detailed description, it can be seen that various features are grouped together in various embodiments for the purpose of making the disclosure coherent. This method of disclosure should not be construed as reflecting an intention to require more features than are expressly recited in the claims. Rather, as reflected in the appended claims, the inventive subject matter lies in fewer than all the features of a single disclosed embodiment. Therefore, the appended claims are thus incorporated into the detailed description, wherein each claim represents itself as a separately claimed subject matter.
Claims
1. A method comprising: A three-dimensional image corresponding to the object is captured via the sensors of a computing device; Detecting regions of interest in the 3D image, wherein the regions of interest correspond to obstacles; Determine the depth from the sensor to the region of interest; The determined depth is compared with the threshold. Based on the comparison between the depth of the obstacle and the threshold, it is determined whether to transmit the 3D image to the dimension annotation module.
2. The method of claim 1, wherein determining whether to transfer the 3D image to the dimensioning module comprises: (i) When the depth is below the threshold, the transmission of the 3D image to the dimension annotation module is suppressed, and (ii) When the depth exceeds the threshold, the 3D image is transmitted to the dimension annotation module to obtain the size of the object; as well as Perform the selected processing action.
3. The method of claim 1, wherein detecting the region of interest includes performing a segmentation operation on the three-dimensional image.
4. The method of claim 3, wherein detecting the region of interest further comprises: Determine the size of the region of interest; as well as The size is determined to be outside a size threshold before the depth is determined.
5. The method of claim 1, wherein the threshold corresponds to the minimum depth setting of the sensor.
6. The method of claim 1, further comprising: In response to suppressing the transmission of the three-dimensional image, it is determined whether the transmission of a predetermined number of consecutive three-dimensional images to the dimension annotation module has been suppressed; as well as When the transmission of the predetermined number of consecutive 3D images to the dimensioning module has been suppressed, a notification is generated via the output of the computing device.
7. The method of claim 6, further comprising: A two-dimensional image is captured substantially simultaneously with the three-dimensional image via a second sensor; as well as The notification is selected based on whether the two-dimensional image depicts at least a portion of the region of interest.
8. The method of claim 1, further comprising: Intensity values associated with the region of interest are captured using the three-dimensional image via the sensor. The intensity value is compared with a second threshold. as well as The processing action is selected based on the comparison between the depth and the threshold, and the comparison between the intensity value and the second threshold.
9. The method of claim 8, wherein selecting the processing action further comprises: When the depth exceeds the threshold and the intensity value is lower than the second threshold, the 3D image is transmitted to the dimension annotation module.
10. A computing device, comprising: sensor; as well as Processor, the processor being configured to: The sensor captures a three-dimensional image corresponding to the object; Detecting regions of interest in the 3D image, wherein the regions of interest correspond to obstacles; Determine the depth from the sensor to the region of interest; The determined depth is compared with the threshold. Based on the comparison between the depth of the obstacle and the threshold, it is determined whether to transmit the 3D image to the dimension annotation module.
11. The computing device of claim 10, wherein the processor is configured to determine whether to transmit the three-dimensional image to the dimensioning module in such a way as: (i) When the depth is below the threshold, the transmission of the 3D image to the dimension annotation module is suppressed, and (ii) When the depth exceeds the threshold, the 3D image is transmitted to the dimensioning module to obtain the dimensions of the object; and Perform the selected processing action.
12. The computing device of claim 10, wherein the processor is configured to detect the region of interest by performing a segmentation operation on the three-dimensional image.
13. The computing device of claim 12, wherein the processor is configured to detect the region of interest in such a manner as follows: Determine the size of the region of interest; and The size is determined to be outside a size threshold before the depth is determined.
14. The computing device of claim 10, wherein the threshold corresponds to a minimum depth setting of the sensor.
15. The computing device of claim 10, wherein the processor is configured to: In response to suppressing the transmission of the 3D images, it is determined whether the transmission of a predetermined number of consecutive 3D images to the dimensioning module has been suppressed; and When the transmission of the predetermined number of consecutive 3D images to the dimensioning module has been suppressed, a notification is generated via the output of the computing device.
16. The computing device of claim 15, wherein the processor is configured to: A two-dimensional image is captured substantially simultaneously with the three-dimensional image via a second sensor; and The notification is selected based on whether the two-dimensional image depicts at least a portion of the region of interest.
17. The computing device of claim 10, wherein the processor is configured to: Intensity values associated with the region of interest are captured using the three-dimensional image via the sensor. The intensity value is compared with a second threshold; and The processing action is selected based on the comparison between the depth and the threshold, and the comparison between the intensity value and the second threshold.
18. The computing device of claim 17, wherein the processor is configured to select the processing action in such a way as: When the depth exceeds the threshold and the intensity value is lower than the second threshold, the 3D image is transmitted to the dimension annotation module.