Method and computing device for mobile dimensioning

By combining a stereo camera and motion sensor on a mobile computing device to determine the tilt angle and gravity vector constraints, the problem of insufficient accuracy in dimension annotation in stereo imaging is solved, and high-precision object dimension detection under complex conditions is achieved.

CN115485724BActive Publication Date: 2026-04-17ZEBRA TECHNOLOGIES CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZEBRA TECHNOLOGIES CORP
Filing Date
2021-04-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Stereo imaging suffers from insufficient detection accuracy in object dimensioning, especially under low point cloud density conditions or gravity vector offset, leading to inaccurate reference plane detection and dimensioning.

Method used

By using a stereo camera assembly and motion sensors on a mobile computing device, the tilt angle of the reference object's supporting surface is determined, and combined with the gravity vector as a constraint, plane fitting is performed to improve the accuracy of the reference surface detection, thereby generating the object's dimensions.

Benefits of technology

It improves the accuracy of dimension annotation under low point cloud density and gravity vector offset conditions, ensuring the accuracy of object surface detection and the reliability of dimension calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115485724B_ABST
    Figure CN115485724B_ABST
Patent Text Reader

Abstract

A method comprising: obtaining a reference stereo image from a stereo camera assembly of a mobile computing device; detecting a reference surface from the reference stereo image; determining a tilt angle of the reference surface relative to an output vector of an auxiliary sensor of the mobile computing device; obtaining a first dimensioning stereo image from the stereo camera assembly; detecting the reference surface in the first dimensioning stereo image based on the tilt angle; detecting a first object surface in the first dimensioning stereo image; and generating dimensions of the object based on the first object surface and the reference surface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the annotation of moving dimensions. Background Technology

[0002] Stereo imaging can be used to dimension objects (such as packages). For example, the surfaces of the object to be dimensioned can be detected from a stereo image, and then the dimensions of those surfaces can be determined. However, under certain conditions, the accuracy of surface detection from a stereo image may be negatively affected, thus reducing the accuracy of the resulting dimensions. Summary of the Invention

[0003] In one embodiment, the present invention is a method comprising: obtaining a reference stereo image from a stereo camera assembly of a mobile computing device; detecting a reference object support surface from the reference stereo image; determining a tilt angle of the reference object support surface based on the reference object support surface detected from the reference stereo image and an output gravity vector of a motion sensor of the mobile computing device, the tilt angle corresponding to the angle between the output gravity vector of the motion sensor of the mobile computing device and the normal vector of the reference object support surface; obtaining a first dimension-annotated stereo image from the stereo camera assembly; detecting the reference object support surface in the first dimension-annotated stereo image using the tilt angle as a constraint; detecting a first object surface in the first dimension-annotated stereo image; and generating the dimensions of the object based on the first object surface and the reference object support surface. Attached Figure Description

[0004] The accompanying drawings (in which the same reference numerals denote the same or functionally similar elements throughout the different views) together with the following detailed description are incorporated into and form part of the specification, and serve to further illustrate embodiments including the concepts of the claimed invention, and to explain the various principles and advantages of those embodiments.

[0005] Figure 1 This is a diagram showing a mobile computing device used for dimensioning objects.

[0006] Figure 2 It shows Figure 1 A diagram showing the rear view of a mobile computing device.

[0007] Figure 3 yes Figure 1 A block diagram of some of the internal hardware components of a mobile computing device.

[0008] Figure 4 This is a flowchart of the method for dimensioning objects.

[0009] Figure 5 It shows Figure 4 The execution diagram of the method is shown in box 405.

[0010] Figure 6 It shows Figure 4 The execution diagram of the method is shown in box 410.

[0011] Figure 7 It shows Figure 4 Another execution diagram of the method in box 410.

[0012] Figure 8 and Figure 9 It shows Figure 4 The execution of the method is shown in block 415 of the diagram.

[0013] Those skilled in the art will understand that the elements in the accompanying drawings are shown for simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some elements in the drawings may be exaggerated relative to other elements to aid in understanding embodiments of the invention.

[0014] The apparatus and method configurations have been indicated in appropriate places in the accompanying drawings by conventional symbols, which show only those specific details relevant to understanding embodiments of the invention, so as not to obscure this disclosure with details that would be obvious to those skilled in the art who benefit from the description herein. Detailed Implementation

[0015] The examples disclosed herein relate to a method comprising: obtaining a reference stereo image from a stereo camera assembly of a mobile computing device; detecting a reference surface from the reference stereo image; determining a tilt angle of the reference surface relative to an output vector of an auxiliary sensor of the mobile computing device; obtaining a first dimension-annotated stereo image from the stereo camera assembly; detecting the reference surface in the first dimension-annotated stereo image based on the tilt angle; detecting a first object surface in the first dimension-annotated stereo image; and generating dimensions of the object based on the first object surface and the reference surface.

[0016] Additional examples disclosed herein relate to a computing device including: a stereo camera assembly; an auxiliary sensor; and a controller connected to the stereo camera assembly and the auxiliary sensor, the controller being configured to: control the stereo camera assembly to capture a reference stereo image; detect a reference surface from the reference stereo image; determine a tilt angle of the reference surface relative to an output vector of the auxiliary sensor; control the stereo camera assembly to capture a first dimension-annotated stereo image; detect the reference surface in the first dimension-annotated stereo image based on the tilt angle; detect a first object surface in the first dimension-annotated stereo image; and generate the dimensions of the object based on the first object surface and the reference surface.

[0017] Further examples disclosed herein relate to a non-transitory computer-readable medium storing instructions executable by a processor of a mobile computing device to: control the stereo camera assembly to capture a reference stereo image; detect a reference surface from the reference stereo image; determine a tilt angle of the reference surface relative to an output vector of an auxiliary sensor; control the stereo camera assembly to capture a first-dimensionally annotated stereo image; and detect the reference surface in the first-dimensionally annotated stereo image based on the tilt angle.

[0018] Detect the surface of a first object in the first dimensioned stereo image; and generate the dimensions of the object based on the first object surface and the reference surface.

[0019] Figure 1 A mobile computing device 100 (also referred to herein as mobile device 100 or simply device 100) is illustrated, which is implemented to capture stereoscopic images and determine the dimensioning of objects presented in the images. For example, device 100 can be operated to capture a first stereoscopic image of object 104, for example, a first side 108 of object 104 is shown. For example, object 104 may be a package or collection of packages in a transportation and logistics facility (e.g., on a pallet). From the first stereoscopic image, device 100 can detect side 108 and a reference surface. The reference surface is a surface on which object 104 is resting, such as ground 110, a ramp, a shelf, or another supporting surface detected by device 100. Device 100 can then be configured to determine the dimensions of object 104, such as height “H” and width “W”, from side 108.

[0020] Device 100 can also be configured to capture a second stereoscopic image, for example, after device 100 is repositioned to face another side 112 of object 104. Device 100 can detect side 112 and ground 110 from the second stereoscopic image and determine the dimensions of object 104, such as height H (e.g., for verifying the height determined from side 108) and depth “D”. The dimensions generated by device 100 can be used to generate a bounding box surrounding object 104 for use by other computing devices associated with device 100 (e.g., optimizing the use of space in a container for transporting object 104, determining the transport cost of object 104, etc.).

[0021] Although the sides 108 and 112 of object 104 are shown as planar surfaces in this example, in other examples, object 104 may include a combination of planar and non-planar (e.g., curved) surfaces.

[0022] As described above, the size of object 104 generated by device 100 depends in part on the accurate detection of ground 110 or other suitable reference surface. For example, a reference surface can be used to aid in the detection of sides 108 and 112 and as a reference from which height H can be measured. The detection of ground 110 can be negatively affected by various factors. One example of these factors is low point cloud density. In particular, a stereo image depicting both object 104 and ground 110 (from which a point cloud is generated for size determination) can depict a small area of ​​ground 110 (especially when object 104 is large). Detecting ground 110 in a stereo image that contains almost no pixels corresponding to ground 110 can lead to inaccurate detection because many candidate planes may match the available pixels corresponding to ground.

[0023] The accuracy of ground detection can be improved by providing input parameters to a plane detection mechanism (such as random sample consistency or RANSAC plane fitting). An example input is a gravity vector determined by device 100, indicating the vertical direction. The gravity vector can be provided as input to the plane fitting operation. For example, the plane fitting operation can be configured to search for planes substantially perpendicular to the gravity vector, which can eliminate some candidate planes that fit the stereo image in other ways and thus increase the likelihood of selecting an accurate plane corresponding to the ground 110.

[0024] However, the above method may lead to inaccurate reference plane detection when the gravity vector is offset (e.g., due to miscalibration of the motion sensor of device 100), or when the ground 110 is not actually perpendicular to gravity. For example, using the gravity vector as input may result in inaccurate detection of slopes or other inclined (i.e., non-horizontal) surfaces.

[0025] Therefore, device 100 implements additional functions such as detecting a reference surface like ground 110, as will be discussed below. Device 100 also generates dimensions as described above, for example, for displaying on display 116 of device 100, transmitting to another computing device, etc.

[0026] Go to Figure 2 The device 100 is shown from the rear to illustrate an example stereo camera assembly, as well as some other components of the device 100. Figure 2 As shown, the device 100 includes a housing 200 that supports various components of the device 100. Among the components supported by the housing 200 are... Figure 1 The display 116 shown may include an integrated touchscreen. In addition to or instead of the aforementioned display 116 and touchscreen, the device 100 may also include other input and / or output components. Examples of such components include speakers, microphones, keypads, etc.

[0027] The device 100 also includes a stereo camera assembly having a first camera 202-1 and a second camera 202-2 spaced apart from each other on the housing 200 of the device 100. Each camera 202 includes a suitable image sensor or a combination of an image sensor, optical components (e.g., a lens), etc. The cameras 202 have respective fields of view (FOV) 204-1 and 204-2 extending away from the rear surface 208 of the device 100 (opposite to the display 116). In the example shown, the FOV 204 is substantially perpendicular to the rear surface 208.

[0028] like Figure 2 As shown, the FOVs 204 overlap, enabling device 100 to identify any object present in either FOV 204 (e.g., Figure 1 Information about the object 104 shown (such as size). Figure 2 The degree of overlap shown is for illustrative purposes only. In other examples, FOV 204 may overlap to a greater or lesser extent than shown.

[0029] Before further discussing the functions implemented by device 100, we will refer to Figure 3 Describe certain components of device 100.

[0030] refer to Figure 3 The diagram shows a block diagram of some components of device 100. In addition to a display (and in this example, an integrated touchscreen) 116 and a camera 202, device 100 includes a dedicated controller (such as processor 300) interconnected with a non-transitory computer-readable storage medium (such as memory 304). Memory 304 includes a combination of volatile memory (e.g., random access memory or RAM) and non-volatile memory (e.g., read-only memory or ROM, electrically erasable programmable read-only memory or EEPROM, flash memory). Processor 300 and memory 304 each include one or more integrated circuits.

[0031] Device 100 also includes a communication interface 308, enabling server 100 to exchange data with other computing devices, for example, via network 312. Other computing devices may include server 316, which may be deployed within the facility where device 100 is deployed. Server 316 may also be deployed remotely from the aforementioned facility.

[0032] Furthermore, device 100 includes motion sensors 320, such as inertial measurement units (IMUs), including a suitable combination of gyroscopes, accelerometers, etc. Motion sensors 320 are configured to provide processor 300 with measurements defining the motion and / or orientation of device 100. Specifically, as discussed below, the motion sensors can provide a gravity vector that at least indicates the orientation of device 100 relative to the vertical direction (i.e., towards the planetary center). Alternatively, processor 300 can generate the gravity vector based on data received from motion sensors 320.

[0033] Memory 304 stores computer-readable instructions that are executed by processor 300. In particular, memory 304 stores dimensioning application 324, which, when executed by processor 300, configures processor 300 to process stereo images captured by camera 202 to detect reference surfaces such as ground 110, and also detects and dimensiones objects (such as object 104) using input parameters derived from the reference surfaces and the aforementioned gravity vector.

[0034] When the processor 300 is configured in this way through the execution of application 324, the processor 300 may also be referred to as a dimensioning controller or simply a controller. Those skilled in the art will understand that, in other embodiments, the functionality implemented by the processor 300 via the execution of application 324 may also be implemented by one or more specially designed hardware and firmware components (such as FPGAs, ASICs, etc.).

[0035] Now go to Figure 4 The functions implemented through device 100 will be discussed in more detail. Figure 4 The dimensioning method 400 is shown, and it will be discussed below in conjunction with the performance of the device 100.

[0036] At box 405, device 100 is configured to capture a reference stereo image using camera 202. As indicated herein, the stereo image is a pair of stereo images captured substantially simultaneously by cameras 202-1 and 202-2 and combined to generate a point cloud or other 3D representation of the scene captured by camera 202.

[0037] A reference stereo image can be captured at frame 405 by device 100 in response to input (e.g., input received at a touchscreen integrated with display 116). For example, the input could be a command to begin the dimensioning process. The reference stereo image is called "reference" because device 100 is configured to detect a reference surface, such as ground 110, from the reference image. That is, although the execution of frame 405 initiates a dimensioning process for the object, the object itself is not detected in the reference image, and it is not actually necessary to capture it in the reference image.

[0038] Go to Figure 5 An example execution of box 405 is shown. Specifically, device 100 is shown oriented such that the field of view 204 of camera 202 primarily or entirely captures a portion of the ground 110. A reference stereo image 500 obtained from the capture operation is also shown. Figure 5 As shown in the image. Although the reference stereoscopic image 500 is in Figure 5 As shown as a single two-dimensional frame, it will be understood that the reference stereo image 500 (and all other stereo images shown herein) can actually be a three-dimensional point cloud generated from a pair of frames captured by camera 202.

[0039] The reference stereo image 500 depicts a portion of the side 108 of object 104 and the ground 110. In other examples, the reference stereo image 500 may completely omit object 104. The processor 300 is configured to detect the ground 110 in the reference stereo image 500, for example, by applying a plane fitting operation (e.g., RANSAC) to the reference stereo image 500. When other surfaces (such as a portion of the side 108 of object 104) are present, the processor 300 may be configured to ignore these surfaces when they occupy a portion of image 500 below a threshold (e.g., 25% of the points in a point cloud).

[0040] The reference surface detected at box 405 can be defined, for example, by points and normal vectors in a reference frame (e.g., a reference frame in which the point cloud is defined). Processor 300 can be configured to evaluate the quality of the detected reference surface at box 405. For example, a plane-fitting operation can generate a confidence level indicating the proportion of points in reference image 500 located on the detected surface, or another suitable metric indicating how well the detected surface fits the points in image 500. When the confidence level falls below a threshold, processor 300 can control display 116 to present a warning or error message and / or prompt to repeat box 405.

[0041] return Figure 4 At box 410, processor 300 is configured to determine the tilt angle of the reference surface detected at box 405 relative to the output vector of the auxiliary sensor. In this example, the auxiliary sensor is combined with... Figure 3 The motion sensor 320 mentioned above has an auxiliary sensor output that is the aforementioned gravity vector. At block 410, the processor 300 is configured to determine the angle between the normal vector defining the reference surface and the gravity vector generated by or based on the output of the motion sensor 320.

[0042] refer to Figure 6The diagram shows a side view of device 100 and a detected reference surface 600 corresponding to ground 110. A normal vector 604 corresponding to the reference surface 600 is also shown. The normal vector 604 and / or other suitable parameters of the reference surface 600 are defined according to reference frame 608. Figure 6 The image also shows a gravity vector 612 generated by motion sensor 320 or by processor 300 based on data from motion sensor 320. (See image for details.) Figure 6 It is evident that the gravity vector 612 is not perfectly perpendicular, for example, due to miscalibration of the motion sensor 320 or other measurement errors.

[0043] The processor 300 is configured to determine an angle 616 between the gravity vector 612 and the normal vector 604 of the reference surface 600. The angle 616 may be defined as, for example, a set of angles (e.g., in each of the XZ, XZ, and YZ planes of the reference frame 608).

[0044] Figure 7 Another example of execution of box 410 is shown. Specifically, a reference surface 700 and a normal vector 704 are shown relative to reference frame 608. Furthermore, a gravity vector 712 is shown as having been generated by motion sensor 320. While gravity vector 712 is vertical, the ground represented by reference surface 700 is not horizontal. Figure 7 In the example, processor 300 is configured to generate angle 716 between normal vector 704 and gravity vector 712 for subsequent use in dimensioning.

[0045] Figure 6 and Figure 7 Both scenarios demonstrate a reference surface that is not perpendicular to the gravity vector. In this case, using the gravity vector alone as input for a plane fitting operation in subsequent dimensioning functions may result in inaccurate reference surface detection, and thus inaccurate dimensioning.

[0046] return Figure 4 At box 415, processor 300 is configured to capture a first dimensioned stereo image via camera 202. The dimensioned stereo image distinguishes itself from the aforementioned reference stereo image by its use in detecting and dimensioning object 104. Therefore, after executing box 410, processor 300 can be configured to generate a prompt on display 116 to guide the operator of device 100 to reorient device 100 to place object 104 within the field of view 204 of camera 202. When device 100 is oriented to capture object 104, processor can capture the first dimensioned stereo image in response to input on a touchscreen or activation of another appropriate input on device 100.

[0047] After capturing the first dimensioned stereo image, the processor 300 is configured to detect the aforementioned reference surface and at least one surface of the object 104. It will be apparent to those skilled in the art that the reference surface is more... Figure 5 The reference image 500 shown occupies a smaller portion of the dimensioned image. To aid in accurate detection of the reference surface, the processor 300 employs a tilt angle defined in frame 410.

[0048] refer to Figure 8 The device 100 is shown as having been repositioned so that the field of view 204 completely surrounds the object 104. In response to the activation of the input, the processor 300 captures a first-dimensional stereo image via the camera 202. Figure 8 The diagram shows a first-dimensional stereoscopic image 800, which depicts the entire side 108 of the object 104 and a portion of the ground 110. It will be apparent by comparing image 800 with image 500 that image 800 depicts a smaller portion of the ground 110 than image 500, and therefore, the accuracy of detecting the ground 110 from image 800 may be reduced compared to detecting the ground 110 from image 500.

[0049] Furthermore, at box 415, processor 300 is configured to detect the reference surface detected at box 405. That is, in this example, processor 300 is configured to detect ground 110 from image 800, for example, based on an appropriate plane fitting operation. However, contrary to the reference surface detection at box 405, processor 300 is configured to apply constraints to the plane fitting operation.

[0050] Specifically, the constraint is based on the tilt angle from box 410. As described above, the tilt angle is measured between the gravity vector (e.g., 612 or 712) and the normal vector 604 or 704 of the reference surface, as detected at box 405. Therefore, at box 415, processor 300 is configured to obtain the current gravity vector from motion sensor 320 (i.e., generated substantially simultaneously with capturing the first dimensioned stereo image). Processor 300 is further configured to generate the aforementioned constraint by combining the gravity vector and tilt angle from box 410.

[0051] Restrict to Figure 9 This shows the gravity vector 612 and the tilt angle 616 generated at box 410. (See diagram.) Figure 9 As can be seen, the combination of gravity vector 612 and tilt angle 616 produces constraint vector 900. Constraint vector 900 is provided as input to the plane fitting operation at box 415. Specifically, processor 300 is configured to search for a reference surface perpendicular to constraint vector 900 (rather than perpendicular to gravity vector 612).

[0052] Figure 9 Two candidate reference surfaces 904 and 908 derived from image 800 are also shown. For example, each of reference surfaces 904 and 908 can fit points in image 800 that depict the ground 110 just as well. However, the use of constraint vector 900 allows processor 300 to select candidate reference surface 904 as corresponding to the ground 110, rather than candidate surface 908 (which might have been selected if only gravity vector 612 was used as input to the plane fitting operation).

[0053] Refer again Figure 4 The processor 300 is also configured to detect at least one surface of the object 104. In this example, such as Figure 8 As shown, device 100 is oriented such that side 108 of object 104 faces device 100, and thus at box 415, processor 300 detects side 108 (instead of side 112 which is not visible in image 800).

[0054] In box 420, after the reference surface and one or more object surfaces have been detected, processor 300 is configured to determine a first subset of the dimensions of object 104. In this example, after the side 108 is identified from image 800, processor 300 is configured to determine the height H and width W of object 104 based on side 108 and reference surface 904 (corresponding to ground 110).

[0055] In some examples, execution of method 400 may terminate after execution at box 420. However, in this example, processor 300 is configured to proceed to box 425 and capture a second dimensioned stereo image. For example, when a reference surface and one or more object surfaces are successfully detected at box 415 and dimensioning is performed at box 420, processor 300 may generate a cue for the operator of device 100 to reposition device 100 to face the side 112 of object 104 and capture an additional frame.

[0056] At box 425, processor 300 controls camera 202 to capture a second dimension-annotated stereo image depicting object 104 (e.g., side 112) and a portion of the ground. Processor 300 is configured to detect ground 110 as a reference surface using the same method described in conjunction with box 425, and to detect side 112 of object 104. At box 430, processor 300 is configured to determine a second subset of the dimensions of object 104, such as height H and depth D.

[0057] In box 435, execute box 430 or box 420 (e.g.) Figure 4(As shown by the dashed line in the diagram) After that, the processor 300 is configured to present the dimensions defined at boxes 420 and / or 430. Presenting the dimensions may include displaying the dimensions on the display 116, transmitting the dimensions to another computing device (e.g., server 316), etc.

[0058] The changes to the aforementioned functions were envisioned. For example, not in Figure 4 The dimensioning process shown in the diagram does not begin with capturing a reference stereo image, but rather captures the reference stereo image after capturing the second dimensioned stereo image. In this embodiment, the point cloud generated from the first dimensioned stereo image captured in box 415 and the second dimensioned stereo image captured in box 425 is stored but not dimensioned. Alternatively, a reference surface is detected in the first and second dimensioned stereo images, and dimensions are then determined from the first and second dimensioned stereo images after capturing the reference stereo image and determining the tilt angle.

[0059] Although the process of dimensioning object 104 as described above is performed at device 100, in other examples, this process may be performed at another computing device, such as server 316. For example, device 100 may capture a reference stereo image and a dimensioned stereo image and transmit the images to server 316, whereby server 316 may detect the reference surface and the object surface, as well as the tilt angle and dimensions. As will be apparent, in this implementation, device 100 may also send a gravity vector 612 or 712 along with the captured stereo image to server 316.

[0060] Specific embodiments have been described in the foregoing specification. However, those skilled in the art will understand that various modifications and changes can be made without departing from the scope of the invention as set forth in the following claims. Therefore, the specification and drawings are to be considered illustrative rather than restrictive, and all such modifications are intended to be included within the scope of this teaching.

[0061] These benefits, advantages, solutions to problems, and any elements(s) that may make any benefit, advantage, or solution occur or become more prominent are not to be construed as key, essential, or necessary features or elements of any or all claims. The invention is defined solely by the appended claims, including any amendments made during the pending examination of this application and all equivalents of these claims in the patent announcement.

[0062] Furthermore, in this document, relational terms such as first and second, top and bottom, etc., may be used individually to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Terms including “(comprises),” “comprising,” “has,” “having,” “includes,” “including,” “contains,” “containing,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes, has, comprises, or contains a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus. Elements beginning with “comprising one,” “having one,” “containing one,” or “covering one” do not exclude the presence of additional identical elements in a process, method, article, or apparatus that includes, has, contains, or covers that element, unless otherwise expressly stated herein. The terms “a” and “an” are defined as one or more unless expressly stated otherwise herein. The terms “basically,” “approximately,” “about,” “approximately,” or any other version of these terms are defined as being as close as understood by those skilled in the art, and in one non-limiting embodiment, these terms are defined as being within 10%, in another within 5%, in yet another within 1%, and in still another within 0.5%. The term “coupled” as used herein is defined as connected, although not necessarily directly connected or mechanically connected. A device or structure “configured” in a certain way is configured at least in this manner, but may also be configured in ways not listed.

[0063] It will be understood that some embodiments may include one or more dedicated processors (or "processing devices"), such as microprocessors, digital signal processors, custom processors, and field-programmable gate arrays (FPGAs), and uniquely stored program instructions (including both software and firmware) that control one or more processors to implement some, most, or all of the functions of the methods and / or apparatuses described herein, in conjunction with certain non-processor circuitry. Alternatively, some or all of the functions may be implemented by a state machine without stored program instructions, or in one or more application-specific integrated circuits (ASICs), wherein each function or some combination of certain functions is implemented as custom logic. Of course, a combination of these two approaches may also be used.

[0064] Furthermore, embodiments can be implemented as computer-readable storage media having computer-readable code stored thereon for programming a computer (e.g., including a processor) to perform the methods described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, hard disks, CD-ROMs, optical storage devices, magnetic storage devices, ROMs (read-only memories), PROMs (programmable read-only memories), EPROMs (erasable programmable read-only memories), EEPROMs (electrically erasable programmable read-only memories), and flash memory. Moreover, it is anticipated that those skilled in the art, while making potentially significant efforts driven by, for example, available time, current technology, and economic considerations, and numerous design choices, will be able to readily generate such software instructions and programs, as well as ICs, with minimal experimentation when guided by the concepts and principles disclosed herein.

[0065] This abstract is provided to allow the reader to quickly determine the nature of the disclosure. This abstract is submitted with the understanding that it is not intended to interpret or limit the scope or meaning of the claims. Furthermore, in the above detailed description, it can be seen that various features are grouped together in various embodiments for the purpose of making the disclosure coherent. This method of disclosure should not be construed as reflecting an intention to require more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive subject matter lies in fewer than all the features of a single disclosed embodiment. Therefore, the following claims are thus incorporated into the detailed description, wherein each claim represents itself as a separately claimed subject matter.

Claims

1. A method for mobile dimensioning, the method comprising: The method includes: Obtain a reference stereo image from the stereo camera component of a mobile computing device; Detect the support surface of the reference object from the reference stereo image; Based on the reference object support surface detected from the reference stereo image and the output gravity vector of the motion sensor of the mobile computing device, the tilt angle of the reference object support surface is determined, the tilt angle corresponding to the angle between the output gravity vector of the motion sensor of the mobile computing device and the normal vector of the reference object support surface; A first-dimensional annotated stereo image is obtained from the stereo camera assembly; Using the tilt angle as a constraint, the reference object support surface in the first dimensioned stereo image is detected; Detect the surface of the first object in the first dimensioned stereo image; and The dimensions of the object are generated based on the surface of the first object and the supporting surface of the reference object.

2. The method as described in claim 1, characterized in that, The method further includes: A second-sized annotated stereo image is obtained from the stereo camera assembly; The reference object support surface in the second dimension-annotated stereo image is detected based on the tilt angle. Detect the surface of the second object in the second dimensioned stereo image; and Further dimensions of the object are generated based on the second object surface and the reference object support surface.

3. The method as described in claim 2, characterized in that, The method includes obtaining the reference stereo image before obtaining the first dimensioned stereo image.

4. The method as described in claim 2, characterized in that, The method includes obtaining the reference stereo image after obtaining the second dimensioned stereo image.

5. The method as described in claim 2, characterized in that, The dimensions include height and width, and the further dimension includes depth.

6. The method as described in claim 1, characterized in that, Using the tilt angle as a constraint includes providing the tilt angle as input to the plane fitting operation.

7. The method as described in claim 1, characterized in that, The method further includes: displaying the size on the display of the mobile computing device.

8. The method as described in claim 1, characterized in that, The reference object support surface includes at least one of the following: ground, slope, or frame.

9. A computing device for moving dimension annotation, characterized in that, The computing device includes: Stereo camera components; Motion sensors; and A controller, connected to the stereo camera assembly and the motion sensor, is configured to: Control the stereo camera assembly to capture a reference stereo image; Detect the support surface of the reference object from the reference stereo image; Based on the reference object support surface detected from the reference stereo image and the output gravity vector of the motion sensor of the mobile computing device, the tilt angle of the reference object support surface is determined, the tilt angle corresponding to the angle between the output gravity vector of the motion sensor of the mobile computing device and the normal vector of the reference object support surface; Control the stereo camera assembly to capture a stereo image with a first-size annotation; Using the tilt angle as a constraint, the reference object support surface in the first dimensioned stereo image is detected; Detect the surface of the first object in the first dimensioned stereo image; and The dimensions of the object are generated based on the surface of the first object and the supporting surface of the reference object.

10. The computing device as claimed in claim 9, characterized in that, The controller is further configured to: Control the stereo camera assembly to capture a second-size-annotated stereo image from the stereo camera assembly; The reference object support surface in the second dimension-annotated stereo image is detected based on the tilt angle. Detect the surface of the second object in the second dimensioned stereo image; as well as Further dimensions of the object are generated based on the second object surface and the reference object support surface.

11. The computing device as claimed in claim 10, characterized in that, The controller is further configured to capture the reference stereo image before capturing the first dimensioned stereo image.

12. The computing device as claimed in claim 10, characterized in that, The controller is further configured to capture the reference stereo image after capturing the second dimensioned stereo image.

13. The computing device as claimed in claim 10, characterized in that, The dimensions include height and width, and the further dimension includes depth.

14. The computing device as claimed in claim 9, characterized in that, In order to use the tilt angle as a constraint, the controller is configured to provide the tilt angle as input to the plane fitting operation.

15. The computing device as claimed in claim 9, characterized in that, The computing device further includes a display, wherein the controller is further configured to represent the size on the display.

16. The computing device as claimed in claim 9, characterized in that, The reference object support surface includes at least one of the ground, a slope, and a frame.

17. A non-transitory computer-readable medium storing instructions executable by a processor of a mobile computing device, the instructions comprising: Control the stereo camera assembly to capture a reference stereo image; Detect the support surface of the reference object from the reference stereo image; Based on the reference object support surface detected from the reference stereo image and the output gravity vector of the motion sensor of the mobile computing device, the tilt angle of the reference object support surface is determined, the tilt angle corresponding to the angle between the output gravity vector of the motion sensor of the mobile computing device and the normal vector of the reference object support surface; Control the stereo camera assembly to capture a stereo image with a first-size annotation; Using the tilt angle as a constraint, the reference object support surface in the first dimensioned stereo image is detected; Detect the surface of the first object in the first dimension-annotated stereo image; as well as The dimensions of the object are generated based on the surface of the first object and the supporting surface of the reference object.

18. The non-transitory computer-readable medium as claimed in claim 17, characterized in that, The instructions further include: Control the stereo camera assembly to capture a second-size-annotated stereo image from the stereo camera assembly; The reference object support surface in the second dimension-annotated stereo image is detected based on the tilt angle. Detect the surface of the second object in the second dimensioned stereo image; and Further dimensions of the object are generated based on the second object surface and the reference object support surface.

Citation Information

Patent Citations

  • Tracking system calibration using object position and orientation

    US20100302378A1

  • Wearable Electronic Device

    US20140139637A1

  • Systems and methods for multiview metrology

    US20140210950A1

  • Estimating a pose of a camera for volume estimation

    US20140355820A1

  • Method for detecting height

    US20170228602A1