Depth Estimation Based on the Bottom Position of an Object
By receiving image data from the camera, identifying the bottom position of the object and calculating the ratio, and estimating the distance between the camera and the object using the distance projection model, the distance measurement difficulty in the prior art that does not rely on imaging hardware is solved, real-time depth calculation on low-cost devices is realized.
Patent Information
- Application Number
- CN202180034017.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-03
- Filing Date
- 2021-05-17
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-05-17
AI Technical Summary
The prior art is difficult to accurately measure the distance between the camera and the object without relying on imaging hardware.
By receiving image data from the camera, identifying the vertical position of the bottom of the object in the image data, calculating the object bottom ratio, and using the distance projection model to estimate the physical distance between the camera and the object based on the object bottom ratio.
Real-time calculation of the depth of objects on low-cost computing devices is achieved, suitable for mobile and power-limited devices, and can provide accurate distance estimation under the conditions of a monocular field of view camera.
Smart Images

Figure CN115516509B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 033,964, filed on June 3, 2020, entitled "Depth Estimation Based on Object Bottom Position", the entire content of which is incorporated herein by reference as if fully set forth in this specification. BACKGROUND OF THE INVENTION
[0003] Stereo cameras can be used to determine the distance or depth associated with an object. Specifically, a stereo camera can capture two or more images of an object simultaneously. The distance or depth can be determined based on the known distance between the image sensors of the stereo camera and the differences between the representations of the object in the two or more simultaneously captured images. Similarly, a camera can be used in combination with a structured - light pattern projector to determine object distance or depth. Specifically, the distance or depth can be determined based on the degree to which the pattern is distorted, scattered, or otherwise altered when projected onto objects at different depths. In each method, depth or distance measurement involves imaging hardware that may not be available on some computing devices. SUMMARY OF THE INVENTION
[0004] Image data generated by a camera can be used to determine the distance between the camera and an object represented within the image data. The estimated distance to the object can be determined by identifying the vertical position of the bottom of the object within the image data and dividing the vertical position by the total height of the image data to obtain an object bottom ratio. A distance projection model can map the object bottom ratio to a corresponding estimate of the physical distance between the camera and the object. The distance projection model can operate under the assumption that the camera is positioned at a specific height within the environment and that the bottom of the object is in contact with the ground of the environment. Additionally, in the case where the image data is generated by a camera oriented at a non - zero pitch angle, an offset calculator can determine an offset to be added to the object bottom ratio to compensate for the non - zero camera pitch.
[0005] In a first example embodiment, a computer-implemented method is provided that includes receiving image data representing an object in an environment from a camera. The method further includes determining a vertical position of a bottom of the object within the image data based on the image data, and determining an object bottom ratio between the vertical position and a height of the image data. The method additionally includes determining an estimate of a physical distance between the camera and the object via a distance projection model and based on the object bottom ratio. The distance projection model defines a mapping between (i) each respective candidate object bottom ratio of a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment. The method further includes generating an indication of the estimate of the physical distance between the camera and the object.
[0006] In a second example embodiment, a computing system is provided that includes: a camera; a processor; and a non-transitory computer-readable storage medium storing instructions that, when executed by the processor, cause the processor to perform operations. The operations include receiving image data representing an object in an environment from the camera. The operations further include determining a vertical position of a bottom of the object within the image data based on the image data, and determining an object bottom ratio between the vertical position and a height of the image data. The operations additionally include determining an estimate of a physical distance between the camera and the object via a distance projection model and based on the object bottom ratio. The distance projection model defines a mapping between (i) each respective candidate object bottom ratio of a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment. The operations further include generating an indication of the estimate of the physical distance between the camera and the object.
[0007] In a third example embodiment, a non-transitory computer-readable storage medium is provided that stores instructions that, when executed by a computing system, cause the computing system to perform operations. The operations include receiving image data representing an object in an environment from the camera. The operations further include determining a vertical position of a bottom of the object within the image data based on the image data, and determining an object bottom ratio between the vertical position and a height of the image data. The operations additionally include determining an estimate of a physical distance between the camera and the object via a distance projection model and based on the object bottom ratio. The distance projection model defines a mapping between (i) each respective candidate object bottom ratio of a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment. The operations further include generating an indication of the estimate of the physical distance between the camera and the object.
[0008] In a fourth example embodiment, a system is provided that includes means for receiving image data representing an object in an environment from a camera. The system further includes means for determining a vertical position of a bottom of the object within the image data based on the image data, and means for determining an object bottom ratio between the vertical position and a height of the image data. The system additionally includes means for determining an estimate of a physical distance between the camera and the object by way of a distance projection model and based on the object bottom ratio. The distance projection model defines a mapping between (i) each respective candidate object bottom ratio of a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment. The system further includes means for generating an indication of the estimate of the physical distance between the camera and the object.
[0009] These and other embodiments, aspects, advantages, and alternatives will become apparent to those of ordinary skill in the art by reading the following detailed description and making appropriate reference to the drawings. Additionally, the present disclosure and other descriptions and drawings provided herein are intended to illustrate embodiments by way of example only, and thus, many variations are possible. For example, structural elements and process steps can be rearranged, combined, distributed, eliminated, or otherwise changed while remaining within the scope of the claimed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A computing system according to an example embodiment is shown.
[0011] Figure 2 A computing device according to an example embodiment is shown.
[0012] Figure 3 A system for estimating object distance according to an example embodiment is shown.
[0013] Figure 4A An optical model according to an example embodiment is shown.
[0014] Figure 4B A mapping between object bottom ratio and physical distance according to an example embodiment is shown.
[0015] Figure 4C 、 Figure 4D 、 Figure 4E and Figure 4F Various model errors according to an example embodiment are shown.
[0016] Figure 5 Compensation for camera pitch according to an example embodiment is shown.
[0017] Figure 6 A use case of a system for estimating object distance according to an example embodiment is shown.
[0018] Figure 7 A flowchart in accordance with an example embodiment is shown. DETAILED DESCRIPTION
[0019] Example methods, devices, and systems are described herein. It should be understood that the terms "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration". Any embodiment or feature described herein as "example", "exemplary", and / or "illustrative" is not necessarily to be construed as preferred or advantageous over other embodiments or features, unless so stated. Accordingly, other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein.
[0020] Thus, the example embodiments described herein are not meant to be limiting. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a variety of different configurations.
[0021] Furthermore, unless the context otherwise implies, features shown in each figure can be used in combination with one another. Accordingly, the figures should generally be regarded as an integral aspect of one or more general embodiments, understanding that not all features shown are necessary for each embodiment.
[0022] Moreover, any listing of elements, boxes, or steps in this specification or claims is for the purpose of clarity. Accordingly, such listing should not be construed as requiring or implying that these elements, boxes, or steps follow a particular arrangement or are to be performed in a particular order. Unless otherwise noted, the figures are not drawn to scale.
[0023] I. OVERVIEW
[0024] A computing device that includes a mobile and / or wearable computing device can include a camera capable of capturing an image of the environment. For example, when a user holds, wears, and / or uses the computing device, the camera can face the environment in front of the user. Accordingly, systems and operations are provided herein that can be used to determine the distance between the camera and an object represented in image data generated by the camera. These systems and operations can be implemented, for example, on a computing device that includes a camera, thus allowing these computing devices to measure the distance to an object within the environment.
[0025] The distance between the camera and an object can be determined by mapping the position of the bottom of the object as represented in the image data generated by the camera to a corresponding physical distance. This approach can provide a computationally inexpensive way to compute the depth of moving and / or stationary objects of interest in real time. Due to the relatively low computational complexity, these systems and operations can be implemented on low-end devices and / or devices with limited power, including mobile phones and / or wearable devices (e.g., smartwatches, dashboard cameras, etc.). Additionally, in some cases, the systems and operations can be used with image data fields generated by a monocular camera that represents the environment from a single viewpoint at a time. Thus, for example, the systems and operations can be used as an alternative to depth measurements based on stereoscopic image data generated by a stereoscopic camera and / or image data that includes a structured light pattern projected by a patterned light projector. However, in other cases, the systems and operations disclosed herein can be used in combination with other methods for distance / depth measurement such as stereoscopic imaging and / or structured light projection.
[0026] The depth measurement process can assume that the camera is set at a known height above the ground within the environment. For example, the camera and / or the computing device housing the camera can be coupled to the user at a specific location on the user's body - such as at the chest, arm, wrist, or waistline - and can thus remain at substantially the same height over time. The depth measurement process can be performed with respect to the images captured by the camera to determine the distances to various objects within the environment. The computing device can generate visual, auditory, tactile, and / or other representations of the distances. Thus, for example, a visually impaired user can use the computing device to navigate within the environment with the help of the depth measurement process performed by the computing device.
[0027] Specifically, in the case where the camera is at a known height, each point at a specific distance on the ground of the environment can be associated with a corresponding position on the image sensor of the camera and is expected to produce an image at that corresponding position. The relationship between the specific distance and its corresponding position on the image sensor can be represented by an empirically determined (e.g., learned and / or trained) mapping that can form part of a distance projection model. Since the mapping assumes that the camera is at a known height, multiple different mappings can be provided as part of the distance projection model, each associated with a different height, to allow for distance measurement as the height of the camera changes.
[0028] To determine the distance between a camera and an object represented in an image captured by the camera, a computing device may be configured to identify the vertical position of the bottom of the object. The bottom of the object can be used because it is expected to be in contact with a point on the ground of the environment, and the mapping of the distance projection model relates the image position to the distance of the point on the ground. The vertical position of the bottom of the object can be divided by the height of the image to obtain an object bottom ratio, which represents the vertical position of the bottom of the object as a fraction of the image height (i.e., the range of the object bottom ratio can be from 0 to 1). By encoding the vertical position as a ratio, rather than, for example, as an absolute number of pixels, the same mapping can be used when the image is downsampled or upsampled. That is, using the object bottom ratio allows the mapping to be resolution invariant for a given image aspect ratio.
[0029] Based on the object bottom ratio, a distance projection model can determine an estimate of the physical distance to the object. Specifically, the distance projection model can select a specific mapping to use based on an indication of the height at which the camera is positioned, the orientation of the camera (e.g., landscape vs. portrait), the aspect ratio of the image, and / or one or more additional camera parameters. That is, each mapping provided by the distance projection model can be associated with a different corresponding set of camera parameters. Thus, depending on which mapping is used, each object bottom ratio can be mapped to a different distance. The distance estimate may be accurate when the actual camera parameters match the camera parameters assumed by the mapping, but the distance estimate may be incorrect when these two sets of camera parameters are different. The error in the estimated distance can be proportional to the difference between the corresponding parameters of the two sets of camera parameters.
[0030] The selected mapping can be used to determine an estimate of the distance to the object by mapping the object bottom ratio to the corresponding physical distance. It is worth noting that since in some cases, the image can be generated by a single monocular field camera and there is no projection of structured light, the specification of camera parameters such as height, image orientation, aspect ratio, and others can be used as an alternative to depth cues that would otherwise be provided by a pair of stereo images or a structured light pattern.
[0031] In addition, when the camera is tilted upward, the position of the bottom of the object may appear to move lower in the image, and when the camera is tilted downward, the position of the bottom of the object may appear to move higher in the image. This apparent shift of the bottom of the object due to camera pitch can be compensated for by an offset calculator. Specifically, the offset calculator can determine an estimated offset expressed in terms of the object bottom ratio based on the product of the tangent of the camera pitch angle and the empirically determined focal length of the camera. Then, the estimated offset can be added to the object bottom ratio, and this sum can be provided as an input to the distance projection model. Adding the estimated offset can have the effect of shifting the vertical position of the bottom of the object back to the position where the bottom would be when the camera is at a zero pitch angle.
[0032] II. Example Computing Device
[0033] Figure 1 An example form factor of computing system 100 is shown. Computing system 100 can be, for example, a mobile phone, a tablet computer, or a wearable computing device. However, other embodiments are possible. Computing system 100 can include various elements such as a body 102, a display 106, and buttons 108 and 110. Computing system 100 can further include a front camera 104, a rear camera 112, a front infrared camera 114, and an infrared pattern projector 116.
[0034] The front camera 104 can be positioned on the side of the body 102 that is typically facing the user during operation (e.g., on the same side as the display 106). The rear camera 112 can be positioned on the side of the body 102 opposite the front camera 104. Referring to the cameras as front and rear is arbitrary, and computing system 100 can include multiple cameras positioned on various sides of the body 102. The front camera 104 and the rear camera 112 can each be configured to capture images in the visible spectrum.
[0035] The display 106 can be capable of representing a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some embodiments, the display 106 can display the current image captured by the front camera 104, the rear camera 112, and / or the infrared camera 114 and / or a digital representation of an image that can be or has recently been captured by one or more of these cameras. Thus, the display 106 can act as a viewfinder for the cameras. The display 106 can also support touchscreen functionality that can adjust settings and / or configurations of any aspect of the computing system 100.
[0036] The front camera 104 may include an image sensor and associated optics, such as a lens. The front camera 104 may provide zoom capabilities or be capable of having a fixed focal length. In other embodiments, interchangeable lenses may be used with the front camera 104. The front camera 104 may have a variable mechanical aperture and mechanical and / or electronic shutters. The front camera 104 is also capable of being configured to capture still images, video images, or both. Additionally, the front camera 104 can represent a monocular, stereo, or multi-view field camera. The rear camera 112 and / or the infrared camera 114 may be arranged similarly or differently. Additionally, one or more of the front camera 104, the rear camera 112, or the infrared camera 114 may be an array of one or more cameras.
[0037] Either or both of the front camera 104 and the rear camera 112 may include or be associated with an illumination component that provides a light field in the visible spectrum to illuminate a target object. For example, the illumination component can provide a flash or constant illumination of the target object. The illumination component can also be configured to provide a light field including one or more of structured light, polarized light, and light having a specific spectral content. Other types of light fields known and used to recover a three-dimensional (3D) model of an object are possible in the context of the embodiments herein.
[0038] The infrared pattern projector 116 may be configured to project an infrared structured light pattern onto a target object. In one example, the infrared projector 116 may be configured to project a dot pattern and / or a flood pattern. Thus, the infrared projector 116 may be used in combination with the infrared camera 114 to determine a plurality of depth values corresponding to different physical features of the target object.
[0039] That is, the infrared projector 116 may project a known and / or predetermined dot pattern onto the target object, and the infrared camera 114 may capture an infrared image of the target object including the projected dot pattern. The computing system 100 may then determine the correspondence between regions in the captured infrared image and specific portions of the projected dot pattern. Given the position of the infrared projector 116, the position of the infrared camera 114, and the position of the region in the captured infrared image corresponding to a specific portion of the projected dot pattern, the computing system 100 may then use triangulation to estimate the depth to the surface of the target object. By repeating this process for different regions corresponding to different portions of the projected dot pattern, the computing system 100 may estimate the depth of various physical features or portions of the target object. In this way, the computing system 100 can be used to generate a three-dimensional (3D) model of the target object.
[0040] The computing system 100 may also include an ambient light sensor that can continuously or from time to time determine the ambient brightness of the scene that cameras 104, 112, and / or 114 can capture (e.g., in terms of visible light and / or infrared light). In some embodiments, the ambient light sensor can be used to adjust the display brightness of the display 106. Additionally, the ambient light sensor can be used to determine the exposure length of one or more of the cameras 104, 112, or 114, or to assist in that determination.
[0041] The computing system 100 can be configured to capture images of a target object using the display 106 and the front camera 104, the rear camera 112, and / or the front infrared camera 114. The captured images can be a plurality of still images or a video stream. Image capture can be triggered by activating the button 108, pressing a soft key on the display 106, or by some other mechanism. Depending on the embodiment, images can be automatically captured at specific time intervals, e.g., when the button 108 is pressed, under appropriate lighting conditions of the target object, when the digital camera device 100 is moved a predetermined distance, or according to a predetermined capture schedule.
[0042] As described above, the functionality of the computing system 100 can be integrated into a computing device such as a wireless computing device, a cellular phone, a tablet computer, a laptop computer, etc. For purposes of example, Figure 2 is a simplified block diagram showing some components of an example computing device 200 that can include a camera assembly 224.
[0043] By way of example and not limitation, the computing device 200 can be a cellular mobile phone (e.g., a smartphone), a still camera, a video camera, a computer (such as a desktop, notebook, tablet, or handheld computer), a personal digital assistant (PDA), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a gaming console, a robotic device, or some other type of device equipped with at least some image capture and / or image processing capabilities. It should be understood that the computing device 200 can represent a physical image processing system, a specific physical hardware platform on which image sensing and processing applications run in software, or some other combination of hardware and software configured to perform image capture and / or processing functions.
[0044] As Figure 2 shown, the computing device 200 can include a communication interface 202, a user interface 204, a processor 206, a data storage device 208, and a camera assembly 224, all of which can be communicatively linked together via a system bus, network, or other connection mechanism 210.
[0045] The communication interface 202 may allow the computing device 200 to communicate with other devices, access networks, and / or transmission networks using analog or digital modulation. Thus, the communication interface 202 may facilitate circuit-switched and / or packet-switched communications, such as plain old telephone service (POTS) communications and / or Internet Protocol (IP) or other packetized communications. For example, the communication interface 202 may include a chipset and an antenna arranged for wireless communication with a radio access network or an access point. Additionally, the communication interface 202 may take the form of or include a wired interface, such as an Ethernet, Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI) port. The communication interface 202 may also take the form of or include a wireless interface, such as Wi-Fi, BLUETOOTH , Global Positioning System (GPS), or wide-area wireless interface (e.g., WiMAX or 3GPP Long Term Evolution (LTE)). However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used on the communication interface 202. Additionally, the communication interface 202 may include multiple physical communication interfaces (e.g., a Wi-Fi interface, a BLUETOOTH interface, and a wide-area wireless interface).
[0046] The user interface 204 may be used to allow the computing device 200 to interact with a human or non-human user, such as receiving input from the user and providing output to the user. Thus, the user interface 204 may include input components, such as a keypad, a keyboard, a touch-sensitive panel, a computer mouse, a trackball, a joystick, a microphone, and the like. The user interface 204 may also include one or more output components, such as a display screen, which may be combined with a touch-sensitive panel, for example. The display screen may be based on CRT, LCD, and / or LED technologies, or other technologies now known or later developed. The user interface 204 may also be configured to generate an auditory output via a speaker, a speaker jack, an audio output port, an audio output device, headphones, and / or other similar devices. The user interface 204 may also be configured to receive and / or capture auditory utterances, noises, and / or signals via a microphone and / or other similar devices.
[0047] In some embodiments, the user interface 204 may include a display screen that serves as a viewfinder for still camera and / or video camera functions supported by the computing device 200 (e.g., in both the visible and infrared spectra). Additionally, the user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and focusing of the camera functions and the capture of images. Some or all of these buttons, switches, knobs, and / or dials may be implemented via a touch-sensitive panel.
[0048] The processor 206 may include one or more general-purpose processors, such as a microprocessor, and / or one or more special-purpose processors, such as a digital signal processor (DSP), a graphics processing unit (GPU), a floating-point unit (FPU), a network processor, or an application-specific integrated circuit (ASIC). In some cases, the special-purpose processor is capable of performing image processing, image alignment and merging images, and other possibilities. The data storage device 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage devices, and may be integrated with the processor 206 in whole or in part. The data storage device 208 may include removable and / or non-removable components.
[0049] The processor 206 is capable of executing program instructions 218 (e.g., compiled or non-compiled program logic and / or machine code) stored in the data storage device 208 to perform the various functions described herein. Accordingly, the data storage device 208 may include a non-transitory computer-readable medium having program instructions stored thereon that, when executed by the computing device 200, cause the computing device 200 to perform any method, process, or operation disclosed in this specification and / or the accompanying drawings. Execution of the program instructions 218 by the processor 206 may cause the processor 206 to use the data 212.
[0050] As an example, the program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device drivers, and / or other modules) installed on the computing device 200 and one or more applications 220 (e.g., a camera function, an address book, email, web browsing, social networking, an audio-to-text function, a text translation function, and / or a game application). Similarly, the data 212 may include operating system data 216 and application data 214. The operating system data 216 may be primarily accessible by the operating system 222, and the application data 214 may be primarily accessible by one or more applications 220. The application data 214 may be arranged in a file system visible or hidden to the user of the computing device 200.
[0051] The applications 220 may communicate with the operating system 222 through one or more application programming interfaces (APIs). These APIs may facilitate, for example, the applications 220 to read and / or write the application data 214, transmit or receive information via the communication interface 202, receive and / or display information on the user interface 204, and so on.
[0052] In some dialects, the application 220 may be abbreviated as "app". Additionally, the application 220 can be downloaded to the computing device 200 through one or more online app stores or app markets. However, the application can also be installed on the computing device 200 in other ways, such as via a web browser or through a physical interface (e.g., USB port) on the computing device 200.
[0053] The camera component 224 can include, but is not limited to, an aperture, a shutter, a recording surface (e.g., photographic film and / or image sensor), a lens, a shutter button, an infrared projector, and / or a visible light projector. The camera component 224 can include components configured to capture images in the visible spectrum (e.g., electromagnetic radiation having a wavelength of 380 to 700 nanometers) and components configured to capture images in the infrared spectrum (e.g., electromagnetic radiation having a wavelength of 701 nanometers to 1 millimeter). The camera component 224 can be at least partially controlled by software executed by the processor 206.
[0054] III. Example Depth Determination System
[0055] Figure 3 An example system that can be used to determine an estimate of the physical distance between a camera and one or more objects within an environment is shown. Specifically, the system 340 can include an object bottom detector 308, an object bottom ratio calculator 310, an offset calculator 312, and a distance projection model 314, each of which can represent a combination of hardware components and / or software components configured to perform the corresponding operations described herein. The system 340 can be configured to receive image data 300 and metadata indicating parameters of the camera as inputs. The metadata can include information about the pose of the camera when the image data 300 is captured, such as the camera pitch 306. The image data 300 can represent one or more objects therein, such as objects 302 to 304. The objects 302 to 304 can include various moving and / or stationary features of the environment, such as humans, animals, vehicles, robotic devices, mailboxes, posts (e.g., lamp posts, traffic light posts, etc.), and / or benches, and other possibilities.
[0056] The object bottom detector 308 can be configured to detect the vertical position of the bottom of an object within the image data 300. In some embodiments, the vertical position can be expressed in pixels. For example, the object bottom detector 308 can determine that the bottom of the object 302 is positioned 250 pixels above the bottom of the image data 300. The object bottom detector 308 can implement one or more algorithms that are configured to (i) detect the object 302 within the image data 300, (ii) detect the bottom of the object 302 within the image data 300 based on the detection of the object 302 within the image data 300, and (iii) determine that the bottom of the object 302 is positioned on the ground of the environment. The object bottom detector 308 can perform commensurate operations with respect to the object 304 and / or any other object represented by the image data 300.
[0057] In the case where the object 302 is a human, one or more algorithms of the object bottom detector 308 can be configured to detect the human, detect the feet and / or shoes of the human (i.e., the bottom of the human within the image data 300), and determine that the feet and / or shoes are in contact with the ground. Similarly, when the object 304 is a vehicle, one or more algorithms of the object bottom detector 308 can be configured to detect the vehicle, detect the wheels and / or tires of the vehicle, and determine that the wheels and / or tires are in contact with the ground. The one or more algorithms can include various image processing algorithms, computer vision algorithms, and / or machine learning algorithms.
[0058] The object bottom detector 308 can be configured to provide the vertical position of the bottom of the object to the object bottom ratio calculator 310. The object bottom ratio calculator 310 can be configured to calculate the ratio between the vertical position of the bottom of the object (e.g., the object 302) and the height of the image data 300. Specifically, the object bottom ratio calculator can implement the function b = v / h, where b is the object bottom ratio, v is the vertical position of the bottom of the object, and h is the height of the image data 300. To this end, the object bottom ratio calculator 310 can determine the orientation of the image data 300 based on the metadata associated with the image data 300. Specifically, the object bottom ratio calculator 310 can determine whether the image data 300 has been acquired in a landscape orientation (i.e., where the longer side of the image data 300 is oriented horizontally) or a portrait orientation (i.e., where the longer side of the image data 300 is oriented vertically). Accordingly, the object bottom ratio calculator can set the value of h based on the orientation of the image data 300. For example, for image data 300 having a resolution of 3840 pixels by 2160 pixels, based on determining that the image data 300 is a portrait image, the height h of the image data 300 can be set to 3480 pixels, or based on determining that the image data 300 is a landscape image, the height h of the image data 300 can be set to 2160.
[0059] The distance projection model 314 can be configured to determine an estimated physical distance 336 between an object (e.g., object 302) represented in the image data 300 and the camera that generated the image data 300. Specifically, the distance projection model 314 can determine the estimated physical distance 336 based on the sum of the object bottom ratio calculated by the object bottom ratio calculator and the estimated offset of the object bottom ratio calculated by the offset calculator 312 to account for the camera pitch 306.
[0060] The offset calculator 312 can be configured to determine the amount or offset by which the object bottom ratio will be shifted / adjusted to account for a non-zero camera pitch angle based on the camera pitch 306. Specifically, the distance projection model 314 can be implemented under the assumption that the image data 300 has been captured when the optical axis of the camera is oriented substantially parallel to the ground in the environment. When the camera is tilted upward to a positive pitch angle, the object bottom ratio calculated by the object bottom ratio calculator 310 decreases relative to the object bottom ratio at zero pitch angle. Similarly, when the camera is tilted downward to a negative pitch angle, the object bottom ratio calculated by the object bottom ratio calculator 310 increases relative to the object bottom ratio at zero pitch angle. Thus, without the offset calculator 312, the estimated physical distance 336 may be underestimated at positive pitch angles and overestimated at negative pitch angles, as Figure 4E and Figure 4F shown and explained with reference to Figure 4E and Figure 4F respectively.
[0061] The offset calculator 312 can thus allow the distance projection model 314 to generate an accurate distance estimate by correcting for the camera pitch 306. The correction process is shown and discussed in more detail in Figure 5 and with reference to Figure 5 respectively. The object bottom ratio determined by the object bottom ratio calculator 310 and the estimated offset calculated by the offset calculator 312 can be added together, and the sum can be provided as an input to the distance projection model 314.
[0062] The distance projection model 314 can include a plurality of mappings 316 to 326, and an estimated physical distance 336 can be determined through one or more of the mappings 316 to 326. Each of the mappings 316 to 326 can associate a plurality of object bottom ratios with a plurality of corresponding physical object distances. For example, mapping 316 can associate object bottom ratios 318 to 322 with corresponding physical object distances 320 to 324. Similarly, mapping 326 can associate object bottom ratios 328 to 332 with corresponding physical object distances 330 to 334. The object bottom ratios associated with the mappings 316 to 326 (e.g., object bottom ratios 318 to 322 and 238 to 332) can be referred to as candidate object bottom ratios because each can be used to determine the estimated physical distance 336.
[0063] Each of the mappings 316 to 326 can be associated with a corresponding set of camera parameters, which can include, for example, the orientation of the image data 300 (i.e., landscape or portrait), the height in the environment where the camera is set when capturing the image data 300, the field of view of the camera used to capture the image data 300 (e.g., as defined by the size of the camera's image sensor and the optical properties of the camera's lens), and / or the aspect ratio of the image data 300, among other possibilities. Thus, one of the mappings 316 to 326 can be selected and used to determine the estimated physical distance 336 based on the values of the camera parameters associated with the image data 300, and these values of the camera parameters can be indicated as part of the metadata associated with the image data 300. Thus, the object bottom ratios 318 to 322 can be similar to, overlap with, or be the same as the object bottom ratios 328 to 332, but the object bottom ratios 318 to 322 can be mapped to a different set of physical object distances than the object bottom ratios 328 to 332. That is, the physical object distances 320 to 324 can be different from the physical object distances 330 to 334, although the two sets may overlap.
[0064] IV. Example Models for Depth Determination
[0065] Figure 4A An example geometric model of a camera is shown. The geometric model can serve as a basis for the mappings 316 to 326 used to generate the distance projection model 314. Specifically, Figure 4A An image sensor 400 and an aperture 402 are shown disposed in an environment including a ground 406. The image sensor 400 and the aperture 402 define an optical axis 404, along Figure 4A which the optical axis 404 extends substantially parallel to the ground 406. The image sensor 400 (i.e., its vertical center) is set at a height H above the ground 406, and the aperture 402 is positioned at a (focal) distance f relative to the image sensor 400.
[0066] Multiple lines are projected from corresponding points on the ground 406 in the environment through the aperture 402 to corresponding points on the image sensor 400. Specifically, the multiple lines include a 1-meter line, a 5-meter line, a 10-meter line, a 20-meter line, a 30-meter line, and an infinite reference line. For example, the 5-meter line (i.e., D = 5 meters) corresponds to a vertical position d relative to the center of the image sensor 400 and creates an image at the vertical position d relative to the center of the image sensor 400, and forms an angle θ with the optical axis 404. The 1-meter line corresponds to the minimum distance between the aperture 402 and an object that can be observable and / or measurable, because this line is projected onto the highest part of the image sensor 400.
[0067] The infinite reference line can correspond to the maximum observable distance in the environment, the distance to the horizon, the distance exceeding a threshold distance value, and / or an infinite distance. For example, when the infinite reference line originates above the ground 406, the infinite reference line can be associated with an infinite distance and thus not associated with a measurable distance along the ground 406. The infinite reference line is Figure 4A shown as approximately coinciding with the optical axis 404. The right part of the infinite reference line is drawn slightly below the optical axis 404, and the left part of the infinite reference line is drawn slightly above the optical axis 404 to visually distinguish the infinite reference line from the optical axis 404. Thus, in the Figure 4A configuration shown, the infinite reference line corresponds to the approximate center of the image sensor 400 and creates an image at the approximate center of the image sensor 400. As the height H of the image sensor 400 increases from the Figure 4A height shown (e.g., when the camera is mounted on an aircraft), the image created by the infinite reference line can move up along the image sensor 400. Similarly, as the height H of the image sensor 400 decreases from the Figure 4A height shown (e.g., when the camera is mounted on a floor cleaning robot), the image created by the infinite reference line can move down along the image sensor 400. Thus, as the height H changes, the infinite reference line may deviate from the optical axis 404. The corresponding positions of the images corresponding to the 1-meter line, 5-meter line, 10-meter line, 20-meter line, and / or 30-meter line on the image sensor 400 can similarly respond to changes in the height H of the image sensor 400.
[0068] Figure 4AThe example geometric model omits some components of the camera, such as the lens, which can be used to generate image data for depth determination. Thus, the example geometric model may not be an accurate representation of some cameras, and explicitly using the geometric model to calculate the object distance may result in an incorrect distance estimate. However, the geometric model shows that there is a non-linear relationship (e.g., tan(θ) = d / f = H / D or d = Hf / D) between the vertical position on the image sensor 400 and the corresponding physical distance along the ground 406 within the environment. Therefore, a non-linear numerical model (e.g., distance projection model 314) can be empirically determined based on training data to correct Figure 4A any inaccuracies in the geometric model and accurately map the position on the image sensor 400 to the corresponding physical distance along the ground 406.
[0069] In addition, Figure 4A the geometric model shows that variations in some camera parameters, including the height H of the camera (e.g., the height of the image sensor 400 and / or the aperture 402), the distance f, the angle θ, the field of view of the camera (defined by the size of the image sensor 400, the lens used to focus light on the image sensor 400, and / or the zoom level produced by the lens), the portion of the image sensor 400 from which the image data is generated (e.g., the aspect ratio of the image data), and / or the orientation of the image sensor 400 (e.g., landscape vs. portrait) can change the relationship (e.g., the mapping) between the physical distance along the ground 406 and the position on the image sensor 400. Therefore, these camera parameters can be taken into account by the non-linear numerical model in order to generate an accurate distance estimate.
[0070] Specifically, each of the mappings 316 to 326 can correspond to a specific set of camera parameters and can generate a distance estimate that is accurate for a camera with that specific set of camera parameters, but may be inaccurate when using a different camera with a different set of camera parameters. Therefore, one of the mappings 316 to 326 can be selected based on the actual set of camera parameters associated with the camera used to generate the image data 300. Specifically, the mapping associated with the camera parameters that most closely match the actual set of camera parameters can be selected.
[0071] For example, the first mapping among mappings 316 to 326 may be associated with a first set of camera parameters corresponding to a first mobile device equipped with a first camera, and the second mapping among mappings 316 to 326 may be associated with a second set of camera parameters corresponding to a second mobile device equipped with a second camera different from the first camera. Thus, the first mapping may be used to measure the distance to an object represented in the image data generated by the first mobile device, and the second mapping may be used to measure the distance to an object represented in the image data generated by the second mobile device. In cases where multiple different mobile devices each use a camera having a similar or substantially identical set of camera parameters, one mapping may be used by multiple different mobile devices. Additionally, since each camera may be positioned at multiple different heights H, each camera may be associated with multiple mappings, each corresponding to a different height.
[0072] Figure 4B A graphical representation showing an example mapping between an object bottom ratio and a physical distance is presented. This mapping can express the vertical position based on the corresponding object bottom ratio rather than expressing the vertical position on the image sensor 400 in terms of pixels. This allows the mapping to remain invariant to the image data resolution. Thus, when the image data is downsampled or upsampled, such a mapping can be used to determine the object distance because the object bottom ratio associated with the object has not changed (assuming the image is not cropped and / or the aspect ratio remains the same).
[0073] Specifically, the user interface (UI) 410 shows multiple horizontal lines corresponding to respective object bottom ratios including 0.0, 0.25, 0.35, 0.24, 0.47, and 0.5. Notably, the horizontal line associated with the object bottom ratio 0.5 is positioned approximately in the middle of the UI 410, dividing the UI 410 into a roughly equal upper half and lower half. The UI 412 shows the same multiple lines as the UI 410, and these lines are marked with corresponding physical distances including 1 meter, 5 meters, 10 meters, 20 meters, 30 meters, and infinity. That is, the object bottom ratios 0.0, 0.25, 0.35, 0.24, 0.47, and 0.5 respectively correspond to physical distances of 1 meter, 5 meters, 10 meters, 20 meters, 30 meters, and the distance associated with the Figure 4A infinite reference line (e.g., infinity). The object bottom ratio can be mapped to the corresponding distance by a function F(b), which can represent one of the mappings 316 to 326 of the distance projection model 314.
[0074] UIs 410 and 412 also display image data including object 414, which may represent a human. Bounding box 416 surrounds object 414. Bounding box 416 may represent the output of a first algorithm implemented by object bottom detector 308 and may be used to define a search region for a second algorithm implemented by object bottom detector 308. For example, bounding box 416 may define a region of interest that has been determined by the first algorithm to contain a representation of a human. Bounding box 416 may be provided as an input to a second algorithm configured to identify the feet and / or shoes of a human in an attempt to identify its bottom. Thus, when looking for the bottom of an object, bounding box 416 may reduce the search space considered by the second algorithm. Additionally, when bounding box 416 is associated with an object label or classification (e.g., human, vehicle, animal, etc.), that label may be used to select an appropriate algorithm for locating the bottom of the object associated with that label. For example, when bounding box 416 is classified as containing a representation of a car, an algorithm for looking for car wheels and / or tires may be selected to search for the bottom of the object rather than an algorithm for looking for human feet and / or shoes.
[0075] UIs 410 and 412 further show a line corresponding to the bottom of object 414. In UI 410, this line is marked with an object bottom ratio of 0.31 (indicating that the bottom of object 414 is located slightly below 1 / 3 of the bottom-up path of UI 410), while in UI 412, it is marked with a distance of 6 meters. The object bottom ratio of 0.31 and, in some cases, the corresponding line thereto may represent the output of object bottom ratio calculator 310 and / or offset calculator 312. The object bottom ratio of 0.31 may be mapped to a corresponding physical distance of 6 meters through function F(b).
[0076] F(b) may be determined based on empirical training data. For example, multiple physical distances may be measured relative to the camera and visually marked within the environment. The camera may be used to capture training image data representing these visually marked distances. When capturing the training image data, the camera may be set at a predetermined height within the environment. Thus, the function or mapping trained based on this training image data may be valid for (i) the same camera or another camera having a similar or substantially identical set of camera parameters and (ii) another camera located at a similar or substantially identical predetermined height for measuring distances. Based on training data obtained using a camera with a different set of camera parameters and / or the same camera located at a different height, a similar process may be used to determine additional functions or mappings.
[0077] In one example, function F(b) may be formulated as a polynomial model F(b) = a 0 + a 1 b 1 + a2 b 2 +a 3 b 3 +…+a n b n , where b represents the object bottom ratio, and a 0 -a n represents an empirically determined coefficient. Based on the training data, multiple object bottom ratios B 训练 =[d 0 =0.5m, d 1 =1.0m,..., d t =20.0m] associated with multiple physical distances D 训练 =[b 0 =0.0, b 1 =0.01,..., b t =0.5] can be used to determine the coefficient a 0 -a n , where A = [a 0 , a 1 ,..., a n . Specifically, a 0 -a n can be calculated by solving A for the equation AB′ 训练 =D 训练 , where B′ 训练 is equal to Therefore, AB′ 训练 =D 训练 can be rewritten as Once the coefficient a 0 -a n is determined based on the training data, the function F(b) can be used to determine the physical distance between the camera and the object based on the object bottom ratio associated with the object. Specifically, AB 观察 =D 估计 , where and D 估计 is a scalar value corresponding to the estimated physical distance 336.
[0078] In other examples, the function F(b) can be implemented as an artificial intelligence (AI) and / or machine learning (ML) model. For example, an artificial neural network (ANN) can be used to implement the mapping between the object bottom ratio and the physical distance. In some embodiments, each set of camera parameters can be associated with a corresponding ANN. That is, each of the mappings 316 to 326 can represent a separate ANN trained using image data captured by a camera with a corresponding set of camera parameters. In other embodiments, a single ANN can implement each of the mappings 316 to 326 simultaneously. To this end, the ANN can be configured to receive at least a subset of the camera parameters as input, which can adjust how the ANN maps the input object bottom ratio to the corresponding physical distance. Thus, the ANN can be configured to map each candidate object bottom ratio to multiple physical distances, and the specific physical distance for a particular object bottom ratio can be selected by the ANN based on the values of the camera parameters.
[0079] Notably, the distance projection model 314 can be configured to determine the distance associated with an object based on a single visible spectrum image captured using a monocular field of view camera without relying on structured light. That is, the distance can be determined without using stereoscopic image data or projecting a predefined pattern onto the environment. Instead, to accurately determine the distance to an object, the distance projection model 314 and / or the offset calculator 312 can estimate the object distance based on camera parameters that define the pose of the camera relative to the environment and / or the optical characteristics of the camera, as well as other aspects of the camera. When the pose of the camera changes and / or a different camera is used, the camera parameters can be updated so that the distance projection model 314 and / or the offset calculator 312 can compensate for such differences, for example, by using an appropriate mapping. However, in some cases, the system 340 can be used in combination with other depth determination methods that rely on stereoscopic image data and / or structured light projection.
[0080] V. Example Model Errors and Error Correction
[0081] Figure 4C 、 Figure 4D 、 Figure 4E and Figure 4F illustrate the errors that can occur when the actual camera parameters deviate from the camera parameters assumed or used by the distance projection model 314. Specifically, Figure 4CThe top of [figure] shows the image sensor 400 moving up from height H to height H′, causing the optical axis 404 to move up a proportional amount, as indicated by line 418. Without this upward movement, the bottom portion of the object 414 closest to the image sensor 400 would create an image on the image sensor 400 at a distance d above the center of the image sensor 400, as indicated by line 422. However, the upward movement causes the image to be created instead at a distance d′ (which is greater than d) above the center of the image sensor 400, as indicated by line 420.
[0082] If the mapping used to calculate the distance D between the aperture 402 and the object 414 corresponds to height H instead of height H′, then the mapping may incorrectly determine that the object 414 is located at a distance D′, as Figure 4C indicated by the bottom portion of [figure], rather than at distance D. Specifically, Figure 4C the bottom portion of [figure] shows the image sensor 400 moving back down so that line 418 coincides with the optical axis 404, and line 420 extends from the same point on the image sensor 400 to a distance D′ on the ground, rather than distance D. The distance D′ is shorter than distance D, resulting in an underestimated distance estimate. By using a mapping that corresponds to the camera height H′ instead of H, this error can be reduced, minimized, or avoided.
[0083] Similarly, Figure 4D the top portion of [figure] shows the image sensor 400 moving down from height H to height H″, causing the optical axis 404 to move down a proportional amount, as indicated by line 424. Without this downward movement, the bottom portion of the object 414 closest to the image sensor 400 would create an image on the image sensor 400 at a distance d above the center of the image sensor 400, as indicated by line 422. However, the downward movement causes the image to be created instead at a distance d″ (which is less than d) above the center of the image sensor 400, as indicated by line 426.
[0084] If the mapping used to calculate the distance D between the aperture 402 and the object 414 corresponds to height H instead of height H″, then the mapping may incorrectly determine that the object 414 is located at a distance D″, as Figure 4D indicated by the bottom portion of [figure], rather than at distance D. Specifically, Figure 4D the bottom portion of [figure] shows the image sensor 400 moving back up so that line 424 coincides with the optical axis 404, and line 426 extends from the same point on the image sensor 400 to a distance D″ on the ground, rather than distance D. The distance D″ is longer than distance D, resulting in an overestimated distance estimate. By using a mapping that corresponds to the camera height H″ instead of H, this error can be reduced, minimized, or avoided.
[0085] In some embodiments, system 340 may be configured to provide a user interface through which the height of a camera including image sensor 400 can be specified. Based on this height specification, a corresponding mapping can be selected from mappings 316 to 326 for use in determining the estimated physical distance 336. Thus, when image sensor 400 is maintained at or near the specified height, Figure 3 system 340 can generate an accurate estimate of the physical distance. However, when image sensor 400 deviates from the specified height, the estimate of the physical distance may be incorrect, and the magnitude of the error may be proportional to the difference between the specified height and the actual height of the camera including image sensor 400.
[0086] In other embodiments, the camera may be equipped with a device configured to measure the height of the camera and thus the height of image sensor 400. For example, the camera may include a light emitter and a detector configured to allow measurement of the height based on the time of flight of light emitted by the light emitter, reflected from the ground, and detected by the light detector. An inertial measurement unit (IMU) can be used to verify that the measured distance is actually the height by detecting the orientation of the camera, the light emitter, and / or the light detector during the time of flight measurement. Specifically, the time of flight measurement result can indicate the height when the light is emitted in a direction parallel to the gravity vector detected by the IMU. Thus, a corresponding mapping can be selected from mappings 316 to 326 based on the height measurement. When a change in the height of the camera is detected, an updated mapping can be selected to keep the height presented by the mapping consistent with the actual height of the camera, thereby allowing accurate distance measurement.
[0087] Figure 4E The top portion shows image sensor 400 tilting upward from a zero pitch angle to a positive pitch angle , causing the optical axis 404 to pitch upward, as indicated by line 428. The height H of the aperture 402 (and the effective height of the camera) may not change due to the upward tilt. Without such an upward tilt, the bottom portion of the object 414 closest to image sensor 400 would create an image on image sensor 400 at a distance d above the center of image sensor 400, as Figure 4C and Figure 4D shown. However, the upward tilt causes the image to be created at a distance s′ (which is greater than d) above the center of image sensor 400, as indicated by line 430.
[0088] If the effect of the pitch angle on the position of the bottom of object 414 on image sensor 400 is not corrected, the distance projection model 314 may incorrectly determine that object 414 is located at distance S′, as Figure 4Eas indicated by the bottom portion of, rather than at distance D. Specifically, Figure 4E the bottom portion of shows the image sensor 400 tilting back downward such that line 428 coincides with the optical axis 404, and line 430 extends from the same point on the image sensor 400 to a distance S′ on the ground rather than distance D. The distance S′ is shorter than distance D, resulting in an underestimated distance estimate. By adding an estimation offset to the object bottom ratio determined for object 414, thereby shifting the object bottom ratio to the case when the pitch angle is zero, this error can be reduced, minimized, or avoided.
[0089] In addition, Figure 4F the top portion of shows the image sensor 400 tilting downward from zero pitch angle to a negative pitch angle α, causing the optical axis 404 to pitch downward as indicated by line 432. The height H of the aperture 402 does not change due to the downward tilt. Without this downward tilt, the bottom portion of the object 414 closest to the image sensor 400 would create an image on the image sensor 400 at a distance d above the center of the image sensor 400, as Figure 4C and Figure 4D shown. However, the downward tilt causes an image to be created instead at a distance s″ (which is less than d) above the center of the image sensor 400, as indicated by line 434.
[0090] If the effect of the pitch angle α on the position of the bottom of the object 414 on the image sensor 400 is not corrected, the distance projection model 314 may incorrectly determine that the object 414 is located at a distance S″, as Figure 4F indicated by the bottom portion of, rather than being located at distance D. Specifically, Figure 4F the bottom portion of shows the image sensor 400 tilting back upward such that line 432 coincides with the optical axis 404, and line 434 extends from the same point on the image sensor 400 to a distance S″ on the ground rather than distance D. The distance S″ is longer than distance D, resulting in an overestimated distance estimate. By adding an estimation offset to the object bottom ratio determined for object 414, thereby shifting the object bottom ratio to the case when the pitch angle α is zero, this error can be reduced, minimized, or avoided.
[0091] VI. Example Pitch Angle Compensation
[0092] Figure 5 shows an example method for compensating for a non - zero pitch angle of a camera. Specifically, Figure 5Shows image sensor 500 and aperture 502 of cameras respectively positioned in orientations 500A and 502A, with a pitch angle of zero such that optical axis 504A extends parallel to the ground in the environment (e.g., perpendicular to the gravity vector of the environment). Infinite reference line 520 is shown projected through image sensor 500 to illustrate the apparent change in the position of infinite reference line 520 relative to image sensor 500 when image sensor 500 is tilted downward from orientation 500A to orientation 500B and / or tilted upward from orientation 500A to orientation 500C. Additionally, in Figure 5 it is shown that infinite reference line 520 coincides with optical axis 504A and thus corresponds to an object bottom ratio of 0.5 when image sensor 500 is in orientation 500A. However, as the height of image sensor 500 changes while image sensor 500 remains in orientation 500A, infinite reference line 520 may deviate from optical axis 504A and may correspond to a different object bottom ratio (e.g., from 0.0 to 1.0, depending on height).
[0093] As image sensor 500 and aperture 502 are tilted upward from orientations 500A and 502A to orientations 500C and 502C respectively, infinite reference line 520 moves upward relative to image sensor 500. During this upward tilt, aperture 502 rotates within reference locus 506 and image sensor 500 moves along focus locus 508 with a radius equal to the focal length f of the camera. Thus, the upward tilt represents a positive change in the pitch angle while maintaining the height H of the camera relative to the ground constant. A positive pitch angle can be considered an elevation angle, while a negative pitch angle can be considered a depression angle.
[0094] When image sensor 500 is in orientation 500C, infinite reference line 520 coincides with the highest part of image sensor 500. Thus, in orientation 500C, infinite reference line 520 moves downward by screen ratio ΔL max仰角 within the corresponding image data 512. Within image data 512, infinite reference line 520 moves downward, not upward, because the image formed on image sensor 500 is upside down (i.e., inverted), and thus the output of image sensor 500 is inverted so that objects appear right side up when the image data is displayed. When infinite reference line 520 coincides with the middle of image sensor 500 when image sensor 500 is in orientation 500A, ΔL max仰角 can be equal to 0.5. However, depending on the height at which image sensor 500 is placed in the environment, ΔL max仰角 can take on other values.
[0095] Similarly, when the image sensor 500 and the aperture 502 are tilted downward from orientations 500A and 502A to orientations 500B and 502B respectively, the infinite reference line 520 moves downward relative to the image sensor 500. During this downward tilt, the aperture 502 rotates within the reference locus 506 and the image sensor 500 moves along the focus locus 508. Thus, the downward tilt represents a negative change in the pitch angle while maintaining the height H of the camera relative to the ground constant. When the image sensor 500 is in orientation 500B, the infinite reference line 520 coincides with the bottommost portion of the image sensor 500. Thus, in orientation 500B, the infinite reference line 520 moves up by the screen ratio ΔL within the corresponding image data 510. max俯角 Since the image formed on the image sensor 500 is inverted, the infinite reference line 520 moves up within the image data 510 instead of moving down.
[0096] When the infinite reference line 520 coincides with the middle of the image sensor 500 when the image sensor 500 is in orientation 500A, ΔL max俯角 can be equal to 0.5. However, depending on the height of the image sensor 500, ΔL max俯角 can take on other values. Regardless of the height at which the image sensor 500 is placed, the sum of ΔL max仰角 and ΔL max俯角 can be equal to 1.0.
[0097] The geometric model 514 shows the orientations 500B and 500C of the image sensor 500 and can be used to determine the mathematical relationships that can be used to compensate for changes in the pitch angle of the camera. Specifically, the geometric model 514 shows that orientation 500B corresponds to a negative pitch angle α max俯角 , the rotation of the optical axis 504A relative to orientation 504B, and the offset by the object bottom ratio associated with the infinite reference line 520 by ΔL max俯角 . Thus, tan(α max俯角 ) = ΔL max俯角 / f and f = ΔL max俯角 / tan(α max俯角 ). Thus, for a rotation of the pitch angle θ between α max俯角 and for the camera (i.e., the image sensor 500 and the aperture 502), the offset Δb of the object bottom ratio is given by Δb = f tan(θ) or equivalently by Δb = (ΔL max俯角 / tan(α max)) Modeling of tan(θ). The offset calculator 312 can use or implement this equation to determine an estimated offset Δb that compensates for the camera pitch 306 associated with the image data 300. Since the estimated offset Δb is calculated based on the object bottom ratio (rather than, for example, the number of pixels), the estimated offset Δb can be directly added to the object bottom ratio calculated by the object bottom ratio calculator 310.
[0098] Notably, for a symmetric camera, α max俯角 and can have the same magnitude, but can indicate different directions of camera pitch. Thus, the corresponding mathematical relationship can be based on rather than based on α as described above max俯角 to determine. Specifically, the orientation 500C corresponds to a positive pitch angle rotation of the optical axis 504A to the orientation 504C and an offset of the object bottom ratio associated with the infinite reference line 520 by ΔL max仰角 . Thus, and Thus, for α max俯角 associated with the image sensor 500 and the rotation of the pitch angle θ between, the offset Δb of the object bottom ratio is modeled by Δb = ftan(θ) or equivalently by . For a positive pitch angle, the estimated offset Δb can be positive (resulting in an increase in the object bottom ratio when summed with the offset), and for a negative pitch angle, the estimated offset Δb can be negative (resulting in a decrease in the object bottom ratio when summed with the offset).
[0099] α max俯角 and values can be determined empirically through a calibration process of the camera. During the calibration process, the camera can be tilted down or up until the infinite reference line 520 moves to the bottom or top of the image sensor 500, respectively, resulting in the offsets shown in the images 510 or 512. That is, calibration can be performed by placing the image sensor 500 and the aperture 502 in the orientations 500B and 502B, respectively, and measuring the values of α max俯角 and ΔL max俯角 , or by placing the image sensor 500 and the aperture 502 in the orientations 500C and 502C, respectively, and measuring the values of and ΔL max仰角 in these orientations. α max俯角 and The determined value may be valid for cameras with similar or substantially identical optical component arrangements, which include similar or substantially identical lenses, similar or substantially identical sensor sizes (i.e., length and width), similar or substantially identical focal lengths f, and / or similar or substantially identical aspect ratios of the generated image data. When one or more of these camera parameters are different, α max俯角 and the value can be re-determined empirically.
[0100] VII. Example Use Cases
[0101] Figure 6 Example use cases of the depth determination models, systems, devices, and techniques disclosed herein are shown. Specifically, Figure 6 it is shown that user 600 wears computing device 602 at approximately chest height. Computing device 602 may correspond to computing system 100 and / or computing device 200 and may include implementations of the camera and system 340. Computing device 602 may be hung around user 600's neck by a tether, cord, strap, or other connecting mechanism. Alternatively, computing device 602 may be connected to user 600's body at different locations and / or by different connecting mechanisms. Thus, when user 600 walks in the environment, computing device 602 and its camera may be positioned at a substantially fixed height above the ground of the environment (which allows for some height variations caused by the movement of user 600). Thus, distance projection model 314 may select the mapping corresponding to this substantially fixed height from mappings 316 to 326 and may thus be used to determine the distance to an object detected within the environment.
[0102] Specifically, the camera on computing device 602 may capture image data representing an environment including object 606 as indicated by field of view 604. Based on this image data, which may correspond to image data 300, Figure 3 system 340 of may be used to determine an estimated physical distance 336 between object 606 and computing device 602, its camera, and / or user 600. Based on estimated physical distance 336, computing device 602 may be configured to generate a representation of physical distance 336. The representation may be visual, auditory, and / or tactile, among other possibilities. Thus, the depth determination techniques discussed herein may be used to assist these users in navigating the environment, for example, by notifying visually impaired individuals of the distances to various objects in the environment.
[0103] For example, the computing device 602 may display an indication of an estimated physical distance 336 that is proximate to a display representation of the object 606 on its display, thereby indicating that the object 606 is horizontally separated from the computing device 602 by the estimated physical distance 336. In another example, the computing device 602 may generate an utterance representing the estimated physical distance 336 via one or more speakers. In some cases, the utterance may also indicate the classification of the object (e.g., human, vehicle, animal, stationary object, etc.) and / or the horizontal direction of the object 606 relative to the vertical centerline of the screen of the computing device 602. Thus, the utterance may be, for example, "box at 2 meters at the 1 o'clock direction", where the 1 o'clock direction uses clock positions to indicate a horizontal direction that is 30 degrees relative to the vertical centerline. In another example, a tactile representation of the estimated physical distance 336 may be generated via vibration of the computing device 602, where the pattern of the vibration encodes information about the distance and orientation of the object 606 relative to the user 600.
[0104] In addition, in some embodiments, the computing device 602 may allow the user 600 to designate a portion of the field of view 604 (i.e., a portion of the display of the computing device 602) as valid and another portion of the field of view 604 as invalid. Based on this designation, the computing device 602 may be configured to generate distance estimates for objects that are at least partially included within the valid portion of the field of view 604 and omit generating such distance estimates for objects that are at least partially not within the valid portion (i.e., objects that are entirely within the invalid portion of the field of view 604). For example, an individual with visual impairments may wish to use the computing device 602 to measure the distance to an object found in front of the user along an intended walking path, but may not be interested in the distance to an object found beside the walking path. Thus, such a user may designate a rectangular portion of the display of the computing device 602 having a height equal to the height of the display and a width less than the width of the display as valid, thereby causing the computing device 602 to ignore objects near the edges of the display represented in the image data.
[0105] Additionally, in some embodiments, the computing device 602 may allow the user 600 to specify the category or type of object for which the distance is to be measured. Based on this specification, the computing device 602 may be configured to generate distance estimates for objects classified into one of the specified categories or types and omit generating such distance estimates for objects not within the specified category or type. For example, an individual with visual impairments may wish to use the computing device 602 to measure the distance to moving objects such as other humans, vehicles, and animals, but may not be interested in the distance to non-moving objects such as benches, lamp posts, and / or mailboxes.
[0106] VIII. Additional Example Operations
[0107] Figure 7A flowchart showing operations related to estimating the distance between an object and a camera is presented. These operations can be performed by one or more of the computing system 100, the computing device 200, the system 340, and / or the computing device 602, and / or various other types of devices or device subsystems. Figure 7 Embodiments of Figure 7 can be simplified by removing any one or more of the features shown therein. Additionally, these embodiments can be combined with features, aspects, and / or embodiments shown in any previous figures or otherwise described herein.
[0108] Block 700 can involve receiving image data representing an object in the environment from the camera.
[0109] Block 702 can involve determining the vertical position of the bottom of the object within the image data based on the image data.
[0110] Block 704 can involve determining the object bottom ratio between the vertical position and the height of the image data.
[0111] Block 706 can involve determining an estimate of the physical distance between the camera and the object through a distance projection model and based on the object bottom ratio. The distance projection model can define a mapping between (i) each corresponding candidate object bottom ratio among a plurality of candidate object bottom ratios and (ii) the corresponding physical distance in the environment.
[0112] Block 708 can involve generating an indication of the estimate of the physical distance between the camera and the object.
[0113] In some embodiments, the mapping can be based on the assumption that the camera is set at a predetermined height within the environment.
[0114] In some embodiments, when the physical height of the camera is higher than the predetermined height when capturing the image data, the estimate of the physical distance between the camera and the object can be underestimated. When the physical height of the camera is lower than the predetermined height when capturing the image data, the estimate of the physical distance between the camera and the object can be overestimated.
[0115] In some embodiments, a specification of the predetermined height can be received through a user interface associated with the camera. Based on the specification of the predetermined height, the distance projection model can be configured by modifying the mapping to assume that the camera is positioned according to the specification of the predetermined height.
[0116] In some embodiments, configuring the distance projection model can include selecting a mapping from a plurality of candidate mappings based on the specification of the predetermined height. Each corresponding mapping among the plurality of candidate mappings can be associated with a corresponding specification of the predetermined height.
[0117] In some embodiments, the distance projection model may include a machine learning model. Configuring the distance projection model may include adjusting at least one input parameter of the machine learning model based on a specification of a predetermined height.
[0118] In some embodiments, the mapping may be based on a geometric model of the camera. The geometric model may include: (i) the camera having a focal length and being set at a predetermined height within the environment, (ii) the optical axis of the camera being oriented substantially parallel to the ground of the environment, and (iii) each respective line of a plurality of lines projecting from a respective point on the image sensor of the camera to a corresponding point on the ground of the environment. Each respective candidate object bottom ratio may be associated with a corresponding physical distance in the environment based on the geometric model.
[0119] In some embodiments, the mapping may include (i) a first mapping corresponding to the longitudinal orientation of the camera and (ii) a second mapping corresponding to the lateral orientation of the camera. Each respective candidate object bottom ratio associated with the first mapping may be between a corresponding vertical position within the longitudinal image data and the height of the longitudinal image data. Each respective candidate object bottom ratio associated with the second mapping may be between a corresponding vertical position within the lateral image data and the height of the lateral image data. The height of the image data may be determined based on the orientation of the camera when the image data is captured. Based on the orientation of the camera when the image data is captured, the first mapping or the second mapping may be selected for use in determining an estimate of the physical distance between the camera and the object.
[0120] In some embodiments, sensor data indicating the pitch angle of the camera may be obtained from one or more sensors associated with the camera. An estimated offset of the object bottom ratio may be determined based on the sensor data indicating the pitch angle of the camera. The estimated offset of the object bottom ratio may account for a change in the vertical position caused by the pitch angle of the camera relative to a zero pitch angle. The sum of the object bottom ratio and the estimated offset may be determined. The distance projection model may be configured to determine an estimate of the physical distance between the camera and the object based on the sum.
[0121] In some embodiments, determining the estimated offset of the object bottom ratio may include determining the product of the estimated focal length of the camera and the tangent of the pitch angle of the camera. A positive pitch angle associated with an upward tilt of the camera may result in an estimated offset having a positive value, such that the sum is higher than the object bottom ratio. A negative pitch angle associated with a downward tilt of the camera may result in an estimated offset having a negative value, such that the sum is lower than the object bottom ratio.
[0122] In some embodiments, estimating the focal length may be based on at least one of the following: (i) determining a maximum pitch angle that offsets an infinite reference line from an initial position on an image sensor of a camera to the top of the image sensor at a first screen ratio, or (ii) determining a minimum pitch angle that offsets the infinite reference line from the initial position on the image sensor to the bottom of the image sensor at a second screen ratio, where the sum of the first screen ratio and the second screen ratio equals one.
[0123] In some embodiments, the range of the bottom ratios of multiple candidate objects may range from (i) zero and a first candidate object bottom ratio corresponding to a minimum measurable physical distance to (ii) a second candidate object bottom ratio corresponding to a maximum measurable physical distance.
[0124] In some embodiments, determining a vertical position of the bottom of an object within image data may include determining a region of interest within the image data corresponding to the position of the object within the image data through one or more object detection algorithms. The bottom of the object may be identified through one or more object bottom detection algorithms and based on the region of interest. Based on identifying the bottom of the object, it may be determined that the bottom of the object is in contact with the ground of the environment. Based on determining that the bottom of the object is in contact with the ground, the vertical position of the bottom of the object within the image data may be determined.
[0125] In some embodiments, additional image data representing additional objects in the environment may be received from the camera. It may be determined that the bottom of the additional object is not visible within the image data based on the additional image data. It may be determined that an additional estimate of the physical distance between the camera and the additional object is below a predetermined value of the minimum measurable physical distance based on determining that the bottom of the additional object is not visible within the image data. An additional indication of the additional estimate of the physical distance between the camera and the additional object may be generated.
[0126] In some embodiments, generating an indication of an estimate of the physical distance between the camera and an object may include one or more of the following: (i) displaying a visual representation of the estimate of the physical distance on a display, (ii) generating an audible utterance representing the estimate of the physical distance, or (iii) generating a tactile representation of the estimate of the physical distance.
[0127] In some embodiments, an assignment of an effective portion of the field of view of the camera may be received. It may be determined that at least a portion of the object is included within the effective portion of the field of view of the camera. An indication of an estimate of the physical distance between the camera and the object may be generated based on determining that at least that portion of the object is included within the effective portion of the field of view of the camera. For an object outside the effective portion of the field of view of the camera, generating an indication of the corresponding estimate of the physical distance between the camera and the corresponding object may be omitted.
[0128] In some embodiments, a selection of one or more object categories from a plurality of object categories can be received. It can be determined that the object belongs to a first object category among the one or more object categories. An indication of an estimated physical distance between the camera and the object can be generated based on determining that the object belongs to the first object category. For an object that does not belong to the first object category, the corresponding indication of the estimated physical distance between the camera and the corresponding object can be omitted.
[0129] IX. Conclusion
[0130] The present disclosure is not limited to the specific embodiments described in this application, which are intended to be illustrative of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from its scope. In addition to the methods and apparatuses described herein, functionally equivalent methods and apparatuses within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.
[0131] The foregoing detailed description has described various features and operations of the disclosed systems, apparatuses, and methods with reference to the accompanying drawings. In the drawings, like symbols generally identify like components, unless the context indicates otherwise. The example embodiments described herein and in the drawings are not meant to be limiting. Other embodiments can be utilized and other changes can be made without departing from the scope of the subject presented herein. It will be readily understood that, as generally described herein and illustrated in the drawings, aspects of the present disclosure can be arranged, substituted, combined, separated, and designed in a variety of different configurations.
[0132] Regarding any and all message flowcharts, scenarios, and flowcharts in the drawings and as discussed herein, each step, block, and / or communication can represent the processing of information and / or the transmission of information according to example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages can be performed not in the order shown or discussed, including substantially simultaneously or in the reverse order, depending on the functions involved. Additionally, more or fewer blocks and / or operations can be used with any of the message flowcharts, scenarios, and flowcharts discussed herein, and these message flowcharts, scenarios, and flowcharts can be combined in part or in whole with each other. Figure 1 and these message flowcharts, scenarios, and flowcharts can be used in part or in whole with each other.
[0133] The steps or blocks representing the processing of information can correspond to circuitry that can be configured to perform the methods or techniques described herein. Alternatively or additionally, the blocks representing the processing of information can correspond to modules, segments, or portions of program code (including related data). The program code can include one or more instructions executable by a processor to implement specific logical operations or actions in the method or technique. The program code and / or related data can be stored on any type of computer-readable medium, such as a storage device including random access memory (RAM), a disk drive, a solid state drive, or other storage media.
[0134] The computer-readable medium can also include non-transitory computer-readable media, such as computer-readable media that stores data for a short period of time, like register memory, processor cache, and RAM. The computer-readable medium can also include non-transitory computer-readable media that stores program code and / or data for a longer period of time. Thus, the computer-readable medium can include secondary or permanent long-term storage devices, such as read-only memory (ROM), optical disks or magnetic disks, solid state drives, compact disc read-only memory (CD-ROM). The computer-readable medium can also be any other volatile or non-volatile storage system. For example, the computer-readable medium can be considered a computer-readable storage medium, or a tangible storage device.
[0135] In addition, the steps or blocks representing the transmission of one or more information can correspond to the transmission of information between software and / or hardware modules in the same physical device. However, other information transmissions can occur between software modules and / or hardware modules in different physical devices.
[0136] The particular arrangements shown in the figures should not be considered restrictive. It should be understood that other embodiments can include more or fewer of each element shown in a given figure. Additionally, some of the elements shown can be combined or omitted. Moreover, example embodiments can include elements not shown in the figures.
[0137] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and are not intended to be restrictive, and their true scope is indicated by the appended claims.
Claims
1. A computer-implemented method, comprising: receiving image data representing an object in an environment from a camera; determining a vertical position of a bottom of the object within the image data based on the image data; determining an object bottom ratio between the vertical position and a height of the image data; determining an estimate of a physical distance between the camera and the object through a distance projection model and based on the object bottom ratio, wherein the distance projection model defines a mapping between (i) each respective candidate object bottom ratio among a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment; and generating an indication of the estimate of the physical distance between the camera and the object.
2. The computer-implemented method according to claim 1, wherein, the mapping is based on an assumption that the camera is positioned at a predetermined height within the environment.
3. The computer-implemented method according to claim 2, wherein, when a physical height of the camera is higher than the predetermined height when capturing the image data, the estimate of the physical distance between the camera and the object is an underestimate, and wherein when the physical height of the camera is lower than the predetermined height when capturing the image data, the estimate of the physical distance between the camera and the object is an overestimate.
4. The computer-implemented method according to claim 2, further comprising: receiving a specification of the predetermined height through a user interface associated with the camera; and configuring the distance projection model based on the specification of the predetermined height by modifying the mapping to assume that the camera is positioned according to the specification of the predetermined height.
5. The computer-implemented method according to claim 4, wherein, configuring the distance projection model comprises: selecting the mapping from a plurality of candidate mappings based on the specification of the predetermined height, wherein each respective mapping among the plurality of candidate mappings is associated with a corresponding specification of the predetermined height.
6. The computer-implemented method according to claim 4, wherein, the distance projection model comprises a machine learning model, and wherein configuring the distance projection model comprises: adjusting at least one input parameter of the machine learning model based on the specification of the predetermined height.
7. The computer-implemented method according to claim 2, wherein, the mapping is based on a geometric model of the camera, wherein the geometric model comprises: (i) the camera has a focal length and is set at the predetermined height within the environment, (ii) an optical axis of the camera is oriented substantially parallel to a ground of the environment, and (iii) each respective line among a plurality of lines projects from a corresponding point on an image sensor of the camera to a corresponding point on the ground of the environment, and wherein each respective candidate object bottom ratio is associated with the corresponding physical distance in the environment based on the geometric model.
8. The computer-implemented method according to claim 1, wherein, The mapping includes (i) a first mapping corresponding to the longitudinal orientation of the camera and (ii) a second mapping corresponding to the lateral orientation of the camera, wherein each respective candidate object bottom ratio associated with the first mapping is between a corresponding vertical position within the longitudinal image data and the height of the longitudinal image data, and wherein each respective candidate object bottom ratio associated with the second mapping is between a corresponding vertical position within the lateral image data and the height of the lateral image data, and wherein the method further comprises: determining the height of the image data based on the orientation of the camera when the image data is captured; and selecting the first mapping or the second mapping based on the orientation of the camera when the image data is captured for use in determining the estimate of the physical distance between the camera and the object.
9. The computer-implemented method according to claim 1, further comprising: obtaining sensor data indicative of the pitch angle of the camera from one or more sensors associated with the camera; determining an estimated offset of the object bottom ratio based on the sensor data indicative of the pitch angle of the camera, the estimated offset taking into account the change in the vertical position caused by the pitch angle of the camera relative to a zero pitch angle; and determining the sum of the object bottom ratio and the estimated offset, wherein the distance projection model is configured to determine the estimate of the physical distance between the camera and the object based on the sum.
10. The computer-implemented method according to claim 9, wherein determining the estimated offset of the object bottom ratio comprises: determining the product of the estimated focal length of the camera and the tangent of the pitch angle of the camera, wherein a positive pitch angle associated with an upward tilt of the camera results in an estimated offset having a positive value such that the sum is higher than the object bottom ratio, and wherein a negative pitch angle associated with a downward tilt of the camera results in an estimated offset having a negative value such that the sum is lower than the object bottom ratio.
11. The computer-implemented method according to claim 10, wherein the estimated focal length is based on at least one of the following: (i) determining the maximum pitch angle that offsets an infinite reference line from an initial position on the image sensor of the camera to the top of the image sensor at a first screen ratio, or (ii) determining the minimum pitch angle that offsets the infinite reference line from the initial position on the image sensor to the bottom of the image sensor at a second screen ratio, wherein the sum of the first screen ratio and the second screen ratio equals one.
12. The computer-implemented method according to claim 1, wherein the range of the plurality of candidate object bottom ratios is from (i) a first candidate object bottom ratio of zero and corresponding to a minimum measurable physical distance to (ii) a second candidate object bottom ratio corresponding to a maximum measurable physical distance.
13. The computer-implemented method according to claim 1, wherein Determining the vertical position of the bottom of the object within the image data includes: Determining a region of interest within the image data corresponding to the position of the object within the image data by one or more object detection algorithms; Identifying the bottom of the object by one or more object bottom detection algorithms and based on the region of interest; Determining that the bottom of the object is in contact with the ground of the environment based on identifying the bottom of the object; and Determining the vertical position of the bottom of the object within the image data based on determining that the bottom of the object is in contact with the ground.
14. The computer-implemented method according to claim 1, further including: Receiving additional image data from the camera representing additional objects in the environment; Determining that the bottom of the additional object is not visible within the image data based on the additional image data; Determining that an additional estimate of the physical distance between the camera and the additional object is below a predetermined value of a minimum measurable physical distance based on determining that the bottom of the additional object is not visible within the image data; and Generating an additional indication of the additional estimate of the physical distance between the camera and the additional object.
15. The computer-implemented method according to claim 1, wherein Generating the indication of the estimate of the physical distance between the camera and the object includes one or more of the following: (i) displaying a visual representation of the estimate of the physical distance on a display, (ii) generating an audible utterance representing the estimate of the physical distance, or (iii) generating a tactile representation of the estimate of the physical distance.
16. The computer-implemented method according to claim 1, further including: Receiving an assignment of an effective portion of the field of view of the camera; Determining that at least a portion of the object is included within the effective portion of the field of view of the camera; and Generating the indication of the estimate of the physical distance between the camera and the object based on determining that at least the portion of the object is included within the effective portion of the field of view of the camera, wherein, for an object outside the effective portion of the field of view of the camera, an indication of the corresponding estimate of the physical distance between the camera and the corresponding object is omitted.
17. The computer-implemented method according to claim 1, further including: Receiving a selection of one or more object categories from a plurality of object categories; Determining that the object belongs to a first object category among the one or more object categories; and Generating the indication of the estimate of the physical distance between the camera and the object based on determining that the object belongs to the first object category, wherein, for an object that does not belong to the first object category, an indication of the corresponding estimate of the physical distance between the camera and the corresponding object is omitted.
18. A computing system, comprising: A camera; A processor; and A non - transitory computer - readable storage medium having instructions stored thereon, which when executed by the processor, cause the processor to perform operations, the operations including: Receiving, from the camera, image data representing an object in the environment; Determining a vertical position of the bottom of the object within the image data based on the image data; Determining an object bottom ratio between the vertical position and the height of the image data; Determining an estimate of the physical distance between the camera and the object through a distance projection model and based on the object bottom ratio, wherein the distance projection model defines a mapping between (i) each respective candidate object bottom ratio of a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment; and Generating an indication of the estimate of the physical distance between the camera and the object.
19. The computing system according to claim 18, further comprising: One or more sensors configured to generate sensor data indicating a pitch angle of the camera, wherein the operations further include: Obtaining the sensor data indicating the pitch angle of the camera from the one or more sensors; Determining an estimated offset of the object bottom ratio based on the sensor data indicating the pitch angle of the camera, the estimated offset taking into account a change in the vertical position caused by the pitch angle of the camera relative to a zero pitch angle; and Determining a sum of the object bottom ratio and the estimated offset, wherein the distance projection model is configured to determine the estimate of the physical distance between the camera and the object based on the sum.
20. A non - transitory computer - readable storage medium having instructions stored thereon, which when executed by a computing system, cause the computing system to perform operations, the operations including: Receiving, from a camera, image data representing an object in the environment; Determining a vertical position of the bottom of the object within the image data based on the image data; Determining an object bottom ratio between the vertical position and the height of the image data; Determining an estimate of the physical distance between the camera and the object through a distance projection model and based on the object bottom ratio, wherein the distance projection model defines a mapping between (i) each respective candidate object bottom ratio of a plurality of candidate object bottom ratios and (ii) a corresponding physical distance in the environment; and Generating an indication of the estimate of the physical distance between the camera and the object.
Citation Information
Patent Citations
Imaging device parameter estimation method and method, device and system using imaging device parameter estimation method
CN102447942A
Automatic scaling of objects based on depth map for image editing
CN105741232A