Attitude Estimation in Three-Dimensional Space
Real-time sparse pose estimation using a rolling shutter camera improves VR and AR imaging stability and comfort by reducing latency in pose estimation, addressing delays in existing technologies.
Patent Information
- Application Number
- JP2023184793
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-06-30
- Filing Date
- 2023-10-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2037-05-17
AI Technical Summary
Existing VR and AR technologies face challenges in generating comfortable, natural, and rich presentations of virtual image elements among real-world elements due to delays in sparse pose estimation, leading to unstable imaging and user discomfort.
Implementing sparse pose estimation by capturing and processing image segments in real-time using a rolling shutter camera to identify sparse points, reducing latency and improving pose estimation accuracy.
Enhances the user experience by minimizing latency and improving the stability of VR and AR imaging, reducing eye fatigue and providing a more pleasant viewing experience.
Smart Images

Figure 0007708832000001 
Figure 0007708832000002 
Figure 0007708832000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 62 / 357,285, filed on June 30, 2016, entitled "ESTIMATING POSE IN 3D SPACE", the contents of which are hereby incorporated by reference in their entirety into this specification.
[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly, to sparse pose estimation in three - dimensional (3D) space.
Background Art
[0003] Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality" or "augmented reality" experiences, where digitally reproduced images or portions thereof are presented to a user in a manner that appears or can be perceived as being real. A virtual reality or "VR" scenario typically involves the presentation of digital or virtual image information without transparency to other actual real-world visual inputs, while an augmented reality or "AR" scenario typically involves the presentation of digital or virtual image information as an augmentation to the visualization of the actual world around the user. For example, referring to FIG. 1, an augmented reality scene 1000 is depicted, where a user of AR technology can see a real-world park-like setting 1100 featuring people, trees, and buildings in the background, and a concrete platform 1120. In addition to these items, a user of AR technology can also "see" virtual content such as a robotic image 1110 standing on the real-world platform 1120 and an avatar character 1130 in the form of a flying cartoon that appears anthropomorphic like a honeybee, even though these elements do not exist within the real world. In summary, the human visual perception system is very complex, and the generation of VR or AR technology that facilitates a comfortable, natural, and rich presentation of virtual image elements among other virtual or real-world image elements is difficult. The systems and methods disclosed herein address various challenges associated with VR and AR technologies. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0004] One aspect of the present disclosure provides sparse pose estimation that is performed as sparse points are captured within an image frame by an image capture device. Thus, the sparse pose estimation can be performed before the entire image frame is captured. In some embodiments, the sparse pose estimation can be refined or updated as the image frame is captured.
[0005] In some embodiments, a system, device, and method for estimating the position of an image capture device within an environment are disclosed. In some implementations, the method may include the step of continuously receiving a first group of a plurality of image segments. The first plurality of image segments may form at least a portion of an image representing the field of view (FOV) from the front of the image capture device, which may include a portion of the environment surrounding the image capture device and a plurality of sparse points. Each sparse point may correspond to a subset of the image segments. The method may also include the step of identifying a first group of sparse points, which may include one or more sparse points that are identified as the first group of a plurality of image segments are received. The method may then include the step of determining, by a position estimation system, the position of the image capture device within the environment based on the first group of sparse points. The method may also include the step of continuously receiving a second group of a plurality of image segments, which may be received from the first group of a plurality of image segments and may form at least another portion of the image. The method may then include the step of identifying a second group of sparse points, which may include one or more sparse points that are identified as the second group of a plurality of image segments are received. The method may then update, by the position estimation system, the position of the image capture device within the environment based on the first and second groups of sparse points.
[0006] In some embodiments, systems, devices, and methods for estimating the position of an image capture device within an environment are disclosed. In some implementations, the method may include the step of continuously receiving a plurality of image segments, which may form an image representing the field of view (FOV) from the front of the image capture device. The FOV may include a portion of the environment surrounding the image capture device and may include a plurality of sparse points. Each sparse point may be identifiable, at least in part, based on a corresponding subset of the plurality of image segments. The method may also include the step of continuously identifying one or more of the plurality of sparse points when each subset of the image segments corresponding to the one or more sparse points is received. The method may then include the step of estimating the position of the image capture device within the environment based on the identified one or more sparse points.
[0007] In some embodiments, systems, devices, and methods for estimating the position of an image capture device within an environment are disclosed. In some implementations, the image capture device may include an image sensor configured to capture an image. The image may be captured via continuously capturing a plurality of image segments representing the field of view (FOV) of the image capture device. The FOV may include a portion of the environment surrounding the image capture device and a plurality of sparse points. Each sparse point may be identifiable, at least in part, based on a corresponding subset of the plurality of image segments. The image capture device may also include a memory circuit configured to store a subset of the image segments corresponding to the one or more sparse points and a computer processor operatively coupled to the memory circuit. The computer processor may be configured to continuously identify one or more of the plurality of sparse points when each subset of the image segments corresponding to the one or more sparse points is received by the image capture device. The computer processor may also be configured to extract the continuously identified one or more sparse points in order to estimate the position of the image capture device within the environment based on the identified one or more sparse points.
[0008] In some embodiments, a system, device, and method for estimating the position of an image capture device within an environment are disclosed. In some implementations, an augmented reality system is disclosed. The augmented reality system may include an imaging device facing outward, computer hardware, and a processor operably coupled to the computer hardware and the imaging device facing outward. The processor may be configured to execute instructions for implementing at least a portion of the methods disclosed herein.
[0009] In some embodiments, a system, device, and method for estimating the position of an image capture device within an environment are disclosed. In some implementations, an autonomous entity is disclosed. The autonomous entity may include an imaging device facing outward, computer hardware, and a processor operably coupled to the computer hardware and the imaging device facing outward. The processor may be configured to execute instructions for implementing at least a portion of the methods disclosed herein.
[0010] In some embodiments, a system, device, and method for estimating the position of an image capture device within an environment are disclosed. In some implementations, a robotic system is disclosed. The robotic system may include an imaging device facing outward, computer hardware, and a processor operably coupled to the computer hardware and the imaging device facing outward. The processor may be configured to execute instructions for implementing at least a portion of the methods disclosed herein.
[0011] The various implementations of the methods and apparatuses within the scope of the appended claims each have several aspects, and no single one of them alone bears the desirable attributes described herein without limiting the scope of the appended claims. Without limiting the scope of the appended claims, some prominent features are described herein.
[0012] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. Neither this summary nor any form of the embodiments for carrying out the invention set forth below purports to define or limit the scope of the subject matter of the invention. This specification also provides, for example, the following items. (Item 1) An imaging system, An image capture device including a lens and an image sensor, wherein the lens is configured to direct light from an environment surrounding the image capture device to the image sensor, and the image sensor is configured to continuously capture a first plurality of image segments of an image based on light from the environment, the image representing a field of view (FOV) of the image capture device, the FOV constituting a portion of the environment and including a plurality of sparse points, is configured to continuously capture a second plurality of image segments, the second plurality of image segments being captured after the first plurality of image segments and forming at least another portion of the image, and an image capture device configured to perform; a non-transitory data storage device configured to continuously receive the first and second pluralities of image segments from the image sensor and store instructions for estimating at least one of a position and an orientation of the image capture device within the environment; at least one hardware processor operably coupled to the non-transitory data storage device, the at least one hardware processor is configured to, in part, identify a first group of sparse points based on a corresponding subset of the first plurality of image segments, the first group of sparse points being identified as the first plurality of image segments are received in the non-transitory data storage device, Determining at least one of the position and orientation of the imaging device in the environment based on the sparse points of the first group; Identifying a second group of sparse points, at least in part, based on corresponding subsets of the second plurality of image segments, wherein the second group of sparse points is identified as the second plurality of image segments are received in the non-transitory data storage device; Updating at least one of the position and orientation of the imaging device in the environment based on the sparse points of the first and second groups; At least one hardware processor configured by instructions for performing the above; An imaging system comprising the above. (Item 2) The imaging system according to item 1, wherein the image sensor is a rolling shutter image sensor. (Item 3) The non-transitory data storage device includes a non-transitory buffer storage device configured to continuously receive the first and second pluralities of image segments as the image segments are captured by the image sensor, and the non-transitory buffer storage device has a storage capacity based at least in part on the number of image segments included in each subset of the image segments. The imaging system according to item 1. (Item 4) The imaging system according to item 1, wherein the sparse points of the first group or the sparse points of the second group include a number of sparse points from 10 to 20. (Item 5) The hardware processor is configured to update at least one of the position and orientation of the image capture device based on the number of recently identified sparse points, and the recently identified sparse points include at least one of the sparse points of the first group, the sparse points of the second group, and one or more of the sparse points of the first and second groups. The imaging system according to item 1. (Item 6) The imaging system according to item 5, wherein the number of the most recently identified sparse points is equal to the number of sparse points in the sparse points of the first group. (Item 7) The imaging system according to item 1, wherein the hardware processor is configured to implement a visual simultaneous localization and mapping (V-SLAM) algorithm. (Item 8) The imaging system according to item 1, wherein the plurality of sparse points are identified based on at least one of a real-world object, a virtual image element, and an invisible indicator projected into the environment. (Item 9) A head-mounted display (HMD) configured to be worn on a user's head, the HMD comprising: A frame; A display supported by the frame and disposed in front of the user's eyes; An outward-facing image capture device disposed on the frame and including a lens and an image sensor, the lens being configured to direct light from an environment surrounding the HMD to the image sensor, the image sensor being configured to continuously capture a plurality of image segments of an image based on the light from the environment, the image representing a field of view (FOV) of the outward-facing image capture device, the FOV constituting a part of the environment and including a plurality of sparse points, each sparse point being at least partially identifiable based on a corresponding subset of the plurality of image segments; A non-transitory data storage device configured to continuously receive the plurality of image segments from the image sensor and store instructions for estimating at least one of a position and an orientation of the HMD within the environment; At least one hardware processor operably coupled to the non-transitory data storage device, the at least one hardware processor comprising: When each subset of the image segments corresponding to the one or more sparse points is received in the non-transitory data storage device, successively identifying one or more of the plurality of sparse points; estimating at least one of the position and orientation of the HMD in the environment based on the one or more identified sparse points; at least one hardware processor configured by instructions for performing; An HMD including. (Item 10) The non-transitory data storage device includes a circular buffer or a rolling buffer, the HMD according to item 9. (Item 11) The plurality of image segments includes at least a first plurality of image segments and a second plurality of image segments, and the image sensor is configured to continuously transmit the first and second image segments to the non-transitory data storage device, the HMD according to item 9. (Item 12) The hardware processor, when a first plurality of image segments corresponding to one or more sparse points of a first group are received, successively identifying one or more of the sparse points of the first group; when a second plurality of image segments corresponding to one or more sparse points of a second group are received, successively identifying one or more of the sparse points of the second group; configured to perform, and the second plurality of image segments are received after the first plurality of image segments, the HMD according to item 11. (Item 13) The hardware processor is configured to estimate at least one of the position and orientation of the HMD based on the one or more identified sparse points of the first group, the HMD according to item 12. (Item 14) The sparse points of the first group or the sparse points of the second group include a number of 2 to 20 sparse points, the HMD according to item 13. (Item 15) The HMD according to item 13, wherein the sparse points of the first group or the second group include a number of sparse points of 10 to 20. (Item 16) The HMD according to item 13, wherein the hardware processor is further configured to update at least one of the position and orientation of the HMD based on one or more sparse points of the identified second group. (Item 17) The HMD according to item 9, wherein the hardware processor is further configured to update at least one of the position and orientation of the HMD when the number of the continuously identified one or more sparse points is identified. (Item 18) The HMD according to item 17, wherein the number of the continuously identified one or more sparse points includes at least one of the sparse points of the one or more sparse points of the first group. (Item 19) The HMD according to item 9, wherein the plurality of sparse points are identified based on at least one of a real-world object, a virtual image element, and an invisible indicator projected into the environment. (Item 20) The hardware processor is further configured to extract the continuously identified one or more sparse points from corresponding subsets of the plurality of image segments, and perform a visual simultaneous localization and mapping (VSLAM) algorithm on the continuously identified one or more sparse points to estimate at least one of the position and orientation of the image capture device The HMD according to item 9, which is configured to perform the above operations.
Brief Description of the Drawings
[0013]
Figure 1
[0014]
Figure 2
[0015]
Figure 3
[0016]
Figure 4
[0017]
Figure 5
[0018]
Figure 6
[0019]
Figure 7
[0020]
Figure 8
[0021]
Figure 9
[0022]
Figure 10
[0023] Throughout the drawings, reference numerals may be reused to indicate correspondence between referenced elements. The provided drawings are not to scale and are provided to illustrate exemplary embodiments described herein and are not intended to limit the scope of the disclosure.
Best Mode for Carrying Out the Invention
[0024] (Overview) With the use of an AR device or other device moving within a three-dimensional (3D) space, the device may need to track its movement through the 3D space and map the 3D space. For example, the AR device may move around in the 3D space either due to the movement of the user or independently of the user (e.g., a robot or other autonomous entity), and map the 3D space and determine one or more of the location, position, or orientation of the device within the 3D space for subsequent processing, for example, to facilitate the display of virtual image elements, especially among virtual and real-world image elements. For example, in order to accurately present virtual and real-world image elements, the device may need to know the location and its orientation within the real world and accurately render the virtual image at a specific location with a specific orientation within the real-world space. In another embodiment, it may be desirable to reproduce the trajectory of the device through the 3D space. Therefore, as the device moves around in the 3D space, it may be desirable to determine in real time the position, location, or orientation of the device within the 3D space (hereinafter collectively referred to as "pose"). In some implementations, sparse pose estimation within the 3D space may be determined, for example, from a continuous stream of image frames from an imaging device included as part of the AR device. Each image frame of the continuous stream may be stored for processing and for inclusion in the sparse pose estimation and to estimate the pose of the device therefrom. However, these techniques can introduce latency in estimating the pose due to the overall transfer of each frame to memory for subsequent processing.
[0025] The present disclosure provides an exemplary device and method configured to estimate the pose of a device (e.g., an autonomous device such as an AR device or a robot) in a 3D space. As an example, the device may perform sparse pose estimation based on receiving a plurality of image frames and estimating the pose of the device from each image frame as the device moves through the 3D space. Each image frame may represent a portion of the 3D space in front of the device indicating the position of the device in the 3D space. In some embodiments, each image frame may include one or more of features or objects that may be represented by sparse points, keypoints, point clouds, or other types of mathematical representations. For each image frame, the image frame may be captured by successively receiving a plurality of image segments that, when combined, constitute the entire image frame. From there, the device may be configured to identify sparse points within the image frame in response to receiving the image segments that include each sparse point. The device may extract a first group of sparse points that includes one or more sparse points. The first group of sparse points may be at least one input to the sparse pose estimation process. Subsequently, the device may identify and extract a second group of sparse points and update the sparse pose estimation based on the second group. In one exemplary implementation, the first group of sparse points may be utilized to estimate the pose of the device prior to the identification of subsequent sparse points (e.g., the second group of sparse points). The subsequent sparse points may become available for use in updating the sparse pose estimation as they are identified.
[0026] Embodiments of methods, devices, and systems are described herein with reference to AR devices, which is not intended to limit the scope of the disclosure. The methods and devices described herein are not limited to AR devices or head-mounted devices. Other devices are also conceivable (e.g., mobile robots, digital cameras, autonomous entities, etc.). Suitable devices include, but are not limited to, those that are movable through 3D space, independently or with user intervention. For example, the methods described herein may be applied to objects moving through 3D space that are tracked by a remote camera. In some embodiments, the processing may also be performed remotely from the object.
[0027] (Exemplary AR Device for Moving within 3D Space) It is desirable for a 3D display to map the real world surrounding the display and reproduce the trajectory of the display through 3D space in order to facilitate a rich presentation that is comfortable and natural, especially for virtual image elements among virtual or real-world image elements. For example, a sparse pose estimation process may be performed to determine a map of the 3D space. If the sparse pose estimation is not performed in real time with minimal latency, the user may experience unstable imaging, harmful eye fatigue, headaches, and generally an unpleasant VR and AR viewing experience. Thus, various embodiments described herein are configured to determine or estimate one or more of the position, location, or orientation of an AR device.
[0028] Figure 2 illustrates an embodiment of a wearable display system 100. The display system 100 includes a display 62 and various mechanical and electronic modules and systems to support the functionality of the display 62. The display 62 may be coupled to a frame 64, which is wearable by a display system user, wearer, or viewer 60 and is configured to position the display 62 in front of the eyes of the viewer 60. The display system 100 can comprise a head-mounted display (HMD) worn on the head of the wearer. An augmented reality display (ARD) can include the wearable display system 100. In some embodiments, a speaker 66 is coupled to the frame 64 and positioned adjacent to the outer ear canal of the user (in some embodiments, another speaker, not shown, is positioned adjacent to the other outer ear canal of the user and may provide stereo / formable sound control). The display system 100 can include one or more outward-facing imaging systems 110 that observe the world (e.g., 3D space) in the environment around the wearer. The display 62 can be operably coupled to a local processing and data module 70 by a communication link 68 such as a wired conductor or wireless connectivity, which can be mounted in various configurations, such as fixed to the frame 64, fixed to a helmet or hat worn by the user, built into headphones, or removably attached to the user 60 in another manner (e.g., in a backpack-style configuration, a belt-coupled configuration).
[0029] The display system 100 may include one or more outward-facing imaging systems 110a or 110b (individually or collectively, hereinafter referred to as "110") disposed on the frame 64. In some embodiments, the outward-facing imaging system 110a can be disposed at a substantially central portion of the frame 64 between the user's eyes. In another embodiment, alternatively or in combination, the outward-facing imaging system 110b can be disposed on one or more sides of the frame adjacent to one or both of the user's eyes. For example, the outward-facing imaging system 110b may be located on both the left and right sides of the user adjacent to both eyes. Exemplary arrangements of the outward-facing camera 110 are provided above, but other configurations are also possible. For example, the outward-facing imaging system 110 may be positioned in any orientation or position relative to the display system 100.
[0030] In some embodiments, the outward-facing imaging system 110 captures an image of a portion of the world in front of the display system 100. The entire region available for viewing or imaging by the viewer may be referred to as the field of regard (FOR). In some implementations, the FOR may include substantially all of the solid angle around the display system 100 (front, back, above, below, or sides of the wearer) since the display can rotate around the environment and image objects surrounding the display. The portion of the FOR in front of the display system may be referred to as the field of view (FOV), and the outward-facing imaging system 110 is sometimes also referred to as an FOV camera. The image obtained from the outward-facing imaging system 110 can be used to identify occlusions in the environment and to estimate poses for use in processes such as an occlusion pose estimation process.
[0031] In some implementations, the outward-facing imaging system 110 may be configured as a digital camera that includes an optical lens system and an image sensor. For example, light from the world in front of the display 62 (e.g., the FOV) may be focused onto the image sensor by the lens of the outward-facing imaging system 110. In some embodiments, the outward-facing imaging system 100 may be configured to operate within the infrared (IR) spectrum, the visible light spectrum, or any other suitable wavelength range or range of wavelengths of electromagnetic radiation. In some embodiments, the imaging sensor may be configured as either a CMOS (complementary metal oxide semiconductor) or a CCD (charge coupled device) sensor. In some embodiments, the image sensor may be configured to detect light within the IR spectrum, the visible light spectrum, or any other suitable wavelength range or range of wavelengths of electromagnetic radiation. In some embodiments, the frame rate of the digital camera may be related to the rate at which image data can be transmitted from the digital camera to a memory or storage unit (e.g., the local processing and data module 70). For example, if the frame rate of the digital camera is 30 Hertz, the data captured by the pixels of the image sensor may be read into memory (e.g., clocked off) every 30 milliseconds. Thus, the frame rate of the digital camera may impose a delay on the storage and subsequent processing of the image data.
[0032] In some embodiments, when the outward-facing imaging system 110 is a digital camera, the outward-facing imaging system 110 may be configured as a global shutter camera or a rolling shutter (also referred to as a progressive scan camera, for example). For example, when the outward-facing imaging system 110 is a global shutter camera, the image sensor may be a CCD sensor configured to capture the entire image frame representing the FOV in front of the display 62 in a single operation. The entire image frame may then be loaded for processing, for example, to perform sparse pose estimation as described herein, into the local processing and data module 70. Thus, in some embodiments, the utilization of the entire image frame may introduce a delay in pose estimation, for example, due to the frame rate and the latency in storing the image as described above. For example, a global shutter digital camera having a 30 hertz frame rate may introduce a 30 millisecond delay before any pose estimation can be performed.
[0033] In other embodiments, when the outward-facing imaging system 110 is configured as a rolling shutter camera, the image sensor may be a CMOS sensor configured to successively capture a plurality of image segments, scan across the scene, and transmit the image data of the captured image segments. When the image segments are combined in the order in which they are captured, they constitute an image frame of the FOV of the outward-facing imaging system 110. In some embodiments, the scanning direction may be horizontal; for example, the outward-facing imaging system 110 may capture a plurality of vertically adjacent image segments horizontally in a leftward or rightward direction. In another embodiment, the scanning direction may be vertical; for example, the outward-facing imaging system 110 may capture a plurality of horizontally adjacent image segments vertically in an upward or downward direction. Each image segment may be successively loaded into the local processing and data module 70 as the individual image segment is captured by the image sensor. Thus, in some embodiments, as described above, the latency due to the frame rate of the digital camera may be reduced or minimized by successively transmitting the image segments as they are captured by the digital camera.
[0034] The local processing and data module 70 may comprise one or more hardware processors and digital memory such as non-volatile memory (e.g., flash memory), both of which may be utilized to assist in the processing, buffering, caching, and storage of data. The data may include a) data captured from sensors such as an image capture device (e.g., the outward-facing imaging system 110), a microphone, an inertial measurement unit (IMU), an accelerometer, a compass, a global positioning system (GPS) unit, a wireless device, and / or a gyroscope (e.g., operably coupled to the frame 64 or otherwise attached to the user 60), and / or b) data that may potentially be obtained and / or processed using the remote processing module 72 and / or the remote data repository 74 for passage to the display 62 after processing or retrieval. The local processing and data module 70 may be operably coupled to the remote processing module 72 and / or the remote data repository 74 via communication links 76 and / or 78, via a wired or wireless communication link or the like, such that these remote modules are available as resources to the local processing and data module 71. Additionally, the remote processing module 72 and the remote data repository 74 may be operably coupled to each other. In some embodiments, the local processing and data module 70 may be operably connected to one or more of an image capture device, a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope. In some other embodiments, one or more of these sensors may be attached to the frame 64 or may be of a stand-alone configuration that communicates with the local processing and data module 70 via a wired or wireless communication path.
[0035] In some embodiments, the local processing and digital memory of data module 70, or a portion thereof, may be configured to store one or more elements of data for a temporary period (e.g., as a non-transitory buffer storage device). For example, the digital memory may be configured to receive a portion or all of the data while the data is being transferred between processes of the local processing and data module 70 and store a portion or all of the data for a short period of time. In some implementations, a portion of the digital memory may be configured as a buffer that continuously receives from the imaging system 110 facing outwards one or more image segments. Thus, the buffer may be a non-transitory data buffer configured to store a set number of image segments prior to the image segments being transmitted to the local processing and data module 70 (or removing the data repository 74) for permanent storage or subsequent processing (as described below with reference to FIGS. 9A and 9B).
[0036] In some embodiments, the remote processing module 72 may comprise one or more hardware processors configured to analyze and process data and / or image information. In some embodiments, the remote data repository 74 may comprise a digital data storage facility, which may be available through the Internet or other networking configurations in a “cloud” resource configuration. In some embodiments, the remote data repository 74 may include one or more remote servers that provide information, e.g., information for generating augmented reality content, to the local processing and data module 70 and / or the remote processing module 72. In some embodiments, all data is stored and all calculations are performed in the local processing and data module 70, enabling complete autonomous use from the remote module.
[0037] Exemplary AR devices are described herein, but it will be understood that the methods and devices disclosed herein are not limited to AR devices or head-mounted devices. For example, other configurations such as mobile robots, digital cameras, autonomous entities, etc. are also conceivable as possibilities. Applicable devices include, but are not limited to, those that can move through 3D space, independently, or with user intervention.
[0038] (Exemplary Trajectory of an AR Device Through 3D Space) FIG. 3 schematically illustrates an imaging device 310 as it moves through a 3D space 300. For example, FIG. 3 shows the imaging device 310 at a plurality of positions 312 (e.g., 312a, 312b, 312c, and 312d) and orientations within an environment 300 as the imaging device 310 moves along a dotted line that schematically represents a trajectory 311. At each position 312, the imaging device 310 may be configured to capture an image frame of the environment 300 at a particular location and orientation that can be used, for example, as a continuous stream of image frames for performing sparse pose estimation. The trajectory 311 may be any trajectory or path of movement through the environment 300. FIG. 3 illustrates four positions 312, but the number of positions can vary. For example, the number of positions 312 may be as few as two positions or any desired number (e.g., 5, 6, 7, etc.) for performing sparse pose estimation with an acceptable level of certainty. In some embodiments, the imaging device 312 may be configured to capture a series of image frames, such as in a video, and each image frame of the video may be utilized for performing sparse pose estimation via the computer vision techniques described herein.
[0039] In some embodiments, the imaging device 310 may be configured as the display system 100 of FIG. 1, the imaging system including an imaging system 110 facing outward, a mobile robot, or a stand-alone imaging device. The imaging device 310 may be configured to capture an image frame at each position 312 depicting a portion of the environment 300 from the front of the imaging device 310 as it moves through the environment 300. As described above, the portion of the environment 300 captured by the imaging device at each position 312 and orientation may be the FOV from the front of the imaging device 310. For example, the FOV at position 312a is schematically illustrated as FOV 315a. Each subsequent position and orientation of the imaging device 310 (e.g., 312b, 312c, and 312d) constitutes a corresponding FOV 315 (e.g., FOV 315b, 315c, and 315d). Computer vision techniques may be performed on each image frame obtained from the imaging device 310 to estimate the pose of the imaging device 310 at each position 312. The pose estimation may be, for example, an input to a sparse point estimation process employed to determine or generate a map (or a portion thereof) of the environment 300 and track the movement of the imaging device 310 through the environment 300.
[0040] The environment 300 may be any 3D space, such as an office room (as shown in FIG. 3), a living room, an outdoor space, etc. The environment 300 may include a plurality of objects 325 (such as furniture, personal items, surrounding structures, textures, detectable patterns, etc.) that are arranged throughout the environment 300. The object 325 may be an individual object that is uniquely identifiable compared to other features in the environment (for example, each wall may not be uniquely identifiable). Further, the object 325 may be a common feature captured in two or more image frames. For example, FIG. 3 illustrates an object 325a (a lamp in this embodiment) located at each position 312 along the corresponding lines of sight 330a-d (shown as dotted lines for illustrative purposes) within the FOV 315 of the imaging device 310. Thus, for each position 312 (such as 312a), the image frame representing each FOV 315 (such as 315a) includes the object 325a imaged along the line of sight 330 (such as 330a).
[0041] The imaging device 310 may be configured to detect and extract a plurality of sparse points 320, and each sparse point 320 (or a plurality of sparse points) corresponds to an object 325 or a part of an object 325, a texture, or a pattern from each image frame representing the FOV 315. For example, the imaging device 310 may extract a sparse point 320a corresponding to the object 325a. In some embodiments, the object 325a may be associated with one or more sparse points 320, and each sparse point 320 may be associated with a different part of the object 325 (such as the corner, top, bottom, side, etc. of the lamp). Therefore, each sparse point 320 may be uniquely identifiable within the image frame. Computer vision techniques can be used to extract and identify each sparse point 320 from the image frame or image segment corresponding to each sparse point 320 (as described in connection with FIGS. 9A and 9B, for example).
[0042] In some embodiments, the sparse points 320 may be utilized to estimate the position, location, or orientation of the imaging device 310 within the environment 300. For example, the imaging device 310 may be configured to extract a plurality of sparse points 320 as inputs to a sparse pose estimation process. Exemplary computer vision techniques used for sparse pose estimation may be a simultaneous localization and mapping (SLAM or V-SLAM, referring to a configuration where the input is image / vision only) process or algorithm. Such exemplary computer vision techniques can be used to output a sparse point representation of the world surrounding the imaging device 310, as will be described in more detail below. In a conventional sparse pose estimation system that uses a plurality of image frames at the position 312, the sparse points 320 can be collected from each image frame, correspondences are calculated between consecutive image frames (e.g., positions 312a - 312b), and the pose change is estimated based on the discovered correspondences. Thus, in some embodiments, the position, orientation, or both the position and orientation of the imaging device 310 can be determined. In some implementations, a 3D map of the locations of the sparse points may be required for the estimation process or may be a byproduct of identifying the sparse points within an image frame or a plurality of image frames. In some embodiments, the sparse points 320 may be associated with one or more descriptors, which may be configured as a digital representation of the sparse points 320. In some embodiments, the descriptors may be configured to facilitate the calculation of correspondences between consecutive image frames. In some embodiments, the pose determination may be performed by an on-board processor of the imaging device (e.g., the local processing and data module 70) or a remote processor of the imaging device (e.g., the remote processing module 72).
[0043] In some embodiments, the computer vision module can be included to operably communicate with the imaging device 310, for example, as part of the local processing and data module 70 or the remote processing module and data repositories 72, 74. The exemplary computer vision module can implement one or more computer vision techniques, as described with reference to the methods 800, 1000 of FIGS. 8 and 10, and can analyze image segments acquired by an outward-facing imaging camera and be used, for example, to identify sparse points, determine poses, etc. The computer vision module can identify objects within the environment surrounding the imaging device 310, such as those described in connection with FIG. 3. The computer vision module can extract sparse points from the image frames as the imaging device moves within the environment and use the sparse points extracted to track and identify objects through various image frames. For example, the sparse points of the first image frame can be compared with the sparse points of the second image frame to track the movement of the imaging device. In some embodiments, one or more of the sparse points of the second image frame can include one or more of the sparse points of the first image frame as reference points for tracking, for example, between the first and second image frames. Image frames such as the third, fourth, fifth, etc. can be similarly used and compared with the preceding and subsequent image frames. The computer vision module can process the sparse points and estimate the position or orientation of the imaging device within the environment based on the identified sparse points.Non-limiting examples of computer vision techniques include Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, Visual Simultaneous Localization and Mapping (vSLAM) techniques, Sequential Bayesian estimators (e.g., Kalman filter, Extended Kalman filter, etc.), Bundle adjustment, Adaptive thresholding (and other thresholding techniques), Iterative Closest Point (ICP), Semi-Global Matching (SGM), Semi-Global Block Matching (SGBM), Feature point histograms, various machine learning algorithms (e.g., Support Vector Machine, k-Nearest Neighbor algorithm, Naive Bayes, Neural Networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), and the like.
[0044] As described above, the current pose estimation process can include a delay when estimating the pose of an imaging device. For example, the frame rate of the imaging device can cause a delay, in part, due to transferring the entire image frame from the imaging device to memory. Without subscribing to any particular scientific theory, sparse pose estimation can be delayed because sparse points are not extracted from the image frame until the entire image frame is read from the imaging device into memory. Thus, the transfer of the entire image frame, based in part on the frame rate capabilities of the imaging device, can be one of the components of the delay incurred in sparse pose estimation. One non-limiting advantage of some of the systems and devices described herein is that the extraction or identification of sparse points for estimating pose can be performed on-the-fly as a portion of the frame of the image is read into the image sensor or memory, and thus, the pose can be estimated at an earlier point in time than would otherwise be possible when using the entire image frame. Further, since only a portion of the frame can be analyzed for keypoints, the processing speed and efficiency can also be increased.
[0045] The foregoing description describes the sparse points 320 in the context of physical objects within the environment 300, which is not intended to be limiting, and other implementations are also conceivable. In some embodiments, the object 325 can refer to any feature of the environment (e.g., a real-world object, a virtual object, an invisible object, or a feature, etc.). For example, the projection device may be configured to project a plurality of indicators, textures, identifiers, etc., which may be visible or invisible, throughout the environment (e.g., projected within the IR spectrum, near-IR spectrum, ultraviolet spectrum, or any other suitable wavelength range or range of wavelengths of electromagnetic radiation). The indicators, textures, identifiers, etc. may be prominent features or shapes detectable by the imaging device 310. The imaging device 310 may be configured to detect these indicators and extract the sparse points 320 from the plurality of indicators. For example, the indicators may be projected onto the walls of the environment within the IR spectrum of electromagnetic radiation, and the imaging device 310 may be configured to operate within the IR spectrum, identify the indicators, and extract the sparse points therefrom. In another embodiment, alternatively or in combination, the imaging device 310 may be included within an AR device configured to display virtual image elements (e.g., on the display 62). The imaging device or AR device may be configured to identify the virtual image elements and extract the sparse points 320 therefrom. The AR device may be configured to use these sparse points 320 to determine the pose of the AR device with respect to the virtual image elements.
[0046] (Examples of shear effects imparted within exemplary image frames and sparse points) As described above, the imaging system 110 facing outward may be implemented as a rolling shutter camera. One non-limiting advantage of a rolling shutter camera is the ability to transmit a portion of the captured scene (e.g., an image segment) while capturing other portions (e.g., not all portions of the image frame are captured exactly simultaneously). However, this can result in distortion of an object moving relative to the camera while the image frame is being captured, because the imaging device may not be in the same position relative to the object throughout the time the image is being captured.
[0047] For example, FIGS. 4A and 4B are diagrams of the rolling shutter effect (e.g., sometimes also referred to herein as "shearing", "skewing", or "distortion") applied to an image of a scene. FIG. 4A schematically illustrates a scene 400a that includes an object 425a (e.g., a square in this example). The scene may be within the FOV of an image capture device (e.g., the outward-facing imaging system 110 of FIG. 2). In the embodiment illustrated in FIG. 4A, the scene may be moving relative to the image capture device in direction 430. FIG. 4B illustrates an image 400b that results from the captured scene 400a and that may be stored in a memory or storage unit (e.g., the local processing and data module 70). As illustrated in FIG. 4B, due to the relative movement of object 425a, the resulting image 400b is a distorted object 425b (e.g., shown as a sheared square or diamond), and the dotted lines of the distorted object are not captured within the resulting image 400b. Without being bound to any particular scientific theory, this may be due to the gradually downward scanning direction of the imaging device, and thus the top of the object is captured first and is less distorted than the bottom of the object.
[0048] FIGS. 5A and 5B are schematic diagrams of the rolling shutter effect applied to a plurality of sparse points included within a FOV (e.g., FOVs 315a, 315b, 315c, or 315d of FIG. 3) captured by an imaging device. For example, as the AR device moves around in 3D space, various sparse points also move relative to the AR device and are distorted in a manner similar to that described above in connection with FIG. 4B and as schematically illustrated in FIG. 5B. FIG. 5A illustrates a scene (which may be similar to scene 300 of FIG. 3) that includes a plurality of sparse points 320 (e.g., 320a, 320b, and 320c). FIG. 4B schematically illustrates the resulting captured image frame that includes distorted sparse points 525 (e.g., 525a, 525b, and 525c). For example, each distorted sparse point 525 is associated with an illustrative corresponding arrow 522. For illustrative purposes only, the size of arrow 522 is proportional to the amount of distortion imparted to sparse point 525. Thus, similar to that described above in connection with FIG. 4B, arrow 522a is smaller than arrow 522e, which may indicate that sparse point 525a associated with arrow 522a is not as highly distorted as compared to sparse point 525e.
[0049] (Exemplary AR Architecture) FIG. 6 is a block diagram of an embodiment of an AR architecture 600. The AR architecture 600 is configured to receive input (e.g., visual input from an outward-facing imaging system 110, input from a room camera, etc.) from one or more imaging systems. The imaging devices not only provide images from the FOV cameras, but they may also be equipped with various sensors (e.g., accelerometers, gyroscopes, temperature sensors, motion sensors, depth sensors, GPS sensors, etc.) to determine the location and various other attributes of the user's environment. This information may further be supplemented with information from stationary cameras within the room that may provide images and / or various cues from different viewpoints.
[0050] The AR architecture 600 may include a plurality of cameras 610. For example, the AR architecture 600 may include the outward-facing imaging system 110 of FIG. 1 configured to input a plurality of images captured from the FOV in front of the wearable display system 100. In some embodiments, the cameras 610 may include a relatively wide field of view, i.e., a passive pair of cameras arranged on the sides of the user's face and a different pair of cameras oriented in front of the user, and may handle a stereoscopic imaging process. However, other imaging systems, cameras, and arrangements are also conceivable.
[0051] The AR architecture 600 may also include a map database 630 that includes map data regarding the world. In one embodiment, the map database 630 may reside, in part, on a user-wearable system (e.g., local processing and data module 70), or in part, on a networked storage location (e.g., remote data repository 74) accessible by a wired or wireless network. In some embodiments, the map database 630 may include real-world map data or virtual map data (e.g., including virtual image elements that define or are overlaid on a real-world environment). In some embodiments, computer vision techniques can be used to produce the map data. In some embodiments, the map database 630 may be an existing map of the environment. In other embodiments, the map database 630 may be populated based on identified sparse points that are read into memory and subsequently stored for comparison and processing against the identified sparse points. In another embodiment, the map database 630 may be an existing map that is dynamically updated based on sparse points identified from one or more image frames (or a portion of a frame for a rolling shutter camera system), alone or in combination. For example, one or more sparse points may be used to identify an object in the environment (e.g., object 325 in FIG. 3) and may be used to populate the map with identifying features of the environment.
[0052] The AR architecture 600 may also include a buffer 620 configured to receive inputs from the camera 610. The buffer 620 may be, for example, a non-transitory data buffer that is separate from or part of a non-transitory data storage device (e.g., the local processing and data module 70 of FIG. 2) and is configured to store image data on a temporary basis. The buffer 620 may then temporarily store some or all of the received inputs. In some embodiments, the buffer 620 may be configured to store one or more portions or segments of the received data (e.g., as described below in connection with FIGS. 9A and 9B) before further processing is performed and the data is moved to another component of the AR architecture 600. In some embodiments, the image data collected by the camera 610 may be loaded into the buffer 620 as the user experiences the wearable display system 100 operating within the environment. Such image data may include an image or a segment of an image captured by the camera 610. The image data representing the image or segment of the image may then be transmitted to and stored in the buffer 620 before being processed by the local processing and data module and transmitted to the display 62 for visualization and presentation to the user of the wearable display system 100. The image data may also be stored, alternatively or in combination, within the map database 630. Or, the data may be removed from the memory (e.g., the local processing and data module 70 or the remote data repository 74) after being stored in the buffer 620. In one embodiment, the buffer 620 may reside, in part, on the user-wearable system (e.g., the local processing and data module 70) or, in part, on a networked storage location (e.g., the remote data repository 74) accessible by a wired or wireless network.
[0053] The AR architecture 600 may also include one or more object recognition devices 650. The object recognition device may be configured to crawl through the received data, identify and / or tag objects, and add information to the objects, for example, via computer vision techniques, using the map database 630. For example, the object recognition device may scan or crawl through the image data or image segments stored in the buffer 620 and identify objects captured in the image data (e.g., object 325 in FIG. 3). Objects identified in the buffer may be tagged or descriptive information may be added thereto with reference to the map database. The map database 630 may include various objects identified over time and between the captured image data and its corresponding objects (e.g., comparison of objects identified in a first image frame and objects identified in subsequent image frames), and may be used to generate the map database 630 or to generate a map of the environment. In some embodiments, the map database 630 may be populated with an existing map of the environment. In some embodiments, the map database 630 is stored on board the AR device (e.g., local processing and data module 70). In other embodiments, the AR device and the map database are interconnected through a network (e.g., LAN, WAN, etc.) and can access a cloud storage device (e.g., remote data repository 74).
[0054] In some embodiments, the AR architecture 600 includes a pose estimation system 640 that is configured to perform a pose estimation process and execute instructions for determining the location and orientation of a wearable computing hardware or device, based at least in part on data stored in buffer 620 and map database 630. For example, position, location, or orientation data may be calculated from data collected by camera 610 as the user experiences the wearable device and moves within the world, as the data is read into buffer 620. For example, based on object information and collection identified from the data and stored in buffer 620, object recognition device 610 may recognize objects 325 and extract these objects as keypoints 320 to a processor (e.g., local processing and data module 70). In some embodiments, keypoints 320 may be extracted and used to determine the pose of the AR device within the associated image frame as one or more image segments of a given image frame are read into buffer 620. The pose estimation may be updated and used to identify additional keypoints as additional image segments of the image frame are read into buffer 620. Optionally, in some embodiments, pose estimation system 640 accesses map database 630, reads out keypoints 320 identified within previously captured image segments or image frames, compares corresponding keypoints 320 between preceding and subsequent image frames as the AR device moves through 3D space, and thereby tracks the movement, position, or orientation of the AR device within the 3D space. For example, referring to FIG. 3, object recognition device 650 may recognize keypoint 320a as lamp 325a within each of a plurality of image frames. The AR device may add certain descriptor information, associate keypoint 320a within one image frame with the corresponding keypoint 320a of other image frames, and store this information within map database 650. Object recognition device 650 may be configured to recognize objects with respect to any number of keypoints 320, e.g., 1, 2, 3, 4, etc.
[0055] Once an object is recognized, the information may be used by the pose estimation system 640 to determine the pose of the AR device. In one embodiment, the object recognition device 650 may identify sparse points corresponding to the image segments as the image segments are received, and then may identify additional sparse points when subsequent image segments of the same image frame are received. The pose estimation system 640 may execute instructions to update the estimate by estimating the pose based on the first identified sparse points and then integrating the subsequently identified sparse points into the estimation process. In another embodiment, alone or in combination, the object recognition device 650 may recognize two sparse points 320a, 320b of two objects (e.g., object 325a shown in FIG. 3 and another object) in the first frame, and then may identify the same two sparse points in the second frame and subsequent frames (e.g., any maximum number of subsequent frames may be considered). Based on a comparison between the sparse points of two or more frames, the pose (e.g., orientation and location) in 3D space may also be estimated or tracked through the 3D space.
[0056] In some embodiments, the accuracy of pose estimation or the reduction of noise in the low pose estimation result can be based on the number of sparse points recognized by the object recognition device 640. For example, in 3D space, the position, location, or orientation of the imaging device can be based on the translational and rotational coordinates in the environment. Such coordinates may include, for example, X, Y, and Z translational coordinates or yaw, roll, and pitch rotational coordinates, as will be described below in connection with FIG. 7. In some embodiments, one sparse point extracted from an image frame may not be able to convey the complete pose of the imaging device. However, a single sparse point can be at least one constraint regarding pose estimation, for example, by providing information related to one or more coordinates. As the number of sparse points increases, the accuracy of pose estimation can be improved, or the noise or error in pose estimation can be reduced. For example, two sparse points can indicate the X and Y positions of the imaging device in 3D space based on the object represented by the sparse points. However, the imaging device may not be able to determine its Z position (e.g., the front or back of the object) or its roll coordinate relative to the object. Thus, in some embodiments, three sparse points may be used to determine the pose, however, any number of sparse points may be used (e.g., 1, 2, 4, 5, 6, 7, 10 or more, etc.).
[0057] In some embodiments, pose determination may be performed by an on-board processor of the AR device (e.g., the local processing and data module 70). The extracted sparse points may be input into a pose estimation system 640 configured to perform computer vision techniques. In some embodiments, the pose estimation system may comprise SLAM or V-SLAM (e.g., refer to the configuration where the input is image / vision only) that is executed by the pose estimation system 640 and may then output a sparse point representation 670 of the world surrounding the AR device. In some embodiments, the pose estimation system 640 may be configured to execute an continuously updated recursive Bayesian estimator (e.g., a Kalman filter). However, the Bayesian estimator is intended as an illustrative example of at least one method for performing pose estimation by the pose estimation system 640, and other methods and processes are also envisioned within the scope of the present disclosure. The system can be configured to find not only the world in which various components exist, but also what the world is composed of. Pose estimation may be a building block that achieves many goals, including populating the map database 630 and using data from the map database 630. In other embodiments, the AR device can be connected to a processor configured to perform pose estimation through a network (e.g., LAN, WAN, etc.) and access a cloud storage device (e.g., the remote data repository 74).
[0058] In some embodiments, one or more remote AR devices may be configured to determine the pose of each AR device based on the pose determination of a single AR device with the AR architecture 600. For example, one or more AR devices may communicate wired or wirelessly with a first AR device that includes the AR architecture 600. The first AR device may perform pose determination based on sparse points extracted from the environment as described herein. The first AR device may also be configured to transmit an identification signal (e.g., an IR signal or other suitable medium) that may be received by one or more remote AR devices (e.g., a second AR device). In some embodiments, the second AR device may attempt to display similar content as the first AR device and receive the identification signal from the first AR device. From the identification signal, the second AR device may extract sparse points and be able to determine its pose relative to the first AR device (e.g., interpret or process the identification signal) without performing pose estimation on the second AR device itself. One non-limiting advantage of this arrangement is that differences in virtual content displayed on the first and second AR devices can be avoided by linking the two AR devices. Another non-limiting advantage of this arrangement is that the second AR system may be able to update its estimated position based on the identification signal received from the first AR device.
[0059] (Examples of the Pose and Coordinate System of the Imaging Device) FIG. 7 is an example of a coordinate system related to the pose of an imaging device. Device 700 may have multiple degrees of freedom. As device 700 moves in different directions, the position, location, or orientation of device 700 will change with respect to the starting position 720. The coordinate system in FIG. 7 shows three translational directions of movement (e.g., the X, Y, and Z directions) that can be used to measure device movement with respect to the starting position 720 of the device and to determine a location in 3D space. The coordinate system in FIG. 7 also shows three angular degrees of freedom (e.g., yaw, pitch, and roll) that can be used to measure device orientation with respect to the starting direction 720 of the device. As shown in FIG. 7, device 700 can also be moved horizontally (e.g., in the X or Z direction) or vertically (e.g., in the Y direction). Device 700 can also tilt forward and backward (e.g., pitch), turn left and right (e.g., yaw), and tilt sideways (e.g., roll). In other implementations, other techniques or angular representations for measuring head pose, such as any other type of Euler angle system, can also be used.
[0060] FIG. 7 illustrates device 700, which may be implemented as, for example, wearable display system 100, an AR device, an imaging device, or any other device described herein. As described throughout this disclosure, device 700 may be used to determine pose. For example, if device 700 is an AR device that includes the AR architecture 600 of FIG. 6, the pose estimation system 640 may use the image segment input to extract sparse points for use in the pose estimation process as described above, track device movement in the X, Y, or Z direction, or track angular movement in yaw, pitch, or roll.
[0061] (Exemplary Routine for Estimating Pose in 3D Space) FIG. 8 is a process flow diagram of an illustrative routine for determining the pose of an imaging device (e.g., the imaging system 110 facing outward in FIG. 2) within a 3D space (e.g., FIG. 3) in which the imaging device moves. Routine 800 describes a way in which a plurality of sparse points can be extracted from an image frame representing a FOV (e.g., FOVs 315a, 315b, 315c, or 315d) to determine one of the position, location, or orientation of the imaging device within the 3D space.
[0062] In block 810, the imaging device may capture an input image regarding the environment surrounding the AR device. For example, the imaging device may successively capture a plurality of image segments of the input image based on light received from the surrounding environment. This may be accomplished through various input devices (e.g., a digital camera on or remote from the AR device). The input is an image representing a FOV (e.g., FOVs 315a, 315b, 315c, or 315d) and may include a plurality of sparse points (e.g., sparse point 320). A FOV camera, sensor, GPS, etc. may communicate information including the image data of the successively captured image segments to the system (block 810) as the image segments are captured by the imaging device.
[0063] At block 820, the AR device may receive an input image. In some embodiments, the AR device may sequentially receive a plurality of image segments that form a portion of the image captured at block 810. For example, as described above, the outward-facing imaging system 110 may be a rolling shutter camera configured to sequentially scan a scene, thereby sequentially capturing a plurality of image segments and causing image data to be sequentially loaded into a memory unit as the data is captured. The information may be stored on a user-wearable system (e.g., local processing and data module 70), or may reside, at least in part, in a networked storage location (e.g., remote data repository 74) accessible by a wired or wireless network. In some embodiments, the information may be temporarily stored in a buffer included within the memory unit.
[0064] At block 830, the AR device may identify one or more sparse points based on the received image segments. For example, an object recognition device may crawl through the image data corresponding to the received image segments and identify one or more objects (e.g., object 325). In some embodiments, the identification of one or more sparse points may be based on the receipt of image segments corresponding to the one or more sparse points, as described below with reference to FIGS. 9A and 9B. The object recognition device may then extract sparse points that can be used as an input for determining pose data (e.g., the pose of the imaging device within 3D space). This information may then be communicated to a pose estimation process (block 840), and the AR device may thus map the AR device through 3D space using a pose estimation system (block 850).
[0065] In various embodiments, routine 800 may be implemented by a hardware processor (e.g., local processing and data module 70 of FIG. 2) configured to execute instructions stored in a memory or storage unit. In other embodiments, a remote computing device (communicating with a display device over a network) with computer-executable instructions may cause the display device to implement aspects of routine 800.
[0066] As described above, the current pose estimation process may include a delay in estimating the pose of the AR device due to transferring data (e.g., extracted sparse points) from an image capture device to the pose estimation system. For example, current implementations may require that an entire image frame be transferred from the image capture device to a pose estimator (e.g., SLAM, VSLAM, or the like). Once the entire image frame has been transferred, the object recognition device is enabled to identify sparse points and extract them to the pose estimator. The transfer of the entire image frame can be one contributing factor to the delay in estimating the pose.
[0067] (Exemplary extraction of sparse points from an image frame) Figures 9A and 9B schematically illustrate an example of extracting one or more sparse points from an image frame based on receiving a plurality of image segments. In some implementations, Figures 9A and 9B may also schematically illustrate an exemplary method of minimizing latency in estimating the pose of an imaging device (e.g., the outward-facing imaging device 110 of FIG. 2) through a 3D space. In some embodiments, Figures 9A and 9B also schematically depict an example of identifying one or more sparse points of an image frame 900. In some implementations, Figures 9A and 9B illustrate an image frame as read from an imaging device into a memory unit by a rolling shutter camera, as described above. The image frame 900 may be captured by an outward-facing imaging system 110 configured as a progressive scan imaging device. The image frame may include a plurality of image segments (sometimes also referred to as scan lines) 905a - 905n that are read into a memory unit (e.g., local processing and data module 70) from the imaging device as the image segments are captured by the imaging device. The image segments may be arranged horizontally (as shown in FIG. 9A) or vertically (not shown). Fifteen image segments are illustrated, but the number of image segments need not be so limited and may be any number of image segments 905a - 905n, depending on the desire for a given application or based on the capabilities of the imaging system. In some implementations, the image segments may be lines (e.g., rows or columns) within a raster scan pattern. For example, the image segments may be rows or columns of pixels within a raster scan pattern of an image captured by the outward-facing imaging device 110. The raster scan pattern may be implemented or executed by a rolling shutter camera, as described throughout the present disclosure.
[0068] Referring again to FIG. 9A, the image frame 900 may include a plurality of image segments 905 that are successively captured and read into the memory unit. The image segments 905 may be combined to represent the field of view (FOV) captured by the imaging device. The image frame 900 may also include a plurality of sparse points 320, as described above with reference to FIG. 3. In some implementations, as illustrated in FIG. 9A, each sparse point 320 may be generated by one or more image segments 905. For example, the sparse point 320a may be generated by a subset 910 of the image segments 905 and may thus be associated therewith. Thus, each sparse point may be identified in response to receiving a subset of the image segments 905 corresponding to each given sparse point when the image segments are received in the memory unit. For example, the sparse point 320a may be identified by an object recognition device (e.g., the object recognition device 650) as soon as the image segments 906a-906n are received in the memory unit of the AR device. The image segments 906a-906n may correspond to the subset 910 of the image segments 905 that represent the sparse point 320a. Thus, the AR device may be able to determine individual sparse points as soon as the corresponding image segments are received from the image capture device (e.g., a progressive scan camera). The subset 910 of the image segments 905 may include the image segments 906a-906n. In some implementations, the number of the image segments 906 may be based on the number of successively received image segments in the vertical direction required to decompose or capture the entire sparse points along the vertical direction. FIG. 9B illustrates seven image segments associated with the sparse point 320a, but this is not necessarily the case, and any number of image segments may be associated with the sparse point 320a and identify the object 325a corresponding to the sparse point 320a (e.g., 2, 3, 4, 5, 6, 8, 9, 10, 11, etc.).
[0069] In one exemplary implementation, sparse point 320 may be identified by implementing a circular buffer or a rolling buffer. For example, the buffer may be similar to buffer 620 of FIG. 6. The buffer may be constructed as part of a memory or storage unit stored on board the AR device (e.g., local processing and data module 70), or may be remote from the AR device (e.g., remote data repository 74). The buffer may be configured to receive image information from an image capture device (e.g., the outward-facing imaging system 110 of FIG. 2). For example, the buffer may continuously receive image data representing image segments from an image sensor as the image sensor captures each sequential image segment. The buffer may also be configured to store a portion of the image data for subsequent processing and identification of the image content. In some embodiments, the buffer may be configured to store one or more image segments, and the number of image segments may be less than the total image frame 900. In some embodiments, the number of image segments stored in the buffer may be a predetermined number, e.g., the number within subset 910. In some embodiments, alternatively or in combination, the buffer may be configured to store a subset 910 of the image segments corresponding to the sparse points. For example, referring to FIG. 9B, sparse point 320a may require a 7×7 pixel window (e.g., 7 rows of pixels present image segment 906, and each image segment includes 7 pixels). In this embodiment, the buffer may be configured to be large enough to store subset 910 of image segment 906, e.g., 7 image segments are illustrated.
[0070] As described above, the buffer may be configured to temporarily store image data. Thus, as new image segments are received from the imaging capture device, older image segments are removed from the buffer. For example, a first image segment 906a may be received, and subsequent image segments may be received in a buffer corresponding to the sparse points 320a. Once all of the image segments 906a - 906n are received, the sparse points 320a may be identified. Subsequently, a new image segment is received (e.g., 906n+1), and the image segment 906a is thereby removed from the buffer. In some embodiments, the segment 906a is moved from the buffer to a storage device (e.g., local processing and data module 70) within digital memory for further processing.
[0071] (Exemplary Routine for Estimating Pose in 3D Space) FIG. 10 is a process flow diagram of an illustrative routine for determining the pose of an imaging device (e.g., the outward-facing imaging system 110 of FIG. 2) within a 3D space (e.g., FIG. 3) in which the imaging device moves. Routine 1000 describes an example of a method by which a first group of sparse points can be extracted from an image frame as image segments corresponding to the sparse points of the first group are received. In various embodiments, the corresponding image segments may be captured prior to capturing the entire image frame representing the FOV of the imaging device. Routine 1000 also describes a method by which subsequent sparse points or a second group of sparse points can be extracted and integrated to update the pose determination. Routine 1000 may be implemented by a hardware processor (e.g., local processing and data module 70 of FIG. 2) that is operably coupled to an outward-facing imaging system (e.g., outward-facing imaging system 110) and a digital memory or buffer as described above. The outward-facing imaging system 110 can comprise a rolling-shutter camera.
[0072] In block 1010, the imaging device may capture an input image regarding the environment surrounding the AR device. For example, the imaging device may successively capture a plurality of image segments of the input image based on the light received from the surrounding environment. This may be achieved through various input devices (e.g., a digital camera on or remote from the AR device). The input is an image frame representing a FOV (e.g., FOV 315a, 315b, 315c, or 315d), which may include a plurality of sparse points (e.g., sparse point 320). FOV cameras, sensors, GPS, etc. may transmit information including the image data of the successively captured image segments to the system (block 1010) as the image segments are captured by the imaging device.
[0073] In block 1020, the AR device may receive the input image. In some embodiments, the AR device may successively receive a first plurality of image segments that form part of the image captured in block 1010. For example, the imaging device may be configured to successively scan the scene, as described above with reference to FIGS. 9A and 9B, thereby successively capturing the first plurality of image segments. The image sensor may also successively read the image data for the memory unit as the data is captured. The information may be stored on a user-wearable system (e.g., local processing and data module 70), or may reside, in part, in a networked storage location (e.g., remote data repository 74) accessible by a wired or wireless network. In some embodiments, the information may be temporarily stored in a buffer included within the memory unit.
[0074] In block 1030, the AR device may identify a first group of sparse points based on receiving a first plurality of image segments (sometimes also referred to as a "pre-list") corresponding to each sparse point. For example, referring to FIGS. 9A and 9B, the AR device may identify one or more sparse points 320 based on receiving a subset 910 of image segments 905 (e.g., the first plurality of image segments) corresponding to one or more sparse points 320 as described above with reference to FIGS. 9A and 9B. The sparse points 320 may be identified as soon as a subset 910 (e.g., image segment 906) of the image segments 905 corresponding to the sparse points 320 is received in a memory unit (e.g., local processing and data module 70).
[0075] In some implementations, the first group of sparse points includes a number (N1) of arbitrary sparse points. The number (N1) may be any number of sparse points selected to estimate the pose of the AR device in the environment. In some embodiments, the number (N1) must be no less than three sparse points. In other embodiments, the number (N1) is 10 to 20 sparse points. One non-limiting advantage of a larger number (N1) is that outlier data points can be rejected, which can provide a certain degree of robustness to noise due to inlier data points in pose determination. For example, the imaging device may be subject to jitter or shock due to events imposed on the physical imaging device, or the recorded scene may be temporarily changed (e.g., a person moves within the foreground). The event may only affect a small group of sparse points within one or more image frames. By using a larger number (N1) of sparse points or updating the pose estimation according to this specification, the noise in the pose estimation due to these outliers or single-instance events can be at least partially reduced.
[0076] In one implementation, the sparse points of the first group may be extracted from an image frame (e.g., by the object recognition device 650) and transmitted to a pose estimation system (e.g., the pose estimation system 640 of FIG. 6) configured to perform pose determination (e.g., SLAM, VSLAM, or the like, as described above) (block 1040). In various embodiments, the sparse points of the first group are transmitted to the pose estimation system in response to identifying a number (N1) of sparse points based on receiving corresponding first plural image segments. Thus, the sparse points of the first group may be transmitted when only a part of the image frame is received because the imaging device does not receive the entire image frame. Subsequent image segments (e.g., second plural image segments acquired after the first plural image segments) will be received as is. In one embodiment, the sparse points of the first group may be extracted as soon as they are each identified based on scanning corresponding subsets of the image segments (e.g., from a storage unit or a part thereof of the AR device, e.g., a buffer). In another embodiment, the sparse points of the first group may be extracted once a number (N1) of sparse points are identified and the sparse points are transmitted in a single process (e.g., from a storage unit or buffer of the AR device).
[0077] In block 1045, the AR device may receive a second plurality of image segments (sometimes also referred to as a "tracking list"). In some embodiments, the AR device may continuously acquire the second plurality of image segments after receiving the first plurality of image segments in block 1020. For example, as described above with reference to FIGS. 9A and 9B, the imaging device continuously scans the scene, thereby continuously capturing the first plurality of image segments (e.g., in block 1020), and subsequently, either after or during block 1030, continuously scans the scene and is configured to acquire the second plurality of image segments. In another embodiment, the second plurality of image segments or a portion thereof may be obtained from a second image captured by the imaging device, and the second image is captured after the first image. The information may be stored on the AR device (e.g., local processing and data module 70), or may reside, in part, in a networked storage location (e.g., remote data repository 74) accessible by a wired or wireless network. In some embodiments, the information may be temporarily stored in a buffer included within the storage unit.
[0078] Referring back to FIG. 10, in block 1050, the AR device may identify a second group of sparse points based on the second plurality of image segments. For example, in one embodiment, the entire image frame has not been received prior to determining the pose in block 1040, and the second plurality of image segments may be received from the imaging device in block 1045. Thus, the AR device may identify one or more new sparse points (e.g., the second group of sparse points) based on receiving the second plurality of image segments corresponding to one or more new sparse points (e.g., as described above with reference to FIGS. 9A and 9B). In another embodiment, the second image may be captured by the imaging device after the first image is captured in block 1010, and the second plurality of image segments may be obtained from the second image. Thus, the AR device may identify one or more new sparse points based on receiving the second plurality of image segments from the second image, which may correspond to the second group of sparse points. In some embodiments, the second group of sparse points may include any number of new sparse points (e.g., 1, 2, 3, etc.). In one implementation, the second group of sparse points may be extracted and incorporated into the pose determination, for example, by communicating the second group of sparse points to the pose estimation system. The following is an exemplary method of incorporating the second group of sparse points into the mapping routine of FIG. 10 along with the first group of sparse points. For example, the exemplary integration methods described herein may be referred to as reintegration, sliding scale integration, or block integration. However, these exemplary integration methods are not intended to be exhaustive. Other methods that may minimize error and reduce latency in pose determination are also conceivable.
[0079] In block 1060, the pose estimation system may be configured to update the pose determination based on the pose determination in block 1040 and the receipt of the second group of sparse points in block 1050.
[0080] One non-limiting advantage of routine 1000 described above can be a reduction in latency resulting from extracting outliers from the image frame prior to the pose estimation process. For example, when the image segments corresponding to those outliers are received in buffer 620, by calculating and identifying the individual outliers, individual or selected groups of outliers may be extracted and processed by the pose estimation system without waiting for the entire image frame to be captured. Thus, pose estimation may be performed considerably before the entire image is transferred to memory and before all outliers can be extracted from the entire image. However, once the first group and all subsequent groups of a particular image frame have been extracted, the entire image frame will then be available for pose estimation.
[0081] In various implementations, the second group of outliers may include a set number of outliers identified after determining the pose at block 1040. In some embodiments, the set number may be one outlier. For example, each time a subsequent outlier is identified, the outlier is transmitted to the pose estimation system and a new pose estimation process is performed at block 1060 to update one or more of the position, location, or orientation of the AR device. The method may sometimes be referred to as a reintegration method. Thus, each subsequently identified outlier may represent a subsequent group of outliers (e.g., groups of second, third, fourth, etc. outliers). In another embodiment, the set number may be any number of subsequently identified outliers (e.g., 2, 3, 4, etc.). For example, when the set number is 3, each time 3 new outliers are identified (e.g., a subsequent group of outliers), the group is transmitted to the pose estimation system at block 1050 and a new pose estimation process is performed at block 1060. The pose estimation process may thus utilize all outliers included within the entire image frame.
[0082] In other implementations, the integration method may be configured to account for the rolling shutter effect, as described above with reference to FIGS. 4A - 5B. For example, the pose estimation process may be performed for a fixed number (N2) of sparse points. This method may sometimes be referred to as a sliding integration method. In this embodiment, the second group of sparse points may include a selected number (k2) of sparse points identified after determining the pose at block 1040. Each time a number (k2) of sparse points can be identified, the pose determination may be updated. However, only the most recent N2 sparse points may be used to update the pose at block 1060. In some embodiments, the method uses the most recent N2 sparse points regardless of the group to which they correspond. For example, if N1 is set to 10, N2 is set to 15, and k2 is set to 5, then the sparse points of the first group include the first 10 sparse points identified at block 1030. Thus, the pose is determined at block 1040 based on the first 10 sparse points. Subsequently, new sparse points are identified, but the pose is not updated. Once 5 new sparse points that make up the sparse points of the second group are identified, the pose may be updated based on the sparse points of the first group (N1) and the second group (k2). If sparse points of a third group are identified (e.g., 5 sparse points following the second group), the pose is updated again at block 1060, however, the update may be based on some of the first group (e.g., sparse points 6 - 10), the second group (e.g., sparse points 11 - 15), and the third group (e.g., sparse points 16 - 21). Thus, the integration is considered as a sliding window or sliding list of sparse points, whereby only the set number of sparse points are used to estimate the pose, and the sparse points used may slide from the first group to the second and third groups. One non - limiting advantage of this method may be that sparse points identified from previously received image segments can be removed from the pose determination at block 1060 as they become old or stale.In some cases, when the AR device is moving with respect to the sparse points, the rolling shutter effect can be reduced by removing the old sparse points and capturing the change in pose between the newly identified sparse points.
[0083] In some embodiments, the previous integration method may be utilized between image frames as, for example, an imaging system 110 facing outward moves between capturing an image frame of FOV 315a in FIG. 3 and capturing an image frame of FOV 315b. For example, a first group of sparse points may be received from an image frame associated with a first position 312a (e.g., FOV 315b), and a second group of sparse points may be received from an image frame associated with a second position 312b (e.g., FOV 315b). A sliding list method may be implemented to reduce the rolling shutter effect between these image frames. However, in some embodiments, it may not be necessary to retain sparse points that exceed the most recent (N2 - 1) from the first frame.
[0084] In another implementation, the pose determination at block 1060 may be performed due to a fixed number or sparsity points of the block. This method may sometimes be referred to as a block integration method. In some embodiments, each group of sparsity points may include the same number of sparsity points as the block. For example, if the block is set to 10, the fixed number (N1) for the first group is 10, and the pose is determined at block 1040 in response to identifying and extracting this first group. Subsequently, the next 10 sparsity points may be included, and a second group may be identified, and the pose is updated at block 1060 using this second group. In some embodiments, this process may continue for multiple groups (e.g., the third, fourth, fifth, etc.). In some embodiments, when an image segment is stored in a buffer (e.g., buffer 620 of FIG. 6), the size of the buffer may be selected and configured to store at least the number of sparsity points that can be included within the block (e.g., the buffer may be selected to have a size configured to store at least 10 sparsity points in the foregoing example). In some embodiments, the buffer may have a limited size configured to store only the number of sparsity points configured within the block.
[0085] Although various embodiments of methods, devices, and systems have been described throughout this disclosure with reference to a head-mounted display device or an AR device, this is not intended to limit the scope of the present application and is merely used as an example for illustrative purposes. The methods and devices described herein are also applicable to other devices such as robots, digital cameras, and other autonomous entities that can implement the methods and devices described herein, map the 3D environment in which the device is located, and track the movement of the device through the 3D environment.
[0086] (Additional aspects) In a first aspect, a method for estimating the position of an image capture device within an environment is disclosed. The method includes receiving a first plurality of image segments successively, wherein the first plurality of image segments form at least a portion of an image representing the field of view (FOV) of the image capture device, the FOV constitutes a portion of the environment surrounding the image capture device and includes a plurality of sparse points, and each sparse point corresponds to a subset of the image segments; identifying a first group of sparse points, wherein the first group of sparse points includes one or more sparse points identified as the first plurality of image segments are received; determining, by a position estimation system, the position of the image capture device within the environment based on the first group of sparse points; receiving a second plurality of image segments successively, wherein the second plurality of image segments are received after the first plurality of image segments and form at least another portion of the image; identifying a second group of sparse points, wherein the second group of sparse points includes one or more sparse points identified as the second plurality of image segments are received; and updating, by the position estimation system, the position of the image capture device within the environment based on the first and second groups of sparse points.
[0087] In a second aspect, the method according to aspect 1, further including successively capturing, at an image sensor of the image capture device, a plurality of image segments.
[0088] In a third aspect, the method according to aspect 1 or 2, wherein the image sensor is a rolling shutter image sensor.
[0089] In a fourth aspect, the method according to any one of aspects 1-3, further including storing, in a buffer, the first and second pluralities of image segments as the image segments are received successively, the buffer having a size corresponding to the number of image segments within a subset of the image segments.
[0090] The method according to any one of aspects 1-4, wherein the fifth aspect further includes a step of extracting the sparse points of the first and second groups for the position estimation system.
[0091] The method according to any one of aspects 1-5, wherein in the sixth aspect, the sparse points of the first group include a certain number of sparse points.
[0092] The method according to aspect 6, wherein in the seventh aspect, the number of sparse points is 10 to 20 sparse points.
[0093] The method according to any one of aspects 1-7, wherein in the eighth aspect, the sparse points of the second group include the number of the second sparse points.
[0094] The method according to any one of aspects 1-8, wherein in the ninth aspect, the update of the position of the image capture device is based on the number of the most recently identified sparse points, and the most recently identified sparse points are at least one of those of the first group, the second group, or one or more of the first group and the second group.
[0095] The method according to aspect 9, wherein in the tenth aspect, the number of the most recently identified sparse points is equal to the number of sparse points in the sparse points of the first group.
[0096] The method according to any one of aspects 1-10, wherein in the eleventh aspect, the position estimation system is configured to perform visual simultaneous localization and mapping (V-SLAM).
[0097] The method according to any one of aspects 1-11, wherein in the twelfth aspect, the plurality of sparse points are extracted based on at least one of real-world objects, virtual image elements, and invisible indicators projected into the environment.
[0098] In a 13th aspect, a method for estimating the position of an image capture device in an environment is disclosed. The method includes receiving a plurality of image segments in succession, where the plurality of image segments form an image representing the field of view (FOV) of the image capture device, the FOV comprising a portion of the environment surrounding the image capture device that includes a plurality of sparse points, each sparse point being identifiable, at least in part, based on a corresponding subset of the image segments of the plurality of image segments; continuously identifying one or more of the plurality of sparse points when each subset of the image segments corresponding to the one or more sparse points is received; and estimating the position of the image capture device in the environment based on the identified one or more sparse points.
[0099] In a 14th aspect, the method according to aspect 13, wherein the step of receiving a plurality of image segments in succession further includes receiving a number of image segments and storing the number of image segments in a buffer.
[0100] In a 15th aspect, the method according to aspect 13 or 14, wherein the step of receiving a plurality of image segments in succession includes receiving at least a first image segment and a second image segment, the first image segment being stored in a buffer.
[0101] In a 16th aspect, the method according to any one of aspects 13 - 15, further including updating the buffer in response to receiving the second image segment, storing the second image segment in the buffer, and removing the first image segment in response to receiving the second image segment.
[0102] In a 17th aspect, the method according to aspect 16, wherein the step of continuously identifying one or more sparse points further includes scanning the image segments stored in the buffer when the buffer is updated.
[0103] In the 18th aspect, when each subset of image segments corresponding to one or more sparse points is received, the step of successively identifying one or more of the plurality of sparse points further includes: when the first plurality of image segments corresponding to one or more sparse points in the first group are received, successively identifying one or more sparse points in the first group; and when the second plurality of image segments corresponding to one or more sparse points in the second group are received, successively identifying one or more sparse points in the second group, wherein the second plurality of image segments are received after the first plurality of image segments. The method according to any one of aspects 13 - 17.
[0104] In the 19th aspect, the step of estimating the position of the image capture device is based on identifying one or more sparse points in the first group, and the first group includes a certain number of sparse points. The method according to any one of aspects 13 - 18.
[0105] In the 20th aspect, the number of sparse points is from 2 to 20. The method according to aspect 19.
[0106] In the 21st aspect, the number of sparse points is from 10 to 20. The method according to aspect 19.
[0107] In the 22nd aspect, the method further includes the step of updating the position of the image capture device based on identifying one or more sparse points in the second group. The method according to any one of aspects 13 - 21.
[0108] In the 23rd aspect, one or more sparse points in the second group include a second number of sparse points. The method according to any one of aspects 13 - 22.
[0109] In the 24th aspect, the method further includes the step of updating the position of the image capture device based on identifying a certain number of successively identified sparse points. The method according to any one of aspects 13 - 23.
[0110] In the 25th aspect, the number of successively identified sparse points is equal to the number of sparse points. The method according to aspect 24.
[0111] In the 26th aspect, the method described in aspect 24, where the number of sparsely identified points in succession includes at least one of the sparsely identified points of the first group of sparsely identified points.
[0112] In the 27th aspect, the method according to any one of aspects 13-26, where a plurality of sparsely identified points are extracted based on at least one of a real-world object, a virtual image element, and an invisible indicator projected into the environment.
[0113] In the 28th aspect, the method further includes the steps of extracting successively identified sparsely identified points from a buffer and transmitting the successively identified sparsely identified points to a visual simultaneous localization and mapping (VSLAM) system, where the VSLAM system estimates the position of an image capture device based on one or more successively identified sparsely identified points, according to any one of aspects 13-27.
[0114] In the 29th aspect, an augmented reality (AR) system is disclosed. The AR system includes an imaging device facing outward, computer hardware, and a processor operably coupled to the computer hardware and the imaging device facing outward and configured to execute instructions for implementing the method according to any one of aspects 1-28.
[0115] In the 30th aspect, the AR system according to aspect 29, where the imaging device facing outward is configured to detect light in the invisible spectrum.
[0116] In the 31st aspect, the AR system according to aspect 29 or 30, where the AR system is configured to display one or more virtual image elements.
[0117] On the 32nd aspect, it further includes a transceiver configured to transmit an identification signal indicating the estimated position of the AR system to the remote AR system, and the remote AR system is configured to update its estimated position based on the received identification signal. The AR system according to any one of aspects 29-31.
[0118] On the 33rd aspect, an autonomous entity is disclosed. The autonomous entity includes an imaging device facing outward, computer hardware, and a processor operably coupled to the computer hardware and the imaging device facing outward and configured to execute instructions for implementing the method according to any one of aspects 1-28.
[0119] On the 34th aspect, the autonomous entity according to aspect 33, wherein the imaging device facing outward is configured to detect light in the invisible spectrum.
[0120] On the 35th aspect, a robot system is disclosed. The robot system includes an imaging device facing outward, computer hardware, and a processor operably coupled to the computer hardware and the imaging device facing outward and configured to execute instructions for implementing the method according to any one of aspects 1-28.
[0121] In a 36th aspect, an image capture device for estimating the position of an image capture device within an environment is disclosed. The image capture device is an image sensor configured to capture an image by successively capturing a plurality of image segments, where the image represents the field of view (FOV) of the image capture device, the FOV constitutes a part of the environment surrounding the image capture device and includes a plurality of sparse points, and each sparse point is at least partially identifiable based on a corresponding subset of the plurality of image segments. The image capture device further includes a memory circuit configured to store a subset of the image segments corresponding to one or more of the sparse points, and a computer processor operably coupled to the memory circuit and configured to successively identify one or more of the plurality of sparse points upon receipt of each subset of the image segments corresponding to the one or more sparse points and to extract the successively identified one or more sparse points to estimate the position of the image capture device within the environment based on the identified one or more sparse points.
[0122] In a 37th aspect, the image capture device of aspect 36 further includes a position estimation system configured to receive the successively identified one or more sparse points and to estimate the position of the image capture device within the environment based on the identified one or more sparse points.
[0123] In a 38th aspect, the image capture device of aspect 36 or 37, wherein the position estimation system is a visual simultaneous localization and mapping (VSLAM) system.
[0124] In a 39th aspect, the image capture device of any one of aspects 36 - 38, wherein the image sensor is configured to detect light within the invisible spectrum.
[0125] In a 40th aspect, the image capture device of any one of aspects 36 - 39 further includes a transceiver configured to transmit an identification signal indicative of its estimated position to a remote image capture device, and the remote image capture device is configured to update its estimated position based on the received identification signal.
[0126] (Other Considerations) The processes, methods, and algorithms described in this specification and / or depicted in the attached figures are each embodied in code modules that are configured to execute specific and particular computer instructions by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware, thereby being fully or partially automated. For example, a computing system can include a general-purpose computer (e.g., a server) or a dedicated computer, a dedicated circuit, etc. programmed with specific computer instructions. The code modules can be installed in a dynamic link library that is compiled and linked into an executable program, or can be written in an interpreted programming language. In some implementations, certain operations and methods can be performed by circuits specific to a given function.
[0127] Furthermore, the functional implementations of the present disclosure are sufficiently mathematically, computationally, or technically complex such that application-specific hardware or one or more physical computing devices or special graphics processing units (utilizing appropriate specialized executable instructions) may be required to implement the functionality, for example, due to or as a result of the amount or complexity of the calculations involved, or to provide substantially real-time pose estimation inputs. For example, a video can contain many frames, each frame can have millions of pixels, and specifically programmed computer hardware is required to process the video data to provide the desired image processing tasks or applications in a commercially reasonable amount of time.
[0128] A code module or any type of data can be stored on any type of non-transitory computer-readable medium, such as a physical computer storage device including a hard drive, solid state memory, random access memory (RAM), read only memory (ROM), optical disk, volatile or non-volatile storage device, combinations of the same, and / or equivalents. The methods and modules (or data) can also be transmitted as data signals generated on various computer-readable transmission media, including wireless-based and wire / cable-based media (e.g., as part of a carrier wave or other analog or digital propagated signal), and can take various forms (e.g., as part of a single or multiplexed analog signal or as multiple discrete digital packets or frames). The results of the disclosed process or process steps can be persistently or otherwise stored within any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.
[0129] Any process, block, state, step, or functionality in the flow diagrams described in and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in a process. The various processes, blocks, states, steps, or functionality can be combined, rearranged, added, deleted, modified, or otherwise changed from the exemplary embodiments provided herein. In some embodiments, additional or different computing systems or code modules may implement some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the associated blocks, steps, or states can be implemented in a suitable other sequence, e.g., continuously, in parallel, or in some other manner. Tasks or events can be added to or removed from the disclosed exemplary embodiments. Further, the separation of the various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems can generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are possible.
[0130] The present process, method, and system can be implemented in a network (or distributed) computing environment. The network environment can include an enterprise-wide computer network, an intranet, a local area network (LAN), a wide area network (WAN), a personal area network (PAN), a cloud computing network, a cloud source computing network, the Internet, and the World Wide Web. The network can be a wired or wireless network or any other type of communication network.
[0131] The systems and methods of the present disclosure each have several innovative aspects, none of which alone contribute to or are required for the desirable attributes disclosed herein. The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Various modifications to the implementations described in the present disclosure may be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other implementations without departing from the spirit or scope of the present disclosure. Accordingly, the claims are not intended to be limited to the implementations shown herein, but should be accorded the widest scope consistent with the present disclosure, the principles, and the novel features disclosed herein.
[0132] Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable sub-combination. Further, a feature may be described above as acting in a certain combination and may further be initially claimed as such, but one or more features from the claimed combination can in some cases be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination. No single feature or group of features is necessary or essential to every embodiment.
[0133] In particular, conditional clauses used herein such as “can,” “could,” “might,” “may,” “e.g.,” and equivalents, generally convey that while one embodiment includes certain features, elements, and / or steps, other embodiments do not, unless specifically stated otherwise or understood otherwise in the context in which they are used. Thus, such conditional clauses are generally not intended to suggest that the features, elements, and / or steps are required in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are to be included or implemented in any particular embodiment, regardless of the author's input or prompting. The terms “comprising,” “including,” “having,” and equivalents are synonyms and are used inclusively in a non-limiting manner, without excluding additional elements, features, acts, operations, etc. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense), and thus, for example, when used to connect a list of elements, the term “or” means one, some, or all of the elements in the list. Additionally, the articles “a,” “an,” and “the” as used in this application and the appended claims should be construed to mean “one or more” or “at least one” unless otherwise defined.
[0134] As used herein, the phrase that refers to a list of items "at least one of" refers to any combination of those items, including a single element. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Connective phrases such as "at least one of X, Y, and Z" are generally understood in a context such that, unless specifically stated otherwise, they are used to convey that an item, term, etc. can be at least one of X, Y, or Z. Thus, such connective phrases are generally not intended to imply that an embodiment requires that at least one of X, at least one of Y, and at least one of Z each be present.
[0135] Similarly, operations may be depicted in the drawings in a particular order, but it should be recognized that this is not necessary for achieving the desired result, such that the operations are performed in the particular order in which they are shown, or in a sequential order, or that all of the illustrated operations be performed. Additionally, the drawings may schematically depict one or more exemplary processes in the form of a flowchart. However, other operations that are not depicted can also be incorporated within the exemplary methods and processes that are schematically illustrated. For example, one or more additional operations can be performed before, after, simultaneously with, or between any of the illustrated operations. In addition, the operations can be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing can be advantageous. Further, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the program components and systems described generally can be integrated together in a single software product or packaged into multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims are performed in a different order and still can achieve the desired result.
Claims
1. A method for estimating the position of an image capture device within an environment, the method comprising: continuously receiving a plurality of image segments, the plurality of image segments forming an image representing the field of view (FOV) of the image capture device, the FOV comprising a portion of the environment surrounding the image capture device that includes a plurality of sparse points, each sparse point being identifiable based at least in part on a corresponding subset of the plurality of image segments, and continuously receiving the plurality of image segments including receiving at least a first image segment and a second image segment; storing the first image segment in a buffer in response to receiving the first image segment; storing the second image segment in the buffer and removing the first image segment from the buffer in response to receiving the second image segment; continuously identifying one or more of the plurality of sparse points when each subset of the plurality of image segments corresponding to the plurality of sparse points is received; estimating the position of the image capture device within the environment based on the one or more continuously identified sparse points comprising a method.
2. The method according to claim 1, wherein the image capture device is a user wearable device and forms part of a user wearable system.
3. The method according to claim 2, wherein the buffer is part of the user wearable device.
4. Continuously identifying one or more of the plurality of sparse points when each subset of the plurality of image segments corresponding to the plurality of sparse points is received further includes continuously identifying one or more sparse points of a first group when a first plurality of image segments corresponding to the one or more sparse points of the first group are received, and continuously identifying one or more sparse points of a second group when a second plurality of image segments corresponding to the one or more sparse points of the second group are received, the second plurality of image segments being received after the first plurality of image segments, the method according to claim 1.
5. Estimating the position of the image capture device is further based on successively identifying one or more sparse points of the first group, the first group including a certain number of sparse points, the method according to claim 4.
6. The method according to claim 5, wherein the number of sparse points is from 2 to 20.
7. The method according to claim 5, wherein the number of sparse points is from 10 to 20.
8. The method according to claim 5, further comprising updating the position of the image capture device based on identifying one or more sparse points of a second group.
9. The method according to claim 8, wherein the one or more sparse points of the second group include a second number of sparse points.
10. The method according to claim 5, further comprising updating the position of the image capture device based on identifying a certain number of the successively identified one or more sparse points.
11. The method according to claim 10, wherein the number of successively identified sparse points is equal to the number of sparse points.
12. The method according to claim 10, wherein the number of successively identified sparse points includes at least one of the sparse points of the first group.
13. The method according to claim 1, wherein the plurality of sparse points are extracted based on at least one of a real-world object, a virtual image element, and an invisible indicator projected into the environment.
14. The method according to claim 1, further comprising extracting the successively identified one or more sparse points from the buffer and transmitting the successively identified one or more sparse points to a visual simultaneous localization and mapping (VSLAM) system, the VSLAM system estimating the position of the image capture device based on the successively identified one or more sparse points.
Citation Information
Patent Citations
Low latency stabilization for head-worn displays
US20160035139A1
Scanning window in hardware for low-power object-detection in images
US20160092735A1