Feature association
By introducing independent projectors to project patterns into passive stereo vision systems and 6DoF systems, the difficulty of determining distance and pose in visually uniform scenes is solved, and more accurate feature association is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-27
AI Technical Summary
Passive stereo vision systems and 6DoF systems struggle to determine distance or pose when faced with visually uniform objects, leading to difficulties in feature association.
A projector independent of the passive stereo vision system and the 6DoF system is used to project the pattern onto the scene, so that the passive stereo vision system can determine the distance and the 6DoF system can determine the pose.
It improves the feature association capabilities of passive stereo vision systems and 6DoF systems in visually uniform scenes, accurately determining distance and pose.
Smart Images

Figure CN121752870A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates in general to feature association. For example, aspects of this disclosure include systems and techniques for improving the ability of a device to associate features captured in an image. Background Technology
[0002] Passive stereo vision systems can use two cameras spaced a predetermined distance apart to capture paired stereo images of a scene. Passive stereo vision systems can correlate features within the images and determine the corresponding depth to points in the scene represented by those features. For example, a passive stereo vision system can determine the distance to a point in the scene based on the location of unique features in each image and the predetermined distance between the cameras.
[0003] Based on Visual Simultaneous Localization and Mapping (VSLAM or SLAM) technology, a six-degrees-of-freedom (6DoF) system can capture successive images of a scene and track the localization of unique features between these images. A 6DoF system can assume that the unique features are fixed and that any changes in the localization of unique features between successive images are based on the movement or reorientation of the 6DoF system. The 6DoF system can calculate changes in its pose based on these changes in the localization of unique features between successive images. Summary of the Invention
[0004] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description presented below.
[0005] Systems and techniques for improved feature association are described. According to at least one example, an apparatus for improved feature association is provided. The apparatus includes: a projector configured to project a pattern onto a scene for feature association by an imaging device that captures an image of the pattern projected onto the scene; wherein the apparatus is separate from the imaging device.
[0006] In another example, a method for improved feature association is provided. The method includes: determining to project a pattern into a scene for feature association by an imaging device that captures an image of the pattern projected into the scene; and projecting the pattern into the scene from a projector separate from the imaging device.
[0007] In another example, an apparatus for improved feature association is provided, the apparatus including at least one memory and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: determine to project a pattern into a scene for feature association by an imaging device that captures an image of the pattern projected into the scene; and project the pattern into the scene from a projector separate from the imaging device.
[0008] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: determine to project a pattern into a scene for feature association by an imaging device that captures an image of the pattern projected into the scene; and to project the pattern into the scene from a projector separate from the imaging device.
[0009] In another example, an apparatus for improved feature association is provided. The apparatus includes: components for determining a pattern to be projected into a scene for feature association by an imaging device, the imaging device capturing an image of the pattern projected into the scene; and components for projecting the pattern into the scene from a projector detached from the imaging device.
[0010] In some aspects, one or more of the devices described herein are, may be part of, or may include: mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices), extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), vehicles (or computing devices or systems of vehicles), smart or connected devices (e.g., Internet of Things (IoT) devices), wearable devices, personal computers, laptop computers, video servers, televisions (e.g., network-connected televisions), robotic devices or systems, or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, the one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level and / or another state) and / or for other purposes.
[0011] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.
[0012] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description
[0013] The following description, with reference to the accompanying drawings, details exemplary examples of this application: Figure 1A This is a perspective view illustrating various aspects of a head-mounted display (HMD) according to the present disclosure; Figure 1B This illustrates various aspects of what is being worn by a user according to this disclosure. Figure 1A A perspective view of a head-mounted display (HMD); Figure 2 These are illustrations of examples of extended reality (XR) systems according to various aspects of this disclosure; Figure 3 This is a block diagram illustrating the architecture of an example XR system according to some aspects of this disclosure; Figure 4 This is a block diagram illustrating the architecture of a simultaneous localization and mapping (SLAM) system according to various aspects of this disclosure; Figure 5 Two images of a scene captured from different camera positions according to various aspects of this disclosure are illustrated; Figure 6 Two images and associated cost functions are illustrated according to various aspects of this disclosure; Figure 7 This is an illustration of an example environment in which pose and / or distance determination can be achieved according to various aspects of the present disclosure; Figure 8 Examples of various aspects according to this disclosure are shown. Figure 7 The projector can project adjusted patterns onto four different scenarios within the scene; Figure 9 This is a block diagram illustrating an example architecture of an example projector according to various aspects of this disclosure; Figure 10 This is an illustration of an example environment in which pose and / or distance determination can be achieved according to various aspects of the present disclosure; Figure 11 This is a flowchart illustrating another example process for achieving pose and / or distance determination according to various aspects of this disclosure; Figure 12This illustrates one aspect of a theme based on a particular aspect.
[0014] Figure 13 This is a block diagram illustrating an example computing device architecture that can implement the various technologies described herein. Detailed Implementation
[0015] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation to provide a thorough understanding of the various aspects of this application. However, it will be apparent, however, that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.
[0016] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with descriptions that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0017] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.
[0018] As described above, a passive stereo vision system can correlate unique features within stereo pairs of images of a scene and determine the corresponding depth to points in the scene represented by these unique features based on the localization of these unique features and the distance between cameras that captured the stereo pairs. For a passive stereo vision system to determine the depth to points in a scene, these points should be represented by visually unique features, enabling the passive stereo vision system to correlate features between stereo pairs. A passive stereo vision system may have difficulty determining the distance between itself and objects lacking visually distinct features. For example, a passive stereo vision system may have difficulty determining the distance between itself and points on a blank wall. A blank wall may be visually uniform and therefore may lack visually unique features. The passive stereo vision system may not be able to correlate the features of the blank wall between stereo pairs of images of the blank wall. Because the passive stereo vision system cannot correlate features, it may be unable to determine the distance between itself and the blank wall.
[0019] Visual Simultaneous Localization and Mapping (VSLAM, or SLAM) is a computational geometry technique used in devices with cameras, such as robots, extended reality (XR) devices (e.g., head-mounted displays (HMDs)), mobile phones, autonomous vehicles, etc. In VSLAM, a device can build and update a map of an unknown environment based on images captured by its cameras. As the device updates the map, it can maintain the device's pose (e.g., position and / or orientation) within the environment. For example, a device can be activated in a specific room of a building and can move throughout the building, capturing images. The device can build a map of the environment and track the location of different objects in the environment by tracking their positions in different images.
[0020] Degrees of freedom (DoF) refer to the number of fundamental ways in which a rigid object can move in three-dimensional (3D) space. In some cases, six different DoFs can be tracked. The six DoFs include three translational degrees of freedom corresponding to translational movement along three vertical axes. These three axes may be referred to as the x-axis, y-axis, and z-axis. The six DoFs also include three rotational degrees of freedom corresponding to rotational movement about three axes, which may be referred to as pitch, yaw, and roll. In this disclosure, the term "pose" can specify position (e.g., described with respect to the three translational degrees of freedom) and orientation (e.g., as described with respect to the three rotational degrees of freedom). Thus, the pose of an object can refer to the object's position and orientation according to the six DoFs.
[0021] In the context of systems that track movement through their environment (such as XR and / or VSLAM systems), a degree of freedom (DOF) can refer to which of the six degrees of freedom the system can track. A 3DoF system typically tracks three rotational DoFs—pitch, yaw, and roll. For example, a 3DoF headset can track a user turning their head left or right, tilting their head up or down, and / or tilting their head left or right. A 6DoF system tracks three translational DoFs and three rotational DoFs. Therefore, a 6DoF system can track a user's forward, backward, lateral, and / or vertical movements in addition to the three rotational DoFs.
[0022] A 6DoF system using VSLAM can compute changes in its pose based on variations in the localization of unique features captured in successive images of a scene. For a 6DoF system to determine its pose changes, the unique features should be fixed in the scene and visually distinct (e.g., so that the 6DoF system can easily identify unique features in successive images, even if the unique features change localization between subsequent images). Similar to passive stereo vision systems, a 6DoF system may struggle to determine its pose when its camera is facing objects that do not have distinct features (e.g., a blank wall). For example, a blank wall may be visually uniform and therefore may lack visually unique features. The 6DoF system may fail to identify and / or correlate unique features between successive images. Because the 6DoF system cannot correlate unique features between successive images, it may be unable to determine its pose based on images of a blank wall.
[0023] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, “Systems and Techniques”) for achieving pose and / or distance determination. The systems and techniques described herein can project patterns (e.g., unique patterns) onto a scene so that a passive stereo vision system can correlate patterns in stereo pairs of images of the scene to determine the distance between the passive stereo vision system and a point in the scene. Additionally or alternatively, the system and techniques can project patterns onto a scene so that a 6DoF system can track patterns in images of the scene to determine the pose of the 6DoF system relative to the scene.
[0024] This system and technology may include a projector. The projector can be independent of both the passive stereo vision system and the 6DoF system. For example, the projector can be decoupled from both the passive stereo vision system and the 6DoF system. Additionally, the projector can project a pattern independently of both the passive stereo vision system and the 6DoF system. For example, the projector can project the pattern onto the scene regardless of whether the passive stereo vision system is present in the scene and / or independently of whether the passive stereo vision system is capturing stereo pairs of images of the pattern. Additionally, the projector can project the pattern onto the scene regardless of whether the 6DoF system is present in the scene and / or independently of whether the 6DoF system is capturing continuous images of the pattern.
[0025] A passive stereo vision system can determine the distance between itself and points in a scene independently of the projector. For example, the passive stereo vision system can determine the distance between itself and points in the scene regardless of whether the projector is projecting a pattern onto the scene. The projector projecting a pattern onto the scene allows the passive stereo vision system to more accurately determine the distance between itself and points in the scene on which the pattern is projected. For example, the projector can project a pattern onto a visually indistinguishable part of the scene. Compared to when the projector is not projecting a pattern onto a visually indistinguishable part of the scene, the passive stereo vision system is better able to correlate features in stereo pairs of visually indistinguishable parts of the scene with the pattern projected onto the scene. Furthermore, because the passive stereo vision system operates independently of the projector, it can determine the distance between itself and points in the scene without prior information about the projector and / or the pattern.
[0026] Additionally or alternatively, a 6DoF system can determine its pose independently of the projector. For example, a 6DoF system can determine its pose independently of whether the projector is projecting a pattern onto the scene. The projector projecting the pattern onto the scene allows the 6DoF system to determine its pose more accurately. For example, the projector can project the pattern onto a visually indistinguishable part of the scene. Compared to when the projector is not projecting the pattern onto a visually indistinguishable part of the scene, the 6DoF system is better able to associate features in a continuous image of the visually indistinguishable part of the scene with the pattern projected onto the scene. Furthermore, because the 6DoF system operates independently of the projector, it can determine its pose without prior information about the projector and / or the pattern.
[0027] Furthermore, the projector can be manipulated independently of the passive stereoscopic vision system and independently of the 6DoF system. For example, the projector can be manipulated (e.g., pointed at) a point or surface in the scene independently of the position where the passive stereoscopic vision system is manipulated and independently of the position where the 6DoF system is manipulated. For example, the projector can be pointed at a visually indistinguishable part of the scene (e.g., a blank wall) independently of the position pointed at by the passive stereoscopic vision system and / or independently of the position pointed at by the 6DoF system to project a pattern onto the visually indistinguishable part of the scene.
[0028] In some respects, the projector can be fixed within the scene. For example, the projector can be positioned within the scene to point towards a visually indistinguishable part of the scene (e.g., a blank wall) to project a pattern onto that visually indistinguishable part. The passive stereoscopic vision system can move within the scene, and the distance between the passive stereoscopic vision system and the visually indistinguishable part of the scene is determined based on stereo pairs of images of the pattern projected onto the originally visually indistinguishable part of the scene. Additionally or alternatively, the 6DoF system can move within the scene, and the pose of the 6DoF system is determined based on successive images of the pattern projected onto the originally visually indistinguishable part of the scene.
[0029] In other respects, the projector can move within the scene. For example, the projector can move within the scene and can be pointed at one or more visually indistinguishable parts of the scene (e.g., one or more blank walls) to project a pattern onto one or more visually indistinguishable parts of the scene as the projector moves through the scene. The passive stereoscopic vision system can move within the scene and determine the distance between the passive stereoscopic vision system and one or more visually indistinguishable parts of the scene based on stereo pairs of images of the pattern projected onto the originally visually indistinguishable parts of the scene. Additionally or alternatively, the 6DoF system can move within the scene and determine the pose of the 6DoF system based on successive images of the pattern projected onto the originally visually indistinguishable parts of the scene.
[0030] Various aspects of this application will be described below with reference to the accompanying drawings. For example, Figures 1A to 4 (And the corresponding text) provides examples of 6DoF systems and technologies. Figure 5 and Figure 6 (And the corresponding text) provides examples of passive stereo vision technology. Figures 7 to 13 Examples of systems and techniques for achieving pose and / or distance determination are provided according to various aspects of this disclosure.
[0031] Specifically, Figure 1A This is a perspective view illustrating various aspects of a head-mounted display (HMD) 100 according to the present disclosure. The HMD 100 may be, for example, an augmented reality (AR) head-mounted device, a virtual reality (VR) head-mounted device, a mixed reality (MR) head-mounted device, an extended reality (XR) head-mounted device, or some combination thereof. The HMD 100 may be... Figure 3 XR system 300, Figure 4 Examples of SLAM systems 400 or combinations thereof, or those that can be implemented Figure 3 XR system 300, Figure 4The SLAM system 400 or a combination thereof. The HMD 100 includes a first camera 102 and a second camera 104 along the front of the HMD 100. The first camera 102 and the second camera 104 may be... Figure 4 The SLAM system 400 includes two cameras from one or more cameras 404. In some examples, the HMD 100 may have only a single camera. In some examples, the HMD 100 may also include one or more additional cameras in addition to the first camera 102 and the second camera 104. In some aspects, the HMD 100 may include one or more additional sensors, such as, for example, an inertial measurement unit (IMU).
[0032] Figure 1B This illustrates various aspects of what is being worn by user 106 according to this disclosure. Figure 1A A perspective view of a head-mounted display (HMD) 100. User 106 wears the HMD 100 on their head and over their eyes. The HMD 100 can capture images using a first camera 102 and a second camera 104. In some examples, the HMD 100 can display one or more display images to the eyes of user 106. The display images may be based on images captured by the first camera 102 and / or the second camera 104. The display images can provide a stereoscopic view of the environment, and in some cases have overlaid information and / or other modifications. For example, the HMD 100 can display a first display image to the left eye of user 106, which is based on an image captured by the first camera 102. The HMD 100 can display a second display image to the right eye of user 106, which is based on an image captured by the second camera 104. For example, the HMD 100 can provide overlay information superimposed on the images captured by the first camera 102 and / or the second camera 104 in the display images.
[0033] HMD 100 can determine the pose of HMD 100, and is therefore provided as an example of a 6DoF system. HMD 100 can determine the pose of HMD 100 using visual simultaneous localization and mapping (VSLAM or SLAM) techniques (e.g., based on successive images captured by first camera 102 and / or second camera 104).
[0034] Figure 2 This is a diagram illustrating an example of an extended reality (XR) system 200 according to various aspects of this disclosure. The XR system 200 may be, for example, an augmented reality (AR) head-mounted device, a virtual reality (VR) head-mounted device, a mixed reality (MR) head-mounted device, an extended reality (XR) head-mounted device, or some combination thereof. The XR system 200 may be... Figure 3 XR system 300, Figure 4Examples of SLAM systems 400 or combinations thereof, or those that can be implemented Figure 3 XR system 300, Figure 4 SLAM systems 400 or combinations thereof.
[0035] As shown in the figure, the XR system 200 includes an XR device 202, an accessory device 204, and a communication link 206 between the XR device 202 and the accessory device 204. In some cases, the XR device 202 typically implements aspects of extended reality (including virtual reality (VR), augmented reality (AR), mixed reality (MR), etc.) display, image capture, and / or view tracking. In some cases, the accessory device 204 typically implements aspects of extended reality computation. For example, the XR device 202 may capture an image of the environment of the user 208 and (e.g., via communication link 206) provide the image to the accessory device 204. The accessory device 204 may render virtual content (e.g., in relation to the captured image of the environment) and (e.g., via communication link 206) provide the virtual content to the XR device 202. The XR device 202 may display the virtual content to the user 208 (e.g., within the user 208's field of view 210).
[0036] Typically, XR device 202 can display virtual content within a field of view 210 for viewing by user 208. In some examples, XR device 202 may include a transparent surface (e.g., optical glass) such that virtual objects can be displayed on the transparent surface (e.g., by projection onto the transparent surface) to overlay virtual content onto real-world objects viewed through the transparent surface (e.g., in a perspective configuration). In some cases, XR device 202 may include a camera and can display both real-world objects (e.g., as frames or images captured by the camera) and virtual objects overlaid on the displayed real-world objects (e.g., in a pass-through configuration). In various examples, XR device 202 may include aspects of a virtual reality headset, smart glasses, a live-feed video camera, a GPU, one or more sensors (e.g., such as one or more inertial measurement units (IMUs), image sensors, microphones, etc.), one or more output devices (e.g., such as speakers, displays, smart glasses, etc.).
[0037] The companion device 204 can render virtual content to be displayed by the companion device 204. In some examples, the companion device 204 may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, or a mobile device acting as a server device), any other computing device and / or combinations thereof.
[0038] The communication link 206 can be based on any suitable wireless protocol, such as, for example, IEEE 802.11 (Wi-Fi), IEEE 802.15, or Bluetooth. ™ The communication link 206 can be a direct wireless connection between the XR device 202 and the companion device 204 in some cases. In other cases, the communication link 206 can be via one or more intermediate devices, such as routers or switches and / or across a network.
[0039] Similar to the HMD 100, the XR system 200 (or its companion device 204) can determine the pose of the user 208 based on six degrees of freedom. Therefore, the XR system 200 is provided as an example of a 6DoF system. The XR system 200 (or its companion device 204) can determine the pose of the user 208 using visual simultaneous localization and mapping (VSLAM or SLAM) techniques (e.g., based on sequential images captured by one or more cameras of the XR device 202).
[0040] Figure 3 This is a diagram illustrating the architecture of an example extended reality (XR) system 300 according to some aspects of this disclosure. The XR system 300 can execute XR applications and implement XR operations. The XR system 300 can be... Figure 1A and Figure 1B HMD 100 and / or Figure 2 An example of an XR system 200, or you can see... Figure 1A and Figure 1B HMD 100 and / or Figure 2 The XR system is implemented in 200.
[0041] In this exemplary example, the XR system 300 includes one or more image sensors 302, accelerometers 304, gyroscopes 306, storage devices 308, input devices 310, displays 312, computing components 314, XR engines 324, image processing engines 326, rendering engines 328, and communication engines 330. It should be noted that... Figure 3 The components 302-330 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include those with... Figure 3The components shown may be more numerous, fewer, or different than those shown. For example, in some cases, the XR system 300 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 3 One or more other software and / or hardware components not shown herein. While various components of the XR system 300 (such as image sensor 302) may be referred to in the singular herein, it should be understood that the XR system 300 may include multiple components discussed herein (e.g., multiple image sensors 302).
[0042] The display 312 may be or may include glass, screen, lens, projector and / or other display mechanism that allows users to see a real-world environment and also allows XR content to be overlaid on, overlapped with, mixed with or otherwise displayed on the real-world environment.
[0043] XR system 300 may include input device 310 or be able to communicate with such input device (wired or wireless). Input device 310 may include any suitable input device, such as a touchscreen, pen or other pointing device, keyboard, mouse, buttons or keys, microphone for receiving voice commands, gesture input device for receiving gesture commands, video game controller, steering wheel, joystick, set of buttons, trackball, remote control, any other input device discussed herein, or any combination thereof. In some cases, image sensor 302 may capture images that can be processed to interpret gesture commands.
[0044] The XR system 300 can also communicate with one or more other electronic devices (wired or wireless). For example, the communication engine 330 can be configured to manage connections and communicate with one or more electronic devices. In some cases, the communication engine 330 may correspond to... Figure 13 The communication interface is 1326.
[0045] In some specific implementations, the image sensor 302, accelerometer 304, gyroscope 306, storage device 308, display 312, computing component 314, XR engine 324, image processing engine 326, and rendering engine 328 can be the same computing device (such as...). Figure 1A and Figure 1BAs part of the HMD 100. For example, in some cases, the image sensor 302, accelerometer 304, gyroscope 306, storage device 308, display 312, computing component 314, XR engine 324, image processing engine 326, and rendering engine 328 may be integrated into the HMD, extended reality glasses, smartphone, laptop computer, tablet computer, gaming system, and / or any other computing device.
[0046] In other specific implementations, the image sensor 302, accelerometer 304, gyroscope 306, storage device 308, display 312, computing component 314, XR engine 324, image processing engine 326, and rendering engine 328 may be part of two or more independent computing devices. For example, in some cases, some components of components 302-330 may be part of or implemented by one computing device, and the remaining components may be part of or implemented by one or more other computing devices. For example, such as in a discrete sensing XR system, XR system 300 may include a first device (such as...) Figure 2 The first device (XR device 202) includes a display 312, an image sensor 302, an accelerometer 304, a gyroscope 306, and / or one or more computing components 314. The XR system 300 may also include a second device (such as...). Figure 2 The second device includes an auxiliary computing component 314 (e.g., implementing an XR engine 324, an image processing engine 326, a rendering engine 328, and / or a communication engine 330). In such examples, the second device may generate virtual content based on information or data (e.g., images, sensor data, such as measurements from accelerometer 304 and gyroscope 306) and may provide the virtual content to the first device for display at the first device. The second device may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server equipment (e.g., an edge or cloud-based server, a personal computer acting as a server equipment, or a mobile device acting as a server equipment), any other computing device, and / or combinations thereof.
[0047] Storage device 308 can be any storage device used for storing data. Furthermore, storage device 308 can store data from any component of the XR system 300. For example, storage device 308 can store data from image sensor 302 (e.g., image or video data), data from accelerometer 304 (e.g., measurements), data from gyroscope 306 (e.g., measurements), data from computing component 314 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from XR engine 324, data from image processing engine 326, and / or data from rendering engine 328 (e.g., output frames). In some examples, storage device 308 may include a buffer for storing frames processed by computing component 314.
[0048] Computing component 314 may be or may include a central processing unit (CPU) 316, a graphics processing unit (GPU) 318, a digital signal processor (DSP) 320, an image signal processor (ISP) 322, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). Computing component 314 may perform various operations such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, localization, pose estimation, map building, content anchoring, content rendering, prediction, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, computing component 314 may implement (e.g., control, operate, etc.) an XR engine 324, an image processing engine 326, and a rendering engine 328. In other examples, computing component 314 may also implement one or more other processing engines.
[0049] Image sensor 302 may include any image and / or video sensor or capture device. In some examples, image sensor 302 may be part of a multi-camera assembly, such as a dual-camera assembly. Image sensor 302 may capture image and / or video content (e.g., raw image and / or video data), which may then be processed by computing component 314, XR engine 324, image processing engine 326, and / or rendering engine 328, as described herein.
[0050] In some examples, image sensor 302 may capture image data and may generate an image (also referred to as a frame) based on that image data and / or provide the image data or frame to XR engine 324, image processing engine 326, and / or rendering engine 328 for processing. The image or frame may include a video frame in a video sequence or a still image. The image or frame may include an array of pixels representing a scene. For example, the image may be: a red-green-blue (RGB) image with red, green, and blue color components per pixel; a lightness, redness, and blueness (YCbCr) image with a lightness component and two chromaticity (redness and blueness) components per pixel; or any other suitable type of color or monochrome image.
[0051] In some cases, the image sensor 302 (and / or other cameras of the XR system 300) may be configured to also capture depth information. For example, in some implementations, the image sensor 302 (and / or other cameras) may include an RGB depth (RGB-D) camera. In some cases, the XR system 300 may include one or more depth sensors (not shown) that are separate from the image sensor 302 (and / or other cameras) and can capture depth information. For example, such depth sensors may acquire depth information independently of the image sensor 302. In some examples, the depth sensor may be physically mounted in the same general location or orientation as the image sensor 302, but may operate at a different frequency or frame rate than the image sensor 302. In some examples, the depth sensor may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in the scene. Depth information can then be obtained by utilizing the geometric deformation of the projected pattern caused by the surface shape of the objects. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).
[0052] The XR system 300 may also include other sensors among its one or more sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 304), one or more gyroscopes (e.g., gyroscope 306), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other positioning-related information to the computing component 314. For example, accelerometer 304 may detect the acceleration of the XR system 300 and may generate an acceleration measurement based on the detected acceleration. In some cases, accelerometer 304 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward) that can be used to determine the positioning or pose of the XR system 300. Gyroscope 306 may detect and measure the orientation and angular velocity of the XR system 300. For example, gyroscope 306 may be used to measure the pitch, roll, and yaw of the XR system 300. In some cases, gyroscope 306 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 302 and / or the XR engine 324 may use measurements obtained by the accelerometer 304 (e.g., one or more translation vectors) and / or measurements obtained by the gyroscope 306 (e.g., one or more rotation vectors) to calculate the pose of the XR system 300. As previously mentioned, in other examples, the XR system 300 may also include other sensors such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye-tracking sensor, a machine vision sensor, a smart scene sensor, a voice recognition sensor, a shock sensor, a vibration sensor, a positioning sensor, a tilt sensor, etc.
[0053] As described above, in some cases, one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure specific forces, angular velocities, and / or orientations of the XR system 300. In some examples, one or more sensors may output measurement information associated with the capture of images by the image sensor 302 (and / or other cameras of the XR system 300) and / or depth information obtained using one or more depth sensors of the XR system 300.
[0054] The XR engine 324 can use the output of one or more sensors (e.g., accelerometer 304, gyroscope 306, one or more IMUs and / or other sensors) to determine the pose of the XR system 300 (also referred to as head pose) and / or the pose of the image sensor 302 (or other cameras of the XR system 300). In some cases, the pose of the XR system 300 and the pose of the image sensor 302 (or other cameras) can be the same. The pose of the image sensor 302 refers to the pose of the image sensor 302 relative to a reference frame (e.g., relative to...). Figure 2The camera pose (210) is determined for the position and orientation of the field of view. In some implementations, the camera pose can be determined for 6 degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, the camera pose can be determined for 3 degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).
[0055] In some cases, a device tracker (not shown) may use measurements from one or more sensors and image data from image sensor 302 to track the pose (e.g., 6DoF pose) of the XR system 300. For example, the device tracker may fuse visual data from image data (e.g., using a visual tracking solution) with inertial data from measurements to determine the position and motion of the XR system 300 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 300, the device tracker may generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map for that scene. 3D map updates may include, for example, but not limited to, new or updated features and / or features or landmarks associated with the scene and / or the 3D map of that scene, positioning updates identifying or updating the position of the XR system 300 within the scene and the 3D map of that scene, etc. The 3D map provides a digital representation of the scene in the real / physical world. In some examples, 3D maps can anchor location-based objects and / or content to real-world coordinates and / or objects. XR system 200 can use mapped scenes (e.g., scenes in the physical world represented by a 3D map and / or scenes associated with that 3D map) to merge the physical and virtual worlds and / or merge virtual content or objects with the physical environment.
[0056] In some aspects, computing component 314 may use a visual tracking solution to determine and / or track the pose of image sensor 302 and / or the XR system 300 as a whole, based on images captured by image sensor 302 (and / or other cameras of XR system 300). For example, in some examples, computing component 314 may perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques. For example, computing component 314 may perform SLAM or may communicate (wired or wirelessly) with a SLAM engine (not shown). SLAM refers to a class of techniques in which a map of the environment (e.g., a map of the environment modeled by system 300) is created while simultaneously tracking the pose of the camera (e.g., image sensor 302) and / or XR system 300 relative to that map. This map may be called a SLAM map and may be three-dimensional (3D). SLAM technology can be performed using color or grayscale image data captured by image sensor 302 (and / or other cameras of XR system 300) and can be used to generate an estimate of 6DoF pose measurement of image sensor 302 and / or XR system 300. Such SLAM technology configured to perform 6DoF tracking can be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 304, gyroscope 306, one or more IMUs and / or other sensors) can be used to estimate, correct, and / or otherwise adjust the estimated pose.
[0057] In some cases, 6DoF SLAM (e.g., 6DoF tracking) can associate features observed from certain input images from image sensor 302 (and / or other cameras) with a SLAM map. For example, 6DoF SLAM can use feature point association from an input image to determine the pose (localization and orientation) of the image sensor 302 and / or XR system 300 for that input image. 6DoF map construction can also be performed to update the SLAM map. In some cases, the SLAM map maintained using 6DoF SLAM can contain 3D feature points triangulated from two or more images. For example, keyframes can be selected from an input image or video stream to represent the observed scene. For each keyframe, a corresponding 6DoF camera pose associated with the image can be determined. The pose of the image sensor 302 and / or XR system 300 can be determined by projecting features from the 3D SLAM map onto the image or video frame and updating the camera pose according to a verified 2D-3D correspondence.
[0058] In an exemplary example, computation component 314 may extract feature points from some input images (e.g., each input image, a subset of input images, etc.) or from each keyframe. Feature points (also referred to as registration points) as used herein are distinctive or identifiable parts of an image, such as a part of a hand, the edge of a table, etc. Features extracted from a captured image may represent different feature points along three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point may have an associated feature location. Feature points in a keyframe may match (be identical to or correspond to) feature points in previously captured input images or keyframes, or may not match. Feature detection may be used to detect feature points. Feature detection may include image processing operations that examine one or more pixels of an image to determine if a feature exists at a particular pixel. Feature detection may be used to process the entire captured image or parts of an image. For each image or keyframe, once a feature has been detected, a local image patch around that feature can be extracted. Any suitable technique can be used to extract features, such as Scale Invariant Feature Transform (SIFT) (which localizes features and generates descriptions of them), Learned Invariant Feature Transform (LIFT), Speeded Robust Features (SURF), Gradient Position Orientation Histogram (GLOH), Orientation Fast Rotation Briefing (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoints (FREAK), KAZE, Accelerated KAZE (AKAZE), Normalized Cross-Correlation (NCC), Descriptor Matching, another suitable technique, or a combination thereof.
[0059] As an illustrative example, computing component 314 can extract feature points corresponding to mobile devices, etc. In some cases, feature points corresponding to mobile devices can be tracked to determine the pose of the mobile device. As described in more detail below, the pose of the mobile device can be used to determine the position of the projection of AR media content that can enhance the media content displayed on the display of the mobile device.
[0060] In some cases, the XR system 300 may also track the user's hands and / or fingers to allow the user to interact with and / or control virtual content in the virtual environment. For example, the XR system 300 may track the pose and / or movement of the user's hands and / or fingertips to identify or translate the user's interaction with the virtual environment. User interaction may include, for example, but not limited to, moving virtual content items, resizing virtual content items, selecting input interface elements in the virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), and providing input through the virtual user interface.
[0061] Figure 4This is a block diagram illustrating the architecture of a Simultaneous Localization and Mapping (SLAM) system 400. In some examples, the SLAM system 400 may be or may include an extended reality (XR) system (such as...). Figure 2 The XR system 400. In some examples, the SLAM system 400 may be a wireless communication device, a mobile device or mobile phone (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, a personal computer, a laptop computer, a server computer, a portable video game console, a portable media player, a camera device, a manned or unmanned ground vehicle, a manned or unmanned aircraft, a manned or unmanned water vehicle, a manned or unmanned underwater vehicle, a manned or unmanned vehicle, an autonomous vehicle, a vehicle, a vehicle's computing system, a robot, another device, or any combination thereof.
[0062] Figure 4 The SLAM system 400 includes or is coupled to each of one or more sensors 402. Sensors 402 may include one or more cameras 404. Each camera 404 may include an image capture device, an image processing device, an image capture and processing system, another type of camera, or a combination thereof. Each camera 404 may respond to light from a specific spectrum. The spectrum may be a subset of the electromagnetic (EM) spectrum. For example, each camera 404 may be a VL camera that responds to the visible (VL) spectrum, an IR camera that responds to the infrared (IR) spectrum, a UV camera that responds to the ultraviolet (UV) spectrum, a camera that responds to light from another part of the electromagnetic spectrum, or some combination thereof.
[0063] Sensor 402 may include one or more other types of sensors besides camera 404, such as one or more of the following: accelerometer, gyroscope, magnetometer, inertial measurement unit (IMU), altimeter, barometer, thermometer, radio detection and ranging (RADAR) sensor, light detection and ranging (LIDAR) sensor, sound navigation and ranging (SONAR) sensor, sound detection and ranging (SODAR) sensor, Global Navigation Satellite System (GNSS) receiver, Global Positioning System (GPS) receiver, BeiDou Navigation Satellite System (BDS) receiver, Galileo receiver, GLONASS receiver, NavIC receiver, Quasi-Zenith Satellite System (QZSS) receiver, Wi-Fi positioning system (WPS) receiver, cellular network positioning system receiver, Bluetooth. ®Beacon positioning receivers, short-range wireless beacon positioning receivers, personal area network (PAN) positioning receivers, wide area network (WAN) positioning receivers, wireless local area network (WLAN) positioning receivers, other types of positioning receivers, other types of sensors discussed herein, or combinations thereof. In some examples, sensor 402 may include... Figure 2 Any combination of sensors for the XR system 200.
[0064] Figure 4 The SLAM system 400 includes a vision-inertial ranging (VIO) tracker 406. The term "vision-inertial ranging" may also be referred to herein as vision ranging. The VIO tracker 406 receives sensor data 426 from sensor 402. For example, sensor data 426 may include one or more images captured by camera 404. Sensor data 426 may include other types of sensor data from sensor 402, such as data from any type of sensor listed herein. For example, sensor data 426 may include IMU data from one or more inertial measurement units (IMUs) of sensor 402.
[0065] After receiving sensor data 426 from sensor 402, VIO tracker 406 performs feature detection, extraction, and / or tracking using VIO tracker 406 feature tracking engine 408. For example, if sensor data 426 comprises one or more images captured by camera 404 of SLAM system 400, VIO tracker 406 can identify, detect, and / or extract features in each image. Features may include visually distinguishable points in the image, such as portions depicting edges and / or corners in the image. VIO tracker 406 may receive sensor data 426 periodically and / or continuously from sensor 402, for example by continuing to receive more images from camera 404 while camera 404 is capturing video, where these images are video frames of the video. VIO tracker 406 may generate descriptors for the features. Feature descriptors may be generated at least in part by generating a description of the feature, such as depicted in a local image patch extracted around the feature. In some examples, the feature descriptor may describe the feature as a set of one or more feature vectors. In some cases, the VIO tracker 406, together with the map building engine 412 and / or the relocation engine 422, can associate multiple features with a map of the environment based on such feature descriptors. The feature tracking engine 408 of the VIO tracker 406 can perform feature tracking by recognizing features in each image that the VIO tracker 406 has previously identified in one or more previous images (in some cases, based on features identified in different images with matching feature descriptors). The feature tracking engine 408 can track changes in one or more locations where the features depicted in each of the different images are located. For example, the feature extraction engine can detect a specific corner of a depicted room on the left side of a first image captured by a first camera in camera 404. The feature extraction engine can detect the same depicted feature (e.g., the same specific corner of the same room) on the right side of a second image captured by the first camera. The feature tracking engine 408 can identify two depictions where the feature detected in the first and second images is the same feature (e.g., the same specific corner of the same room), and the feature appears in two different locations in the two images. VIO tracker 406 can determine that the first camera has moved, for example, if the feature (e.g., a specific corner of a room) depicts a static part of the environment, based on the same feature appearing on the left side of the first image and the right side of the second image.
[0066] VIO tracker 406 may include sensor integration engine 410. Sensor integration engine 410 may use sensor data from other types of sensors 402 (other than camera 404) to determine information that can be used by feature tracking engine 408 when performing feature tracking. For example, sensor integration engine 410 may receive IMU data from the IMU of sensor 402 (e.g., it may be included as part of sensor data 426). Sensor integration engine 410 may determine, based on the IMU data in sensor data 426, that the SLAM system 400 has rotated 15 degrees clockwise between the acquisition or capture of a first image by the first camera in camera 404 and the acquisition or capture of a second image. Based on this determination, sensor integration engine 410 may identify that a feature depicted at a first location in the first image is expected to appear at a second location in the second image, and that the second location is expected to be located at a predetermined distance to the left of the first location (e.g., a predetermined number of pixels, inches, centimeters, millimeters, or another distance measure). Feature tracking engine 408 may take this expectation into account when tracking features between the first and second images.
[0067] Based on feature tracking performed by feature tracking engine 408 and / or sensor integration performed by sensor integration engine 410, VIO tracker 406 can determine 3D feature localization 428 of a specific feature. 3D feature localization 428 may include one or more 3D feature localizations and may also be referred to as 3D feature points. 3D feature localization 428 may be a set of coordinates along three distinct axes perpendicular to each other, such as an X-coordinate along the X-axis (e.g., in the horizontal direction), a Y-coordinate along the Y-axis perpendicular to the X-axis (e.g., in the vertical direction), and a Z-coordinate along the Z-axis perpendicular to both the X-axis and Y-axis (e.g., in the depth direction). VIO tracker 406 can also determine one or more keyframes 430 (hereinafter referred to as keyframes 430) corresponding to a specific feature. A keyframe corresponding to a specific feature (from one or more keyframes 430) may be an image in which the specific feature is clearly depicted. In some examples, a keyframe corresponding to a specific feature (from one or more keyframes 430) may be an image in which the specific feature is clearly depicted. In some examples, a keyframe corresponding to a specific feature may be an image that reduces the uncertainty of the 3D feature localization 428 of that specific feature when considered by the feature tracking engine 408 and / or sensor integration engine 410 for determining the 3D feature localization 428. In some examples, the keyframe corresponding to a specific feature may also include data associated with the pose 436 of the SLAM system 400 and / or camera 404 during the capture of that keyframe. In some examples, the VIO tracker 406 may transmit 3D feature localization 428 and / or keyframe 430 corresponding to one or more features to the map building engine 412. In some examples, the VIO tracker 406 may receive map tiles 432 from the map building engine 412. The VIO tracker 406 may perform feature extraction on the information within the map tiles 432 for feature tracking using the feature tracking engine 408.
[0068] Based on feature tracking performed by feature tracking engine 408 and / or sensor integration performed by sensor integration engine 410, VIO tracker 406 can determine the pose 436 of SLAM system 400 and / or camera 404 during the capture of each image in the images in sensor data 426. Pose 436 may include the position of SLAM system 400 and / or camera 404 in 3D space, such as a set of coordinates (e.g., X, Y, and Z coordinates) along three distinct axes perpendicular to each other. Pose 436 may include the orientation of SLAM system 400 and / or camera 404 in 3D space, such as pitch, roll, yaw, or some combination thereof. In some examples, VIO tracker 406 may transmit pose 436 to relocalization engine 422. In some examples, VIO tracker 406 may receive pose 436 from relocalization engine 422.
[0069] The SLAM system 400 also includes a map-building engine 412. The map-building engine 412 generates a 3D map of the environment based on 3D feature localization 428 and / or keyframes 430 received from the VIO tracker 406. The map-building engine 412 may include a map densification engine 414, a keyframe remover 416, a bundle adjuster 418, and / or a loop closure detector 420. The map densification engine 414 performs map densification, increasing the number and / or density of 3D coordinates describing the map geometry in some examples. The keyframe remover 416 removes keyframes and / or adds keyframes in some cases. In some examples, the keyframe remover 416 removes keyframes 430 corresponding to areas in the map to be updated and / or areas with low confidence values. In some examples, the bundle adjuster 418 refines the 3D coordinates describing the scene geometry, parameters of relative motion, and / or optical characteristics of the image sensor used to generate the frames according to an optimality criterion involving the corresponding image projections of all points. The loop closure detector 420 can identify when the SLAM system 400 has returned to a previously mapped area, and can use this information to update map tiles and / or reduce uncertainty in certain 3D feature points or other points in the map geometry. The map building engine 412 can output map tiles 432 to the VIO tracker 406. Map tiles 432 can represent 3D portions or subsets of the map. Map tiles 432 can include map tiles 432 representing new, previously unmapped areas of the map. Map tiles 432 can include map tiles 432 representing updates (or modifications or revisions) to previously mapped areas of the map. The map building engine 412 can output map information 434 to the relocation engine 422. Map information 434 can include at least a portion of the map generated by the map building engine 412. Map information 434 can include one or more 3D points constituting the geometry of the map, such as one or more 3D feature locations 428. Map information 434 can include one or more keyframes 430 corresponding to certain features and certain 3D feature locations 428.
[0070] The SLAM system 400 also includes a relocalization engine 422. The relocalization engine 422 can perform relocalization, for example, when the VIO tracker 406 fails to identify more than a threshold number of features in an image, and / or when the VIO tracker 406 loses tracking of the pose 436 of the SLAM system 400 within a map generated by the map building engine 412. The relocalization engine 422 can perform relocalization by performing extraction and matching using an extraction and matching engine 424. For example, the extraction and matching engine 424 can extract features from an image captured by the camera 404 of the SLAM system 400 when the SLAM system 400 is in its current pose 436, and can match the extracted features with features depicted in different keyframes 430, identified by 3D feature localization 428, and / or identified in map information 434. By matching these extracted features with previously identified features, the relocalization engine 422 can identify the pose 436 of the SLAM system 400 as the pose 436 when the previously identified features were visible to the camera 404 of the SLAM system 400, and therefore similar to one or more previous poses 436 when the previously identified features were visible to the camera 404. In some cases, the relocalization engine 422 may perform relocalization based on a wide baseline map or the distance between the current camera location and the camera location when the features were initially captured. The relocalization engine 422 may receive information about the pose 436 (e.g., information about one or more recent poses of the SLAM system 400 and / or the camera 404) from the VIO tracker 406, and may use this information to make its relocalization determination. Once the relocalization engine 422 has relocalized the SLAM system 400 and / or the camera 404 and thus determined the pose 436, the relocalization engine 422 may output the pose 436 to the VIO tracker 406.
[0071] In some examples, VIO tracker 406 may modify the image in sensor data 426 before performing feature detection, extraction, and / or tracking on the modified image. For example, VIO tracker 406 may rescale and / or resample the image. In some examples, rescaling and / or resampling the image may include reducing the image size, downsampling, rescaling, and / or resampling it once or multiple times. In some examples, VIO tracker 406 modifying the image may include, for example, converting the image from color to grayscale or from color to black and white by desaturating colors in the image, removing certain color channels, reducing the color depth in the image, replacing colors in the image, or combinations thereof. In some examples, VIO tracker 406 modifying the image may include VIO tracker 406 masking certain areas of the image. Dynamic objects may include objects that have a changing appearance between one image and another. For example, a dynamic object may be an object that moves within an environment, such as a person, vehicle, or animal. A dynamic object may be an object that has a changing appearance at different times, such as a display screen that can display different things at different times. Dynamic objects can be objects whose appearance changes based on the pose of camera 404, such as reflective surfaces, prisms, or specular surfaces that reflect, refract, and / or scatter light differently depending on the positioning of camera 404 relative to the dynamic object. VIO tracker 406 can detect dynamic objects using face detection, face recognition, face tracking, object detection, object recognition, object tracking, or combinations thereof. VIO tracker 406 can detect dynamic objects using one or more artificial intelligence algorithms, one or more trained machine learning models, one or more trained neural networks, or combinations thereof. VIO tracker 406 can mask one or more dynamic objects in an image by overlaying a mask over a region of the image that includes a depiction of one or more dynamic objects. The mask can be an opaque color (such as black). The region can be a bounding box with a rectangular or other polygonal shape. The region can be determined on a pixel-by-pixel basis.
[0072] As previously noted, passive stereo vision systems can use two cameras spaced a predetermined distance apart to capture stereo pairs of images of a scene. For example, in a passive stereo vision system, two cameras can be positioned from different perspectives of the same scene, with each camera capturing images of the scene substantially simultaneously. The system can determine depth information (e.g., a depth map of the scene) of the scene based on the images captured by the two cameras (which may be referred to as stereo pairs). Depth information may include the depth of objects in the scene (e.g., the distance between the camera (or a point relative to the camera) and the object).
[0073] For example, if the scene captured in a pair of stereo images includes objects, a pixel representing a point on an object in an image from one camera may have a corresponding pixel representing the same point on the same object in an image from a second camera. However, because the images were taken by cameras from different perspectives of the same scene, the localization of the pixel corresponding to a point on an object in the first image may differ from the localization of the pixel corresponding to the same point on an object in the second image. By matching corresponding pixels in the two images and calculating the distance between these corresponding pixels, the relative depth of points on the object within the scene can be determined. For example, in some cases, the closer the object is to the camera, the greater the distance between corresponding pixels within the images.
[0074] Figure 5 Two images of a single scene 502 captured from different camera positions according to various aspects of this disclosure are illustrated, namely image 506 and image 508 (in...). Figure 5 This is also represented as image I. L and Image I R Different camera positions are marked as the left "origin" O. L And the right "origin" O R They were offset by a distance T x Due to offset T x The same point P of object 504 appears in both images 506 (I L ) and 508 (I R Different pixel positions p within ) L and p R Location. As can be seen, with image 508 (I R Point P in ) R Corresponding image 508 (I) R x-axis coordinates in ) R Along polar line 510 from coordinate x L The disparity d is offset, where the coordinate x L Corresponding to image 506 (I) L The location of point P in the image is determined. This parallax (also known as phase difference) of the pixel location can be used to determine the approximate distance from the camera to point P on object 504 in scene 502. By understanding the geometry of the stereo camera and applying this analysis to each point in the image, a depth map of the scene can be generated.
[0075] To determine the disparity d, the system can, for example, include the pixel location p. L The pixel window of the pixel at and around the image 508 (I) R Multiple pixel windows in the image are compared to determine the image 508 (I R The pixel position p in ) R Corresponding to image 506 (I) LThe pixel position p in ) L .about Figure 6 An example of this window-based comparison technique is described. For instance, a passive stereo vision system can determine image 508 (I R Polar line 510 in (). Polar line 510 can be drawn from the origin O. L The ray projected onto point P is defined as shown in image 506 (I). R This was observed in [the study]. Passive stereo vision systems can include [the data] located at pixel position p. L The pixel window of the pixel at and around the point is compared with a window of similar size along the epipolar line 510.
[0076] Figure 6 Two images according to various aspects of this disclosure are illustrated, including image 602 (which may be a "right image" or a "reference image") and image 604 (which may be a "left image"), along with an associated cost function 614. To compare windows between image 602 and image 604, a pixel window 606 from image 602 can be selected. The pixel window 606 from image 602 can be compared with one or more pixel windows from image 604. In some cases, window 606 can be compared with windows of similar size (e.g., all windows of similar size) along the epipolar line 612 of image 604.
[0077] Figure 6 The cost function 614 shown represents the similarity between window 606 and windows of similar size along the epipolar line 612 of image 604, as a function of disparity. The similarity between windows can be based on the similarity between the corresponding red, green, blue, and / or intensity (or brightness) values of the pixels included in the respective windows. The lower the value of the cost function 614 for a particular disparity, the higher the similarity between window 606 and a window in image 602 at the corresponding disparity. For example, the cost function 614 includes two minimum values, c1 and c2. The minimum value c1 corresponds to disparity d1, which corresponds to the comparison between window 606 and candidate window 608 of image 604. The minimum value c2 corresponds to disparity d2, which corresponds to the comparison between window 606 and candidate window 610 of image 604.
[0078] A disparity map can be a two-dimensional map of disparity. A two-dimensional map can be compared with an image (e.g., Figure 5The image 506 is related to this. For example, a two-dimensional disparity map may include the same (or in some cases substantially the same) resolution as the corresponding image, where each pixel of the image has a corresponding disparity value. In an exemplary example, a disparity map can be generated by determining the corresponding disparity for each pixel in a plurality of pixels (e.g., all or most pixels) of the image (e.g., by scanning an epipolar window across a stereo pair of images and determining the disparity for each pixel in the plurality of pixels). Each value of the disparity map may represent the disparity (e.g., Figure 5 The disparity d). Depth maps can be scene-based (e.g., Figure 5 The 3D geometry of scene 502 is derived from the disparity map, which includes the distances between cameras that captured the image (e.g., Figure 5 Distance T X ).
[0079] A depth map can be a representation of three-dimensional information (e.g., depth information). For example, a depth map can be a two-dimensional map representing depth values (e.g., pixel values). The values of a depth map can correspond to a corresponding image (e.g., Figure 5 The pixels in image 506. For example, the depth map may have the same or substantially the same resolution as the corresponding image, where each depth value of the depth map represents the origin (e.g., Figure 5 Origin L ) and points (e.g., Figure 5 The depth or distance between points P). In some cases, each pixel in the depth map may have a depth value. Because the depth map is based on the disparity map, in some cases, each pixel in the disparity map may have a disparity value.
[0080] Including distances known to the other (e.g., Figure 5 T X A system or device with two cameras can realize passive stereo vision technology and can be a passive stereo vision system. For example, Figure 1A and Figure 1B The HMD 100 includes a first camera 102 and a user 106, and can be an example of a passive stereo vision system. Other examples of passive stereo vision systems include robots, vehicles, and cameras.
[0081] Figure 7This is an illustration of an example environment 700 in which pose and / or distance determination can be implemented according to various aspects of the present disclosure. For example, a projector 702 can project a pattern 704 onto a scene 708 (including onto a surface 706 of the scene 708). A 6DoF system 710 can operate in the environment 700 and can capture successive images of the scene 708. The 6DoF system 710 can determine its pose based on successive images of the scene 708. It is possible for the 6DoF system 710 to determine its pose based on the pattern 704 being projected onto the scene 708. For example, the 6DoF system 710 may be able to determine its pose more accurately and / or more quickly based on successive images of the scene 708 including the pattern 704, compared to a case where the successive images do not include the pattern 704. Additionally or alternatively, a passive stereo vision system 712 can operate in the environment 700 and can capture stereo paired images of the scene 708. The passive stereo vision system 712 can determine the distance between itself and points in scene 708 based on stereo pair images. This allows the passive stereo vision system 712 to determine this distance based on pattern 704 projected onto scene 708. For example, compared to stereo pair images excluding pattern 704, the passive stereo vision system 712 may be able to determine the distance between itself and individual points in scene 708 more accurately and / or more quickly based on stereo pair images of scene 708 including pattern 704.
[0082] The 6DoF system 710 can be any suitable 6DoF system capable of determining the pose of the 6DoF system 710 based on successive images of scene 708. The 6DoF system 710 can be, for example, a head-mounted display (e.g., Figure 1A and Figure 1B HMD 100), XR systems (e.g., Figure 2 The XR system 200 or robot may be included therein.
[0083] The passive stereo vision system 712 can be any suitable passive stereo vision system capable of determining the distance between the passive stereo vision system 712 and corresponding points in the scene 708 based on stereo pair images of the scene 708. The passive stereo vision system 712 can be, for example, a head-mounted display (e.g., Figure 1A and Figure 1B HMD 100), XR systems (e.g., Figure 2 The XR system 200), robot, vehicle or camera, or may be included therein.
[0084] Projector 702 may be a projector or emitter capable of projecting or emitting electromagnetic radiation onto scene 708 to generate pattern 704. The electromagnetic radiation may have any suitable wavelength, including, for example, visible light (of any color), near-infrared light, infrared light, or any combination thereof. In some aspects, projector 702 may pattern the electromagnetic radiation to generate pattern 704, for example, by projecting light through one or more digital micromirror devices (DMDs) or liquid crystal displays (LCDs). Additionally or alternatively, projector 702 may use one or more lasers and / or mirrors to generate pattern 704. Projector 702 may be or may include a digital light processing (DLP) projector, an LCD projector, a light-emitting diode (LED) projector, a liquid crystal on silicon (LCOS) projector, and / or a laser projector. In some cases, projector 702 may include multiple projectors, for example, a group of projectors. Projectors may be located in one position and point in the same or different directions (e.g., to cover a wider field of view). Additionally or alternatively, projectors may be located at different locations throughout the environment.
[0085] Projector 702 may include a pattern generator capable of generating pattern 704 (e.g., Figure 9 The pattern generator 906) and the projection module (e.g., the pattern generator 704) that can project the pattern 704 onto the scene 708. Figure 9 The projection module 904). Pattern 704 may include uniquely arranged dots or shapes. In this disclosure, the term "unique" may refer to a pattern (or part of a pattern) that is visually different from other patterns in the scene. For example, a "unique pattern" may include unique and / or distinguishable elements or shapes (such as dots, lines, etc.) that, individually or together with scene content, form blocks that can be detected and / or matched when compared between images captured by two cameras. Thus, pattern 704 may be unique relative to scene 708. Additionally or alternatively, pattern 704 may include unique portions at various points of scene 708. For example, as Figure 7 As illustrated, pattern 704 may include different unique portions at different points on surface 706. In some aspects, pattern 704 may consist of points arranged in one or more patterns. Points may have any shape (e.g., circles, squares, triangles, or stars). Points may all have the same shape, or the points of pattern 704 may have different shapes. Additionally or alternatively, pattern 704 may include lines (e.g., lines extending across surface 706).
[0086] Pattern 704 may have a uniform color (or wavelength or combination of wavelengths). Alternatively, dots or portions of pattern 704 may have a different color (or wavelength or combination of wavelengths) than other dots or portions. For example, dots on a first side of surface 706 may have a first color, and dots on a second side of surface 706 may have a different second color. Additionally or alternatively, dots on the top side of each group of dots may have a third color, and dots on the bottom side of each group of dots may have a fourth color.
[0087] Pattern 704 can make surface 706 appear visually distinct. For example, without pattern 704, surface 706 can be substantially visually uniform. For instance, if 6DoF system 710 were to capture successive images of surface 706 without pattern 704, 6DoF system 710 might not be able to correctly correlate features of the successive images, and 6DoF system 710 might not be able to accurately perform visual simultaneous localization and mapping (VSLAM or SLAM) techniques. However, if 6DoF system 710 were to capture successive images of surface 706 on which pattern 704 is projected, 6DoF system 710 might be able to correlate points of the successive images to perform VSLAM techniques. Similarly, if passive stereo vision system 712 were to capture stereo pairs of surface 706 without pattern 704, passive stereo vision system 712 might not be able to correctly correlate features of the stereo pairs, and passive stereo vision system 712 might not be able to accurately determine distances to points on surface 706. However, if the passive stereo vision system 712 is to capture stereo pairs of images of a surface 706 on which the pattern 704 is projected, the passive stereo vision system 712 may be able to correlate the features of the stereo pairs of images to determine the depth of the points.
[0088] Projector 702 can generate pattern 704 and project it onto surface 706 in a manner where pattern 704 is fixed relative to surface 706. Pattern 704 (projected onto surface 706) can remain constant. This consistency of pattern 704 relative to surface 706 allows 6DoF system 710 to determine its pose based on a series of images of pattern 704 on surface 706. Additionally or alternatively, the consistency of pattern 704 relative to surface 706 allows passive stereo vision system 712 to determine the distance between passive stereo vision system 712 and a point on surface 706.
[0089] In some cases, the user can place the projector 702 relative to the surface 706 within the scene 708. For example, the user can position the projector 702 to project the pattern 704 onto the surface 706. Furthermore, the user can adjust the pattern 704 based on the scene 708 and / or the surface 706. For example, the user can adjust the intensity of the light from the pattern 704, the wavelength of the light from 704, the sparsity of the dots in the pattern 704, the size of the dots in the pattern 704, and / or the shape of the dots in the pattern 704.
[0090] For example, Figure 8 Four scenarios are illustrated whereby, according to various aspects of this disclosure, the projector 702 can project an adjusted pattern 704 onto a scene 708. For example, in a first scenario 802, the projector 702 can be positioned on the floor 810 of the scene 708 and can project a sparse pattern 812 onto a surface 706. In a second scenario 804, the projector 702 can be positioned on the floor 810 of the scene 708 and can project a dense pattern 814 onto the surface 706. In a third scenario 806, the projector 702 can be positioned on a wall (e.g., surface 706) of the scene 708 and can project a dense pattern 814 onto the surface 706 (albeit from an angle different from the angle at which the projector 702 projects the dense pattern 814 onto the surface 706 in the second scenario 804). In the fourth scenario 808, the projector 702 can be positioned on the ceiling 816 of the scene 708 and can project a very dense pattern 818 onto the surface 706 and / or other surfaces of the scene 708.
[0091] return Figure 7 In other cases, projector 702 may determine to project pattern 704 onto scene 708 and / or surface 706. For example, projector 702 may include a camera (e.g., Figure 9 (camera 908) and image analyzer (e.g., Figure 9 (Image analyzer 910). Projector 702 can use camera 908 to capture one or more images of scene 708 and determine that surface 706 includes visually indistinguishable portions. Projector 702 can determine to project pattern 704 onto surface 706 such that visually indistinguishable portions become visually distinct. For example, projector 702 can capture an image of surface 706 (which may be a blank wall). Projector 702 can identify surface 706 as visually indistinguishable within the scene. Projector 702 can determine pattern 704 and can determine to project pattern 704 onto surface 706.
[0092] Furthermore, the projector 702 can capture an image of the surface 706 on which the pattern 704 is projected, and determine whether the pattern 704 makes visually indistinguishable portions of the surface 706 visually distinct. For example, the projector 702 can determine whether the surface 706 on which the pattern 704 is projected is visually distinct enough for the 6DoF system 710 to determine its pose and / or for the passive stereo vision system 712 to determine its distance from the surface 706. In some aspects, the projector 702 can compare portions (e.g., windows or features) of the image of the surface 706 (on which the pattern 704 is projected) with other portions of the image to determine whether the surface 706 (on which the pattern 704 is projected) is visually distinct enough. If the surface 706 (on which the pattern 704 is projected) is not visually distinct enough, the projector 702 can adjust the pattern 704 and project the adjusted pattern onto the surface 706. The projector 702 can adjust the intensity of the light in the pattern 704, the wavelength of the light in the pattern 704, the sparsity of the dots in the pattern 704, the size of the dots in the pattern 704, and / or the shape of the dots in the pattern 704.
[0093] Additionally or alternatively, the projector 702 can be connected from another system or device (e.g., using a communication module, such as...) Figure 9 The communication module 912 receives instructions on the surface 706 (e.g., an indication that the surface 706 is visually indistinguishable or includes visually indistinguishable portions). For example, the 6DoF system 710 may determine that it has difficulty determining its pose based on the surface 706, and may send an indication of this difficulty to the projector 702. Additionally or alternatively, the passive stereo vision system 712 may determine that it has difficulty determining the distance between itself and the surface 706, and may send an indication of this difficulty to the projector 702. The projector 702 may project a pattern 704 onto the surface 706, and / or adjust the pattern 704 in response to such indications.
[0094] In some aspects, projector 702 can generate pattern 704 to encode information such as location information (e.g., coordinates, such as latitude and longitude or local coordinates), time information (e.g., time of day), or messages (e.g., labels, instructions, and / or warnings). This information can be decoded by an image processing system. For example, 6DoF system 710 and / or passive stereo vision system 712 can decode this information. 6DoF system 710 and / or passive stereo vision system 712 can use this information. For example, a robot can capture an image of surface 706, identify pattern 704, and decode the information encoded by pattern 704. This information may include, for example, instructions on how to navigate within environment 700. The robot can navigate according to these instructions.
[0095] Additionally or alternatively, the projector 702 may modify this information. For example, if the information is time information, the projector 702 may update the time information over time. As another example, if the information is a message, the projector 702 may modify the message in response to different messages provided by the user.
[0096] Figure 9 This is a block diagram illustrating an example architecture of an example projector 902 according to various aspects of this disclosure. The projector 902 may be... Figure 7 and / or Figure 8 Example of projector 702.
[0097] Projector 902 includes projection module 904 that can project patterns (e.g., unique patterns) onto the environment (e.g., onto a surface of the environment). Projector 902 may include one or more light sources (e.g., lamps, bulbs, or lasers) and / or one or more patterning modules (e.g., mirrors or LCDs).
[0098] In some aspects, the projector 902 may include a pattern generator 906 capable of generating patterns. The pattern generator 906 may be implemented by one or more processors. In other aspects, the projector 902 may receive patterns from another source (e.g., via a communication module 912).
[0099] In some aspects, projector 902 may include camera 908, which can capture one or more images of the environment of projector 902. In some cases, projector 902 may be configured to use camera 908 to scan the environment. In other aspects, projector 902 may not include camera 908.
[0100] In some aspects, projector 902 may include image analyzer 910, which can analyze images captured by camera 908. Image analyzer 910 may be implemented by one or more processors (e.g., one or more processors implementing pattern generator 906). Image analyzer 910 can analyze the image to determine whether the environment includes visually indistinguishable parts. In other aspects, projector 902 may not include image analyzer 910.
[0101] In some aspects, projector 902 may include communication module 912, which can receive indications of one or more visually indistinguishable parts of the environment. For example, a 6DoF system or a passive stereoscopic vision system may send indications of one or more visually indistinguishable parts of the environment to projector 902 via communication module 912. In other aspects, projector 902 may not include communication module 912.
[0102] If the image analyzer 910 determines that there are visually indistinguishable parts of the environment and / or if the projector 902 receives an indication of visually indistinguishable parts of the environment via the communication module 912, the projector 902 may determine to project a pattern onto the visually indistinguishable parts. Additionally or alternatively, the projector 902 may determine to adjust the projected pattern based on the determined visually indistinguishable parts and / or the visually indistinguishable parts indicated by the received indication.
[0103] Figure 10 This is an illustration of an example environment 1000 in which pose and / or distance determination can be implemented according to various aspects of the present disclosure. For example, projector 1002 can project pattern 1004 onto scene 1008 (including onto surface 1006 of scene 1008). 6DoF system 710 can capture successive images of scene 1008 (including pattern 1004 projected onto surface 1006) and can determine the pose of 6DoF system 710 based on the captured successive images. Additionally or alternatively, passive stereo vision system 712 can capture stereo pairs of scene 1008 (including pattern 1004 projected onto surface 1006) and can determine the distance between passive stereo vision system 712 and a point on surface 1006 based on the captured stereo pairs.
[0104] Figure 10 The projector 1002 can be used with Figure 7The projector 1002 is identical to, substantially similar to, and / or can perform the same or substantially the same operations as, the projector 702. However, although the projector 702 is described as fixed relative to the scene 708, the projector 1002 can move or be moved relative to the scene 1008. For example, the projector 1002 can be positioned on a moving object (e.g., a robot or a drone). The projector 1002 can move within the environment 1000. Despite the movement of 1002, the projector 1002 can project a pattern 1004 onto the surface 1006. In some aspects, the projector 1002 can project a pattern 1004 such that, despite the movement of the projector 1002 within the environment 1000, the pattern 1004 is constant with respect to the surface 1006.
[0105] Figure 11 This is a flowchart illustrating a process 1100 for implementing pose and / or distance determination according to various aspects of this disclosure. One or more operations of process 1100 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capability to perform process 1100. One or more operations of process 1100 may be implemented as software components that execute and run on one or more processors.
[0106] At box 1102, a computing device (or one or more components thereof) can cause a projector to project a pattern onto a scene for feature association by an imaging device that captures an image of the pattern projected onto the scene. The projector can be separate from the imaging device. For example, projector 702 can project pattern 704 onto scene 708 so that an imaging device (e.g., a 6DoF system 710 or a passive stereo vision system 712) performs feature association based on the image of pattern 704 projected onto scene 708.
[0107] In some respects, the imaging device can be configured to determine the distance between the imaging device and a point in the scene based on an image of a pattern projected onto the scene. For example, the passive stereo vision system 712 can be configured to determine the distance between the passive stereo vision system 712 and a point in the scene 708 based on an image of pattern 704 captured by the passive stereo vision system 712 when pattern 704 appears in scene 708.
[0108] In some aspects, the imaging device can be configured to determine the distance between the imaging device and a point in the scene, regardless of whether the projector is projecting a pattern. For example, the imaging device can be configured to determine the distance between the imaging device and an object or surface in the scene even if the projector is not projecting a pattern onto the scene. As another example, the imaging device can be configured to determine this distance even if the imaging device is capturing an image of another part of the scene (e.g., a part excluding the projected pattern). For example, a passive stereo vision system 712 can be configured to determine the distance between the passive stereo vision system 712 and a point in the scene 708, regardless of whether the projector 702 is projecting a pattern 704 onto the scene 708.
[0109] In some aspects, the imaging device can be configured to determine the distance between the imaging device and a point in the scene without prior information about the pattern. For example, a passive stereo vision system 712 can determine the distance between itself and a point in the scene 708 without prior information about the pattern 704. For example, the passive stereo vision system 712 can match features of the pattern 704 projected into the scene 708 between images captured by itself without prior information about the pattern 704. For example, the passive stereo vision system 712 may not have information about the pattern 704 (e.g., the shape or pattern of the pattern 704), or even information about whether the pattern 704 is projected into the scene 708.
[0110] In some aspects, the imaging device may be or may include a passive stereo vision system configured to correlate features of patterns in stereo pairs of images of a scene to determine the distance between the passive stereo vision system and a point in the scene. For example, the imaging device may be a passive stereo vision system 712, which can capture stereo pairs of images of scene 708 and correlate features between the stereo pairs to determine the distance between the passive stereo vision system 712 and a point in scene 708.
[0111] In some respects, the imaging device can be configured to determine its pose relative to the scene based on an image of a pattern projected onto the scene. For example, the 6DoF system 710 can determine its pose relative to the scene 708 based on an image of a pattern 704 projected onto the scene 708 captured by the 6DoF system 710.
[0112] In some aspects, the imaging device can be configured to determine its pose regardless of whether the projector is projecting a pattern. For example, the imaging device can be configured to determine its pose relative to the scene even if the projector is not projecting a pattern onto the scene. As another example, the imaging device can be configured to determine its pose even if it is capturing an image of another part of the scene (e.g., a part excluding the projected pattern). For example, the 6DoF system 710 can be configured to determine its pose regardless of whether the projector 702 is projecting a pattern 704 onto the scene 708.
[0113] In some aspects, the imaging device can be configured to determine its pose without prior information about the pattern. For example, the 6DoF system 710 can determine its pose without prior information about the pattern 704. For example, the 6DoF system 710 can match features of the pattern 704 projected into the scene 708 between images captured by the 6DoF system 710 without prior information about the pattern 704. For example, the 6DoF system 710 may not have information about the pattern 704 (e.g., the shape or pattern of the pattern 704), or even information about whether the pattern 704 is projected into the scene 708.
[0114] In some aspects, the imaging device may be or may include a six-degrees-of-freedom (6DoF) system configured to correlate features of a pattern in sequential images of a scene to determine the pose of the 6DoF system relative to the scene. For example, the imaging device may be a 6DoF system 710, and the 6DoF system 710 may be configured to capture sequential images of a scene 708 and correlate features of a pattern 704 projected onto the scene 708 to determine the pose of the 6DoF system 710 relative to the scene 708.
[0115] In some respects, a projector can project a pattern onto a scene without receiving communication from an imaging device. For example, projector 702 can project pattern 704 onto scene 708 without receiving any communication (e.g., any instructions, requests, etc.) from either the 6DoF system 6710 or the passive stereo vision system 712.
[0116] In some respects, a projector can project a pattern onto a scene regardless of whether the imaging device captures an image of the pattern projected onto the scene. For example, projector 702 can project pattern 704 onto scene 708, regardless of whether the 6DoF system 710 or the passive stereoscopic vision system 712 captures an image of scene 708. Furthermore, projector 702 may not have any information about the presence of the 6DoF system 710 and / or the passive stereoscopic vision system 712 in the environment of scene 708 or whether the 6DoF system 710 and / or the passive stereoscopic vision system 712 is capturing an image of scene 708.
[0117] In some respects, the projector can project a pattern onto a scene independently of the imaging device. For example, projector 702 can project pattern 704 onto scene 708 independently of 6DoF system 710 and / or passive stereoscopic vision system 712. For example, projector 702 can project pattern 704 onto scene 708 without any communication from 6DoF system 710 and / or passive stereoscopic vision system 712, regardless of the presence of 6DoF system 710 and / or passive stereoscopic vision system 712 in the environment of projector 702, regardless of whether 6DoF system 710 and / or passive stereoscopic vision system 712 is operating in the environment of projector 702, and regardless of whether 6DoF system 710 and / or passive stereoscopic vision system 712 is capturing an image of scene 708.
[0118] In some respects, the projector may be movable relative to the imaging device. For example, the projector 702 may be movable (and / or manipulated) relative to the 6DoF system 710 and / or the passive stereoscopic vision system 712. For example, the projector 702 may be movable (and / or reoriented) separately from the 6DoF system 710 and / or the passive stereoscopic vision system 712. For example, the projector 702 may be movable independently.
[0119] In some aspects, the projector can project a pattern onto the scene at a first time, wherein the imaging device has a first pose relative to the projector at the first time; and project the pattern onto the scene at a second time, wherein the imaging device has a second pose relative to the projector at the second time. For example, projector 702 can project pattern 704 onto scene 708 at a first time. At the first time, 6DoF system 710 and / or passive stereo vision system 712 may have a first pose relative to projector 702. Projector 702 can project pattern 704 onto scene 708 at a second time. At the second time, 6DoF system 710 and / or passive stereo vision system 712 may have a first pose relative to projector 702. For example, the pose of 6DoF system 710 and / or passive stereo vision system 712 relative to projector 702 may change over time.
[0120] In some respects, a projector can project an image so that the image is fixed relative to the scene. For example, projector 702 can project an image 704 so that the image 704 remains fixed relative to the scene 708.
[0121] In some aspects, the projector can be configured to be fixed relative to the scene, and the imaging device can be configured to move relative to the scene. For example, projector 702 can be fixed relative to scene 708. For example, projector 702 may include legs and / or a stand, or may be able to be attached to a wall or ceiling. 6DoF system 710 and / or passive stereo vision system 712 can be configured to be movable, for example, they can be wearable or able to be attached to a mobile system (e.g., a robot).
[0122] In some respects, the projector can be configured to move independently of the imaging device relative to the scene. For example, projector 1002 can be configured to move relative to scene 1008. In such cases, the projector can project a pattern such that the pattern is fixed relative to the scene. For example, projector 1002 can project a pattern 1004 such that although 1002 moves relative to scene 1008, pattern 1004 remains fixed relative to scene 1008.
[0123] In some respects, a computing device (or one or more components thereof) can enable a camera to capture an image of a scene. The computing device (or one or more components thereof) can determine visually indistinguishable portions of the scene based on the image; and enable a projector to project a pattern onto the visually indistinguishable portions of the scene. For example, a camera 908 of projector 902 can capture an image of scene 708. An image analyzer 910 of projector 902 can analyze the image to determine visually indistinguishable portions of scene 708. Projector 902 can project a pattern 704 onto the visually indistinguishable portions of scene 708.
[0124] In some aspects, a computing device (or one or more components thereof) can enable a camera to capture an image of a scene. The computing device (or one or more components thereof) can analyze a pattern represented in the image of the scene and projected onto the scene, and cause a projector to adjust the intensity used to project the pattern; adjust the wavelength used to project the pattern; adjust the sparsity of the dots in the pattern; adjust the size of the dots in the pattern; and / or adjust the shape of the dots in the pattern in response to the analysis. For example, a camera 908 of projector 902 can capture an image of scene 708. An image analyzer 910 of projector 902 can analyze the image. The image analyzer 910 can determine how the pattern 704 projected onto scene 708 appears. Projector 902 can adjust the pattern 704 or how the pattern 704 is projected onto scene 708 based on the analysis, for example, so that the pattern 704 makes scene 708 include more visually distinct parts or makes visually indistinguishable parts of scene 708 appear more visually distinct.
[0125] In some aspects, the computing device (or one or more components thereof) can generate patterns. For example, projector 902 may include pattern generator 906 capable of generating pattern 704. In some aspects, the projector can modify the pattern. For example, pattern generator 906 of projector 902 can modify pattern 704.
[0126] In some aspects, the pattern can encode information. For example, pattern 704 can encode information. In some aspects, the information can be or may include location information; time information; or a message. For example, pattern 704 can encode location information (e.g., relative to scene 708 and / or relative to the world coordinate system). As another example, pattern 704 can encode time information (e.g., time of day and / or date). As another example, pattern 704 can encode a message (e.g., an instruction or warning).
[0127] In some respects, the pattern can encode the first information, and the projector can be further configured to change the pattern to encode the second information. For example, at a first moment, the projector 702 can project a pattern 704 that encodes the first information. The projector 702 can change the pattern 704 to encode the second information. Then, the projector 702 can project the changed pattern 704 that encodes the second information.
[0128] Figure 12This is a flowchart illustrating a process 1200 for implementing pose and / or distance determination according to various aspects of this disclosure. One or more operations of process 1200 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capability to perform process 1200. One or more operations of process 1200 may be implemented as software components that execute and run on one or more processors.
[0129] At box 1202, a computing device (or one or more components thereof) may determine to project a pattern into a scene for feature association by an imaging device that captures an image of the pattern projected into the scene. For example, projector 702 may determine to project pattern 704 into scene 708 such that 6DoF system 710 and / or passive stereo vision system 712 can capture an image of pattern 704 in scene 708 and associate features of pattern 704 in the image.
[0130] At frame 1204, a computing device (or one or more components thereof) enables a projector to project a pattern onto a scene. The projector may be detached from the imaging device. For example, projector 702 may project pattern 704 onto scene 708. Projector 702 may be detached from 6DoF system 710 and / or passive stereoscopic vision system 712.
[0131] In some aspects, a computing device (or one or more components thereof) can enable a camera to capture an image of a scene and determine visually indistinguishable portions of the scene based on the image. Projecting a pattern onto the scene can be or may include projecting the pattern onto visually indistinguishable portions of the scene. For example, camera 908 can capture an image of scene 708. Image analyzer 910 can determine visually indistinguishable portions of scene 708. Projector 902 can project pattern 704 onto visually indistinguishable portions of scene 708.
[0132] In some aspects, a computing device (or one or more components thereof) can enable a camera to capture an image of a scene. The computing device (or one or more components thereof) can analyze a pattern represented in the image of the scene and projected onto the scene, and cause a projector to adjust the intensity used to project the pattern; adjust the wavelength used to project the pattern; adjust the sparsity of the dots in the pattern; adjust the size of the dots in the pattern; and / or adjust the shape of the dots in the pattern in response to the analysis. For example, a camera 908 of projector 902 can capture an image of scene 708. An image analyzer 910 of projector 902 can analyze the image. The image analyzer 910 can determine how the pattern 704 projected onto scene 708 appears. Projector 902 can adjust the pattern 704 or how the pattern 704 is projected onto scene 708 based on the analysis, for example, so that the pattern 704 makes scene 708 include more visually distinct parts or makes visually indistinguishable parts of scene 708 appear more visually distinct.
[0133] In some examples, as previously noted, the methods described herein (e.g., Figure 11 Process 1100 Figure 12 The process 1200 and / or other methods described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more of these methods may be performed by... Figure 7 and / or Figure 8 Projector 702 Figure 9 Projector 902, Figure 10 The projector 1002 or another system or device performs the operation. In another example, these methods (e.g., Figure 11 Process 1100 Figure 12 One or more of the processes 1200 and / or other methods described herein may be performed by Figure 13 The computing device architecture 1300 shown is implemented wholly or partially. For example, it has Figure 13The computing device of the illustrated computing device architecture 1300 may include or be included in components of projector 702, projector 902, and / or projector 1002, and may implement the operation of process 1100, process 1200, and / or other processes described herein. In some cases, the computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.
[0134] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0135] Processes 1100, 1200, and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0136] Additionally, processes 1100, 1200, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or implemented in combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0137] Figure 13 Example computing device architecture 1300 illustrates example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 1300 may include, implement Figure 7 and / or Figure 8 Projector 702 Figure 9 Projector 902 and / or Figure 10 The projector 1002 may be any or all of them, or may be included therein. Additionally or alternatively, the computing device architecture 1300 may be configured to perform process 1100, process 1200 and / or other processes described herein.
[0138] The components of computing device architecture 1300 are shown to communicate electrically with each other using a connection 1312, such as a bus. Example computing device architecture 1300 includes a processing unit (CPU or processor) 1302 and a computing device connection 1312 that couples various computing device components, including computing device memories 1310 (such as read-only memory (ROM) 1308 and random access memory (RAM) 1306), to the processor 1302.
[0139] The computing device architecture 1300 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1302. The computing device architecture 1300 may copy data from memory 1310 and / or storage device 1314 to cache 1304 for fast access by the processor 1302. In this way, the cache can provide performance improvements by avoiding latency for the processor 1302 while waiting for data. These and other modules may control or be configured to control the processor 1302 to perform various actions. Other computing device memory 1310 may also be available. Memory 1310 may include various different types of memory with different performance characteristics. The processor 1302 may include any general-purpose processor and hardware or software services configured to control the processor 1302 (such as services 1 1316, 2 1318, and 3 1320 stored in storage device 1314), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1302 may be a standalone system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.
[0140] To enable user interaction with the computing device architecture 1300, input device 1322 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 1324 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker equipment, etc. In some cases, multi-mode computing devices allow users to provide multiple types of input to communicate with computing device architecture 1300. Communication interface 1326 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the basic features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0141] Storage device 1314 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 1306, read-only memory (ROM) 1308, and hybrid forms thereof. Storage device 1314 may include services 1316, 1318, and 1320 for controlling processor 1302. Other hardware or software modules are envisioned. Storage device 1314 may be connected to computing device connection 1312. In one aspect, a hardware module performing a specific function may include software components for performing that function stored in a computer-readable medium connected to necessary hardware components such as processor 1302, connection 1312, output device 1324, etc.
[0142] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.
[0143] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.
[0144] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.
[0145] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.
[0146] The various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow diagrams, structural diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, process, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0147] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.
[0148] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, magnetic disks or optical disks, USB devices equipped with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0149] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.
[0150] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or intercalation cards. By way of additional examples, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.
[0151] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0152] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the various inventive concepts can be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including such variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.
[0153] Those skilled in the art will understand that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols without departing from the scope of this description.
[0154] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0155] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0156] Claim language or other languages that state "at least one of" and / or "one or more of" in a set indicate that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may refer to A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.
[0157] Claims using phrases such as "at least one processor, configured to," "at least one processor configured to," "one or more processors, configured to," or "one or more processors configured to," or other languages, indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim stating "at least one processor, configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks to perform operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim stating "at least one processor, configured to: X, Y, and Z" may mean that any single processor can perform only a subset of operations X, Y, and Z.
[0158] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.
[0159] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).
[0160] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.
[0161] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.
[0162] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration). Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.
[0163] The exemplary aspects of this disclosure include: Aspect 1. An apparatus comprising: a projector configured to project a pattern onto a scene for feature association by an imaging device, the imaging device capturing an image of the pattern projected onto the scene; wherein the apparatus is separate from the imaging device.
[0164] Aspect 2. The apparatus according to aspect 1, wherein the imaging device is configured to determine the distance between the imaging device and a point in the scene based on the image of the pattern projected onto the scene.
[0165] Aspect 3. The apparatus according to any one of Aspect 1 or 2, wherein the imaging device is configured to determine the distance between the imaging device and the point in the scene, regardless of whether the projector projects the pattern.
[0166] Aspect 4. The apparatus according to any one of Aspects 2 or 3, wherein the imaging device is configured to determine the distance between the imaging device and the point in the scene without having prior information about the pattern.
[0167] Aspect 5. The apparatus according to aspect 1, wherein the imaging device includes a passive stereo vision system configured to correlate features of the pattern in a stereo pair of images of the scene to determine the distance between the passive stereo vision system and a point in the scene.
[0168] Aspect 6. The apparatus according to aspect 1, wherein the imaging device is configured to determine the pose of the imaging device relative to the scene based on the image of the pattern projected onto the scene.
[0169] Aspect 7. The apparatus according to aspect 6, wherein the imaging device is configured to determine the pose of the imaging device regardless of whether the projector projects the pattern.
[0170] Aspect 8. The apparatus according to any one of Aspects 6 or 7, wherein the imaging device is configured to determine the pose of the imaging device without having prior information about the pattern.
[0171] Aspect 9. The apparatus according to Aspect 1, wherein the imaging device includes a six-degree-of-freedom (6DoF) system configured to correlate features of the pattern in sequential images of the scene to determine the pose of the 6DoF system relative to the scene.
[0172] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein the projector is configured to project the pattern onto the scene without receiving communication from the imaging device.
[0173] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the projector is configured to project the pattern onto the scene, regardless of whether the imaging device captures an image of the pattern projected onto the scene.
[0174] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the projector is configured to project the pattern onto the scene independently of the imaging device.
[0175] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein the projector is movable relative to the imaging device.
[0176] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein the projector is configured to: project the pattern onto the scene at a first time, wherein the imaging device has a first pose relative to the projector at the first time; and project the pattern onto the scene at a second time, wherein the imaging device has a second pose relative to the projector at the second time.
[0177] Aspect 15. The apparatus according to any one of Aspects 1 to 14, wherein the projector is configured to project the pattern such that the pattern is fixed relative to the scene.
[0178] Aspect 16. The apparatus according to any one of Aspects 1 to 15, wherein the projector is configured to be fixed relative to the scene, and wherein the imaging device is configured to move relative to the scene.
[0179] Aspect 17. The apparatus according to any one of aspects 1 to 16, wherein the projector is configured to move separately from the imaging device relative to the scene.
[0180] Aspect 18. The apparatus according to any one of Aspects 1 to 17, the apparatus further comprising: a camera configured to capture an image of the scene; and at least one processor configured to: determine visually indistinguishable portions of the scene based on the image; and cause the projector to project the pattern onto the visually indistinguishable portions of the scene.
[0181] Aspect 19. The apparatus according to any one of Aspects 1 to 18, the apparatus further comprising: a camera configured to capture an image of the scene; and at least one processor configured to analyze a pattern represented in the image of the scene and projected onto the scene, wherein, in response to analyzing the pattern, the projector is configured to perform at least one of: adjusting the intensity for projecting the pattern; adjusting the wavelength for projecting the pattern; adjusting the sparsity of the dots in the pattern; adjusting the size of the dots in the pattern; or adjusting the shape of the dots in the pattern.
[0182] Aspect 20. The apparatus according to any one of aspects 1 to 19, the apparatus further comprising a pattern generator configured to generate the pattern.
[0183] Aspect 21. The apparatus according to any one of aspects 1 to 20, wherein the pattern encodes information.
[0184] Aspect 22. The apparatus according to aspect 21, wherein the information includes at least one of: location information; time information; or a message.
[0185] Aspect 23. The apparatus according to any one of aspects 1 to 22, wherein the projector is further configured to change the pattern.
[0186] Aspect 24. The apparatus according to any one of Aspects 1 to 23, wherein the pattern encodes first information, and wherein the projector is further configured to change the pattern to encode second information.
[0187] Aspect 25. A method comprising: determining to project a pattern into a scene for feature association by an imaging device, the imaging device capturing an image of the pattern projected into the scene; and projecting the pattern into the scene from a projector detached from the imaging device.
[0188] Aspect 26. The method according to aspect 25, the method further comprising: capturing an image of the scene; and determining visually indistinguishable portions of the scene based on the image; wherein projecting the pattern into the scene includes projecting the pattern onto the visually indistinguishable portions of the scene.
[0189] Aspect 27. The method according to any one of Aspects 25 or 26, the method further comprising: capturing an image of the scene; analyzing a pattern represented in the image of the scene and projected onto the scene; and, in response to analyzing the pattern, performing at least one of the following: adjusting the intensity for projecting the pattern; adjusting the wavelength for projecting the pattern; adjusting the sparsity of the points of the pattern; adjusting the size of the points of the pattern; or adjusting the shape of the points of the pattern.
[0190] Aspect 28. The method according to any one of aspects 25 to 27, the method further comprising generating the pattern.
[0191] Aspect 29. The method according to any one of Aspects 25 to 28, wherein the pattern encodes information.
[0192] Aspect 30. The method according to aspect 29, wherein the information includes at least one of: location information; time information; or a message.
[0193] Aspect 31. The method according to any one of aspects 25 to 30, wherein the method further comprises changing the pattern.
[0194] Aspect 32. The method according to any one of Aspects 25 to 31, wherein the pattern encodes first information, and the method further comprises changing the pattern to encode second information.
[0195] Aspect 33. An imaging apparatus comprising: two cameras configured to capture stereo paired images of a scene; and at least one processor configured to correlate features of the stereo paired images to determine distances between the two cameras and points in the scene corresponding to the features; wherein the points in the scene are illuminated by a pattern, and wherein the pattern is projected by a projector separate from the imaging apparatus.
[0196] Aspect 34. An imaging device comprising: a camera configured to capture sequential images of a scene; and at least one processor configured to track features across the sequential images to determine the pose of the imaging device relative to the scene; wherein the features correspond to points in the scene illuminated by a pattern, and wherein the pattern is projected by a projector separate from the imaging device.
[0197] Aspect 35. The apparatus according to aspect 1, wherein the imaging device includes a three-degree-of-freedom (3DoF) system configured to correlate features of the pattern in sequential images of the scene to determine the location of the 3DoF system relative to the scene.
Claims
1. An apparatus, the apparatus comprising: A projector configured to project a pattern onto a scene for feature association by an imaging device, the imaging device capturing an image of the pattern projected onto the scene; The device thereon is separate from the imaging equipment.
2. The apparatus of claim 1, wherein the imaging device is configured to determine the distance between the imaging device and a point in the scene based on the image of the pattern projected onto the scene.
3. The apparatus of claim 2, wherein the imaging device is configured to determine the distance between the imaging device and the point in the scene, regardless of whether the projector projects the pattern.
4. The apparatus of claim 2, wherein the imaging device is configured to determine the distance between the imaging device and the point in the scene without having prior information about the pattern.
5. The apparatus of claim 1, wherein the imaging device includes a passive stereo vision system configured to correlate features of the pattern in a stereo pair image of the scene to determine the distance between the passive stereo vision system and a point in the scene.
6. The apparatus of claim 1, wherein the imaging device is configured to determine the pose of the imaging device relative to the scene based on the image of the pattern projected onto the scene.
7. The apparatus of claim 6, wherein the imaging device is configured to determine the pose of the imaging device regardless of whether the projector projects the pattern.
8. The apparatus of claim 6, wherein the imaging device is configured to determine the pose of the imaging device without having prior information about the pattern.
9. The apparatus of claim 1, wherein the imaging device includes a six-degree-of-freedom (6DoF) system configured to correlate features of the pattern in sequential images of the scene to determine the pose of the 6DoF system relative to the scene.
10. The apparatus of claim 1, wherein the projector is configured to project the pattern onto the scene without receiving communication from the imaging device.
11. The apparatus of claim 1, wherein the projector is configured to project the pattern onto the scene, regardless of whether the imaging device captures an image of the pattern projected onto the scene.
12. The apparatus of claim 1, wherein the projector is configured to project the pattern onto the scene independently of the imaging device.
13. The apparatus of claim 1, wherein the projector is movable relative to the imaging device.
14. The apparatus of claim 1, wherein the projector is configured to: The pattern is projected onto the scene at a first instant, wherein the imaging device has a first pose relative to the projector at the first instant; and The pattern is projected onto the scene at a second time, wherein the imaging device has a second pose relative to the projector at the second time.
15. The apparatus of claim 1, wherein the projector is configured to project the pattern such that the pattern is fixed relative to the scene.
16. The apparatus of claim 1, wherein the projector is configured to be fixed relative to the scene, and wherein the imaging device is configured to move relative to the scene.
17. The apparatus of claim 1, wherein the projector is configured to move separately from the imaging device relative to the scene.
18. The apparatus of claim 1, further comprising: A camera, configured to capture images of the scene; and At least one processor, said at least one processor being configured to: Based on the image, determine the visually indistinguishable parts of the scene; and The projector projects the pattern onto the visually indistinguishable portion of the scene.
19. The apparatus of claim 1, further comprising: A camera, configured to capture images of the scene; and At least one processor is configured to analyze the pattern represented in the image of the scene and projected into the scene. In response to analyzing the pattern, the projector is configured to perform at least one of the following: Adjust the intensity used to project the pattern; Adjust the wavelength used to project the pattern; Adjust the sparsity of the dots in the pattern; Adjust the size of the points in the pattern; or Adjust the shape of the points in the pattern.
20. The apparatus of claim 1, further comprising a pattern generator configured to generate the pattern.
21. The apparatus of claim 1, wherein the pattern encodes information.
22. The apparatus of claim 21, wherein the information includes at least one of the following: Location information; Time information; or information.
23. The apparatus of claim 1, wherein the projector is further configured to change the pattern.
24. The apparatus of claim 1, wherein the pattern encodes first information, and wherein the projector is further configured to change the pattern to encode second information.
25. A method comprising: The pattern is determined to be projected into a scene for feature association by an imaging device, which captures an image of the pattern projected into the scene; as well as The pattern is projected onto the scene from a projector that is separate from the imaging device.
26. The method of claim 25, further comprising: Capture images of the scene; as well as Based on the image, determine the visually indistinguishable parts of the scene; Projecting the pattern into the scene includes projecting the pattern onto the visually indistinguishable portion of the scene.
27. The method of claim 25, further comprising: Capture images of the scene; Analyze the pattern represented in the image of the scene that is projected into the scene; as well as In response to analyzing the pattern, at least one of the following is performed: Adjust the intensity used to project the pattern; Adjust the wavelength used to project the pattern; Adjust the sparsity of the dots in the pattern; Adjust the size of the points in the pattern; or Adjust the shape of the points in the pattern.
28. The method of claim 25, further comprising generating the pattern.
29. The method of claim 25, wherein the pattern encodes information.
30. The method of claim 29, wherein the information comprises at least one of the following: Location information; Time information; or information.