Close IMU-camera coupling for dynamic bend estimation
By rigidly mounting IMU sensors in flexible AR/VR devices, changes in spatial relationships are dynamically measured, solving the display error problem caused by frame bending and achieving accurate virtual content display and saving computing resources.
Patent Information
- Application Number
- CN202480018363.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-15
- Filing Date
- 2024-03-14
- Publication Date
- 2025-11-07
AI Technical Summary
The bending of the frame of flexible AR/VR devices causes errors in the spatial relationship between the camera device and the display, affecting the accurate display of virtual content. Existing methods are time-consuming or result in bulky devices.
By rigidly mounting the IMU sensor to the camera device, changes in spatial relationships are dynamically measured, frame bending is predicted and estimated, and display errors are corrected.
It achieves accurate display of virtual content while maintaining device flexibility and ergonomics, reducing computing resources and workload.
Smart Images

Figure CN120917403A_ABST
Abstract
Description
[0001] CLAIM OF PRIORITY
[0002] This application claims the benefit of priority to U.S. Patent Application Serial No. 18 / 184,333, filed March 15, 2023, which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0003] The subject matter disclosed herein relates generally to visual tracking systems. Specifically, the present disclosure presents systems and methods for mitigating bending effects in visual-inertial tracking systems. BACKGROUND
[0004] Augmented reality (AR) devices enable a user to observe a scene while also seeing relevant virtual content that is aligned with items, images, objects, or the environment in the field of view of the device. Virtual reality (VR) devices provide a more immersive experience than AR devices. VR devices occlude the user’s view with virtual content that is displayed based on the positioning and orientation of the VR device.
[0005] Both AR and VR devices rely on motion tracking systems that track the pose (e.g., orientation, positioning, location) of the device. Motion tracking systems are typically factory calibrated (based on predefined relative positioning between cameras and other sensors) to accurately display virtual content at the desired location relative to its environment. However, due to mechanical stresses in the AR / VR device (e.g., bending of the eyewear frame), the factory calibration parameters can drift over time when the user wears the AR / VR device. BRIEF DESCRIPTION OF DRAWINGS
[0006] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number are often duplicated in several places throughout the detailed description.
[0007] Figure 1 is a block diagram illustrating an environment for operating a display device, according to one example embodiment.
[0008] Figure 2 is a block diagram illustrating a display device, according to one example embodiment.
[0009] Figure 3 is a block diagram illustrating a visual tracking system, according to one example embodiment.
[0010] Figure 4 is a block diagram illustrating a dynamic bending estimation module, according to one example embodiment.
[0011] Figure 5 is a block diagram illustrating a corrected frame, according to one example embodiment.
[0012] Figure 6 Rigid component coupling and non-rigid component coupling are shown in accordance with one example implementation.
[0013] Figure 7 An example trajectory of a display device is shown in accordance with one example implementation.
[0014] Figure 8 is a flowchart showing a method for adjusting a frame in accordance with one example implementation.
[0015] Figure 9 is a flowchart showing a method for processing a frame in accordance with one example implementation.
[0016] Figure 10 is a flowchart showing a method for detecting spatial changes based on proximity sensors in accordance with one example implementation.
[0017] Figure 11 A network environment in which a head wearable device can be implemented is shown in accordance with one example implementation.
[0018] Figure 12 is a block diagram showing a software architecture in which the present disclosure can be implemented in accordance with example implementations.
[0019] Figure 13 is a diagrammatic representation of a machine in the form of a computer system within which a set of instructions can be executed to cause the machine to perform any one or more of the methodologies discussed herein, in accordance with one example implementation. DETAILED DESCRIPTION
[0020] The following description describes systems, methods, techniques, instruction sequences, and computing machine program products that show example implementations of the present subject matter. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of various implementations of the present subject matter. It will be apparent, however, to one skilled in the art that implementations of the present subject matter can be practiced without some or all of these specific details. Examples are merely representative. Structure (e.g., structural components such as modules) is optional and can be combined or subdivided, and operations (e.g., in procedures, algorithms, or other functions) can vary in sequence or be combined or subdivided, unless otherwise stated.
[0021] The term “augmented reality” (AR) is used herein to refer to an interactive experience of a real-world environment where physical objects residing in the real world are “augmented” or enhanced by computer-generated digital content (also referred to as virtual content or synthetic content). AR can also refer to a system that enables a combination of real and virtual worlds, real-time interaction, and 3D registration of virtual objects and real objects. A user of an AR system perceives virtual content that appears to be connected with or interacting with physical objects of the real world.
[0022] The term “virtual reality” (VR) is used herein to refer to a simulated experience of a virtual world environment that is completely different from the real-world environment. Computer-generated digital content is displayed in the virtual world environment. VR also refers to a system that enables a user of the VR system to be fully immersed in the virtual world environment and to interact with virtual objects presented in the virtual world environment.
[0023] The term “AR application” is used herein to refer to a computer-operated application that implements an AR experience. The term “VR application” is used herein to refer to a computer-operated application that implements a VR experience. The term “AR / VR application” refers to a computer-operated application that implements a combination of an AR experience or a VR experience.
[0024] The term “visual tracking system” is used herein to refer to a computer-operated application or system that enables the system to track visual features identified in images captured by one or more cameras of the visual tracking system. The visual tracking system constructs a model of a real-world environment based on the tracked visual features. Non-limiting examples of visual tracking systems include: visual simultaneous localization and mapping systems (VSLAM) and visual-inertial odometry (VIO) systems. VSLAM can be used to construct a map of a target from an environment or scene based on one or more cameras of the visual tracking system. VIO (also referred to as a visual-inertial tracking system) determines a latest pose (e.g., position and orientation) of a device based on data acquired from multiple sensors of the device (e.g., optical sensors, inertial sensors).
[0025] The term “inertial measurement unit” (IMU) is used herein to refer to a device that can report inertial states of a moving body, including acceleration, velocity, orientation, and position of the moving body. An IMU enables tracking of movement of a body by integrating acceleration and angular velocity measured by the IMU. An IMU can also refer to a combination of an accelerometer and a gyroscope that can determine and quantify linear acceleration and angular velocity, respectively. Values obtained from the IMU gyroscope can be processed to obtain pitch, roll, and heading of the IMU, and thereby of a body associated with the IMU. Signals from the accelerometer of the IMU can also be processed to obtain velocity and displacement of the IMU.
[0026] The term "flexible device" is used herein to refer to a device that can be bent without breaking. Non-limiting examples of flexible devices include: a head-mounted device such as glasses, a flexible display device such as AR / VR glasses, or any other wearable device that can be bent without breaking to fit a user's body part.
[0027] Both AR and VR applications allow a user to access information, for example, in the form of virtual content that is presented in a display of an AR / VR display device (also referred to as a display device, a flexible device, a flexible display device). The rendering of the virtual content can be based on the positioning of the display device relative to a physical object or relative to a frame of reference (external to the display device) so that the virtual content appears correctly in the display. For AR, the virtual content appears to align with a physical object as perceived by the user and the camera of the AR display device. The virtual content appears to be attached to the physical world (e.g., a physical object of interest). To do this, the AR display device detects the physical object and tracks the pose of the positioning of the AR display device relative to the physical object. The pose identifies the positioning and orientation of the display device relative to a frame of reference or relative to another object. For VR, the virtual object appears at a location that is based on the pose of the VR display device. Thus, the virtual content is refreshed based on the latest pose of the device. A vision tracking system at the display device determines the pose of the display device.
[0028] A flexible device that includes a vision tracking system can operate with one or more cameras mounted on the flexible device. For example, one camera is mounted to the left temple of the frame of the flexible device and another camera is mounted to the right temple of the frame of the flexible device. The flexible device can be bent to accommodate different user head sizes. The flexible device can be designed for a less rigid, more ergonomic and visually appealing frame design. However, such a flexible design also results in a change in spatial relationships between different components (e.g., such as the display, the cameras, the IMU, and the projector) over time. Such spatial relationships can change during normal operation by simply putting them on, walking around, or touching the frame. The frame bending and the change in spatial relationships cause undesirable shifts in the stereo images (away from the out-of-the-box calibrated configuration), which results in an unrealistic AR experience (errors in sensing the environment from the stereo images with the shifts). Therefore, accurate knowledge of the spatial relationships of the key system components of the display device is important for an accurate AR experience.
[0029] Existing methods to mitigate frame bending include:
[0030] • Highly accurate out-of-the-box calibration of each device: This can be time consuming.
[0031] • Increase rigidity of the eyewear, leading to specialized frame structures and clunky, uncomfortable designs.
[0032] • Allow device models to estimate spatial correlations using best-fit methods: this leads to additional computational requirements.
[0033] The present application describes a flexible device in which each IMU sensor is rigidly mounted to a corresponding camera in a multi-camera device to measure spatial correlations in the display device during runtime. The IMU sensors are used to dynamically measure changes in spatial relationships between each sensing group (e.g., IMU1 + Camera 1 and IMU2 + Camera 2). The IMU data can be used to predict and estimate frame bending prior to processing the camera frames. Because each camera is rigidly connected to its own IMU, the flexible device does not need to estimate spatial changes between each camera and the corresponding IMU (thereby eliminating uncertainty from internal IMU state estimation). The IMU sensors and cameras that have been rigidly connected allow for more flexible and ergonomic frame designs while maintaining close-to-optimal AR experiences.
[0034] In one example implementation, a method includes forming a plurality of sensor groups of an augmented reality (AR) display device, wherein one sensor group of the plurality of sensor groups includes a camera that is tightly coupled to a corresponding IMU (inertial measurement unit) sensor, a spatial relationship between the camera and the corresponding IMU sensor being predefined; accessing sensor group data from the plurality of sensor groups; estimating spatial relationships between the plurality of sensor groups based on the sensor group data; and displaying virtual content in a display of the AR display device based on the spatial relationships between the plurality of sensor groups.
[0035] Accordingly, one or more of the methods described herein help to address the technical problem of inaccurately displaying stereoscopic images generated from a bendable device. In other words, bending of a flexible device causes errors in spatial relationships between cameras and a display. The presently described methods provide improvements to the operation of the functionality of a computing device by correctly estimating spatial correlations of components of a bent flexible device. Accordingly, one or more of the methods described herein can eliminate the need for certain workloads or computational resources. Examples of such computational resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.
[0036] Figure 1is a network diagram illustrating an environment 100 suitable for operating a display device 106, in accordance with some example embodiments. The environment 100 includes a user 102, a display device 106, and a physical object 104. The user 102 operates the display device 106. The user 102 can be a human user (e.g., a human), a machine user (e.g., a computer configured by a software program to interact with the display device 106), or any suitable combination thereof (e.g., a human assisted by a machine or a machine supervised by a human). The user 102 is associated with the display device 106.
[0037] The display device 106 includes a flexible device. In one example, the flexible device includes a computing device having a display such as eyeglasses. In one example, the display includes a screen that displays images captured by a camera of the display device 106. In another example, the display of the device can be transparent, for example in the lens of wearable computing eyeglasses.
[0038] The display device 106 includes an AR application that generates virtual content based on images detected with a camera of the display device 106. For example, the user 102 can direct the camera of the display device 106 to capture an image of the physical object 104. The AR application generates virtual content corresponding to an identified object (e.g., the physical object 104) in the image and presents the virtual content in the display of the display device 106.
[0039] The display device 106 includes a visual tracking system 108. The visual tracking system 108 tracks a pose (e.g., a position and an orientation) of the display device 106 relative to a real-world environment 110 using, for example, optical sensors (e.g., depth-enabled 3D cameras, image cameras), inertial sensors (e.g., gyroscopes, accelerometers, magnetometers), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors. The visual tracking system 108 can include a VIO system. In one example, the display device 106 displays virtual content based on the pose of the display device 106 relative to the real-world environment 110 and / or the physical object 104.
[0040] Figure 1 Any of the machines, databases, or devices shown can be implemented in a general-purpose computer modified (e.g., configured or programmed) by software to be a special-purpose computer to perform one or more of the functions described herein for that machine, database, or device. For example, one or more of the following may Figures 8 to 10A computer system capable of implementing any one or more of the methodologies described herein is discussed. As used herein, a "database" is a data storage resource and can store data structured as a text file, a table, a spreadsheet, a relational database (e.g., an object- relational database), a triple store, a hierarchical data store, or any suitable combination thereof. Moreover, Figure 1 Any two or more of the machines, databases, or devices illustrated in FIG. 6 can be combined into a single machine, and the functions described herein for any single machine, database, or device can be subdivided among multiple machines, databases, or devices.
[0041] The display device 106 can operate over a computer network. The computer network can be any network that enables communication among or between machines, databases, and devices. Thus, the computer network can be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. The computer network can include one or more portions that constitute a private network, a public network (e.g., the Internet), or any suitable combination thereof.
[0042] Figure 2 is a block diagram illustrating modules (e.g., components) of the display device 106 according to some example embodiments. The display device 106 includes a sensor 202, a display 204, a processor 208, and a storage device 206. Examples of the display device 106 include a wearable computing device (e.g., smart glasses).
[0043] The sensor 202 includes, for example, a proximity sensor 230, an IMUC 228, an IMU A 214 rigidly connected to / mounted to a camera A 212, an IMU B 226 rigidly connected to / mounted to a camera B 224. Due to the rigid connection, the spatial relationship between the camera A 212 and the IMU A 214 is already predefined and constant. Moreover, mounting each IMU to the corresponding camera enables more precise and accurate pixel measurement control (e.g., scan line triggering) and precise triggering of pixel acquisition across the stereo baseline for accurate stereo matching and depth inference. Examples of the cameras include a color camera, a thermal camera, a depth sensor camera, a grayscale camera, a global shutter tracking camera. Each IMU sensor includes a combination of a gyroscope, an accelerometer, or a magnetometer.
[0044] Additional standalone IMUs can be placed where deformation is also to be estimated at lower precision. For example, the IMUC 228 can be rigidly connected to another component (e.g., the display 204). Information about the deformation is also used to estimate the spatial correlation of other components. Accurate tracking of the spatial correlation among different components enables the display device 106 to maintain the best user experience.
[0045] The proximity sensor 230 is configured to detect whether the display device 106 is being worn. Other examples of the sensor 202 include a position sensor (e.g., near field communication, GPS, Bluetooth, Wi-Fi), an audio sensor (e.g., a microphone), or any suitable combination thereof. Note that the sensor 202 described herein is for purposes of illustration and thus the sensor 202 is not limited to the sensors described above.
[0046] The display 204 includes a screen or monitor configured to display images generated by the processor 208. In one example implementation, the display 204 can be transparent such that the user 102 can see through the display 204 (in an AR use case). In another example implementation, the display 204 covers the eyes of the user 102 and blocks the entire field of view of the user 102 (in a VR use case). In another example, the display 204 includes a touch screen display configured to receive user input via contact on the touch screen display.
[0047] The processor 208 includes an AR application 210, a visual tracking system 108, and a dynamic bending estimation module 216. The AR application 210 uses computer vision to detect and recognize the physical environment or physical objects 104. The AR application 210 retrieves a virtual object (e.g., a 3D object model) based on the recognized physical objects 104 or physical environment. The AR application 210 renders the virtual object in the display 204. In one example, the AR application 210 includes a local rendering engine that generates a visualization of the virtual object overlaid on (e.g., superimposed on or otherwise displayed contemporaneously with) images of the physical objects 104 captured by the camera A 212 or the camera B 224. The visualization of the virtual object can be manipulated by adjusting the positioning (e.g., its physical location, orientation, or both) of the physical objects 104 relative to the camera A 212 / camera B 224. Similarly, the visualization of the virtual object can be manipulated by adjusting the pose of the display device 106 relative to the physical objects 104.
[0048] The visual tracking system 108 estimates the pose of the visual tracking system 108. For example, the visual tracking system 108 uses image data (from the camera A 212 and the camera B 224) and corresponding inertial data (from the IMU A 214 and the IMU B 226) to track the position and pose of the display device 106 relative to a frame of reference (e.g., the real-world environment 110). In one example, the visual tracking system 108 includes a VIO system as previously described above.
[0049] The dynamic bend estimation module 216 accesses IMU data and / or camera data from the visual tracking system 108 to estimate changes in spatial motion between components (e.g., camera A 212, camera B 224, display 204). From the IMU data / camera data, the dynamic bend estimation module 216 can predict and estimate frame bending prior to processing camera frames from camera A 212 and camera B 224 by the AR application 210 and / or the visual tracking system 108. Examples of bend estimations include pitch, roll, and yaw bends. The dynamic bend estimation module 216 corrects for the offset to mitigate any display errors from bending. More details are described below with respect to Figure 4 Example components of the dynamic bend estimation module 216 are described in more detail.
[0050] The storage device 206 stores virtual object content 218, factory calibration data 222, and bend data 220. The factory calibration data 222 includes predefined values for spatial relationships between components (e.g., camera A 212 and camera B 224). The bend data 220 includes values for estimated bends (pitch-roll bias and yaw bend offset) of the display device 106. The virtual object content 218 includes, for example, a database of visual references (e.g., images) and corresponding experiences (e.g., three-dimensional virtual objects, interactive features of three-dimensional virtual objects).
[0051] Any one or more of the modules described herein can be implemented using hardware (e.g., a processor of a machine) or a combination of hardware and software. For example, any module described herein can configure a processor to perform the operations described herein for that module. Moreover, any two or more of these modules can be combined into a single module, and the functions described herein for a single module can be divided among multiple modules. Furthermore, according to various example embodiments, modules described herein as being implemented within a single machine, database, or device can be distributed across multiple machines, databases, or devices.
[0052] Figure 3 A visual tracking system 108 according to one example embodiment is shown. The visual tracking system 108 includes, for example, a VIO system 302 and a depth map system 304. The VIO system 302 accesses inertial sensor data from sensor group 306, which includes an IMU A 214 rigidly / tightly coupled to camera A 212, and sensor group 308, which includes an IMU B 226 rigidly / tightly coupled to camera B 224.
[0053] The VIO system 302 determines a pose (e.g., position, location, orientation) of the display device 106 relative to a frame of reference (e.g., the real-world environment 110). In one example implementation, the VIO system 302 estimates the pose of the display device 106 based on a 3D map of feature points from images captured with the sensor groups 306 and 308.
[0054] The depth map system 304 accesses image data from the camera A 212 and generates a depth map based on VIO data (e.g., feature point depths) from the VIO system 302. For example, the depth map system 304 generates a depth map based on depths of matching features between a left image (generated by the left camera) and a right image (generated by the right camera). In another example, the depth map system 304 generates a depth map based on triangulation of element differences in stereo images.
[0055] Figure 4 FIG. 4 is a block diagram illustrating a dynamic bend estimation module 216, according to one example implementation. The dynamic bend estimation module 216 includes a spatial relationship module 402, a bend estimation module 404, and a mitigation module 406.
[0056] The spatial relationship module 402 estimates spatial correlations between sensor groups. A sensor group includes a combination of zero or more IMUs, zero or more cameras, or a combination of previous devices coupled to a non-sensor component (e.g., a display 204, a projector, or an actuator). For example, a sensor group can include an IMU, a camera, and a component. Figure 4 Example combinations of sensor groups are illustrated: the sensor group 306 (including the camera A 212 rigidly coupled to the IMU A 214), the sensor group 308 (including the camera B 224 rigidly coupled to the IMU B 226), the sensor group 408 (including the display 204 rigidly coupled to the IMU C 228), the sensor group 414 (including the camera C 410 rigidly coupled to the actuator 418), the sensor group 416 (including the camera D 412). The spatial relationship module 402 accesses image data and / or IMU sensor data from the sensor groups. The spatial relationship module 402 retrieves factory calibration data 222 for the sensors from the storage device 206.
[0057] The spatial relationship module 402 accesses IMU data (from IMU A 214, IMU B 226, and IMU C 228) and image data (from camera A 212, camera B 224, camera C 410, camera D 412) and estimates spatial relationships between sensor groups (e.g., sensor group 306, sensor group 308, sensor group 408, camera C 410, camera D 412). The spatial relationship module 402 estimates changes in relative positioning / location between IMUs and / or cameras.
[0058] In one example, the spatial relationship module 402 accesses image data and IMU data from sensor group 306, sensor group 308 to estimate changes in relative positioning / location of each sensor group. For example, the spatial relationship module 402 uses a combination of image data and IMU data to detect that sensor group 306 (camera A 212 / IMU A 214) has moved closer to or farther away from sensor group 308 (camera B 224 / IMU B 226). The spatial relationship module 402 detects that sensor group 306 has moved closer to or farther away from sensor group 308 based on factory calibration data 222 and sensor data from the combination of camera A 212 and IMU A 214 and the combination of camera B 224 and IMU B 226 from an initial spatial positional relationship between sensor group 306 and sensor group 308.
[0059] In another example, the spatial relationship module 402 detects that display 204 is located closer to or farther away from camera A 212 based on sensor data from sensor group 306 and sensor data from sensor group 408. The sensor data from sensor group 306 includes image data from camera A 212 and / or IMU data from IMU A 214. The sensor data from sensor group 408 includes IMU data from IMU C 228.
[0060] In another example, the spatial relationship module 402 detects that actuator 418 is located closer to or farther away from camera B 224 based on sensor data from sensor group 308 and sensor data from sensor group 414. The sensor data from sensor group 308 includes image data from camera B 224 and / or IMU data from IMU B 226. The sensor data from sensor group 414 includes image data from camera C 410.
[0061] In another example, the spatial relationship module 402 detects, based on sensor data from the sensor group 408 and sensor data from the sensor group 416, that the camera D 412 is located closer or further away from the display 204. In another example, the spatial relationship module 402 detects that the orientation between the camera D 412 and the display 204 has changed. Thus, the spatial relationship module 402 detects linear and / or rotational displacements. The sensor data from the sensor group 416 includes image data from the camera D 412. The sensor data from the sensor group 408 includes IMU data from the IMUC 228.
[0062] The spatial relationship module 402 starts operating at runtime using the factory calibration data 222 as a starting point. The spatial relationship module 402 updates the spatial positional relationships between sensor groups (e.g., the sensor group 306, the sensor group 308, and the sensor group 408, the camera D 412, the sensor group 416) based on dynamic IMU data / image data from each IMU (e.g., the IMU A 214, the IMU B 226, the IMUC 228) and / or camera (e.g., the camera A 212, the camera B 224, the camera C 410, the camera D 412).
[0063] The bend estimation module 404 updates the spatial relationships of non-instrumented components (e.g., the actuator 418, the display 204, temporarily disabled sensor groups). For example, the bend estimation module 404 estimates mechanical deformations (e.g., bending of the display device 106) based on changes in the spatial relationships between components based on the IMU data / image data.
[0064] In one example, the bend estimation module 404 estimates the magnitude and direction of the bending of the eyewear frame of the display device 106: the temples of the eyewear frame splay outward, the distance between the cameras decreases, the orientation of the cameras deviates by x degrees, the relative position between the display 204 and the actuator 418 changes, etc.... Since the IMUs have a high sampling rate and are rigidly mounted to the non-instrumented components, the bend estimation module 404 can use the IMU / camera sensor data to measure high dynamic mechanical deformations and predict spatial changes between image frames. The bend estimation module 404 can predict / estimate the frame bending before processing the camera frames. Advantages of such bend prediction include: more accurate changes in search range, more robust measurement tracking, system state closer to the true state, resulting in smaller linearization errors and smaller corrections. Other advantages include faster processing (e.g., smaller search range, fewer estimator / optimizer iterations).
[0065] Once the spatial relationship module 402 and the bend estimation module 404 determine the updated spatial relationship between the sensor groups, the mitigation module 406 compares the updated spatial relationship between the sensor groups to the expected spatial relationship (based on the factory calibration data 222). The mitigation module 406 uses the difference between the updated spatial relationship and the factory calibrated relationship to calculate the amount of bend or misalignment that has occurred. The mitigation module 406 mitigates this bend or misalignment by adjusting the display or other components of the HMD in real-time. For example, if the display is misaligned, the mitigation module 406 can adjust the display image to compensate for the misalignment so that the user perceives a properly aligned virtual environment.
[0066] In one example, the mitigation module 406 uses the spatial correlation data from the spatial relationship module 402 and the bend estimation module 404 to perform operations to mitigate spatial changes between the sensor groups due to mechanical deformation of the display device 106. Examples of mitigation functions include: (1) depth correction, (2) rendering correction, and (3) feature prediction and feature tracking. Depth correction involves adjusting detected depths in images based on the relative spatial changes between the camera devices of the sensor groups. Rendering correction involves adjusting the rendering of virtual objects based on the relative spatial changes between the sensor groups. Feature prediction involves using spatial correlation data and information from previous frames to predict the location of objects or features in a current frame. Feature tracking involves using spatial correlation data and image data to identify and follow specific objects or features through a sequence of image or video frames. In one example, the mitigation module 406 is able to predict and track the location of sparse features based on the updated bend information and IMU data.
[0067] In one example, the mitigation module 406 uses the deformation estimate to correct for bends that cause pitch, roll, or yaw deviations of the display device 106. The mitigation module 406 determines a corrected coordinate system based on the deformation estimate. In other examples, the mitigation module 406 adjusts / realigns image data from the camera devices based on the bend estimate. In other examples, the mitigation module 406 determines whether to correct for bends that cause rotational deviations (e.g., yaw / pitch-roll offsets) of the display device 106. The mitigation module 406 minimizes any deformation based on the bend estimate provided by the bend estimation module 404. For example, the mitigation module 406 is able to minimize yaw deviations between VIO depth and stereo depth algorithm results by correcting the yaw estimate. The corrected configuration is then communicated to the AR application 210 for displaying content based on the corrected configuration.
[0068] To reduce computational complexity, spatial correlations can be selected dynamically when an estimation is to be made. For example, the proximity sensor 230 detects that the user is wearing the display device 106 to determine whether the spatial correlations between components have changed. Accordingly, the dynamic bend estimation module 216 operates in response to detecting that the user is wearing the display device 106. Accordingly, computational workload can be reduced to conserve power while maintaining high quality and accuracy of the AR experience.
[0069] Figure 5 is a block diagram illustrating a corrected depth frame according to one embodiment. The camera A 212 generates a frame t-1 504 at time t-1. The dynamic bend estimation module 216 determines a bend estimate (e.g., a bend estimate proximate to T 502) based on IMU data between the time at which the camera generates the frame t-1 504 and the time at which the camera generates a subsequent frame (e.g., frame t 508 at time t). Prior to processing the frame t 508, the mitigation module 406 applies a correction at 506.
[0070] Figure 6 Rigid component coupling and non-rigid component coupling are illustrated according to one example embodiment. The non-rigid setup 602 illustrates that the IMU 606 is flexibly coupled to the camera 608. Accordingly, the spatial relationship between the IMU 606 and the camera 608 can change due to physical deformation of the display device 106 due to the user wearing the display device 106, as illustrated in the arrangement 614. The physical deformation causes misalignment between the IMU 606 and the camera 608 and results in inaccurate / offset sensor measurements 610.
[0071] The rigid setup 604 illustrates that the IMU 606 is rigidly coupled to the camera 608. Accordingly, the spatial relationship between the IMU 606 and the camera 608 remains constant regardless of physical deformation of the display device 106 due to the user wearing the display device 106, as illustrated in the arrangement 616. IMU data from the IMU 606 can be used to predict the deformation and correct for any misalignment between the cameras for accurate sensor measurements 612.
[0072] Figure 7 Example user trajectories 702 of the display device 106 worn by the user 102 are illustrated according to one example embodiment. The IMU and camera are used to extract tracking information. Group 1 camera samples 708 are based on group 1 IMU and camera sensors. Group 2 camera samples 712 are based on group 2 IMU and camera sensors.
[0073] Because the IMUs have a high sampling rate and are rigidly mounted to the corresponding camera, their data is used to measure high dynamic mechanical deformations and to predict the spatial changes of the camera between image frames. For example, sensor group 2 trajectory 704 includes several IMU samples (e.g., group 2 IMU sample 714) between group 2 camera sample 712 and group 2 camera sample 716. Similarly, sensor group 1 trajectory 706 includes several IMU samples (e.g., group 1 IMU sample 710) between group 1 camera sample 708 and group 1 camera sample 718.
[0074] Figure 8 is a flowchart illustrating a method 800 for adjusting a frame according to one example implementation. The operations in method 800 can be performed by dynamic bend estimation module 216 using components (e.g., modules, engines) described above with respect to FIG. 4. Thus, spatial relationship module 402 is described by way of example with respect to method 800. However, it should be appreciated that at least some of the operations of method 800 can be deployed on various other hardware configurations, or performed by similar components residing elsewhere. Figure 4
[0075] In block 802, spatial relationship module 402 accesses factory calibration data for display device 106. In block 804, spatial relationship module 402 accesses (current) IMU data and / or image data for each sensor group. In block 806, spatial relationship module 402 determines updated spatial relationships between each sensor group based on the factory calibration data and the IMU data / image data. In block 808, bend estimation module 404 determines a frame bend estimate based on the updated spatial relationships. In block 810, mitigation module 406 generates a spatial change mitigation based on the frame bend estimate. In block 812, mitigation module 406 adjusts camera frames based on the spatial change mitigation. In another example, mitigation module 406 adjusts processing of camera frames based on the functionality described above with respect to mitigation module 406 in FIG. 4. Figure 4
[0076] It is noted that other implementations can use different ordering, additional or fewer operations, and different nomenclature or terminology to accomplish similar functionality. In some implementations, various operations can be performed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein were selected to illustrate some principles of operation in a simplified form.
[0077] Figure 9 is a flowchart illustrating a method 900 for processing image frames according to one example implementation. The operations in method 900 can be performed by dynamic bend estimation module 216 using components (e.g., modules, engines) described above with respect to FIG. 4. Thus, spatial relationship module 402 is described by way of example with respect to method 900. However, it should be appreciated that at least some of the operations of method 900 can be deployed on various other hardware configurations, or performed by similar components residing elsewhere. Figure 4 The operations of method 900 are illustrated in an example by reference to space relationship module 402. However, it will be understood that at least some of the operations of method 900 can be performed by other components in various other hardware configurations, or performed by similar components residing at other locations.
[0078] In block 902, prior to processing camera frames, bend estimation module 404 predicts a frame bend estimate based on the spatial relationship. In block 904, mitigation module 406 generates a camera spatial variation adjustment. At block 906, AR application 210 processes camera frames based on the camera spatial variation adjustment.
[0079] Figure 10 is a flowchart illustrating a method 1000 for detecting spatial variation based on proximity sensors, according to one example implementation. The operations in method 1000 can be performed by dynamic bend estimation module 216 using the components described above with respect to Figure 4 The operations of method 1000 are illustrated in an example by reference to space relationship module 402. However, it will be understood that at least some of the operations of method 1000 can be performed by other components in various other hardware configurations, or performed by similar components residing at other locations.
[0080] In block 1002, space relationship module 402 detects a spatial relationship change based on a triggering event determined by proximity sensors 230. In block 1004, space relationship module 402 determines spatial relationships between each sensor group in response to detecting the spatial relationship change. In block 1006, bend estimation module 404 determines a frame bend estimate based on the spatial relationships. In block 1008, mitigation module 406 generates a camera spatial variation mitigation. In block 1010, AR application 210 adjusts image frames based on the camera spatial variation mitigation.
[0081] System with head wearable device
[0082] Figure 11 A network environment 1100 in which a head wearable device 1102 can be implemented is shown, according to one example implementation. Figure 11 is a high-level functional block diagram of an example head wearable device 1102 communicatively coupling a mobile client device 1138 and a server system 1132 via various networks 1140.
[0083] The head wearable device 1102 includes a camera, such as at least one of a visible light camera 1112, an infrared emitter 1114, and an infrared camera 1116. The client device 1138 can be able to connect with the head wearable device 1102 using both the communication 1134 and the communication 1136. The client device 1138 is connected to the server system 1132 and the network 1140. The network 1140 can include any combination of wired and wireless connections.
[0084] The head wearable device 1102 also includes two image displays of the optical assembly’s image display 1104. The two image displays include one image display associated with a left side of the head wearable device 1102 and one image display associated with a right side of the head wearable device 1102. The head wearable device 1102 also includes an image display driver 1108, an image processor 1110, low-power low-power circuitry 1126, and high-speed circuitry 1118. The optical assembly’s image display 1104 is used to present images and video to a user of the head wearable device 1102, including images that can include a graphical user interface.
[0085] The image display driver 1108 commands and controls the image displays in the optical assembly’s image display 1104. The image display driver 1108 can deliver image data directly to the image displays in the optical assembly’s image display 1104 for presentation or can have to convert the image data into a signal or data format suitable for delivery to the image display devices. For example, the image data can be video data formatted according to a compression format (e.g., H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc.), and still image data can be formatted according to a compression format (e.g., Portable Network Group (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable image file format (Exif), etc.).
[0086] As described above, the head wearable device 1102 includes a frame and a stem (or temple) extending from a side of the frame. The head wearable device 1102 also includes a user input device 1106 (e.g., a touch sensor or a push button), including an input surface on the head wearable device 1102. The user input device 1106 (e.g., a touch sensor or a push button) is used to receive input selections from a user that manipulate a graphical user interface of a presented image.
[0087] Figure 11The components shown for the head-worn wearable device 1102 are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in blocks, frames, hinges, or nose bridges of the head-worn wearable device 1102. The left and right sides may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, camera lenses, or any other corresponding visible light or light-capturing elements that can be used to capture data, including images of scenes with unknown objects.
[0088] The head-mounted wearable device 1102 includes a memory 1122 that stores instructions for performing a subset or all of the functions described herein. The memory 1122 may also include a storage device.
[0089] like Figure 11 As shown, the high-speed circuit system 1118 includes a high-speed processor 1120, a memory 1122, and a high-speed wireless circuit system 1124. In this example, an image display driver 1108 is coupled to the high-speed circuit system 1118 and operated by the high-speed processor 1120 to drive the left and right image displays in the image display 1104 of the optical components. The high-speed processor 1120 can be any processor capable of managing the operation of any general computing system required for high-speed communication and the head-worn wearable device 1102. The high-speed processor 1120 includes the processing resources required to manage high-speed data transmission over communication 1136 to a wireless local area network (WLAN) using the high-speed wireless circuit system 1124. In some examples, the high-speed processor 1120 executes an operating system (e.g., a LINUX operating system or another such operating system for the head-worn wearable device 1102), and the operating system is stored in the memory 1122 for execution. Among other duties, the high-speed processor 1120, which executes the software architecture of the head-worn wearable device 1102, manages data transmission with the high-speed wireless circuit system 1124. In some examples, the high-speed wireless circuit system 1124 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard (also referred to herein as Wi-Fi). In other examples, other high-speed communication standards can be implemented using the high-speed wireless circuit system 1124.
[0090] The low-power wireless circuit system 1130 and high-speed wireless circuit system 1124 of the head-mounted wearable device 1102 may include a short-range transceiver (Bluetooth). TMThe transceiver includes a wireless wide area network transceiver, a wireless local area network transceiver, or a wide area network transceiver (e.g., cellular or WiFi). The client device 1138, which includes a transceiver communicating via communications 1134 and communications 1136, can be implemented using the architectural details of the head-mounted wearable device 1102, as can other elements of the network 1140.
[0091] Memory 1122 includes any storage device capable of storing various data and applications, including camera data generated by the left and right infrared imaging devices 1116 and image processor 1110, and images for display generated by image display driver 1108 on an image display in an optical component image display 1104, etc. While memory 1122 is shown as integrated with high-speed circuitry 1118, in other examples, memory 1122 may be a separate, independent component of head-mounted wearable device 1102. In some such examples, wiring may provide a connection from image processor 1110 or low-power processor 1128 to memory 1122 via a chip including high-speed processor 1120. In other examples, high-speed processor 1120 may manage addressing of memory 1122, such that low-power processor 1128 will activate high-speed processor 1120 whenever a read or write operation involving memory 1122 is required.
[0092] like Figure 11 As shown, the low-power processor 1128 or high-speed processor 1120 of the head-mounted wearable device 1102 may be coupled to a camera device (visible light camera 1112; infrared emitter 1114 or infrared camera 1116), an image display driver 1108, a user input device 1106 (e.g., a touch sensor or a press button), and a memory 1122.
[0093] The head-mounted wearable device 1102 is connected to a host computer. For example, the head-mounted wearable device 1102 is paired with a client device 1138 via communication 1136, or connected to a server system 1132 via a network 1140. For example, the server system 1132 may be one or more computing devices as part of a service or network computing system, including a processor, memory, and a network communication interface to communicate with the client device 1138 and the head-mounted wearable device 1102 via the network 1140.
[0094] Client device 1138 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via network 1140, communication 1134, or communication 1136. Client device 1138 may also store at least a portion of the instructions for generating binaural audio content in the memory of client device 1138 to implement the functions described herein.
[0095] The output components of the head-wearable device 1102 include visual components, e.g., a display (e.g., a liquid crystal display (LCD), a plasma display panel (PDP), a light emitting diode (LED) display, a projector, or a waveguide). The image display of the optical assembly is driven by an image display driver 1108. The output components of the head-wearable device 1102 also include acoustic components (e.g., a speaker), haptic components (e.g., a vibratory motor), other signal generators, and so forth. The input components of the head-wearable device 1102, client device 1138, and server system 1132, e.g., user input devices 1106, can include alphanumeric input components (e.g., a keyboard, a touchscreen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), pointing components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instrument), tactile input components (e.g., a physical button, a touchscreen that provides locations and forces of touches or touch gestures, or other tactile input components), audio input components (e.g., a microphone), and the like.
[0096] The head-wearable device 1102 can optionally include additional peripheral device elements. Such peripheral device elements can include biometric sensors, additional sensors, or display elements integrated with the head-wearable device 1102. For example, the peripheral device elements can include any I / O components, including output components, motion components, positioning components, or any other such elements described herein.
[0097] For example, the biometric components include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body poses, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components include acceleration sensor components (e.g., accelerometer), gravitation sensor components, orientation sensor components (e.g., gyroscope), and the like. The positioning components include positioning sensor components for generating positioning coordinates (e.g., a Global Positioning System (GPS) receiver component), Wi-Fi or Bluetooth® TM transceivers, altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like. Such positioning system coordinates can also be received from the client device 1138 via the low-power wireless circuitry 1130 or the high-speed wireless circuitry 1124 through the communication 1136.
[0098] Figure 12is a block diagram 1200 illustrating software architecture 1204, which can be installed on any one or more of the devices described herein. The software architecture 1204 is supported by hardware such as machine 1202 that includes processors 1220, memory 1226, and I / O components 1238. In this example, the software architecture 1204 can be conceptualized as a stack of layers, where each layer provides particular functionality. The software architecture 1204 includes layers such as an operating system 1212, libraries 1210, frameworks 1208, and applications 1206. Operationally, the applications 1206 make function calls through the software stack using, for example, an application programming interface (API) 1250 and receive message 1252 in response to the function calls.
[0099] The operating system 1212 manages hardware resources and provides common services. The operating system 1212 includes, for example, a kernel 1214, services 1216, and drivers 1222. The kernel 1214 acts as an abstraction layer between the hardware and the other software layers. For example, the kernel 1214 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functionality. The services 1216 can provide other common services for the other software layers. The drivers 1222 are responsible for controlling or interfacing with the underlying hardware. For instance, the drivers 1222 can include display drivers, camera drivers, Bluetooth® drivers, flash or low-power drivers, flash drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), storage device drivers (e.g., Advanced Technology Attachment (ATA) or NVM Express (NVMe) drivers), audio drivers, power management drivers, and so forth.
[0100] The libraries 1210 provide a low-level common infrastructure used by the applications 1206. The libraries 1210 can include system libraries 1218 (e.g., C standard library) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, and the like. In addition, the libraries 1210 can include API libraries 1224, such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., an OpenGL framework used to render two-dimensional (2D) and three-dimensional (3D) graphics on a display), database libraries (e.g., SQLite to provide various relational database functions), web libraries (e.g., WebKit to provide web browsing functionality), and the like. The libraries 1210 also include a wide variety of other libraries 1228 to provide many other APIs to the applications 1206.
[0101] The framework 1208 provides more abstracted infrastructure that is used by the applications 1206. For example, the framework 1208 provides various graphical user interface (GUI) functions, high-level resource management, and high-level location services. The framework 1208 can provide a broad spectrum of other APIs that can be used by the applications 1206, some of which can be specific to a particular operating system or platform.
[0102] In an example implementation, the applications 1206 can include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a game application 1248, and a broad assortment of other applications such as a third-party application 1240. The applications 1206 are programs that execute functions defined in the programs. Various programming languages can be employed to create one or more of the applications 1206, structured in a variety of manners, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a specific example, the third-party application 1240 (e.g., an application developed using the ANDROID TM or IOS TM software development kit (SDK) by an entity other than the vendor of the particular platform) can be mobile software running on a mobile operating system such as IOS TM , ANDROID TM , In this example, the third-party application 1240 can make API calls 1250 to the operating system 1212 through a
[0103] Figure 13is a diagrammatic representation of the machine 1300 within which instructions 1308 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1300 to perform any one or more of the methodologies discussed herein can be executed. For example, the instructions 1308 can cause the machine 1300 to execute any one or more of the methods described herein. The instructions 1308 transform the general, non-programmed machine 1300 into a particular machine 1300 programmed to carry out the described and illustrated functions in the manner described. The machine 1300 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1300 can operate in the capacity of a server machine or a client machine in server-client network environments, or as a peer machine in peer-to-peer (or distributed) network environments. The machine 1300 can comprise, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1308, sequentially or otherwise, that specify actions to be taken by machine 1300. Further, while only a single machine 1300 is illustrated, the term “machine” shall also be taken to include a collection of machines that individually or jointly execute the instructions 1308 to perform any one or more of the methodologies discussed herein.
[0104] The machine 1300 can include processors 1302, memory 1304, and I / O components 1342, which each can communicate with one another by way of a bus 1344. In example embodiments, the processors 1302 (e.g., central processing units (CPUs), reduced instruction set computing (RISC) processors, complex instruction set computing (CISC) processors, graphics processing units (GPUs), digital signal processors (DSPs), ASICs, radio-frequency integrated circuits (RFICs), other processors, or any suitable combination thereof) can include, for example, a processor 1306 and a processor 1310 that execute the instructions 1308. The term “processor” is intended to include multi-core processors that can include two or more independent processors (sometimes referred to as “cores”) that can execute instructions contemporaneously. Although FIG. 13 shows multiple processors 1302, the machine 1300 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof. Figure 13 Although FIG. 13 shows multiple processors 1302, the machine 1300 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0105] The storage 1304 includes a main memory 1312, a static memory 1314, and a storage unit 1316, which can be accessed via the bus 1344 by the processor 1302. The main memory 1314, the static memory 1314, and the storage unit 1316 store the instructions 1308 embodying any one or more of the methodologies or functions described herein. The instructions 1308 can also reside, completely or
[0106] The I / O components 1342 can include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. The specific I / O components 1342 that are included in the particular machine will depend on the type of machine. For example, portable machines such as mobile phones will likely include a touch input device or other such input mechanisms, while a headless server machine will likely not include such a touch input device. It will be appreciated that the I / O components 1342 can include many other components that are not specifically shown in FIG. 13. In various example embodiments, the I / O components 1342 can include output components 1328 and input components 1330. The output components 1328 can include visual components (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth. The input components 1330 can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., a physical button, a touch screen that provides location and / or force of touches or touch gestures, or other tactile input components), audio input components (e.g., a microphone), and the like. Figure 13
[0107] In other example implementations, the I / O components 1342 can include biometric components 1332, motion components 1334, environmental components 1336, or positioning components 1338, among a variety of other components. Biometric components 1332, for example, include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body poses, or eye-tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components 1334 include acceleration sensor components (e.g., accelerometer), gravitation sensor components, orientation sensor components (e.g., gyroscope), and the like. Environmental components 1336 include, for example, illumination sensor components (e.g., photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometer), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to a surrounding physical environment. The positioning components 1338 include location sensor components (e.g., a GPS receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like.
[0108] Communication can be implemented using a wide variety of technologies. The I / O components 1342 further include communication components 1340 operable to couple components (e.g., low energy), components, and other communication components to provide communication via other modalities. The devices 1322 can be another machine or any of a wide array of peripheral devices (e.g., a peripheral device coupled via a USB).
[0109] Moreover, the communication components 1340 can detect identifiers or include components operable to detect identifiers. For example, the communication components 1340 can include radio frequency identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes), or acoustic detection components (e.g., microphones to identify tagged audio signals). In addition, a variety of information can be derived via the communication components 1340, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via cellular signal triangulation, location via the detection of NFC beacon signals that can indicate a particular location, and so forth.
[0110] The various memories (e.g., memory 1304, main memory 1312, static memory 1314, and / or the memory of processor 1302) and / or storage unit 1316 can store one or more sets of instructions and data structures (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. These instructions (e.g., instructions 1308), when executed by processor 1302, cause various operations to implement the disclosed embodiments.
[0111] The instructions 1308 can be transmitted or received over the network 1320 via the network interface device (e.g., a network interface component included in the communication components 1340) using a transmission medium and any one of a number of well-known transfer protocols (e.g., hypertext transfer protocol (HTTP)). Similarly, instructions 1308 can be transmitted or received using a transmission medium via the coupling 1326 (e.g., a peer-to-peer coupling) to device 1322.
[0112] While implementations have been described with reference to particular embodiments, it will be apparent to those of ordinary skill in the art that various modifications and changes can be made to these implementations without departing from the broader spirit and scope of the disclosure. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate various implementations of the present disclosure and together with the description serve to explain the principles of the subject matter. The illustrated implementations are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other implementations can be utilized and derived therefrom, such that structural and logical substitutions and changes can be made without departing from the scope of the disclosure. Thus, the detailed description is not to be taken in a limiting sense and the scope of various implementations is defined by the appended claims and their equivalents.
[0113] Such implementations of the inventive subject matter can be referred to herein, individually and / or collectively, by the term "application" merely for convenience and without intending to voluntarily limit the scope of this application to any single implementation or inventive concept if there is more than one. Thus, although specific implementations have been illustrated and described herein, it should be appreciated that any arrangement can be substituted for the specific implementations shown. This disclosure is intended to cover any and all changes or modifications within the scope of the various implementations. Combinations of the above embodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description.
[0114] The abstract of the disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or the meaning of the claims. In addition, in the foregoing detailed description, it can be seen that various features are grouped together in a single embodiment for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments require more features than are explicitly recited in each claim. Rather, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby expressly incorporated into this detailed description, with each claim acting as a separate embodiment of the disclosure.
[0115] Examples
[0116] Example 1 is a method comprising: forming a plurality of sensor groups of an augmented reality (AR) display device, wherein one of the plurality of sensor groups comprises a combination of a camera, an IMU (inertial measurement unit), and a component, each of the camera, the IMU, and the component are tightly coupled to each other, a spatial relationship between the camera, the IMU sensor, or the component is predefined; accessing sensor group data from the plurality of sensor groups; estimating a spatial relationship between the plurality of sensor groups based on the sensor group data; and displaying virtual content in a display of the AR display device based on the spatial relationship between the plurality of sensor groups.
[0117] In Example 2, the subject matter of Example 1 includes: accessing factory calibration data indicating a static spatial relationship or a dynamic spatial relationship among the plurality of sensor groups, wherein the spatial relationship between each sensor group is predefined, wherein estimating the spatial relationship between the plurality of sensor groups is based on the factory calibration data.
[0118] In Example 3, the subject matter of Example 2 includes: wherein, wherein one of the plurality of sensor groups comprises one of: an IMU sensor tightly coupled to the camera, an IMU sensor tightly coupled to the component, and a camera tightly coupled to the component, wherein the component comprises one of a display component, a projector, an illuminator, a LIDAR component, or an actuator.
[0119] In Example 4, the subject matter of Example 3 includes: estimating a spatial relationship between a first sensor group and a second sensor group of the plurality of sensor groups based on the sensor group data; and estimating a bend of the AR display device based on the spatial relationship between the first sensor group and the second sensor group.
[0120] In Example 5, the subject matter of Example 4 includes: wherein estimating the spatial relationship between the plurality of sensor groups is based on a combination of factory calibration data and image data and IMU data from each sensor group.
[0121] In Example 6, the subject matter of Examples 1-5 includes: wherein estimating the spatial relationship between the plurality of sensor groups further comprises: fusing data from the plurality of sensor groups; and correcting one or more sensor data based on the fused data.
[0122] In Example 7, the subject matter of Examples 1-6 includes capturing a first set of image frames from a set of cameras of an AR display device; accessing IMU data between the first set of image frames and a second set of image frames, wherein the second set of image frames is generated after the first set of image frames; estimating a first spatial relationship between each camera of the set of cameras for the first set of image frames; estimating a second spatial relationship between each camera of the set of cameras for the second set of image frames using the IMU data; and processing the second set of image frames based on the first spatial relationship and the second spatial relationship.
[0123] In Example 8, the subject matter of Example 7 includes wherein processing the second set of image frames includes adjusting a predicted position of virtual content in the second set of image frames based on the second spatial relationship between each camera of the set of cameras for the second set of image frames.
[0124] In Example 9, the subject matter of Examples 1-8 includes wherein the AR display device includes a proximity sensor, wherein the method includes detecting a triggering event based on proximity data from the proximity sensor, wherein estimating the spatial relationship between the plurality of sensor groups is performed in response to detecting the triggering event.
[0125] In Example 10, the subject matter of Examples 1-9 includes wherein the AR display device includes an eyewear frame.
[0126] Example 11 is a computing device comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to perform operations comprising: forming a plurality of sensor groups of an augmented reality (AR) display device, wherein one of the plurality of sensor groups comprises a combination of a camera, an IMU (inertial measurement unit), and a component, each of the camera, the IMU, and the component are tightly coupled to each other, a spatial relationship between the camera, the IMU sensor, or the component is predefined; accessing sensor group data from the plurality of sensor groups; estimating a spatial relationship between the plurality of sensor groups based on the sensor group data; and displaying virtual content in a display of the AR display device based on the spatial relationship between the plurality of sensor groups.
[0127] In Example 12, the subject matter of Example 11 includes wherein the operations comprise accessing factory calibration data indicating a static spatial relationship or a dynamic spatial relationship among the plurality of sensor groups, wherein the spatial relationship between each sensor group is predefined, wherein estimating the spatial relationship between the plurality of sensor groups is based on the factory calibration data.
[0128] In Example 13, the subject matter of Example 12 includes, wherein, wherein one of the plurality of sensor groups includes one of: an IMU sensor tightly coupled to a camera, an IMU sensor tightly coupled to a component, and a camera tightly coupled to a component, wherein the component includes one of a display component, a projector, a luminaire, a LIDAR component, or an actuator.
[0129] In Example 14, the subject matter of Example 13 includes, wherein the operations include: estimating a spatial relationship between a first sensor group and a second sensor group of the plurality of sensor groups based on the sensor group data; and estimating a bend of the AR display device based on the spatial relationship between the first sensor group and the second sensor group.
[0130] In Example 15, the subject matter of Example 14 includes, wherein estimating the spatial relationship between the plurality of sensor groups is based on a combination of factory calibration data and image data and IMU data from each sensor group.
[0131] In Example 16, the subject matter of Examples 11-15 includes, wherein estimating the spatial relationship between the plurality of sensor groups further includes: fusing data from the plurality of sensor groups; and correcting one or more sensor data based on the fused data.
[0132] In Example 17, the subject matter of Examples 11-16 includes, wherein the operations include: capturing a first set of image frames from a set of cameras of the AR display device; accessing IMU data between the first set of image frames and a second set of image frames, wherein the second set of image frames is generated after the first set of image frames; estimating a first spatial relationship between each camera of the set of cameras for the first set of image frames; estimating a second spatial relationship between each camera of the set of cameras for the second set of image frames with the IMU data; and processing the second set of image frames based on the first spatial relationship and the second spatial relationship.
[0133] In Example 18, the subject matter of Example 17 includes, wherein processing the second set of image frames includes: adjusting a predicted position of virtual content in the second set of image frames based on the second spatial relationship between each camera of the set of cameras for the second set of image frames.
[0134] In Example 19, the subject matter of Examples 11-18 includes, wherein the AR display device includes a proximity sensor, wherein the method includes: detecting a triggering event based on proximity data from the proximity sensor, wherein estimating the spatial relationship between the plurality of sensor groups is performed in response to detecting the triggering event.
[0135] Example 20 is a non-transitory computer-readable storage medium including instructions that when executed by a computer cause the computer to: form a plurality of sensor groups of an augmented reality (AR) display device, wherein one of the plurality of sensor groups includes a combination of a camera, an IMU (inertial measurement unit), and a component, each of the camera, the IMU, and the component being tightly coupled to each other, a spatial relationship between the camera, the IMU sensor, or the component being predefined; access sensor group data from the plurality of sensor groups; estimate a spatial relationship between the plurality of sensor groups based on the sensor group data; and display virtual content in a display of the AR display device based on the spatial relationship between the plurality of sensor groups.
[0136] Example 21 is at least one machine readable medium including instructions that when executed by processing circuitry cause the processing circuitry to perform operations for implementing any of Examples 1 to 20.
[0137] Example 22 is an apparatus comprising means for implementing any of Examples 1 to 20.
[0138] Example 23 is a system implementing any of Examples 1 to 20.
[0139] Example 24 is a method implementing any of Examples 1 to 20.
Claims
1. A method comprising: forming a plurality of sensor groups of an augmented reality (AR) display device, wherein one sensor group of the plurality of sensor groups comprises a combination of a camera, an IMU (inertial measurement unit), and a component, each of the camera, the IMU, and the component being tightly coupled to each other, a spatial relationship between the camera, the IMU sensor, or the component being predefined; accessing sensor group data from the plurality of sensor groups; estimating a spatial relationship between the plurality of sensor groups based on the sensor group data; and displaying virtual content in a display of the AR display device based on the spatial relationship between the plurality of sensor groups.
2. The method of claim 1, further comprising: accessing factory calibration data indicative of a static spatial relationship or a dynamic spatial relationship among the plurality of sensor groups, wherein a spatial relationship between each sensor group is predefined, wherein estimating the spatial relationship between the plurality of sensor groups is based on the factory calibration data.
3. The method of claim 2, further comprising: wherein one sensor group of the plurality of sensor groups comprises one of: an IMU sensor tightly coupled to the camera, an IMU sensor tightly coupled to the component, and the camera tightly coupled to the component, wherein the component comprises one of a display component, a projector, an illuminator, a LIDAR component, or an actuator.
4. The method of claim 3, further comprising: estimating a spatial relationship between a first sensor group and a second sensor group of the plurality of sensor groups based on the sensor group data; and wherein estimating the spatial relationship between the plurality of sensor groups is based on a combination of the factory calibration data and image data and IMU data from each sensor group. estimating the spatial relationship between the plurality of sensor groups further comprises:
5. The method of claim 4, wherein, fusing data from the plurality of sensor groups; and 6. The method of claim 1, wherein, correcting one or more sensor data based on the fused data.
7. The method of claim 1, further comprising: capturing a first set of image frames from a set of cameras of the AR display device; accessing IMU data between the first set of image frames and a second set of image frames, wherein the second set of image frames is generated after the first set of image frames; estimating a first spatial relationship between each camera of the set of cameras for the first set of image frames; estimating a second spatial relationship between each camera of the set of cameras for the second set of image frames with the IMU data; and processing the second set of image frames based on the first spatial relationship and the second spatial relationship. processing the second set of image frames comprises: 8. The method of claim 7, wherein, adjust a predicted position of the virtual content in the second set of image frames based on the second spatial relationship between each camera of the set of cameras for the second set of image frames.
9. The method of claim 1, wherein, the AR display device includes a proximity sensor, wherein the method includes detecting a trigger event based on proximity data from the proximity sensor, wherein estimating the spatial relationship between the plurality of sensor groups is in response to detecting the trigger event.
10. The method of claim 1, wherein, the AR display device includes an eyewear frame.
11. A computing device comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to perform operations comprising: forming a plurality of sensor groups of an augmented reality (AR) display device, wherein one of the plurality of sensor groups includes a combination of a camera, an IMU (inertial measurement unit), and a component, each of the camera, the IMU, and the component being tightly coupled to one another, a spatial relationship between the camera, the IMU sensor, or the component being predefined; accessing sensor group data from the plurality of sensor groups; estimating a spatial relationship between the plurality of sensor groups based on the sensor group data; and displaying virtual content in a display of the AR display device based on the spatial relationship between the plurality of sensor groups.
12. The computing device of claim 11, wherein, the operations include: accessing factory calibration data indicating a static spatial relationship or a dynamic spatial relationship among the plurality of sensor groups, wherein a spatial relationship between each sensor group is predefined, wherein estimating the spatial relationship between the plurality of sensor groups is based on the factory calibration data.
13. The computing device of claim 12, wherein, one of the plurality of sensor groups includes one of: an IMU sensor tightly coupled to the camera, an IMU sensor tightly coupled to the component, and the camera tightly coupled to the component, wherein the component includes one of a display component, a projector, an illuminator, a LIDAR component, or an actuator.
14. The computing device of claim 13, wherein, the operations include: estimating a spatial relationship between a first sensor group and a second sensor group of the plurality of sensor groups based on the sensor group data; and estimating a bend of the AR display device based on the spatial relationship between the first sensor group and the second sensor group.
15. The computing device of claim 14, wherein, estimating the spatial relationship between the plurality of sensor groups is based on a combination of the factory calibration data and image data and IMU data from each sensor group.
16. The computing device of claim 11, wherein, estimating the spatial relationship between the plurality of sensor groups further includes: fusing data from the plurality of sensor groups; and correcting one or more sensor data based on the fused data.
17. The computing device of claim 11, wherein, the operations include: capturing a first set of image frames from a set of cameras of the AR display device; accessing IMU data between the first set of image frames and a second set of image frames, wherein the second set of image frames is generated after the first set of image frames; estimate a first spatial relationship between each camera of the set of cameras for the first set of image frames; estimate a second spatial relationship between each camera of the set of cameras for the second set of image frames using the IMU data; and process the second set of image frames based on the first spatial relationship and the second spatial relationship.
18. The computing device of claim 17, wherein, processing the second set of image frames includes: adjust a predicted position of the virtual content in the second set of image frames based on the second spatial relationship between each camera of the set of cameras for the second set of image frames.
19. The computing device of claim 11, wherein, the AR display device includes a proximity sensor, wherein the method includes detecting a trigger event based on proximity data from the proximity sensor, wherein estimating the spatial relationship between the plurality of sensor groups is performed in response to detecting the trigger event.
20. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to: forming a plurality of sensor groups of an augmented reality (AR) display device, wherein one of the plurality of sensor groups comprises a combination of a camera, an IMU (inertial measurement unit), and a component, each of the camera, the IMU, and the component being tightly coupled to each other, a spatial relationship between the camera, the IMU sensor, or the component being predefined; access sensor group data from the plurality of sensor groups; estimate a spatial relationship between the plurality of sensor groups based on the sensor group data; and display virtual content in a display of the AR display device based on the spatial relationship between the plurality of sensor groups.