Augmented reality using split architecture
Patent Information
- Application Number
- CN202280080922.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-08
- Filing Date
- 2022-12-07
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-12-07
Smart Images

Figure CN118302806B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This disclosure is a continuation of U.S. Patent Application No. 17 / 643,233, filed on December 8, 2021, and claims priority to that application. Technical Field
[0003] This disclosure relates to augmented reality, and more specifically, to a head-motion-based user interface such that the user interface appears to be registered in the real world when the user's head moves. Background Technology
[0004] Augmented reality user interfaces (UIs) can render content within a perspective overlay that appears layered on the user's real-world view (i.e., the optical perspective display). As the user's head moves, the graphics displayed on the optical perspective display may appear fixed in the real world (i.e., a world-locked UI). World-locked UIs are compatible with augmented reality (AR) glasses. For example, semi-transparent graphics displayed on AR glasses can reduce their impact on the user's view and may not require the user to constantly shift focus between the displayed graphics and the real world. Furthermore, spatially registering graphics with the real world can provide intuitive information. Therefore, world-locked UIs are particularly useful for applications requiring high cognitive loads, such as navigation. For example, world-locked UIs can be used in AR glasses to implement turn-by-turn navigation and / or destination recognition. Summary of the Invention
[0005] In at least one aspect, this disclosure generally describes a method for displaying AR elements on an augmented reality (AR) display. The method includes receiving an initial two-dimensional (2D) texture of the AR element at an AR device, the initial 2D texture being rendered at a computing device physically separated but (communicatively) coupled (in a split architecture) to the AR device. The method also includes distorting the initial 2D texture of the AR element at the AR device (by the AR device's processor) to generate a registered 2D texture of the AR element; and triggering the display of the registered 2D texture of the AR element on the AR display of the AR device.
[0006] In another aspect, this disclosure generally describes an AR glasses system comprising an inertial measurement unit (IMU) configured to collect IMU data, a camera configured to capture camera data, a wireless interface configured to transmit first information to and receive the first information from a computing device via a wireless communication channel, an AR display configured to display second information to a user of the AR glasses, and a processor configured by software to display AR elements on the AR display. For this purpose, the processor is configured to transmit IMU data and camera data to the computing device, such that the computing device can calculate initial (e.g., high-resolution) pose data based on the IMU data and camera data, estimate a first pose based on the initial (high-resolution) pose data and a latency estimate corresponding to rendering, and render a two-dimensional (2D) texture of the AR elements based on the first pose. The processor is then configured to receive the initial (high-resolution) pose data, the first pose, and the 2D texture of the AR elements from the computing device. The processor is configured to calculate corrected (high-resolution) pose data based on the IMU data and the initial (high-resolution) pose data. The processor is also configured to estimate a second pose based on corrected (high-resolution) pose data and to warp the 2D texture of the AR element based on a comparison of the second pose with the first pose. The processor is also configured to trigger the display of the warped 2D texture of the AR element on an AR display.
[0007] In one aspect, an augmented reality (AR) glasses includes: an inertial measurement unit (IMU) configured to collect IMU data; a camera configured to capture camera data; a wireless interface configured to transmit first information to and receive first information from a computing device via a wireless communication channel; an AR display configured to display second information to a user of the AR glasses; and a processor software-configured to: transmit IMU data and camera data to the computing device; receive initial high-resolution pose data, a first pose, and a two-dimensional (2D) texture of an AR element from the computing device; calculate corrected high-resolution pose data based on the IMU data and the initial high-resolution pose data; estimate a second pose based on the corrected high-resolution pose data; distort the 2D texture of the AR element based on a comparison of the second pose and the first pose; and trigger the display of the distorted 2D texture of the AR element on the AR display. The computing device may be configured to: calculate the initial high-resolution pose data based on the IMU data and camera data; estimate the first pose based on the initial high-resolution pose data and a latency estimate corresponding to rendering; and render the 2D texture of the AR element based on the first pose.
[0008] In another aspect, this disclosure generally describes a system, such as a split-architecture augmented reality system, comprising a communication-coupled computing device and AR glasses. In the split-architecture, the computing device is configured to compute initial (e.g., high-resolution) pose data, estimate rendering latency, estimate a first pose based on the rendering latency and the initial (high-resolution) pose data, and render a 2D texture of an AR element based on the first pose. In the split-architecture, the AR glasses are configured to collect inertial measurement unit (IMU) data and camera data; compute corrected (high-resolution) pose data based on the IMU data and the initial (high-resolution) pose data; and estimate a second pose based on the corrected (high-resolution) pose data. The AR glasses are also configured to compare the second pose and the first pose; distort the 2D texture of the AR element based on the comparison between the second pose and the first pose; and display the distorted 2D texture of the AR element on an AR display of the AR glasses.
[0009] The foregoing illustrative description of the invention, as well as other exemplary objects and / or advantages of this disclosure and the implementation methods of this disclosure, will be further explained in the following detailed description and accompanying drawings. Attached Figure Description
[0010] Figure 1A The UI for world-locking from a first-person perspective of augmented reality glasses is shown according to an implementation of this disclosure.
[0011] Figure 1B The UI for world-locking augmented reality glasses from a second-person perspective is shown in accordance with the implementation of this disclosure.
[0012] Figure 2 This is a flowchart illustrating a method for displaying world-locked AR elements on a display according to an implementation of this disclosure.
[0013] Figure 3 This illustrates possible implementations according to this disclosure. Figure 2 A flowchart detailing the rendering process.
[0014] Figure 4 This illustrates possible implementations according to this disclosure. Figure 2 A flowchart detailing the distortion process.
[0015] Figure 5 This is a perspective view of AR glasses according to a possible implementation of this disclosure.
[0016] Figure 6 This illustrates a possible split architecture for possible implementations according to this disclosure.
[0017] Figure 7A flowchart illustrating a method for augmented reality using a split architecture according to a possible implementation of this disclosure is shown.
[0018] The components in the accompanying drawings are not necessarily drawn to scale relative to each other. In several views, similar reference numerals indicate the corresponding parts. Detailed Implementation
[0019] This disclosure describes an augmented reality approach using a split architecture. Specifically, this disclosure relates to rendering a world-locked user interface (UI) for an optical see-through display on augmented reality (AR) glasses. Rendering graphics to make them appear world-locked (i.e., anchored) to a point in space while the user's head moves freely may require a high rendering rate to prevent lag between rendering and user head movement that could distract the user and / or disorient them. Furthermore, registration of graphics with the real world requires iterative measurement of the user's head position and orientation (i.e., pose) as part of the rendering process. A world-locked UI can pose challenges to the limited processing and / or power resources of AR glasses. Therefore, the measurement and rendering processes can be split between the AR glasses and another computing device (e.g., a mobile phone, laptop, tablet, etc.). This split-processing approach is called a split architecture.
[0020] The split architecture utilizes computing devices (e.g., mobile computing devices) with more processing and power resources than AR glasses to perform computationally complex rendering processes, while leveraging the AR glasses to perform less computationally complex rendering processes. Therefore, the split architecture can facilitate world-locked UIs for optical see-through displays on AR glasses without exhausting the AR glasses' processing / power resources.
[0021] Split architectures require communication between the mobile computing device and the AR glasses. This communication may ideally be wireless. Wireless communication may have high latency (e.g., 300 milliseconds (ms)) that can vary over a wide range (e.g., 100 ms). This latency can make rendering difficult because, when displaying graphics, rendering requires predicting the user's head position / orientation (i.e., pose). The latency variation caused by the wireless communication channel can make pose prediction in a split architecture less accurate. This disclosure includes systems and methods for mitigating the latency impact of wireless channels on world-locked user interfaces (UIs) for optical see-through displays on augmented reality (AR) glasses. The disclosed systems and methods can have the technical effect of providing AR elements (e.g., graphics) on the AR display that appear locked to real-world locations with less jitter and lag in response to movement. Furthermore, this disclosure describes systems and methods for distributing processing among multiple devices to alleviate the processing / power burden on AR glasses. Processing distribution can have the technical effect of extending the functionality of the AR glasses to perform applications such as navigation within the limited resources of this device (e.g., processing power, battery capacity).
[0022] Figures 1A to 1B The UI for world-locking augmented reality glasses in a split architecture with a computing device, according to an implementation of this disclosure, is shown. Figure 1A This shows the environment 100 as seen by the user from a first-person perspective (i.e., first viewpoint) through AR glasses. Figure 1B This illustrates the same environment 100 as seen by the user through AR glasses from a second perspective (i.e., second viewpoint). The viewpoint (i.e., the angle of view) can be determined based on the user's head position (i.e., x, y, z) and / or orientation (i.e., yaw, pitch, roll). The combination of position and orientation can be referred to as the user's head pose. For the first perspective (i.e., ... Figure 1A The user's head is in the first pose 101, while for the second perspective (i.e.) Figure 1B The user's head is in the second pose 102.
[0023] As shown in the figure, the optical see-through display (i.e., head-up display, AR display) of AR glasses is configured to display AR elements. AR elements can include any combination of one or more graphics, text, and images; AR elements can be static or animated (e.g., animation, video). Information associated with AR elements can be stored in memory as 3D assets. 3D assets can be in a file format (e.g., .OBJ format) that includes information describing the AR element in three dimensions. 3D assets can be rendered as two-dimensional (2D) images based on a defined viewpoint. 2D images that include necessary modifications (e.g., distortion) to depict them as if seen from a single viewpoint are called 2D textures (i.e., textures). AR elements can also include information describing where they should be anchored in the environment.
[0024] Here, the AR element is a perspective graphic of arrow 105, which is transformed into a 2D texture and overlaid on the user's view of environment 100. This texture is world-locked (i.e., anchored, registered) to a location within environment 100. Arrow 105 is world-locked to its position corresponding to the corridor, and its display guides the user along the corridor to help navigate to their destination. Arrow 105 is world-locked because its position relative to the corridor does not change as the user's posture changes (i.e., as the user's viewpoint changes). As part of a navigation application running on AR glasses, AR elements can be world-locked to a single location.
[0025] Figure 2 This is a flowchart illustrating a method 200 for displaying world-locked AR elements on a display of AR glasses in a split architecture. World-locked AR elements include determining the user's pose and rendering the AR object as a 2D texture as seen from the viewpoint. Due to the complexity of rendering, rendering may include latency (i.e., delay, bottleneck). To accommodate this latency, rendering may include estimating the position of the pose at the end of rendering based on measurements taken at the start of rendering. The accuracy of this estimation may affect the quality of world-locking. Poor estimation may result in position jitter and / or hysteresis of AR elements when repositioning in response to pose changes. It should be noted that, in addition to the user's head pose, the determined user pose may also depend on body position and / or eye (or binocular) position, and while the principles of the disclosed techniques can be adapted and / or extended to utilize such additional / additional pose information, the description of this disclosure is limited to the user's head pose.
[0026] A user's head pose can be described using six degrees of freedom (6DOF), which includes position (x, y, z) in a three-axis coordinate system and rotation (pitch, roll, yaw) in the same three-axis coordinate system. AR glasses can be configured for 6DOF tracking to provide pose information related to head pose at various times. For example, 6DOF tracking can include continuously streamed, timestamped head pose information.
[0027] 6DOF tracking can be performed by a 6DOF tracker 210 configured to receive measurements from sensors on the AR glasses. For example, the 6DOF tracker (i.e., a 6DOF estimator) can be coupled to the AR glasses' inertial measurement unit (IMU 201). IMU 201 may include a combination of at least accelerometers, gyroscopes, and magnetometers for measuring position and acceleration along each of the three dimensions. Individually, the positioning resolution provided by IMU 201 may be insufficient to accurately lock AR elements in the world. For example, IMU 201 may not provide accurate depth information about the environment, which could help in realistically rendering AR elements within the environment. Accordingly, the 6DOF tracker can also be coupled to the AR glasses' camera 202. Camera 202 can be configured to capture images of the user's field of view, which can be analyzed to determine the surface depth relative to the user within the field of view. This depth information can be used to improve the accuracy of the determined user pose. 6DOF tracking can be highly accurate when both IMU and camera data are used to calculate pose, but it consumes a lot of power, especially when looping at the rate required to capture fast movements (i.e., rapid head movements, rapid environmental changes) and using a camera.
[0028] At the first time point (t1), the 6DOF tracker outputs 6DoF information (i.e., 6DoF(t1)). This 6DoF information can be used to render AR elements based on the expected viewing position of anchor points in the environment at a location on the display after rendering 220. Therefore, rendering can include calculating the viewpoint (i.e., pose) of the rendered texture.
[0029] Figure 3 A flowchart illustrating a method for rendering a 2D texture based on 6DoF information according to a possible implementation of this disclosure is shown. Rendering 220 is a computationally complex process that may have a considerably long latency period, which, when longer than the head movement, can cause a significant lag in the displayed position. Therefore, rendering may include estimating (i.e., predicting) the first pose (P1) of the head at the end of rendering. In other words, 6DoF information (i.e., 6DoF(t1)) and an estimate of the latency period (Δt) are used. 估计The first pose (P1) can be estimated using time 221 for rendering. For example, 6DOF tracking obtained at a first time (t1) can be processed to determine the trajectory of head movement. The head movement can then be projected along the trajectory determined for the estimated time delay to estimate the first pose (P1). The accuracy of the first pose depends on the accuracy of the determined trajectory (i.e., the 6DoF measurement) and the estimation of the time delay. Rendering 220 may also include determining 222 the position on the display relative to the real-world anchor of the AR element and determining the viewpoint from which the anchor is viewed. Rendering may also include generating 223 a 2D texture of the 3D asset. The 3D asset can be retrieved from memory and processed to determine the rendered 2D texture (T1). Processing may include transforming the 3D asset into a 2D image with perspective features so that it appears as seen from the viewpoint of the first pose.
[0030] In reality, the actual rendering latency (Δt) 实际 This may differ from the estimated time delay (Δt). 估计 The pose estimation is different. Therefore, at the second time point (t2) after rendering ends, the actual pose of the user's head may not be equal to the estimated first pose (P1). Consequently, the rendering may show a position that does not match the expected anchor point in the real environment. To compensate for inaccurate pose estimation, the method also includes a temporal warp (i.e., warp 230) of the texture after rendering. Warp 230 involves shifting and / or rotating the rendered texture to register it at the appropriate viewpoint. Warp can also be referred to as a transformation. Since the shifting / rotation of warp 230 is likely much simpler computationally than rendering, it can be performed much faster, allowing the correction to not add any significant latency that could cause noticeable artifacts (e.g., jitter, lag) in the display. Therefore, in a split architecture, warp can be performed on the AR glasses while rendering can be performed on the computing device.
[0031] Figure 4A method for warping a rendered texture according to a possible implementation of this disclosure is illustrated. Warping 230 includes estimating 231 a second pose (P2) based on 6DOF information (6DoF(t2)) obtained at a second time (t2) after rendering. In other words, the second pose (P2) corresponds to the actual pose of the user's head after rendering 220. Warping 230 may also include comparing a first pose (P1) (i.e., the estimated pose after rendering) with the second pose (P2) (i.e., the actual pose after rendering) to determine the amount and / or type of warping required to register the rendered AR element with the second pose. For example, warping may include calculating 232 a transformation matrix (i.e., a warping transformation matrix) representing the warping from the first pose (P1) and the second pose (P2). Warping 230 may also include applying (e.g., multiplying) the rendered texture (T1) to (e.g., multiplying) the warping transformation matrix (W) to transform 233 the rendered texture into a warped 2D texture (i.e., the registered 2D texture) (T2). This transformation process may be referred to as a perspective transformation. The registered 2D texture corresponds to the latest (and more accurate) pose information captured at the second time.
[0032] Pose information for warping can be captured periodically. Therefore, in some implementations, the estimation of the second pose can be triggered by a synchronization signal (VSYNC) associated with the AR glasses' display. In some implementations, the timing derived from the synchronization signal (VSYNC) can provide a delay estimate that can be used to estimate the second pose. For example... Figure 2 As shown, after distortion 230, the method 200 for displaying world-locked AR elements on the display of AR glasses may include displaying 240 a registered texture on the display of AR glasses.
[0033] When the first pose (P1) matches the second pose (P2), no warp is required. In this case, warp 230 can be skipped (e.g., not triggered) or a unit warp transformation matrix can be applied. For example, the first pose can be matched with the second pose when there is no head movement during rendering, or when the estimated latency period used to generate the first pose (P1) (i.e., the estimated latency) matches the actual latency of the rendering.
[0034] Warping can be performed at any time after rendering. Therefore, rendering can be repeated at the applied rendering rate, while warping can be repeated at a higher rate (e.g., the display rate). Because the processes can run independently, warping is asynchronous with rendering. Therefore, warping can be called asynchronous time warping (ATW).
[0035] Figure 5This is a perspective view of AR glasses according to a possible implementation of this disclosure. AR glasses 500 are configured to be worn on a user's head and face. AR glasses 500 include a right temple 501 and a left temple 502 supported by the user's ears. AR glasses also include a nose bridge portion 503 supported by the user's nose, such that a left lens 504 and a right lens 505 can be positioned in front of the user's left and right eyes, respectively. The various parts of the AR glasses can be collectively referred to as the frame of the AR glasses. The frame of the AR glasses may contain electronics for enabling functionality. For example, the frame may include a battery, a processor, memory (e.g., a non-transitory computer-readable medium), and electronics for supporting sensors (e.g., cameras, depth sensors, etc.) and interface devices (e.g., speakers, displays, network adapters, etc.). The AR glasses can be displayed and sensed relative to a coordinate system 530. The coordinate system 530 can be aligned with the user's head posture when wearing the AR glasses. For example, the user's eyes may be along a line in the horizontal (e.g., x-direction) direction of the coordinate system 530.
[0036] Users wearing AR glasses can experience information displayed within one or more lenses, allowing them to view virtual elements within their natural field of vision. Therefore, AR glasses 500 may also include a heads-up display (i.e., AR display, see-through display) configured to display visual information at one or more lenses of the AR glasses. As shown, the heads-up display can present AR data (e.g., images, graphics, text, icons, etc.) on a portion 515 of one or more lenses of the AR glasses, allowing the user to view the AR data when looking through the lenses of the AR glasses. In this way, the AR data can overlap with the user's view of the environment. The portion 515 may include part or all of one or more lenses of the AR glasses.
[0037] AR glasses 500 may include a camera 510 (e.g., an RGB camera, a FOV camera) pointing to a camera field of view that overlaps with the user's natural field of view when wearing the glasses. In possible implementations, the AR glasses may also include a depth sensor 511 (e.g., LiDAR, structured light, time-of-flight, depth camera) pointing to a depth sensor field of view that overlaps with the user's natural field of view when wearing the glasses. Data from the depth sensor 511 and / or the FOV camera 510 can be used to measure depth within the user's (i.e., the wearer's) field of view (i.e., the region of interest). In possible implementations, the camera field of view and the depth sensor field of view can be calibrated such that the depth (i.e., range) of objects in the image from the FOV camera 510 can be determined in the depth image, where pixel values correspond to depth measured at locations corresponding to pixel positions.
[0038] The AR glasses 500 may also include eye-tracking sensors. The eye-tracking sensors may include a right-eye camera 520 and a left-eye camera 521. The right-eye camera 520 and left-eye camera 521 may be located within the lens portion of the frame, such that when the AR glasses are worn, the right FOV 522 of the right-eye camera includes the user's right eye, while the left FOV 523 of the left-eye camera includes the user's left eye.
[0039] AR glasses 500 may also include one or more microphones. These microphones may be spaced apart on the frame of the AR glasses. For example... Figure 5 As shown, the AR glasses may include a first microphone 531 and a second microphone 532. The microphones can be configured to operate together as a microphone array. The microphone array can be configured to apply sound localization to determine the direction of sound relative to the AR glasses.
[0040] AR glasses may also include a left speaker 541 and a right speaker 542 configured to transmit audio to a user. Alternatively, transmitting audio to a user may include transmitting audio to a listening device (e.g., a hearing aid, earbuds, etc.) via a wireless communication link 545. For example, AR glasses may transmit audio to a left wireless earbud 546 and a right wireless earbud 547.
[0041] The size and shape of AR glasses can affect the resources available for power and processing. Therefore, AR glasses can communicate wirelessly with other devices. Wireless communication can facilitate device-shared processing, thereby mitigating the impact of devices on the available resources of AR glasses. The process of using AR glasses for one part of the processing and using another device for a second part can be called a split architecture.
[0042] A split architecture allows for the advantageous allocation of resources based on the device's capabilities. For example, when AR glasses and a mobile phone are in a split architecture, the phone's faster processor and larger battery can be used to compute complex processes, while the AR glasses' sensors and display can be used to sense the user and display AR elements to the user.
[0043] Figure 6A possible split-architecture implementation according to this disclosure is illustrated. As shown, AR glasses 500 may include a wireless interface (i.e., a wireless module) that can be configured to wirelessly communicate with other devices (i.e., be communicatively coupled to other devices). Wireless communication can be performed via wireless communication channel 601. Wireless communication can use various wireless protocols, including (but not limited to) WiFi, Bluetooth, ultra-wideband, and mobile technologies (4G, 5G). Other devices communicating with the AR glasses may include (but are not limited to) a smartwatch 610, a mobile phone 620, a laptop 630, a cloud network 640, a tablet computer 650, etc. The split-architecture may include AR glasses that wirelessly communicate with non-mobile computing devices or mobile computing devices, and in some implementations may include AR glasses that wirelessly communicate with two or more computing devices. While these variations are within the scope of this disclosure, a specific implementation of a split-architecture involving AR glasses and a single mobile computing device (such as a mobile phone 620 (i.e., a smartphone)) is described in detail.
[0044] Return to Figure 2 A split architecture can be used to implement methods for displaying world-locked AR elements on the display of AR glasses. For example, while user measurements and measurements of the user environment can be performed by the AR device's IMU 201 and camera 202, it might be desirable to perform 6DoF tracking 210 and rendering 220 on the phone, at least because the computational complexity of such tracking and rendering might exceed the resources of the AR glasses. Instead, it might be desirable to implement warping 230 and display 240 on the AR glasses, as it might be desirable to minimize the latency between warping 230 and display 240 of AR elements on the AR glasses' display, at least because doing so can help prevent display-related artifacts (e.g., jitter, lag).
[0045] One technical problem with these functions (i.e., steps, processes, operations) of the method 200 for splitting the AR glasses 500 and the mobile phone 620 is related to wireless communication. Wireless communication can introduce large, highly variable latency. For example, the latency in a split architecture can be hundreds of milliseconds (e.g., 300 ms) compared to tens of milliseconds (e.g., 28 ms) in a non-split architecture. Large and highly variable latency can reduce estimation accuracy, resulting in artifacts in the display of AR elements. This disclosure describes a method for making the generation of world-locked AR elements on a split architecture more accurate, which may have the technical effect of minimizing artifacts associated with their display. The implementation of displaying a world-locked AR element on a display will be discussed, but it should be noted that the principles of the disclosed method can be extended to accommodate the simultaneous display of multiple world-locked AR elements.
[0046] Figure 7A method for displaying world-locked elements on a display (i.e., an AR display) of an AR device (e.g., AR glasses) according to an implementation of this disclosure is illustrated. Method 700 illustrates a split architecture, wherein a first portion of the operation of the method (i.e., computing device thread 701) is executed on a computing device (e.g., a mobile phone) (e.g., by the processor of the computing device), while a second portion of the operation of the method (i.e., AR glasses thread 702) is executed on the AR glasses. The computing device and the AR glasses are physically separated and each is configured to exchange information via a wireless communication channel 703. A flowchart of method 700 also illustrates the information (e.g., metadata) exchanged between the two devices via the wireless communication channel 703.
[0047] As shown in the figure, in the split architecture, the AR glasses are configured to collect sensor data, which can be used to determine the user's (i.e., head) position / orientation (i.e., posture). The AR glasses can be configured to collect (i.e., measure) IMU data using the AR glasses' IMU and capture image and / or range data using the AR glasses' camera.
[0048] Method 710's AR glasses thread 702 includes collecting 710 IMU / camera data. This IMU / camera data collection can be triggered by a computing device. For example, an application running on a mobile phone can request the AR glasses to begin sending IMU data streams and camera data streams from the AR glasses. Therefore, in this method, the AR glasses can transmit the collected IMU / camera data 715 from the AR glasses to the computing device. Data transmission can include data streaming or periodic measurements.
[0049] The computing device may include a high-resolution 6DoF tracker (i.e., a 6DoF estimator) configured to output position / orientation data (i.e., pose data) based on received IMU / camera data. The pose data is high-resolution, at least because it is partially based on camera data. High-resolution pose data (i.e., Hi-Res pose data) can correspond to high-resolution measurements of a user's head pose. In this disclosure, high resolution means a resolution higher than low resolution, for example, where low-resolution pose data may be based solely on IMU data. Furthermore, as used herein, "high resolution" implies higher accuracy (i.e., higher fidelity) than "low resolution." In other words, high-resolution pose data (e.g., captured by high-resolution tracking) can be more accurate than low-resolution pose data (e.g., captured by low-resolution tracking).
[0050] The computing device thread 701 of method 700 includes calculating 720 high-resolution pose data based on received IMU / camera data 715. The high-resolution pose data can be included in pose synchronization metadata 725 transmitted back to the AR glasses. Transmission can occur periodically or on request. Therefore, method 700 can include periodically transmitting pose synchronization metadata 725 from the computing device to the AR glasses. This transmission can allow the AR glasses to obtain high-resolution position / or orientation measurements without performing potentially computationally complex high-resolution pose estimations of their own. The high-resolution pose data received at the AR glasses can be based on IMU data and / or camera data captured momentarily before rendering.
[0051] As previously mentioned, the position / orientation data and estimated latency can be used to estimate the pose (i.e., the first pose). Therefore, the computational device thread 701 of method 700 may also include the latency of estimation 730 and rendering 750. Latency estimation can be performed for each repetition (i.e., loop) of method 700, and this latency estimation can vary between different loops. For example, the latency of the current loop of the method can be increased or decreased from a previous value to minimize the error in the latency of the previous loop. As will be discussed later, this error can be fed back from the AR glasses as latency feedback 735 (i.e., feedback).
[0052] The computing device thread 701 of method 700 may further include estimating a first pose (P1) of user 740 based on high-resolution pose data and estimated latency. As previously described, the first pose (P1) may be the expected head position / orientation at the end of the estimated latency period, such that the rendering latency does not introduce errors in the display of the rendered 2D texture, such as errors in the display position and / or display viewpoint of the rendered 2D texture on the display. After estimating the first pose (P1), the computing device thread 701 of method 700 may further include rendering a 2D texture 750 (T1) based on the first pose.
[0053] The rendered 2D texture (T1) and the first pose (P1) can be included in the rendering synchronization metadata 745 transferred from the computing device to the AR glasses, allowing the glasses to receive the rendered texture without performing the computations associated with rendering 750. Therefore, method 700 also includes transferring the rendering synchronization metadata 745 from the computing device to the AR glasses. The transfer can be triggered by a new pose and / or the rendered 2D texture (T1). Method 700 includes receiving a 2D texture of an AR element at the AR device (AR glasses in this example), the 2D texture being rendered by a computing device coupled to the AR device, wherein the computing device and the AR device are physically separated.
[0054] AR glasses can be used to perform warping because, as previously mentioned, it is a relatively simple operation compared to rendering, and because it is closely related to the display performed by the AR glasses. As mentioned above, warping transformation (i.e., warping) requires estimating a second pose (P2) of the user (e.g., head) after rendering. AR glasses do not include high-resolution 6DoF trackers because their computational burden can be high. Instead, AR glasses can include low-resolution trackers for measuring the user's (e.g., head) position / orientation (i.e., pose). Low-resolution 6DoF trackers can be configured to compute low-resolution (Lo-Res) pose data from IMU data collected by the AR glasses' IMU. By not using camera data to compute 6DoF data, low-resolution 6DoF trackers can save resources by eliminating image processing associated with pose estimation. The result is that the resolution of the pose data is lower than the resolution estimated using camera data. Therefore, the AR glasses thread 702 of method 700 includes computing 755 low-resolution pose data (i.e., Lo-Res pose data) based on the received IMU data collected by the AR glasses.
[0055] Low-resolution pose data can be used to correct high-resolution pose data transmitted from a computing device. When high-resolution pose data is received at the AR glasses, it may be inaccurate (i.e., outdated). Inaccuracy may be due to latency associated with communication on wireless channel 703 and / or latency associated with rendering. For example, at the end of rendering, the computing device may transmit the high-resolution pose data used for rendering to the AR glasses. Low-resolution pose data includes accurate (i.e., up-to-date) information about the head's position and orientation. For example, low-resolution pose data may be collected after rendering is complete. Therefore, low-resolution pose data can include up-to-date pose information about the user, which can correct inaccuracies in (older) high-resolution pose data. In other words, high-resolution pose data may be based on IMU data and camera data captured at a first time before rendering, while the IMU data of low-resolution pose data may be captured at a second time after rendering. Therefore, correcting high-resolution pose data may involve modifying the high-resolution pose data captured at the first time using IMU data captured at the second time to generate corrected high-resolution pose data that corresponds to the user's pose at the second time.
[0056] Method 700's AR glasses thread 702 includes a computationally-based correction of low-resolution pose data to high-resolution pose data 760. The result is corrected high-resolution pose data that represents the user's head pose at a later time closer to the display time. In other words, the high-resolution 6DoF data may correspond to the pose at a first time (t1) (i.e., before rendering), the low-resolution 6DoF data may correspond to the pose at a second time (t2) (i.e., after rendering), and the corrected high-resolution 6DoF data may be high-resolution 6DoF data adapted from the first time (t1) to the second time (t2).
[0057] The AR glasses thread 702 of method 700 may further include estimating 765 a second pose (P2) of the user based on corrected high-resolution pose data. As previously described, the first pose (P1) may be the expected position / orientation of the head at the end of the estimated latency period, while the second pose (P2) may be the actual position / orientation of the head at the end of the actual latency period. The AR glasses thread 702 of method 700 may further include comparing 770 the first pose (P1) and the second pose (P2) to evaluate the estimate of the latency period. For example, if the first pose (P1) and the second pose (P2) match, the estimated latency period matches the actual latency period, and there is no latency error in the estimate. However, if the first pose (P1) and the second pose (P2) do not match, there is an error in the estimated latency period. This error can be corrected by latency feedback 735 transmitted to the computing device (e.g., during the next rendering cycle). Latency feedback 735 may correspond to the error between the estimated latency calculated based on the comparison between the first pose (P1) and the second pose (P2) and the actual latency. Latency feedback can be used to reduce or increase the estimated latency of subsequent rendering cycles. The amount of increase or decrease in estimated latency can be determined based on an algorithm configured to minimize the error between the estimated latency and the actual latency.
[0058] After estimating the second pose, the AR glasses thread 702 of method 700 may further include warping the rendered 2D texture (t1) 775. For example, warping 775 may include calculating a 232 warp transformation matrix from the first pose (P1) and the second pose (P2) (i.e., based on a comparison between P1 and P2). Warping 230 may also include applying (e.g., multiplying) the rendered 2D texture (T1) from the computing device to the warp transformation matrix (W) to transform the rendered 2D texture (T1) into a registered 2D texture (T2). The registered 2D texture (T2) corresponds to the latest (and more accurate) pose information captured after rendering.
[0059] After generating the registered 2D texture (T2), the AR glasses thread 702 of method 700 may also include displaying the registered 2D texture (T2) at 780. The registered 2D texture (T2) may include information that helps determine the position of AR elements displayed on the AR glasses' display. Although rendering and warping are independent operations, the metadata exchanged in the split architecture still helps in enabling the operation.
[0060] The exchanged metadata may include pose synchronization (i.e., sync) metadata 725. Pose synchronization metadata may include an estimate of the pose (i.e., high-resolution 6DoF data) associated with a timestamp. Pose synchronization metadata 725 may also include estimates to help correct 6DoF data at the AR glasses. For example, pose synchronization metadata may include estimated device velocity, estimated IMU bias, estimated IMU intrinsic parameters, and estimated camera extrinsic parameters. Pose synchronization metadata may be sent periodically (e.g., at 10 Hz).
[0061] The exchanged metadata may include rendering synchronization (i.e., sync) metadata 745. Rendering synchronization metadata may include a pose timestamp for rendering, a pose for rendering, a demo timestamp, and the rendered frame (i.e., a 2D texture). Rendering synchronization metadata may be sent periodically (e.g., 20 Hz).
[0062] The rendering process on the computing device can be repeated at a first rate (i.e., looped), while the warping process on the AR glasses can be repeated at a second rate. The first rate may not be equal to the second rate. In other words, these processes may be asynchronous.
[0063] Figure 7 The illustrated process can be performed by an augmented reality system with a split architecture comprising a computing device and AR glasses. The computing device and AR glasses can each have one or more processors and non-transitory computer-readable storage so as to be configured to perform the processes associated with rendering and warping described so far. The computing device and AR glasses can perform certain operations in parallel (i.e., simultaneously), such as... Figure 7 As shown. The metadata described so far can help explain and compensate for timing differences between (asynchronous) operations performed by each device.
[0064] In this system, the processor of the computing device can be configured by software instructions to receive IMU / camera data and calculate high-resolution pose data based on the received IMU / camera data. The processor of the computing device can also be configured by software instructions to receive latency feedback and estimate the rendering latency based on the latency feedback. The processor of the computing device can also be configured by software instructions to estimate a first pose (P1) based on the latency and the high-resolution pose data, and render a 2D texture (T1) based on the first pose (P1).
[0065] In this system, the AR glasses' processor can be configured by software instructions to calculate corrected high-resolution pose data based on received IMU data and received high-resolution pose data, and estimate a second pose (P2) based on the corrected high-resolution pose data. The AR glasses' processor can also be configured by software instructions to compare the second pose (P2) with the first pose (P1), and distort the 2D texture of the AR elements received from the computing device based on the comparison. The AR glasses' processor can also be configured by software instructions to transmit the distorted 2D texture of the AR elements to the AR display of the AR glasses.
[0066] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as would be understood by one of ordinary skill in the art. Methods and materials similar to or equivalent to those described herein may be used to practice or test this disclosure. As used in this specification and the appended claims, the singular forms “a” and “the” include the plural indicating objects, unless the context clearly indicates otherwise. The term “comprising” and its variations are used synonymously with the term “including” and its variations and are open-ended, non-limiting terms. The terms “optional” or “optionally” as used herein mean that a feature, event, or situation subsequently described may or may not occur, and mean that the description includes instances where said feature, event, or situation occurs and instances where said feature, event, or situation does not occur. A range may be expressed herein as from “about” one specific value and / or to “about” another specific value. When such a range is expressed, one aspect includes from one specific value and / or to another specific value. Similarly, when a value is expressed as an approximation, the specific value will be understood to form another aspect by using the antecedent “about”. It should be further understood that each endpoint of the range is valid relative to the other endpoint, and is independent of the other endpoint.
[0067] While certain features of the described implementations have been described herein, those skilled in the art will now conceive of numerous modifications, substitutions, alterations, and equivalents. Therefore, it should be understood that the appended claims are intended to cover all such modifications and alterations falling within the scope of the implementations. It should be understood that they are presented by way of example only and not limitation, and various changes in form and detail are possible. Any part of the apparatus and / or method described herein may be combined in any combination, except for mutually exclusive combinations. The implementations described herein may include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.
[0068] As used in this specification, the singular form may include the plural form unless the context clearly indicates otherwise. Spatial relative terms (e.g., above, above, upper, below, under, lower, etc.) are intended to include different orientations of the device in use or operation other than those depicted in the figures. In some implementations, the relative terms above and below may respectively include vertically above and vertically below. In some implementations, the term "adjacent" may include laterally adjacent or horizontally adjacent.
Claims
1. A method for displaying AR elements on an augmented reality (AR) display: An initial 2D image of the AR element is received at the AR device, and the initial 2D image is rendered at a computing device coupled to the AR device in a split architecture, wherein the computing device and the AR device are physically separated. The initial 2D image of the AR element is distorted at the AR device, wherein, The distortion includes: Receive a first pose from the computing device, the first pose being based on an estimate of the rendering latency; Estimate the second posture; The distortion transformation is calculated based on the comparison between the first pose and the second pose; The distortion transformation is applied to the initial 2D image of the AR element to generate a distorted 2D image of the AR element; and Triggering the display of the distorted 2D image of the AR element on the AR display of the AR device; and Feedback is sent to the computing device to update the estimated latency of the rendering, wherein the feedback is based on a comparison of the first pose and the second pose.
2. The method of claim 1, wherein the initial 2D image of the AR element is received via a wireless communication channel between the AR device and the computing device.
3. The method according to claim 1, wherein the AR device is AR glasses.
4. The method according to claim 1, wherein the computing device is a mobile phone.
5. The method according to claim 1, further comprising: Collect sensor data at the AR device; as well as The sensor data is transmitted from the AR device to the computing device via a wireless communication channel.
6. The method of claim 5, wherein the sensor data includes: Inertial Measurement Unit (IMU) data and camera data.
7. The method according to claim 5, further comprising: Determine the time delay feedback at the AR device; The delay feedback is transmitted from the AR device to the computing device via the wireless communication channel; as well as The rendering latency is estimated at the computing device based on the latency feedback, and the initial 2D image of the AR element rendered at the computing device is based on the first pose determined using the sensor data and the latency feedback from the AR device.
8. The method of claim 7, wherein determining the delay feedback at the AR device comprises: The first gesture is received at the AR device from the computing device; Estimate the second pose at the AR device; Compare the first pose and the second pose at the AR device; as well as The time delay feedback is determined at the AR device based on the comparison.
9. The method of claim 1, wherein calculating the warp transformation based on a comparison of the first pose and the second pose comprises: Compare the first posture and the second posture; as well as The distortion transformation matrix is calculated based on the comparison.
10. The method of claim 9, wherein the distortion at the AR device comprises: Receive the initial 2D image of the AR element from the computing device; as well as The initial 2D image of the AR element is applied to the warp transformation matrix to generate a registered 2D image of the AR element.
11. The method of claim 1, wherein estimating the second pose at the AR device comprises: Inertial measurement unit (IMU) data and camera data are collected at the AR device; The IMU data and the camera data are transmitted from the AR device to the computing device; High-resolution pose data is received from the computing device, the high-resolution pose data being calculated at the computing device based on the IMU data and the camera data; Low-resolution pose data is calculated at the AR device based on the IMU data; The low-resolution pose data is used at the AR device to correct the high-resolution pose data from the computing device to calculate the corrected high-resolution pose data; as well as The second pose is estimated based on the corrected high-resolution pose data.
12. The method of claim 10, wherein the registered 2D image of the AR element is world-locked to the location in the user environment on the AR display.
13. An augmented reality (AR) glasses, comprising: An inertial measurement unit (IMU) configured to collect IMU data; A camera, configured to capture camera data; A wireless interface configured to transmit first information to and receive first information from a computing device via a wireless communication channel; An AR display configured to show second information to the user of the AR glasses; as well as Processor, the processor being configured by software to: The IMU data and the camera data are transmitted to the computing device, which is configured to: Calculate initial high-resolution pose data based on the IMU data and the camera data; The first pose is estimated based on the initial high-resolution pose data and the estimated latency corresponding to the rendering. as well as Render a 2D image of the AR element based on the first pose; Receive the initial high-resolution pose data, the first pose, and the 2D image of the AR element from the computing device; Calculate the corrected high-resolution pose data based on the IMU data and the initial high-resolution pose data; Estimate the second pose based on the corrected high-resolution pose data; The 2D image of the AR element is distorted based on a comparison between the second pose and the first pose; Trigger the display of a distorted 2D image of the AR element on the AR display; as well as Feedback is sent to the computing device to update the estimated latency of the rendering, wherein the feedback is based on a comparison of the first pose and the second pose.
14. The AR glasses of claim 13, wherein the distorted 2D image of the AR element is locked onto the AR display of the AR glasses.
15. The AR glasses of claim 14, wherein, as part of a navigation application running on the AR glasses, the distorted 2D image of the AR element is locked to a location by the world.
16. The AR glasses according to any one of claims 13 to 15, wherein, in order to calculate the corrected high-resolution pose data based on the IMU data and the initial high-resolution pose data, the AR glasses are configured to: The initial high-resolution pose data is received from the computing device, the initial high-resolution pose data being based on the IMU data and the camera data captured in a first-time event prior to rendering; The IMU data was collected immediately after rendering. as well as The initial high-resolution pose data captured at the first time is modified using the IMU data captured at the second time to generate the corrected high-resolution pose data corresponding to the pose at the second time.
17. A split-architecture augmented reality (AR) system, comprising: Computing device, the computing device being configured to: Calculate the initial high-resolution pose data; Estimate the rendering latency; The first pose is estimated based on the rendering latency and the initial high-resolution pose data; as well as Render a 2D image of the AR element based on the first pose; as well as AR glasses, which are communicatively coupled to the computing device, are configured to: Collect inertial measurement unit (IMU) data and camera data; Calculate the corrected high-resolution pose data based on the IMU data and the initial high-resolution pose data; Estimate the second pose based on the corrected high-resolution pose data; Compare the second posture with the first posture; The 2D image of the AR element is distorted based on the comparison between the second pose and the first pose; as well as The distorted 2D image of the AR element is displayed on the AR display of the AR glasses. In order to estimate the rendering latency, the computing device is configured to: Receive feedback from the AR glasses, the feedback being based on the comparison between the second pose and the first pose; as well as The estimated delay is updated based on the feedback.
18. The system of claim 17, wherein the computing device and the AR glasses are communicatively coupled via a wireless communication channel.
19. The system according to any one of claims 17 to 18, wherein the computing device is a mobile phone or a tablet computer.
Citation Information
Patent Citations
Continuous time warp and binocular time warp for virtual and augmented reality display systems and methods
CN109863538A
Optimized Display Image Rendering
US20210118357A1