3D Depth Bending Correction

By estimating and correcting the pitch-roll and yaw deviations of flexible devices using the VIO stereo matching method, the depth sensing error caused by the bending of flexible devices is solved, improving the display accuracy of virtual content and reducing the consumption of computing resources.

CN117321634BActive Publication Date: 2026-04-03SNAP INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

The bending of flexible devices causes depth sensing errors in stereoscopic vision tracking systems, which are difficult to correct effectively with existing technologies, affecting the accurate display of virtual content.

Method used

By using a visual inertial odometry (VIO) stereo matching method, pitch-roll and yaw deviations of a flexible stereo depth device are estimated and compensated by projecting stereo VIO features onto a modified coordinate system.

Benefits of technology

It reduces depth sensing errors in flexible devices, improves the display accuracy of virtual content, and reduces the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117321634B_ABST
    Figure CN117321634B_ABST
Patent Text Reader

Abstract

A method for correcting the bending of a flexible device. In one aspect, the method includes: accessing feature data of a first stereo frame generated by a stereo optical sensor of the flexible device, the feature data being generated based on a visual inertial odometry (VIO) system of the flexible device; accessing depth map data of the first stereo frame, the depth map data being generated based on a depth map system of the flexible device; estimating pitch-roll and yaw errors based on the feature data and depth map data of the first stereo frame; and generating a second stereo frame after the first stereo frame, the second stereo frame being based on the pitch-roll and yaw errors of the first stereo frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority requirements

[0002] This application claims priority to U.S. Patent Application Serial No. 17 / 480,405, filed September 21, 2021, and U.S. Provisional Patent Application Serial No. 63 / 188,815, filed May 14, 2021, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] The subject matter disclosed herein generally relates to visual tracking systems. Specifically, this disclosure presents systems and methods for mitigating bending effects in visual-inertial tracking systems. Background Technology

[0004] Augmented reality (AR) devices allow users to observe a scene while seeing related virtual content that aligns with items, images, objects, or the environment within the device's field of view. Virtual reality (VR) devices offer a more immersive experience than AR devices. VR devices obscure the user's field of view with virtual content displayed based on the VR device's positioning and orientation.

[0005] Both AR and VR devices rely on motion tracking systems that track the device's posture (e.g., orientation, orientation, position). Motion tracking systems are typically factory-calibrated (based on predefined relative positioning between the camera and other sensors) to accurately display virtual content at a desired location relative to its environment. However, when a user wears an AR / VR device, the factory calibration parameters can drift over time due to mechanical stress and temperature variations within the device. Attached Figure Description

[0006] To facilitate the identification of any particular element or action being discussed, one or more of the highest-order digits in the reference numerals indicate the drawing number in which the element was first introduced.

[0007] Figure 1 This is a block diagram illustrating an environment for operating an AR / VR display device according to an example implementation.

[0008] Figure 2 This is a block diagram illustrating an AR / VR display device according to an example implementation.

[0009] Figure 3 This is a block diagram illustrating a visual tracking system according to an example implementation.

[0010] Figure 4 This is a block diagram illustrating a bending correction module according to an example implementation.

[0011] Figure 5This is a block diagram illustrating a pitch-roll bending module according to an example implementation.

[0012] Figure 6 This is a block diagram illustrating a yaw bending module according to an example implementation.

[0013] Figure 7 This is a block diagram illustrating a corrected depth frame according to one embodiment.

[0014] Figure 8 This is a flowchart illustrating a method for adjusting a coordinate system according to an example implementation.

[0015] Figure 9 This is a flowchart illustrating a method for adjusting yaw bias according to an example implementation.

[0016] Figure 10 This is a flowchart illustrating a method for updating the total yaw curvature estimate according to an example implementation.

[0017] Figure 11 The diagram illustrates misalignment errors caused by bending of a flexible device according to one embodiment.

[0018] Figure 12 A pitch-roll misalignment according to one embodiment is shown.

[0019] Figure 13 The projection of a two-dimensional image according to one embodiment shows a depth misalignment.

[0020] Figure 14 A network environment in which a head-mounted wearable device can be implemented is shown according to an example implementation.

[0021] Figure 15 This is a block diagram illustrating a software architecture in which the present disclosure can be implemented according to an example embodiment.

[0022] Figure 16 It is a schematic representation of a machine in the form of a computer system according to an example implementation, in which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein. Detailed Implementation

[0023] The following description illustrates systems, methods, techniques, sequences of instructions, and computer program products that demonstrate exemplary embodiments of the subject matter. In this description, numerous specific details are set forth for illustrative purposes to provide an understanding of various embodiments of the subject matter. However, it will be apparent to those skilled in the art that embodiments of the subject matter can be practiced without some or more of these specific details. The examples merely represent possible variations. Unless explicitly stated otherwise, structures (e.g., structural components, such as modules) are optional and can be combined or subdivided, and operations (e.g., in processes, algorithms, or other functions) can vary in sequence or be combined or subdivided.

[0024] The term "augmented reality" (AR) is used in this article to refer to interactive experiences in real-world environments where physical objects existing in the real world are "enhanced" or strengthened by computer-generated digital content (also known as virtual or synthetic content). AR can also refer to systems that enable the combination of the real and virtual worlds, real-time interaction, and 3D registration of virtual and real objects. In AR systems, users perceive virtual content that appears to be attached to or interacts with real-world physical objects.

[0025] The term "virtual reality" (VR) is used in this article to refer to a simulated experience of a virtual world environment that is completely different from the real world environment. Computer-generated digital content is displayed in the virtual world environment. VR also refers to systems that allow users to be fully immersed in a virtual world environment and interact with virtual objects presented within that environment.

[0026] The term "AR application" is used herein to refer to a computer application that enables AR experiences. The term "VR application" is used herein to refer to a computer application that enables VR experiences. The term "AR / VR application" refers to a computer application that enables either AR or a combination of AR and VR experiences.

[0027] The term "visual tracking system" is used herein to refer to a computer-operated application or system that enables the system to track visual features identified in images captured by one or more camera devices of the visual tracking system. The visual tracking system builds a model of the real-world environment based on the tracked visual features. Non-limiting examples of visual tracking systems include visual simultaneous localization and mapping (VSLAM) systems and visual inertial odometry (VIO) systems. VSLAM can be used to construct a target from an environment or scene based on one or more camera devices of the visual tracking system. VIO (also known as a visual inertial tracking system) determines the latest posture (e.g., localization and orientation) of the device based on data acquired from multiple sensors (e.g., optical sensors, inertial sensors) of the device.

[0028] The term "Inertial Measurement Unit" (IMU) is used herein to refer to a device capable of reporting the inertial state of a moving subject, including the subject's acceleration, velocity, orientation, and position. An IMU tracks the subject's motion by integrating acceleration and angular velocity measurements taken by the IMU. An IMU can also refer to a combination of accelerometers and gyroscopes, which can accordingly determine and quantify linear acceleration and angular velocity. Values ​​obtained from the IMU's gyroscopes can be processed to obtain the IMU's pitch, roll, and yaw, thereby obtaining the pitch, roll, and yaw of the subject associated with the IMU. Signals from the IMU's accelerometers can also be processed to obtain the IMU's velocity and displacement.

[0029] The term "flexible device" is used herein to refer to a device that can bend without breaking. Non-limiting examples of flexible devices include: head-mounted devices such as glasses, flexible display devices such as AR / VR glasses, or any other wearable device that can bend without breaking to fit a part of a user's body.

[0030] Both AR and VR applications allow users to access information, such as in the form of virtual content presented on the display of an AR / VR display device (also known as a display device, flexible device, or flexible display device). The presentation of virtual content can be based on the positioning of the display device relative to a physical object or relative to a frame of reference (outside the display device), ensuring that the virtual content appears correctly on the display. For AR, the virtual content appears aligned with the physical object perceived by the user and the camera mechanism of the AR display device. The virtual content appears attached to the physical world (e.g., a physical object of interest). To do this, the AR display device detects the physical object and tracks the pose of the AR display device relative to the physical object. The pose identifies the positioning and orientation of the display device relative to a frame of reference or relative to another object. For VR, the virtual object appears at a location based on the pose of the VR display device. Therefore, the virtual content is refreshed based on the device's latest pose. A visual tracking system at the display device determines the pose of the display device.

[0031] Flexible devices, including vision tracking systems, can operate on stereo vision using two cameras mounted on the flexible device. For example, one camera is mounted on the left temple of the flexible device's frame, and the other on the right temple. However, the flexible device can be bent to accommodate different user head sizes. Therefore, bending can cause undesirable offsets in the stereo image (away from the factory-calibrated configuration), which can lead to errors when using the offset stereo image for depth sensing.

[0032] One approach to compensate for warp is to match features from each frame on both the left and right images and optimize a symmetric model (where the camera is rotated symmetrically, leaving only three angles to be solved). However, such per-frame optimization incurs additional computation time. Furthermore, per-frame optimization is better suited for outdoor environments where distant features can be used for anchoring. Conversely, for indoor settings, objects are typically closer to flexible equipment, so yaw estimation may be unstable.

[0033] This application describes a method for estimating and compensating for the curvature of flexible stereo-to-depth devices by using VIO stereo matching to validate corrections. In one example, the method includes validating the corrections by “projecting” stereo VIO features from the “native” camera device to the modified coordinate system, and triggering curvature compensation processing when the VIO matches are not located on the same raster line.

[0034] In one example implementation, a method for correcting the curvature of a flexible device is described. In one aspect, the method includes: accessing feature data of a first stereo frame generated by a stereo optical sensor of the flexible device, wherein the feature data is generated based on the VIO system of the flexible device; accessing depth map data of the first stereo frame, wherein the depth map data is generated based on the depth map system of the flexible device; estimating pitch-roll bias and yaw bias based on the feature data and depth map data of the first stereo frame; and generating a second stereo frame after the first stereo frame, the second stereo frame being based on the pitch-roll bias and yaw bias of the first stereo frame.

[0035] As a result, one or more methods described herein help address the technical problem of inaccurate depth sensing from stereo extraction using flexible devices. In other words, the bending of the flexible device leads to errors in depth sensing. The methods described herein provide improvements to the operation of computing device functionality by correcting the depth map from a bent flexible stereo depth device. Therefore, one or more methods described herein can avoid the need for certain effort or computational resources. Examples of such computational resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.

[0036] Figure 1 This is a network diagram illustrating an environment 100 suitable for operating an AR / VR display device 106 according to some example embodiments. Environment 100 includes a user 102, the AR / VR display device 106, and physical objects 104. The user 102 operates the AR / VR display device 106. The user 102 can be a human user (e.g., a human), a machine user (e.g., a computer configured by software programs to interact with the AR / VR display device 106), or any suitable combination thereof (e.g., a machine-assisted human or a machine supervised by a human). The user 102 is associated with the AR / VR display device 106.

[0037] AR / VR display device 106 includes a flexible device. In one example, the flexible device includes a computing device with a display, such as a smartphone, tablet, or wearable computing device (e.g., a watch or glasses). The computing device may be handheld or removably mounted to the head of user 102. In one example, the display includes a screen that displays images captured by a camera device of AR / VR display device 106. In another example, the device's display may be transparent, for example, within the lenses of wearable computing glasses. In other examples, the display may be opaque, partially transparent, or partially opaque. In still other examples, the display may be worn by user 102 to cover user 102's field of vision.

[0038] AR / VR display device 106 includes an AR application that generates virtual content based on images detected by a camera device of AR / VR display device 106. For example, user 102 can instruct the camera device of AR / VR display device 106 to capture an image of a physical object 104. The AR application generates virtual content corresponding to the identified object (e.g., physical object 104) in the image and presents the virtual content on the display of AR / VR display device 106.

[0039] AR / VR display device 106 includes a visual tracking system 108. The visual tracking system 108 uses, for example, optical sensors (e.g., a 3D camera with depth capabilities, an image camera), inertial sensors (e.g., a gyroscope, an accelerometer), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors to track the posture (e.g., positioning and orientation) of the AR / VR display device 106 relative to the real-world environment 110. The visual tracking system 108 may include a VIO system. In one example, the AR / VR display device 106 displays virtual content based on the posture of the AR / VR display device 106 relative to the real-world environment 110 and / or physical object 104.

[0040] Figure 1 Any of the machines, databases, or devices shown can be implemented in a general-purpose computer that has been software-modified (e.g., configured or programmed) to perform one or more of the functions described herein for that machine, database, or device. For example, see below. Figures 8 to 10 This discussion focuses on computer systems capable of implementing any one or more of the methods described herein. As used herein, "database" is a data storage resource and can store data structured as text files, tables, spreadsheets, relational databases (e.g., object-relational databases), triplet storage, hierarchical data storage, or any suitable combination thereof. Furthermore, Figure 1 Any two or more of the machines, databases or devices shown may be combined into a single machine, and the functionality described herein for any single machine, database or device may be subdivided among multiple machines, databases or devices.

[0041] AR / VR display device 106 can operate via a computer network. The computer network can be any network that enables communication between or within machines, databases, and devices. Therefore, the computer network can be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. The computer network may include one or more components constituting a private network, a public network (e.g., the Internet), or any suitable combination thereof.

[0042] Figure 2 This is a block diagram illustrating the modules (e.g., components) of an AR / VR display device 106 according to some example embodiments. The AR / VR display device 106 includes a sensor 202, a display 204, a processor 208, and a storage device 206. Examples of the AR / VR display device 106 include wearable computing devices, desktop computers, in-vehicle computers, tablet computers, navigation devices, portable media devices, or smartphones.

[0043] Sensor 202 includes, for example, optical sensors 212 (e.g., imaging devices such as color imaging devices, thermal imaging devices, depth sensors, and one or more grayscale, global shutter tracking imaging devices) and inertial sensors 214 (e.g., gyroscopes, accelerometers). Other examples of sensor 202 include proximity sensors or position sensors (e.g., near-field communication, GPS, Bluetooth, Wi-Fi), audio sensors (e.g., microphones), or any suitable combination thereof. Note that sensor 202 described herein is for illustrative purposes and therefore sensor 202 is not limited to the sensors described above.

[0044] Display 204 includes a screen or monitor configured to display images generated by processor 208. In one example implementation, display 204 may be transparent or semi-opaque, allowing user 102 to view through it (in an AR use case). In another example implementation, display 204 covers user 102's eyes and obstructs user 102's entire field of vision (in a VR use case). In yet another example, display 204 includes a touchscreen display configured to receive user input via touch on a touchscreen display.

[0045] Processor 208 includes AR / VR application 210, visual tracking system 108, and bending correction module 216. AR / VR application 210 uses computer vision to detect and identify the physical environment or physical object 104. AR / VR application 210 retrieves virtual objects (e.g., 3D object models) based on the identified physical object 104 or physical environment. AR / VR application 210 renders the virtual objects in display 204. For AR applications, AR / VR application 210 includes a local rendering engine that generates visualizations of the virtual objects overlaid (e.g., superimposed on or otherwise displayed in cooperation with) an image of the physical object 104 captured by optical sensor 212. The visualization of the virtual objects can be manipulated by adjusting the positioning of the physical object 104 relative to optical sensor 212 (e.g., its physical location, orientation, or both). Similarly, the visualization of the virtual objects can be manipulated by adjusting the pose of AR / VR display device 106 relative to physical object 104. For VR applications, AR / VR application 210 displays virtual objects on display 204 at a position (in display 204) determined based on the pose of AR / VR display device 106.

[0046] The visual tracking system 108 estimates the pose of the AR / VR display device 106. For example, the visual tracking system 108 uses image data and corresponding inertial data from the optical sensor 212 and the inertial sensor 214 to track the position and pose of the AR / VR display device 106 relative to a reference frame (e.g., the real-world environment 110). In one example, the visual tracking system 108 includes the VIO system previously described above.

[0047] The bending correction module 216 accesses VIO data from the vision tracking system 108 to estimate pitch-roll bias and yaw bending bias. The bending correction module 216 corrects for these biases to mitigate any depth sensing errors resulting from bending. In one example implementation, the bending correction module 216 estimates and compensates for the bending of the flexible stereo depth device by validating the correction using stereo matching of the VIO. The bending correction module 216 validates the correction by "projecting" stereo VIO features from the "native" camera device to the corrected coordinate system. When the VIO match is not on the same grid line, the bending correction module 216 triggers bending compensation processing. See below for further details. Figure 4 A sample component of the bending correction module 216 is described in more detail.

[0048] Storage device 206 stores virtual object content 218 and bias data 220. Bias data 220 includes estimated pitch-roll and yaw curve bias values ​​for AR / VR display device 106. Virtual object content 218 includes a database of, for example, visual references (e.g., images) and corresponding experiences (e.g., 3D virtual objects, interactive features of 3D virtual objects).

[0049] Any one or more modules described herein may be implemented using hardware (e.g., a machine's processor) or a combination of hardware and software. For example, any module described herein may configure a processor to perform the operations described herein for that module. Furthermore, any two or more of these modules may be combined into a single module, and the functionality described herein for a single module may be subdivided among multiple modules. Additionally, modules described herein as being implemented within a single machine, database, or device may be distributed across multiple machines, databases, or devices, depending on various example implementations.

[0050] Figure 3 A visual tracking system 108 according to an example implementation is shown. The visual tracking system 108 includes, for example, a VIO system 302 and a depth map system 304. The VIO system 302 accesses inertial sensor data from an inertial sensor 214 and images from an optical sensor 212.

[0051] VIO system 302 determines the pose (e.g., position, orientation, orientation) of AR / VR display device 106 relative to a reference frame (e.g., real-world environment 110). In one example implementation, VIO system 302 estimates the pose of AR / VR display device 106 based on a 3D map of feature points from images captured by optical sensor 212 and inertial sensor data captured by inertial sensor 214.

[0052] Depth mapping system 304 accesses image data from optical sensor 212 and generates a depth map based on VIO data (e.g., feature point depth) from VIO system 302. For example, depth mapping system 304 generates a depth map based on the depth of matching features between the left image (generated by the left camera device) and the right image (generated by the right camera device). In another example, depth mapping system 304 is based on triangulation of element disparities in a stereo image.

[0053] Figure 4 This is a block diagram illustrating a bend correction module 216 according to an example embodiment. The bend correction module 216 includes a pitch-roll bend module 402, a yaw bend module 404, and a mitigation module 406.

[0054] The pitch-roll bending module 402 determines whether the mitigation module 406 should correct the bending that is causing pitch or roll deviations in the flexible device. In one example, the pitch-roll bending module 402 projects stereo VIO features from the optical sensor 212 onto a modified coordinate system. When the VIO feature matches are not on the same grid line, the pitch-roll bending module 402 triggers the mitigation module 406.

[0055] The yaw bending module 404 determines whether the mitigation module 406 should correct the bending that is causing yaw deviation in the flexible device. In one example, the yaw bending module 404 estimates the yaw deviation by accessing a 3D landmark with a wide baseline determined by VIO to obtain a time-consistent result.

[0056] Mitigation module 406 minimizes pitch-roll and yaw deviations based on estimates provided by pitch-roll curve module 402 and yaw curve module 404. For example, mitigation module 406 can minimize the yaw deviation between VIO depth and stereo depth algorithm results by correcting the yaw estimate. The corrected configuration is then transmitted to AR / VR application 210 for displaying content based on the corrected configuration.

[0057] Figure 5 This is a block diagram illustrating a pitch-roll bending module according to an example embodiment. The pitch-roll bending module 402 includes a stereo frame access module 502, a feature projection module 504, a feature alignment module 506, and a pitch-roll deviation estimation module 508.

[0058] The stereo frame access module 502 accesses a first image from a first camera device (e.g., on the left) and a second image from a second camera device (e.g., on the right) of the optical sensor 212. In one example, the stereo frame access module 502 determines the stereo VIO features of the first stereo frame.

[0059] The feature projection module 504 accesses stereo VIO data (e.g., 3D points, pose) from the VIO system 302. In one example, the feature projection module 504 projects stereo VIO features from the optical sensor 212 onto a modified two-dimensional coordinate system.

[0060] Feature alignment module 506 determines whether corresponding stereo VIO features (from the left and right images) lie on the same grid line. For example, feature alignment module 506 determines that a VIO feature from the left image is not aligned on the same grid line as the corresponding VIO feature from the right image. In this case, feature alignment module 506 triggers compensation processing at pitch-roll deviation estimation module 508.

[0061] The pitch-roll bias estimation module 508 estimates the pitch-roll bias based on the misalignment of the grid lines between the left and right VIO features relative to one of the right or left images. The module calculates the pitch-roll bias based on the misalignment in the stereo VIO features projected in a two-dimensional coordinate system. The estimated pitch-roll bias is then provided to the mitigation module 406 for correction.

[0062] Figure 6 This is a block diagram illustrating a yaw curve module 404 according to an example implementation. The yaw curve module 404 includes a disparity map module 602, a yaw deviation estimation module 608, and a yaw update module 610.

[0063] The disparity map module 602 accesses VIO data from the VIO system 302 and depth data from the depth map system 304. The disparity map module 602 identifies the disparity or misalignment (for feature points) between the VIO data and the depth data. In one example implementation, the disparity map module 602 includes a 3D landmark projection module 604 and a landmark disparity calculation module 606.

[0064] The 3D landmark projection module 604 projects 3D feature points from VIO data onto a two-dimensional coordinate system (e.g., depth data for each feature point). For example, the 3D landmark projection module 604 uses VIO data from the VIO system 302 to identify 3D landmarks in the first stereo frame. The 3D landmark projection module 604 projects the 3D landmarks onto a two-dimensional disparity map. The two-dimensional disparity map indicates the 2D location of the landmarks and their corresponding depth values.

[0065] The landmark disparity calculation module 606 determines the depth misalignment between the stereo depth from the depth data and the VIO depth from the VIO data for each feature point or landmark. In one example, the landmark disparity calculation module 606 calculates the depth deviation value for each landmark in the first stereo frame.

[0066] The yaw deviation estimation module 608 estimates the yaw deviation based on the depth misalignment of feature points. In one example, the yaw deviation estimation module 608 calculates the yaw deviation value based on the depth deviation value of each landmark.

[0067] The yaw update module 610 corrects the yaw deviation based on the depth misalignment and provides the updated configuration settings to the depth map system 304. The updated configuration settings include the yaw deviation estimate. In one example, the yaw update module 610 updates the corrected map based on the total yaw curvature estimate and requests the depth map system 304 to calculate the depth map based on the corrected map.

[0068] The following pseudocode illustrates an example implementation of the operation of the disparity map module 602:

[0069] 1. Call [FeatureCount, RotationMean, RotationVariance] = EstimateYawRotationBias()

[0070] 2. If FeatureCount is greater than the threshold

[0071] 3. Call UpdateYaw(RotationMean, RotationVariance)

[0072] The following pseudocode illustrates an example implementation of the operation of the yaw deviation estimation module 608:

[0073] Yaw estimation process EstimateYawRotationBias():

[0074] Set featureCount = 0

[0075] Let rotationMean = 0

[0076] Let rotationMeanSquared = 0

[0077] For each feature from VIO:

[0078] If the feature is invalid, proceed to the next feature.

[0079] #Match it with an existing disparity map

[0080] [X,Y]=projectFeatureToDepthSystemCoordinates()

[0081] Assign sampledDisparityVal to the points X, Y.

[0082] #Convert depth from VIO to equivalent parallax

[0083] Let f be the focal length of the camera device.

[0084] Let B be the baseline between the two camera devices.

[0085] Let vioDepth be the depth value of the VIO feature.

[0086] Distribute vioDisparity = f * B / vioDepth

[0087] # Calculate the equivalent rotation (achieved using the equation above)

[0088] Assign EquivalentRotation=ComputeEquivalentRotation(X,d1=sampledDisparityVal,d2=vioDisparity)

[0089] If abs(EquivalentRotation) is less than the threshold

[0090] Assign featureCount = featureCount + 1

[0091] Assign rotationMean=rotationMea+EquivalentRotation

[0092] RotationMeanSquared =

[0093] rotationMeanSquared+EquivalentRotation*EquivalentRotation

[0094] Assign rotationMean=rotationMean / featureCount

[0095] Assign rotationMeanSquared rotationMeanSquared / featureCount

[0096] Assign RotationVariance=rotationMeanSquared-rotationMean*rotationMean;

[0097] Returns [featureCount, rotationMean, RotationVariance]

[0098] The following pseudocode illustrates an example implementation of the yaw update module 610's operation:

[0099] Yaw update process UpdateYaw():

[0100] 1. Assign FilteredYawBias=Filter(RotationMean,RotationVariance)

[0101] 2. If FilteredYawBias is greater than the threshold

[0102] 3. Update total yaw curvature estimate

[0103] Figure 7 This is a block diagram illustrating a corrected depth frame according to one embodiment. VIO t-1 702 shows VIO data from the VIO system 302, captured at time t-1. Depth frame t-1 704 shows stereo depth data for the same frame captured at time t-1. The warp correction module 216 determines a corrected yaw 706 based on an estimated yaw deviation (based on a comparison of depth data from feature points in the VIO data and stereo depth data). The depth map system 304 receives updated configuration settings including the estimated yaw deviation and corrects its yaw deviation. The depth map system 304 generates a depth frame (depth frame t 708) at time t based on the corrected yaw deviation.

[0104] Figure 8 This is a flowchart illustrating a method 800 for adjusting a stereoscopic VIO feature according to an example embodiment. The operations in method 800 can be performed by the bending correction module 216 as described above. Figure 4 The described components (e.g., modules, engines) are used to perform this action. Therefore, method 800 is described by way of example with reference to pitch-roll bend module 402. However, it should be understood that at least some operations of method 800 can be deployed on various other hardware configurations or performed by similar components residing elsewhere.

[0105] In block 802, the pitch-roll bend module 402 accesses the previous stereo frame. In block 804, the pitch-roll bend module 402 determines the stereo VIO features from the previous stereo frame. In block 806, the pitch-roll bend module 402 projects the stereo VIO features onto the depth coordinate system. In decision block 808, the pitch-roll bend module 402 determines whether the stereo VIO features are aligned. In block 810, the pitch-roll bend module 402 calculates the pitch-roll offset. In block 812, the pitch-roll bend module 402 adjusts the depth coordinate system in the current stereo frame using the pitch-roll offset.

[0106] It should be noted that other implementations may use different sequencing, more or fewer operations, and different nomenclature or terminology to accomplish similar functionality. In some implementations, various operations may be executed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein have been chosen to illustrate some operational principles in a simplified form.

[0107] Figure 9 This is a flowchart illustrating a method 900 for adjusting yaw deviation according to an example embodiment. The operations in method 900 can be performed by the bend correction module 216 as described above. Figure 4The described components (e.g., modules, engines) perform the operations. Therefore, method 900 is described by way of example with reference to yaw bending module 404. However, it should be understood that at least some operations of method 900 can be deployed on various other hardware configurations or performed by similar components residing elsewhere.

[0108] In block 902, the yaw curve module 404 accesses the 3D landmark at time t from the VIO. In block 904, the yaw curve module 404 generates a 2D projection of the 3D landmark onto the 2D disparity map. In block 906, the yaw curve module 404 provides the depth of the 3D landmark at time t from the depth map onto the disparity map. In block 908, the yaw curve module 404 calculates the disparity value between the landmark and other landmarks at time t. In block 910, the yaw curve module 404 estimates the yaw deviation for each landmark based on the corresponding disparity value. In block 912, the yaw curve module 404 adjusts the yaw deviation based on the estimated yaw deviation.

[0109] Figure 10 This is a flowchart illustrating a method 1000 for updating a total yaw curvature estimate according to an example implementation. The operations in method 1000 can be performed by the curvature correction module 216 as described above. Figure 2 The described components (e.g., modules, engines) are used to perform this operation. Therefore, method 1000 is described by way of example with reference to the bending correction module 216. However, it should be understood that at least some operations of method 1000 can be deployed on various other hardware configurations or performed by similar components residing elsewhere.

[0110] In block 1002, the yaw curve module 404 calculates the disparity value for each landmark in the previous frame at time t-1. In block 1004, the yaw curve module 404 calculates the yaw deviation value based on the disparity value of each landmark in the previous frame at time t-1. In block 1006, the yaw curve module 404 calculates the mean / variance of the yaw deviation values. In block 1008, the yaw curve module 404 filters the yaw deviation values. In decision block 1010, the yaw curve module 404 determines whether the filtered yaw deviation values ​​exceed a yaw deviation threshold. In block 1012, the yaw curve module 404 updates the total yaw curve estimate for the next frame at time t.

[0111] Figure 11 The following examples illustrate misalignment errors caused by the bending of flexible devices. Example 1102 shows feature counterparts that are not located on the same grid line due to pitch / roll bending. Example 1104 shows z-bias 1106 caused by yaw bending.

[0112] Figure 12The example illustrates pitch-roll misalignment according to one embodiment. Example 1202 shows corresponding features (between the left and right sides) that are not on the same grid line due to curvature. Example 1204 shows corresponding features that are on the same grid line.

[0113] Figure 13 The diagram illustrates depth misalignment on a projected two-dimensional map according to one embodiment. Example 1302 shows a disparity map that varies over time. Example 1302 indicates that the disparity (corresponding to each feature) is relatively small over time (within a threshold). Example 1304 indicates that the disparity remains outside the yaw threshold over time, and the yaw deviation is updated.

[0114] Systems with head-worn devices

[0115] Figure 14 A network environment 1400 in which a head-mounted wearable device 1402 can be implemented is shown according to an example implementation. Figure 14 This is a high-level functional block diagram of an example head-mounted wearable device 1402, which is communicatively coupled to a mobile client device 1438 and a server system 1432 via various networks 1440.

[0116] The head-mounted wearable device 1402 includes at least one of a visible light camera 1412, an infrared emitter 1414, and an infrared camera 1416. A client device 1438 is capable of connecting to the head-mounted wearable device 1402 using both communication 1434 and communication 1436. The client device 1438 is connected to a server system 1432 and a network 1440. The network 1440 may include any combination of wired and wireless connections.

[0117] The head-worn wearable device 1402 also includes two image displays of the optical component image display 1404. These two image displays include one image display associated with the left side of the head-worn wearable device 1402 and one image display associated with the right side. The head-worn wearable device 1402 also includes an image display driver 1408, an image processor 1410, low-power circuitry 1426, and high-speed circuitry 1418. The image display 1404 of the optical component is used to present images and videos to the user of the head-worn wearable device 1402, including images that may include a graphical user interface.

[0118] The image display driver 1408 commands and controls the image display 1404 of the optical component to display images. The image display driver 1408 can directly deliver image data to the image display 1404 of the optical component for display, or it can convert the image data into a signal or data format suitable for delivery to the image display device. For example, the image data can be video data formatted according to compression formats such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, ​​etc., and still image data can be formatted according to compression formats such as Portable Network Group (PNG), Joint Photographic Experts Group (JPEG), Tag Image File Format (TIFF), or Exchangeable Image File Format (Exif), etc.

[0119] As described above, the head-wearable device 1402 includes a frame and a stem (or temple) extending laterally from the frame. The head-wearable device 1402 also includes a user input device 1406 (e.g., a touch sensor or a press button), comprising an input surface on the head-wearable device 1402. The user input device 1406 (e.g., a touch sensor or a press button) is used to receive input selections from the user for manipulating a graphical user interface of a presented image.

[0120] Figure 14 The components shown for the head-worn wearable device 1402 are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in blocks, frames, hinges, or nose bridges of the head-worn wearable device 1402. The left and right sides may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, camera lenses, or any other corresponding visible light or light-capturing elements that can be used to capture data, including images of scenes with unknown objects.

[0121] The head-mounted wearable device 1402 includes a memory 1422 that stores instructions for performing a subset or all of the functions described herein. The memory 1422 may also include a storage device.

[0122] like Figure 14As shown, the high-speed circuit 1418 includes a high-speed processor 1420, a memory 1422, and a high-speed wireless circuit 1424. In this example, an image display driver 1408 is coupled to the high-speed circuit 1418 and operated by the high-speed processor 1420 to drive the left and right image displays of the image display 1404 of the optical components. The high-speed processor 1420 can be any processor capable of managing the operation of any general computing system required for high-speed communication and the head-worn wearable device 1402. The high-speed processor 1420 includes the processing resources required to manage high-speed data transmission over communication 1436 to a wireless local area network (WLAN) using the high-speed wireless circuit 1424. In some examples, the high-speed processor 1420 executes an operating system (such as a LINUX operating system or another such operating system for the head-worn wearable device 1402), and the operating system is stored in the memory 1422 for execution. Among other duties, the high-speed processor 1420, which executes the software architecture of the head-worn wearable device 1402, manages data transmission with the high-speed wireless circuit 1424. In some examples, the high-speed wireless circuit 1424 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard (also referred to herein as Wi-Fi). In other examples, other high-speed communication standards can be implemented using the high-speed wireless circuit 1424.

[0123] The low-power wireless circuitry 1430 and high-speed wireless circuitry 1424 of the head-mounted wearable device 1402 may include a short-range transceiver (Bluetooth). TM The transceiver includes wireless wide area network, local area network, or wide area network transceivers (e.g., cellular or WiFi). Client device 1438, which includes transceivers communicating via communications 1434 and communications 1436, can be implemented using details of the architecture of head-mounted wearable device 1402, as can other elements of network 1440.

[0124] Memory 1422 includes any storage device capable of storing various data and applications, including camera data generated by the left and right infrared imaging devices 1416 and image processor 1410, and images for display generated by image display driver 1408 on image display 1404 of optical components. While memory 1422 is shown as integrated with high-speed circuitry 1418, in other examples, memory 1422 may be a separate, independent component of head-mounted wearable device 1402. In some such examples, the circuitry may be provided by lines connecting the image processor 1410 or low-power processor 1428 to memory 1422 via a chip including high-speed processor 1420. In other examples, high-speed processor 1420 may manage addressing of memory 1422 such that low-power processor 1428 will activate high-speed processor 1420 whenever a read or write operation involving memory 1422 is required.

[0125] like Figure 14 As shown, the low-power processor 1428 or high-speed processor 1420 of the head-mounted wearable device 1402 may be coupled to a camera device (visible light camera 1412; infrared emitter 1414, or infrared camera 1416), an image display driver 1408, a user input device 1406 (e.g., a touch sensor or a press button), and a memory 1422.

[0126] The head-mounted device 1402 is connected to a host computer. For example, the head-mounted wearable device 1402 is paired with a client device 1438 via communication 1436, or connected to a server system 1432 via a network 1440. The server system 1432 may be one or more computing devices as part of a service or network computing system, and for example, it includes a processor, memory, and a network communication interface to communicate with the client device 1438 and the head-mounted wearable device 1402 via the network 1440.

[0127] Client device 1438 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via network 1440, communication 1434, or communication 1436. Client device 1438 may also store in its memory at least a portion of instructions for generating dual-channel audio content to implement the functions described herein.

[0128] The output components of the head-worn wearable device 1402 include visual components such as displays (such as liquid crystal displays (LCDs), plasma display panels (PDPs), light-emitting diode (LED) displays, projectors, or waveguides). The image display of the optical components is driven by an image display driver 1408. The output components of the head-worn wearable device 1402 also include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. The input components (e.g., user input devices 1406) of the head-worn wearable device 1402, client device 1438, and server system 1432 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide the position and force of touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.

[0129] The head-mounted wearable device 1402 may optionally include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-mounted wearable device 1402. For example, peripheral device elements may include any I / O components, including output components, motion components, position components, or any other such components described herein.

[0130] For example, biometric components include those that detect expressions (e.g., hand gestures, facial expressions, voice expressions, body posture, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identify people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion components include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Positioning components include position sensor components (e.g., Global Positioning System (GPS) receiver components), WiFi or Bluetooth™ transceivers that generate positioning system coordinates, altitude sensor components (e.g., altimeters or barometers that detect air pressure, from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates can also be received from client device 1438 via low-power wireless circuit 1430 or high-speed wireless circuit 1424 via communication 1436.

[0131] Figure 15This is a block diagram 1500 illustrating a software architecture 1504, which can be installed on any one or more of the devices described herein. The software architecture 1504 is supported by hardware such as machine 1502, which includes a processor 1520, memory 1526, and I / O components 1538. In this example, the software architecture 1504 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1504 includes layers such as an operating system 1512, libraries 1510, frameworks 1508, and applications 1506. Operationally, application 1506 invokes API call 1550 through the software stack and receives message 1552 in response to API call 1550.

[0132] Operating system 1512 manages hardware resources and provides public services. Operating system 1512 includes, for example, kernel 1514, services 1516, and drivers 1522. Kernel 1514 serves as an abstraction layer between hardware and other software layers. For example, kernel 1514 provides memory management, processor management (e.g., scheduling), component management, networking and security settings, and other functions. Services 1516 can provide other public services to other software layers. Driver 1522 is responsible for controlling or interfacing with the underlying hardware. For example, driver 1522 may include a display driver, a camera driver, etc. or Low-power drives, flash drives, serial communication drives (e.g., Universal Serial Bus (USB) drives), Drivers, audio drivers, power management drivers, etc.

[0133] Library 1510 provides low-level public infrastructure used by application 1506. Library 1510 may include system library 1518 (e.g., the C standard library), which provides functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 1510 may include API library 1524, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., OpenGL frameworks for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite, which provides various relational database functions), web libraries (e.g., WebKit, which provides web browsing capabilities), and so on. Library 1510 may also include various other libraries 1528 to provide many other APIs to application 1506.

[0134] Framework 1508 provides high-level common infrastructure for use by Application 1506. For example, Framework 1508 provides various graphical user interface (GUI) functions, high-level resource management, and high-level location services. Framework 1508 can provide a wide range of other APIs that can be used by Application 1506, some of which may be specific to a particular operating system or platform.

[0135] In an example implementation, application 1506 may include a home application 1536, a contacts application 1530, a browser application 1532, a book reader application 1534, a location application 1542, a media application 1544, a messaging application 1546, a game application 1548, and a variety of other applications such as third-party application 1540. Application 1506 is a program that performs the functions defined in the program. One or more applications 1506 can be created using various programming languages, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In a particular example, third-party application 1540 (e.g., used by an entity other than a vendor of a particular platform using Android) TM or iOS TM Applications developed using a Software Development Kit (SDK) can be used on platforms such as iOS. TM ANDROID TM , Mobile software running on the phone's mobile operating system or another mobile operating system. In this example, a third-party application 1540 can invoke API calls 1550 provided by the operating system 1512 to facilitate the functions described herein.

[0136] Figure 16 This is a schematic representation of machine 1600, within which instructions 1608 (e.g., software, program, application, app, or other executable code) can be executed to cause machine 1600 to perform any or more of the methods discussed herein. For example, instructions 1608 can cause machine 1600 to perform any or more of the methods described herein. Instructions 1608 transform the general, unprogrammed machine 1600 into a specific machine 1600 programmed to perform the described and illustrated functions in the described manner. Machine 1600 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1600 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1600 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 1608 specifying actions to be taken by machine 1600. Furthermore, although only a single machine 1600 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 1608 to perform any one or more of the methods discussed herein.

[0137] Machine 1600 may include processor 1602, memory 1604, and I / O components 1642 configured to communicate with each other via bus 1644. In an example embodiment, processor 1602 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 1606 and processor 1610 that execute instruction 1608. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 16 Multiple processors 1602 are shown, but machine 1600 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0138] Memory 1604 includes main memory 1612, static memory 1614, and memory cells 1616, all of which are accessible by processor 1602 via bus 1644. Main memory 1604, static memory 1614, and memory cells 1616 store instructions 1608 embodying any one or more of the methods or functions described herein. Instructions 1608 may also reside wholly or partially in main memory 1612, in static memory 1614, in machine-readable medium 1618 within memory cells 1616, within at least one processor in processor 1602 (e.g., within the processor's cache memory), or in any suitable combination thereof during execution by machine 1600.

[0139] I / O component 1642 may include a wide variety of components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurement results, etc. The specific I / O component 1642 included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones may include touch input devices or other such input mechanisms, while headless server machines are unlikely to include such touch input devices. It should be understood that I / O component 1642 may be included in... Figure 16 Many other components are not shown. In various example embodiments, I / O component 1642 may include output component 1628 and input component 1630. Output component 1628 may include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., a speaker), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. Input component 1630 may include alphanumeric input components (e.g., a keyboard, a touchscreen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), haptic input components (e.g., a physical button, a touchscreen or other haptic input component that provides the position and / or force of a touch or touch gesture), audio input components (e.g., a microphone), etc.

[0140] In other example implementations, I / O component 1642 may include biometric component 1632, motion component 1634, environmental component 1636 or positioning component 1638, and various other components. For example, biometric component 1632 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1634 includes accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), and the like. Environmental component 1636 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an hearing sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting hazardous gas concentrations to ensure safety or measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Positioning component 1638 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer for detecting air pressure, from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), etc.

[0141] A wide variety of technologies can be used to implement communication. I / O component 1642 also includes communication component 1640, which is operable to couple machine 1600 to network 1620 or device 1622 via coupling 1624 and coupling 1626, respectively. For example, communication component 1640 may include a network interface component or another suitable device to interface with network 1620. In other examples, communication component 1640 may include wired communication component, wireless communication component, cellular communication component, near field communication (NFC) component, etc. Components (e.g.) (Low energy consumption) Components and other communication components for providing communication via other modes. Device 1622 can be any peripheral device from another machine or various peripheral devices (e.g., a peripheral device coupled via USB).

[0142] Furthermore, the communication component 1640 can detect identifiers or include components operable to detect identifiers. For example, the communication component 1640 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes) or an acoustic detection component (e.g., a microphone for identifying audio signals from tags). Additionally, various information can be derived via the communication component 1640, such as location via Internet Protocol (IP) geolocation, etc. The location of signal triangulation, the location of NFC beacon signals that can be detected to indicate a specific location, and so on.

[0143] Various memories (e.g., memory 1604, main memory 1612, static memory 1614, and / or the memory of processor 1602) and / or storage units 1616 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1608) cause various operations to implement the disclosed embodiments when executed by processor 1602.

[0144] Instructions 1608 can be sent or received over network 1620 via a transmission medium using a network interface device (e.g., a network interface component included in communication component 1640) and using any of a plurality of known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 1608 can be sent or received to device 1622 via a transmission medium using coupling 1626 (e.g., peer-to-peer coupling).

[0145] Although embodiments have been described with reference to specific example embodiments, it will be apparent that various modifications and changes can be made to these embodiments without departing from the broader scope of this disclosure. Therefore, the specification and drawings should be considered illustrative rather than restrictive. The accompanying drawings, which form a part of this invention, illustrate specific embodiments in which the subject matter can be practiced by way of illustration rather than limitation. The illustrated embodiments have been described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other embodiments can be utilized and derived therefrom, allowing for structural and logical substitutions and changes without departing from the scope of this disclosure. Therefore, the specific embodiments should not be construed as restrictive, and the scope of the various embodiments is defined only by the appended claims and the full scope of their equivalents.

[0146] These embodiments of the subject matter of this invention may be referred to herein individually and / or collectively by the term "invention," merely for convenience, and if more than one invention or inventive concept is disclosed, it is not intended to voluntarily limit the scope of this application to any single invention or inventive concept. Therefore, although specific embodiments have been shown and described herein, it should be understood that any arrangement calculated to achieve the same purpose may substitute for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of the various embodiments. Combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art upon review of the foregoing description.

[0147] An abstract of this disclosure is provided to allow the reader to quickly determine the nature of this technical disclosure. The abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen in the foregoing detailed description, various features are combined in a single embodiment for the purpose of simplifying this disclosure. This approach of the disclosure should not be construed as reflecting an intention to require more features than expressly stated in each claim. Rather, as reflected in the appended claims, the subject matter of the invention lies in fewer than all features of a single disclosed embodiment. Therefore, the claims are hereby incorporated into the detailed description, wherein each claim is considered an independent, separate embodiment.

[0148] Example

[0149] Example 1 is a method for correcting the curvature of a flexible device, comprising: accessing feature data of a first stereo frame generated by a stereo optical sensor of the flexible device, the feature data being generated based on a visual inertial odometry (VIO) system of the flexible device; accessing depth map data of the first stereo frame, the depth map data being generated based on a depth map system of the flexible device; estimating pitch-roll and yaw deviations based on the feature data and the depth map data of the first stereo frame; and generating a second stereo frame after the first stereo frame, the second stereo frame being based on the pitch-roll and yaw deviations of the first stereo frame.

[0150] Example 2 includes Example 1, wherein estimating the pitch-roll bias further includes: determining a stereo VIO feature of the first stereo frame; projecting the stereo VIO feature onto a two-dimensional coordinate system; determining that the projected stereo VIO feature is not aligned in the two-dimensional coordinate system; and, in response to determining that the projected stereo VIO feature is not aligned in the two-dimensional coordinate system, calculating the pitch-roll bias based on the misalignment of the projected stereo VIO feature in the two-dimensional coordinate system.

[0151] Example 3 includes Example 2, wherein determining that the projected stereo VIO features are not aligned in the two-dimensional coordinate system further includes: identifying a first feature in the left frame of the first stereo frame; identifying a second feature in the right frame of the first stereo frame, the second feature corresponding to the first feature; and determining that the first feature and the second feature are located on different grid lines in the two-dimensional coordinate system.

[0152] Example 4 includes Example 2, wherein generating the second stereo frame further includes: identifying a first feature in the left frame of the second stereo frame; identifying a second feature in the right frame of the second stereo frame, the second feature corresponding to the first feature; and correcting the position of the first feature in the left frame or the second feature in the right frame based on the pitch-roll deviation of the first stereo frame.

[0153] Example 5 includes Example 1, wherein estimating the yaw deviation further includes: using the VIO system to identify three-dimensional landmarks in the first stereo frame; projecting the three-dimensional landmarks onto a two-dimensional disparity map, the two-dimensional disparity map indicating the two-dimensional position of the landmarks and their corresponding depth values; calculating a depth deviation value for each landmark in the first stereo frame; and calculating a yaw deviation value based on the depth deviation value for each landmark.

[0154] Example 6 includes Example 5, wherein projecting the three-dimensional landmark further includes: using the VIO system to depict the first landmark of the first stereo frame on the two-dimensional disparity map with a first depth value; and using the depth map data to depict the first landmark of the first stereo frame on the two-dimensional disparity map with a second depth value, wherein the depth deviation value of the first landmark is the difference between the first depth value and the second depth value.

[0155] Example 7 includes Example 5, and further includes: calculating statistics on yaw deviation values ​​of a plurality of landmarks in the first stereo frame; calculating filtered yaw deviation values ​​based on the statistics on yaw deviation values; determining that the filtered yaw deviation values ​​exceed a yaw deviation threshold; updating the total yaw curvature estimate of the second stereo frame in response to determining that the filtered yaw deviation values ​​exceed the yaw deviation threshold; updating a correction map based on the total yaw curvature estimate; and calculating a depth map based on the correction map.

[0156] Example 8 includes Example 1, wherein the VIO system is configured to identify the posture of the flexible device based on sensor data from inertial and optical sensors of the flexible device.

[0157] Example 9 includes Example 1, wherein the depth map system determines the depth of pixels in the first stereo frame based on triangulation of pixels depicted in the right and left sides of the first stereo frame.

[0158] Example 10 includes Example 1, and further includes: generating virtual content in the first stereo frame; and adjusting the display of the virtual content in the second stereo frame based on the pitch-roll and yaw deviations of the first stereo frame.

[0159] Example 11 is a computing device including a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access feature data of a first stereo frame generated by a stereo optical sensor of a flexible device, the feature data being generated based on a visual inertial odometry (VIO) system of the flexible device; access depth map data of the first stereo frame, the depth map data being generated based on a depth map system of the flexible device; estimate pitch-roll and yaw deviations based on the feature data and the depth map data of the first stereo frame; and generate a second stereo frame after the first stereo frame, the second stereo frame being based on the pitch-roll and yaw deviations of the first stereo frame.

[0160] Example 12 includes Example 11, wherein estimating the pitch-roll bias further includes: determining a stereo VIO feature of the first stereo frame; projecting the stereo VIO feature onto a two-dimensional coordinate system; determining that the projected stereo VIO feature is not aligned in the two-dimensional coordinate system; and, in response to determining that the projected stereo VIO feature is not aligned in the two-dimensional coordinate system, calculating the pitch-roll bias based on the misalignment of the projected stereo VIO feature in the two-dimensional coordinate system.

[0161] Example 13 includes Example 12, wherein determining that the projected stereo VIO features are not aligned in the two-dimensional coordinate system further includes: identifying a first feature in the left frame of the first stereo frame; identifying a second feature in the right frame of the first stereo frame, the second feature corresponding to the first feature; and determining that the first feature and the second feature are located on different grid lines in the two-dimensional coordinate system.

[0162] Example 14 includes Example 12, wherein generating the second stereo frame further includes: identifying a first feature in the left frame of the second stereo frame; identifying a second feature in the right frame of the second stereo frame, the second feature corresponding to the first feature; and correcting the position of the first feature in the left frame or the second feature in the right frame based on the pitch-roll deviation of the first stereo frame.

[0163] Example 15 includes Example 11, wherein estimating the yaw deviation further includes: using the VIO system to identify three-dimensional landmarks in the first stereo frame; projecting the three-dimensional landmarks onto a two-dimensional disparity map, the two-dimensional disparity map indicating the two-dimensional position of the landmarks and their corresponding depth values; calculating a depth deviation value for each landmark in the first stereo frame; and calculating a yaw deviation value based on the depth deviation value for each landmark.

[0164] Example 16 includes Example 15, wherein projecting the three-dimensional landmark further includes: using the VIO system to depict the first landmark of the first stereo frame on the two-dimensional disparity map with a first depth value; and using the depth map data to depict the first landmark of the first stereo frame on the two-dimensional disparity map with a second depth value, wherein the depth deviation value of the first landmark is the difference between the first depth value and the second depth value.

[0165] Example 17 includes Example 15, wherein the instructions further configure the apparatus to: calculate statistics of yaw deviation values ​​for a plurality of landmarks in the first stereo frame; calculate filtered yaw deviation values ​​based on the statistics of the yaw deviation values; determine that the filtered yaw deviation values ​​exceed a yaw deviation threshold; update the total yaw curvature estimate of the second stereo frame in response to determining that the filtered yaw deviation values ​​exceed the yaw deviation threshold; update a correction map based on the total yaw curvature estimate; and calculate a depth map based on the correction map.

[0166] Example 18 includes Example 11, wherein the VIO system is configured to identify the posture of the flexible device based on sensor data from inertial and optical sensors of the flexible device.

[0167] Example 19 includes Example 11, wherein the depth map system determines the depth of pixels in the first stereo frame based on triangulation of pixels depicted in the right and left sides of the first stereo frame.

[0168] Example 20 is a non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to: access feature data of a first stereo frame generated by a stereo optical sensor of a flexible device, the feature data being generated based on a visual inertial odometry (VIO) system of the flexible device; access depth map data of the first stereo frame, the depth map data being generated based on a depth map system of the flexible device; estimate pitch-roll and yaw deviations based on the feature data and the depth map data of the first stereo frame; and generate a second stereo frame after the first stereo frame, the second stereo frame being based on the pitch-roll and yaw deviations of the first stereo frame.

Claims

1. A method for correcting the bending of a flexible device, comprising: Access feature data of a first stereo frame generated by a stereo optical sensor of the flexible device, the feature data being generated based on the visual inertial odometry (VIO) system of the flexible device; Access the depth map data of the first stereo frame, the depth map data being generated based on the depth map system of the flexible device; The pitch-roll bias is estimated based on the feature data and depth map data of the first stereo frame by the following operation: projecting the stereo VIO features of the first stereo frame onto a two-dimensional coordinate system, and calculating the pitch-roll bias based on the misalignment of the stereo VIO features projected in the two-dimensional coordinate system. as well as A second stereo frame is generated after the first stereo frame, the second stereo frame being based on the pitch-roll deviation of the first stereo frame.

2. The method according to claim 1, wherein, The estimation of the pitch-roll deviation also includes: Determine the stereo VIO features of the first stereo frame; and It was determined that the projected stereo VIO feature was not aligned in the two-dimensional coordinate system.

3. The method according to claim 2, wherein, Determining that the projected stereo VIO feature is not aligned in the two-dimensional coordinate system also includes: Identify the first feature in the left frame of the first stereo frame; Identify a second feature in the right frame of the first stereo frame, the second feature corresponding to the first feature; and The first feature and the second feature are determined to be located on different grid lines in the two-dimensional coordinate system.

4. The method according to claim 2, wherein, Generating the second stereo frame also includes: Identify the first feature in the left frame of the second stereo frame; Identify a second feature in the right frame of the second stereoscopic frame, the second feature corresponding to the first feature; and The position of the first feature in the left frame or the second feature in the right frame is corrected based on the pitch-roll deviation of the first stereo frame.

5. The method according to claim 1, further comprising: The yaw deviation is estimated based on the feature data and depth map data of the first stereo frame using the following operations: The VIO system is used to identify the three-dimensional landmarks in the first stereo frame; The three-dimensional landmark is projected onto a two-dimensional parallax map, which indicates the two-dimensional position of the landmark and its corresponding depth value. Calculate the depth deviation value of each landmark in the first stereo frame; as well as The yaw deviation value is calculated based on the depth deviation value for each landmark.

6. The method according to claim 5, wherein, The projected three-dimensional landmarks also include: The VIO system is used to depict the first landmark of the first stereo frame on the two-dimensional disparity map with a first depth value; and The first landmark of the first stereo frame is depicted on the two-dimensional disparity map using the depth map data with a second depth value. The depth deviation value of the first landmark is the difference between the first depth value and the second depth value.

7. The method according to claim 5, further comprising: Calculate the statistics of yaw deviation values ​​of multiple landmarks in the first stereo frame; The filtered yaw deviation value is calculated based on the statistics of the yaw deviation value; The filtered yaw deviation value is determined to exceed the yaw deviation threshold; In response to determining that the filtered yaw deviation value exceeds the yaw deviation threshold, the total yaw curvature estimate of the second stereo frame is updated; The correction map is updated based on the total yaw curvature estimate; as well as A depth map is calculated based on the corrected map.

8. The method according to claim 1, wherein, The VIO system is configured to identify the posture of the flexible device based on sensor data from inertial and optical sensors.

9. The method according to claim 1, wherein, The depth mapping system determines the depth of pixels in the first stereo frame based on triangulation of pixels depicted in the right and left sides of the first stereo frame.

10. The method according to claim 1, further comprising: Virtual content is generated in the first stereoscopic frame; as well as The display of the virtual content in the second stereo frame is adjusted based on the pitch-roll and yaw deviations of the first stereo frame.

11. A computing device, comprising: One or more processors; as well as A memory storing instructions that, when executed by the one or more processors, configure the device to perform operations including: Access feature data of a first stereo frame generated by a stereo optical sensor of the flexible device, the feature data being generated based on the visual inertial odometry (VIO) system of the flexible device; Access the depth map data of the first stereo frame, the depth map data being generated based on the depth map system of the flexible device; The pitch-roll bias is estimated based on the feature data and depth map data of the first stereo frame by the following operations: projecting the stereo VIO features of the first stereo frame onto a two-dimensional coordinate system, and calculating the pitch-roll bias based on the misalignment of the stereo VIO features projected in the two-dimensional coordinate system; and A second stereo frame is generated after the first stereo frame, the second stereo frame being based on the pitch-roll deviation of the first stereo frame.

12. The computing device according to claim 11, wherein, The estimation of the pitch-roll deviation also includes: Determine the stereo VIO features of the first stereo frame; and It was determined that the projected stereo VIO feature was not aligned in the two-dimensional coordinate system.

13. The computing device according to claim 12, wherein, Determining that the projected stereo VIO feature is not aligned in the two-dimensional coordinate system also includes: Identify the first feature in the left frame of the first stereo frame; Identify a second feature in the right frame of the first stereo frame, the second feature corresponding to the first feature; and The first feature and the second feature are determined to be located on different grid lines in the two-dimensional coordinate system.

14. The computing device according to claim 12, wherein, Generating the second stereo frame also includes: Identify the first feature in the left frame of the second stereo frame; Identify a second feature in the right frame of the second stereoscopic frame, the second feature corresponding to the first feature; and The position of the first feature in the left frame or the second feature in the right frame is corrected based on the pitch-roll deviation of the first stereo frame.

15. The computing device according to claim 11, wherein, The operation also includes: The yaw deviation is estimated based on the feature data and depth map data of the first stereo frame using the following operations: The VIO system is used to identify the three-dimensional landmarks in the first stereo frame; The three-dimensional landmark is projected onto a two-dimensional parallax map, which indicates the two-dimensional position of the landmark and its corresponding depth value. Calculate the depth deviation value of each landmark in the first stereo frame; and The yaw deviation value is calculated based on the depth deviation value for each landmark.

16. The computing device according to claim 15, wherein, The projected three-dimensional landmarks also include: The VIO system is used to depict the first landmark of the first stereo frame on the two-dimensional disparity map with a first depth value; and The first landmark of the first stereo frame is depicted on the two-dimensional disparity map using the depth map data with a second depth value. The depth deviation value of the first landmark is the difference between the first depth value and the second depth value.

17. The computing device according to claim 15, wherein, The operation also includes: Calculate the statistics of yaw deviation values ​​of multiple landmarks in the first stereo frame; The filtered yaw deviation value is calculated based on the statistics of the yaw deviation value; The filtered yaw deviation value is determined to exceed the yaw deviation threshold; In response to determining that the filtered yaw deviation value exceeds the yaw deviation threshold, the total yaw curvature estimate of the second stereo frame is updated; The correction map is updated based on the total yaw curvature estimate; and A depth map is calculated based on the corrected map.

18. The computing device according to claim 11, wherein, The VIO system is configured to identify the posture of the flexible device based on sensor data from inertial and optical sensors.

19. The computing device according to claim 11, wherein, The depth mapping system determines the depth of pixels in the first stereo frame based on triangulation of pixels depicted in the right and left sides of the first stereo frame.

20. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform an operation, the operation comprising: Access feature data of a first stereo frame generated by a stereo optical sensor of the flexible device, the feature data being generated based on the visual inertial odometry (VIO) system of the flexible device; Access the depth map data of the first stereo frame, the depth map data being generated based on the depth map system of the flexible device; The pitch-roll bias is estimated based on the feature data and depth map data of the first stereo frame by the following operation: projecting the stereo VIO features of the first stereo frame onto a two-dimensional coordinate system, and calculating the pitch-roll bias based on the misalignment of the stereo VIO features projected in the two-dimensional coordinate system. as well as A second stereo frame is generated after the first stereo frame, the second stereo frame being based on the pitch-roll deviation of the first stereo frame.

Citation Information

Patent Citations

  • 3D mapping with flexible camera rig

    US20150316767A1