Bundle adjustment using epipolar constraints

Bundle adjustment with epipolar constraints addresses deformation issues in augmented reality headsets by refining camera parameters and environmental models, enhancing calibration accuracy and efficiency.

JP7781161B2Active Publication Date: 2025-12-05MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023535837
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-14
Filing Date
2021-12-03
Publication Date
2025-12-05
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Augmented reality devices, particularly lightweight and wearable headsets, experience small but rapid deformations due to user movements, leading to inaccuracies in camera calibration and environmental mapping, which current methods struggle to address efficiently.

Method used

Implementing bundle adjustment using epipolar constraints to refine camera extrinsic parameters and environmental models by jointly optimizing reprojection and epipolar errors, allowing for real-time calibration and deformation correction during device use.

Benefits of technology

Enhances the accuracy and efficiency of camera calibration in augmented reality systems by accounting for deformations, improving the precision of environmental modeling and device positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007781161000012
    Figure 0007781161000012
  • Figure 0007781161000013
    Figure 0007781161000013
  • Figure 0007781161000014
    Figure 0007781161000014
Patent Text Reader

Abstract

A method, system, and apparatus for performing bundling adjustment using epipolar constraints. The method includes receiving image data from a headset for a particular pose. The method includes, at least in part, identifying at least one key point in a three-dimensional model of an environment represented in a first image and a second image, and performing a bundle adjustment. The bundle adjustment is performed by jointly optimizing a reprojection error for the at least one key point and an epipolar error for the at least one key point. Results of the bundle adjustment are used to at least one of: (i) update the three-dimensional model; (ii) determine a position of the headset at the particular pose; or (iii) determine extrinsic parameters of the first camera and the second camera.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification relates generally to image processing in extended reality systems, such as virtual, mixed, or augmented reality systems. [Background technology]

[0002] Augmented reality ("AR") and mixed reality ("MR") devices can include multiple sensors. Some examples of sensors are cameras, accelerometers, gyroscopes, global positioning system receivers, and magnetometers, e.g., compasses.

[0003] The AR device can receive data from multiple sensors, combine the data, and determine an output related to the user. For example, the AR device can receive gyroscope and camera data from separate sensors and use the received data to present content on a display. The AR device can use the sensor data, e.g., the camera data, to generate an environment map and use the environment map to present content on a display. Summary of the Invention [Means for solving the problem]

[0004] A computer vision system can use sensor data to generate an environmental model of the environment in which a device, e.g., the computer vision system, is located, estimate the location of the device within the environment, or both. For example, a computer vision system can use data from multiple sensors to generate an environmental model of the environment in which the device is located. The sensors can include a depth sensor, a camera, an inertial measurement unit, or a combination of two or more of these.

[0005] Augmented reality headsets can use a map or three-dimensional (“3D”) model of the environment to provide 3D information corresponding to a view of the environment. A computer vision system can use a simultaneous localization and mapping (“SLAM”) process to both update the environment model and determine an estimated location of the device within the environment. The device's location can include position data, orientation data, or both. As part of the SLAM process, the computer vision system can use bundle adjustment, a set membership process, a statistical process, or another suitable process. For example, as part of the SLAM process, the computer vision system can determine locations for three-dimensional points in the environment model that represent observable points in the environment. The observable points can represent portions of objects in the environment. The computer vision system can then use bundle adjustment to refine the locations of the three-dimensional points in the environment model, e.g., using additional data, updated data, or both, to make more accurate predictions of the locations of the observable points.

[0006] Bundle adjustment is a process for simultaneously refining a 3D model of the environment, the pose of the camera that captured the images, and / or the camera's extrinsic parameters using a set of images from different viewpoints. In bundle adjustment, errors between camera-related data, such as reprojection errors, are minimized.

[0007] Bundle adjustment using epipolar constraints can be used to perform highly accurate online calibration of camera sensor extrinsic characteristics in any camera-based SLAM system. When performed as part of bundle adjustment, online calibration can be sensitive to the weighting scheme employed to correct for the extrinsic characteristics. Epipolar constraints can be used to estimate the probability of deformation as the cross product of rotations and translations between two cameras or between one camera and a reference point. Thus, deformation errors can be recovered to achieve higher accuracy, efficiency, or both, for example, with a higher update rate.

[0008] Bundle adjustment using epipolar constraints can be used to more accurately estimate deformation between multiple camera sensors on flexible devices. Lightweight and wearable augmented reality devices, such as headsets, can be prone to small but rapid deformations over time. Applying epipolar constraints to bundle adjustment can address the problem of estimating that deformation in an efficient manner based on geometric constraints. The estimated deformations can be used to generate updated camera parameters, such as extrinsic features. The updated camera parameters can be used in multi-camera triangulation-based SLAM systems.

[0009] The described system can use epipolar constraints to perform bundle adjustment while the augmented reality system is being used by a user. For example, bundle adjustment can be performed for a system, e.g., an extended reality system, simultaneously as the system captures image data, generates extended reality output data, and displays the output data on a headset or other display. Some extended reality systems, e.g., augmented reality systems, include or are provided as wearable devices, such as headsets, on which cameras and other sensors are mounted. As a result, these systems are often moved during use as users walk, turn their heads, or perform other movements. These movements often change the forces and stresses on the device, which can cause temporary and / or permanent deformations in the wearable device. Bundle adjustment can be performed automatically while the augmented reality system is worn and moving, as the system determines it is necessary. This can result in improved performance for highly deformable systems that may undergo large amounts of bending, rotation, and other movements.

[0010] Performing bundle adjustment using epipolar constraints can improve the accuracy of camera calibration in an augmented reality headset. By applying epipolar constraints, the system can account for deformations between the cameras as they move from one pose to another. Thus, changes in the relative position or relative rotation between the cameras can be taken into account in the bundle adjustment.

[0011] In one general aspect, a method includes receiving image data from a headset for a particular pose of the headset, the image data including (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image and (ii) an epipolar error for the at least one key point based on the first image and the second image; and using a result of the bundle adjustment to perform at least one of (i) updating the three-dimensional model, (ii) determining a position of the headset at the particular pose, or (iii) determining extrinsic parameters of the first camera and the second camera.

[0012] In some implementations, the method includes providing an output for display by a headset based on the updated three-dimensional model.

[0013] In some implementations, epipolar error is the result of deformations in the headset that cause deviations from the headset calibration.

[0014] In some implementations, the method includes determining a set of extrinsic parameters for the first camera and the second camera based on the first image and the second image.

[0015] In some implementations, the extrinsic parameters include translation and rotation, which together indicate the relationship of the first camera or the second camera to a reference position on the headset.

[0016] In some implementations, the method includes receiving images from first and second cameras at each of a plurality of different poses along a path of movement of the headset, and using results of the optimization involving epipolar error to determine different extrinsic parameters for the first and second cameras for at least some of the different poses.

[0017] In some implementations, the method includes, at least in part, identifying a plurality of key points in a three-dimensional model of an environment represented in the first image and the second image, respectively, and jointly optimizing an error across each of the plurality of key points.

[0018] In some implementations, jointly optimizing the error includes minimizing the total error across each of the multiple key points, where the total error includes a combination of the reprojection error and the epipolar error.

[0019] In some implementations, the method includes receiving, from the headset, second image data for multiple poses of the headset; identifying, at least in part, at least one second key point in a three-dimensional model of the environment represented in the second image data; performing a bundle adjustment for each of the multiple poses by jointly optimizing (i) a reprojection error for the at least one key point based on the second image data and (ii) an epipolar error for the at least one key point based on the second image data; and using results of the bundle adjustment for each of the multiple poses to perform at least one of: (i) updating the three-dimensional model; (ii) determining another position of the headset in each of the multiple poses; or (iii) determining other extrinsic parameters of the first camera and the second camera in each of the multiple poses.

[0020] In some implementations, the method includes receiving, from a headset, first image data related to a first pose of the headset and second image data related to a second pose of the headset, wherein a deformation of the headset occurs between the first pose of the headset and the second pose of the headset; identifying at least one key point in a three-dimensional model of the environment represented, at least in part, in the first image data and the second image data; performing bundle adjustment using the first image data and the second image data by jointly optimizing at least an epipolar error (a) related to the at least one key point and (b) representing the deformation of the headset that occurred between the first pose of the headset and the second pose of the headset; and using a result of the bundle adjustment to perform at least one of: (i) updating the three-dimensional model; (ii) determining a first position of the headset in the first pose or a second position of the headset in the second pose; or (iii) determining first extrinsic parameters of a first camera and a second camera in the first pose or a second extrinsic parameters of a first camera and a second camera in the second pose.

[0021] In some implementations, the method includes performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for at least one key point based on the first image and the second image, (ii) an epipolar error for at least one key point based on the first image and the second image, and (iii) an error based on factory calibration data for the headset.

[0022] In some implementations, the method includes using the results of the bundle adjustment to update a set of poses for the headset.

[0023] In some implementations, using the results of the bundle adjustment includes updating the 3D model, including updating the positions of one or more key points in the 3D model.

[0024] In some implementations, using the results of the bundle adjustment includes determining a position for a particular pose, including determining a position for the headset relative to a three-dimensional model.

[0025] In some implementations, the first image and the second image were captured substantially simultaneously.

[0026] Other embodiments of these aspects include corresponding systems, apparatus, and computer programs encoded on computer storage devices configured to perform the actions of the methods. One or more computer systems can be so configured by software, firmware, hardware, or a combination thereof installed on the system that, when operated, causes the system to perform the actions. One or more computer programs can be so configured by having instructions that, when executed by a data processing device, cause the device to perform the actions.

[0027] The subject matter described herein can be implemented in various embodiments and may provide one or more of the following advantages: In some implementations, the bundle adjustment process described herein is faster, uses fewer computer resources, e.g., processor cycles or memory, or both, or a combination thereof, compared to other systems; In some implementations, the camera calibration process described herein can make adjustments to account for deformations of a device that includes a camera. These adjustments can improve the accuracy of calculations performed on the device, such as a three-dimensional model of the environment in which the device is located, the predicted position of the device within the environment, extrinsic parameters related to the camera included in the device, or a combination thereof.

[0028] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. The present specification also provides, for example, the following: (Item 1) 1. A method comprising: receiving image data from a headset relating to a particular pose of the headset, the image data comprising (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image, and (ii) an epipolar error for the at least one key point based on the first image and the second image; using the results of the bundle adjustment to at least one of: (i) updating the three-dimensional model; (ii) determining a position of the headset in the particular pose; or (iii) determining extrinsic parameters of the first camera and the second camera; A method comprising: (Item 2) Item 10. The method of item 1, comprising providing an output for display by the headset based on the updated three-dimensional model. (Item 3) Item 10. The method of item 1, wherein the epipolar error is a result of deformation of the headset causing a deviation from the headset calibration. (Item 4) Item 10. The method of item 1, comprising determining a set of extrinsic parameters for the first camera and the second camera based on the first image and the second image. (Item 5) Item 10. The method of item 1, wherein the extrinsic parameters include a translation and a rotation that together indicate a relationship of the first camera or the second camera to a reference position on the headset. (Item 6) receiving images from the first and second cameras at each of a plurality of different poses along a path of movement of the headset; using results of the optimization with the epipolar error to determine different extrinsic parameters for the first and second cameras for at least some of the different poses; Item 1. The method according to item 1, comprising: (Item 7) identifying a plurality of key points in the three-dimensional model of the environment, the plurality of key points being at least partially represented in the first image and the second image, respectively; jointly optimizing an error across each of the plurality of key points; Item 1. The method according to item 1, comprising: (Item 8) Item 8. The method of item 7, wherein jointly optimizing the error includes minimizing a total error across each of the plurality of key points, the total error comprising a combination of the reprojection error and the epipolar error. (Item 9) receiving second image data from the headset relating to a plurality of poses of the headset; identifying, at least in part, at least one second key point in the three-dimensional model of the environment represented in the second image data; performing a bundle adjustment for each of the plurality of poses by jointly optimizing (i) a reprojection error for the at least one key point based on the second image data, and (ii) an epipolar error for the at least one key point based on the second image data; using results of the bundle adjustment for each of the plurality of poses to at least one of: (i) updating the three-dimensional model; (ii) determining another position of the headset in each of the plurality of poses; or (iii) determining other extrinsic parameters of the first camera and the second camera in each of the plurality of poses; Item 1. The method according to item 1, comprising: (Item 10) receiving, from the headset, first image data relating to a first pose of the headset and second image data relating to a second pose of the headset, wherein a deformation of the headset occurs between the first pose of the headset and the second pose of the headset; identifying at least one second key point in the three-dimensional model of the environment represented, at least in part, in the first image data and in the second image data; performing the bundle adjustment using the first image data and the second image data by jointly optimizing at least the epipolar error (a) with respect to at least a second key point and (b) representing a deformation of the headset that occurs between a first pose and a second pose of the headset; using results of the bundle adjustment to at least one of: (i) updating the three-dimensional model; (ii) determining a first position of the headset in the first pose or a second position of the headset in the second pose; or (iii) determining first extrinsic parameters of the first camera and the second camera in the first pose or second extrinsic parameters of the first camera and the second camera in the second pose; Item 1. The method according to item 1, comprising: (Item 11) Item 10. The method of claim 1, comprising performing a bundle adjustment using the first image and the second image by jointly optimizing (i) the reprojection error for the at least one key point based on the first image and the second image, (ii) the epipolar error for the at least one key point based on the first image and the second image, and (iii) an error based on factory calibration data for the headset. (Item 12) Item 10. The method of item 1, comprising updating a set of poses for the headset using results of the bundle adjustment. (Item 13) Item 1. The method of item 1, wherein using the results of the bundle adjustment includes updating the three-dimensional model, which includes updating the positions of one or more key points in the three-dimensional model. (Item 14) Item 10. The method of item 1, wherein using the results of the bundle adjustment includes determining a position of the headset in the particular pose relative to the three-dimensional model. (Item 15) Item 10. The method of item 1, wherein the first image and the second image were captured substantially simultaneously. (Item 16) A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to: receiving image data from a headset relating to a particular pose of the headset, the image data comprising (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image, and (ii) an epipolar error for the at least one key point based on the first image and the second image; using the results of the bundle adjustment to at least one of: (i) updating the three-dimensional model; (ii) determining a position of the headset in the particular pose; or (iii) determining extrinsic parameters of the first camera and the second camera; 10. A non-transitory computer storage medium for performing operations including: (Item 17) Item 17. The computer storage medium of item 16, further comprising providing an output for display by the headset based on the updated three-dimensional model. (Item 18) Item 17. The computer storage medium of item 16, wherein the epipolar error is a result of deformation of the headset that causes a deviation from the headset calibration. (Item 19) Item 17. The computer storage medium of item 16, further comprising determining a set of extrinsic parameters for the first camera and the second camera based on the first image and the second image. (Item 20) 1. A system comprising one or more computers and one or more storage devices, the one or more storage devices having instructions stored thereon that, when executed by the one or more computers, cause the one or more computers to: receiving image data from a headset relating to a particular pose of the headset, the image data comprising (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image, and (ii) an epipolar error for the at least one key point based on the first image and the second image; using the results of the bundle adjustment to at least one of: (i) updating the three-dimensional model; (ii) determining a position of the headset in the particular pose; or (iii) determining extrinsic parameters of the first camera and the second camera; The system is operable to cause operations to be performed, including: [Brief explanation of the drawings]

[0029] [Figure 1]FIG. 1 depicts an example system in which a device updates its model of the environment using bundle adjustment with epipolar constraints.

[0030] [Figure 2] FIG. 2 depicts an example of the projection of key points on an augmented reality device.

[0031] [Figure 3] FIG. 3 depicts an example of a tracked pose of an augmented reality device.

[0032] [Figure 4] FIG. 4 depicts an exemplary system for performing bundle adjustment using epipolar constraints.

[0033] [Figure 5] FIG. 5 is a flow diagram of a process for updating a model of an environment using bundle adjustment with epipolar constraints.

[0034] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION

[0035] Detailed Description 1 depicts an example system 100 in which a device updates a model of an environment using bundle adjustment with epipolar constraints. Although FIG. 1 is described with reference to an augmented reality headset 102 as the device, any other suitable computer vision system can be used instead of or in addition to the augmented reality headset 102. For example, the augmented reality headset 102 can be any other suitable type of extended reality headset, such as a mixed reality headset.

[0036] The augmented reality headset 102 or another device can use bundle adjustment to update various data, including a three-dimensional ("3D") model 122 of the physical environment in which the augmented reality headset 102 is located, extrinsic parameters of the augmented reality headset's 102 camera, an estimated position of the augmented reality headset 102 within the physical environment, or any combination thereof. By performing bundle adjustment, the augmented reality headset 102 can determine a more accurate 3D model 122, more accurate extrinsic parameters, a more accurate estimated device position, or any combination thereof.

[0037] Bundle adjustment is a process for simultaneously refining a 3D model of the environment, the pose of the cameras that captured the images, and / or the extrinsic parameters of the cameras using a set of images from different viewpoints. In bundle adjustment, errors between cameras, such as reprojection errors, are minimized.

[0038] The augmented reality headset 102 can repeat the process of generating updated extrinsic parameters based on the received camera data for multiple physical locations or poses of the augmented reality headset 102. For example, as the augmented reality headset 102 moves through a physical environment along a path, the headset 102 can generate updated extrinsic parameters, updates to the 3D model 122, updated estimated device positions, or a combination thereof, for multiple positions along the path.

[0039] The augmented reality headset 102 includes a right camera 104 and a left camera 106. The augmented reality headset 102 may optionally include a center camera 105, a depth sensor 108, or both. As the augmented reality headset 102 moves through a physical environment, the augmented reality headset 102 receives image data 110 captured by the cameras 104, 106. For example, when the augmented reality headset 102 is in a particular physical location or pose within the physical environment, the cameras 104, 106 may capture particular image data 110 related to the particular physical location. The image data 110 may be any suitable image data, such as a first image captured by the camera 104 and a second image captured by the camera 106, or other data representing images captured by separate cameras. In some implementations, the image data 110 relates to an image in a sequence of video images. For example, the image data 110 may relate to a frame in the video sequence.

[0040] In FIG. 1 , deformation 101 occurs in headset 102. Deformation 101 may occur, for example, when headset 102 moves from a first pose to a second pose. In some examples, movement of headset 102 from a first pose to a second pose may cause deformation 101. In some examples, calculation drift, for example, due to rounding error, may occur over time, causing deformation 101. The deformation may cause differences in camera parameters from the headset calibration. For example, a result of the deformation is that, over time, the calibration parameters for the headset no longer represent the actual calibration of the headset.

[0041] The deformation 101 may cause a change in the relative position between the camera 104 and a reference position on the headset 102. The reference position may be, for example, the right camera 104, the center camera 105, the left camera 106, or the depth sensor 108. In a first pose, the camera 104 may be positioned at a first position 103, represented in FIG. 1 by a dashed circular line. When the augmented reality headset 102 moves to a second pose, the camera 104 may move to a second position, represented by a solid circular line. Thus, movement of the augmented reality headset 102 causes a deformation 101 in the headset 102, regardless of the physical deformation 101 or the calculated drift deformation. This deformation may result in a change in the relative position between the camera 104 and a reference position, for example, the camera 106.

[0042] The augmented reality headset 102 may be a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described herein are implemented. In some implementations, one or more of the components described with reference to the augmented reality headset 102 may be included in a separate system, such as on a server that communicates with the augmented reality headset 102 using a network. The network (not shown) may be a local area network ("LAN"), a wide area network ("WAN"), the Internet, or a combination thereof. The separate system may use a single server computer or multiple server computers operating in conjunction with each other, including, for example, a set of remote computers deployed as a cloud computing service.

[0043] The augmented reality headset 102 may include several different functional components, including a bundle adjustment module 116. The bundle adjustment module 116 may include one or more data processing devices. For example, the bundle adjustment module 116 may include one or more data processors and instructions that cause the one or more data processors to perform the operations discussed herein.

[0044] The various functional components of the augmented reality headset 102 may be installed on one or more computers as separate functional components or as different modules of the same functional component. For example, the bundle adjustment module 116 can be implemented as a computer program installed on one or more computers at one or more locations coupled together through a network. For example, in a cloud-based system, these components can be implemented by individual computing nodes of a distributed computing system.

[0045] The cameras 104, 106 of the headset 102 capture image data 110. The image data 110 and factory calibration data 112, e.g., coarse-grained extrinsic parameters, are input to a bundle adjustment module 116. The bundle adjustment module 116 performs bundle adjustment to jointly minimize the error of the headset 102 across multiple key points. The bundle adjustment module 116 applies epipolar constraints and updates the camera extrinsic parameters 120. The bundle adjustment module 116 can also generate updates to the 3D model 122, the estimated position of the headset 102, or both. The bundle adjustment module 116 can output the updated 3D model 122. The updated 3D model, or a portion of the updated 3D model, can be presented on the display of the headset 102.

[0046] The bundle adjustment module 116 can perform bundle adjustment, for example, in response to receiving data or a change in headset configuration, or at predetermined intervals, for example, in time, physical distance, or frames. For example, the bundle adjustment module 116 may perform bundle adjustment every image frame of a video, every other image frame, every third image frame, every tenth image frame, etc. In some examples, the bundle adjustment module 116 can perform bundle adjustment based on receiving image data 110 for multiple poses. In some examples, the bundle adjustment module 116 can perform bundle adjustment based on movement of the headset 102 over a threshold distance. In some examples, the bundle adjustment module 116 can perform bundle adjustment based on a change in the position of the headset from one area of ​​the environment to another area of ​​the environment. For example, the bundle adjustment module 116 may perform bundle adjustment when the headset 102 moves from one room in a physical location to another room in the physical location. The headset 102 can determine the amount of movement based, for example, on applying SLAM techniques to determine the position of the headset 102 within the environment.

[0047] In some embodiments, the bundle adjustment module 116 can perform a bundle adjustment in response to detecting an error. For example, the headset 102 may detect a deformation error between the cameras 104, 106. In response to the headset 102 detecting a deformation error above a specified threshold, the bundle adjustment module 116 can perform a bundle adjustment.

[0048] FIG. 2 depicts an example of a projection of a key point on an augmented reality headset 102. The headset 102 includes a right camera 104 and a left camera 106. The image planes of the cameras 104, 106 are illustrated as rectangles in FIG. 2. The cameras 104, 106 each capture an image of the environment depicting the key point x1. The cameras 104, 106 may capture the images of the environment simultaneously or nearly simultaneously. Although illustrated in FIG. 2 as having only two cameras, in some implementations the headset may include additional cameras.

[0049] Key point x 1206 represents a point in the environment. For example, key point x 1206 may represent a point on the edge of a piece of furniture or a point at the corner of a doorway. Key point x 1206 may be the coordinates of a corresponding two-dimensional ("2D") observation or any geometric property that is a visual descriptor of the key point.

[0050] The system can project a representation of the key point x1 206 onto the cameras 104, 106, e.g., the image plane for the cameras as represented in the 3D model. The key point x1 206 is projected onto point y1 of the left camera 106 and onto point z1 of the right camera 104. Points y1 and z2 can each represent a pixel of an individual camera.

[0051] The extrinsic parameter θ is a camera parameter that is external to the camera and can vary relative to a reference point, e.g., another camera. The extrinsic parameter θ defines the location and orientation of the camera relative to the reference point. For example, the extrinsic parameter θ in FIG. 2 indicates the position transformation between the cameras 104, 106 when one of the cameras 104, 106 is the reference point. For example, the extrinsic parameter can indicate the translational and rotational relationship between the camera 104 and the camera 106.

[0052] The headset 102 can capture images from the cameras 104, 106 at different poses along a path of headset movement. In some implementations, the device can receive first image data from the cameras 104, 106 relating to a first pose of the headset and second image data relating to a second pose of the headset. A transformation can occur between the first pose of the headset and the second pose of the headset. An example is described with reference to FIG. 3.

[0053] 3 depicts examples of tracked poses 302a-b of the augmented reality headset 102. In the example of FIG. 3, the augmented reality headset 102 moves from a first pose 302a to a second pose 302b. In the first pose 302a, the headset 102 moves in accordance with the extrinsic parameter θ a In the second pose 302b, the headset 102 has an extrinsic parameter θ b θ a and θ b The difference between indicates a deformation of the headset 102. The deformation may be caused by movement of the headset 102 from the first pose 302a to the second pose 320b, by calculated drift, or by other types of deformation, whether physical or calculated.

[0054] In the first pose 302a, the headset 102 is adjusted to the external parameter θ a The right camera 104 and the left camera 106 each have a key point x i where i represents the index of each key point. i contains the key point x1. The representation of the key point x1 is projected onto the cameras 104, 106 at the first pose 302a. The key point x1 is located at point y 1a Point z of the top and right cameras 104 1a is projected upwards.

[0055] The headset 102 moves from a first pose 302a to a second pose 302b, represented by a translation T. The translation T can include translation, rotation, or both of the headset 102. For example, a user wearing the headset can move to different locations, rotate the headset up or down, and tilt or rotate the headset between the first pose 302a and the second pose 302b. This movement of the headset 102 is represented by the translation T.

[0056] In the second pose 302b, the headset 102 adjusts the extrinsic parameter θ b θ a and θ b The difference or deformation between x and x may be caused by a translation T. In the second pose 302b, the right camera 104 and the left camera 106 are aligned with the key points x, including the key point x1. i The key point x1 is captured at point y of the left camera 106. 1b Point z of the top and right cameras 104 1b is projected upwards.

[0057] For a point to be used in calibrating two cameras, such as camera 104 and camera 106, the point should be within the field of view of both cameras. For example, a point in the field of view of camera 104 may fall along a line in the field of view of camera 106. The line can be considered the epipolar line of the point. The distance between the epipolar line in the field of view of camera 106 and its corresponding point in that field of view of camera 104 is the epipolar error.

[0058] An epipolar constraint, such as an epipolar error, can be applied to the bundle adjustment based on the image plane points y, z and the extrinsic parameter θ for the poses 302a and 302b. The epipolar constraint can be used to represent the transformation between two cameras that capture the same projected point. In this way, observation of the shared point can be used to correct the extrinsic parameter θ of the cameras.

[0059] As an example, the key point x2 is the point y of the left camera 106 at the pose 302b. 2b The key point x2 is projected from the position x2' to the point z of the right camera 104 at the pose 302b. 2b The key point x2' in the field of view of the camera 104 is projected onto the point z 2b For example, it falls along line 304 from key point x2' to key point x2'. The epipolar error can be determined from the distance between line 304 and key point x2, which is the actual location relative to key point x2'.

[0060] Epipolar error e i can be defined for each key point with index i as shown in Equation 1 below.

number

[0061] In equation 1, y i represents the projection of the key point with index i on the first camera, and z i represents the projection of the key point with index i on the second camera. The symbol E represents the fundamental matrix consisting of extrinsic parameters. The fundamental matrix E can be defined by Equation 2 below.

number

[0062] In equation 2, the symbol [ka] represents the translation matrix between the first and second orientations. The symbol R represents the rotation matrix between the first and second orientations. When the extrinsic parameters are accurate, the epipolar error e i is zero. For example, when the calibrated extrinsic parameters of a headset represent the actual relationships between the physical components of the headset, the extrinsic parameters are accurate. The epipolar error e i When is zero, Equation 3 below applies.

number

[0063] Each key point x i is the point y ia , z ia and point y at pose 302b ib , z ib During bundle adjustment, the key point x i The epipolar error for each point can be used as a constraint to refine the pose, the calibrated extrinsic parameters, the 3D model of the environment, or a combination of two or more of these. The bundle adjustment is i and the observed pixel y i , z i An exemplary process for performing bundle adjustment using epipolar constraints is described with reference to FIG.

[0064] 4 depicts an example system 400 for performing bundle adjustment using epipolar constraints. In system 400, headset 102 performs bundle adjustment using bundle adjustment module 116. In some implementations, another device or combination of devices can perform bundle adjustment using bundle adjustment module 116.

[0065] In the system 400, the right camera 104 and the left camera 106 capture images of an environment. The right camera 104 outputs image data 110a to the bundle adjustment module 116. The left camera 104 outputs image data 110b to the bundle adjustment module 116. The image data 110a and the image data 110b may be image data representing images of the environment captured at approximately the same time. The bundle adjustment module 116 receives the image data 110a, 110b as input. The bundle adjustment module may also receive a 3D model 122 of the environment as input.

[0066] The bundle adjustment module 116 can determine the reprojection error based on the image data 110a, 110b and the 3D model 122. The reprojection error r i can be expressed by Equation 4 below:

number

[0067] In equation 4, r i is the 3D point x i represents the reprojection error with respect to [ka] Based on 3D model 122, 3Dx i represents the expected projection of the point. i is the 3D point x i represents the actual projection of

[0068] In some implementations, the system 400 may include, for example, x i represents the value for the point depicted in the image captured by the right camera, and y i When y represents a value for a point depicted in the image captured by the left camera, one or more of equations (1), (3), or (4) with values ​​from different cameras can be used. For example, the system can calculate y i The value of x i One or more of these equations may be used, with the value of x substituted in. In some embodiments, the system i The value of y i One or more of these equations can be used, with values ​​of

[0069] In some implementations, the system 400 may use different equations to determine corresponding values ​​for different cameras. For example, the system may use the reprojection error r to determine the reprojection error for the right camera. iand the reprojection error r with respect to the left camera with respect to index i can be used. yi Equation 5 below can be used to determine

number

[0070] In equation 5, [ka] is the 3D point x with index i on the right camera i represents the projection of y i is the projected 3d point x i represents a 2D observation in the left image of

[0071] In some implementations, the system 400 can determine the reprojection error for both cameras. For example, the system 400 can determine the reprojection error r for the first camera, e.g., the right camera. i and the reprojection error r with respect to a second different camera, e.g., the left camera. yi can be determined.

[0072] The reprojection error can provide a measure of accuracy by quantifying how closely an estimate of a 3D model point reproduces the true projection of the point in the environment. The reprojection error is the distance between the projection of a key point found in the 3D model 122 and the corresponding projection of the key point in the real-world environment projected onto the same image.

[0073] In addition to the image data 110a, 110b, the bundle adjustment module 116 may receive as input factory calibration data 112. The factory calibration data 112 may include coarse-grained extrinsic parameters for the right camera 104 and the left camera 106.

[0074] The coarse-grained extrinsic parameters for the camera can indicate an initial translation and rotation between a particular camera and a reference position. The reference position can be, for example, a reference position on the headset 102. The reference position can be a reference camera or a reference sensor, such as a reference depth sensor, or another suitable reference position on the headset 102. The data term u for the factory calibration data can be represented by Equation 6 below:

number

[0075] In Equation 4, θ represents the factory-calibrated extrinsic parameters of the camera, and the symbol θ represents the calculated extrinsic parameters of the camera.

[0076] The bundle adjustment module 116 can also receive as input a tracked headset pose 412. The tracked headset pose 412 can be, for example, an estimated location of the first pose 302a and the second pose 302b within the environment. The tracked headset pose can include several poses along a path of movement of the headset.

[0077] The bundle adjustment module 116 can use the image data 110a, 110b, the factory calibration data 112, and the tracked headset pose 412 as inputs to the bundle adjustment process. In the bundle adjustment process, the bundle adjustment module 116 can perform a joint optimization 420 of multiple potential error sources based on the received inputs. Specifically, the bundle adjustment module can perform a joint optimization of the reprojection error, the error based on the epipolar constraint, e.g., the epipolar error, and the factory calibration error. The total error [ka] The joint optimization of these three error sources to determine σ can be expressed by Equation 7 below.

number

[0078] Each sum in Equation 7 represents a source of error summed over a set of i key points. Specifically, the first sum represents the reprojection error, the second sum represents the factory calibration error, and the third sum represents the epipolar error. Joint optimization can be performed by minimizing the combined error e(T,X,θ) across multiple key points.

[0079] By including one or more epipolar constraints in the joint optimization, the system can improve the accuracy of the error calculation. The epipolar constraint terms can rely on raw observations of pixel projections from key points to correct for deformations of the headset 102. Thus, the epipolar constraint terms can enhance the bundle adjustment module by minimizing sources of error that would otherwise be unaccounted for, ignored, or both.

[0080] The combined error e(T,X,θ) can be optimized, e.g., minimized, by inserting estimates for the rotation R and translation T and obtaining the output of Equation 7. The bundle adjustment module 116 can select the rotation and translation matrices that result in the smallest combined error as the optimized rotation and translation for all key points. In other words, the goal of bundle adjustment is to find values ​​of the rotation R and translation T that minimize the error across all translation and rotation matrices for key points in the set of key points.

[0081] In response to optimizing the error, the bundle adjustment module 116 can use the optimized rotation and translation matrices to update values, parameters, or both for the headset 102. For example, the bundle adjustment module 116 can output updates to the 3D model of the environment 416, update the calibrated extrinsic parameters 414, output updates to the position of the headset 418 within the environment, or any combination thereof.

[0082] For example, the bundle adjustment module 116 can apply updates to the 3D model of the environment 416 to the 3D model 122. Updating the 3D model can include updating the positions of one or more key points within the 3D model. The updated 3D model or a portion of the updated 3D model can then be presented on the display 422 of the headset 102.

[0083] The bundle adjustment module 116 can also update the position of a particular pose of the headset 102. Updating the position of the headset 102 can include determining the position of the headset 102 relative to the 3D model 122. In some examples, the bundle adjustment module 116 can update a series of poses or a path of poses of the headset 102 based on the bundle adjustment.

[0084] The bundle adjustment module 116 can update the extrinsic parameters for the right camera 104, the left camera 106, or both. The bundle adjustment module 116 may determine different extrinsic parameters for each camera. The bundle adjustment module 116 can output or otherwise use extrinsic parameters for multiple different poses.

[0085] The updated parameters determined by the bundle adjustment module 116, such as the updated extrinsic parameters 414, can be applied to future image data captured by the cameras 104, 106. For example, the updated extrinsic parameters can be provided to the SLAM system for future use in determining the position of the headset 102 within the environment. In some embodiments, the updated extrinsic parameters 414 may be used to overwrite the current extrinsic parameters of the headset 102. The updated extrinsic parameters 414 may be specific to each position or pose of the headset 102.

[0086] Using the updated extrinsic parameters 414, accuracy of the 3D model, the headset pose, or both can be improved. Due to bundle adjustment using epipolar constraints, the updated extrinsic parameters can be optimized across all key points for each pose of the headset 102. Therefore, minimizing the epipolar constraints during the optimization process along with the bundle adjustment problem can result in more accurate online calibration of the headset compared to other systems.

[0087] 5 is a flow diagram of a process 500 for performing bundle adjustment using epipolar constraints. For example, process 500 can be used by a device such as augmented reality headset 102 from system 100.

[0088] The device receives image data from the headset for a particular pose of the headset, the image data including (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset (502). In some examples, the device can receive images from the first and second cameras at each of a plurality of different poses along a path of movement of the headset. In some examples, the device can receive first image data from the headset for the first pose of the headset and second image data for the second pose of the headset. A transformation can occur between the first pose of the headset and the second pose of the headset.

[0089] The device can receive image data from one or more cameras. The one or more cameras can be integrated into the device or another device. In some examples, the device can receive a first portion of the image data from a first camera and a second, different portion of the image data from a second, different camera. The first image and the second image can be captured simultaneously or nearly simultaneously by the first camera and the second camera, respectively. This can enable synchronization of the first image and the second image. In some implementations, the second image can be captured within a threshold period of time from the capture of the first image.

[0090] The device identifies (504) at least one key point in a three-dimensional model of the environment represented, at least in part, in the first image and the second image. In some embodiments, the device may identify multiple key points. The device may identify the at least one key point, for example, by receiving data indicative of the at least one key point received from another device or system. The device may identify the at least one key point by performing one or more calculations to identify the key point. For example, the device may analyze the image or multiple images to determine points depicted in each of the image or multiple images. The image may be a first image or a second image. The multiple images may include a first image and a second image.

[0091] The device performs bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for at least one key point based on the first image and the second image and (ii) an epipolar error for at least one key point based on the first image and the second image (506). In some embodiments, the epipolar error represents a deformation of the headset that causes the headset to deviate from calibration.

[0092] In some embodiments, the device can identify multiple key points and jointly optimize an error across each of the multiple key points. Jointly optimizing the error across the multiple key points can include minimizing a total error across each of the multiple key points. The total error can include a combination of reprojection error and epipolar error for two or more of the multiple key points. The total error can include a factory calibration error.

[0093] In some examples, the device can perform bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for at least one key point based on the first image and the second image, (ii) an epipolar error for at least one key point based on the first image and the second image, and (iii) an error based on a factory calibration of the headset.

[0094] The device uses the results of the bundle adjustment to at least one of (i) update the three-dimensional model, (ii) determine a position of the headset in a particular pose, or (iii) determine extrinsic parameters of the first and second cameras 508. Determining the position of the particular pose can include determining a position of the headset relative to the three-dimensional model.

[0095] In some examples, the device can use the results of the bundle adjustment to update the three-dimensional model and provide an output for display by the headset based on the updated three-dimensional model. In some examples, updating the model includes updating the positions of one or more key points in the model.

[0096] In some examples, the device can determine a set of extrinsic parameters for a particular pose based on the first image and the second image. The device can use optimization with epipolar error to determine different extrinsic parameters for the first and second cameras for at least some of a plurality of different poses. The device can update a series of poses or a pose path for the headset based on the bundle adjustment.

[0097] The device can use the results of the bundle adjustment to determine the position of the headset in each of a plurality of poses, determine extrinsic parameters of the first and second cameras in each of the plurality of poses, or both. For example, the device can determine first extrinsic parameters of the first and second cameras in the first pose and second extrinsic parameters of the first and second cameras in the second pose. The first extrinsic parameters may differ from the second extrinsic parameters. In some examples, the difference between the first and second extrinsic parameters may be due to a deformation of the headset that occurred between the first and second poses.

[0098] In some examples, the extrinsic parameters include a translation and a rotation that indicate the relationship of the first camera or the second camera to a reference position on the headset. The reference position may be, for example, another camera different from the camera for which the extrinsic parameters are being determined, a depth sensor, an inertial measurement unit, or another suitable position on the headset. In some implementations, the extrinsic parameters for each camera are relative to the same reference position on the headset. In some implementations, the extrinsic parameters for at least some of the cameras are relative to different reference positions on the headset.

[0099] The device can provide updated camera extrinsic parameters for use as input to a simultaneous localization and mapping process that determines an updated environmental model of the environment and an estimated position of the device within the environment.

[0100] The order of steps in process 500 described above is illustrative only, and bundle adjustment can be performed in a different order. In some implementations, process 500 can include additional steps, fewer steps, or some of the steps can be divided into multiple steps. For example, the device can receive image data, identify at least one key point, and perform bundle adjustment without using the results of the bundle adjustment. In these examples, the device can provide the results of the bundle adjustment to another device, for example, a headset when the device is part of a system of one or more computers that performs the bundle adjustment.

[0101] Although several implementations have been described, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. For example, various forms of the flows shown above may be used, with steps reordered, added, or removed.

[0102] Embodiments and functional operations of the subject matter described herein can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more of these. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by or to control the operation of a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these.

[0103] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. The apparatus can also be or further include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Optionally, in addition to hardware, the apparatus can include code that creates an execution environment for a computer program, such as processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0104] A computer program, which may also be referred to or described as a program, software, software application, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data, e.g., in one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple cooperating files, e.g., files storing one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers, located at one site or distributed across multiple sites and interconnected by a communications network.

[0105] The processes and logic flows described herein can be implemented by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be implemented by, and apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0106] A computer suitable for executing a computer program includes, by way of example, a general-purpose or special-purpose microprocessor or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or be operatively coupled to receive data from or transfer data to, or both. However, a computer need not have such devices. A computer can also be embedded within another device, such as a mobile phone, smartphone, personal digital assistant (PDA), mobile audio or video player, game console, global positioning system (GPS) receiver, or portable storage device, such as a universal serial bus (USB) flash drive, to name but a few.

[0107] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0108] To provide for interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, e.g., an LCD (liquid crystal display), OLED (organic light-emitting diode), or other monitor, for displaying information to a user, and a keyboard and pointing device, e.g., a mouse or trackball, by which the user may provide input to the computer. Other types of devices can be used to provide interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, voice, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's device in response to a request received from the web browser.

[0109] Embodiments of the subject matter described herein can be implemented in a computing system that includes a back-end component, e.g., as a data server, or includes a middleware component, e.g., an application server, or includes a front-end component, e.g., a client computer having a graphical user interface or web browser through which a user may interact with an implementation of the subject matter described herein, or includes any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.

[0110] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on separate computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., HyperText Markup Language (HTML) pages, to user devices, e.g., for the purpose of displaying data to and receiving user input from users interacting with the user devices, acting as clients. Data generated at the user devices, e.g., results of user interactions, can be received from the user devices at the server.

[0111] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Also, while features may be described above as operative in a combination and even initially claimed as such, one or more features from a claimed combination can, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0112] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, and that all illustrated operations be performed, to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Also, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.

[0113] Specific embodiments of the present invention have been described. Other embodiments are within the scope of the following claims. For example, the steps recited in the claims, described herein, or depicted in the figures can be performed in a different order and still achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0114] Claims

Claims

1. A method implemented by one or more computers, the method comprising: receiving image data from a headset relating to a particular pose of the headset, the image data comprising (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image, (ii) an epipolar error for the at least one key point based on the first image and the second image, and (iii) a factory calibration error, wherein the factory calibration error includes a change in relative position between the first camera and a reference position on the headset; using the results of the bundle adjustment to at least one of (i) update the three-dimensional model, (ii) determine the position of the headset in the particular pose, or (iii) determine extrinsic parameters of the first camera and the second camera; A method comprising:

2. The method of claim 1 , further comprising providing an output for display by the headset based on the updated three-dimensional model.

3. The method of claim 1 , wherein the epipolar error is a result of deformation of the headset that causes the headset to deviate from calibration.

4. The method of claim 1 , comprising determining a set of extrinsic parameters for the first camera and the second camera based on the first image and the second image.

5. The method of claim 1 , wherein the extrinsic parameters include a translation and a rotation that together indicate a relationship of the first camera or the second camera to a reference position on the headset.

6. receiving images from the first and second cameras at each of a plurality of different poses along a path of movement of the headset; determining different extrinsic parameters for the first and second cameras for at least some of the different poses using results of the joint optimization of the reprojection error, the epipolar error, and the factory calibration error; The method of claim 1 , comprising:

7. identifying a plurality of key points in the three-dimensional model of the environment, the plurality of key points being at least partially represented in the first image and the second image, respectively; jointly optimizing an error across each of the plurality of key points; The method of claim 1 , comprising:

8. 8. The method of claim 7, wherein jointly optimizing the error comprises minimizing a total error across each of the plurality of key points, the total error comprising a combination of the reprojection error and the epipolar error.

9. receiving second image data from the headset relating to a plurality of poses of the headset; identifying, at least in part, at least one second key point in a three-dimensional model of the environment represented in the second image data; performing a bundle adjustment for each of the plurality of poses by jointly optimizing (i) a reprojection error for the at least one key point based on the second image data and (ii) an epipolar error for the at least one key point based on the second image data; using results of the bundle adjustment for each of the plurality of poses to at least one of (i) updating the three-dimensional model, (ii) determining another position of the headset in each of the plurality of poses, or (iii) determining other extrinsic parameters of the first camera and the second camera in each of the plurality of poses; The method of claim 1 , comprising:

10. The receiving includes receiving, from the headset, first image data relating to a first posture of the headset and second image data relating to a second posture of the headset, wherein a deformation of the headset occurs between the first posture of the headset and the second posture of the headset; said identifying includes, at least in part, identifying at least one key point in a three-dimensional model of the environment represented in the first image data and in the second image data; The jointly optimizing includes jointly optimizing (i) the reprojection error for the at least one key point based on the first image and the second image, (ii) the epipolar error for the at least one key point based on the first image and the second image, the epipolar error representing a deformation of the headset that occurs between a first pose and a second pose of the headset, and (iii) the factory calibration error; 2. The method of claim 1, wherein using the results includes using the results of the bundle adjustment to perform at least one of: (i) updating the three-dimensional model; (ii) determining a first position of the headset in the first pose or a second position of the headset in the second pose; or (iii) determining first extrinsic parameters of the first camera and the second camera in the first pose or second extrinsic parameters of the first camera and the second camera in the second pose.

11. The method of claim 1 , comprising using the results of the bundle adjustment to update a set of poses of the headset.

12. The method of claim 1 , wherein using the results of the bundle adjustment comprises updating the three-dimensional model, including updating the positions of one or more key points in the three-dimensional model.

13. The method of claim 1 , wherein using the results of the bundle adjustment includes determining a position of the headset at the particular pose relative to the three-dimensional model.

14. The method of claim 1 , wherein the first image and the second image were captured substantially simultaneously.

15. A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to: receiving image data from a headset relating to a particular pose of the headset, the image data comprising (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image, (ii) an epipolar error for the at least one key point based on the first image and the second image, and (iii) a factory calibration error, wherein the factory calibration error includes a change in relative position between the first camera and a reference position on the headset; using the results of the bundle adjustment to at least one of (i) update the three-dimensional model, (ii) determine the position of the headset in the particular pose, or (iii) determine extrinsic parameters of the first camera and the second camera; A non-transitory computer storage medium for performing operations including:

16. 16. The computer storage medium of claim 15, further comprising providing an output for display by the headset based on the updated three-dimensional model.

17. 16. The computer storage medium of claim 15, wherein the epipolar error is a result of deformation of the headset that causes the headset to deviate from calibration.

18. The computer storage medium of claim 15 , further comprising determining a set of extrinsic parameters for the first camera and the second camera based on the first image and the second image.

19. 1. A system comprising one or more computers and one or more storage devices, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to: receiving image data from a headset relating to a particular pose of the headset, the image data comprising (i) a first image from a first camera of the headset and (ii) a second image from a second camera of the headset; identifying, at least in part, at least one key point in a three-dimensional model of an environment represented in the first image and the second image; performing a bundle adjustment using the first image and the second image by jointly optimizing (i) a reprojection error for the at least one key point based on the first image and the second image, (ii) an epipolar error for the at least one key point based on the first image and the second image, and (iii) a factory calibration error, wherein the factory calibration error includes a change in relative position between the first camera and a reference position on the headset; using the results of the bundle adjustment to at least one of (i) update the three-dimensional model, (ii) determine the position of the headset in the particular pose, or (iii) determine extrinsic parameters of the first camera and the second camera; The system is operable to cause operations to be performed, including:

Citation Information

Patent Citations

  • Augmented reality display method, attitude information determination method and device

    JP2020529065A

  • Calibration of stereo cameras and handheld object

    US20180330521A1