Efficient Location Identification Based on Multiple Feature Types
The method addresses the inefficiencies of conventional pose determination in cross-reality systems by using quadratic polynomial equations and meta-variables to reduce computational burden and enhance accuracy in camera localization, facilitating efficient rendering of virtual content.
Patent Information
- Application Number
- JP2022552439
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2021-03-02
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2041-03-02
AI Technical Summary
Conventional methods for determining the pose of a camera relative to a map in machine vision systems, such as cross-reality systems, are computationally intensive and require significant effort to implement, particularly when using points and lines for location determination, leading to high computational burden and inefficiency.
A method involving the development of correspondences between images and map features, converting them into a set of quadratic polynomial equations, solving for a rotation matrix, and calculating a translation matrix using techniques like Cayley-Gibbs-Rodriguez parameterization and meta-variables to reduce computational burden and increase accuracy.
The proposed method significantly reduces computational requirements and improves accuracy in determining the camera's pose relative to a map, enabling efficient and accurate rendering of virtual content in cross-reality systems, even in large environments.
Smart Images

Figure 0007701932000084 
Figure 0007701932000085 
Figure 0007701932000086
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 085,994, filed on September 30, 2020, under Attorney Docket No. M1450.70054US01, titled "EFFICIENT LOCALIZATION BASED ON MULTIPLE FEATURE TYPES", and U.S. Provisional Patent Application No. 62 / 984,688, filed on March 3, 2020, under Attorney Docket No. M1450.70054US00, titled "POSE ESTIMATION USING POINT AND LINE CORRESPONDENCE", each of which is hereby incorporated by reference in its entirety under 35 U.S.C. § 119(e).
[0002] This application generally relates to machine vision systems such as cross - reality systems.
Background Art
[0003] Localization is performed in some machine vision systems to associate the location of a device equipped with a camera for capturing an image of a 3D environment with a location within a map of the 3D environment. New images captured by the device can be matched to a portion of the map. The spatial transformation between the new images of the matching portion of the map can indicate the "pose" of the device relative to the map.
[0004] A form of localization can be performed during map creation. The location of new images relative to existing portions of the map can enable those new images to be integrated into the map. New images can be used to expand the map, represent portions of the 3D environment that were not previously mapped, or update the representation of portions of the 3D environment that were previously mapped.
[0005] Location-specific results may be used in various ways in various machine vision systems. In a robotic system, for example, the location of a target or an obstacle may be defined relative to the coordinates of a map. Once a robotic device is located relative to the map, it may be guided towards the target along a route that avoids obstacles. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0006] Aspects of the present application relate to methods and apparatuses for providing location determination. The techniques described herein may be used together, separately, or in any suitable combination.
[0007] The inventors understand that points and lines can be used, separately or together, for location determination within a cross-reality (XR) or robotic system. Typically, the resulting problems are handled individually, and multiple algorithms, for example, algorithms related to different numbers N of correspondences (e.g., minimum problem (N = 3) and least squares problem (N>3)) and different configurations (planar and non-planar configurations), are implemented within the location determination or robotic system. The inventors understand that a lot of effort may be required to implement these algorithms.
[0008] In some aspects, location determination may be used within an XR system. In such a system, a computer may control a human user interface to create a cross-reality environment in which some or all of the XR environment is generated by the computer as perceived by the user. These XR environments may be virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) environments, in which some or all of the XR environment may be generated by the computer. The data generated by the computer may describe virtual objects that may be rendered, for example, to be perceived as part of the user's physical world so that the user may interact with the virtual objects. The user may experience these virtual objects as a result of the data being rendered through a user interface device such as a head-mounted display device that allows the user to see both virtual content and objects within the physical world simultaneously.
[0009] To realistically render virtual content, the XR system may construct a representation of the physical world around the user of the system. This representation may be constructed, for example, by processing images obtained using sensors on wearable devices that form part of the XR system. The locations of both physical and virtual objects may be represented relative to a map in which a user device within the XR system may be located. Location determination enables the user device to render virtual objects taking into account the location of physical objects. Also, multiple user devices may be enabled to render virtual content such that their individual users share the same experience of the virtual content within a 3D environment.
[0010] Conventional approaches to location - specific involve, along with a map, storing a set of feature points derived from an image of a 3D environment. The feature points may be selected for inclusion in the map based on their ease of identification and their likelihood of representing persistent objects such as corners of rooms or large furniture. Location - specific involves the steps of selecting feature points from a new image and identifying matching feature points within the map. The identification is based on the step of finding a transformation that aligns the set of feature points from the new image with the matching feature points within the map.
[0011] The step of finding a suitable transformation is computationally intensive and is often carried out by attempting to calculate a transformation that selects a group of feature points in the new image and aligns that group of feature points against each of a plurality of groups of feature points from the map. The step of attempting to calculate a transformation may use a non - linear least - squares approach, which may involve the step of calculating a Jacobean matrix, which is used to iteratively reach the transformation. This calculation may be repeated for a plurality of groups of feature points within the map and possibly one or a plurality of groups of feature points in the new image and may reach a transformation that is approved as providing a suitable match.
[0012] One or more techniques may be applied to reduce the computational burden of such matching. For example, RANSAC is a process in which the matching process is carried out in two stages. In the first stage, a rough transformation between the new image and the map can be identified based on the processing of a plurality of groups, each with a small number of feature points. The rough alignment is used as a starting point for calculating a more refined transformation that achieves a suitable alignment between larger groups of feature points.
[0013] Some aspects relate to a method for determining the pose of a camera relative to a map based on one or more images captured using the camera, where the pose is represented as a rotation matrix and a translation matrix. The method may include steps of developing correspondences between one or more images and combinations of points and / or lines within the map, converting the correspondences into a set of three quadratic polynomial equations, solving the set of equations for the rotation matrix, and calculating the translation matrix based on the rotation matrix.
[0014] In some embodiments, the combinations of points and / or lines may be dynamically determined based on characteristics of one or more images.
[0015] In some embodiments, the method may further include a step of refining the pose by minimizing a cost function.
[0016] In some embodiments, the method may further include a step of refining the pose by using a damped Newton step.
[0017] In some embodiments, the step of converting the correspondences into a set of three quadratic polynomial equations includes steps of deriving a set of constraints from the correspondences, forming a closed-form expression for the translation matrix, and forming a parameterization of the rotation matrix using 3D vectors.
[0018] In some embodiments, the step of converting the correspondences into a set of three quadratic polynomial equations further includes a step of noise removal by rank approximation.
[0019] In some embodiments, the step of solving the set of equations for the rotation matrix includes a step of using a hidden variable method.
[0020] In some embodiments, the step of forming a parameterization of a rotation matrix using a 3D vector includes the step of using a Cayley-Gibbs-Rodriguez (CGR) parameterization.
[0021] In some embodiments, the step of forming a closed-form representation of a translation matrix includes the step of forming a system of linear equations using a set of constraints.
[0022] Some aspects relate to a method for determining a camera's pose relative to a map based on one or more images captured using a camera, where the pose is represented as a rotation matrix and a translation matrix. The method may include the steps of developing a plurality of correspondences between one or more images and a combination of points and / or lines within the map, representing the correspondences as a set of over-determined systems of equations in a plurality of variables, formatting the set of over-determined systems of equations as a minimal set of equations in meta-variables, where each meta-variable represents a group of the plurality of variables, calculating values of the meta-variables based on the minimal set of equations, and calculating the pose from the meta-variables.
[0023] In some embodiments, the combination of points and / or lines may be determined dynamically based on characteristics of one or more of the images.
[0024] In some embodiments, the step of calculating the pose from the meta-variables includes the step of calculating a rotation matrix and the step of calculating a translation matrix based on the rotation matrix.
[0025] In some embodiments, the step of calculating a translation matrix based on the rotation matrix includes the step of calculating the translation matrix from equations that represent a plurality of correspondences and are linear with respect to the translation matrix, based on the rotation matrix.
[0026] In some embodiments, the step of calculating the translation matrix includes deriving a set of constraints from the correspondence, forming a closed-form expression of the translation matrix, and using the set of constraints to form a system of linear equations.
[0027] Some aspects relate to a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method. The method may include developing correspondences between one or more images and combinations of points and / or lines in a map, converting the correspondences into a set of equations of three quadratic polynomials, solving a set of equations for a rotation matrix, and calculating a translation matrix based on the rotation matrix.
[0028] In some embodiments, the points and / or lines in one or more images may be two-dimensional features, and the corresponding features in the map may be three-dimensional features.
[0029] Some aspects relate to a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method. The method may include developing a plurality of correspondences between one or more images and combinations of points and / or lines in a map, representing the correspondences as a set of overdetermined equations in a plurality of variables, formatting the set of overdetermined equations as a minimal set of equations in meta-variables, where each meta-variable represents a group of the plurality of variables, calculating values of the meta-variables based on the minimal set of equations, and calculating a pose from the meta-variables.
[0030] Some aspects relate to a portable electronic device comprising a camera configured to capture one or more images of a 3D environment and at least one processor configured to execute computer-executable instructions. The computer-executable instructions may comprise instructions for determining information about points and / or combinations of lines in one or more images of the 3D environment, transmitting the information about the points and / or combinations of lines in one or more images to a location service to determine the pose of the camera relative to a map, and receiving from the location service the pose of the camera relative to the map, represented as a rotation matrix and a translation matrix, to determine the pose of the camera relative to the map based on the one or more images.
[0031] In some embodiments, the location service is implemented on the portable electronic device.
[0032] In some embodiments, the location service is implemented on a server remote from the portable electronic device, and the information about the points and / or combinations of lines in one or more images is transmitted to the location service via a network.
[0033] In some embodiments, the step of determining the pose of the camera relative to the map comprises the steps of developing a correspondence between one or more images and combinations of points and / or lines in the map, converting the correspondence into a set of equations of three quadratic polynomials, solving the set of equations for the rotation matrix, and calculating the translation matrix based on the rotation matrix.
[0034] In some embodiments, the combinations of points and / or lines are determined dynamically based on the characteristics of one or more images.
[0035] In some embodiments, the step of determining the pose of the camera relative to the map further includes refining the pose by minimizing a cost function.
[0036] In some embodiments, the step of determining the pose of the camera relative to the map further includes refining the pose by using a damped Newton step.
[0037] In some embodiments, the step of converting the correspondence into a set of three quadratic polynomial equations includes deriving a set of constraints from the correspondence, forming a closed-form representation of the translation matrix, and forming a parameterization of the rotation matrix using 3D vectors.
[0038] In some embodiments, the step of converting the correspondence into a set of three quadratic polynomial equations further includes removing noise by rank approximation.
[0039] In some embodiments, the step of solving a set of equations for the rotation matrix includes using a hidden variable method.
[0040] In some embodiments, the step of forming a parameterization of the rotation matrix using 3D vectors includes using a Cayley-Gibbs-Rodriguez (CGR) parameterization.
[0041] In some embodiments, the step of forming a closed-form representation of the translation matrix includes forming a system of linear equations using a set of constraints.
[0042] In some embodiments, the step of determining the pose of the camera relative to the map includes the steps of developing correspondences between one or more images and combinations of points and / or lines in the map, representing the correspondences as a set of over-determined equations in a plurality of variables, formatting the set of over-determined equations as a minimal set of equations in meta-variables, where each meta-variable represents a group of the plurality of variables, calculating the values of the meta-variables based on the minimal set of equations, and calculating the pose from the meta-variables.
[0043] In some embodiments, the combinations of points and / or lines are determined dynamically based on one or more characteristics of the image.
[0044] In some embodiments, the step of calculating the pose from the meta-variables includes calculating a rotation matrix and calculating a translation matrix based on the rotation matrix.
[0045] In some embodiments, the step of calculating the translation matrix based on the rotation matrix includes calculating the translation matrix from equations that represent a plurality of correspondences based on the rotation matrix and are linear with respect to the translation matrix.
[0046] In some embodiments, the step of calculating the translation matrix includes deriving a set of constraints from the correspondences, forming a closed-form expression of the translation matrix, and forming a system of linear equations using the set of constraints.
[0047] In some embodiments, the points and lines in one or more images are two-dimensional features, and the corresponding features in the map are three-dimensional features.
[0048] Some aspects relate to a method for determining the pose of a camera relative to a map based on one or more images of a 3D environment captured by the camera, the method including: determining information about points and / or combinations of lines within one or more images of the 3D environment; transmitting the information about the points and / or combinations of lines within one or more images to a location service to determine the pose of the camera relative to the map; and receiving from the location service the pose of the camera relative to the map, represented as a rotation matrix and a translation matrix.
[0049] Some aspects relate to a non-transitory computer-readable medium comprising computer-executable instructions for execution by at least one processor, the computer-executable instructions including: determining information about points and / or combinations of lines within one or more images of a 3D environment; transmitting the information about the points and / or combinations of lines within one or more images to a location service to determine the pose of the camera relative to the map; and receiving from the location service the pose of the camera relative to the map, represented as a rotation matrix and a translation matrix, based on one or more images of the 3D environment captured by the camera.
[0050] The foregoing description is provided by way of example and not intended to be limiting. The present invention provides, for example, the following. (Item 1) A method for determining the pose of a camera with respect to a map based on one or more images captured using the camera, wherein the pose is represented as a rotation matrix and a translation matrix, and the method comprises: developing correspondences between the one or more images and combinations of points and / or lines in the map; converting the correspondences into a set of equations of three quadratic polynomials; solving the set of equations for the rotation matrix; calculating the translation matrix based on the rotation matrix and including. (Item 2) The method according to item 1, wherein the combination of points and / or lines is dynamically determined based on the characteristics of the one or more images. (Item 3) The method according to item 1, further comprising refining the pose by minimizing a cost function. (Item 4) The method according to item 1, further comprising refining the pose by using a damped Newton step. (Item 5) Converting the correspondences into a set of equations of three quadratic polynomials comprises: deriving a set of constraints from the correspondences; forming a closed-form expression of the translation matrix; forming a parameterization of the rotation matrix using 3D vectors and including. (Item 6) Converting the correspondences into a set of equations of three quadratic polynomials further comprises removing noise by rank approximation. (Item 7) Solving the set of equations for the rotation matrix comprises using a hidden variable method. (Item 8) Forming a parameterization of the rotation matrix using 3D vectors comprises using a Cayley-Gibbs-Rodriguez (CGR) parameterization. (Item 9) Forming a closed-form expression of the translation matrix comprises forming a system of linear equations using the set of constraints. (Item 10) A method for determining the pose of a camera with respect to a map based on one or more images captured using the camera, wherein the pose is represented as a rotation matrix and a translation matrix, and the method comprises: developing a plurality of correspondences between the one or more images and combinations of points and / or lines in the map; representing the correspondence as a well-determined set of equations in a plurality of variables; formatting the well-determined set of equations as a minimal set of equations of meta-variables, each of the meta-variables representing a group of the plurality of variables; calculating values of the meta-variables based on the minimal set of equations; calculating the pose from the meta-variables; A method comprising: (Item 11) The method according to item 10, wherein the combination of the points and / or lines may be dynamically determined based on the one or more image characteristics. (Item 12) Calculating the pose from the meta-variables includes: calculating the rotation matrix; calculating the translation matrix based on the rotation matrix; The method according to item 11, comprising: (Item 13) Calculating the translation matrix based on the rotation matrix includes calculating the translation matrix from an equation that represents the plurality of correspondences based on the rotation matrix and is linear with respect to the translation matrix. The method according to item 11. (Item 14) Calculating the translation matrix includes: deriving a set of constraints from the correspondences; forming a closed-form expression of the translation matrix; forming a system of linear equations using the set of constraints; The method according to item 12, comprising: (Item 15) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method, the method comprising: developing correspondences between one or more images and a combination of points and / or lines in a map; converting the correspondences into a set of equations of three quadratic polynomials; solving a set of equations regarding the rotation matrix; calculating a translation matrix based on the rotation matrix; A non-transitory computer-readable storage medium comprising: (Item 16) The points and / or lines in the one or more images are two-dimensional features, and the corresponding features in the map are three-dimensional features. The non-transitory computer-readable storage medium according to item 15. (Item 17) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method, the method comprising: Developing a plurality of correspondences between one or more images and combinations of points and / or lines in a map; Representing the correspondences as a set of overdetermined equations in a plurality of variables; Formatting the set of overdetermined equations as a minimal set of equations in meta-variables, each meta-variable representing a group of the plurality of variables; Calculating values of the meta-variables based on the minimal set of equations; Calculating the pose from the meta-variables A non-transitory computer-readable storage medium comprising the steps above. (Item 18) A portable electronic device, A camera configured to capture one or more images of a 3D environment, At least one processor configured to execute computer-executable instructions, the computer-executable instructions comprising: Determining information about combinations of points and / or lines in the one or more images of the 3D environment; Sending information about combinations of points and / or lines in the one or more images to a location service to determine a pose of the camera relative to a map; Receiving from the location service the pose of the camera relative to the map represented as a rotation matrix and a translation matrix Instructions for determining, based on the one or more images, a pose of the camera relative to a map, comprising the steps above At least one processor comprising the instructions above, and A portable device comprising the at least one processor. (Item 19) The location service is implemented on the portable electronic device of Item 18. (Item 20) The location service is implemented on a server remote from the portable electronic device, and information about combinations of points and / or lines in the one or more images is sent to the location service via a network. The portable device of Item 18. (Item 21) Determining the pose of the camera relative to the map comprises: developing a correspondence between the one or more images and a combination of points and / or lines in the map; converting the correspondence into a set of equations of three quadratic polynomials; solving a set of equations for the rotation matrix; calculating the translation matrix based on the rotation matrix; The portable device according to item 19 or 20, comprising: (Item 22) The portable device according to item 21, wherein the combination of points and / or lines is dynamically determined based on characteristics of the one or more images. (Item 23) The portable device according to item 21, wherein determining the pose of the camera relative to the map further includes refining the pose by minimizing a cost function. (Item 24) The portable device according to item 21, wherein determining the pose of the camera relative to the map further includes refining the pose by using a damped Newton step. (Item 25) Converting the correspondence into a set of equations of three quadratic polynomials includes: deriving a set of constraints from the correspondence; forming a closed-form representation of the translation matrix; forming a parameterization of the rotation matrix using 3D vectors; The portable device according to item 21, comprising: (Item 26) The portable device according to item 21, wherein converting the correspondence into a set of equations of three quadratic polynomials further includes removing noise by rank approximation. (Item 27) The portable device according to item 21, wherein solving the set of equations for the rotation matrix includes using a hidden variable method. (Item 28) The portable device according to item 25, wherein forming a parameterization of the rotation matrix using 3D vectors includes using a Cayley-Gibbs-Rodriguez (CGR) parameterization. (Item 29) The portable device according to item 25, wherein forming a closed-form representation of the translation matrix includes forming a system of linear equations using the set of constraints. (Item 30) Determining the pose of the camera relative to the map includes: developing a correspondence between the one or more images and a combination of points and / or lines in the map; representing the correspondence as a set of over-determined systems of equations in a plurality of variables; Formatting the set of dominant equations of the equation as the minimum set of equations of the meta-variables, where each of the meta-variables represents a group of the plurality of variables, and calculating values of the meta-variables based on the minimum set of equations, calculating the pose from the meta-variables The portable device according to item 19 or 20, comprising: (Item 31) The portable device according to item 30, wherein the combination of the points and / or lines is dynamically determined based on the one or more image characteristics. (Item 32) Calculating the pose from the meta-variables includes: calculating the rotation matrix, calculating the translation matrix based on the rotation matrix The portable device according to item 30, comprising: (Item 33) Calculating the translation matrix based on the rotation matrix includes calculating the translation matrix from an equation that represents the plurality of correspondences based on the rotation matrix and is linear with respect to the translation matrix. The portable device according to item 32. (Item 34) Calculating the translation matrix includes: deriving a set of constraints from the correspondences, forming a closed-form expression of the translation matrix, forming a system of linear equations using the set of constraints The portable device according to item 32, comprising: (Item 35) The points and lines in the one or more images are two-dimensional features, the corresponding features in the map are three-dimensional features, The portable device according to item 30. (Item 36) A method for determining the pose of a camera with respect to a map based on one or more images of a 3D environment captured by the camera, comprising: determining information about a combination of points and / or lines in the one or more images of the 3D environment, sending information about the combination of points and / or lines in the one or more images to a location service to determine the pose of the camera with respect to the map, receiving from the location service the pose of the camera with respect to the map represented as a rotation matrix and a translation matrix A method, comprising: (Item 37) A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising computer-executable instructions for execution by at least one processor, the computer-executable instructions determining information about points and / or combinations of lines in one or more images of a 3D environment transmitting the information about the points and / or combinations of lines in one or more images to a location service to determine a pose of a camera relative to a map receiving, from the location service, the pose of the camera relative to the map represented as a rotation matrix and a translation matrix instructions for determining a pose of a camera relative to a map based on one or more images of a 3D environment captured by the camera, the instructions including A non-transitory computer-readable medium comprising the same.
Brief Description of the Drawings
[0051] The accompanying drawings are not intended to be drawn to scale. In the drawings, each same or substantially same component illustrated in the various figures is represented by like numerals. For purposes of clarity, not all components are labeled in all of the drawings.
[0052]
Figure 1
[0053]
Figure 2
[0054]
Figure 3
[0055]
Figure 4
[0056]
Figure 5A
[0057]
Figure 5B
[0058]
Figure 6A
[0059]
Figure 6B
[0060]
Figure 7
[0061]
Figure 8
[0062]
Figure 9
[0063]
Figure 10
[0064]
Figure 11
[0065]
Figure 12
[0066]
Figure 13
[0067]
Figure 14A
[0068]
Figure 14B
[0069]
Figure 14C
[0070]
Figure 14D
[0071]
Figure 15A
[0072]
Figure 15B
[0073]
Figure 16A
[0074]
Figure 16B
[0075]
Figure 16C
[0076]
Figure 16D
[0077]
Figure 17A
[0078]
Figure 17B
[0079]
Figure 17C
[0080]
Figure 17D
[0081]
Figure 18
[0082]
Figure 19
Chemical formula
[0083]
Figure 20A
[0084]
Figure 20B
[0085]
Figure 21A
[0086]
Figure 21B
[0087]
Figure 22A
[0088]
Figure 22B
[0089]
Figure 23A
[0090]
Figure 23B
[0091]
Figure 24A
[0092]
Figure 24B
[0093]
Figure 24C
[0094]
Figure 24D
[0095]
Figure 25A
[0096]
Figure 25B
[0097]
Figure 25C
[0098]
Figure 25D
[0099]
Figure 26A
[0100]
Figure 26B
[0101]
Figure 26C
[0102]
Figure 26D
[0103]
Figure 27A
[0104]
Figure 27B
[0105]
Figure 27C
[0106]
Figure 27D
[0107]
Figure 28-1
Figure 28-2
[0108]
Figure 29A
[0109]
Figure 29B
[0110]
Figure 29C
[0111]
Figure 30
[0112]
Figure 31
[0113]
Figure 32
[0114] Detailed Description What is described in this specification is a method and apparatus for efficiently and accurately calculating the pose between a device containing a camera and the coordinate frame of other image information. The other image information may act as a map such that the step of determining the pose locates the device relative to the map. The map may represent, for example, a 3D environment. The device containing the camera may be, for example, an XR system, an autonomous vehicle, or a smartphone. The step of locating these devices relative to the map enables the device to perform location-based functions such as rendering virtual content aligned with the physical world, navigation, or rendering location-based content.
[0115] The pose may be calculated by finding correspondences between at least one set of features extracted from an image obtained using the camera and features stored within the map. The correspondences may be based, for example, on the determination that corresponding features are likely to represent the same structure within the physical world. Once corresponding features within the image and the map are identified, an attempt is made to determine a transformation that aligns the corresponding features with little or no error calculated. Such a transformation indicates the pose between the image and the reference frame of the features provided by the map. Since the image is correlated with the location of the camera at the time the image was obtained, the calculated pose also indicates the pose of the camera, and further, the device containing the camera, relative to the reference frame of the map.
[0116] The inventors recognized and appreciated the true value of an algorithm that provides a unified solution, meaning that it can be used to solve the problems resulting from one algorithm for all results, and can significantly reduce the coding effort for software architecture design regardless of whether the features are based on points, lines, or a combination of both. Further, the experimental results described in this specification show that an algorithm that provides a unified solution can achieve better or comparable performance compared to previous studies from both the perspective of accuracy and runtime.
[0117] Calculating a pose has conventionally required a large amount of computational resources, such as processing power or battery power in the case of a portable device. Any two corresponding features can provide constraints regarding the calculated pose. However, considering noise or other errors, conventionally, a set of features contains more constraints than degrees of freedom for the transformation to be calculated. Finding a solution in this case can involve the step of calculating the solution of an overdetermined system of equations. Conventional techniques for solving an overdetermined system can employ a least squares approach, which is a known iterative approach for finding a solution that provides a transformation having a low overall squared error when satisfying all the constraints.
[0118] In many practical devices, the computational burden is exacerbated by the fact that the step of finding a pose may require attempting to calculate a transformation between multiple corresponding sets of features. For example, two structures in the physical world can give rise to two sets of similar features, which may seemingly correspond. However, the calculated transformation can have a relatively high error to the extent that their seemingly corresponding features are ignored for pose calculation. The calculation can be repeated for other seemingly corresponding sets of features until a transformation is calculated with a relatively low error. Alternatively, or in addition, since a set of features in an image can seemingly correspond incorrectly to a set of features in a map, the calculated transformation cannot be approved as a solution unless there is sufficient similarity among the transformations calculated for multiple sets of features obtained from different parts of the image or different images.
[0119] Techniques as described herein can reduce the computational burden of calculating a pose. In some embodiments, the computational burden can be reduced by reformulating a well-determined set of equations into a minimal set of equations that can be solved with a lower computational burden than solving a least squares problem. The minimal set of equations can be represented from the perspective of meta-variables, each representing a group of variables within the well-determined set of equations. Once a solution is obtained with respect to the meta-variables, elements of the transformation between feature sets can be calculated from the meta-variables. The elements of the transformation can be, for example, a rotation matrix and a translation vector.
[0120] The use of meta-variables can, for example, enable the problem to be solved to be represented as a set with low-order polynomials of decimals, which can be solved more efficiently than a full least squares problem. Some or all of the polynomials can have a low degree of about 2. In some embodiments, there are at least such polynomials of about 3, which can enable a solution to be reached with a relatively low calculation.
[0121] A lower computational burden and / or increased accuracy in calculating a pose can be brought about by selecting a set of features for which the correspondence is less likely to be incorrect in that regard. The image features used to calculate the pose are often image points and represent small areas of the image. The feature points can be represented, for example, as rectangular regions with sides extending over 3 or 4 pixels of the image. For some systems, using points as features can lead to a proper solution in many scenarios. However, in other scenarios, using lines as features can be more likely to lead to a proper solution, which may require fewer trials to calculate a suitable transformation compared to using points as features. Thus, the overall computational burden can be reduced by using lines as features. Techniques as described herein can be used to efficiently calculate a pose when lines are used as features.
[0122] In some systems, an efficient solution may be more likely to result from using features that are combinations of features and lines. The number or ratio of each type of feature leading to an efficient solution can vary based on the scenario. A system configured to calculate a pose based on a corresponding set of features with an arbitrary mix of feature types can be enabled such that the mix of feature types is selected to increase the likelihood of finding a solution with a reduced computational burden from multiple trials of finding a solution. Techniques as described herein may be used to efficiently calculate a pose when an arbitrary mix of points and lines is used as features.
[0123] These techniques can be used, alone or in combination, to reduce the computational burden and / or increase the accuracy of localization, leading to more efficient or accurate operation of many types of devices. For example, during the operation of an XR system that may contain multiple components that can move relative to each other, there can be multiple scenarios where the coordinate frame of one component can be related to the coordinate frame of another component. Such a relationship that defines the relative pose of two components can be developed through a localization process. In the localization process, information represented within the coordinate frame of one component (e.g., a portable XR device) is transformed to match corresponding information represented within the coordinate frame of another component (e.g., a map). The transformation may be used to relate a location defined within the coordinate frame of one component to a location within the coordinate frame of the other, and vice versa.
[0124] The location identification techniques described in this specification may be used to provide XR scenes. The XR system thus provides useful examples of the degree of computational efficiency where pose calculation techniques can be applied in practice. To provide a realistic XR experience to multiple users, the XR system must ascertain the location of the user within the physical world in order to correctly correlate the location of virtual objects to real objects. The inventors recognized and appreciated the true value of methods and apparatuses that are computationally efficient and fast in identifying the location of XR devices even within large and very large environments (e.g., neighborhood, city, country, world).
[0125] The XR system may construct a map of the environment in which the user device operates. The environmental map may be created from image information collected using sensors that are part of the XR device worn by a user of the XR system. Each XR device may develop a local map of its physical environment by integrating information from one or more images collected as the device operates. In some embodiments, the coordinate system of the local map is tied to the location and / or orientation of the device when the device first begins to scan the physical world (e.g., starts a new session). The location and / or orientation of the device may change from session to session as the user interacts with the XR system, regardless of whether different sessions are associated with different users, each with their own wearable device with sensors that scan the environment, or the same user using the same device at different times.
[0126] The XR system may implement one or more techniques to enable persistent operations across sessions based on persistent spatial information. The techniques may, for example, enable the persistent spatial information to be created, stored, and read by any of a plurality of users of the XR system, providing XR scenes for a single or multiple users that are more computationally efficient and immersive. When shared by multiple users, the persistent spatial information provides a more immersive experience as it enables multiple users to experience virtual content at the same location with respect to the physical world. Even when used by a single user, the persistent spatial information may enable the head pose on the XR device to be quickly restored and reset in a computationally efficient manner.
[0127] The persistent spatial information may be represented by a persistent map. The persistent map may be stored in a remote storage medium (e.g., the cloud). A wearable device worn by a user may, after being turned on, read an appropriate map previously created and stored from the persistent storage device. The previously stored map may be based on data about the environment collected using sensors on the user's wearable device during a previous session. Reading the stored map may enable the use of the wearable device without completing a scan of the physical world using sensors on the wearable device. Alternatively, or in addition, the device may similarly read an appropriate stored map in response to entering a new area of the physical world.
[0128] The stored map may be represented in a standard format to which a local reference frame on each XR device can be related. In a multi-device XR system, a stored map accessed by one device may be created and stored by another device and / or may be constructed by aggregating data about the physical world collected by sensors on multiple wearable devices that pre-exist within at least a portion of the physical world represented by the stored map.
[0129] In some embodiments, persistent spatial information may be represented in a way that can be easily shared among users and among distributed components, including applications.
[0130] The standard map may provide information about the physical world and may be formatted, for example, as a Persistent Coordinate Frame (PCF). The PCF may be defined based on a set of features recognized within the physical world. The features may be selected such that they are likely to be the same for each user session of the XR system. The PCF may be sparse so that they can be efficiently processed and transferred and may provide less than all of the available information about the physical world.
[0131] Techniques for processing persistent spatial information may also include creating a dynamic map based on the local coordinate system of one or more devices. These maps may be sparse maps representing the physical world, with features such as points or edges or other structures that appear as lines detected in the images used to form the map. The standard map may be formed by merging a plurality of maps created by one or more XR devices.
[0132] The relationship between the reference map and the local map for each device may be determined through a localization process. The localization process may be performed on each XR device based on a set of reference maps that are selected and sent to the device. Alternatively, or in addition, the localization service may be provided on a remote processor such that it can be implemented within the cloud.
[0133] For example, two XR devices having access to the same stored map may both be localized with respect to the stored map. Once localized, the user device may render virtual content having a location defined by reference to the stored map by converting that location to a reference frame maintained by the user device. The user device may use this local reference frame to control the display of the user device and render the virtual content at the defined location.
[0134] The XR system may be configured to create, share, and use persistent spatial information with low usage of computational resources and / or short latency in order to provide a more immersive user experience. To support these operations, the system may use techniques for efficient comparison of spatial information. Such comparison may occur, for example, as part of localization, in which a set of features from the local device are matched to a set of features within the reference map. Similarly, in a map merge process, an attempt may be made to match one or more sets of features in a tracking map from the device to corresponding features within the reference map.
[0135] The techniques described herein may be used together or separately with many types of devices, including wearable or portable devices with limited computational resources, for many types of scenarios, to provide augmented or mixed reality scenarios. In some embodiments, the techniques may be implemented by one or more services that form part of the XR system.
[0136] AR System Overview
[0137] Figures 1 and 2 illustrate scenes with virtual content that are displayed in conjunction with a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3 - 6B illustrate an exemplary AR system that can operate in accordance with the techniques described herein and includes one or more processors, a memory, sensors, and a user interface.
[0138] Referring to FIG. 1, an outdoor AR scene 354 is depicted, and to a user of AR technology, a physical world park-like setting 356 is visible that features people, trees, buildings in the background, and a concrete platform 358. In addition to these items, a user of AR technology also "sees" a robot image 357 standing on the physical world concrete platform 358 and an avatar character 352 in the form of a flying comic that appears to be an anthropomorphic representation of a bumblebee, but these elements (e.g., avatar character 352 and robot image 357) do not exist within the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is difficult to produce AR technology that promotes a comfortable, natural, and rich presentation of virtual image elements among other virtual or physical world image elements.
[0139] Such AR scenarios can be achieved using a system that enables a user to place AR content within the physical world, determines the location within a map of the physical world where the AR content is placed, saves the AR scenario so that the placed AR content can be reloaded for display within the physical world, for example, during different AR experience sessions, and constructs a map of the physical world based on tracking information that enables multiple users to share the AR experience. The system can construct and update a digital representation of the surface of the physical world around the user. This representation may be used, in whole or in part, for placing virtual objects, in physics-based interactions, and for virtual character path planning and navigation, or for other operations where information about the physical world is used, and appears to be occluded by physical objects between the user and the rendered location of the virtual content so as to render the virtual content.
[0140] FIG. 2 depicts another example of an indoor AR scene 400 according to some embodiments and shows an exemplary use case of the XR system. The exemplary scene 400 is a living room having a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology may also perceive virtual objects such as an image on the wall behind the sofa (i.e., as in 402), a bird flying through the door (i.e., as in 404), a deer peeking out from the bookshelf, and an ornament in the form of a windmill placed on the coffee table (i.e., as in 406).
[0141] Regarding the image on the wall, AR technology requires information not only about the surface of the wall but also about objects and surfaces in the room, such as the shape of a lamp, that occlude the image in order to correctly render the virtual object. Regarding a flying bird, AR technology requires information about all objects and surfaces around the room in order to avoid objects and surfaces or, if the bird collides, to render the bird using realistic physics so that it bounces back. Regarding a deer, AR technology requires information about surfaces such as the floor or a coffee table in order to calculate where the deer should be placed. Regarding a windmill, the system can identify that it is an object separate from the table and can determine that it is movable, while the corner of a shelf or the corner of a wall can be determined to be stationary. Such specificities may be used in determining the parts of the scene that are used or updated in each of the various actions.
[0142] Virtual objects may be placed within a previous AR experience session. When a new AR experience session starts in the living room, AR technology requires that the virtual object be accurately displayed in the previously placed location and be realistically visible from different viewpoints. For example, the windmill should be displayed as standing on a book rather than floating above the table in different locations without the book. Such floating can occur if the location of the user in the new AR experience session is not accurately pinpointed within the living room. As another example, if the user is viewing the windmill from a viewpoint different from the viewpoint when the windmill was placed, AR technology requires the corresponding side of the displayed windmill.
[0143] The scene may be presented to the user via a system that includes a plurality of components, including a user interface that can stimulate one or more user perceptions such as vision, hearing, and / or touch. Additionally, the system may include one or more sensors that can measure parameters of the physical portion of the scene, including the position and / or movement of the user within the physical portion of the scene. Further, the system may include one or more computing devices with associated computer hardware such as memory. These components may be integrated within a single device or may be distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated within a wearable device.
[0144] Figure 3 is a schematic diagram 300 depicting an AR system 502 configured to provide an experience of AR content that interacts with a physical world 506, according to some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by the user as part of a headset such that the user can wear the display across their eyes, such as a pair of goggles or glasses. At least a portion of the display may be transparent such that the user can observe the see-through reality 510. The see-through reality 510 may correspond to the portion of the physical world 506 within the current viewpoint of the AR system 502, which may correspond to the user's viewpoint when the user wears a headset that incorporates both the display and sensors of the AR system and obtains information about the physical world.
[0145] The AR content may also be presented on the display 508, overlaid on the see-through reality 510. To provide an accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 may include a sensor 522 configured to capture information about the physical world 506.
[0146] Sensor 522 may include one or more depth sensors that output a depth map 512. Each depth map 512 may have a plurality of pixels that may each represent the distance to a surface within the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensor and may create a depth map. Such depth maps may be updated as fast as the depth sensor can form a new image, which may be hundreds or thousands of times per second. However, the data may be noisy and incomplete and may have holes that are shown as black pixels on the illustrated depth maps.
[0147] The system may include other sensors such as an image sensor. The image sensor may obtain monocular or stereoscopic information that may be processed to represent the physical world in other ways. For example, the image may be processed within the world reconstruction component 516 to create a mesh that represents connected portions of objects within the physical world. For example, metadata about such objects, including color and surface texture, may also be obtained using the sensors and stored as part of the world reconstruction.
[0148] The system may also obtain information about the user's head pose with respect to the physical world. In some embodiments, the system's head pose tracking component may be used to calculate the head pose in real time. The head pose tracking component may represent, for example, the user's head pose within a coordinate frame with six degrees of freedom, including translations along three perpendicular axes (e.g., forward / backward, up / down, left / right) and rotations about the three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, sensor 522 may include an inertial measurement unit that may be used to calculate and / or determine head pose 514. The head pose 514 for the depth map may indicate, for example, the current viewpoint of the sensor that captures the depth map with six degrees of freedom, but the head pose 514 may also be used for other purposes such as associating image information with a particular portion of the physical world or associating the position of a display worn on the user's head with the physical world.
[0149] In some embodiments, the head pose information may be derived by methods other than an IMU, such as from the analysis of objects in an image captured using a camera worn on the user's head. For example, the head pose tracking component may calculate the relative position and orientation of the AR device with respect to the physical object based on the visual information captured by the camera and the inertial information captured by the IMU. The head pose tracking component may then calculate the pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device with respect to the physical object with the features of the physical object. In some embodiments, the comparison may be made by identifying features in an image captured using one or more of sensors 522 that are stable over time such that changes in the position of these features in the image captured over time can be associated with changes in the user's head pose.
[0150] The inventors have realized techniques for operating an XR system to provide an XR scene for a more immersive user experience, such as estimating a head pose at a frequency of 1 kHz, with a low usage of computing resources connected to an XR device, which can be configured with, for example, four video graphic array (VGA) cameras operating at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, the computing power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and a network with a bandwidth of less than 100 Mbps. These techniques relate to reducing the processing required to generate and maintain a map and estimate a head pose, and to providing and consuming data with low computational overhead. The XR system may calculate its pose based on matched visual features. U.S. Patent Application No. 16 / 221,065, published as Application No. 2019 / 0188474, describes hybrid tracking and is incorporated herein by reference in its entirety.
[0151] In some embodiments, the AR device may construct a map from features such as points and / or lines that are recognized in successive images within a series of image frames captured as the user moves through the physical world with the AR device. Each image frame may be obtained from a different pose as the user moves, but the system may adjust the orientation of the features of each successive image frame and match the orientation of the initial image frame by matching the features of the successive image frames with the previously captured image frames. Translation of the successive image frames may be used to align each successive image frame and match the orientation of the previously processed image frames such that points and lines representing the same feature will match corresponding feature points and feature lines from the previously collected image frames. The frames within the resulting map may have a common orientation established when the first image frame is added to the map. The map may be used to determine the pose of the user within the physical world by matching features from the current image frame with the set of feature points and lines within the common reference frame. In some embodiments, the map may be referred to as a tracking map.
[0152] In addition to enabling tracking of the pose of the user within the environment, the map may enable other components of the system, such as the world reconstruction component 516, to determine the location of physical objects relative to the user. The world reconstruction component 516 may receive the depth map 512 and the head pose 514 and any other data from the sensors and may integrate that data into the reconstruction 518. The reconstruction 518 may be more complete and less noisy than the sensor data. The world reconstruction component 516 may update the reconstruction 518 using the spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0153] The reconstruction 518 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same part of the physical world or different parts of the physical world. In the illustrated embodiment, on the left side of the reconstruction 518, a part of the physical world is presented as a global surface, and on the right side of the reconstruction 518, a part of the physical world is presented as a mesh.
[0154] In some embodiments, the map maintained by the head pose component 514 may be sparse relative to other maps of the physical world that can be maintained. Instead of providing information about location and other characteristics of the surface as possibilities, the sparse map may indicate the location of interest, which can be reflected as points and / or lines in the image that result from visually distinct structures such as corners or edges. In some embodiments, the map may include image frames as captured by the sensor 522. These frames may be reduced to features that can represent the location of interest. Along with each frame, information about the user's pose from which the frame was obtained may also be stored as part of the map. In some embodiments, all images obtained by the sensor may or may not be stored. In some embodiments, the system may process the images as they are collected by the sensor and select a subset of the image frames for further calculations. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may add a new image frame to the map, for example, based on the overlap with previous image frames already added to the map or based on an image frame that contains a sufficient number of features determined to be likely to represent stationary objects. In some embodiments, the selected image frame or group of features from the selected image frame may serve as a keyframe for the map, which is used to provide spatial information.
[0155] In some embodiments, the amount of data processed when constructing a map may be reduced by constructing a sparse map with a set of mapped points and keyframes, and / or by splitting the map into blocks to enable per-block updates. The mapped points and / or lines may be associated with points of interest and / or lines in the environment. The keyframes may include information selected from camera capture data. U.S. Patent Application No. 16 / 520,582 (published as Publication No. 2020 / 0034624) describes steps for determining and / or evaluating a localization map and is hereby incorporated herein by reference in its entirety.
[0156] The AR system 502 may integrate sensor data from multiple viewpoints of the physical world over time. The pose of the sensor (e.g., position and orientation) may be tracked as the device including the sensor is moved. As the frame pose of the sensor and how it relates to other poses are understood, these multiple viewpoints of the physical world may each be fused together into a single combined reconstruction of the physical world, which may serve as an abstraction layer for the map and provide spatial information. The reconstruction may be more complete and less noisy than the original sensor data by using spatial and temporal averaging (i.e., averaging of data from multiple viewpoints over time) or any other suitable method.
[0157] In the embodiment illustrated in FIG. 3, the map represents a portion of the physical world in which a user of a single wearable device is present. In that scenario, the head pose associated with a frame in the map may be represented as a local head pose that indicates the orientation relative to the initial orientation for the single device at the start of the session. For example, the head pose may be tracked relative to the initial head pose when the device is turned on or otherwise operated to scan the environment and construct a representation of that environment.
[0158] In combination with content that characterizes that part of the physical world, the map may include metadata. The metadata may indicate, for example, the capture time of sensor information used to form the map. The metadata may alternatively or additionally indicate the location of the sensor at the capture time of the information used to form the map. The location may be represented directly, using information from a GPS chip etc., or indirectly, using a wireless (e.g., Wi-Fi) signature etc. that indicates the strength of a signal received from one or more wireless access points while the sensor data was being collected, and / or using an identifier such as the BSSID of a wireless access point to which the user device was connected while the sensor data was being collected.
[0159] The reconstruction 518 may be used for AR functions such as the production of a surface representation of the physical world for occlusion processing or physics-based processing. This surface representation may change as the user moves or as objects within the physical world change. The side of the reconstruction 518 may be used, for example, by a component 520 that produces a changing global surface representation within world coordinates that can be used by other components.
[0160] AR content may be generated based on this information by an AR application 504 etc. The AR application 504 may be, for example, a game program that implements one or more functions based on information about the physical world such as visual occlusion, physics-based interactions, and environment inference. This may be done by querying data in a different format from the reconstruction 518 produced by the world reconstruction component 516 to implement these functions. In some embodiments, the component 520 may be configured to output an update when the representation within the area of interest of the physical world changes. The area of interest may be set to approximate a part of the physical world within the vicinity of the user of the system, such as a part within the user's field of view, or projected (predicted / decided) to enter the user's field of view.
[0161] The AR application 504 may use this information to generate and update AR content. The virtual part of the AR content may be presented on the display 508 in combination with see-through reality 510 to create a realistic user experience.
[0162] In some embodiments, the AR experience may be part of a system that may include remote processing and / or remote data storage devices, a wearable display device, and / or, in some embodiments, other wearable display devices worn by other users, and may be provided to the user through an XR device. FIG. 4 illustrates, for illustrative convenience, an example of a system 580 (hereinafter referred to as the "system 580") that includes a single wearable display device. The system 580 includes a head-mounted display device 562 (hereinafter referred to as the "display device 562") and various mechanical and electronic modules and systems that support the functions of the display device 562. The display device 562 may be coupled to a frame 564, which is wearable by a user or viewer 560 (hereinafter referred to as the "user 560") of the display system and is configured to position the display device 562 in front of the eyes of the user 560. According to various embodiments, the display device 562 may be a sequential display. The display device 562 may be monocular or binocular. In some embodiments, the display device 562 may be an example of the display 508 in FIG. 3.
[0163] In some embodiments, speaker 566 is coupled to frame 564 and positioned proximate to the ear canal of user 560. In some embodiments, another speaker, not shown, is positioned adjacent to the other ear canal of user 560 to provide stereo / adjustable sound control. Display device 562 is operably coupled to local data processing module 570 by means such as a wired conductor or wireless connectivity 568, which may be mounted in a variety of configurations, such as fixed to a helmet or hat worn by user 560, fixed to a headset built into frame 564, or otherwise removably attached to user 560 (e.g., in a backpack configuration, in a belt attachment configuration).
[0164] Local data processing module 570 may include a digital memory such as a processor and non-volatile memory (e.g., flash memory), both of which may be utilized to assist in the processing, caching, and storage of data. The data includes a) data captured from sensors such as an image capture device (e.g., a camera), microphone, inertial measurement unit, accelerometer, compass, GPS unit, wireless device, and / or gyroscope (e.g., operably coupled to frame 564 or otherwise attachable to user 560), and / or b) data that may potentially be obtained and / or processed using remote processing module 572 and / or remote data repository 574 for passage to display device 562 after processing or retrieval.
[0165] In some embodiments, the wearable device may communicate with remote components. The local data processing module 570 may be operably coupled to the remote processing module 572 and the remote data repository 574 via communication links 576, 578, such as a wired or wireless communication link, respectively, such that these remote modules 572, 574 are operably coupled to each other and available as resources to the local data processing module 570. In further embodiments, in addition to or instead of the remote data repository 574, the wearable device may be able to access cloud-based remote data repositories and / or services. In some embodiments, the head pose tracking component described above may be implemented, at least in part, within the local data processing module 570. In some embodiments, the world reconstruction component 516 in FIG. 3 may be implemented, at least in part, within the local data processing module 570. For example, the local data processing module 570 may be configured to execute computer-executable instructions and generate a map and / or a physical world representation, at least in part, based on at least a portion of the data.
[0166] In some embodiments, the processing may be distributed across local and remote processors. For example, local processing may be used to construct a map (e.g., a tracking map) on the user's device based on sensor data collected using sensors on the user's device. Such a map may be used by an application on the user's device. Additionally, previously created maps (e.g., reference maps) may be stored in a remote data repository 574. If a suitable stored or persistent map is available, it may be used instead of, or in addition to, a tracking map created locally on the device. In some embodiments, the tracking map may be geolocated with respect to a stored map such that the correspondence can be oriented with respect to the position of the wearable device at the time the user turned the system on, and with respect to one or more persistent features, and the reference map can be oriented. In some embodiments, the persistent map may be loaded onto the user's device and enable the rendering of virtual content without the latency associated with scanning the location where the user's device constructs a tracking map of the user's complete environment from sensor data obtained during the scan. In some embodiments, the user's device may access a remote persistent map (e.g., stored in the cloud) without having to download the persistent map onto the user's device.
[0167] In some embodiments, spatial information may be communicated from the wearable device to a remote service such as a cloud service configured to identify the device and store it in a map maintained on the cloud service. According to one embodiment, the location identification process may occur within the cloud, returning a transformation that matches the device location to an existing map such as a reference map and links virtual content to the wearable device location. In such embodiments, the system can avoid communicating the map from the remote resource to the wearable device. Other embodiments are configured for both device-based and cloud-based location identification and can enable functionality, for example, when network connectivity is unavailable or the user chooses not to enable cloud-based location identification.
[0168] Alternatively, or in addition, the tracking map may be merged with previously stored maps to extend those maps or improve their quality. The process for determining whether a suitable previously created environmental map is available and / or for merging the tracking map with one or more stored environmental maps may be performed within the local data processing module 570 or the remote processing module 572.
[0169] In some embodiments, the local data processing module 570 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of the local data processing module 570 but enable a smaller device. In some embodiments, the world reconstruction component 516 may use a computational budget less than that of a single advanced RISC machine (ARM) core so that the remaining computational budget of the single ARM core can be accessed for other uses such as mesh extraction, etc., and generate a physical world representation in real time over an unspecified space.
[0170] In some embodiments, the remote data repository 574 may include a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, enabling fully autonomous use from the remote module. In some embodiments, all data is stored and all or most of the computations are performed within the remote data repository 574, enabling a smaller device. The world reconstruction may be stored, for example, wholly or partially, within this repository 574.
[0171] In an embodiment where data is remotely stored and accessible via a network, the data may be shared by multiple users of the augmented reality system. For example, the user device may upload its tracking map and expand it into a database of environmental maps. In some embodiments, the upload of the tracking map occurs at the end of a user session with the wearable device. In some embodiments, the upload of the tracking map may occur persistently, semi-persistently, intermittently, at a predefined time, after a predefined period from the previous upload, or when triggered by an event. The tracking map uploaded by any user device may be used to extend or improve a previously stored map, regardless of whether it is based on data from that user device or any other user device. Similarly, the persistent map downloaded to the user device may be based on data from that user device or any other user device. Thus, a high-quality environmental map may be readily available to the user to improve the experience using the AR system.
[0172] In a further embodiment, the download of the persistent map may be limited and / or avoided based on a localization performed on a remote resource (e.g., in the cloud). In such a configuration, the wearable device or other XR device communicates to the cloud service feature information (e.g., positioning information regarding the device at the time when a feature represented in the feature information is sensed) combined with the pose information. One or more components of the cloud service may match the feature information with an individual stored map (e.g., a reference map) and generate a transformation between the coordinate systems of the tracking map maintained by the XR device and the reference map. Each XR device having the tracking map localized with respect to the reference map may accurately render virtual content at a location defined with respect to the reference map based on its own tracking.
[0173] In some embodiments, the local data processing module 570 is operably coupled to the battery 582. In some embodiments, the battery 582 is a removable power source such as a commercially available battery. In other embodiments, the battery 582 is a lithium-ion battery. In some embodiments, the battery 582 enables the user 560 to operate the system 580 for longer time periods without having to connect to a power source, charge a lithium-ion battery, or shut down the system 580 and replace the battery, including both an internal lithium-ion battery chargeable by the user 560 during non-operation of the system 580 and a removable battery.
[0174] FIG. 5A illustrates a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as “environment 532”). Information captured by the AR system along the user's movement path may be processed into one or more tracking maps. The user 530 positions the AR display system at location 534, and the AR display system records ambient information about the passable world for location 534 (e.g., a digital representation of real objects in the physical world that can be memorized and updated as real objects in the physical world change). That information may be stored as a pose in combination with an image, feature, directional audio input, or other desired data. Location 534 is aggregated, for example, as part of a tracking map, for data input 536 and processed at least by the passable world module 538, which may be implemented, for example, by processing on the remote processing module 572 of FIG. 4. In some embodiments, the passable world module 538 may include a head pose component 514 and a world reconstruction component 516 such that the processed information can indicate the location of an object in the physical world in combination with other information about the physical object used when rendering virtual content.
[0175] The passable world module 538 determines, at least in part, where and how the AR content 540 can be placed within the physical world, as determined from the data input 536. The AR content is "placed" within the physical world by presenting both a representation of the physical world and the AR content via a user interface, and the AR content is rendered as if it were interacting with objects within the physical world, and the objects within the physical world are presented as if the AR content were obscuring the user's view of those objects when appropriate. In some embodiments, the AR content may be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from the reconstruction (e.g., reconstruction 518) and determining the shape and position of the AR content 540. As an example, the fixed element may be a table, and the virtual content may be positioned to appear on that table. In some embodiments, the AR content may be placed within a structure within the field of view 544, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be persisted with respect to a model 546 (e.g., a mesh) of the physical world.
[0176] As described, the fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element in the physical world that can be stored within the passable world module 538 such that the user 530 can perceive content on the fixed element 542 each time it is visible to the user 530 without the system having to map to the fixed element 542. The fixed element 542 may thus be a mesh model that is stored by the passable world module 538 for future reference by multiple users, even though it is determined from a previous modeling session or from a different user. Thus, the passable world module 538 recognizes the environment 532 from a previously mapped environment and can display AR content without the user 530's device first mapping all or part of the environment 532, saving computational processes and cycles and avoiding latency for any rendered AR content.
[0177] The mesh model 546 of the physical world may be created by the AR display system, interact with the AR content 540, and the appropriate surfaces and metrics for display can be stored by the passable world module 538 for future retrieval by the user 530 or other users without having to recreate the model completely or partially. In some embodiments, the data input 536 provides the passable world module 538 with input such as the geographical location, user identification, and current activity as to which of one or more fixed elements 542 are available, the AR content 540 last placed on the fixed element 542, and whether the same content should be displayed (such AR content being "persistent" content regardless of whether the user is viewing a particular passable world model).
[0178] Even in embodiments where an object is considered to be fixed (e.g., a kitchen table), the passable world module 538 may update those objects in the physical world model at any time to account for possible changes in the physical world. The models of fixed objects may be updated very infrequently. Other objects in the physical world may be considered to be moving or otherwise not fixed (e.g., a kitchen chair). To render the AR scene with a realistic feel, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update the fixed objects. To enable accurate tracking of all objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.
[0179] FIG. 5B is a schematic illustration of the viewing optics assembly 548 and associated components. In some embodiments, two eye-tracking cameras 550 are directed toward the user's eyes 549 to detect metrics of the user's eyes 549, such as eye shape, eyelid occlusion, pupil direction, and glints on the user's eyes 549.
[0180] In some embodiments, one of the sensors is a depth sensor 551, such as a time-of-flight sensor, that emits signals into the world and detects the reflections of those signals from neighboring objects to determine the distance to a given object. The depth sensor may, for example, quickly determine whether an object has entered the user's field of view as a result of either the movement of those objects or a change in the user's pose. However, information about the positions of objects within the user's field of view may alternatively or additionally be collected using other sensors. Depth information may be obtained, for example, from a stereoscopic image sensor or a plenoptic sensor.
[0181] In some embodiments, the world camera 552 records, maps, and / or otherwise creates a model of the environment 532 and detects inputs that can affect the AR content, capturing a view wider than the periphery. In some embodiments, the world camera 552 and / or the camera 553 may be grayscale and / or color image sensors, which may output grayscale and / or color image frames at a fixed time interval. The camera 553 may further capture a physical world image within the user's field of view at a specific time. The pixels of the frame-based image sensor may be sampled iteratively even if their values are invariant. The world camera 552, the camera 553, and the depth sensor 551 each have individual fields of view 554, 555, and 556, respectively, and collect and record data from a physical world scene such as the physical world environment 532 depicted in FIG. 34A.
[0182] The inertial measurement unit 557 may determine the movement and orientation of the visual optics assembly 548. In some embodiments, the inertial measurement unit 557 may provide an output indicating the direction of gravity. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 551 is operably coupled to the eye tracking camera 550 as a confirmation of the measured focus adjustment for the actual distance being viewed by the user's eye 549.
[0183] The visual optical system assembly 548 may include some of the components illustrated in FIG. 34B, and it should be understood that it may include components instead of or in addition to the illustrated components. In some embodiments, for example, the visual optical system assembly 548 may include two world cameras 552 instead of four. Alternatively, or in addition, cameras 552 and 553 need not capture visible light images of their full fields of view. The visual optical system assembly 548 may include other types of components. In some embodiments, the visual optical system assembly 548 may include one or more dynamic vision sensors (DVSs), the pixels of which may respond asynchronously to relative changes in light intensity exceeding a threshold.
[0184] In some embodiments, the visual optical system assembly 548 may not include a depth sensor 551 based on time-of-flight information. In some embodiments, for example, the visual optical system assembly 548 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of incident light, from which depth information may be determined. For example, a plenoptic camera may include an image sensor overlaid with a transmissive diffraction mask (TDM).
[0185] Alternatively, or in addition, the plenoptic camera may include an image sensor containing angle-sensing pixels and / or phase-detection autofocus pixels (PDAFs) and / or a microlens array (MLA). Such sensors may serve as a depth information source instead of or in addition to the depth sensor 551.
[0186] Also, it should be understood that the component configuration in FIG. 5B is provided as an example. The visual optics assembly 548 may include components with any suitable configuration, which may be set to provide the user with a practical maximum field of view for a particular set of components. For example, if the visual optics assembly 548 has one world camera 552, the world camera may be installed within the central region of the visual optics assembly instead of on the side.
[0187] Information from sensors within the visual optics assembly 548 may be coupled to one or more than one of the processors within the system. The processor may generate data that can be rendered to make the user perceive the virtual content interacting with objects in the physical world. That rendering may be implemented in any suitable manner, including the step of generating image data depicting both physical and virtual objects. In other embodiments, the physical and virtual content may be depicted in one scene by modulating the opacity of a display device through which the user sees through the physical world. The opacity may be controlled to create the appearance of the virtual object and block the view of objects in the physical world that are occluded by the virtual object from the user. In some embodiments, the image data may be modified to be perceived by the user as interacting realistically with the physical world when the virtual content is viewed through the user interface (e.g., clipping the content and taking occlusion into account), and may include only the virtual content.
[0188] The location on the visual optics assembly 548 where content can be displayed to create an impression of an object at a particular location may depend on the physics of the visual optics assembly. Additionally, the user's head pose and the direction the user's eyes are looking with respect to the physical world will affect the location within the physical world content that will be displayed at a particular location on the visual optics assembly where the content will appear. Sensors such as those described above may supply information from which a processor receiving the sensor input can collect and / or calculate this information so that the object can be rendered on the visual optics assembly 548 and the location where the desired appearance for the user should be created can be calculated.
[0189] Regardless of how the content is presented to the user, a model of the physical world may be used so that the characteristics of virtual objects, including the shape, position, motion, and visibility of the virtual objects, can be correctly calculated, which may be affected by physical objects. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 518.
[0190] The model may be created from data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected by multiple users, which may be aggregated within a computing device remote from all users (and may be "in the cloud").
[0191] The model may be created, at least in part, by a world reconstruction system, such as world reconstruction component 516 of FIG. 3, described in further detail in FIG. 6A. The world reconstruction component 516 may include a perception module 660 that can generate, update, and store a representation for a portion of the physical world. In some embodiments, the perception module 660 may represent a portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world, includes surface information, and may indicate whether a surface exists within the volume represented by the voxel. The voxels may be assigned a value indicating whether the corresponding volume has been determined to contain the surface of a physical object, to be empty, or not yet measured using the sensor, and thus whether its value is unknown. It should be understood that the value indicating a voxel determined to be empty or unknown need not be explicitly stored, and the voxel values may be stored in computer memory in any suitable manner, including not storing information regarding voxels determined to be empty or unknown.
[0192] In addition to generating information for the persistent world representation, the perception module 660 may identify and output an indication of a change in the area surrounding the user of the AR system. Such an indication of a change may trigger other functions, such as triggering an update to the volumetric data stored as part of the persistent world, or triggering component 604 to generate and update AR content.
[0193] In some embodiments, the perception module 660 may identify changes based on a signed distance function (SDF) model. The perception module 660 may be configured to receive sensor data such as, for example, a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a may directly provide SDF information, and an image may be processed to arrive at SDF information. The SDF information represents distances from sensors used to capture that information. Since those sensors may be part of a wearable unit, the SDF information may represent the physical world from the perspective of the wearable unit and thus the user's perspective. The head pose 660b may enable the SDF information to be associated with voxels within the physical world.
[0194] In some embodiments, the perception module 660 may generate, update, and store a representation for a portion of the physical world that is within the perception range. The perception range may be determined at least in part based on the reconstruction range of the sensors, which may be determined at least in part based on the limits of the observation range of the sensors. As a specific example, an active depth sensor that operates using active IR pulses may reliably operate over a certain range of distances and may create an observation range of the sensors that can be several centimeters or tens of centimeters to several meters.
[0195] The world reconstruction component 516 may include additional modules that may interact with the perception module 660. In some embodiments, the persistent world module 662 may receive a representation of the physical world based on data obtained by the perception module 660. The persistent world module 662 may also include representations of the physical world in various formats. For example, the module may include stereo information 662a. For example, volumetric metadata such as voxels 662b may be stored along with a mesh 662c and a plane 662d. In some embodiments, other information such as a depth map may also be stored.
[0196] In some embodiments, a representation of the physical world, such as that illustrated in FIG. 6A, may provide relatively dense information about the physical world as compared to a sparse map, such as a tracking map based on feature points and / or lines, as described above.
[0197] In some embodiments, the perception module 660 may include modules that generate representations for the physical world in various formats, such as, for example, a mesh 660d, a plane, and a semantics 660e. Representations for the physical world may be stored across local and remote memory media. Representations for the physical world may be described in different coordinate frames, for example, depending on the location of the memory media. For example, a representation for the physical world stored within a device may be described in a coordinate frame local to the device. Representations for the physical world may have counterparts stored within the cloud. The counterparts within the cloud may be described in a coordinate frame shared by all devices within the XR system.
[0198] In some embodiments, these modules may generate a representation based on data within the perception range of one or more sensors at the time the representation is generated, data captured at previous times, and information within the persistent world module 662. In some embodiments, these components may act on depth information captured using a depth sensor. However, the AR system may include a vision sensor and may generate such a representation by analyzing monocular or binocular vision information.
[0199] In some embodiments, these modules may act on regions of the physical world. Those modules may be triggered to update a sub-region of the physical world when the perception module 660 detects a change in the physical world within that sub-region. Such changes may be detected, for example, by detecting a new surface within the SDF model 660c or by other criteria, such as a change in the values of a sufficient number of voxels representing the sub-region.
[0200] The world reconstruction component 516 may include a component 664 that can receive a representation of the physical world from the perception module 660. The component 664 may include a visual occlusion 664a, a physics-based interaction 664b, and / or an environment inference 664c. Information about the physical world may be pulled by these components according to, for example, usage requests from an application. In some embodiments, the information may be pushed to the consuming component via, for example, an indication of a change in a pre-identified area or a change in the representation of the physical world within the perception range. The component 664 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interaction, and environment inference.
[0201] In response to a query from the component 664, the perception module 660 may send a representation of the physical world in one or more formats. For example, when the component 664 indicates that the usage is for visual occlusion or physics-based interaction, the perception module 660 may send a representation of the surface. When the component 664 indicates that the usage is for environment inference, the perception module 660 may send a mesh, plane, and semantics of the physical world.
[0202] In some embodiments, the perception module 660 may include a component that provides format information to the component 664. An example of such a component may be a raycasting component 660f. The consuming component (e.g., the component 664) may query for information about the physical world from a particular perspective, for example. The raycasting component 660f may select from one or more representations of the physical world data within the field of view from that perspective.
[0203] In some embodiments, the components of the passable world model may be distributed, with some parts being executed locally on the XR device and some parts being executed remotely, such as on a network connected to a server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud can affect the functionality of the XR system and the user experience. For example, by distributing processing to the cloud, reducing the processing on the local device can enable a longer battery life and reduce the heat generated on the local device. However, distributing much more processing to the cloud can create unacceptable wait times that cause an undesirable user experience.
[0204] FIG. 6B depicts a distributed component architecture 600 configured for spatial computing, according to some embodiments. The distributed component architecture 600 may include a passable world component 602 (e.g., PW538 in FIG. 5A), a Lumin OS 604, an API 606, an SDK 608, and an application 610. The Lumin OS 604 may include a Linux-based kernel with custom drivers compatible with the XR device. The API 606 may include an application programming interface that provides XR applications (e.g., application 610) access to the spatial computing features of the XR device. The SDK 608 may include a software development kit that enables the creation of XR applications.
[0205] One or more components within architecture 600 may create and maintain a model of possible worlds. In this example, sensor data is collected on a local device. Processing of that sensor data may be performed, at least in part, locally on the XR device and, at least in part, within the cloud. PW538 may include an environmental map that is created, at least in part, based on data captured by an AR device worn by a plurality of users. During a session of an AR experience, an individual AR device (such as the wearable device described above in connection with FIG. 4) may create a tracking map, which is one type of map.
[0206] In some embodiments, the device may include a component that constructs both a sparse map and a dense map. The tracking map may serve as the sparse map. The dense map may include surface information, which may be represented by a mesh or depth information. Alternatively, or in addition, the dense map may include a higher level of information derived from surface or depth information such as the location and / or characteristics of planes and / or other objects.
[0207] The sparse map and / or the dense map may persist for reuse by the same device and / or for sharing with other devices. Such persistence may be achieved by storing the information in the cloud. The AR device may send the tracking map to the cloud and merge it, for example, with an environmental map selected from previously stored persistent maps in the cloud. In some embodiments, the selected persistent map may be sent from the cloud to the AR device for merging. In some embodiments, the persistent maps may be oriented with respect to one or more persistent coordinate frames. Such maps may serve as reference maps since they can be used by any of a plurality of devices. In some embodiments, a model of the traversable world may consist of, or be created from, one or more reference maps. The device may perform some operations based on a coordinate frame local to the device, but may use the reference map by determining the transformation between that coordinate frame local to the device and the reference map.
[0208] The reference map may arise as a tracking map (TM). The tracking map may be persisted, for example, such that the reference frame of the tracking map becomes a persistent coordinate frame. Thereafter, once a device accessing the reference map determines the transformation between its local coordinate system and the coordinate system of the reference map, it may use the information in the reference map to determine the location of the objects represented in the reference map within the physical world surrounding the device.
[0209] Therefore, the reference map, the tracking map, or other maps may have a similar format, but for example, the places where they are used or stored are different. FIG. 7 depicts an exemplary tracking map 700 according to some embodiments. In this example, the tracking map represents the features of interest as points. In other embodiments, lines may be used instead of or in addition to points. The tracking map 700 may provide a floor plan 706 of the physical objects in the corresponding physical world represented by the points 702. In some embodiments, the map points 702 may represent the features of a physical object that may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. The features may be derived from a processed image such as obtained using sensors of a wearable device in an augmented reality system. The features may be derived, for example, by processing an image frame output by a sensor and identifying the features based on large gradients or other suitable criteria in the image. Further processing may limit the number of features in each frame. For example, the processing may select features that are likely to represent persistent objects. One or more heuristics may be applied for this selection.
[0210] The tracking map 700 may include data regarding the points 702 collected by the device. For each image frame with data points added to the tracking map, the pose may be stored. The pose may represent the orientation from which the image frame was captured such that the feature points in each image frame can be spatially correlated to the tracking map. The pose may be determined by positioning information such as can be derived from sensors such as IMU sensors on the wearable device. Alternatively, or in addition, the pose may be determined by matching a subset of the features in the image frame to features already in the tracking map. A transformation between the matching subsets of features may be calculated, which indicates the relative pose between the image frame and the tracking map.
[0211] Since much of the information collected using the sensor is likely to be redundant, not all of the feature points and image frames collected by the device can be retained as part of the tracking map. In some embodiments, a relatively small subset of features from the image frame may be processed. Those features can be clearly different, such as arising from sharp corners or edges. Additionally, only features from a certain frame may be added to the map. Those frames may be selected based on one or more criteria such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality metric for the features within the frame. Image frames not added to the tracking map may be discarded or used to revise the location of features. As a further alternative, data from multiple image frames, represented as a set of features, may be retained, but only features from a subset of those frames may be designated as key frames, which are used for further processing.
[0212] The key frame may be processed to produce a key rig 704. The key frame may be processed to produce a three-dimensional set of feature points and stored as the key rig 704. Such processing may involve, for example, comparing image frames derived simultaneously from two cameras and stereoscopically determining the 3D position of feature points. Metadata such as pose may be associated with these key frames and / or key rigs. The key rig may subsequently be used when localizing the device with respect to the map based on newly acquired images from the device.
[0213] The environmental map may have any of a plurality of formats, depending on, for example, the storage location of the environmental map, including, for example, the local storage device and the remote storage device of the AR device. For example, the map in the remote storage device may have a higher resolution than the map in the local storage device on the wearable device when the memory is limited. To transmit the higher resolution map from the remote storage device to the local storage device, the map may be downsampled or otherwise converted to an appropriate format, such as by reducing the number of poses per area of the physical world stored in the map and / or the number of feature points stored per pose. In some embodiments, a slice or portion of the high resolution map from the remote storage device may be transmitted to the local storage device, and the slice or portion is not downsampled.
[0214] The database of environmental maps may be updated as new tracking maps are created. To determine which of potentially a very large number of environmental maps in the database should be updated, the updating step may include efficiently selecting one or more environmental maps stored in the database that are related to the new tracking map. The one or more selected environmental maps may be ranked by relevance, and one or more of the highest ranked maps may be selected for processing to merge the higher ranked selected environmental maps with the new tracking map to create one or more updated environmental maps. When the new tracking map represents a portion of the physical world for which there is no existing environmental map to update across, the tracking map may be stored in the database as a new environmental map.
[0215] Remote Positioning
[0216] Various embodiments can utilize remote resources to facilitate a persistent and consistent cross-reality experience among individual users and / or groups of users. The advantages of the operation of an XR device using a canonical map as described herein can be achieved without downloading a set of canonical maps. This advantage can be achieved, for example, by transmitting feature and pose information to a remote service that maintains a set of canonical maps. A device that requires virtual content to be positioned at locations defined relative to a canonical map may receive from the remote service one or more conversions between the features and the canonical map. Those conversions are used to position virtual content at one or more locations defined relative to a canonical map on the device, while maintaining information about the location of those features within the physical world, or alternatively, to identify locations within the physical world defined relative to the canonical map.
[0217] In some embodiments, spatial information is captured by an XR device and communicated to a remote service such as a cloud-based service, which uses the spatial information to localize the XR device relative to a canonical map used by an application or other component of the XR system and to define the location of virtual content relative to the physical world. Once localized, a conversion that links the tracking map maintained by the device to the canonical map can be communicated to the device.
[0218] In some embodiments, a camera and / or a portable electronic device comprising a camera may be configured to capture and / or determine information about features (e.g., a combination of points and / or lines) and transmit the information to a remote service such as a cloud-based device. The remote service may use the information to determine the pose of the camera. The pose of the camera may be determined, for example, using the methods and techniques described herein. In some examples, the pose may include a rotation matrix and / or a translation matrix. In some examples, the pose of the camera may be represented relative to any of the maps described herein.
[0219] The transformation may be used, in conjunction with the tracking map, to determine the position at which virtual content defined relative to the reference map should be rendered, or alternatively, to identify a location within the physical world defined relative to the reference map.
[0220] In some embodiments, the result returned to the device from the location service may be one or more transformations that associate the uploaded features with a portion of the reference map. Those transformations may be used within the XR device, in conjunction with the tracking map, to identify the location of virtual content, or alternatively, to identify a location within the physical world. In embodiments where persistent spatial information such as PCF is used to define locations relative to the reference map, the location service may download to the device the transformation between the features and one or more PCFs after location success.
[0221] In some embodiments, the location service may further return the pose of the camera to the device. In some embodiments, the result returned to the device from the location service may associate the pose of the camera relative to the reference map.
[0222] As a result, the network bandwidth consumed by communication between the XR device and the remote service for performing location determination can be reduced. The system can thus support frequent location determination and enable each device interacting with the system to quickly obtain information for positioning virtual content or performing other location-based functions. As the device moves within the physical environment, it may repeatedly request updated location determination information. Additionally, the device may frequently obtain updates to the location determination information, expand the map, or increase its accuracy, such as through the merging of additional tracking maps when the reference map changes.
[0223] FIG. 8 is a schematic diagram of an XR system 6100. During a user session, the user device that displays cross-reality content can appear in various forms. For example, the user device can be a wearable XR device (e.g., 6102) or a handheld mobile device (e.g., 6104). As discussed above, these devices can be configured with software such as applications or other components and / or be wired-connected and can generate local location information (e.g., a tracking map) that can be used to render virtual content on their individual displays.
[0224] Virtual content positioning information may be defined relative to global location information, which may be formatted, for example, as a reference map containing one or more persistent coordinate frames (PCFs). A PCF may be a set of features within the map that can be used to locate the map. A PCF may be selected, for example, based on a process that identifies the set of features as being easily recognizable and likely to persist across user sessions. According to some embodiments, such as the embodiment shown in FIG. 8, system 6100 is configured with a cloud-based service that supports the functionality of virtual content and its display on a user device, where the location is defined relative to a PCF within the reference map.
[0225] In one example, the positioning function is provided as a cloud-based service 6106. The cloud-based service 6106 may be implemented on any of a plurality of computing devices, from which computing resources may be allocated to one or more services running within the cloud. Those computing devices may be interconnected to each other and to devices such as wearable XR device 6102 and handheld device 6104 in an accessible manner. Such connections may be provided via one or more networks.
[0226] In some embodiments, the cloud-based service 6106 is configured to receive descriptor information from individual user devices and "locate" them relative to a reference map or maps that match the devices. For example, the cloud-based positioning service matches the received descriptor information to descriptor information regarding an individual reference map. The reference map may be created using techniques as described above, by merging maps provided by one or more devices having image sensors or other sensors that obtain information about the physical world.
[0227] However, it is not a requirement that the reference map be created by the devices accessing them, and thus the map may be created by a map developer who can publish them, for example, by making the map available to a location service 6106.
[0228] FIG. 9 is an exemplary process flow that can be executed by a device to use a cloud-based service to identify the location of a device using a reference map and receive conversion information that defines one or more conversions between the device local coordinate system and the coordinate system of the reference map.
[0229] According to one embodiment, process 6200 can start at 6202 using a new session. Starting a new session on the device can start the capture of image information and construct a tracking map for the device. Additionally, the device may send a message, register with the server of the location service, and prompt the server to create a session for that device.
[0230] Once a new session is established, process 6200 can continue at 6204 to capture a new frame of the device's environment. Each frame can be processed at 6206 to select from the frames in which features are captured. The features may be of one or more types such as feature points and / or feature lines.
[0231] Feature extraction in 6206 may include adding pose information to the features extracted in 6206. The pose information may be the pose within the local coordinate system of the device. In some embodiments, the pose may be with respect to a reference point in the tracking map, which may be with respect to the origin of the tracking map of the device. Regardless of the format, the pose information may be added to each feature or each set of features such that the location service may use the pose information for calculating a transformation that may be returned to the device in response to matching the features to features in the stored map.
[0232] Process 6200 may continue to decision block 6207, where a decision is made as to whether to request location. In some embodiments, location accuracy is improved by performing location for each of a plurality of image frames. Location is considered successful only when there is sufficient correspondence between the results calculated for a sufficient number of the plurality of image frames. Thus, a location request may be sent only when sufficient data can be captured to achieve location success.
[0233] One or more criteria may be applied to determine whether to request location. The criteria may include the passage of time such that the device may request location after a certain threshold amount of time. For example, if location has not been attempted within a certain threshold amount of time, the process may continue from decision block 6207 to action 6208, where location is requested from the cloud. The threshold amount of time may be, for example, between 10 and 30 seconds such as 25 seconds. Alternatively, or in addition, location may be triggered by the movement of the device. The device executing process 6200 may use the IMU and / or its tracking map to track its movement and initiate location in response to detecting movement that exceeds a threshold distance from where the device last requested location. The threshold distance may be, for example, between 1 and 10 meters such as 3 to 5 meters.
[0234] Regardless of how the location determination is triggered, once triggered, process 6200 may proceed to act 6208, where the device sends a request for the location service that includes data used by the location service to perform the location determination. In some embodiments, data from multiple image frames may be provided for the location determination attempt. The location service may not consider a location determination successful, for example, unless the features in the multiple image frames yield consistent location determination results. In some embodiments, process 6200 may include storing a set of features and the added pose information in a buffer. The buffer may be, for example, a circular buffer that stores a set of features extracted from the most recently captured frame. Thus, the location determination request may be sent along with some sets of features accumulated in the buffer.
[0235] The device may transfer the contents of the buffer to the location service as part of the location determination request. Other information may also be transmitted, along with the feature points and the added pose information. For example, in some embodiments, geographic information may be transmitted, which may assist in selecting the map against which to attempt the location determination. The geographic information may include, for example, GPS coordinates or a wireless signature associated with a device tracking map or the current persistent pose.
[0236] In response to the request sent at 6208, the cloud location service may process the set of features and locate the device in a reference map or other persistent map maintained by the service. For example, a cloud-based location service may generate a transformation based on the pose of the set of features transmitted from the device relative to the matching features in the reference map. The location service may return the transformation to the device as the location determination result. This result may be received at block 6210.
[0237] Regardless of how the transformation is formatted, in act 6212, the device may use these transformations to calculate, for the virtual content, the location where it should be rendered with respect to any of the PCFs as defined by an application or other component of the XR system. This information may alternatively or additionally be used on the device to perform any location-based operations where the location is defined based on the PCF.
[0238] In some scenarios, the location service may not be able to match the features sent from the device to any stored reference map, or may not be able to match a sufficient number of sets of features communicated with the request for the location service to consider that a location determination success has occurred. In such scenarios, rather than returning the transformation to the device as described above in connection with act 6210, the location service may indicate to the device that the location determination has failed. In such scenarios, process 6200 may branch in decision block 6209 to act 6230, and the device may take one or more actions for failure handling. These actions may include increasing the size of a buffer that holds the set of features transmitted for location determination. For example, if the location service does not consider a location determination success unless three sets of features match, the buffer size may be increased from five to six, increasing the likelihood that three of the sets of features transmitted may be matched to the reference map maintained by the location service.
[0239] In some embodiments, the reference map maintained by the location service may contain PCFs that are pre-identified and memorized. Each PCF may be represented by a plurality of features that may include a mixture of feature points and feature lines for each image frame processed at 6206. Thus, the location service may identify the reference map using a set of features that match the set of features transmitted with the location request, and may calculate the transformation between the coordinate frame represented by the pose transmitted with the request for location and one or more PCFs.
[0240] In the illustrated embodiment, the location result may be represented as a transformation that aligns the coordinate frame of the set of extracted features to the selected map. This transformation may be returned to the user device, where it may be applied as either a forward or inverse transformation to relate the locations defined relative to the shared map to the coordinate frame used by the user device, or vice versa. The transformation may, for example, enable the device to render virtual content for its user at locations relative to the physical world defined within the coordinate frame of the map to which the device is located.
[0241] Pose Estimation Using 2D / 3D Point and Line Correspondences
[0242] The pose of the set of features relative to other image information can be calculated in many scenarios, including XR systems, to localize the device relative to the map. FIG. 10 illustrates a method 1000 that can be implemented to calculate such a pose. In this example, method 1000 calculates the pose for any mixture of feature types. The features may be, for example, all feature points or all feature lines or a combination of feature points and feature lines. Method 1000 may be implemented, for example, as part of the process illustrated in FIG. 9, where the pose calculated therein is used to localize the device relative to the map.
[0243] The processing for method 1000 may begin once an image frame is captured for processing. In block 1010, a mixture of feature types may be determined. In some embodiments, the extracted features may be points and / or lines. In some embodiments, the device may be configured to select a certain mixture of feature types. The device may be programmed, for example, to select a set percentage of features as points and the remaining features as lines. Alternatively, or in addition, the pre-configuration may be based on ensuring at least a certain number of points and a certain number of lines within the set of features from the image.
[0244] Such a selection may be guided, for example, by one or more metrics indicative of the likelihood that the feature will be recognized in subsequent images of the same scene. Such metrics may be based, for example, on the characteristics of the physical structure giving rise to such features and / or locations within the physical environment. The corners of a photo frame mounted on a window or wall, for example, may result in feature points with high scores. As another example, the corners of a room or the edges of a staircase may result in feature points with high scores. Such metrics may be used to select the best features within an image, or may be used to select an image for which further processing is performed, with the further processing being performed, for example, only on images with a number of features exceeding a threshold with high scores.
[0245] In some embodiments, the selection of features may be made such that the same number or mix of points and lines are selected for all images. Image frames that do not supply a defined mix of features may be discarded, for example. In other scenarios, the selection may be dynamic, based on the visual characteristics of the physical environment. The selection may be guided, for example, based on the magnitude of the metric assigned to the detected features. For example, in a small room with a monochromatic wall and few furnishings, there may be few physical structures that give rise to feature points with large metrics. FIG. 11 illustrates, for example, an environment in which location trials based on feature points are likely to fail. Similar results can occur in environments with structures that give rise to a large number of similar feature points. In those environments, the mix of selected features may include more lines than points. Conversely, in large or outdoor spaces, there may be many structures that give rise to feature points with few straight edges, such that the mix of features will be biased towards points.
[0246] In block 1020, the determined mix of features may be extracted from the image frame and processed. It should be understood that blocks 1010 and 1020 need not be performed in the order shown, and the processing may be dynamic such that the processes of selecting features and determining the mix can occur in parallel. Techniques for processing an image and identifying points and / or lines may be applied in block 1020 to extract features. Additionally, one or more criteria may be applied to limit the number of features extracted. The criteria may include the total number of features included within the set of extracted features or a quality metric regarding the features.
[0247] The process may then proceed to block 1030, where a correspondence between the features extracted from the image and other image information such as a previously stored map is determined. The correspondence may be determined, for example, based on visual similarity and / or descriptor information associated with the features. These correspondences may be used to generate a set of constraints on the transformation that defines the pose of the extracted features relative to the features from the other image information. In a localization embodiment, these correspondences are between a selected set of features in the image captured using the camera on the device and the stored map.
[0248] In some embodiments, the image used as input for pose estimation is a 2D image. Thus, the image features are 2D. The other image information may represent the features in 3D. For example, a key rig as described above may have 3D features constructed from a plurality of 2D images. Correspondences can still be determined despite the different dimensions. FIG. 12 illustrates, for example, that correspondences can be determined by projecting the 3D features into the 2D plane of the image from which the 2D features were extracted.
[0249] Regardless of the manner in which the set of features is extracted therein, the process proceeds to block 1040, where the pose is calculated. This pose may, for example, serve as a result of a localization attempt in an XR system as described above.
[0250] According to some embodiments, any step of method 1000 may be performed on the device described herein and / or on a remote service such as those described herein.
[0251] In some embodiments, the processing at block 1040 may be selected based on a mixture of feature types extracted from the image frame. In other embodiments, the processing may be general purpose such that the same software can be executed, for example, for an arbitrary mixture of points and lines.
[0252] The step of estimating the camera pose using 2D / 3D point or line correspondences, called the PnPL problem, is a fundamental problem in computer vision with many applications such as Simultaneous Localization and Mapping (SLAM), Structure from Motion (SfM), and Augmented Reality. The PnPL algorithms described herein can be complete, robust, and efficient. Here, a "complete" algorithm can mean that the algorithm can handle all potential inputs regardless of the mixture of feature types so that the same process can be applied in any scenario and can be applied in any scenario.
[0253] According to some embodiments, general processing may be achieved by programming the system to calculate the pose from a set of correspondences by converting a least squares problem into a minimization problem.
[0254] Conventional methods for solving the PnPL problem do not provide a complete algorithm that is as accurate and efficient as the individual customized solutions for each problem. The inventors recognize that by using one algorithm to solve multiple problems, the effort in algorithm implementation can be significantly reduced.
[0255] According to some embodiments, the localization method may include the step of using a complete, accurate, and efficient solution to the PnPL problem. According to some embodiments, the method may also be able to solve the PnP and PnL problems as specific cases of the PnPL problem. In some embodiments, the method may be able to solve multiple multiple types of problems including minimization problems (e.g., P3L, P3P, and / or PnL) and / or least squares problems (e.g., PnL, PnP, PnPL). For example, the method may be able to solve any of the P3L, P3P, PnL, PnP, and PnPL problems. In the literature, there are custom solutions for each problem, but in practice, it is too laborious to implement a specific solution for each problem.
[0256] FIG. 13 is an example of a process that can be general and can lead to the conversion of a problem that can be conventionally solved as a least squares problem to a minimal problem. FIG. 13 is a flowchart illustrating a method 1300 for efficient pose estimation according to some embodiments. The method 1300 may be implemented, for example, in block 1030 in FIG. 10, for example, in correspondence. The method may start with the step of obtaining 2×(m + n) constraints (act 1310) on the premise of n 2D / 3D point correspondences and m 2D / 3D line correspondences.
[0257] The method 1300 may include the step of restructuring a set of constraints (act 1320) and the step of obtaining an equation system using a partial linearization method. The method may further include the step of solving the equation system and obtaining a rotation matrix (act 1330), and the step of obtaining t, which is a translation vector, using a closed form of the rotation matrix and t (act 1340). Both the rotation matrix and the translation vector can define a pose. According to some embodiments, any step of the method 1300 may be implemented on the device described herein and / or on a remote service such as those described herein.
[0258] Integrated solution for pose estimation using 2D / 3D points and line correspondences
[0259] According to some embodiments, solving the PnPL problem means estimating the camera pose (i.e., R and t) using N 2D / 3D point correspondences (i.e.,
Chem.
Chem.
[0260] In an exemplary embodiment of the method of 1300, the following notation may also be used: [ka] and M 2D / 3D line correspondences [ka] and estimating the camera pose (i.e., R and t) using P i =[x i ,y i ,z i ] T may represent a 3D point, and p i =[u i ,v i ] T may represent the corresponding 2D pixel in the image. i can represent 3D lines, and l i can represent the corresponding 2D line. i 1 and Q i2 may be used to represent L i and may be used for two pixels q i 1 and q i 2 to represent l i For simplicity of notation, we use normalized pixel coordinates.
[0261] According to some embodiments, in act 1310, the step of obtaining 2×(m + n) constraints, assuming n 2D / 3D point correspondences and m 2D / 3D line correspondences, may include the step of using the point correspondences, where the i-th 2D / 3D point correspondence
Chem.
Math.
[0262] According to some embodiments, in act 1310 of method 1300, the step of obtaining 2×(m + n) constraints further includes the step of multiplying both sides of the equation by the denominator in (1), resulting in the following.
Math.
Chem.
Number
Number
Number
[0263] According to some embodiments, in act 1320 of method 1300, the step of reconfiguring the set of constraints may include the step of generating a quadratic system using the constraints, which is an expression of R using Cayley-Gibbs-Rodriguez parameterization and a closed form of t.
[0264] The M = 2×(n + m) constraints as (4) are obtained assuming n 2D / 3D point correspondences and m 2D / 3D line correspondences. For the i-th constraint, the following may be defined.
Number
Number
[0265] (7) is linear with respect to t, so the closed form of t can be described as follows.
Number
Number
[0266] The solution for R can then be determined. A three-dimensional vector s, which is the Cayley-Gibbs-Rodriguez (CGR) parameterization, may be used to represent R as follows.
Number
Chemistry
[0267] Substituting (10) into (9) and expanding (6), the resulting system is as follows.
Number
Chemistry
Number
Number
Chemistry
Chemistry
Chemistry
Number
[0268] Lemma 1: The rank of \(H\) is less than 9 for data without noise. Proof: Equation (13) is a homogeneous linear system. \(r\) with nine elements is a non-trivial solution of (13). Therefore, \(H\) should be singular; otherwise, this homogeneous system would have only the zero (or trivial) solution, which contradicts the fact that \(r\) is a solution of (13). Theorem 1: The rank of \(A\) in (11) is less than 9 for data without noise. Proof: The CGR representation in (10), \(r\) in (13), and
Chem.
Math.
Chem.
[0269] According to some embodiments, rank approximation may be used for noise removal. Matrix \(A\) may have rank degradation. In some embodiments, generally,
Chem.
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
[0270] According to some embodiments, in act 1320 of method 1300, the step of obtaining an equation system using a partial linearization method may include the step of using a partial linearization method to convert a PnPL problem into an essential minimum formula (EMF), and the step of generating an equation system. In some embodiments, the partial linearization method may
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Math.
Chem.
Chem.
Chem.
Chem.
Math.
[0271] Equation (17) may be rewritten as follows.
Math.
Math.
[0272] According to some embodiments, the step of solving the system of equations and obtaining the rotation matrix (act 1330) may include obtaining the rotation matrix by solving a system of equations where the equations are in the form of (19). According to some embodiments, the step of obtaining t using the closed form of the rotation matrix and t (act 1340) may include obtaining t from (8) after solving for s.
[0273] Exemplary results
[0274] 14-17 are schematic views of experimental results of embodiments of the efficient localization method compared to other known PnPL solvers. FIGS. 14A-14D show the average and median rotation and translation errors of different PnPL solvers, including OPnPL and evxpnpl, as described in "Accurate and linear time pose estimation from points and lines: European Conference on Computer Vision", Alexander Vakhitov, Jan Funke, and Francesc Moreno Noguer, Springer, 2016 and "CvxPnPL: A unified convex solution to the absolute pose estimation problem from point and line correspondences" by Agostinho, Sergio, Joao Gomes, and Alessio Del Bue, 2019 (both incorporated herein by reference in their entirety).
[0275] FIG. 14A shows the median rotation error of different PnPL algorithms in degrees. FIG. 14B shows the median translation error of different PnPL algorithms in percentage. FIG. 14C shows the average rotation error of different PnPL algorithms in degrees. FIG. 14D shows the average translation error of different PnPL algorithms in percentage. In FIGS. 14A-D, the pnpl curves 40100A-D show the errors in rotation and translation using the method described herein according to some embodiments. The OPnPL curves 40200A-D and the cvxpnpl curves 40300A-D show errors in percentage and degrees that are consistently higher than those of the pnpl curves 40100.
[0276] FIG. 15A is a schematic diagram of the computation time of different PnPL algorithms. FIG. 15B is a schematic diagram of the computation time of different PnPL algorithms. The computation time to solve the PnPL problem using the method described herein is represented by 50100A-B, and the OPnPL curves 50200A-B and the cvxpnpl curves 50300A-B show consistently higher computation times than the methods including embodiments of the algorithms described herein.
[0277] FIG. 16A shows the number of instances of a range of errors versus the logarithmic error of the PnPL solution compared to the P3P and UPnP solutions for the PnP problem according to some embodiments described herein.
[0278] FIG. 16B shows a box plot of the PnPL solution compared to the P3P and UPnP solutions for the PnP problem according to some embodiments described herein.
[0279] Figure 16C shows the average rotational error in radians of the PnPL solution for the PnP problem, compared to the P3P and UPnP solutions, according to some embodiments described herein. The PnPL solution for the PnP problem, according to some embodiments described herein, has an error 60100C, which can be seen to be less than the error for the UPnP solution 60200C.
[0280] Figure 16D shows the average positional error in meters of the PnPL solution for the PnP problem, compared to the P3P and UPnP solutions, according to some embodiments described herein. The PnPL solution for the PnP problem, according to some embodiments described herein, has an error 60100D, which can be seen to be less than the error for the UPnP solution 60200D.
[0281] Figures 17A-D show the average and median rotation and translation errors of different PnL algorithms, including OAPnL, DLT, LPnL, Ansar, Mirzaei, OPnPL, and ASPnL. OAPnL is described in "A Robust and Efficient Algorithm for the PnL problem Using Algebraic Distance to Approximate the Reprojection Distance," by Zhou, Lipu, et al., 2019, which is hereby incorporated by reference in its entirety. DLT is described in “Absolute pose estimation from line correspondences using direct linear transformation. Computer Vision and Image Understanding” by Pibyl, B., Zemk, P., and Adk, M., 2017, which is hereby incorporated by reference in its entirety. LPnL is described in “Pose estimation from line correspondences: A complete analysis and a series of solutions” by Xu, C., Zhang, L., Cheng, L., and Koch, R., 2017, which is hereby incorporated by reference in its entirety. Ansar is described in “Linear pose estimation from points or lines” by Ansar, A., and Daniilidis, K., 2003, which is hereby incorporated by reference in its entirety. Mirzaei is described in “Globally optimal pose estimation from line correspondences” by Mirzaei, F. M., and Roumeliotis, S. I., 2011, which is hereby incorporated by reference in its entirety.As described herein, OPnPL is addressed in "Accurate and linear time pose estimation from points and lines: European Conference on Computer Vision". As described herein, aspects of ASPnL are described in "Pose estimation from line correspondences: A complete analysis and a series of solutions".
[0282] FIG. 17A shows the median rotation error of different PnL algorithms in degrees. FIG. 17B shows the median translation error of different PnL algorithms in percentage. FIG. 17C shows the mean rotation error of different PnL algorithms in degrees. FIG. 17D shows the mean translation error of different PnL algorithms in percentage. Curves 70100A-D show the median and mean rotation and translation errors of the PnPL solution using the method described herein.
[0283] Pose Estimation Using Feature Lines
[0284] In some embodiments, instead of or in addition to the general approach, an efficient process may be applied to calculate the pose when only lines are selected as features. FIG. 18 illustrates a method 1800 that is an alternative to method 1000 in FIG. 10. Similar to method 1000, method 1800 may start with steps of determining a feature mixture in blocks 1810 and 1820 and extracting features using that mixture. In the process in block 1810, the feature mixture may include only lines. For example, only lines may be selected in the environment as illustrated in FIG. 11.
[0285] Similarly, in block 1830, the correspondence may be determined as described above. From these correspondences, the pose may be calculated in sub-process 1835. In this embodiment, the process may branch depending on whether the feature includes at least one point. If so, the pose may be estimated using a technique capable of solving the pose based on a set of features that includes at least one point. A general algorithm as described above may be applied, for example, in box 1830.
[0286] Conversely, if the set of features includes only lines, the process may be performed by an algorithm that delivers accurate and efficient results in that case. In this embodiment, the process branches to block 3000. Block 3000 may solve the perspective-n-line (PnL) problem as described below. Since lines often exist and can serve as easily recognizable features, specifically providing a solution for the feature set using only lines in an environment where pose estimation may be desired can provide an efficiency or accuracy advantage for a device operating in such an environment.
[0287] According to some embodiments, any step of method 1800 may be performed on the device described herein and / or on a remote service such as those described herein.
[0288] As described herein, in the special case of the PnPL problem, it includes the perspective-n-line (PnL) problem, and the pose of the camera can be estimated from several 2D / 3D line correspondences. The PnL problem can be described as a line correspondence of the PnP problem, as described in "A direct least-squares (dls) method for pnp" by Hesch, J.A., Roumeliotis, S.I., International Conference on Computer Vision, "Upnp: An optimal o (n) solution to the absolute pose problem with universal applicability. In: European Conference on Computer Vision." by Kneip, L., Li, H., Seo, Y., "Revisiting the pnp problem: A fast, general and optimal solution" In: Proceedings of the IEEE by Kuang, Y., Sugimoto, S., Astrom, K., Okutomi, M., all of which are hereby incorporated by reference in their entirety.
[0289] The PnL problem is a fundamental problem in computer vision and robotics with many applications, including simultaneous localization and mapping (SLAM), structure from motion (SfM), and augmented reality (AR). Generally, the camera pose can be determined from several N 2D-3D line correspondences, where N ≥ 3. When the number N of line correspondences is 3, the problem can be called a minimal problem, also known as the P3L problem. When the number N of correspondences is greater than 3, the problem can be known as a least-squares problem. The minimal problem (e.g., N = 3) and the least-squares problem (e.g., N > 3) are generally solved in different ways. Solutions for both the minimal and least-squares problems play important roles in various robot and computer vision tasks. Due to its importance, much effort has been made to solve both problems.
[0290] Regarding the PnL problem, the conventional methods and algorithms proposed generally use different algorithms to solve the minimization problem (P3L problem) and the least squares problem. For example, in a conventional system, the minimization problem is formulated as a system of equations, while the least squares problem is formulated as a minimization problem. By upgrading the minimization problem to a least squares problem, theoretically, other least squares solutions that can handle the minimum case result in inefficient minimum solutions, and since the minimum solution is required to be launched multiple times in the RANSAC framework, it is impractical for use in real-time applications (e.g., as described in "Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography” by Fischler, M.A., Bolles, R.C., which is incorporated herein by reference in its entirety).
[0291] Other conventional systems that address the least squares problem as the minimum solution are also inefficient for use in real-time applications. The solution to the minimization problem generally leads to an eighth-degree polynomial, while in this document, the least squares problem solution, described as the Generalized Minimum Formula (GMF), requires solving a more complex system of equations.
[0292] By dealing with the least squares as the minimum solution, the conventional system is inefficient when solving the minimum solution by dealing with a more complex system of equations than required for the least squares solution. For example, the Mirzaei algorithm (as described, for example, in 'Optimal estimation of vanishing points in a Manhattan world. In: 2011 International Conference on Computer Vision' by Mirzaei, F.M., Roumeliotis, S.I., which is incorporated herein by reference in its entirety) requires finding the roots of three fifth-degree polynomial equations, and the algorithm described in 'A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance' results in an equation of a 27th-degree univariate polynomial, as described in 'Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence' (which is incorporated herein by reference in its entirety), and 'Camera pose estimation from lines: a fast, robust and general method. Machine Vision and Applications 30(4), 603{614 (2019)' by Wang, P., Xu, G., Cheng, Y., Yu, Q. (which is incorporated herein by reference in its entirety) proposes a subset-based solution, which requires solving an equation of a 15th-degree univariate polynomial.
[0293] As described herein, the minimal (P3L) problem generally requires solving an eighth-order univariate equation and thus has up to eight solutions, except for some specific geometric configurations (e.g., as described in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence" by Xu, C., Zhang, L., Cheng, L., Koch, R.). One widely adopted strategy for the minimal (P3L) problem is to simplify the problem by means of some geometric transformations (e.g., as described in "Determination of the attitude of 3d objects from a single perspective view. IEEE transactions on pattern analysis and machine intelligence", "Pose determination from line-to-plane correspondences: existence condition and closed-form solutions. IEEE Transactions on Pattern Analysis & Machine Intelligence" Chen, H.H., "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence", "Camera pose estimation from lines: a fast, robust and general method. Machine Vision and Applications 30" by Wang, P., Xu, G., Cheng, Y., Yu, Q.).
[0294] Specifically, aspects of the cited references discuss several specific intermediate coordinate systems for reducing the number of unknowns that result in univariate equations. The problem with these methods is that the transformation can involve some numerically unstable operations with respect to a certain configuration, such as the denominator of the fraction in equation (4) of "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence" by Xu, C., Zhang, L., Cheng, L., Koch, R., which can be a very small value. In the aspect of "A stable algebraic camera pose estimation for minimal configurations of 2d / 3d point and line correspondences. In: Asian Conference on Computer Vision" by Zhou, L., Ye, J., Kaess, M., quaternions are used to parameterize rotations and an algebraic solution for the P3L problem is introduced.Some studies have focused on specific configurations of the P3L problem, such as three lines forming a Z shape (e.g., as described in "A new method for pose estimation from line correspondences. Acta Automatica Sinica" 2008, by Li-Juan, Q., Feng, Z., which is incorporated herein by reference in its entirety), or the planar three-line junction problem (e.g., as described in ’The planar three-line junction perspective problem with application to the recognition of polygonal patterns. Pattern recognition 26(11), 1603{1618 (1993)’ by Caglioti, V., which is incorporated herein by reference in its entirety), or the P3L problem with a known vertical direction (e.g., as described in ’Camera pose estimation based on pnl with a known vertical direction. IEEE Robotics and Automation Letters 4(4), 3852{3859 (2019)’ by Lecrosnier, L., Boutteau, R., Vasseur, P., Savatier, X., Fraundorfer, F., which is incorporated herein by reference in its entirety).
[0295] Initial research on solutions to the least squares PnL problem mainly focused on the error function formula and iterative solutions. Liu et al. (’Determination of camera location from 2-d to 3-d line and point correspondences. IEEE Transactions on pattern analysis and machine intelligence 12(1), 28{37 (1990)’ by Liu, Y., Huang, T.S., Faugeras, O.D., incorporated herein by reference in its entirety) studied the constraints from 2D-3D point and line correspondences and separated the estimation of rotation and translation. Kumar and Hanson (’Robust methods for estimating pose and a sensitivity analysis. CVGIP: Image understanding 60(3), 313{342 (1994)’ by Kumar, R., Hanson, A.R., incorporated herein by reference in its entirety) proposed optimizing both rotation and translation in an iterative method. They presented a sampling-based method to obtain an initial estimate. Later research (e.g., ’Pose estimation using point and line correspondences. Real-Time Imaging 5(3), 215{230 (1999)’ by Dornaika, F., Garcia, C. and as described in Iterative pose computation from line correspondences (1999), both incorporated herein by reference in their entirety) proposed starting the iteration from poses estimated by weak perspective or pseudo-perspective camera models. The accuracy of the iterative algorithm depends on the quality of the initial solution and the parameters of the iterative algorithm. There is no guarantee that the iterative method will converge.As for most 3D vision problems, linear formulas play an important role (e.g., as described in 'Multiple view geometry in computer vision. Cambridge university press (2003)' by Hartley, R., Zisserman, A., which is incorporated herein by reference in its entirety). Direct Linear Transformation (DLT) provides a simple method for calculating poses (e.g., as described in 'Multiple view geometry in computer vision. Cambridge university press (2003)' by Hartley, R., Zisserman, A.). This method requires at least six line correspondences. Pribyl et al. (e.g., as described in 'Camera pose estimation from lines using pln ucker coordinates. arXiv preprint arXiv:1608.02824 (2016)' by Pribyl, B., Zemcik, P., Cadik, M.) introduced a new DLT method based on the plucker coordinates of 3D lines, which requires at least nine lines. In subsequent research (e.g., as described in 'Absolute pose estimation from line correspondences using direct linear transformation. Computer Vision and Image Understanding 161, 130{144 (2017)' by Pribyl, B., Zemcik, P., Cadik, M.), they combined two DLT methods, which showed improved performance and reduced the minimum number of line correspondences to five.By exploring the similarities between the constraints derived from the PnP and PnL problems, the EPnP algorithm is extended to solve the PnL problem (e.g., as described in ’Accurate and linear time pose estimation from points and lines. In: European Conference on Computer Vision. pp. 583{599. Springer (2016)’ and “Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence” by Xu, C., Zhang, L., Cheng, L., Koch, R.). The EPnP-based PnL algorithm is applicable for N = 4, but it is unstable and requires specific processing for the planar PnL problem (i.e., all lines lie in a plane) when N is small. The linear formula ignores the unknown constraints. This leads to less accurate results and limits its usability. To solve the above problems, a method based on the polynomial formula has been proposed. Ansar et al. (’Linear pose estimation from points or lines. IEEE Transactions on Pattern Analysis and Machine Intelligence 25(5), 578{589 (2003)’ by Ansar, A., Daniilidis, K.) adopted a quadratic system to represent the constraints and presented a linearization approach to solve this system. Their algorithm is applicable for N ≧ 4, but it is too slow when N is large.Motivated by the RPnP algorithm, a subset-based PnL approach was proposed in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence" by Xu, C., Zhang, L., Cheng, L., Koch, R. and "Camera pose estimation from lines: a fast, robust and general method. Machine Vision and Applications 30". They divide N line correspondences into N - 2 triplets, and each triplet is a P3L problem. Then they minimize the sum of quadratic polynomials derived from each P3L problem. The subset-based PnL approach will be time-consuming when N is large, as shown in "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance" (incorporated herein by reference in its entirety). Using Grobner basis techniques (as described in, for example, 'Using algebraic geometry, vol. 185. Springer Science & Business Media (2006)' by Cox, D.A., Little, J., O’shea, D., incorporated herein by reference in its entirety), it is possible to directly solve the polynomial system. This leads to a series of direct minimization methods.In the literature, CGR (e.g., as described in 'Optimal estimation of vanishing points in a manhattan world. In: 2011 International Conference on Computer Vision. pp. 2454{ 2461. IEEE (2011)' by Mirzaei, F.M., Roumeliotis, S.I. and 'Globally optimal pose estimation from line correspondences. In: 2011 IEEE International Conference on Robotics and Automation. pp. 5581{5588. IEEE (2011)' by Mirzaei, F.M., Roumeliotis, S.I.), which is incorporated herein by reference in its entirety) and quaternions (e.g., as described in 'Accurate and linear time pose estimation from points and lines. In: European Conference on Computer Vision. pp. 583{599. Springer (2016)' by Vakhitov, A., Funke, J., Moreno-Noguer, F., which is incorporated herein by reference in its entirety) have been adopted to parameterize rotation, which has led to polynomial cost functions. Then, Grobner basis techniques are used to solve the first optimality condition of the cost function. The Grobner basis technique can encounter numerical problems (for example, as described in ’Using algebraic geometry, vol. 185. Springer Science & Business Media (2006)’ by Cox, D.A., Little, J., O’shea, D. and ’Fast and stable polynomial equation solving and its application to computer vision. International Journal of Computer Vision 84(3), 237{256 (2009)’ by Byrod, M., Josephson, K., Astrom, K., which are hereby incorporated by reference in their entirety), as described in “A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance”, Zhou et al. introduced a hidden variable polynomial solver. They showed improved accuracy but were still significantly slower than most of the linear formula-based algorithms. The PnL problem has several extensions for certain applications. Some applications involve multiple cameras. Lee (for example, as described in ’A minimal solution for non-perspective pose estimation from line correspondences. In: European Conference on Computer Vision. pp. 170{185. Springer (2016)’ by Lee, G.H., which is hereby incorporated by reference in its entirety) proposed a closed-form P3L solution for multi-camera systems. Recently, Hichem (for example, as described in its entirety by reference As described in 'A direct least-squares solution to multi-view absolute and relative pose from 2d-3d perspective line pairs. In: Proceedings of the IEEE International Conference on Computer Vision Workshops (2019)' by Abdellali, H., Frohlich, R., Kato, Z., which is incorporated herein by reference in its entirety, a direct least-squares solution to the PnL problem of a multi-camera system was proposed. In some applications, the vertical direction is obtained from a sensor (e.g., an IMU). This can be used as a prior value for pose estimation (e.g., as described in 'Camera pose estimation based on pnl with a known vertical direction. IEEE Robotics and Automation Letters 4(4), 3852{3859 (2019)' and 'Absolute and relative pose estimation of a multi-view camera system using 2d-3d line pairs and vertical direction. In: 2018 Digital Image Computing: Techniques and Applications (DICTA). pp. 1{8. IEEE (2018)' by Abdellali, H., Kato, Zm, which are incorporated herein by reference in their entireties).The PnL solution for a single camera can be extended to a multi-camera system (e.g., as described in ’A direct least-squares solution to multi-view absolute and relative pose from 2d-3d perspective line pairs. In: Proceedings of the IEEE International Conference on Computer Vision Workshops (2019)’), so this paper focuses on the PnL problem for a single camera.
[0296] A desirable PnL solution is that it is accurate and efficient for any possible input considered. As described above, algorithms based on linear formulas are generally unstable or infeasible for small N, require specific processing, or even do not function for the case of a plane. On the other hand, algorithms based on polynomial formulas can achieve better accuracy and are applicable to a wider range of PnL inputs, but are more computationally demanding. Furthermore, they lack an integrated solution for minimum and least-squares problems. Therefore, conventionally, there has been a significant room for improvement over state-of-the-art PnL solutions such as those provided by the techniques of this specification.
[0297] According to some embodiments, the localization method may include a complete, accurate, and efficient solution to the perspective-n-line (PnL) problem. In some embodiments, the least-squares problem may be transformed into a general minimum formula (GMF), which can have the same form as the minimum problem by a novel hidden variable method. In some embodiments, the Gram-Schmidt process may be used to avoid special cases in the transformation.
[0298] Figure 30 is a flowchart illustrating a method 3000 for efficient position identification according to some embodiments. The method may start with a step (act 3010) of determining a set of correspondences of the extracted features, assuming a set of n 2D / 3D point correspondences and m 2D / 3D line correspondences, and a step (act 3020) of obtaining 2N constraints. The method 3000 may include a step (act 3030) of reconstructing the set of constraints using a partial linearization method to obtain a system of equations. The method may further include a step (act 3040) of solving the system of equations to obtain a rotation matrix, and a step (act 3050) of obtaining t using the rotation matrix and a closed form of t.
[0299] According to some embodiments, any step of method 3000 may be performed on the device described herein and / or on a remote service such as those described herein.
[0300] According to some embodiments, the 2N constraints of act 3020 of method 3000 may include two constraints, each describable in the form l
Chem.
[0301] FIG. 19 is an exemplary schematic diagram of constraints from
Chem.
Chem.
Chem.
Math.
[0302] Regarding the rotation R and the translation t, there may be a total of 6 degrees of freedom. Each line correspondence
Chem.
Math.
[0303] According to some embodiments, the step of restructuring the set of constraints in act 3020 of method 3000 may include generating a quadratic system by using a constraint, namely, an expression of R and a closed form of t using the Cayley-Gibbs-Rodriguez (CGR) parameterization. In some embodiments, CGR may be used to represent R, as discussed, for example, in “A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance”. For example, a 3D vector may be shown as s = [S1, S2, S3]. According to some embodiments, the expression of R using CGR parameterization may be in a form described by the following equation (3’). In (3’), I3 may be a 3×3 identity matrix, and x is the skew matrix of the 3D vector s. In (3’),
Chem.
Math.
[0304] According to some embodiments, the closed form of t in act 3020 may be in the form of τ = -(B T B)B T Ar. In some embodiments, the closed form of t may be derived by first substituting (3’) into (1’), multiplying both sides by the term (1 + S T S), and resulting in the following. [Mathematics]
[0305] Second, as follows, in (4’), [Chemistry] Expand the terms and derive polynomials in s and t. [Mathematics] where a ij is a 10-dimensional vector in s and t, and (1 + s T s) is a cubic polynomial.
[0306] Equation (5’) defines the following, [Mathematics] (5’) may be simplified by rewriting it as follows. [Mathematics]
[0307] Assuming N 2D-3D correspondences, 2N equations can be had as (7’). Stacking the 2N equations of (7’) can give the following. [Mathematics] where A = [a 11 , a 12 , …, a N1 , a N2 T and B = [l1, l1, …, l N , l N TThat is. Regarding the following, (8’) can be treated as a system of linear equations in τ, and a closed-form solution can be obtained.
Number
[0308] According to some embodiments, the quadratic system of act 3020 may be a quadratic system in s1, s2, and s3, and may be in the following form.
Number
[0309] According to some embodiments, the step of obtaining the system of equations in act 3020 of method 3000 using the partial linearization method may include the step of converting the PnL problem into a general minimum formula (GMF) using the partial linearization method and the step of generating the system of equations.
[0310] In some embodiments, the partial linearization method divides the monomials in r defined in (5’) into two groups r3 = [s1 2 , s2 2 , s3 2 T and r7 = [s1s2, s1s3, s2s3, s2, s3, 1] T and may further include the step of appropriately dividing the matrix K in (10’) into K3 and K7, and further rewriting (10’) as follows.
Number
Number
[0311] Wherein, the elements of r3 can be treated as individual unknowns. According to some embodiments, the method may require that the matrix K3 for r3 is of full rank. According to some embodiments, the closed-form solution for r3 with respect to r7 may be described as follows.
Number
[0312] Wherein, for the equation (13’), -(K3 T K3) -1 K3 T K7 can represent a 3×7 matrix. According to some embodiments, when K9 (the K in (10’)) is of maximum rank, r3 may be arbitrarily selected. According to some embodiments, the matrix K9 (i.e., the K in (10’)) may be rank-deficient with respect to an arbitrary number of 2D-3D line correspondences for noise-free data. In some embodiments, when K9 (i.e., the K in (10’)) is rank-deficient, an input can cause K3 to be rank-deficient or approximately rank-deficient for a fixed choice of r3.
[0313] According to some embodiments, K3 may be determined by a Gram-Schmidt process with column pivoting, selecting three independent columns from K9 to generate K3.
Number
[0314] Equation (16’) may be used, and the i-th, j-th, and k-th columns of K are selected such that they are K3, and the corresponding monomials can form r3. The remaining columns may be selected to form K7, and the corresponding monomials may form r7. According to some embodiments, equation (16’) may be solved using other polynomial solvers.
[0315] (The notation in (13’) is C7=(K3 T K3) -1 K3T It may be simplified to K7, and (13’) may be rewritten as follows
Number
[0316] The above system of equations includes three quadratic equations in s1, s2, and s3. Each of the three quadratic equations may have the following form
Number
[0317] According to some embodiments, the step of solving the system of equations to obtain the rotation matrix (act 3030) may include the step of obtaining the rotation matrix by solving the system of equations where the equations are in the form of (15’). According to some embodiments, the system of equations may be solved using the Grobner basis approach. According to some embodiments, the system of equations may be solved using the methods and approaches described by Kukelova et al. (e.g., as described in “Efficient intersection of three quadrics and applications in computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition” by Kukelova, Z., Heller, J., Fitzgibbon, A., which is incorporated herein by reference in its entirety), and the approach described by Zhou may be used to improve the stability
[0318] According to some embodiments, a hidden variable method may be used to solve the system of equations (14’). In some embodiments, a customized hidden variable method may be used to solve the system of equations. For example, the customized hidden variable method is described in “Using algebraic geometry, vol. 185. Springer Science & Business Media (2006)”. In some embodiments, the customized hidden variable method may be implemented by treating what is known in (15’) as constants. For example, s3 may be treated as a constant while s1 and s2 are treated as unknowns so that the system of equations (15’) can be described in the following form.
Number
Number
[0319] When s0 = 1, F i = f i That is, thus, the determinant J of the Jacobian matrix of F0, F1, and F2 may be described as follows.
Number
[0320] J can be a cubic homogeneous equation in s0, s1, and s2, and its coefficients are polynomials in s3. All partial derivatives of J with respect to s0, s1, and s2, F i can be quadratic homogeneous equations in s0, s1, and s2 with the same form as, that is, the following.
Number
[0321] q ij (s3) can be a polynomial in s3. At all non-trivial solutions where F0 = F1 = F2 = 0, G0 = G1 = G2 = 0 (for example, as described in
[10] ). Therefore, they can be combined to form a new homogeneous system for s0, s1, and s2 as in (21’).
Number
[0322] Q(s3) can be a 6×6 matrix, and its elements are polynomials in s3 and u = [s1 2 , s1s2, s2 2 , s0s1, s0s2, s0 2 T Based on linear algebra theory, the homogeneous linear system (21’) can have non-trivial solutions if and only if det(Q(s3)) = 0, where det(Q(s3)) = 0 is an 8th-degree polynomial in s3, which is in the same form as the GMF. Up to 8 solutions can exist.
[0323] According to some embodiments, after obtaining s3, s3 can be back-substituted into (21’) to derive a system of linear homogeneous equations for u. According to some embodiments, s1 and s2 can be calculated through the linear system (21’) by back-substituting s3 into (21’) and setting s0 = 1.
[0324] According to some embodiments, the step of obtaining the rotation matrix in method 3000 (act 3030) may include calculating R using (3’) once s1, s2, and s3 are obtained. According to some embodiments, τ may be calculated by (6’). According to some embodiments, the step of obtaining t (act 3030) may include obtaining t using equation (9’).
[0325] According to some embodiments, an iterative method may be used to refine the solution, as described, for example, in "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance", "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence" and "Camera pose estimation from lines: a fast, robust and general method. Machine Vision and Applications 30". The solution may be refined by minimizing a cost function that is a sixth-degree polynomial in s and t (as described, for example, in "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance").In some embodiments, a damped Newton step may be used to refine the solution (e.g., as described in "Revisiting the pnp problem: A fast, general and optimal solution. In: Proceedings of the IEEE International Conference on Computer Vision" by Zheng, Y., Kuang, Y., Sugimoto, S., Astrom, K., Okutomi, M. and "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance", which are hereby incorporated by reference in their entirety). Specifically, for the k-th step, the Hessian H of the cost function with respect to s and t. k and the gradient g k are calculated. Then, the solution is [s k+1 , t k+1 = [s k , t k - (H k + λI6) -1 g k . Where λ is adjusted at each step according to the Levenberg / Marquardt algorithm (e.g., as described in "The levenberg - marquardt algorithm: implementation and theory. In: Numerical analysis" by More, J.J., which is hereby incorporated by reference in its entirety) to reduce the cost at each step. The solution with the minimum cost can be regarded as the solution.
[0326] According to some embodiments, the PnL solution described herein is applicable to N≧3 two-dimensional / three-dimensional line correspondences. In some embodiments, the method of solving the PnL problem may include four steps. In some embodiments, the first step may include the step of combining 2N constraints (4') into three equations (15'). In some embodiments, the three equations (15'), which are a system of equations, may be solved by a hidden variable method to restore the rotation R and the translation t. According to some embodiments, the PnL solution may be further refined by a damped Newton step. FIG. 31 shows an exemplary algorithm 3100 for solving the PnL problem according to some embodiments.
[0327] The computational complexity of step 2 (action 3120) and step 3 (action 3130) of algorithm 3100 is O(1) because it is independent of the number of correspondences. The main computational cost of step 1 is for solving the linear least squares problems (9') and (13'). The main computational cost of step 4 is for calculating the sum of the squared distance functions. The computational complexity of these steps increases linearly with N. In short, the computational complexity of algorithm 3100 is O(N).
[0328] According to some embodiments, the components of the algorithm for the solution of the PnL problem described herein are referred to as MinPnL. FIGS. 24-27 show a comparison of the MinPnL algorithm with previous P3L and least squares PnL algorithms, according to some embodiments. The algorithms being compared for solving the P3L and least squares PnL algorithms are, for the P3L problem, three recent studies AlgP3L (as described, for example, in “A stable algebraic camera pose estimation for minimal configurations of 2d / 3d point and line correspondences. In: Asian Conference on Computer Vision”), RP3L (as described, for example, in “Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence”), and SRP3L (as described, for example, in ’A novel algebraic solution to the perspective-threeline pose problem. Computer Vision and Image Understanding p. 102711 (2018)’ by Wang, P., Xu, G., Cheng, Y., which is incorporated herein by reference in its entirety), and for the least squares problem, OAPnL, SRPnL (for example, ’A novel algebraic solution to the perspective-threeline pose problem. Computer Vision and Image Understanding p.102711 (as described in (2018)), ASPnL (e.g., as described in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence"), Ansar (e.g., as described in 'Linear pose estimation from points or lines. IEEE Transactions on Pattern Analysis and Machine Intelligence 25(5), 578{589 (2003)'), Mirzaei (e.g., as described in 'Optimal estimation of vanishing points in a manhattan world. In: 2011 International Conference on Computer Vision'), LPnL DLT (e.g., as described in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence"), DLT Combined Lines (e.g., as described in 'Camera pose estimation from lines using pln ucker coordinates. arXiv preprint arXiv:1608.02824 (2016)'), DLT Plucker Lines (e.g., "Absolute pose estimation from line correspondences using direct linear transformation.including LPnL Bar LS (as described, for example, in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence"), LPnL Bar ENull (as described, for example, in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence"), cvxPnPL (as described, for example, in "’Cvxpnpl: A unified convex solution to the absolute pose estimation problem from point and line correspondences"), OPnPL, and EPnPL Planar (as described, for example, in "Accurate and linear time pose estimation from points and lines. In: European Conference on Computer Vision.").
[0329] In FIGS. 24-27, the following metrics (e.g., as described in previous studies "Absolute pose estimation from line correspondences using direct linear transformation. Computer Vision and Image Understanding" and "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance") are used to measure the estimation error. Specifically, R gt and t gt are the ground truth rotation and translation, and
Chem.
Chem.
Chem.
[0330] For FIGS. 24-26, synthetic data is used to evaluate the performance of different algorithms. The polynomial solver for the equation system (15') is first compared with the influence of the Gram-Schmidt process, and then MinPnL is compared with the state-of-the-art P3L and least squares PnL algorithms.
[0331] The synthetic data used for the purpose of comparison in FIGS. 24-26 is generated in the same manner as described in "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance", "The planar three-line junction perspective problem with application to the recognition of polygonal patterns. Pattern recognition", "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence", and "Camera pose estimation from lines: a fast, robust and general method. Machine Vision and Applications 30" (incorporated herein by reference). Specifically, the camera resolution may be set to 640×480 pixels and the focal length may be set to 800. Euler angles α, β, γ may be used to generate the rotation matrix. For each trial, the camera is randomly placed within [-10m; 10m] 3 within the cube, and the Euler angles are uniformly sampled from α, γ ∈ [0°, 360°] and β ∈ [0°, 180°]. Then, N 2D / 3D line correspondences are randomly generated. The endpoints of the 2D lines are first generated randomly, and then the 3D endpoints are generated by projecting the 2D endpoints into 3D space. The depth of the 3D endpoints is within [4m; 10m]. These 3D endpoints are then transformed into the world frame.
[0332] Histograms and box plots may be used to compare the estimation errors. The histogram may be used to present the main distribution of the errors, while the box plot may be used to better show large errors. In a box plot, the center mark of each box indicates the median, and the lower and upper edges indicate the 25th and 75th percentiles, respectively. The whiskers extend up to + / -2.7 standard deviations, and errors outside this range are plotted individually using the "+" symbol. The numerical stability of the Hidden Variable (HV) polynomial solver is compared with the Grobner, E3Q3, and RE3Q3 algorithms using 10,000 trials (as described, for example, in "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance").
[0333] Figures 20A - B show the results. It is clear that the hidden variable solver is more stable than other algorithms. The algorithms described in "Efficient solvers for minimal problems by syzygy - based reduction. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition", "Upnp: An optimal o(n) solution to the absolute pose problem with universal applicability. In: European Conference on Computer Vision", and "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance" generate large errors. Like the Grobner method, E3Q3 and RE3Q3 all involve steps to calculate the inverse of a matrix, and they may encounter numerical problems, which can lead to these large errors.
[0334] One important step of the method described herein is to rearrange Kr = 0(10’) as K3r3=-K7r7(13’). There are 84 options for r3. Different options may have different impacts on numerical stability. Respectively, MinPn_s i 2 , MinPnL_s i s i , and MinPnL_s i consider three options for r3, namely, [S1 2 , S2 2 , S3 2 , [s1s2, s1s3, s2s3] and [s1, s2, s3]. For this comparison, the corresponding number of N is increased from 4 to 20, and the standard deviation of the noise is set to 2 pixels. For each N, 1,000 trials are conducted to test the performance.
[0335] Figures 23A-B demonstrate the results. Figure 23A shows a comparison of the average rotational error in degrees between different P3L algorithms. Figure 23B shows a box-and-whisker plot of the rotational error between different P3L algorithms. The fixed selection of r3 can encounter numerical problems when K3 approximates a singular matrix. The Gram-Schmidt process, used in some embodiments of the solution for the algorithms described herein, solves this problem and can thus produce more stable results.
[0336] MinP3L, which is a solution to the P3L problem as described herein, may be compared with previous P3L algorithms, including AlgP3L (as described, for example, in "A stable algebraic camera pose estimation for minimal configurations of 2d / 3d point and line correspondences. In: Asian Conference on Computer Vision"), RP3L (as described, for example, in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence"), and SRP3L. To compare performance fairly, the results are without iterative refinement since the algorithms being compared do not have refinements. The numerical stability of the different algorithms, i.e., the estimation error without noise, must be considered. 10,000 trials were conducted to test accuracy. Figures 22A - B show the results. Figure 22A shows a box - and - whisker plot of the rotational error of an embodiment of the algorithm described herein and the algorithms AlgP3L, RP3L, and SRP3L. Figure 22B shows a box - and - whisker plot of the translational error of an embodiment of the algorithm described herein and the previous algorithms AlgP3L, RP3L, and SRP3L. The rotational and translational errors of MinP3L implemented using the methods and techniques described herein are 10 -5Smaller. All other algorithms result in large errors, as indicated by the longer tails in the box-and-whisker plot of FIG. 22. Next, the behavior of the P3L algorithm is examined under varying noise levels. Gaussian noise is added to the endpoints of the 2D lines. The standard deviation increases from 0.5 to 5 pixels. FIGS. 23A-B show the results. FIG. 23A shows the mean rotational error of certain embodiments of the algorithms described herein and the previous algorithms AlgP3L, RP3L, and SRP3L. FIG. 23B shows the mean translational error of certain embodiments of the algorithms described herein and the previous algorithms AlgP3L, RP3L, and SRP3L.
[0337] The MinP3L algorithm, implemented using the techniques described herein, exhibits stability. As in the noise-free case, the algorithms being compared (e.g., as described in "A stable algebraic camera pose estimation for minimal configurations of 2d / 3d point and line correspondences. In: Asian Conference on Computer Vision", "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence") each have longer tails than the algorithms developed using the techniques described herein. This can be caused by numerically unstable behavior within these algorithms.
[0338] As discussed in the references "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance", "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence", and "Camera pose estimation from lines: a fast, robust and general method. Machine Vision and Applications 30", two configurations of 2D line segments were considered, including the case where they are matched (e.g., the 2D line segments are uniformly distributed throughout the image) and the case where they are not matched (e.g., the 2D line segments are constrained within [0,160]×[0,120]). The following results are from 500 independent trials.
[0339] In the first experiment, the performance of the PnL algorithm is examined with respect to the correspondence of variables. The standard deviation of the Gaussian noise added to the 2D line endpoints is set to 2 pixels. In the second experiment, the situation of increasing noise levels is examined. σ is stepped in 0.5 pixel increments from 0.5 pixel to 5 pixels, and N is set to 10. Figures 24A-D and 25A-D show the mean and median errors. Figure 24A shows the mean rotation error of different PnL algorithms. Figure 24B shows the mean translation error of different PnL algorithms. Figure 24C shows the median rotation error of different PnL algorithms. Figure 24D shows the median translation error of different PnL algorithms. Figure 25A shows the mean rotation error of different PnL algorithms. Figure 25B shows the mean translation error of different PnL algorithms. Figure 25C shows the median rotation error of different PnL algorithms. Figure 25D shows the median translation error of different PnL algorithms.
[0340] Typically, solutions based on polynomial formulas are more stable than linear solutions. Other algorithms clearly provide larger errors. Additionally, the performance of the PnL algorithm in a planar configuration is also considered (i.e., when all 3D lines lie in a plane). Planar configurations are widely present in artificial environments. However, many PnL algorithms are infeasible with respect to planar configurations, as shown in "A robust and efficient algorithm for the pnl problem using algebraic distance to approximate the reprojection distance". Here, as shown in FIGS. 26A-D and 27A-D, it is compared with five PnL algorithms. FIG. 26A shows the average rotational error of different PnL algorithms. FIG. 26B shows the average translational error of different PnL algorithms. FIG. 26C shows the median rotational error of different PnL algorithms. FIG. 26D shows the median translational error of different PnL algorithms. FIG. 27A shows the average rotational error of different PnL algorithms. FIG. 27B shows the average translational error of different PnL algorithms. FIG. 27C shows the median rotational error of different PnL algorithms. FIG. 27D shows the median translational error of different PnL algorithms.
[0341] MinPnL, implemented using the techniques and methods described herein, achieves the best results. cvxPnPL and ASPnL (e.g., as described in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence") produce large errors that are out of range.
[0342] Some of the methods and techniques described herein for finding the pose of a camera using features can function even when feature points and feature lines lie in the same plane.
Example
[0343] Example
[0344] Actual data was also used to evaluate the PnL algorithm. The MPI and VGG datasets were used to evaluate performance. They contain a total of 10 datasets, and their characteristics are listed in Table 1. Here, since the ground truth translation is [0;0;0] in some cases, in the simulation, instead of the relative error, the absolute translation error
Chemical formula
[0345] Using Matlab 2019a, the computation time of the PnL algorithm was determined on a 3.1 HZ intel i7 laptop. The results from 500 independent trials are illustrated in FIGS. 29A-C. The algorithms Ansar and cvxPnPL are slow and thus not shown within the graph range. As can be seen from FIGS. 29A-C, LPnLBarLS is the fastest among those tested, however, it is not stable. As shown above, OAPnL and the algorithms according to the embodiments described herein are generally the two most stable algorithms. As shown in FIG. 29B, the algorithm according to the embodiments described herein is approximately twice as fast as OAPnL.The MinPnL algorithm has a startup time that is similar compared to the linear algorithms DLTCombined (e.g., as described in "Absolute pose estimation from line correspondences using direct linear transformation. Computer Vision and Image Understanding") and DLT Plucker (e.g., as described in "Camera pose estimation from lines using pln ucker coordinates. arXiv preprint"), is slightly faster than LPnL Bar ENull (e.g., as described in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence") when N is within 100, and is faster than LPnL DLT (e.g., as described in "Pose estimation from line correspondences: A complete analysis and a series of solutions. IEEE transactions on pattern analysis and machine intelligence") when N is large.
[0346] Figure 29A is a schematic diagram of the calculation times of many algorithms.
[0347] Figure 29B is a schematic diagram of the calculation time of an embodiment of the algorithm described herein compared to the calculation times of algorithms involving polynomial systems.
[0348] Figure 29C is a schematic diagram of the calculation time of an embodiment of the algorithm described herein compared to the calculation times of algorithms based on linear transformation.
[0349] Further Considerations
[0350] FIG. 32 shows a schematic representation of a machine in an exemplary form of computer system 1900, where a set of instructions for causing the machine to carry out any one or more of the methodologies discussed herein may be executed in accordance with some embodiments. In alternative embodiments, the machine may operate as a stand-alone device or may be connected (e.g., networked) to other machines. Further, although only a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set of instructions (or multiple sets) to carry out any one or more of the methodologies discussed herein.
[0351] Exemplary computer system 1900 includes a processor 1902 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), main memory 1904 (e.g., read only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), and static memory 1906 (e.g., flash memory, static random access memory (SRAM), etc.), which communicate with each other via bus 1908.
[0352] Computer system 1900 may further include a disk drive unit 1916 and a network interface device 1920.
[0353] The disk drive unit 1916 includes a machine-readable medium 1922 on which is stored a set 1924 of one or more instructions (e.g., software) that embody any one or more of the methodologies or functions described herein. The software may also reside, in whole or in part, within the main memory 1904 and / or within the processor 1902 during its execution by the computer system 1900, main memory 1904, and processor 1902, and may likewise constitute a machine-readable medium.
[0354] The software may also be transmitted or received via the network 18, via the network interface device 1920.
[0355] The computer system 1900 includes a driver chip 1950 that drives a projector and is used to generate light. The driver chip 1950 includes its own data storage device 1960 and its own processor 1962.
[0356] Although the machine-readable medium 1922 is shown as a single medium in the exemplary embodiment, the term "machine-readable medium" should be construed to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store a set of one or more instructions. The term "machine-readable medium" should also be construed to include any medium that is capable of storing, encoding, or carrying a set of instructions for machine execution and that causes a machine to perform any one or more of the methodologies of the present invention. The term "machine-readable medium" should, therefore, be construed to include, without limitation, solid-state memory, optical and magnetic media, and carrier wave signals.
[0357] According to various embodiments, the communication network 1928 may be a local area network (LAN), a cellular network, a Bluetooth® network, the Internet, or any other such network.
[0358] Although some aspects of some embodiments have been described so far, it should be understood that various modifications, corrections, and improvements will readily occur to those skilled in the art.
[0359] As an example, embodiments are described in relation to an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied within an MR environment, or more generally, within other XR environments and VR environments.
[0360] As another example, embodiments are described in relation to devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.
[0361] Such modifications, corrections, and improvements are intended to be part of this disclosure and are intended to be within the spirit and scope of this disclosure. Further, while the advantages of this disclosure are shown, it should be understood that not all embodiments of this disclosure include all of the advantages described. Some embodiments may not implement any of the features described as advantageous in this specification and in some cases. Therefore, the foregoing description and drawings are merely examples.
[0362] The foregoing embodiments of the present disclosure can be implemented in any of a number of ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether provided within a single computer or distributed among multiple computers. Such processors include, by way of example, one or more processors within an integrated circuit component, including commercially available integrated circuit components known in the art such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors, etc. In some embodiments, the processor may be implemented within a custom circuit such as an ASIC or within a semi-custom circuit resulting from configuring a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores such that a subset of one or those cores may constitute the processor. However, the processor may be implemented using a circuit in any suitable format.
[0363] Furthermore, it should be understood that the computer can be embodied in any of several forms such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, the computer may be embodied in a device that is generally not considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any suitable portable or stationary electronic device.
[0364] In addition, the computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual presentation of output, or a speaker or other sound generating device for audible presentation of output. Examples of input devices that can be used for a user interface include a keyboard, and pointing devices such as a mouse, touchpad, and digitizing tablet. As another example, the computer may receive input information through speech recognition or in other audible formats. In the illustrated embodiment, the input / output devices are shown physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated within the same unit as the processor or other elements of the computing device. For example, a keyboard can be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated through a wireless connection.
[0365] Such computers may be interconnected by one or more networks in any suitable form, including local area networks or wide area networks such as corporate networks or the Internet. Such networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0366] Moreover, the various methods and processes outlined in this specification may be encoded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software may be written using any of several suitable programming languages and / or programming or scripting tools, and may also be compiled as executable machine language code or intermediate code to be executed on a framework or virtual machine.
[0367] In this aspect, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy (registered trademark) disks, compact disk (CD), optical disk, digital video disk (DVD), magnetic tape, flash memory, circuit configurations within a field programmable gate array or other semiconductor device, or other tangible computer storage media) encoded with one or more programs that perform a method of implementing the various embodiments of the present disclosure discussed above when executed on one or more computers or other processors. As is apparent from the foregoing examples, a computer-readable storage medium can retain information for a sufficient time to provide computer-executable instructions in a non-transitory fashion. Such a computer-readable storage medium or media can be transportable such that the one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage medium" includes only computer-readable media that can be regarded as a manufactured (i.e., a manufactured article) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.
[0368] As described above, the terms "program" or "software" are used herein in a general sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the present disclosure.
[0369] In addition, according to one aspect of the present embodiment, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but may be distributed in a modular fashion among several different computers or processors to implement various aspects of the present disclosure.
[0370] Computer-executable instructions may be in many forms such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.
[0371] Also, data structures may be stored on a computer-readable medium in any suitable form. For simplicity of illustration, a data structure may be shown as having fields that are related through locations within the data structure. Such relationships may also be achieved by allocating storage for fields with locations within the computer-readable medium that convey relationships between the fields. However, any suitable mechanism, including the use of pointers, tags, or other mechanisms for establishing relationships between data elements, may be used to establish relationships between the information in the fields of a data structure.
[0372] Various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and thus, its use is not limited to the details and arrangements of the components described in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined with aspects described in other embodiments in any manner.
[0373] Also, the present disclosure may be embodied as a method for which examples are provided. The acts performed as part of the method may be ordered in any suitable way. Thus, in the illustrative embodiments, although shown as consecutive acts, embodiments may be constructed in which acts are performed in a different order than illustrated, including performing some acts simultaneously.
[0374] The use of ordinal terms such as "first", "second", "third", etc. in the claims to modify claim elements does not by itself imply any priority, precedence, or order of one claim element over another, or the temporal order in which acts of a method are performed, but the ordinal terms are used only as labels to distinguish one claim element having a certain name from another element having the same name for the purpose of distinguishing claim elements.
[0375] Also, the phrases and terminology used herein are for the purpose of explanation and should not be regarded as limiting. The use of "comprising", "having", "including", "containing", "accompanying", and variations thereof herein means including the recited items, their equivalents, and additional items.
[0376] Some values are described as being derived by "minimizing" or "optimizing". It should be understood that words such as "minimizing" and "optimizing" may involve steps to find the minimum or maximum possible value, but this is not necessarily the case. Rather, these results may be achieved by finding the minimum or maximum value based on practical constraints, such as the execution of successive iterations of a process until the change between iterations of a certain order is below a threshold.
Claims
**Claim 1** A method for determining the pose of a camera relative to a map based on one or more images captured using the camera, wherein the pose is represented as a rotation matrix and a translation matrix, the method comprising: developing correspondences between the one or more images and a combination of points and / or lines within the map; converting the correspondences into a set of three quadratic polynomial equations, the conversion comprising: determining a set of constraints based on the correspondences between the one or more images and the combination of points and / or lines within the map; obtaining the set of three quadratic polynomial equations by reconstructing the set of constraints based on the correspondences using a partial linearization method; and solving the set of three quadratic polynomial equations for the rotation matrix; calculating the translation matrix based on the rotation matrix. A method comprising the above. **Claim 2** The method according to claim 1, wherein the combination of points and / or lines is dynamically determined based on characteristics of the one or more images. **Claim 3** The method according to claim 1, further comprising refining the pose by minimizing a cost function. **Claim 4** The method according to claim 1, further comprising refining the pose by using a damped Newton step. **Claim 5** The conversion of the correspondences into the set of three quadratic polynomial equations comprises: deriving a set of constraints from the correspondences; forming a closed-form expression for the translation matrix; forming a parameterization of the rotation matrix using 3D vectors. A method according to claim 1, comprising the above. **Claim 6** The method according to claim 1, wherein the conversion of the correspondences into the set of three quadratic polynomial equations further comprises noise removal by order approximation. **Claim 7** The method according to claim 1, wherein solving the set of three quadratic polynomial equations for the rotation matrix comprises using a hidden variable method. **Claim 8** The method according to claim 1, wherein forming a parameterization of the rotation matrix using 3D vectors comprises using a Cayley-Gibbs-Rodriguez (CGR) parameterization. **Claim 9** Forming the closed-form expression of the translation matrix includes forming a system of linear equations using the set of constraints, the method of claim 5. **Claim 10** The point and / or line in the one or more images are two-dimensional features, The corresponding feature in the map is a three-dimensional feature, the method of claim 1. **Claim 11** Converting the correspondence into a set of equations of three quadratic polynomials includes representing the correspondence as a well-determined system of equations in a plurality of variables, and formatting the well-determined system of equations as a set of equations of three quadratic polynomials the method of claim 1. **Claim 12** The set of equations of three quadratic polynomials is an equation of meta-variables, and in the equation of meta-variables, each of the meta-variables represents a group of the plurality of variables, the method of claim 11. **Claim 13** Solving the set of equations of three quadratic polynomials for the rotation matrix includes calculating the values of the meta-variables, and calculating the pose from the meta-variables the method of claim 12. **Claim 14** The point and / or line in the one or more images are two-dimensional features, Expanding the correspondence between the one or more images and the combination of points and / or lines in the map includes expanding the correspondence between the two-dimensional features and the three-dimensional features in the map the method of claim 1. **Claim 15** A portable electronic device comprising a camera configured to capture one or more images, a processor, and a non-transitory computer-readable medium storing instructions that, when executed by the processor, cause the processor to determine the pose of the camera relative to a map based on the one or more images captured by the camera, the pose being represented as a rotation matrix and a translation matrix, and the instructions cause the processor to perform the method of any of claims 1 to 14, the non-transitory computer-readable medium a portable electronic device. **Claim 16** A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method, the method comprising: Developing correspondences between one or more images and combinations of points and / or lines in a map; Converting the correspondences into a set of three quadratic polynomial equations, the converting comprising: Determining a set of constraints based on the correspondences between the one or more images and the combinations of points and / or lines in the map; Obtaining the set of three quadratic polynomial equations by restructuring the set of constraints based on the correspondences using a partial linearization method; Including; Solving the set of three quadratic polynomial equations for a rotation matrix; Calculating a translation matrix based on the rotation matrix; A non-transitory computer-readable storage medium including.
17. The points and / or lines in the one or more images are two-dimensional features, The corresponding features in the map are three-dimensional features, the non-transitory computer-readable storage medium according to claim 16.
18. The combinations of points and / or lines are dynamically determined based on the characteristics of the one or more images, the non-transitory computer-readable storage medium according to claim 16.
19. The non-transitory computer-readable storage medium according to claim 16, further comprising refining the pose by minimizing a cost function.
20. The non-transitory computer-readable storage medium according to claim 16, further comprising refining the pose by using a damped Newton step.
21. The converting the correspondences into the set of three quadratic polynomial equations comprises: Deriving a set of constraints from the correspondences; Forming a closed-form expression of the translation matrix; Forming a parameterization of the rotation matrix using 3D vectors; Including, the non-transitory computer-readable storage medium according to claim 16.
22. The converting the correspondences into the set of three quadratic polynomial equations further comprises removing noise by rank approximation, the non-transitory computer-readable storage medium according to claim 16.
23. Solving the set of equations of the three quadratic polynomials for the rotation matrix involves using a hidden variable method, the non-transitory computer-readable storage medium according to claim 16.
24. Forming the parameterization of the rotation matrix using a 3D vector involves using a Cayley-Gibbs-Rodriguez (CGR) parameterization, the non-transitory computer-readable storage medium according to claim 16.
25. Forming the closed-form representation of the translation matrix involves using the set of constraints to form a system of linear equations, the non-transitory computer-readable storage medium according to claim 21.
26. A portable electronic device, A camera configured to capture one or more images of a 3D environment, At least one processor configured to execute computer-executable instructions, the computer-executable instructions comprising: Determining information about points and / or combinations of lines in the one or more images of the 3D environment, Sending the information about the points and / or combinations of lines in the one or more images to a location service to determine the pose of the camera relative to a map, Receiving from the location service the pose of the camera relative to the map represented as a rotation matrix and a translation matrix Instructions for determining the pose of the camera relative to the map based on the one or more images, including At least one processor comprising Comprising The location service is implemented on the portable electronic device, Determining the pose of the camera relative to the map involves: Developing a correspondence between the one or more images and the points and / or combinations of lines in the map, Converting the correspondence into a set of equations of three quadratic polynomials, the converting comprising: Determining a set of constraints based on the correspondence between the one or more images and the points and / or combinations of lines in the map, Obtaining the set of equations of the three quadratic polynomials by reconfiguring the set of constraints based on the correspondence using a partial linearization method Including solving the set of equations of the three quadratic polynomials for the rotation matrix; calculating the translation matrix based on the rotation matrix; A portable electronic device including the above.
27. The portable electronic device according to claim 26, wherein the combination of the points and / or lines is dynamically determined based on the one or more image characteristics.
28. The portable electronic device according to claim 26, wherein determining the pose of the camera with respect to the map further includes refining the pose by minimizing a cost function.
29. The portable electronic device according to claim 26, wherein determining the pose of the camera with respect to the map further includes refining the pose by using a damped Newton step.
30. Converting the correspondence into a set of equations of the three quadratic polynomials includes: deriving a set of constraints from the correspondence; forming a closed-form representation of the translation matrix; forming a parameterization of the rotation matrix using 3D vectors; The portable electronic device according to claim 26, including the above.
31. The portable electronic device according to claim 26, wherein converting the correspondence into a set of equations of the three quadratic polynomials further includes noise removal by rank approximation.
32. The portable electronic device according to claim 26, wherein solving the set of equations of the three quadratic polynomials for the rotation matrix includes using a hidden variable method.
33. The portable electronic device according to claim 30, wherein forming the parameterization of the rotation matrix using the 3D vectors includes using a Cayley-Gibbs-Rodriguez (CGR) parameterization.
34. The portable electronic device according to claim 30, wherein forming the closed-form representation of the translation matrix includes forming a system of linear equations using the set of constraints.
Citation Information
Patent Citations
Three-dimensional coordinate calculation device, three-dimensional coordinate calculation method, and three-dimensional coordinate calculation program
JP2016070674A
Image processing device, image processing method, and image processing program
WO2017022033A1