Map optimization method, electronic device, and non-transitory computer-readable storage medium

By generating a virtual camera and using global bundle adjustment to optimize the map, the problem of map offset caused by errors in point cloud maps is solved, improving the accuracy of position tracking and the matching accuracy of the virtual environment.

CN116433753BActive Publication Date: 2025-11-21HTC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310018088.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-16
Filing Date
2023-01-06
Publication Date
2025-11-21
Estimated Expiration
2043-01-06

AI Technical Summary

Technical Problem

Point cloud maps are prone to map shifts due to accumulated errors during map generation in inside-out tracking methods, which affects the accuracy of position tracking.

Method used

By generating multiple virtual cameras, global bundle adjustment is used to optimize the map, adjust the estimated coordinates and camera poses to reduce reprojection errors, and combine the projection errors of real keyframes and virtual frames to optimize the position of map points.

Benefits of technology

It improves map accuracy and location tracking accuracy, ensuring a precise match between the virtual and real environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116433753B_ABST
    Figure CN116433753B_ABST
Patent Text Reader

Abstract

The map optimization method includes: identifying first and second map points from a map using first and second estimated coordinates generated based on first and second markers, respectively; generating a plurality of virtual cameras from the first and second estimated coordinates, the virtual cameras being controlled by virtual poses and including a plurality of optical axes intersecting at first and second intersection coordinates, a distance between the first and second intersection coordinates being a first distance value of a plurality of distance values, the virtual cameras providing virtual frames indicating that the first and second markers are observed at the first and second intersection coordinates, respectively; performing a global bundle adjustment to optimize the map, including adjusting the first and second estimated coordinates to reduce a sum of a plurality of re-projection errors calculated from real key frames and the virtual frames, and the electronic device is configured to be positioned according to the optimized map. The map optimization method provided by the present disclosure can reduce the map drift caused by the accumulated errors in the map generation process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an inside-out tracking technique, and in particular, to a map optimization method, an electronic device, and a non-transitory computer-readable storage medium. BACKGROUND

[0002] Inside-out tracking is a commonly used position tracking method in virtual reality (VR) and related technologies, which is used to track the position of a head-mounted device (HMD) and a controller. In the inside-out tracking method, the required sensors (e.g., cameras) are installed on the HMD, while the outside-in tracking method requires sensors (e.g., infrared light towers) placed in a fixed position. Inside-out tracking can be achieved through simultaneous localization and mapping (SLAM). SLAM uses images captured by cameras on the HMD to obtain depth information of the real-world environment, and then constructs a point cloud map. The HMD can use the point cloud map to render a virtual environment, track the position of a virtual object, and / or perform self-localization. However, the point cloud map is prone to map drift caused by accumulated errors in the map generation process. SUMMARY

[0003] The present disclosure provides a map optimization method for an electronic device storing a plurality of distance values and a map, the map optimization method comprising the following steps: identifying a first map point from the map using a first estimated coordinate generated based on a first marker, and identifying a second map point from the map using a second estimated coordinate generated based on a second marker; generating a plurality of virtual cameras based on the first estimated coordinate and the second estimated coordinate, the plurality of virtual cameras being controlled by a virtual pose and comprising a plurality of optical axes intersecting at a first intersection coordinate and a second intersection coordinate, a distance between the first intersection coordinate and the second intersection coordinate being a first distance value of the plurality of distance values, wherein the plurality of virtual cameras provide a plurality of virtual frames indicating that the first marker and the second marker are observed at the first intersection coordinate and the second intersection coordinate, respectively; and performing a global bundle adjustment to optimize the map, comprising adjusting the first estimated coordinate and the second estimated coordinate to reduce a sum of a plurality of re-projection errors calculated based on a plurality of real key frames and the plurality of virtual frames, wherein the electronic device is configured to perform localization based on the optimized map.

[0004] In some embodiments of the map optimization method, the plurality of real keyframes are captured by the electronic device at a plurality of camera poses, and the plurality of camera poses are stored in the electronic device, wherein the step of performing global bundle adjustment to optimize the map further comprises the steps of: adjusting the plurality of camera poses and the virtual poses to reduce the sum of the plurality of re-projection errors.

[0005] In some embodiments of the map optimization method, the step of performing global bundle adjustment to optimize the map further comprises calculating the sum of the plurality of re-projection errors, wherein calculating the sum of the plurality of re-projection errors comprises the steps of: projecting the first estimated coordinates onto the plurality of real keyframes and the plurality of virtual frames having the plurality of first pixel points associated with the first map point to generate a plurality of first projected points; calculating each first re-projection error between each first pixel point and a corresponding one of the plurality of first projected points; projecting the second estimated coordinates onto the plurality of real keyframes and the plurality of virtual frames having the plurality of second pixel points associated with the second map point to generate a plurality of second projected points; calculating each second re-projection error between each second pixel point and a corresponding one of the plurality of second projected points; and adding the plurality of first re-projection errors, the plurality of second re-projection errors, and a plurality of other re-projection errors of a plurality of other map points of the map to generate the sum of the plurality of re-projection errors.

[0006] In some embodiments of the map optimization method, the number of the plurality of virtual cameras is twelve, the plurality of optical axes of one half of the plurality of virtual cameras intersect at a first intersection point coordinate, and the plurality of optical axes of the other half of the plurality of virtual cameras intersect at a second intersection point coordinate.

[0007] In some embodiments of the map optimization method, when the plurality of virtual cameras are generated, the first intersection point coordinate, the second intersection point coordinate, the first estimated coordinate, and the second estimated coordinate are aligned on an imaginary straight line with each other.

[0008] In some embodiments of the map optimization method, when the plurality of virtual cameras are generated, the distance between the first estimated coordinate and the first intersection point coordinate is equal to the distance between the second estimated coordinate and the second intersection point coordinate.

[0009] In some embodiments of the map optimization method, the first distance value represents a real-world distance between the first marker and the second marker.

[0010] In some embodiments of the map optimization method, each distance value represents a real-world distance between two of the plurality of markers, the plurality of markers including a first marker and a second marker, wherein the electronic device is configured to generate a plurality of sets of virtual cameras, the plurality of virtual cameras being one of the plurality of sets of virtual cameras, wherein during the performing of the global bundle adjustment, each set of virtual cameras is split into two sets, the two sets being separated from each other by a corresponding one of the plurality of distance values to incorporate the plurality of distance values into the global bundle adjustment.

[0011] The present disclosure provides an electronic device including a storage circuit and a computing circuit. The storage circuit is configured to store a plurality of distance values and a map. The computing circuit is configured to: identify a first map point from the map using a first estimated coordinate generated based on a first marker, and identify a second map point from the map using a second estimated coordinate generated based on a second marker; generate a plurality of virtual cameras based on the first estimated coordinate and the second estimated coordinate, the plurality of virtual cameras being controlled by a virtual pose and including a plurality of optical axes intersecting at a first intersection coordinate and a second intersection coordinate, a distance between the first intersection coordinate and the second intersection coordinate being a first distance value of the plurality of distance values, wherein the plurality of virtual cameras provide a plurality of virtual frames indicating that the first marker and the second marker are observed at the first intersection coordinate and the second intersection coordinate, respectively; and perform a global bundle adjustment to optimize the map, including adjusting the first estimated coordinate and the second estimated coordinate to reduce a sum of a plurality of re-projection errors calculated based on a plurality of real key frames and the plurality of virtual frames, wherein the electronic device is configured to perform positioning based on the optimized map.

[0012] In some embodiments of the electronic device, the plurality of real key frames are captured by the electronic device at a plurality of camera poses, respectively, and the storage circuit is configured to store the plurality of camera poses, wherein when the global bundle adjustment is performed to optimize the map, the computing circuit is further configured to adjust the plurality of camera poses and the virtual pose to reduce the sum of the plurality of re-projection errors.

[0013] In some embodiments of the electronic device, when performing global bundle adjustment to optimize the map, the computing circuitry is configured to perform the following steps to compute the sum of the plurality of re-projection errors: project the first estimated coordinates onto the plurality of real keyframes and the plurality of virtual frames having the plurality of first pixel points associated with the first map point to generate a plurality of first projected points; compute each first re-projection error between each first pixel point and a corresponding one of the plurality of first projected points; project the second estimated coordinates onto the plurality of real keyframes and the plurality of virtual frames having the plurality of second pixel points associated with the second map point to generate a plurality of second projected points; compute each second re-projection error between each second pixel point and a corresponding one of the plurality of second projected points; and sum the plurality of first re-projection errors, the plurality of second re-projection errors, and a plurality of other re-projection errors of a plurality of other map points of the map to generate the sum of the plurality of re-projection errors.

[0014] In some embodiments of the electronic device, the number of the plurality of virtual cameras is twelve, the plurality of optical axes of one half of the plurality of virtual cameras intersect at the first intersection coordinates, and the plurality of optical axes of another half of the plurality of virtual cameras intersect at the second intersection coordinates.

[0015] In some embodiments of the electronic device, when the plurality of virtual cameras are generated, the first intersection coordinates, the second intersection coordinates, the first estimated coordinates, and the second estimated coordinates are aligned on an imaginary straight line with each other.

[0016] In some embodiments of the electronic device, when the plurality of virtual cameras are generated, a distance between the first estimated coordinates and the first intersection coordinates is equal to a distance between the second estimated coordinates and the second intersection coordinates.

[0017] In some embodiments of the electronic device, the first distance value represents a real-world distance between the first marker and the second marker.

[0018] In some embodiments of the electronic device, each distance value represents a real-world distance between two of the plurality of markers, the plurality of markers including the first marker and the second marker, wherein the computing circuitry is configured to generate a plurality of sets of virtual cameras, the plurality of virtual cameras being one of the plurality of sets of virtual cameras, wherein during performing the global bundle adjustment, each set of virtual cameras is split into two sets, the two sets being separated from each other by a corresponding one of the plurality of distance values to incorporate the plurality of distance values into the global bundle adjustment.

[0019] The present disclosure provides a non-transitory computer-readable storage medium storing a plurality of computer-readable instructions for controlling an electronic device. The electronic device includes a computing circuit and a storage circuit storing a plurality of distance values and a map. When the plurality of computer-readable instructions are executed by the computing circuit, the computing circuit is configured to: identify a first map point from the map using a first estimated coordinate generated based on a first marker and a second map point from the map using a second estimated coordinate generated based on a second marker; generate a plurality of virtual cameras based on the first estimated coordinate and the second estimated coordinate, the plurality of virtual cameras being controlled by a virtual pose and including a plurality of optical axes intersecting at a first intersection coordinate and a second intersection coordinate, a distance between the first intersection coordinate and the second intersection coordinate being a first distance value of the plurality of distance values, wherein the plurality of virtual cameras provide a plurality of virtual frames indicating that the first marker and the second marker are observed at the first intersection coordinate and the second intersection coordinate, respectively; and perform a global bundle adjustment to optimize the map, including adjusting the first estimated coordinate and the second estimated coordinate to reduce a sum of a plurality of re-projection errors calculated based on a plurality of real keyframes and the plurality of virtual frames, wherein the electronic device is configured to perform positioning based on the optimized map.

[0020] In some embodiments of the non-transitory computer-readable storage medium, the plurality of real keyframes are captured by the electronic device at a plurality of camera poses, and the plurality of camera poses are stored in the electronic device. The performing the global bundle adjustment to optimize the map includes adjusting the plurality of camera poses and the virtual pose to reduce the sum of the plurality of re-projection errors.

[0021] In some embodiments of the non-transitory computer-readable storage medium, the performing the global bundle adjustment to optimize the map includes calculating the sum of the plurality of re-projection errors. The calculating the sum of the plurality of re-projection errors includes: projecting the first estimated coordinate onto the plurality of real keyframes and the plurality of virtual frames having a plurality of first pixel points associated with the first map point to generate a plurality of first projected points; calculating each first re-projection error between each first pixel point and a corresponding one of the plurality of first projected points; projecting the second estimated coordinate onto the plurality of real keyframes and the plurality of virtual frames having a plurality of second pixel points associated with the second map point to generate a plurality of second projected points; calculating each second re-projection error between each second pixel point and a corresponding one of the plurality of second projected points; and adding the plurality of first re-projection errors, the plurality of second re-projection errors, and a plurality of other re-projection errors of a plurality of other map points of the map to generate the sum of the plurality of re-projection errors.

[0022] In some embodiments of the non-transitory computer-readable storage medium, the number of the plurality of virtual cameras is twelve, the optical axes of half of the plurality of virtual cameras intersect at a first intersection point coordinate, and the optical axes of the other half of the plurality of virtual cameras intersect at a second intersection point coordinate.

[0023] It is to be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further explanation of the disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0024] The foregoing and other objects, features, and advantages of the disclosure will be more apparent from the following detailed description, which proceeds with reference to the accompanying drawings.

[0025] Figure 1 A simplified functional block diagram of an electronic device according to an embodiment of the disclosure;

[0026] Figure 2 A schematic diagram of a real-world environment according to an embodiment of the disclosure;

[0027] Figure 3 A partial mapping schematic diagram according to an embodiment of the disclosure;

[0028] Figure 4 A flowchart of a map optimization method according to an embodiment of the disclosure;

[0029] Figure 5 A schematic diagram of a constraint used in a map optimization method according to an embodiment of the disclosure;

[0030] Figure 6 A schematic diagram of a virtual frame provided by a virtual camera according to an embodiment of the disclosure;

[0031] Figure 7 A schematic diagram of a sum of re-projection errors according to an embodiment of the disclosure;

[0032] Figure 8 A schematic diagram of a plurality of constraints used in a map optimization method according to an embodiment of the disclosure; and

[0033] Figure 9 A schematic diagram of a Hessian matrix according to an embodiment of the disclosure.

[0034] LEGEND

[0035] 10: tracking module

[0036] 11: partial mapping module

[0037] 12: marker detection module

[0038] 13: global optimization module

[0039] 14: closed loop module

[0040] 100: electronic device

[0041] 110: computing circuitry

[0042] 112: SLAM module

[0043] 120: storage circuitry

[0044] 122: map

[0045] 124_1-124_n: distance values

[0046] 130: communication circuitry

[0047] 140: camera circuitry

[0048] 150: display circuitry

[0049] 200: environment

[0050] 210-230: markers

[0051] 400: map optimization method

[0052] S410-S440: steps

[0053] Fo1-Fo3: camera poses

[0054] Kf1-Kf3: real keyframes

[0055] Kv1-Kv12: virtual frames

[0056] Px1, Px2, Pv1, Pv2: pixel points

[0057] Px1', Px2', Pv1': projected points

[0058] Mp1-Mp3: map points

[0059] O1, O2: ideal coordinates

[0060] O1'-O3': estimated coordinates

[0061] L1, L2: imaginary straight lines

[0062] R1, R2: virtual poses

[0063] S1-S4: intersection coordinates

[0064] V1-V24: virtual cameras

[0065] D1-D3: directions

[0066] A, B, C, D: sub-matrices DETAILED DESCRIPTION

[0067] The following disclosure provides many different embodiments, or examples, for implementing different characteristics of the provided subject matter. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments. It should be apparent, however, that the embodiments can be practiced in a wide variety of embodiments or examples, and that not all embodiments or examples

[0068] Figure 1 FIG. 1 is a simplified functional block diagram of an electronic device 100 according to an embodiment of the present disclosure. The electronic device 100 includes a computing circuit 110, a storage circuit 120, a communication circuit 130, a camera circuit 140, and a display circuit 150. The computing circuit 110 includes a Simultaneous Localization and Mapping (SLAM) module 112. In some embodiments, the SLAM module 112 is used to construct a three-dimensional (3D) point cloud map 122 (hereinafter referred to as "map 122"), which can be stored in the storage circuit 120. The computing circuit 110 can use the camera circuit 140 to capture features of the real world and recognize these features through information stored in the map 122, and then determine the precise position of the electronic device 100 through an inside-out tracking technique.

[0069] In some embodiments, the electronic device 100 includes an optical see-through system and / or a video see-through system for providing an Augmented Reality (AR) environment. The optical see-through system can project images of virtual objects into the field of view of a user (e.g., through the display circuit 150) while directly observing a real-world environment (e.g., through a transparent lens), thereby augmenting the real-world environment perceived by the user with virtual objects. The video see-through system captures images of a real-world environment (e.g., through the camera circuit 140) and provides these images to the user (e.g., through the display circuit 150) such that images of virtual objects can be superimposed onto the images of the real-world environment while directly observing the real world. In some embodiments, the electronic device 100 can provide a fully immersive Virtual Reality (VR) environment to the user (e.g., through the display circuit 150).

[0070] The map 122 can be used by the computing circuit 110 to render a virtual environment and / or virtual objects, e.g., a virtual apartment with virtual furniture based on the user's actual apartment. In some embodiments, the electronic device 100 is a standalone head-mounted device (HMD) with all the required hardware integrated in a single device. In other embodiments, part of the communication circuit 130, the camera circuit 140, and the display circuit 150 can be integrated in the HMD, while the other part of the communication circuit 130, the computing circuit 110, and the storage circuit 120 can be implemented in a device with logical computing capability, such as a personal computer, a laptop, etc. The two parts of the communication circuit 130 can communicate through wired or wireless means.

[0071] In some embodiments, the computing circuit 110 can include one or more processors (e.g., one or more general purpose single- or multi-chip processors), digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), other programmable logic devices, or a combination of such devices. The camera circuit 140 includes one or more cameras that can be mounted on the HMD. The display circuit 150 can be implemented through a see-through holographic display, a liquid crystal display (LCD), an active-matrix organic LED (AMOLED) display, etc.

[0072] Please refer to Figure 1 and Figure 2 wherein Figure 2 is a schematic diagram of a real-world environment 200 according to an embodiment of the present disclosure. Conventional mapping techniques, such as binocular depth estimation and monocular depth estimation, can have low accuracy in estimating features with long distances, and thus in some embodiments, the map 122 can be generated by performing multiple short-range local mapping. To generate the map 122 of the environment 200, the electronic device 100 can prompt the user wearing the HMD to move in the environment 200 to capture multiple images through the camera circuit 140. The tracking module 10 of the SLAM module 112 is configured to recognize pre-set types of real-world features from the captured images. In some embodiments, the pre-set types of real-world features can include the markers 210-230, edges and corners of furniture and / or appliances, etc., but the present disclosure is not limited thereto.

[0073] The images captured by the camera circuit 140 are referred to as "frames" in the following paragraphs. One or more frames captured by the camera circuit 140 at a certain time instance can be selected to form a "key frame", where the key frame contains a six-dimensional (6D) camera pose of the camera circuit 140 at the certain time instance. The recognition of real-world features performed by the tracking module 10 is used to select the frames for forming the key frames. In addition, the local mapping module 11 of the SLAM module 112 can generate a plurality of map points of the map 122 based on the key frames.

[0074] In some embodiments, the markers 210-230 can be one-dimensional barcodes, two-dimensional barcodes, or any combination thereof.

[0075] Figure 3 A local mapping diagram according to an embodiment of the present disclosure is shown. As the user moves in the environment 200, the six-dimensional camera poses Fo1-Fo3 (hereinafter referred to as "camera poses Fo1-Fo3") of the camera circuit 140 are stored in the storage circuit 120. Real key frames Kf1-Kf3 are captured by the camera circuit 140 when the camera circuit 140 is at the camera poses Fo1-Fo3, respectively. For ease of illustration, each of the real key frames Kf1-Kf3 is represented by one of the frames, but the present disclosure is not limited thereto. Figure 2 The markers 210-220 are recorded as pixel points Px1-Px2 in the real key frames Kf1-Kf3, respectively. In an ideal case, the local mapping module 11 can generate map points Mp1-Mp2 corresponding to the markers 210-220, respectively, according to the camera poses Fo1-Fo3 and the real key frames Kf1-Kf3. Since the focal length of the camera circuit 140, the camera poses Fo1-Fo3, and the pixel coordinates of the pixel points Px1-Px2 in the image plane are known parameters to the electronic device 100, the local mapping module 11 can calculate the three-dimensional ideal coordinates O1-O2 of the map points Mp1-Mp2 by triangulation.

[0076] However, mechanical errors of the gyroscope, the accelerometer, the motor, and the like, and noises in the real keyframes Kf1-Kf3 can affect the local mapping. Therefore, in actual situations, the local mapping module 11 generates map points Mp1-Mp2 with three-dimensional estimated coordinates O1'-O2' based on the camera poses Fo1-Fo3 and the real keyframes Kf1-Kf3, where the estimated coordinates O1'-O2' can deviate from the ideal coordinates O1-O2. The error of the estimated coordinates O1'-O2' and the ideal coordinates O1-O2 can be observed by projecting the estimated coordinates O1'-O2' onto the real keyframes Kf1-Kf3 to form projection points Px1'-Px2' corresponding to the estimated coordinates O1'-O2', respectively. The local mapping module 11 can perform a local bundle adjustment to adjust the estimated coordinates O1'-O2' and the stored camera poses Fo1-Fo3 to reduce the linear distance between the pixels Px1-Px2 and the projection points Px1'-Px2' in each of the real keyframes Kf1-Kf3 (i.e., to reduce the re-projection error of the map points Mp1-Mp2). Specifically, to perform the local bundle adjustment, the re-projection error of the map points Mp1-Mp2 can be described by a function of a non-linear least squares problem. A Gauss-Newton algorithm can be used to find increments to adjust each estimated coordinate to reduce the output value of such a function (i.e., the re-projection error). After a number of iterations, when the re-projection error is small enough, the estimated coordinates O1'-O2' have been sufficiently optimized.

[0077] During the multiple local mapping processes, the local mapping module 11 can generate a plurality of general map points, which are generated based on edges and corners of furniture and / or home appliances, instead of the markers 210-230. For brevity, the pixels corresponding to the general map points are omitted in the real keyframes Kf1-Kf3. In addition, the local mapping module 11 reduces the re-projection error of the general map points in the local bundle adjustment.

[0078] Figure 4This is a flowchart of a map optimization method 400 according to an embodiment of this disclosure. After multiple local mapping operations, SLAM module 112 can reduce the overall offset of map 122 by executing map optimization method 400, which includes global bundle adjustment. Any combination of features of map optimization method 400 or any other method described in this disclosure may be embodied in instructions stored in a non-transitory computer-readable storage medium. When these instructions are executed (e.g., by computing circuitry 110), they may cause some or all of these methods to be performed. It should be understood that any method discussed in this disclosure may contain more or fewer steps than shown in the flowchart, and these steps may be performed in any appropriate order.

[0079] Map optimization method 400 uses the data stored in the map. Figure 1 The distance values ​​124_1 to 124_n in the storage circuit 120 are used to optimize the map 122. Each of the distance values ​​124_1 to 124_n represents the real-world distance between two markers. For example, distance value 124_1 could indicate that markers 210 and 220 are 150 cm apart in environment 200. In some embodiments, the distance values ​​124_1 to 124_n can be input by a user to the electronic device 100 through any suitable user interface. In other embodiments, the distance values ​​124_1 to 124_n are generated by marker-attached devices and transmitted to the electronic device 100, wherein these devices can measure the distance between them using Time-of-Flight (ToF) technology, which is based on Wi-Fi, Bluetooth Low Energy (BLE), ultrasound, ZigBee, or any combination thereof.

[0080] Please refer to Figure 4 and Figure 5 , Figure 5 This is a schematic diagram of the constraints used in a map optimization method 400 according to an embodiment of this disclosure. In step 410, the marker detection module 12 of the SLAM module 112 identifies two map points generated based on markers from a plurality of map points on the map 122. For example, map points Mp1 to Mp2 generated based on markers 210 to 220 are identified from the map 122.

[0081] In step 420, the global optimization module 13 of the SLAM module 112 generates virtual cameras V1-V12 to optimize the estimated coordinates O1'-O2'. The virtual cameras V1-V12 are treated as real cameras in the global bundle adjustment described below. In other words, information related to the virtual cameras V1-V12 and information from the camera circuit 140 are used in a similar manner in the global bundle adjustment.

[0082] The optical axes of the virtual cameras V1-V6 are set to intersect at the intersection point coordinate S1. The optical axes of the virtual cameras V7-V12 are set to intersect at the intersection point coordinate S2. Specifically, the virtual cameras V1-V2 are configured face-to-face along a direction D1; the virtual cameras V3-V4 are configured face-to-face along a direction D2; and the virtual cameras V5-V6 are configured face-to-face along a direction D3. The virtual cameras V7-V8 are configured face-to-face along the direction D1; the virtual cameras V9-V10 are configured face-to-face along the direction D2; and the virtual cameras V11-V12 are configured face-to-face along the direction D3. In some embodiments, the directions D1-D3 are perpendicular to each other, but the present disclosure is not limited thereto. The virtual cameras V1-V12 have a fixed relative positional relationship with each other and share a six-dimensional virtual pose R1. Therefore, when the virtual pose R1 is adjusted in the global bundle adjustment, the virtual cameras V1-V12 can move and / or rotate with the virtual pose R1. In other words, the virtual pose R1 is used to control the configuration of the virtual cameras V1-V12.

[0083] Please refer to Figure 6 , Figure 6 FIG. 6 shows a schematic diagram of virtual frames Kv1-Kv12 provided by the virtual cameras V1-V12 according to an embodiment of the present disclosure. The virtual frames Kv1-Kv12 can be regarded as being generated by the virtual cameras V1-V12 at the virtual pose R1, and form virtual key frames of the virtual cameras V1-V12 used in the global bundle adjustment. The virtual frames Kv1-Kv6 provided by the virtual cameras V1-V6 form a cube that contains the intersection point coordinate S1, which is the center of the cube in some embodiments. The virtual frames Kv1-Kv6 are used to indicate to the global bundle adjustment algorithm that the marker 210 is observed by the virtual cameras V1-V6 at the intersection point coordinate S1. Specifically, six pixel points Pv1 corresponding to the marker 210 are formed at the intersection of the virtual frames Kv1-Kv6 and the imaginary straight lines from the positions of the virtual cameras V1-V6 to the intersection point coordinate S1.

[0084] Similarly, virtual frames Kv7-Kv12 provided by virtual cameras V7-V12 form a cube containing the intersection point coordinates S2, which in some embodiments is the center of this cube. Virtual frames Kv7-Kv12 are used to indicate to the global bundle adjustment algorithm that marker 220 is observed by virtual cameras V7-V12 at the intersection point coordinates S2. Specifically, six pixels Pv2 corresponding to marker 220 are formed at the intersection of virtual frames Kv7-Kv12 and an imaginary straight line from the position of virtual cameras V7-V12 to the intersection point coordinates S2.

[0085] The distance between the intersection point coordinates S1 and S2 is set to a distance value of 124_1 (i.e., the distance between markers 210 and 220) and remains unchanged during global bundle adjustment. Therefore, the constraint that correctly indicates the real-world distance between markers 210 and 220 is incorporated into the global bundle adjustment to aid in the optimization of the estimated coordinates O1' to O2'. In some embodiments, when virtual cameras V1 to V12 are initially generated, the intersection point coordinates S1 to S2 and the estimated coordinates O1' to O2' are aligned with each other on the imaginary line L1. In some embodiments, when virtual cameras V1 to V12 are initially generated, the intersection point coordinates S1 to S2 are located between the estimated coordinates O1' to O2', wherein the distance between the estimated coordinate O1' and the intersection point coordinate S1 is equal to the distance between the estimated coordinate O2' and the intersection point coordinate S2.

[0086] like Figure 6 As shown, in some embodiments, to maintain the distance 124_1 between the intersection point coordinates S1 and S2, virtual cameras facing the same direction are set to be spaced 124_1 apart, or virtual frames facing the same direction are set to be spaced 124_1 apart. For example, virtual cameras V4 and V10 facing the second direction D2 are spaced 124_1 apart. As another example, virtual frames Kv2 and Kv8 facing the first direction D1 are spaced 124_1 apart.

[0087] Please refer to Figure 4 and Figure 7 , Figure 7This is a schematic diagram illustrating the calculation of the sum of reprojection errors according to an embodiment of this disclosure. In steps S430-S440, the global optimization module 13 performs global bundle adjustment. In step S430, the sum of reprojection errors for all map points in map 122 is calculated. It is worth noting that the global optimization module 13 does not generate virtual cameras for general map points not generated based on markers 210-230. The reprojection errors of these general map points are related to the observations of camera circuitry 140, but not to the observations of virtual cameras. Those skilled in the art to which this disclosure pertains will understand the method used to calculate the reprojection errors of general map points; therefore, for the sake of brevity, a detailed description is omitted here. The following paragraphs will explain how the reprojection errors of map points generated based on markers 210-230 (e.g., map points Mp1 and Mp2 identified in step S410) are calculated.

[0088] Let's take the calculation of the reprojection error of map point Mp1 as an example. Figure 7 As shown, because the real-world keyframes Kf1-Kf3 and the virtual frames Kv1-Kv6 contain pixels Px1 and Pv1 related to map point Mp1, the global optimization module 13 projects the estimated coordinates O1' of map point Mp1 onto the real-world keyframes Kf1-Kf3 and the virtual frames Kv1-Kv6. Therefore, the projected point Px1' is generated on the real-world keyframes Kf1-Kf3, while the projected point Pv1' is generated on the virtual frames Kv1-Kv6. The projected point Pv1' is... Figure 7 The center is represented by a hollow circle for the sake of simplicity. Figure 7 The image only shows the reference features of pixel Pv1 and the projection point Pv1' on virtual frames Kv2 and Kv5. Next, the global optimization module 13 calculates the reprojection error (i.e., linear distance) between each pixel Px1 and the corresponding projection point Px1' on the same real keyframe, and also calculates the reprojection error between each pixel Pv1 and the corresponding projection point Pv1' on the same virtual frame.

[0089] Since the reprojection error of map points Mp1 to Mp2 can be calculated in a similar way, for the sake of simplicity... Figure 7 The estimated coordinates O2' of map point Mp2 and its corresponding virtual frames V7 to V12 are omitted. In short, the reprojection error of map point Mp2 is calculated by projecting the estimated coordinates O2' onto real-world keyframes Kf1 to Kf3 and virtual frames V7 to V12 containing pixels Px2 and Pv2 associated with map point MP2, thereby generating multiple projection points. Next, the global optimization module 13 calculates the reprojection error between each pixel Px2 and its corresponding projection point on the same real-world keyframe, and simultaneously calculates the reprojection error between each pixel Pv2 and its corresponding projection point on the same virtual frame.

[0090] The global optimization module 13 can add the re-projection error of the map point Mpl, the re-projection error of the map point Mp2, and the re-projection errors of other map points (e.g., general map points) to generate a total sum of re-projection errors. In other words, the global optimization module 13 can add the re-projection errors of all map points of the map 122 to generate the total sum of re-projection errors.

[0091] In step S440, the total sum of re-projection errors is reduced by adjusting the estimated coordinates of all map points in the map 122, including adjusting the estimated coordinates O1' ~ O2'. The total sum of re-projection errors can be represented as a function of a non-linear least squares problem. A Gauss-Newton algorithm can be used to find increments to adjust each estimated coordinate to reduce the output value (i.e., the re-projection error) of such a function.

[0092] Steps S430 ~ S440 can be repeatedly performed multiple times until the total sum of re-projection errors is less than or equal to a preset threshold value stored in the storage circuit 120. From Figure 6 and Figure 7 It can be seen that the estimated coordinate O1' is adjusted close to the intersection coordinate S1 during steps S430 ~ S440, because the intersection coordinate S1 is the position that minimizes the re-projection errors on the virtual frames Kv1 ~ Kv6. Similarly, the estimated coordinate O2' is adjusted close to the intersection coordinate S2 to minimize the re-projection errors on the virtual frames Kv7 ~ Kv12. Therefore, the distance between the estimated coordinates O1' ~ O2' is limited to be close to the distance value 124_1 after multiple iterations.

[0093] In some embodiments, the camera poses Fo1 ~ Fo3 and the virtual pose R1 are adjusted in step S440 to reduce the total sum of re-projection errors. It should be understood that the projection points are located at the intersection of the keyframes and the imaginary straight lines from the poses to the estimated coordinates, so adjusting one or both of the poses and the estimated coordinates helps to reduce the re-projection errors.

[0094] It is worth noting that general map points generated based on real-world features other than markers 210-230 are prone to accumulating errors during local mapping and, due to the lack of precise real-world constraints (e.g., lack of precise distance and / or depth information), are difficult to adjust to accurate positions using traditional global bundle adjustment algorithms. By performing map optimization method 400 and the global bundle adjustment included therein, real-world constraints (e.g., distance value 124_1) make the distance between map points Mp1-Mp2 substantially equal to the real-world distance between markers 210-220. Since the map points in map 122 are interconnected, optimizing the positions of map points Mp1-Mp2 will correspondingly optimize the positions of adjacent map points. Therefore, through map 122, electronic device 100 can perform accurate indoor positioning and / or render a virtual environment to precisely match the real-world environment.

[0095] Please refer to Figure 8 , Figure 8 This is a schematic diagram illustrating the various constraints used in a map optimization method 400 according to an embodiment of this disclosure. In some embodiments, step S410 may be performed multiple times to generate multiple sets of virtual cameras based on distance values ​​124_1 to 124_n, wherein the virtual cameras V1 to V12 mentioned above are one of the multiple sets of virtual cameras. Each set of virtual cameras is controlled by a virtual pose (i.e., shares a virtual pose) and is further divided into two groups, which are separated from each other by a corresponding one of the distance values ​​124_1 to 124_n. Each group contains six virtual cameras, similar to the virtual cameras V1 to V12 described above.

[0096] like Figure 8 As shown, for example, after the electronic device 100 generates map points Mp1 to Mp3 corresponding to markers 210 to 230 through local mapping and local bundle adjustment, the global optimization module 13 executes steps S410 to S420 multiple times to generate virtual cameras V1 to V12 based on distance value 124_1, and to generate multiple virtual cameras V13 to V24 based on distance value 124_2. Distance value 124_2 represents... Figure 2 The real-world distance between markers 220 and 230. Therefore, the optical axes of virtual cameras V13 to V24 are set to intersect at the intersection point coordinates S3 to S4, which are 124_2 apart, and the virtual frame-directed global bundle adjustment algorithm provided by virtual cameras V13 to V24 indicates that markers 220 and 230 are observed at the intersection point coordinates S3 to S4.

[0097] In some embodiments, when generating virtual cameras V13 to V24, the intersection point coordinates S3 to S4, the estimated coordinates O2' of map point Mp2, and the estimated coordinates O3' of map point Mp3 are aligned with each other on the imaginary straight line L2. In some embodiments, when generating virtual cameras V13 to V24, the intersection point coordinates S3 to S4 are located between the estimated coordinates O2' and O3', and the distance between the estimated coordinates O2' and the intersection point coordinates S3 is equal to the distance between the estimated coordinates O3' and the intersection point coordinates S4.

[0098] In some embodiments, the number of virtual camera groups can be equal to the number of distance values ​​124_1 to 124_n stored in the storage circuit 120. In other words, the global optimization module 13 can generate n virtual cameras to incorporate n real-world constraints (e.g., distance values ​​124_1 to 124_n) into the global bundle adjustment algorithm to optimize map 122, where n is a positive integer. Steps S430 to S440 can be performed using all the information provided by the multiple virtual cameras.

[0099] like Figure 8 As shown, for example, virtual cameras V13 to V24 are controlled by virtual pose R2, and virtual poses R1 and R2 can be adjusted during the execution of steps S430 to S440 so that the intersection point coordinates S2 to S3 and the estimated coordinates O2' are close to each other, thereby reducing the reprojection error of map point Mp2 on the virtual frames provided by virtual cameras V7 to V12 and V13 to V18.

[0100] Please refer to Figure 9 , Figure 9 This is a schematic diagram of a Hessian matrix according to an embodiment of this disclosure. In step S440, the following Equation 1 needs to be calculated when using the Gauss-Newton algorithm, where "H" is the Hessian matrix, "x" is the increment, and "b" is the Jacobian matrix.

[0101] Hx = b (Formula 1)

[0102] The Hansen matrix comprises submatrices A, B, C, and D. Since the Hansen matrix is ​​typically quite large, the Schur complement is frequently used to reduce the computational complexity of Equation 1, but it only applies when submatrices A and B are diagonal matrices. To ensure that submatrices A and B are diagonal matrices, map optimization method 400 incorporates a constraint from keyframes to map points as a real-world constraint, rather than a constraint from map points to map points. For example, as... Figure 6As shown, the way the distance value 124_1 is brought into the global bundle adjustment algorithm is by setting two corresponding virtual frames (or two corresponding positions of a virtual camera) that are separated by the distance value 124_1, rather than directly setting the intersection coordinates S1-S2 to be separated by the distance value 124_1. Thus, the map optimization method 400 helps to generate an accurate three-dimensional map with less computation time.

[0103] In addition, in some embodiments, Figure 1 The SLAM module 112 further includes a loop closure module 14 configured to determine whether the electronic device 100 returns to an area where the electronic device 100 was previously located after moving a predetermined distance or after a predetermined time. When the loop closure module 14 detects a loop closure, the loop closure module 14 merges a new portion of the map 122 with an old portion of the map 122 having similar map points to reduce deformation of the map 122.

[0104] In some embodiments, the tracking module 10, the local mapping module 11, the landmark detection module 12, the global optimization module 13, and the loop closure module 14 can be implemented by hardware, software, or any combination thereof, as appropriate.

[0105] In the description and claims, certain terms have been used for brevity, clarity, and understanding. No unnecessary limitations are to be implied therefrom beyond the require limitations set forth in the claims. The terms "a" and "an," as used in the context of this specification, are to be construed to cover both the singular and the plural, unless otherwise indicated. The term "or" as used in the context of this specification is to be construed as inclusive or unless otherwise indicated. The terms "associated with" and "associated therewith" as used in the context of this specification are to be construed as inclusive of both direct and indirect associations.

[0106] The term "and / or" as used in the context of this specification is to be construed as inclusive of any one or more of the associated listed items. In addition, the term "one" or "a" as used in the context of this specification is to be construed as including "at least one" unless otherwise indicated.

[0107] Those skilled in the relevant technology will appreciate that various changes and modifications can be made to the structures of the present disclosure without departing from the spirit and scope of the present disclosure. In light of the above, the present disclosure is intended to cover the changes and modifications in the present disclosure, as long as they fall within the scope of the claims and their equivalents.

Claims

1. A map optimization method, characterized in that, An electronic device suitable for storing multiple distance values ​​and a map, comprising: A first map point is identified from the map using a first estimated coordinate generated based on a first marker, and a second map point is identified from the map using a second estimated coordinate generated based on a second marker; Multiple virtual cameras are generated based on the first estimated coordinates and the second estimated coordinates. The multiple virtual cameras are controlled by a virtual pose and include multiple optical axes that intersect at a first intersection point coordinate and a second intersection point coordinate. The distance between the first intersection point coordinate and the second intersection point coordinate is a first distance value among the multiple distance values. The multiple virtual cameras provide multiple virtual frames, which are used to indicate that the first mark and the second mark are observed at the first intersection point coordinate and the second intersection point coordinate, respectively. as well as Perform a global bundle adjustment to optimize the map, including adjusting the first estimated coordinates and the second estimated coordinates to reduce the sum of multiple reprojection errors calculated based on multiple real keyframes and multiple virtual frames, wherein the electronic device is used for positioning based on the optimized map.

2. The map optimization method as described in claim 1, characterized in that, These multiple real-world keyframes are captured by the electronic device at multiple camera poses, and these multiple camera poses are stored in the electronic device. The global bundle adjustment to optimize the map further includes: Adjust the poses of the multiple cameras and the virtual pose to reduce the sum of the multiple reprojection errors.

3. The map optimization method as described in claim 1, characterized in that, Performing this global bundle adjustment to optimize the map further includes calculating the sum of the multiple reprojection errors, wherein calculating the sum of the multiple reprojection errors includes: The first estimated coordinates are projected onto the plurality of real-world keyframes and the plurality of virtual frames that have a plurality of first pixels associated with the first map point, to generate a plurality of first projection points; Calculate each first reprojection error between each first pixel and its corresponding one of the plurality of first projection points; The second estimated coordinates are projected onto the plurality of real-world keyframes and the plurality of virtual frames that have a plurality of second pixels associated with the second map point, to generate a plurality of second projection points; Calculate each second projection error between each second pixel and its corresponding one of the plurality of second projection points; as well as The multiple first reprojection errors, the multiple second reprojection errors, and the multiple other reprojection errors of the map's other map points are added together to generate the sum of the multiple reprojection errors.

4. The map optimization method as described in claim 1, characterized in that, The number of virtual cameras is twelve. Half of the multiple optical axes of the multiple virtual cameras intersect at the first intersection point coordinates, and the other half of the multiple optical axes of the multiple virtual cameras intersect at the second intersection point coordinates.

5. The map optimization method as described in claim 1, characterized in that, When the multiple virtual cameras are generated, the coordinates of the first intersection point, the coordinates of the second intersection point, the first estimated coordinates, and the second estimated coordinates are aligned with each other on an imaginary straight line.

6. The map optimization method as described in claim 5, characterized in that, When the multiple virtual cameras are generated, the distance between the first estimated coordinate and the first intersection point coordinate is equal to the distance between the second estimated coordinate and the second intersection point coordinate.

7. The map optimization method as described in claim 1, characterized in that, The first distance value represents a real-world distance between the first marker and the second marker.

8. The map optimization method as described in claim 1, characterized in that, Each distance value represents a real-world distance between two of a plurality of markers, including the first marker and the second marker. The electronic device is used to generate multiple sets of virtual cameras, and these multiple virtual cameras are one of the multiple sets of virtual cameras. During the global bundle adjustment, each virtual camera group is divided into two groups, which are separated from each other and at a distance from one of the corresponding distance values, so as to bring the distance values ​​into the global bundle adjustment.

9. An electronic device, characterized in that, Include: A storage circuit for storing multiple distance values ​​and a map; and A computing circuit, used for: A first map point is identified from the map using a first estimated coordinate generated based on a first marker, and a second map point is identified from the map using a second estimated coordinate generated based on a second marker; Multiple virtual cameras are generated based on the first estimated coordinates and the second estimated coordinates. The multiple virtual cameras are controlled by a virtual pose and include multiple optical axes that intersect at a first intersection point coordinate and a second intersection point coordinate. The distance between the first intersection point coordinate and the second intersection point coordinate is a first distance value among the multiple distance values. The multiple virtual cameras provide multiple virtual frames, which are used to indicate that the first mark and the second mark are observed at the first intersection point coordinate and the second intersection point coordinate, respectively. as well as Perform a global bundle adjustment to optimize the map, including adjusting the first estimated coordinates and the second estimated coordinates to reduce the sum of multiple reprojection errors calculated based on multiple real keyframes and multiple virtual frames, wherein the electronic device is used for positioning based on the optimized map.

10. The electronic device as claimed in claim 9, characterized in that, The multiple real-world keyframes are captured by the electronic device at multiple camera poses, and the storage circuit is used to store these multiple camera poses. When performing global bundle adjustment to optimize the map, the computational circuit is further used for: Adjust the poses of the multiple cameras and the virtual pose to reduce the sum of the multiple reprojection errors.

11. The electronic device as claimed in claim 9, characterized in that, When the global bundle adjustment is performed to optimize the map, the computational circuitry performs the following steps to calculate the sum of the multiple reprojection errors: The first estimated coordinates are projected onto the plurality of real-world keyframes and the plurality of virtual frames that have a plurality of first pixels associated with the first map point, to generate a plurality of first projection points; Calculate each first reprojection error between each first pixel and its corresponding one of the plurality of first projection points; The second estimated coordinates are projected onto the plurality of real-world keyframes and the plurality of virtual frames that have a plurality of second pixels associated with the second map point, to generate a plurality of second projection points; Calculate each second projection error between each second pixel and its corresponding one of the plurality of second projection points; as well as The multiple first reprojection errors, the multiple second reprojection errors, and the multiple other reprojection errors of the map's other map points are added together to generate the sum of the multiple reprojection errors.

12. The electronic device as claimed in claim 9, characterized in that, The number of virtual cameras is twelve. Half of the multiple optical axes of the multiple virtual cameras intersect at the first intersection point coordinates, and the other half of the multiple optical axes of the multiple virtual cameras intersect at the second intersection point coordinates.

13. The electronic device as claimed in claim 9, characterized in that, When the multiple virtual cameras are generated, the coordinates of the first intersection point, the coordinates of the second intersection point, the first estimated coordinates, and the second estimated coordinates are aligned with each other on an imaginary straight line.

14. The electronic device as claimed in claim 13, characterized in that, When the multiple virtual cameras are generated, the distance between the first estimated coordinate and the first intersection point coordinate is equal to the distance between the second estimated coordinate and the second intersection point coordinate.

15. The electronic device as claimed in claim 9, characterized in that, The first distance value represents a real-world distance between the first marker and the second marker.

16. The electronic device as claimed in claim 9, characterized in that, Each distance value represents a real-world distance between two of a plurality of markers, including the first marker and the second marker. The computing circuit is used to generate multiple sets of virtual cameras, and these multiple virtual cameras are one of the multiple sets of virtual cameras. During the global bundle adjustment, each virtual camera group is divided into two groups, which are separated from each other and at a distance from one of the corresponding distance values, so as to bring the distance values ​​into the global bundle adjustment.

17. A non-transitory computer-readable storage medium, characterized in that, The system stores multiple computer-readable instructions for controlling an electronic device, which includes a computing circuit and a storage circuit. The storage circuit stores multiple distance values ​​and a map. When the multiple computer-readable instructions are executed by the computing circuit, the computing circuit is used to: A first map point is identified from the map using a first estimated coordinate generated based on a first marker, and a second map point is identified from the map using a second estimated coordinate generated based on a second marker; Multiple virtual cameras are generated based on the first estimated coordinates and the second estimated coordinates. The multiple virtual cameras are controlled by a virtual pose and include multiple optical axes that intersect at a first intersection point coordinate and a second intersection point coordinate. The distance between the first intersection point coordinate and the second intersection point coordinate is a first distance value among the multiple distance values. The multiple virtual cameras provide multiple virtual frames, which are used to indicate that the first mark and the second mark are observed at the first intersection point coordinate and the second intersection point coordinate, respectively. as well as Perform a global bundle adjustment to optimize the map, including adjusting the first estimated coordinates and the second estimated coordinates to reduce the sum of multiple reprojection errors calculated based on multiple real keyframes and multiple virtual frames, wherein the electronic device is used for positioning based on the optimized map.

18. The non-transitory computer-readable storage medium as described in claim 17, characterized in that, The multiple real-world keyframes are captured by the electronic device at multiple camera poses, and these multiple camera poses are stored in the electronic device. The global bundle adjustment to optimize the map includes: Adjust the poses of the multiple cameras and the virtual pose to reduce the sum of the multiple reprojection errors.

19. The non-transitory computer-readable storage medium as described in claim 17, characterized in that, Performing this global bundle adjustment to optimize the map involves calculating the sum of the multiple reprojection errors, wherein calculating the sum of the multiple reprojection errors includes: The first estimated coordinates are projected onto the plurality of real-world keyframes and the plurality of virtual frames that have a plurality of first pixels associated with the first map point, to generate a plurality of first projection points; Calculate each first reprojection error between each first pixel and its corresponding one of the plurality of first projection points; The second estimated coordinates are projected onto the plurality of real-world keyframes and the plurality of virtual frames that have a plurality of second pixels associated with the second map point, to generate a plurality of second projection points; Calculate each second projection error between each second pixel and its corresponding one of the plurality of second projection points; as well as The multiple first reprojection errors, the multiple second reprojection errors, and the multiple other reprojection errors of the map's other map points are added together to generate the sum of the multiple reprojection errors.

20. The non-transitory computer-readable storage medium as described in claim 17, characterized in that, The number of virtual cameras is twelve. Half of the multiple optical axes of the multiple virtual cameras intersect at the first intersection point coordinates, and the other half of the multiple optical axes of the multiple virtual cameras intersect at the second intersection point coordinates.

Citation Information

Patent Citations

  • HMD calibration with direct geometric modeling

    CN105320271A

  • Slam map joining method and system

    EP3886048A1