Position estimation method
The method enhances Visual SLAM accuracy for self-position estimation in vehicles by dividing images into frames, estimating camera pose and depth, and correcting poses with landmark positions, thereby reducing computational load.
Patent Information
- Application Number
- JP2024004009
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-28
AI Technical Summary
Existing Visual SLAM methods for self-position estimation in vehicles face challenges in achieving high accuracy while minimizing computational load.
A method involving dividing captured images into frame images, estimating camera pose and depth using depth-based Visual SLAM, generating map points, and correcting camera pose using landmark positions to reduce computational load.
Accurate self-position estimation is achieved with reduced computational resources by separating map point generation from Visual SLAM processing and utilizing landmark corrections.
Smart Images

Figure 2025110207000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of position estimation methods.
Background Art
[0002] As a method of this kind, for example, a method of updating a part of a high-resolution map based on a plurality of image frames captured by a camera equipped on a vehicle has been proposed (see Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] For example, SLAM (Simultaneous Localization and Mapping) may be used to estimate the self-position of a moving object such as an automobile. As one aspect of SLAM, Visual SLAM using an image sensor (for example, a camera) has been proposed. For example, for the purpose of improving the estimation accuracy related to Visual SLAM, a process with a relatively high computational load may be performed.
[0005] In view of the above circumstances, the present invention has been made, and an object thereof is to propose a position estimation method capable of accurately estimating the self-position using Visual SLAM while suppressing an increase in the computational load.
Means for Solving the Problems
[0006] The position estimation method according to one aspect of the present invention includes a step of dividing an image captured by an in-vehicle camera into a plurality of frame images, a step of estimating a camera pose and a depth related to the in-vehicle camera for each key frame included in the plurality of frame images by depth-based Visual SLAM, a step of acquiring the estimated camera pose and the estimated depth of each key frame, and a step of generating map points using the estimated depth.
Brief Description of Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Modes for Carrying Out the Invention
[0008] Embodiments related to the position estimation method will be described with reference to FIGS. 1 to 4. In FIG. 1, the vehicle 1 includes an in-vehicle camera 11, a GPS (Global Positioning System) 12, and a position estimation device 100. The in-vehicle camera 11 may be at least one of a monocular camera, a depth camera (e.g., an RGB-D camera), and a stereo camera. The vehicle 1 may be provided with one camera or a plurality of cameras as the in-vehicle camera 11. Note that the in-vehicle camera 11 may be installed inside the vehicle 1 (e.g., in the passenger compartment) or on the outer surface of the vehicle 1 (e.g., on the roof). Note that various existing modes can be applied to the GPS 12 (e.g., a GPS receiver or a GPS module). Therefore, the detailed description of the GPS 12 will be omitted.
[0009] The position estimation device 100 includes an arithmetic device 110 and a storage device 120. The arithmetic device 110 may include at least one of, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), and a TPU (Tensor Processing Unit). The storage device 120 may include at least one of, for example, a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and an optical disk array.
[0010] The in-vehicle camera 11 may generate an image (in other words, a moving image) by photographing the surroundings of the vehicle 1. The generated image may be stored in the storage device 120 of the position estimation device 100.
[0011] The operation of the position estimation device 100 will be described with reference to the flowchart of FIG. 2. The arithmetic unit 110 of the position estimation device 100 divides the video generated by the in-vehicle camera 11 into a plurality of frame images. In FIG. 2, the arithmetic unit 110 performs processing related to Visual SLAM on the plurality of frame images (step S1). In the processing of step S1, the arithmetic unit 110 may estimate the camera pose of the in-vehicle camera 11 for each key frame included in the plurality of frame images. Here, the camera pose means the position and orientation of the in-vehicle camera 11 in the SLAM coordinate system related to the in-vehicle camera 11. Specific examples of the algorithm related to Visual SLAM include, but are not limited to, DROID-SLAM and ORB-SLAM. The "key frame" may mean one or more frame images extracted according to the criteria related to Visual SLAM among the plurality of frame images. The criteria related to Visual SLAM may be, for example, that the number of common feature points between one frame image and another frame image that is temporally later than the one frame image is equal to or less than a predetermined number. In addition, the arithmetic unit 110 may thin out frame images that do not meet the criteria related to Visual SLAM (for example, frame images with little change in features from one frame image) from the plurality of frame images. In this case, the "key frame" may mean the frame images that remain without being thinned out.
[0012] In the process of estimating the camera pose of the in-vehicle camera 11 in the processing of step S1, the depth of each key frame is estimated. The camera pose estimated in the processing of step S1 is appropriately referred to as the "estimated camera pose". Similarly, the depth estimated in the processing of step S1 is appropriately referred to as the "estimated depth". As a result of the processing of step S1, the arithmetic unit 110 acquires the estimated camera pose and the estimated depth for each key frame. For example, the arithmetic unit 11 may associate a key frame with the estimated camera pose and the estimated depth related to the key frame. Note that the camera pose may be represented by, for example, the position and orientation of the in-vehicle camera 11.
[0013] Next, the arithmetic unit 110 generates map points using the estimated camera pose and the estimated depth (step S2). The process of step S2 will be described with reference to FIGS. 3 and 4.
[0014] In FIG. 3, the arithmetic unit 110 selects two key frames that are temporally close (step S201). For example, when arranging a plurality of key frames in time series, the arithmetic unit 110 may select a key frame at a certain point in time and a key frame immediately after the key frame as the two key frames.
[0015] Of the two key frames selected in the process of step S201, the key frame earlier in time is appropriately referred to as the "left frame", and the key frame later in time is appropriately referred to as the "right frame". For example, as shown in FIG. 4, it is assumed that there are a plurality of key frames KF1, KF2, KF3,..., KFn. For example, in the process of step S201, when key frames KF1 and KF2 are selected, the key frame KF1 may be referred to as the left frame, and the key frame KF2 may be referred to as the right frame. For example, in the process of step S201, when key frames KF2 and KF3 are selected, the key frame KF2 may be referred to as the left frame, and the key frame KF3 may be referred to as the right frame.
[0016] By selecting two key frames in the process of step S201, the arithmetic unit 110 acquires the estimated camera pose and the estimated depth related to the two selected key frames (see "estimated camera pose of the left frame", "estimated camera pose of the right frame", "estimated depth of the left frame", and "estimated depth of the right frame" in FIG. 3).
[0017] The arithmetic unit 110 performs feature point matching (step S202) using the two key frames selected in the process of step S201 and the estimated camera poses related to the two key frames. For example, the arithmetic unit 110 extracts feature points from each of the two selected key frames. Note that existing feature point extraction methods such as SuperPoint may be applied to the extraction of feature points. Next, the arithmetic unit 110 matches the feature points extracted from one of the two selected key frames with the feature points extracted from the other key frame. Note that existing matching methods such as SuperGlue may be applied to the matching of feature points.
[0018] For example, in FIG. 4, assume that the key frame KF1 includes feature points FP11, FP12, and FP13. The feature points FP11, FP12, and FP13 are assumed to be feature points corresponding to points P1, P2, and P11 in the three-dimensional space, respectively. Assume that the key frame KF2 includes feature points FP21, FP22, and FP23. The feature points FP21, FP22, and FP23 are assumed to be feature points corresponding to points P1, P2, and P3 in the three-dimensional space, respectively.
[0019] For example, in the process of step S202, the arithmetic unit 110 may associate the feature point FP11 included in the key frame KF1 with the feature point FP21 included in the key frame KF2. Also, in the process of step S202, the arithmetic unit 110 may associate the feature point FP12 included in the key frame KF1 with the feature point FP22 included in the key frame KF2. Note that the feature point corresponding to the feature point FP13 included in the key frame KF1 is not included in the key frame KF2. Therefore, the arithmetic unit 110 does not have to associate the feature point FP13 with the feature points included in the key frame KF2.
[0020] For example, in the process of step S202, the arithmetic unit 110 may extract the feature points FP11 and FP21 associated with each other, and the feature points FP12 and FP22, as a feature point pair. The feature point pair extracted in the process of step S202 may be referred to as a "corresponding pixel pair group".
[0021] Next, the arithmetic unit 110 calculates the three-dimensional coordinates of each of the two feature points associated with each other in the process of step S202, using the estimated depth related to the two key frames selected in the process of step S201 (step S203). Note that the three-dimensional coordinates calculated in step S203 may be coordinates in the SLAM coordinate system.
[0022] Next, the arithmetic unit 110 may calculate the distance between the two feature points based on the three-dimensional coordinates of each of the two feature points associated with each other. Then, the arithmetic unit 110 determines whether or not the calculated distance is less than or equal to a predetermined value (step S204). In other words, the arithmetic unit 110 determines whether or not the calculated distance is short.
[0023] Note that the predetermined value is a value for determining whether or not the two feature points associated with each other correspond to the same point in the three-dimensional space. The predetermined value may be a fixed value determined in advance, or may be a variable value according to some physical quantity or parameter. The predetermined value may be set, for example, as follows. First, when an object is photographed with different camera poses, the range in which the distance between two feature points each included in the two images and corresponding to a part of the one object can take may be obtained. The maximum value of the obtained range may be set as the predetermined value.
[0024] For example, when the feature point FP11 and the feature point FP21 shown in FIG. 4 are associated with each other, since the feature points FP11 and FP21 correspond to the point P1, in the process of step S204, it may be determined that the distance is equal to or less than a predetermined value (in other words, the distance is close). If, as shown in FIG. 4, the feature point FP11 and the feature point FP22 are associated with each other (that is, when the process of step S202 described above is inaccurate), the distance between the feature point FP11 corresponding to the point P1 and the feature point FP22 corresponding to the point P2 becomes relatively large. In this case, in the process of step S204, it may be determined that the distance is not equal to or less than a predetermined value (in other words, the distance is far). For example, even when the feature point FP11 corresponding to the point P1 and the feature point FP21 shown in FIG. 4 are associated with each other, if the estimated depth related to at least one of the key frames KF1 and KF2 is incorrect, the distance between the feature point FP11 and the feature point FP21 may become relatively large. In this case, in the process of step S204, it may be determined that the distance is not equal to or less than a predetermined value (in other words, the distance is far).
[0025] In the process of step S204, when it is determined that the distance is equal to or less than a predetermined value (step S204: Yes), the arithmetic unit 110 may select, as map points, the points corresponding to the two feature points associated with each other. On the other hand, in the process of step S204, when it is determined that the distance is not equal to or less than a predetermined value (step S204: No), the arithmetic unit 110 may discard the two feature points associated with each other. The arithmetic unit 110 may perform the processes of steps S203 and S204 for all of the extracted feature point pairs in the process of step 202.
[0026] After that, the arithmetic unit 110 determines whether or not the above-described processing has been completed for all of the plurality of key frames included in the video generated by the in-vehicle camera 11 (step S205). In the process of step S205, if it is determined that the processing has been completed (step S205: Yes), the process of step S3 in FIG. 2 is performed. On the other hand, in the process of step S205, if it is determined that the processing has not been completed (step S205: No), the process of step S201 is performed.
[0027] For example, for the key frames KF1 and KF2 shown in FIG. 4, after the processes of steps S201 to S204 are performed, if it is determined in the process of step S205 that the processing has not been completed (step S205: No), in the process of step S201, the arithmetic unit 110 may select the key frames KF2 and KF3.
[0028] For example, in FIG. 4, it is assumed that the key frame KF3 includes the feature points FP31, FP32, and FP33. It is assumed that the feature points FP31, FP32, and FP33 correspond to the points P2, P3, and P12 in the three-dimensional space, respectively. For example, when the processes of steps S201 to S204 are performed for the key frames KF1 and KF2, the points P1 and P2 in FIG. 4 may be selected as map points. After that, when the processes of steps S201 to S204 are performed for the key frames KF2 and KF3, the points P2 and P3 in FIG. 4 may be selected as map points.
[0029] Here, the feature point corresponding to the point P3 is not included in the key frame KF1. Therefore, the point P3 is not selected as a map point when the processes of steps S201 to S204 are performed for the key frames KF1 and KF2. After that, when the processes of steps S201 to S204 are performed for the key frames KF2 and KF3, the point P3 may be selected (added) as a new map point.
[0030] By repeatedly performing the processes of steps S201 to S205 described above, the correspondence relationship between map points common to two or more key frames and the key frames on which the map points are projected is collected.
[0031] Returning to FIG. 2, the arithmetic unit 110 performs bundle adjustment based on the map points generated in the process of step S2 (step S3). In the process of step S3, the position and angle (i.e., posture) of the in-vehicle camera 11 are adjusted so as to minimize the deviation (i.e., reprojection error) between the reprojection position of the map point and the position of the feature point corresponding to the map point in the key frame. Note that since various existing modes can be applied to bundle adjustment, the detailed description thereof is omitted.
[0032] After the process of step S3, the arithmetic unit 110 may extract landmarks that appear in two or more key frames from the landmark database. The landmark database may be stored in the storage device 120. Here, the landmark database may store the three-dimensional coordinates of features (e.g., traffic lights, etc.) that are landmarks registered on the map and the positions on the frame image when the landmarks are photographed by the in-vehicle camera 11. For example, the three-dimensional coordinates of the landmarks registered on the map may be represented by "latitude, longitude, altitude". The arithmetic unit 110 may convert the three-dimensional coordinates of the landmarks registered on the map into coordinates in the SLAM coordinate system based on, for example, the position of the vehicle 1 (or the in-vehicle camera 11) represented by "latitude, longitude, altitude" detected by the GPS 12 and the position of the in-vehicle camera 11 in the SLAM coordinate system.
[0033] The landmark database may be constructed as follows. For example, for each frame image of the video captured by the in-vehicle camera 11, an object detection algorithm may be applied to detect the detected objects related to the landmarks in each frame image. Next, by using an object tracking algorithm, when the same object appears on each frame image, it may be specified that they are the same object. Next, the three-dimensional coordinates of the detected object may be obtained by using the camera pose of each frame image. Next, the three-dimensional coordinates of the detected object may be associated with the three-dimensional coordinates of the landmarks already registered in the map. By such processing, the landmark database may be constructed.
[0034] The arithmetic device 110 treats the landmark in the same way as the map point described above, and corrects the position and angle (i.e., the pose) of the in-vehicle camera 11 adjusted in the process of step S3 by performing bundle adjustment based on the landmark (step S4).
[0035] Referring to the following formula (1), the position estimation method will be further described.
[0036]
Equation
[0037] In formula (1), "E" indicates an edge, "t i " indicates the camera position in frame i, "r i " indicates the camera pose in frame i, "p i " indicates the position of the map point in frame i, "F" indicates a set of frames, "S i " indicates a set of feature points reflected on the image of frame i, "q ij " indicates the coordinates on the image of feature point j in frame i, "proj(t i ,r i ,p j )" projects feature point j with camera pose t i ,r ishows the coordinates reprojected onto the image using, and "K" i " represents the set of landmarks projected onto the image of frame i, and "q" ik " represents the coordinates of landmark k on the image in frame i, and "proj(t" i , r" i , p" k )" represents the coordinates obtained by reprojecting landmark k onto the image using camera pose t i , r" i .
[0038] In the processing of step S3 described above, bundle adjustment is performed to minimize the first term on the right side of equation (1). Note that in the processing of step S3, since bundle adjustment is performed on the map points generated in the processing of step S2, the second term on the right side of equation (1) is not considered. In the processing of step S4 described above, based on the landmarks, bundle adjustment is performed to minimize the second term on the right side of equation (1).
[0039] (Technical Effect) In the above-described embodiment, the camera pose (in other words, the position and orientation of in-vehicle camera 11) related to the key frame estimated by depth-based Visual SLAM is corrected using the position information of the landmarks on the map. Therefore, the accuracy of the corrected camera pose is higher than the accuracy of the camera pose estimated by Visual SLAM. That is, according to the position estimation method according to the above-described embodiment, an accurate self-position can be obtained. In addition, compared with a method that performs optimization by adjusting the entire depth, for example, the position estimation method can reduce the optimization parameters, so that resource saving and calculation time shortening can be achieved. Therefore, according to the position estimation method, self-position estimation can be accurately performed using Visual SLAM while suppressing an increase in the calculation load.
[0040] As is clear from the flowchart of FIG. 2, the processing related to Visual SLAM and the processing related to the generation of map points are performed separately. If, as part of the processing related to Visual SLAM, the processing related to the generation of map points is performed, it is necessary to change the processing related to the generation of map points according to the algorithm related to Visual SLAM. On the other hand, if the processing related to the generation of map points is independent of the processing related to Visual SLAM, even if the algorithm related to Visual SLAM is changed, it is not necessary to change the processing related to the generation of map points. For this reason, the position estimation method is applicable, for example, to existing position estimation devices that use Visual SLAM. That is, an existing position estimation device can be updated to a device that executes the position estimation method.
[0041] In Visual SLAM, a technique called loop closing, which corrects the position when a point that has been photographed once is photographed again, may be used. However, when a vehicle (for example, vehicle 1) is running, it is rare to take a driving route in which loops frequently occur. That is, it is difficult to apply loop closing when estimating the self-position of an in-vehicle camera (for example, in-vehicle camera 11). As described above, in the position estimation method, since the camera pose is corrected using the position information of the landmark, the self-position of the in-vehicle camera can be accurately estimated without applying loop closing.
[0042] (Modification example) In the above-described embodiment, depth-based Visual SLAM has been described, but it is not limited thereto. For example, Visual SLAM may be feature-based Visual SLAM. That is, in the processing of step S1 described above, at least one of, for example, features and feature vectors may be estimated instead of depth.
[0043] Regarding the process of step S2 in this case, it will be described with reference to the flowchart of FIG. 5. In FIG. 5, the arithmetic unit 110 selects two keyframes that are temporally close (step S211). By selecting two keyframes in the process of step S211, the arithmetic unit 110 obtains the estimated camera poses and estimated feature amounts related to the two selected keyframes (see "estimated camera pose of the left frame", "estimated camera pose of the right frame", "estimated feature amount of the left frame", and "estimated feature amount of the right frame" in FIG. 5).
[0044] Using the two keyframes selected in the process of step S211 and the estimated camera poses related to the two keyframes, the arithmetic unit 110 performs feature point matching (step S212).
[0045] Next, using the estimated feature amounts related to the two keyframes selected in the process of step S211, the arithmetic unit 110 calculates the three-dimensional coordinates of each of the two feature points associated with each other in the process of step S212 (step S213).
[0046] Next, based on the three-dimensional coordinates of each of the two feature points associated with each other, the arithmetic unit 110 may calculate the distance between the two feature points. Then, the arithmetic unit 110 determines whether the calculated distance is less than or equal to a predetermined value (step S214). In other words, the arithmetic unit 110 determines whether the calculated distance is short.
[0047] In the process of step S204, if it is determined that the distance is less than or equal to the predetermined value (step S214: Yes), the arithmetic unit 110 may select the points corresponding to the two feature points associated with each other as map points. On the other hand, in the process of step S214, if it is determined that the distance is not less than or equal to the predetermined value (step S214: No), the arithmetic unit 110 may discard the two feature points associated with each other. The arithmetic unit 110 may perform the processes of steps S213 and S214 for all of the extracted feature point pairs in the process of step 212.
[0048] Thereafter, the arithmetic unit 110 determines whether or not the above-described processing has been completed for all of the plurality of key frames included in the video generated by the in-vehicle camera 11 (step S215). In the process of step S215, if it is determined that the processing has been completed (step S215: Yes), the process of step S3 in FIG. 2 is performed. On the other hand, in the process of step S215, if it is determined that the processing has not been completed (step S215: No), the process of step S211 is performed.
[0049] Aspects of the invention derived from the embodiments and modifications described above will be described below.
[0050] A position estimation method according to an aspect of the invention includes a step of dividing a video captured by an in-vehicle camera into a plurality of frame images, a step of estimating a camera pose and a depth related to the in-vehicle camera for each key frame included in the plurality of frame images by depth-based Visual SLAM, a step of acquiring the estimated camera pose and the estimated depth of each key frame, and a step of generating map points using the estimated depth.
[0051] The position estimation method may include a step of estimating the position and orientation of the in-vehicle camera based on a reprojection point of the map point and a detection position of the map point. The position estimation method may further include a step of correcting the estimated position and orientation based on a reprojection point of a landmark registered in a map and a position of the landmark in a key frame.
[0052] The present invention is not limited to the above-described embodiments, and can be appropriately modified without departing from the gist or idea of the invention read from the claims and the entire specification, and a position estimation method involving such a modification is also included in the technical scope of the present invention.
Description of Reference Numerals
[0053] 1… Vehicle, 11… On-vehicle camera, 100… Position estimation device, 110… Arithmetic device, 120… Storage device
Claims
1. A step of dividing an image captured by an in-vehicle camera into a plurality of frame images; A step of estimating a camera pose and a depth related to the in-vehicle camera for each key frame included in the plurality of frame images by depth-based Visual SLAM; A step of obtaining the estimated camera pose and the estimated depth of each key frame; A step of generating map points using the estimated depth; A position estimation method including the above.
2. Including a step of estimating the position and orientation of the in-vehicle camera based on a reprojection point of the map point and a detection position of the map point. The position estimation method according to Claim 1.
3. Including a step of correcting the estimated position and orientation based on a reprojection point of a landmark registered in a map and a position of the landmark in a key frame. The position estimation method according to Claim 2.
Citation Information
Patent Citations
Point Cloud Registration System for Autonomous Vehicles
JP2021514886A
Cited By
Depth measurement method and system based on multi-frame vision
CN121190538A