Estimation device, method, and program
The estimation device employs a quantum-inspired machine to optimize feature point selection, addressing the challenge of maintaining accuracy and reducing load in self-position estimation systems, thereby improving processing efficiency.
Patent Information
- Application Number
- JP2023220687
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-09
AI Technical Summary
Existing self-position estimation systems face challenges in maintaining processing accuracy while reducing the processing load, particularly when using multiple cameras, leading to difficulties in real-time processing.
An estimation device utilizing a quantum-inspired machine to solve a combination optimization problem for selecting an optimal combination of feature points from images, enabling parallel processing of self-position estimation and map creation.
This approach effectively reduces processing load while maintaining processing accuracy by optimizing feature point selection, enhancing the efficiency of self-position estimation and map creation.
Smart Images

Figure 2025103344000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an estimation device, method, and program.
Background Art
[0002] In the technology described in Patent Document 1, a self-position estimation device provided in a vehicle calculates the reliability of cameras using feature points extracted from each image acquired by a plurality of cameras arranged at different positions of the vehicle, and selects a camera to be used for self-position estimation based on the reliability.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the technology described in Patent Document 1, before executing the self-position estimation process, the reliability of a plurality of cameras is calculated using feature points extracted from the images acquired by each camera. For this reason, the processing load of reliability tends to increase according to the number of cameras. As a result, there has been a problem that real-time processing of self-position estimation becomes difficult. Therefore, a technology that can achieve both suppression of an increase in processing load and maintenance of processing accuracy has been demanded.
Means for Solving the Problems
[0005] According to one embodiment of the present disclosure, an estimation device (100) is provided. This estimation device is an estimation device that performs self-position estimation and map creation using an image obtained by a camera (200) mounted on a moving body (300) to photograph the environment around the moving body. This estimation device includes a SLAM processing unit (51) that executes the process of self-position estimation using the image and the process of map creation in parallel, and a combination determination unit (52) that causes a quantum-inspired machine (60) to execute a process of minimizing an objective function formulated as a combination optimization problem of selecting an optimal combination from among a plurality of feature points extracted from the image.
[0006] According to the estimation device of the above embodiment, by solving the combination optimization problem in which the quantum-inspired machine selects an optimal combination from among a plurality of feature points extracted from the image, an optimal combination can be selected from among a plurality of feature points extracted from the image. Therefore, it is possible to achieve both suppression of an increase in processing load and maintenance of processing accuracy.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0008] A. First Embodiment: A1. Outline of the Estimation Device: As shown in FIG. 1, the estimation device 100 is mounted on the vehicle 300. Based on the image captured by the camera 200 provided in the vehicle 300, the estimation device 100 uses Visual SLAM (Simultaneous Localization and Mapping) technology to execute the process of estimating its own position and the process of creating a map representing the surrounding environment.
[0009] The vehicle 300 includes a vehicle control device (not shown) for controlling each part of the vehicle 300 and an actuator group (not shown) including one or more actuators that are driven according to the control of the vehicle control device. The actuator group includes an actuator of a driving device for accelerating the vehicle 300, an actuator of a steering device for changing the traveling direction of the vehicle 300, and an actuator of a braking device for decelerating the vehicle 300. The driving device includes a battery, a traveling motor driven by the power of the battery, and driving wheels rotated by the traveling motor.
[0010] The camera 200 is provided, for example, at the upper part of the front glass of the vehicle 300. The camera 200 is a monocular camera that can capture a color image. The camera 200 captures a determined range outside the vehicle 300 at a determined frame rate. In the embodiment, the camera 200 captures a determined range in front of the vehicle 300. The camera 200 outputs the acquired image to the estimation device 100. Each image corresponds to a "frame" that constitutes a moving image.
[0011] The estimation device 100 is a computer including a memory 20, an input / output interface 30, a CPU (Central Processing Unit) 50 as a processor, and a quantum inspired machine 60. When the estimation device 100 is provided in the vehicle 300, the functions of this computer are realized by an ECU (Electronic Control Unit) responsible for the driving control of the vehicle 300.
[0012] The memory 20 and the input / output interface 30 are connected to the CPU 50 via the bus 90. The quantum-inspired machine 60 is connected to the CPU 50 via the bus 91. The input / output interface 30 is connected to the camera 200 via Ethernet (registered trademark).
[0013] The memory 20 stores various programs and various data used for various processes executed by the CPU 50. In this embodiment, the memory 20 stores the program PG1, the map point DB 21, and the key frame DB 22.
[0014] The map point DB 21 stores information on map points, which are feature points among the feature points included in the image for which the position in the world coordinate system has been obtained. The feature points refer to points of characteristic portions in the image such as edges and corners. Fig. 2 shows an explanatory diagram of the feature points extracted from the image. In Fig. 2, the feature points extracted from the image IM1 are surrounded by circles. Fig. 3 shows an explanatory diagram of the map points corresponding to the feature points. In Fig. 3, the map points are represented by diamonds. Using the map points, the change in the pose of the camera 200, that is, the movement trajectory of the vehicle 300, is estimated. The pose represents the position and orientation of the camera 200 represented using 6DoF (six degrees of freedom).
[0015] In the map point DB 21, as information on the map points, the position of the map points in the world coordinate system and the feature amount of the map points are stored in association with each other. The feature amount is calculated by an algorithm for feature point extraction using the pixel value of the feature point, the pixel values in the vicinity of the feature point, and the like.
[0016] The key frame DB 22 stores information on key frames, which are important images (frames) among the plurality of images acquired by the camera 200. In the key frame DB 22, as information on the key frames, the number for identifying the key frames, the information on the pose of the camera 200 when the key frames are captured, and the like are stored in association with each other.
[0017] By executing the program stored in the memory 20, the CPU 50 realizes various functions. For example, the CPU 50 stores the image output by the camera 200 in the memory 20. In the embodiment, the CPU 50 functions as the SLAM processing unit 51 and the optimization processing unit 52 by executing the program PG1 stored in the memory 20.
[0018] The quantum inspired machine 60 is a type of quantum annealing machine that executes an algorithm for simulating the behavior of quantum on a classical computer. The form for realizing the function of the quantum inspired machine 60 is arbitrary. For example, the function of the quantum inspired machine 60 is provided in an FPGA (Field Programmable Gate Array), a GPU (Graphics Processing Unit), or the estimation device 100, and is realized by another arithmetic device different from the CPU 50.
[0019] The SLAM processing unit 51 uses a plurality of images acquired by the camera 200 to perform self-position estimation and map creation. The function possessed by the SLAM processing unit 51 is also referred to as the "SLAM processing function". The optimization processing unit 52 causes the quantum inspired machine 60 to execute a process of minimizing an objective function formulated as a combinatorial optimization problem for the selection of feature points and the like. The optimization processing unit 52 is also called the "combinatorial determination unit". The function possessed by the optimization processing unit 52 is also referred to as the "optimization processing function".
[0020] A2. Processing flow: FIG. 4 shows a flowchart of the process executed by the SLAM processing unit 51. For example, when the vehicle 300 starts traveling, the process shown in FIG. 4 is started.
[0021] Reference 1: Raul Mur-Artal, et al., “ORB-SLAM: A Versatile and Accurate Monocular SLAM System.”, IEEE Transactions on Robotics, vol. 31, no. 5, pp. 1147-1163, October 2015., [online], [Searched on December 12, 2023], Internet URL: https: / / doi.org / 10.1109 / TRO.2015.2463671
[0022] In this embodiment, self-position estimation and map creation by ORB-SLAM described in Reference 1 are executed. When using ORB-SLAM, a map is created that includes information on the positions of a plurality of feature points extracted from an image (frame) in world coordinates, information on a plurality of key frames that observed those feature points, and the like. The map also includes information on the estimated pose of the camera at the time of acquisition of each key frame.
[0023] In step S10, the SLAM processing unit 51 executes an initialization process. For example, various data (map point DB 21, key frame DB 22, etc.) stored in the memory 20 are initialized.
[0024] In step S20, the SLAM processing unit 51 starts a tracking thread. In the tracking thread, the SLAM processing unit 51 executes a process (tracking) to estimate how much the pose of the camera 200 has changed between the image acquired immediately before and the image acquired this time by analyzing the image acquired by the camera 200. Details of the process in the tracking thread will be described later.
[0025] In step S30, the SLAM processing unit 51 activates a local mapping thread. In the local mapping thread, the SLAM processing unit 51 executes a process (local mapping) of creating a map using key frames. The generated map is a map representing the surrounding environment including objects around the vehicle 300. This map is also referred to as an environmental map. Details of the processing in the mapping thread will be described later.
[0026] In step S40, the SLAM processing unit 51 activates a loop closing thread. In the loop closing thread, the SLAM processing unit 51 continuously checks whether loop closing is possible, and if loop closing is possible, executes loop closing. Loop closing is a process of correcting the acquired key frames, the positions of map points related to the corresponding key frames in the world coordinate system, etc., from when the vehicle 300 was previously at that location until the present, using the deviation between the estimated pose when the vehicle 300 was at the same location and the current estimated pose of the vehicle 300 when it is recognized that the vehicle 300 has returned to the same location where it has been before.
[0027] In this way, the processing in the tracking thread, the mapping thread, and the loop closing thread is executed in parallel. The parallel execution of these processes by the SLAM processing unit 51 is also referred to as the "SLAM processing step". The estimated pose of the camera 200 is handled as the pose of the vehicle 300, for example, in the travel control of the vehicle 300.
[0028] In step S50, the SLAM processing unit 51 determines whether to end the process. The SLAM processing unit 51 determines not to end the process until it receives an end instruction from a higher-level application program, for example (step S50; NO), and waits. The higher-level application program is, for example, an application program that controls the autonomous driving of the vehicle 300. When the SLAM processing unit 51 determines to end the process (step S50; YES), it executes the process of step S60.
[0029] In step S60, the SLAM processing unit 51 stops each of the tracking thread, the local mapping thread, and the loop closing thread. Thereafter, the process shown in FIG. 4 is terminated.
[0030] FIG. 5 shows a flowchart showing a series of processes executed in the tracking thread. Each time a new frame is input, the process shown in FIG. 5 is started. A series of processes executed in the tracking thread are, except for some parts, based on the description in the above-mentioned reference 1.
[0031] In step S201, the SLAM processing unit 51 extracts feature points and feature amounts of each feature point from the image (frame) acquired by the camera 200. Here, the feature point ORB feature amount is extracted as the feature amount (ORB Extraction). For this purpose, first, a plurality of corners are detected by a detector using the FAST (Features from accelerated segment test) algorithm. The plurality of detected corners are treated as a plurality of feature points. Further, the ORB feature amount is calculated for each feature point.
[0032] In step S202, the SLAM processing unit 51 estimates the initial pose of the camera 200 (Initial Pose Estimation).
[0033] If not lost in the previous frame, the initial pose of the camera 200 is estimated assuming a constant velocity motion model from the previous frame to the current frame. Losing in the previous frame means that the self-position estimation using the previous frame has failed. The self-position estimation fails when, for example, the number of matches between the feature points observed in the frame acquired from the camera and the feature points recorded on the map is small.
[0034] If it has been lost in the previous frame, the initial pose of the camera 200 is estimated by Global Relocalization. In Global Relocalization, the positional relationship between the current frame and all the key frames that have been processed so far is checked, a search is performed for all the registered key frames, and a return is achieved.
[0035] In step S203, the SLAM processing unit 51 executes tracking of the map (local map) (Track Local Map). The feature points of the map (local map) are projected onto the current frame, and the correspondence relationship of the map points is searched for. The pose of the camera is finally optimized using all the map points discovered within the current frame.
[0036] In step S204, in order to remove some of the plurality of feature points extracted in step S201, combinatorial optimization processing for feature point selection is executed. This step is also referred to as the "optimization processing step". The processing of step S204 is not described in the above reference 1 and is a characteristic processing in this embodiment. By reducing the number of feature points to be processed, it is possible to reduce the amount of calculation of the processing and speed up the processing.
[0037] The optimization processing unit 52 causes the quantum-inspired machine 60 to execute combinatorial optimization processing in order to determine which of the plurality of feature points extracted from the current frame to leave. The feature points selected by the combinatorial optimization processing are left as processing targets. On the other hand, the feature points not selected by the combinatorial optimization processing are excluded from the targets of subsequent processing. Also, the feature quantities calculated from the feature points not selected, and the map points corresponding to the feature points not selected are excluded from the targets of subsequent processing. Details of the combinatorial optimization processing will be described later.
[0038] In step S205, the SLAM processing unit 51 determines whether the current frame becomes a new key frame (New KeyFrame Decision) according to whether the current frame meets a predetermined condition. The above is a series of processes in the tracking thread.
[0039] FIG. 6 shows a flowchart of a series of processes executed in the local mapping thread. When a new key frame is determined (see step S205 in FIG. 5), the process shown in FIG. 6 is started. A series of processes executed in the local mapping thread are based on the description in the above reference 1 except for some parts.
[0040] In step S301, a process for adding a new key frame is executed (KeyFrame Insertion). Specifically, the SLAM processing unit 51 reflects the relationship of common map points between the new key frame (hereinafter, key frame Ki) and other key frames in the first graph (Covisibility Graph). The first graph (Covisibility Graph) is a graph having key frames as nodes. In the first graph (Covisibility Graph), nodes representing two key frames having common map points are connected by edges. Also, in the first graph (Covisibility Graph), the number of common map points is represented as the weight of the edge.
[0041] In step S302, decimation of map points is executed (Recent Map Points Culling) so that the map (local map) does not include outliers of map points. Whether to leave the map points in the map (local map) is determined according to whether a predetermined condition is satisfied.
[0042] In step S303, in order to remove some of the plurality of map points, combinatorial optimization processing for the selection of map points is executed. This step is also referred to as the "optimization processing step". The processing of step S303 is not described in the above reference 1 and is a characteristic processing in this embodiment. By reducing the number of map points to be processed, it is possible to reduce the amount of calculation of the processing and speed up the processing.
[0043] The optimization processing unit 52 causes the quantum inspired machine 60 to execute combinatorial optimization processing in order to determine which of the plurality of map points to leave. The map points selected by the combinatorial optimization processing are left as processing targets. On the other hand, the map points not selected by the combinatorial optimization processing are excluded from the targets of subsequent processing. Details of the combinatorial optimization processing will be described later.
[0044] In step S304, the SLAM processing unit 51 introduces new map points (New Map Point Creation). New map points are generated by triangulation using the ORB feature amounts of other key frames (hereinafter referred to as key frame Kc) connected to the key frame Ki in the first graph (Covisibility Graph). For each non-matching ORB feature amount in the key frame Ki, it is searched whether it matches the map points in other key frames. The matching that does not satisfy the epipolar constraint is discarded.
[0045] In step S305, the SLAM processing unit 51 optimizes the key frame Ki that is the target of the current processing, all the key frames connected in the first graph (Covisibility Graph), and the map points detected by these key frames by local bundle adjustment (Local Bundle Adjustment).
[0046] In step S306, the SLAM processing unit 51 performs pruning of redundant keyframes (Local Keyframe Culling). Whether a keyframe is retained or not is determined according to whether it meets a predetermined condition.
[0047] In step S307, in order to remove some of the multiple poses estimated from multiple images (frames), combinatorial optimization processing for pose selection is performed. This step is also referred to as the "optimization processing step". The processing of step S307 is not described in the above Reference 1 and is a characteristic processing in this embodiment. By reducing the number of estimated poses, it is possible to reduce the computational amount of processing and speed up the processing.
[0048] The optimization processing unit 52 causes the quantum-inspired machine 60 to execute combinatorial optimization processing in order to determine which of the multiple estimated poses to retain. The poses selected by the combinatorial optimization processing are retained as processing targets. On the other hand, the poses not selected by the combinatorial optimization processing are excluded from the targets of subsequent processing. Details of the combinatorial optimization processing will be described later. The above is a series of processing in the local mapping thread.
[0049] A3. Combinatorial optimization processing for feature point selection: Subsequently, the combinatorial optimization processing for feature point selection (refer to step S204 in FIG. 5) will be described. In this combinatorial optimization processing, feature points are selected based on the following viewpoints. (11) Reduce the number of feature points as much as possible. (12) Reduce the feature points representing moving objects. (13) Retain the feature points with large feature amounts. (14) Retain the feature points detected at positions away from the central part of the image. (15) Select feature points so that there is no bias in the positions of the selected feature points within the image.
[0050] Note that conventionally, processing for limiting feature points has been performed based on the above viewpoints (11) to (14). The viewpoint of (15) above is characteristic of this embodiment.
[0051] An annealing machine including the quantum inspired machine 60 solves a combinatorial optimization problem as a quadratic unconstrained binary optimization (QUBO) problem. Therefore, it is necessary to represent the objective function in QUBO form. The Hamiltonian, which is the objective function, is expressed as in Equation (M1).
[0052]
Equation
[0053] The Hamiltonian represents energy, and the state in which the value of the Hamiltonian becomes small corresponds to the optimal solution when solving the combinatorial optimization problem with the annealing machine. Equation (M1) is also called a cost function. Also, solving the problem using the annealing machine corresponds to obtaining the most stable spin state (combination of spin directions) of the Ising model. σ i and σ j correspond to the spin directions. In Equation (M1), σ i and σ j are binary variables that take either "0" or "1". σ i = 1 represents the selected state, and σ i = 0 represents the unselected, i.e., removed state.
[0054] Here, as the first term (linear term) h i on the right side of Equation (M1), an equation for achieving the objective in the above viewpoints (11) to (14) is set. As the second term (quadratic term) J ij on the right side of Equation (M1), an equation for achieving the objective in the above viewpoint (15) is set.
[0055] In order to achieve the objective in the above viewpoint (11), it is desirable to minimize the value represented by Equation (M2).
[0056] [Number]
[0057] The subscript i is an integer that identifies the feature point. Here, it is assumed that the number of feature points detected in the target frame is n. N and n are integer values. However, N and n do not necessarily match. Equation (M2) represents the sum of σ of the same number as the number of feature points. As described above, σ i represents the sum of. As described above, σ i is a binary variable that takes either "0" or "1". In order to reduce the number of feature points as much as possible, the value represented by Equation (M2) is minimized. By using Equation (M2), it is possible to reduce the amount of computation and speed up the processing.
[0058] In order to achieve the objective from the viewpoint of (12) above, it is desirable to minimize the value represented by Equation (M3).
[0059] [Number]
[0060] The movement vector (motion vector) of feature point i from the previous frame is obtained, for example, by the optical flow method based on the position vector in the image coordinate system representing the position of feature point i in the previous frame and the position vector in the image coordinate system representing the position of feature point i in the current frame. As shown in Equation (M3), the angle difference of the movement is obtained by taking the inner product of the two vectors. This is because when estimating the movement vector using the epipolar constraint, the position vector in the image coordinate system representing the position of feature point i in the previous frame, the movement speed, and the rudder angle are used. Equation (M3) represents the sum of the angle differences of the movement. In order to reduce the feature points representing the moving object, the value represented by Equation (M3) is minimized. By using Equation (M3), the occurrence of errors in self-position estimation can be suppressed.
[0061] In order to achieve the object from the viewpoint of the above (13), it is desirable to minimize the value represented by the formula (M4).
[0062] [Number]
[0063] feature i represents the magnitude of the feature amount of the feature point i. In the present embodiment, the feature amount of the feature point i is the ORB feature amount. σ i is a binary variable that takes "0" or "1". The formula (M4) is σ i = the sum of feature i when "0". That is, it represents the sum of the feature amounts that were not selected. In order to leave the feature points with large feature amounts, the value represented by the formula (M4) is minimized. By using the formula (M4), it is possible to improve the accuracy of self-position estimation.
[0064] In order to achieve the object from the viewpoint of the above (14), it is desirable to minimize the value represented by the formula (M5).
[0065] [Number]
[0066] The formula (M5) represents the sum of the magnitudes of the distances between the pixels in the central part of the image and the feature point i when σ i = "0". That is, it represents the sum of the magnitudes of the distances between the feature points i that were not selected and the pixels in the central part. In order to leave the feature points detected at positions far from the central part of the image, the value represented by the formula (M5) is minimized. By using the formula (M5), it is possible to improve the accuracy of self-position estimation.
[0067] To achieve the object from the perspective of the above (15), it is desirable to minimize the value represented by Expression (M6). Expression (M6) represents the sum of the distances between feature point i and another feature point j different from feature point i. By using Expression (M6), the accuracy of self-position estimation can be improved. Since Expression (M6) acts in the direction of leaving map points, it is given a negative sign. In order to select feature points so that there is no bias in the positions of the selected feature points within the image, the value representing Expression (M6) is minimized.
[0068]
Number
[0069] Expression (M7) is an objective function in QUBO form in the combinatorial optimization process for the selection of feature points. In Expression (M7), the expressions representing the above Expressions (M2) to (M6) are respectively multiplied by weighting coefficients A1 to E1, and the sum is taken. This objective function formulates a combinatorial optimization problem of selecting an optimal combination from among a plurality of feature points.
[0070]
Number
[0071] Also, the first term (linear term) on the right side of Expression (M7) is added to the Hamiltonian, which is the objective function, only when σ i = 1. The second term (quadratic term) on the right side of Expression (M7) is added to the Hamiltonian, which is the objective function, only when σ i = 1 and σ j = 1. Although the viewpoints of (11) to (14) have been conventionally considered, in the present embodiment, the purpose is to select feature points that can further improve the accuracy of self-position estimation and map creation by considering the viewpoint of (15).
[0072] Conventionally, feature points detected at positions far from the center of the image were preferentially selected, but among the selected ones, there were sometimes feature points with small feature amounts. In this case, the accuracy of self-position estimation decreases.
[0073] On the other hand, according to the above method, as represented by formula (M5), feature points detected at positions far from the center of the image are left, and as represented by formula (M4), feature points with large feature amounts are left. Also, as represented by formula (M6), while selecting feature points so that there is no bias in the positions of each feature point to be processed within the image, feature points with large feature amounts are left. Feature points with large feature amounts can be selected well-balancedly from within the image. In this way, a decrease in the accuracy of self-position estimation can be suppressed. Also, it is known that an annealing machine including the quantum-inspired machine 60 can quickly calculate the solution to the optimization problem represented by the annealing model. Therefore, it is possible to achieve both maintaining the processing accuracy and suppressing an increase in the processing load.
[0074] A4. Combinatorial optimization processing for map point selection: Next, the combinatorial optimization processing for map point selection (refer to step S303 in FIG. 6) will be described. In this combinatorial optimization processing, map points are selected based on the following viewpoints. (21) Reduce the number of map points as much as possible. (22) Leave map points with a large leverage ratio. (23) Leave map points referred to by multiple poses. (24) Select map points so that there is no bias in the positions of the selected map points within the map space.
[0075] Note that conventionally, processing for limiting map points has also been performed based on the above viewpoints (21) to (23). The viewpoint of (24) is characteristic of this embodiment. The viewpoint of (24) means selecting map points so that the map points are sparsely distributed in the map space.
[0076] Here, h in the first term (linear term) on the right side of the above formula (M1) i is used to set a formula for achieving the purpose from the viewpoints of the above (21) to (23). J in the second term on the right side of formula (M1) ij is used to set a formula for achieving the purpose from the viewpoint of the above (24).
[0077] To achieve the purpose from the viewpoint of the above (21), it is desirable to minimize the value represented by formula (M8).
[0078]
Number
[0079] The subscript i is an integer that identifies a map point. Assume that the number of detected map points is n. N and n are integer values. However, N and n do not necessarily match. Formula (M8) represents the sum of the same number of σ i as the number of map points. σ i is a binary variable that takes either "0" or "1". To reduce the number of map points as much as possible, the value represented by formula (M8) is minimized. By using formula (M8), the reduction of the amount of calculation and the acceleration of processing can be achieved.
[0080] To achieve the purpose from the viewpoint of the above (22), it is desirable to minimize the value represented by formula (M9).
[0081]
Number
[0082] Leverage is generally said to be an index indicating the influence that each observed value has on the estimated value. In the present embodiment, leverage is an index indicating the result of self-position estimation for a map point. Therefore, as f(i) in Expression (M9), a mathematical formula representing leverage can be adopted. Leverage is a diagonal component of the hat matrix (projection matrix). As a data matrix for obtaining the hat matrix, a position vector in the map space coordinate system of the map point is used. Since Expression (M9) acts in the direction of leaving map points, it is given a negative sign. In order to leave map points with large leverage, the value represented by Expression (M9) is minimized. By using Expression (M9), the occurrence of an error in self-position estimation can be suppressed.
[0083] In order to achieve the objective from the above viewpoint (23), it is desirable to minimize the value represented by Expression (M10).
[0084]
Number
[0085] covis(i) represents the number of one or more poses that refer to the target map point i. Since Expression (M10) acts in the direction of leaving map points, it is given a negative sign. In order to leave map points referred to by a plurality of poses, the value represented by Expression (M10) is minimized. By using Expression (M10), the occurrence of an error in self-position estimation can be suppressed.
[0086] In order to achieve the objective from the above viewpoint (24), it is desirable to minimize the value represented by Expression (M11). Expression (M11) represents the sum of the magnitudes of the distances between map point i and other map points j in the map space coordinate system. Since Expression (M11) acts in the direction of leaving map points, it is given a negative sign. Expression (M11) is minimized so that the selected map points are sparsely distributed in the map space. By using Expression (M11), the accuracy of self-position estimation can be improved.
[0087]
Mathematics
[0088] Equation (M12) is the objective function in QUBO form for the combinatorial optimization process of map point selection combinations. In Equation (M12), the above Equations (M8) to (M11) are respectively multiplied by weighting coefficients A2 to D2, and the sum is taken. This objective function formulates a combinatorial optimization problem for selecting the optimal combination of map points.
[0089]
Mathematics
[0090] Also, the first term (linear term) on the right side of Equation (M12) is added to the Hamiltonian, which is the objective function, only when σ i = 1. The second term (quadratic term) on the right side of Equation (M12) is added to the Hamiltonian, which is the objective function, only when σ i = 1 and σ j = 1. Although the viewpoints in (21) to (23) have been conventionally considered, in this embodiment, by considering the viewpoint in (24), the purpose is to select map points that can further improve the accuracy of self-position estimation and map creation.
[0091] Conventionally, map points have been selected so that the map points selected in the map space are sparsely located. Among the selected map points, there are those that have no common relationship with other map points or have almost no common relationship with other map points. In this case, the accuracy of self-position estimation will decrease.
[0092] On the one hand, according to the above method, a map point having a relationship common with other map points is selected as represented by formulas (M9) to (M10). A map point having a relationship common with other map points means, for example, a map point estimated from feature points detected in a frame that can be regarded as the same as a map point estimated from feature points detected in another frame. Also, the map points are selected so that the selected map points are sparsely arranged as represented by formula (M11). In this way, a decrease in the accuracy of self-position estimation can be suppressed. Also, it is known that an imaging machine including the quantum-inspired machine 60 can quickly calculate the solution of an optimization problem represented by an Ising model. Therefore, it is possible to achieve both maintenance of processing accuracy and suppression of an increase in processing load.
[0093] A5. Combinatorial optimization processing for pose selection: Subsequently, combinatorial optimization processing for pose selection (refer to step S307 in FIG. 6) will be described. In this combinatorial optimization processing, poses are selected based on the following viewpoints. (31) Reduce the number of poses as much as possible. (32) Reduce poses with a small number of visible map points. (33) Leave pairs of poses with large motion parallax among two or more poses that can see one or more common map points.
[0094] Note that conventionally, processing for limiting poses has been performed based on the above viewpoints (31) to (32). The viewpoint of (33) is characteristic of this embodiment. The viewpoint of (33) means leaving pairs of poses that can see one or more common map points and having a large angle formed by the respective line-of-sight directions with respect to the one or more common map points. In other words, the viewpoint of (33) means selecting pairs of poses such that the motion parallax of the selected pairs of poses (two poses that can see one or more common map points) is larger than the motion parallax of the unselected pairs of poses. Here, for h in the first term (linear term) on the right side of the above formula (M1) i as, a formula for achieving the objective from the viewpoints of the above (31) to (32) is set. For J in the second term on the right side of formula (M1) ij as, a formula for achieving the objective from the viewpoint of the above (33) is set.
[0095] To achieve the objective from the viewpoint of the above (31), it is desirable to minimize the value represented by formula (M13).
[0096]
Equation
[0097] The subscript i is an integer that identifies a pose. Assume that the number of estimated poses is n. N and n are integer values. However, N and n do not necessarily match. Formula (M13) represents the sum of the same number of σ as the number of poses. As described above, σ i is a binary variable that takes "0" or "1". To reduce the number of poses as much as possible, the value represented by formula (M13) is minimized. By using formula (M13), the reduction of the amount of calculation and the acceleration of processing can be achieved. i is a binary variable that takes "0" or "1". To reduce the number of poses as much as possible, the value represented by formula (M13) is minimized. By using formula (M13), the reduction of the amount of calculation and the acceleration of processing can be achieved.
[0098] To achieve the objective from the viewpoint of the above (32), it is desirable to minimize the value represented by formula (M14).
[0099]
Equation
[0100] count(i) represents the number of map points that pose i can see. Formula (M14) represents the sum of the reciprocals of the number of map points that pose i can see. To reduce the poses with a small number of map points that can be seen, the value represented by formula (M14) is minimized. By using formula (M14), the occurrence of errors in self-position estimation can be suppressed.
[0101] In order to achieve the object from the viewpoint of the above (33), it is desirable to minimize the value represented by the formula (M15).
[0102]
Number
[0103] The formula (M15) represents the direction vector in the map space coordinate system of the map point m i seen from the pose x. k Since the formula (M15) is an expression that acts in the direction of leaving the pose, it is given a negative sign.
[0104] The formula (M16) is an objective function in the QUBO format in the combinatorial optimization process for pose selection. In the formula (M16), the weight coefficients A3 to C3 are multiplied by the above formulas (M13) to (M15), respectively, and the sum is taken. This objective function formulates a combinatorial optimization problem of selecting an optimal combination from among a plurality of estimated poses.
[0105]
Number
[0106] Also, the first term (linear term) on the right side of the formula (M16) is added to the Hamiltonian, which is the objective function, only when σ i = 1. The second term (quadratic term) on the right side of the formula (M16) is added to the Hamiltonian, which is the objective function, only when σ i = 1 and σ j = 1. The viewpoints of (31) to (32) have been conventionally considered, but in the present embodiment, the purpose is to select a pose that can further improve the accuracy of self-position estimation and map creation by considering the viewpoint of (33).
[0107] FIG. 7 shows an explanatory diagram for comparing the effects of the present embodiment with the conventional case. In FIG. 7, the estimated poses are indicated by circles, and the map points are indicated by diamonds. Also, the circles representing the poses with a large number of viewed map points are hatched with cross-hatching, and the circles representing the poses with a small number of viewed map points are hatched with single hatching. A large number of viewed map points means that the number of viewed map points is equal to or greater than a preset threshold. A small number of viewed map points means that the number of viewed map points is less than a preset threshold.
[0108] Conventionally, poses with a small number of viewable map points were removed. As shown in the upper right of FIG. 7, conventionally, pose P1 and pose P2 were excluded. On the other hand, in the present embodiment, as shown in the lower right of FIG. 7, pose P1 and pose P2 are retained despite having a small number of viewed map points.
[0109] As shown in the upper right of FIG. 7, the motion parallax θa during the movement from pose P1 to pose P12 following pose P11, which is subsequent to the excluded pose P1, is larger than the motion parallax θc during the movement from pose P11 to pose P12. Motion parallax refers to the parallax generated by the movement of the observer's viewpoint or the monitored object. In SLAM for estimating the self-position based on motion parallax, the larger the motion parallax, the more likely the estimation accuracy of the self-position is to improve.
[0110] However, conventionally, even when the motion parallax between two poses that can view the same map point is large, they may be removed. If poses with a small number of viewable map points are uniformly removed as in the conventional case, the accuracy of self-position estimation will decrease.
[0111] On the one hand, according to the method of this embodiment, as represented by the formula (M14), poses with a small number of visible map points are reduced. On the other hand, as represented by the formula (M15), a pair of poses (two poses that can see one or more common map points) is selected such that the motion parallax of the selected pair of poses is larger than the motion parallax of the pair of poses that was not selected. Therefore, a pair of poses that can see one or more common map points and in which the angle formed by the respective line-of-sight directions with respect to the one or more common map points is large remains. In this way, a decrease in the accuracy of self-position estimation can be suppressed. Also, it is known that an annealing machine including the quantum-inspired machine 60 can quickly calculate the solution to the optimization problem represented by the annealing model. Therefore, it is possible to achieve both maintaining the processing accuracy and suppressing an increase in the processing load.
[0112] B. Other Embodiments: B1. Other Embodiment 1: In the above embodiment, an example in which the combinatorial optimization process for selecting each of the feature points, the map points, and the poses is executed has been described. However, not all of these optimization processes need to be executed. For example, only the combinatorial optimization process for selecting the feature points may be executed. Or only the combinatorial optimization process for selecting the map points may be executed. Or only the combinatorial optimization process for selecting the poses may be executed. Or two of the three optimization processes may be executed.
[0113] B2. Other Embodiment 2: Also, the combinatorial optimization process for selecting feature points (see step S204 in FIG. 5) does not necessarily have to be executed every time a frame is input. The combinatorial optimization process may be executed at a specific timing. The number of input frames may be counted, and the combinatorial optimization process may be executed every determined number of frames. The process shown in FIG. 5 starts every time a new frame is input. However, for example, the process of step S204 may be executed every 20 frames. In other cases, when FIG. 5 is executed, the process of step S204 may be skipped.
[0114] B3. Other Embodiment 3: Also, the combinatorial optimization process for selecting map points (see step S303 in FIG. 6) does not necessarily have to be executed every time a new key frame is determined. The combinatorial optimization process may be executed at a specific timing. The number of newly determined key frames may be counted, and the combinatorial optimization process may be executed every determined number of key frames. The process shown in FIG. 6 starts when a new key frame is determined. However, for example, the process of step S303 may be executed every 10 key frames. In other cases, when FIG. 6 is executed, the process of step S303 may be skipped.
[0115] Also, regarding the combinatorial optimization process for selecting poses, similarly, it does not necessarily have to be executed every time a new key frame is determined. The combinatorial optimization process may be executed at a specific timing.
[0116] B4. Other Embodiment 4: The combinatorial optimization process for selecting feature points (see step S204 in FIG. 5) may be executed, for example, before the tracking of the local map (see step S203 in FIG. 5). The combinatorial optimization process for selecting feature points (see step S204 in FIG. 5) may be executed, for example, after the extraction of feature points and feature amounts (see step S201 in FIG. 5).
[0117] B5. Other Embodiment 5: Also, the combinatorial optimization process for map point selection (refer to step S303 in FIG. 6) may be executed, for example, before the decimation of map points (refer to step S302 in FIG. 6), or may be executed after the introduction of new map points (refer to step S304 in FIG. 6).
[0118] Also, the combinatorial optimization process for pose selection (refer to step S307 in FIG. 6) may be executed, for example, before the decimation of key frames (refer to step S306 in FIG. 6).
[0119] B6. Other Embodiment 6: In the above embodiment, an example of setting the objective function shown in formula (M7) for the selection of feature points has been described. However, the objective function may be expressed by a formula other than formula (M7). For example, as h in the first term (linear term) on the right side of formula (M1), an objective function in which a part of formulas (M2) to (M5) is adopted may be set. i
[0120] In the above embodiment, an example of setting the objective function shown in formula (M12) for the selection of map points has been described. However, the objective function may be expressed by a formula other than formula (M12). For example, as h in the first term (linear term) on the right side of formula (M1), an objective function in which a part of formulas (M8) to (M10) is adopted may be set. i
[0121] In the above embodiment, an example of setting the objective function shown in formula (M16) for the selection of poses has been described. However, the objective function may be expressed by a formula other than formula (M16). For example, as h in the first term (linear term) on the right side of formula (M1), an objective function in which a part of formulas (M13) to (M14) is adopted may be set. i
[0122] B7. Other Embodiment 7: The estimation device 100 may further include an output unit that outputs the estimated self-position and the generated map to, for example, a vehicle control device. The vehicle control device can use the estimated self-position and the generated map for, for example, the control of the automatic driving of the vehicle 300.
[0123] B8. Other Embodiment 8: In the above embodiment, an example of the process of self-position estimation and map creation using ORB-SLAM described in Reference 1 was explained. However, the technique of selecting feature points and the like using the combinatorial optimization process in the present disclosure is applicable to other SLAM algorithms based on feature points (feature quantity-based), which are non-direct methods. As another SLAM algorithm, for example, there is PTAM (Parallel Tracking and Mapping).
[0124] B9. Other Embodiment 9: In the embodiment, an example of the vehicle 300 is given as the moving body, but the moving body is not limited to a vehicle. The moving body means an object that can move. Examples of the moving body include an AGV (Automatic Guided Vehicle) and an AMR (Autonomous Mobile Robot).
[0125] The estimation device and its method described in the present disclosure may be implemented by a dedicated computer configured by a processor and a memory programmed to execute one or more functions embodied by a computer program. Alternatively, the estimation device and its method described in the present disclosure may be implemented by a dedicated computer provided by configuring a processor with one or more dedicated hardware logic circuits. Or, the estimation device and its method described in the present disclosure may be implemented by one or more dedicated computers configured by a combination of a processor and a memory programmed to execute one or more functions and a processor configured by one or more hardware logic circuits. Also, the computer program may be stored in a computer-readable non-transitory tangible recording medium as instructions executable by a computer.
[0126] The present disclosure is not limited to the above-described embodiments, and can be realized in various configurations without departing from the gist thereof. For example, the technical features in the embodiments corresponding to the technical features in each form described in the summary of the invention can be appropriately replaced or combined in order to solve some or all of the above-described problems or to achieve some or all of the above-described effects. Also, if the technical feature is not described as essential in this specification, it can be appropriately deleted.
Description of Reference Numerals
[0127] 20... Memory, 21... Map Point DB, 22... Key Frame DB, 30... Input / Output Interface, 50... CPU, 51... SLAM Processing Unit, 52... Optimization Processing Unit, 60... Quantum Inspired Machine, 90, 91... Bus, 100... Estimation Device, 200... Camera, 300... Vehicle, PG1... Program
Claims
1. An estimation device (100) that performs self-position estimation and map creation using an image captured by a camera (200) mounted on a moving body (300) of the surrounding environment of the moving body, a SLAM processing unit (51) that executes in parallel the process of self-position estimation using the image and the process of map creation, a combination determination unit (52) that causes a quantum inspired machine (60) to execute a process of minimizing an objective function formulated as a combination optimization problem of selecting an optimal combination from among a plurality of feature points extracted from the image, The estimation device comprising.
2. The estimation device according to claim 1, wherein the objective function is represented in QUBO (Quadratic Unconstrained Binary Optimization) form, and the quadratic term of the equation formulated in the QUBO form is set so as to satisfy the condition of not biasing the positions in the image of the plurality of feature points included in the combination in the combination of the selected feature points, Estimation device.
3. The estimation device according to claim 2, wherein the linear term of the formulated equation is set so as to satisfy the condition of reducing the number of the plurality of feature points included in the selected combination, Estimation device.
4. The estimation device according to claim 3, wherein the linear term of the formulated equation is further set so as to satisfy the condition that the feature amounts of the plurality of feature points included in the selected combination are larger than the feature amounts of the feature points not selected, Estimation device.
5. The estimation device according to claim 1, wherein the combination determination unit, in order to determine a combination of map points representing the coordinates of the feature points in the world coordinate system, causes the quantum inspired machine to execute a process of minimizing the objective function formulated as the combination optimization problem of selecting an optimal combination from among the plurality of map points, Estimation device.
6. The estimation device according to claim 5, wherein the objective function is represented in QUBO (Quadratic Unconstrained Binary Optimization) form, The quadratic term of the equation formulated in the QUBO form is set so as to satisfy the condition of not biasing the positions in the map space of a plurality of the map points included in the combination of the selected map points. Estimation device. **Claim 7** The estimation device according to claim 6, The linear term of the formulated equation is set so as to satisfy the condition of reducing the number of a plurality of the map points included in the selected combination. Estimation device. **Claim 8** The estimation device according to claim 1, The combination determination unit causes the quantum-inspired machine to execute a process of minimizing an objective function formulated as the combination optimization problem of selecting an optimal combination from among a plurality of estimated poses of the camera representing changes in the pose of the camera accompanying the movement of the mobile body. Estimation device. **Claim 9** The estimation device according to claim 8, The objective function is represented in the QUBO (Quadratic Unconstrained Binary Optimization) form, The quadratic term of the equation formulated in the QUBO form is set so as to satisfy the condition that the motion parallax is large in two or more poses that can view one or more common map points among a plurality of the selected estimated poses. Estimation device. **Claim 10** The estimation device according to claim 9, The linear term of the formulated equation is set so that a plurality of the estimated poses included in the selected combination satisfy the condition of being able to view more map points than the estimated poses not selected. Estimation device. **Claim 11** The estimation device according to any one of claims 1 to 10, further comprising an output unit that outputs the estimated self-position and the generated map. Estimation device. **Claim 12** A method for performing self-position estimation and map creation, an SLAM processing step of executing in parallel a self-position estimation process and a map creation process using an image obtained by a camera mounted on a mobile body for photographing the environment around the mobile body; an optimization processing step of selecting an optimal combination by causing a quantum-inspired machine to execute a process of minimizing an objective function formulated as a combination optimization problem of selecting an optimal combination from among a plurality of feature points extracted from the image.
13. A program for causing a computer that performs self-position estimation and map creation to execute: A SLAM processing function that executes in parallel a self-position estimation process using an image obtained by a camera mounted on a moving body to photograph the environment around the moving body and a map creation process; An optimization processing function that selects an optimal combination by causing a quantum inspired machine to execute a process of minimizing an objective function formulated as a combination optimization problem of selecting an optimal combination from among a plurality of feature points extracted from the image; A program for causing the computer to implement the above.
Citation Information
Patent Citations
Position estimating device
JP2021043486A