Information processing device, self-localization method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2026-04-01
Smart Images

Figure 0007838646000001 
Figure 0007838646000002 
Figure 0007838646000003
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information processing device, a self-localization method, and a non-temporary computer-readable medium. [Background technology]
[0002] In recent years, services that rely on robots moving autonomously have become widespread. For a robot to move autonomously, it is necessary for the robot to perceive its surrounding environment and estimate its own position with high accuracy. Therefore, VSLAM (Visual Simultaneous Localization and Mapping), which simultaneously creates an environmental map of the surroundings from video footage captured by the robot and estimates its own position by referring to the created environmental map, is being considered. In typical VSLAM, the same point captured in multiple videos is recognized as a feature point in the multiple images (still images) that make up those videos, and the position of the camera that captured the video is estimated from the difference between the images of that feature point. Since the position of the camera on the robot is fixed, if the camera position can be estimated, the robot's position can be estimated. Camera position estimation using VSLAM is done by estimating the 3D position of feature points contained in multiple images, and then using the difference between the 2D position obtained by projecting the estimated 3D position onto the image and the position of the camera that captured the video. Since such VSLAM requires immediate processing, it is necessary to reduce the processing load.
[0003] Patent Document 1 describes the configuration of an autonomous mobile device that estimates its own position by obtaining the correspondence between feature points contained in the image information stored in the memory unit and feature points extracted from the captured image. Furthermore, Patent Document 1 describes the process of thinning out the images stored in the memory unit according to the number of corresponding feature points obtained during estimation.
[0004] Patent Document 2 describes the configuration of an information processing device that extracts feature points from an input image and detects the position and orientation of an imaging device that captured the input image based on the extracted feature points. The information processing device in Patent Document 2 changes the number of feature points extracted from the input image based on the processing time required to detect the position and orientation of the imaging device from the input image. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2020-57187 [Patent Document 2] Japanese Patent Publication No. 2021-9557 [Overview of the project] [Problems that the invention aims to solve]
[0006] However, the autonomous mobile device disclosed in Patent Document 1 increases the number of images to be downsampled and decreases the number of images to be stored as the number of corresponding feature points acquired during estimation exceeds a threshold. In this case, there is a problem that the accuracy of self-localization deteriorates as the number of images used for self-localization decreases. Furthermore, in the information processing device disclosed in Patent Document 2, as the processing load of processes other than the process for detecting the position and orientation of the imaging device increases, the processing time required to detect the position and orientation of the imaging device also increases. In this case, the number of feature points extracted from the input image by the information processing device decreases, which is a problem that deteriorates the accuracy of self-localization.
[0007] One of the purposes of this disclosure is to provide an information processing device, a self-localization method, and a non-temporary computer-readable medium that can prevent deterioration of the accuracy of self-localization when reducing the processing load, in light of the above-mentioned problems. [Means for solving the problem]
[0008] An information processing device according to a first aspect of this disclosure includes: a detection unit that detects a plurality of new feature points from a first image; a identification unit that identifies corresponding feature points from the plurality of new feature points that correspond to known feature points associated with a three-dimensional position included in at least one management image used to generate an environmental map; and an estimation unit that estimates the position and orientation of an imaging device that took the first image using the corresponding feature points, wherein the detection unit changes the number of new feature points it detects from a target image for which the position and orientation of the imaging device are to be estimated, according to the number of corresponding feature points.
[0009] A self-localization method according to a second aspect of the present disclosure detects a plurality of new feature points from a first image, identifies corresponding feature points from among the plurality of new feature points that correspond to known feature points associated with a three-dimensional position included in at least one management image used to generate an environmental map, estimates the position and orientation of the imaging device that took the first image using the corresponding feature points, and changes the number of new feature points to be detected from the target image for which the position and orientation of the imaging device are to be estimated according to the number of corresponding feature points.
[0010] A program according to a third aspect of this disclosure causes a computer to detect a plurality of new feature points from a first image, identify corresponding feature points among the plurality of new feature points that correspond to known feature points associated with a three-dimensional position included in at least one control image used to generate an environmental map, estimate the position and orientation of the camera that took the first image using the corresponding feature points, and change the number of new feature points to be detected from the target image for which the position and orientation of the camera is to be estimated according to the number of corresponding feature points. [Effects of the Invention]
[0011] This disclosure provides an information processing device, a self-localization method, and a non-temporary computer-readable medium that can prevent deterioration of the accuracy of self-localization when reducing the processing load. [Brief explanation of the drawing]
[0012] [Figure 1] It is a configuration diagram of the information processing apparatus according to Embodiment 1. [Figure 2] It is a diagram showing the flow of the self-position estimation process according to Embodiment 1. [Figure 3] It is a configuration diagram of the information processing apparatus according to Embodiment 2. [Figure 4] It is a diagram for explaining the feature point matching process according to Embodiment 2. [Figure 5] It is a diagram for explaining the feature point classification process according to Embodiment 2. [Figure 6] It is a diagram showing the flow of the update process of the target number of feature points according to Embodiment 2. [Figure 7] They are configuration diagrams of the information processing apparatus according to each embodiment.
Modes for Carrying Out the Invention
[0013] (Embodiment 1) Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. A configuration example of the information processing apparatus 10 according to Embodiment 1 will be described using FIG. 1. The information processing apparatus 10 may be a computer device that operates by a processor executing a program stored in a memory. The information processing apparatus 10 may be, for example, a server device.
[0014] The information processing device 10 includes a detection unit 11, a specification unit 12, and an estimation unit 13. The detection unit 11, the specification unit 12, and the estimation unit 13 may be software or modules whose processing is performed by a processor executing a program stored in memory. Alternatively, the detection unit 11, the specification unit 12, and the estimation unit 13 may be hardware such as a circuit or chip. In Figure 1, the detection unit 11, the specification unit 12, and the estimation unit 13 are shown as being included in one information processing device 10, but the detection unit 11, the specification unit 12, and the estimation unit 13 may each be located on different computer devices. Alternatively, one of the components of the detection unit 11, the specification unit 12, and the estimation unit 13 may be located on different computer devices. Computers including the detection unit 11, the specification unit 12, and the estimation unit 13 may communicate with each other via a network.
[0015] The detection unit 11 detects a plurality of new feature points from the first image. The first image may be an image taken by an imaging device mounted on a moving object such as a vehicle. The imaging device mounted on the moving object may generate an image by taking pictures in the direction of travel of the moving object or around the moving object while the moving object is moving. The detection unit 11 may receive the image taken by the imaging device via a network. Alternatively, if the imaging device is used as an integral part of the information processing device 10, that is, if the imaging device is included in the information processing device 10 or the imaging device is connected to the information processing device 10, the detection unit 11 may acquire the image without going through a network. Alternatively, the first image may be an image received from another information processing device, etc., via a network.
[0016] The shooting device may be, for example, a camera, or a device having camera functionality. A device having camera functionality may be, for example, a mobile terminal such as a smartphone. The image may be, for example, a still image. Alternatively, the image may be a frame image that makes up a video. Furthermore, the multiple images may be a dataset or data record representing multiple still images, such as multiple frame images that make up a video. Alternatively, the multiple images may be frame images extracted from multiple frame images that make up a video.
[0017] The moving object may be, for example, an autonomously moving robot or vehicle. Autonomous movement means that the robot or vehicle operates by a control device mounted on it, without direct human control of the vehicle.
[0018] Novel feature points may be detected using, for example, SIFT, SURF, ORB, AKAZE, etc. Novel feature points may also be represented using two-dimensional coordinates, which are camera coordinates defined in the imaging device.
[0019] The identification unit 12 identifies known feature points included in at least one management image used to generate the environment map, which are associated with a 3D position, and novel feature points corresponding to these known feature points as corresponding feature points.
[0020] An environmental map is a three-dimensional map that shows the environment around the imaging device using three-dimensional information. Three-dimensional information may also be referred to as 3D information, three-dimensional coordinates, etc. The environmental map includes map information showing the environment around the imaging device, as well as information about the position and orientation of the imaging device. The orientation of the imaging device may, for example, be information about the tilt of the imaging device. The environmental map is generated by identifying the shooting locations where multiple images were taken and reconstructing the three-dimensional positions of feature points recorded on the images. In other words, the environmental map includes information about the three-dimensional position or three-dimensional coordinates of feature points in images taken using the imaging device. For example, an environmental map may be generated by performing SfM (Structure from Motion) using multiple images. SfM calculates all feature points from a series of already acquired two-dimensional images (or frames) and estimates matching feature points from multiple images that are in a temporal order. Furthermore, SfM accurately estimates the three-dimensional position or orientation of the camera that took each frame based on the difference in position on the two-dimensional plane in the frame in which each feature point appears. The control image is the image used when performing SfM. Alternatively, the environment map may be created by accumulating images previously estimated using VSLAM. In this case, the control image will be the image input to VSLAM with its 3D position estimated.
[0021] Known feature points are feature points included in the management image and represented using two-dimensional coordinates. The three-dimensional position associated with a known feature point may also be represented using three-dimensional coordinates, for example. Corresponding feature points may be feature points that have the same or similar features as a known feature point. Corresponding feature points may also be described as feature points that match a known feature point. In other words, the identification unit 12 may be described as identifying or extracting new feature points that match a known feature point from among a plurality of new feature points.
[0022] The estimation unit 13 estimates the position and orientation of the imaging device that captured the first image using corresponding feature points. The estimation unit 13 may estimate the position and orientation of the imaging device that captured the first image by, for example, performing VSLAM. Estimating the position and orientation of the imaging device that captured the first image may also mean estimating the position and orientation of the mobile body on which the imaging device is mounted.
[0023] Here, the detection unit 11 estimates the position and orientation of the imaging device that took each image, even for images taken at a time later than when the first image was taken. Images taken at a time later than when the first image was taken are referred to as target images for which the position and orientation of the imaging device are estimated.
[0024] The detection unit 11 changes the number of new feature points to detect from the target image according to the number of corresponding feature points used to estimate the position and orientation of the imaging device that took the first image. For example, if the number of corresponding feature points used to estimate the position and orientation of the imaging device that took the first image (hereinafter simply referred to as "number of corresponding feature points") is greater than a predetermined number, the number of new feature points to detect from the target image may be reduced from the currently set number. If the number of corresponding feature points is less than a predetermined number, the number of new feature points to detect from the target image may be increased from the currently set number.
[0025] A predetermined number of corresponding feature points can be considered a sufficient number of corresponding feature points to estimate the position and orientation of the imaging device that captured the target image. A sufficient number of corresponding feature points to estimate the position and orientation of the imaging device that captured the target image means a number of corresponding feature points that allows for highly accurate estimation of the position and orientation of the imaging device that captured the target image. Therefore, if the number of corresponding feature points is greater than a predetermined number, the identification unit 12 can identify a sufficient number of corresponding feature points to estimate the position and orientation of the imaging device that captured the target image, even if the number of new feature points detected from the target image is reduced. Furthermore, by reducing the number of new feature points detected from the target image, the processing load related to the identification of corresponding feature points and the processing load related to the estimation of the position and orientation of the imaging device using the corresponding feature points can be reduced.
[0026] On the other hand, if the number of corresponding feature points is less than a predetermined number, it can be said that a sufficient number of corresponding feature points have not been identified to estimate the position and orientation of the imaging device that captured the target image. Therefore, if the number of corresponding feature points is less than a predetermined number, the identification unit 12 increases the number of corresponding feature points used to estimate the position and orientation of the imaging device that captured the target image by increasing the number of new feature points detected from the target image. As a result, the estimation unit 13 can improve the accuracy of estimating the position and orientation of the imaging device with respect to the target image.
[0027] Alternatively, if the number of corresponding feature points used to estimate the position and orientation of the imaging device for multiple target images, including the first image, is increasing, the detection unit 11 may reduce the number of new feature points detected from the target image from the currently set number. Furthermore, if the number of corresponding feature points used to estimate the position and orientation of the imaging device for multiple target images, including the first image, is decreasing, the detection unit 11 may increase the number of new feature points detected from the target image from the currently set number.
[0028] Next, the flow of the self-position estimation process performed in the information processing device 10 according to Embodiment 1 will be explained using Figure 2. Self-position estimation is the process of estimating the position and orientation of the imaging device that captured the target screen.
[0029] First, the detection unit 11 detects multiple new feature points from the first image (S11). Next, the identification unit 12 identifies corresponding feature points from among the multiple new feature points that correspond to known feature points associated with a 3D position included in at least one control image used to generate the environment map (S12).
[0030] Next, the estimation unit 13 estimates the position and orientation of the imaging device that captured the first image using the corresponding feature points (S13). Then, the detection unit 11 changes the number of new feature points to detect from the target image for which the position and orientation of the imaging device are to be estimated, according to the number of corresponding feature points (S14).
[0031] As described above, the information processing device 10 according to Embodiment 1 changes the number of new feature points detected from the target image for which the position and orientation of the imaging device are to be estimated, according to the number of corresponding feature points. As a result, the information processing device 10 can maintain the accuracy of estimating the position and orientation of the imaging device that captured the target image, while also reducing the processing load.
[0032] (Embodiment 2) Next, an example of the configuration of the information processing device 20 according to Embodiment 2 will be described using Figure 3. The information processing device 20 may be a computer device, similar to the information processing device 10. The information processing device 20 has the same configuration as the information processing device 10, with the addition of an environment map generation unit 21, a feature point management unit 22, an acquisition unit 23, a feature point determination unit 24, and a detection count management unit 25. In the following description, detailed explanations of the same configuration and functions as the information processing device 10 in Figure 1 will be omitted.
[0033] The environmental map generation unit 21, feature point management unit 22, acquisition unit 23, feature point determination unit 24, and detection count management unit 25 may be software or modules whose processing is performed by the processor executing a program stored in memory. Alternatively, the environmental map generation unit 21, feature point management unit 22, acquisition unit 23, feature point determination unit 24, and detection count management unit 25 may be hardware such as circuits or chips. Alternatively, the feature point management unit 22 and detection count management unit 25 may be memory that stores data.
[0034] The information processing device 20 uses multiple images captured by the imaging device to estimate the position and orientation of the imaging device that captured each image in real time. For example, the information processing device 20 estimates the position and orientation of the imaging device that captured each image in real time by executing VSLAM. The information processing device 20 may also be used when correcting the position and orientation of an autonomously moving robot. In estimating the position and orientation of an autonomously moving robot, images captured in real time by the moving robot are compared with environmental images in the environmental map that are similar to the images captured in real time. Environmental images correspond to management images. The comparison between the images captured in real time and the environmental images is performed using feature points contained in each image. The position and orientation of the robot are estimated and corrected based on the comparison results. Here, the estimation and correction of the robot's position and orientation are performed by VSLAM. Furthermore, in this disclosure, the term "robot" is not limited to a form that constitutes a device as long as it can move, and broadly includes, for example, robots that mimic humans or animals, and transport vehicles that move using wheels (e.g., Automated Guided Vehicles). The transport vehicle may be, for example, a forklift.
[0035] The environmental map generation unit 21 may generate an environmental map by performing SfM using multiple images captured by the imaging device. If the information processing device 20 has a camera function, the environmental map generation unit 21 may generate an environmental map using images captured by the information processing device 20. Alternatively, the environmental map generation unit 21 may receive images captured by an imaging device, which is a different device from the information processing device 20, via a network or the like, and generate an environmental map.
[0036] The environment map generation unit 21 outputs the environment map and the multiple images used to generate the environment map to the feature point management unit 22. At this time, the environment map generation unit 21 may not output the image information directly to the feature point management unit 22, but may output only the information regarding the feature points detected in the images to the feature point management unit 22. The feature point management unit 22 manages the environment map and images received from the environment map generation unit 21. Alternatively, the feature point management unit 22 manages the information regarding the feature points received from the environment map generation unit 21. Furthermore, the feature point management unit 22 manages each image received from the environment map generation unit 21 in association with the position and orientation of the camera that captured each image. In addition, the feature point management unit 22 manages each image in association with the 3D coordinates of the feature points on the environment map within each image. The images managed by the feature point management unit 22 may be called keyframes. Here, keyframes can also be described as frame images that can serve as the starting point for the series of image processing described below. Furthermore, the three-dimensional coordinates of feature points within a keyframe on the environment map may be referred to as landmarks.
[0037] The acquisition unit 23 acquires multiple frame images that constitute an image or video captured by the imaging device. The acquisition unit 23 acquires images captured by the imaging device mounted on an autonomously moving robot in virtually real time. In other words, the acquisition unit 23 acquires images captured by the imaging device in real time in order to estimate the position and orientation of the autonomously moving robot or imaging device in real time. In the following description, the images acquired by the acquisition unit 23 will be referred to as real-time images.
[0038] The detection unit 11 detects feature points in the real-time image according to the target number of feature points to be detected. The detection unit 11 detects feature points in the real-time image in a manner that approaches the target number. Specifically, the detection unit 11 may detect the same number of feature points as the target number, or it may detect a number of feature points within a predetermined range that includes the target number. In other words, the detection unit 11 may detect more feature points than the target number, or it may detect fewer feature points than the target number. The difference between the number of feature points detected by the detection unit 11 and the target number is set to a value that is sufficiently small compared to the target number. In other words, the difference between the number of feature points detected by the detection unit 11 and the target number is set to a number that is recognized as an error with respect to the target number. The target number may be changed for each real-time image on which feature points are to be detected. Alternatively, the target number may be changed for multiple real-time images on which feature points are to be detected. In other words, the same target number may be applied to multiple real-time images.
[0039] The identification unit 12 identifies new feature points (new feature points) from among multiple feature points (new feature points) extracted from the real-time image that match feature points (known feature points) managed by the feature point management unit 22. Specifically, the identification unit 12 compares the feature vectors of the known feature points and the feature vectors of the new feature points and matches feature points that are close in distance indicated by the vectors. The identification unit 12 may also extract several images from among multiple images managed by the feature point management unit 22 and identify new feature points that match the known feature points contained in each image.
[0040] Here, the feature point matching process performed by the identification unit 12 will be explained using Figure 4. Figure 4 shows the feature point matching process using keyframes 60 and real-time images 50. u1, u2, and u3 in keyframes 60 are known feature points, and t1, t2, and t3 in real-time images 50 are new feature points detected by the detection unit 11. The identification unit 12 identifies t1, t2, and t3 as new feature points that match u1, u2, and u3, respectively. In other words, t1, t2, and t3 are corresponding feature points that correspond to u1, u2, and u3, respectively.
[0041] Furthermore, q1 is the 3D coordinate associated with the known feature point u1, and represents the landmark of the known feature point u1. q2 represents the landmark of the known feature point u2, and q3 represents the landmark of the known feature point u3.
[0042] Since the new feature point t1 matches the known feature point u1, the 3D coordinates of the new feature point t1 become landmark q1. Similarly, the 3D coordinates of the new feature point t2 become landmark q2, and the 3D coordinates of the new feature point t3 become landmark q3.
[0043] The estimation unit 13 estimates the position and orientation of the imaging device that captured the real-time image 50 using known feature points u1, u2, and u3, new feature points t1, t2, and t3, and landmarks q1, q2, and q3. Specifically, the estimation unit 13 first assumes the position and orientation of the imaging device 30 that captured the real-time image 50. The feature point detection unit 23 projects the positions of q1, q2, and q3 onto the real-time image 50, assuming that q1, q2, and q3 were captured at the assumed position and orientation of the imaging device 30. The estimation unit 13 repeatedly changes the position and orientation of the imaging device 30 that captured the real-time image 50 and projects the positions of q1, q2, and q3 onto the real-time image 50. The estimation unit 13 estimates the position and orientation of the imaging device 30 as the position and orientation at which the difference between the positions of q1, q2, and q3 projected onto the real-time image 50 and the feature points t1, t2, and t3 within the real-time image 50 is minimized.
[0044] Here, using Figure 5, the feature point classification process performed by the feature point determination unit 24 will be explained. The feature point determination unit 24 determines the positions where q1, q2, and q3 are projected onto the real-time image 50 as t'1, t'2, and t'3, assuming that the position and orientation of the imaging device 30 that captured the real-time image 50 are those estimated by the estimation unit 13. In other words, t'1, t'2, and t'3 are the positions of q1, q2, and q3 within the real-time image 50 when the imaging device 30 captured the real-time image 50 at the position and orientation estimated by the estimation unit 13. The dotted circles within the real-time image 50 in Figure 5 indicate t'1, t'2, and t'3.
[0045] Here, the feature point determination unit 24 determines the distance between t1 and t'1 associated with landmark q1, the distance between t2 and t'2 associated with landmark q2, and the distance between t3 and t'3 associated with landmark q1. In Figure 5, t'1 and t'3 are shown to be in substantially the same position as t1 and t3, or to be at or below a predetermined distance from t1 and t3. Figure 5 also shows that the position of t'2 is different from that of t2 and is at or above a predetermined distance from t2.
[0046] The fact that the position of t'2 is different from that of t2 indicates that t2 is misaligned from the position of landmark q2 which should be displayed in the real-time image 50. In other words, it indicates that the matching accuracy of the new feature point t2 to the known feature point u2 is low. On the other hand, the fact that the positions of t'1 and t'3 are substantially the same as those of t1 and t3 indicates that t1 and t3 coincide with landmarks q1 and q3 which should be displayed in the real-time image 50. In other words, it indicates that the matching accuracy of the new feature points t1 and t3 to the known feature points u1 and u3 is high. The feature point determination unit 24 refers to t2 in Figure 5 as a low-precision feature point, and t1 and t3 as high-precision feature points. In other words, the feature point determination unit 24 classifies t2 in Figure 5 as a low-precision feature point, and t1 and t3 as high-precision feature points. High-precision feature points may also be called inlier feature points, and low-precision feature points may be called outlier feature points.
[0047] The feature point determination unit 24 outputs to the detection count management unit 25 the number of high-precision feature points and the number of low-precision feature points of the new feature points used to estimate the position and orientation of the imaging device 30 in the real-time image 50. Alternatively, the feature point determination unit 24 may output only the number of high-precision feature points to the detection count management unit 25.
[0048] The detection number management unit 25 calculates the target number of feature points (target feature point number) to be detected from the real-time image acquired by the acquisition unit 23, using the number of high-precision feature points received from the feature point determination unit 24. The target number of feature points f n-i , n-i , n-i , n-i , n-i , n-i , n-i to be detected from the n-th real-time image acquired by the acquisition unit 23 may be calculated, for example, using the following formula 1.
[0049] (Formula 1) f n = f n-1 +α×(I - i n-i )
[0050] “I” represents the target number of high-precision feature points, and i n-i represents the number of high-precision feature points in the previous frame (previous image). Also, α represents a coefficient and is a number greater than 0. α may be the same value in both cases of I - i n-i ≧0 and I - i n-i <0, or may be different values in the case of I - i n-i ≧0 and the case of I - i n-i <0.
[0051] For example, when I - i n-i ≧0, α = 2, and when I - i n-i <0, α = 1 may be used. The case of I - i n-i ≧0 means that the number of specified high-precision feature points has not reached the target number of high-precision feature points, and the case of I - i n-i <0 means that the number of specified high-precision feature points has exceeded the target number of high-precision feature points. Making the value of α in the case of I - i n-i ≧0 larger than the value of α in the case of I - i n-i <0 means increasing the increase amount of the target number of feature points when the number of specified high-precision feature points has not reached the target number of high-precision feature points. Also, making the value of α in the case of I - i n-i ≧0 larger than the value of α in the case of I - i n-iMaking the value of α greater than the value when it is <0 means reducing the decrease in the target number of feature points when the number of identified high-precision feature points exceeds the target number of feature points. In other words, it means that the focus is on improving the accuracy of estimating the position and orientation of the imaging device rather than reducing the processing load of the position and orientation estimation process of the imaging device.
[0052] On the other hand, Ii n-i The value of α when ≥ 0 is Ii n-i The value of α may be smaller than the value when <0. This means that the focus is on reducing the workload of estimating the position and orientation of the imaging device rather than improving the estimation accuracy of the imaging device's position and orientation estimation process.
[0053] Furthermore, the coefficient α is the target feature score f n And, high-precision feature points i n We find the correlation function between and and the high-precision feature points i n Function g(i n ) may be determined using ). Alternatively, the coefficient α may be determined using a function that holds the estimation results of the position and orientation of the image for a certain period of time and takes the amount of change in position and orientation between the most recent real-time images as variables. The amount of change is, for example, velocity, and the function that takes the amount of change as a variable may be a function that takes velocity as a variable.
[0054] Furthermore, if the number of target feature points is too large, the processing load for estimating the position and orientation of the imaging device increases, and if the number of target feature points is too small, the accuracy of estimating the position and orientation of the imaging device decreases. For this reason, a maximum and minimum value may be set for the number of target feature points, and a value between the minimum and maximum value may be used for the target feature points.
[0055] The target number of high-precision feature points may be set to an arbitrary value by the administrator of the information processing device 20. For example, the administrator of the information processing device 20 may set the target number of feature points to a value they deem appropriate. Alternatively, the target number of high-precision feature points may be determined using machine learning. For example, the target number of high-precision feature points may be determined using a learning model that has learned the relationship between the number of high-precision feature points and the processing load or position and orientation estimation accuracy of the information processing device 20.
[0056] Next, the flow of the update process for the target feature point count according to Embodiment 2 will be explained using Figure 6. First, the detection unit 11 detects feature points from the real-time image acquired by the acquisition unit 23 according to the target feature point count (S21). Next, the identification unit 12 identifies new feature points from among the new feature points extracted from the real-time image that match known feature points managed by the feature point management unit 22 (S22).
[0057] Next, the feature point determination unit 24 classifies new feature points that match known feature points managed by the feature point management unit 22 into high-precision feature points and low-precision feature points, and determines the number of high-precision feature points (S23).
[0058] Next, the detection count management unit 25 determines whether the number of high-precision feature points is equal to or greater than the target number of high-precision feature points (S24). If the detection count management unit 25 determines that the number of high-precision feature points is equal to or greater than the target number of high-precision feature points (YES determination in S24), it updates the target number of feature points to decrease (S25). If the detection count management unit 25 determines that the number of high-precision feature points is less than the target number of high-precision feature points (NO determination in S24), it updates the target number of feature points to increase (S26).
[0059] As described above, the information processing device 20 according to Embodiment 2 identifies new feature points included in the real-time image that match known feature points. Furthermore, the information processing device 20 identifies high-precision feature points from among the new feature points that match known feature points, where the distance to the projected point obtained by projecting the 3D position of the known feature point onto the real-time image is shorter than a predetermined distance. Furthermore, the detection count management unit 25 determines the target number of feature points to extract from the real-time image according to the number of high-precision feature points.
[0060] Generally, the accuracy of position and orientation estimation improves as the number of high-precision feature points used to estimate the position and orientation of the imaging device increases. It is assumed that a certain number of high-precision feature points are included in the new feature points extracted from the real-time image. In this case, as the number of high-precision feature points increases, the number of new feature points extracted from the real-time image also increases. Therefore, the number of feature points used for position and orientation estimation accuracy also increases, and the processing load on the information processing device 20 increases. For this reason, when the number of high-precision feature points exceeds the target number of high-precision feature points, it is considered that the position and orientation estimation accuracy can be maintained at a sufficiently high level, and the target number of feature points extracted from the real-time image can be reduced. As a result, it is possible to maintain the position and orientation estimation accuracy while preventing an increase in the processing load for position and orientation.
[0061] Figure 7 is a block diagram showing an example configuration of the information processing device 10 and information processing device 20 (hereinafter referred to as "information processing device 10, etc.") described in the above-described embodiment. Referring to Figure 7, the information processing device 10, etc. includes a network interface 1201, a processor 1202, and a memory 1203. The network interface 1201 may be used to communicate with a network node. The network interface 1201 may include, for example, a network interface card (NIC) compliant with the IEEE 802.3 series. IEEE stands for Institute of Electrical and Electronics Engineers.
[0062] The processor 1202 reads and executes software (computer programs) from the memory 1203, thereby performing the processing of the information processing device 10, etc., as described using a flowchart in the above embodiment. The processor 1202 may be, for example, a microprocessor, an MPU, or a CPU. The processor 1202 may include multiple processors.
[0063] Memory 1203 is composed of a combination of volatile and non-volatile memory. Memory 1203 may also include storage located away from the processor 1202. In this case, the processor 1202 may access memory 1203 via an I / O (Input / Output) interface, which is not shown.
[0064] In the example shown in Figure 7, memory 1203 is used to store a group of software modules. The processor 1202 can read these software modules from memory 1203 and execute them, thereby enabling the information processing device 10 and the like, as described in the above embodiment.
[0065] As explained with reference to Figure 7, each of the processors in the information processing device 10, etc., in the above-described embodiment executes one or more programs that include a set of instructions for causing the computer to perform the algorithm described with reference to the drawings.
[0066] In the examples described above, the program includes a set of instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more of the functions described in the embodiments. The program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically or otherwise propagating signals.
[0067] Furthermore, the technical concepts in this disclosure are not limited to the embodiments described above, and may be modified as appropriate without departing from the spirit of the invention. [Explanation of symbols]
[0068] 10 Information Processing Devices 11 Detection Unit 12 Specific part 13 Estimation part 20 Information Processing Devices 21 Environmental Map Generation Unit 22 Feature Point Management Department 23 Acquisition Department 24 Feature Point Determination Unit 25 Detection Count Management Department 30 Imaging device 50 Real-time Images 60 keyframes
Claims
1. A detection unit that detects multiple new feature points from the first image, A selection unit identifies, among the plurality of new feature points, corresponding feature points that correspond to known feature points whose three-dimensional positions are associated with those included in at least one management image used to generate the environment map, The system includes an estimation unit that estimates the position and orientation of the imaging device that captured the first image using the corresponding feature points, The detection unit is An information processing device that changes the number of new feature points detected from a target image for which the position and orientation of the imaging device are to be estimated, according to the number of corresponding feature points.
2. The detection unit is The information processing apparatus according to claim 1, wherein if the number of corresponding feature points is greater than the target number, the number of new feature points detected from the target image is reduced to a number currently set, and if the number of corresponding feature points is less than the target number, the number of new feature points detected from the target image is increased to a number currently set.
3. The specified part is, The distance between the projection point obtained by projecting the three-dimensional position associated with the known feature point onto the first image and the corresponding feature point is shorter than a predetermined distance, thereby identifying high-precision feature points and low-precision feature points where the distance between the projection point and the corresponding feature point is longer than the predetermined distance. The detection unit is The information processing apparatus according to claim 1 or 2, wherein the number of new feature points detected from the target image for which the position and orientation of the imaging device are to be estimated is changed according to the number of high-precision feature points.
4. The detection unit is The information processing apparatus according to claim 3, wherein if the number of high-precision feature points is greater than the target number of high-precision feature points, the number of new feature points detected from the target image is reduced to a number currently set, and if the number of high-precision feature points is less than the target number of high-precision feature points, the number of new feature points detected from the target image is increased to a number currently set.
5. The detection unit is The information processing device according to claim 4, wherein if the number of high-precision feature points is greater than the target number of high-precision feature points, the value obtained by subtracting the target number of high-precision feature points from the number of high-precision feature points is subtracted from the currently set number of detected new feature points, and if the number of high-precision feature points is less than the target number of high-precision feature points, the value obtained by subtracting the number of high-precision feature points from the target number of high-precision feature points is added to the currently set number of detected new feature points.
6. The detection unit is If the number of high-precision feature points is greater than the target number of high-precision feature points, the value obtained by subtracting the target number of high-precision feature points from the number of high-precision feature points is multiplied by a first coefficient and subtracted from the currently set number of detected new feature points; if the number of high-precision feature points is less than the target number of high-precision feature points, the value obtained by subtracting the number of high-precision feature points from the target number of high-precision feature points is multiplied by a second coefficient and added to the currently set number of detected new feature points; the first coefficient and the second coefficient are positive numbers, and the first coefficient is smaller than the second coefficient, as described in claim 5.
7. The detection unit is The information processing apparatus according to claim 1 or 2, wherein the number of new feature points detected from the target image is changed within the range of the maximum and minimum values of the number of new feature points detected from the target image.
8. The information processing device is Multiple novel feature points are detected from the first image, Of the aforementioned multiple novel feature points, we identify the corresponding feature points that correspond to known feature points whose 3D positions are associated with those included in at least one management image used to generate the environment map. Using the aforementioned corresponding feature points, the position and orientation of the imaging device that captured the first image are estimated. A self-localization method that changes the number of new feature points detected from a target image for which the position and orientation of the imaging device are to be estimated, according to the number of corresponding feature points.
9. Multiple novel feature points are detected from the first image, Of the aforementioned multiple novel feature points, we identify the corresponding feature points that correspond to known feature points whose 3D positions are associated with those included in at least one management image used to generate the environment map. Using the aforementioned corresponding feature points, the position and orientation of the imaging device that captured the first image are estimated. A program that causes a computer to change the number of new feature points detected from the target image used to estimate the position and orientation of the imaging device, according to the number of corresponding feature points.
Citation Information
Patent Citations
Image processor, image processing method and image processing program
JP2018036901A
Autonomous moving apparatus, autonomous moving method and program
JP2020057187A
Information processing device, information processing method, and program
JP2021009557A
Image processing device, image processing method, and image processing program
WO2017022033A1
Information processing device, information processing method, and program
WO2020095541A1