Keypoint generation program, apparatus and method using three-dimensional points, and image matching and camera position and orientation determination program

The keypoint generation program addresses the challenges of image localization by generating target keypoints for query images based on reliable 3D points, enhancing matching accuracy and reducing computational requirements, thus improving the efficiency of image localization.

JP7692388B2Active Publication Date: 2025-06-13KDDI CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022083000
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-06-13
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Existing image localization techniques face challenges in accurately determining the position and orientation of a camera related to a query image, especially when images are taken from significantly different viewpoints or under varying illuminations, leading to difficulties in keypoint matching and image search processing.

Method used

A keypoint generation program that determines highly reliable images within an image group, converts these images' keypoints into 3D points, and generates target keypoints for the query image without using machine learning models like transformers, enabling more suitable image matching and camera pose determination.

Benefits of technology

This approach allows for more accurate and robust image matching and camera position/orientation determination, overcoming the limitations of existing techniques by reducing calculation time and improving real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692388000005
    Figure 0007692388000005
  • Figure 0007692388000006
    Figure 0007692388000006
  • Figure 0007692388000007
    Figure 0007692388000007
Patent Text Reader

Abstract

To provide a program that can generate key points of an objective image that can be used for matching with a target image without using a machine learning model such as a transformer, which requires a huge amount of computing time.SOLUTION: The program generates key points of an objective image that can be used for matching the objective image included in a certain image group with a target image. The program causes a computer to function as: means for determining a highly reliable image in the image group that is considered easier to match with the target image compared to the objective image in terms of key points to be matched or camera information; means for transforming the key points of the highly reliable image into three-dimensional (3D) points in a camera coordinate system using camera information of the highly reliable image to generate objective 3D points in the camera coordinate system pertaining to the objective image from the 3D points; and means for generating key points for the objective image using the camera information pertaining to the objective image from the objective 3D points.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image localization technique for performing visual localization from an image.

Background Art

[0002] The image localization technique is, for example, an essential technique for an autonomous vehicle or an autonomous mobile robot to know its current position, and is currently being actively developed.

[0003] What is important in this technique is the determination process of similarity (similarity) between a query image as a target image for which the position is to be specified and an image in an image database. Here, generally, in order to improve the similarity determination accuracy, two types of similarity, so-called global similarity and local feature similarity, are used.

[0004] Among these, the global similarity is used when searching for a group of more similar images for a given position from an image database. Next, a matching process is performed between each image retrieved using the other local similarity and the query image, and thereby, the camera pose of the camera related to the query image, specifically its position and orientation, is estimated.

[0005] For example, Patent Documents 1 and 2 disclose a typical image localization technique in which first, a group of images related to a query image is retrieved from an image database, then a keypoint matching process is performed between each retrieved image and the query image using the keypoints of the images, and finally, the camera pose of the camera related to the query image is obtained using this matching result.

[0006] In order to further improve the accuracy of similarity determination between images, for example, in the technologies disclosed in Patent Document 2 and Non-Patent Document 1 above, universal features applicable to both image search processing and keypoint matching processing are adopted. Specifically, in the technology of Patent Document 2, global features are fused with local features used in keypoint matching processing and then used. Also, in the technology of Non-Patent Document 1, hierarchical features compatible with both local features and global features are adopted.

[0007] Furthermore, for example, the technology disclosed in Non-Patent Document 2 aims to improve the accuracy of matching processing based on local features by using a deep learning model, the Transformer. By using the Transformer here, it is said that keypoint matching processing between images that are significantly different in terms of viewpoint and illuminance can be performed more accurately.

Prior Art Documents

Patent Documents

[0008]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0009]

Non-Patent Document 1

[0010] However, even with the prior art as described above, there still remains an essential gap between the purpose of image search processing and the purpose of keypoint matching processing. Therefore, it has not been possible to solve the problem that in many cases, it is difficult to finally perform good image localization.

[0011] In fact, in conventional image search processing, efforts have been made to find images with higher relevance to a query image from an image database. For example, images containing the same target (such as a landmark like a building) as the query image are searched as much as possible. Therefore, among the retrieved image group, there is a high possibility that it contains, for example, "images as a result of photographing the same target from significantly different viewpoints or under significantly different illuminations (compared to the query image)".

[0012] On the other hand, in conventional keypoint matching processing, local features in an image, specifically features at the pixel level, are focused on, and matching is performed using local similarity. Therefore, for example, for the "image" as described above and the query image, corresponding keypoints cannot be obtained sufficiently in the first place. As a result, the robustness and accuracy of the keypoint matching processing decrease. Also, as a result, it becomes difficult to improve the accuracy of specifying the position of the query image (estimating the position and orientation of the camera).

[0013] In this regard, even with the universal feature amounts disclosed in Patent Document 2 and Non-Patent Document 1, in reality, optimization of both image search processing and keypoint matching processing has not been achieved. Also, according to the technique using a transformer disclosed in Non-Patent Document 2, it is indeed possible to improve the accuracy of matching processing between images with significantly different viewpoints, for example. However, this technique usually requires an enormous amount of calculation time for matching processing and cannot be applied at all to applications that require real-time performance, such as utilizing the image position specification result with an autonomous mobile body.

[0014] Therefore, an object of the present invention is to provide a keypoint generation program, apparatus, and method that can generate keypoints of a target image that can be used for matching with a target image (query image) without using a machine learning model such as a transformer that requires a huge amount of calculation time. Another object is to provide an image matching program that can perform more suitable image matching by performing such keypoint generation processing, and a camera position and orientation determination program that can determine the position and orientation of a camera related to a more suitable target image.

Means for Solving the Problems

[0015] According to the present invention, there is provided a keypoint generation program for generating keypoints of a target image that can be used for matching between a target image and a target image included in an image group, image reliability determination means for determining a highly reliable image that is a highly reliable image included in the image group and is considered to be more easily matched with the target image than the target image from the viewpoint of matching keypoints or camera information; target 3D point generation means for converting the keypoints of the highly reliable image into 3D points in the camera coordinate system related to the highly reliable image using the camera information related to the highly reliable image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points; target keypoint generation means for generating keypoints for the target image using the camera information related to the target image from the target 3D points A keypoint generation program is provided that causes a computer to function.

[0016] In the keypoint generation program according to the present invention, the image reliability determination means determines a plurality of the highly reliable images, The target 3D point generation means generates the target 3D points related to each of the plurality of highly reliable images using the keypoints in each of the plurality of highly reliable images that match the target image. The target keypoint generation means preferably generates a plurality of the keypoints for the target image from the target 3D points corresponding to each of the plurality of highly reliable images.

[0017] Also, it is preferable that the image group is one of a plurality of image groups generated by classifying the images included in a certain image group based on the position and orientation of the camera related to the image.

[0018] Furthermore, as an embodiment of the keypoint generation program according to the present invention, the image reliability determination means determines the highly reliable image based on the ratio or number of keypoints that match the target image in the images included in the image group, or based on the distance between the position of the camera related to the images included in the image group and the representative position of the camera related to the image group.

[0019] Also, in this embodiment, when determining the highly reliable image, the image reliability determination means preferably determines whether to be based on the ratio or number of the matching keypoints or based on the distance according to the relative relationship between the position of the camera related to the images included in the image group and the representative position of the camera related to the image group.

[0020] According to the present invention, there is also provided an image matching program for performing matching between a target image and a target image included in an image group, When the image group is one of certain a plurality of image groups generated by classifying the images included in the image group based on the position and orientation of the camera related to the image and is one of , for each of the image groups, an image reliability determination means for determining a highly reliable image included in the image group, which is considered to be more easily matched with the target image than the target image from the viewpoint of matching keypoints or camera information For each of the image groups, target 3D point generation means for converting the key points of the highly reliable image into 3D points in the camera coordinate system related to the highly reliable image using the camera information related to the highly reliable image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points; For each of the image groups, target key point generation means for generating key points for the target image using the camera information related to the target image from the target 3D points; Matching means for generating a dense descriptor for the target image using the generated key points of the target image, and performing matching between the target image and the target image using the dense descriptor An image matching program that causes a computer to function is provided.

[0021] According to the present invention, further, an image groups included in target A camera position and orientation determination program for determining the position and orientation of a camera related to a target image using an image, When the image group is one of certain A plurality of image groups generated by classifying images included in an image group based on the position and orientation of the camera related to the image and is one of , For each of the image groups, image reliability determination means for determining a highly reliable image that is a highly reliable image included in the image group and is considered to be more easily matched with the target image than the target image from the viewpoint of matching key points or camera information; For each of the image groups, target 3D point generation means for converting the key points of the highly reliable image into 3D points in the camera coordinate system related to the highly reliable image using the camera information related to the highly reliable image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points; For each of the image groups, target key point generation means for generating key points for the target image using the camera information related to the target image from the target 3D points; Based on the matching result between the target image and the target image using the generated key points, the position and orientation of the camera are derived for each of the image groups, and based on the matching result between the high-reliability image and the target image, the position and orientation of the camera are derived for each of the image groups, and from the derived positions and orientations of the plurality of cameras, a camera position and orientation determination means for determining the position and orientation of the camera related to the target image A camera position and orientation determination program is provided to cause a computer to function.

[0022] As an embodiment of the camera position and orientation determination program according to the present invention, it is also preferable that the camera position and orientation determination means generates a dense descriptor for the target image using the key points of the generated target image, performs matching between the target image and the target image using the dense descriptor, and generates a matching result between the target image and the target image.

[0023] Also, as another embodiment of the camera position and orientation determination program according to the present invention, the camera position and orientation determination means performs a weighted average with individual weights added to the position and orientation of the camera derived based on the matching result between the high-reliability image and the target image and the position and orientation of the camera derived based on the matching result between the target image and the target image, and determines the result as the position and orientation of the camera related to the target image. It is also preferable.

[0024] According to the present invention, also, a key point generation device that generates key points of the target image that can be used for matching between the target image and the target image included in the image group, An image reliability determination means for determining a high-reliability image that is a high-reliability image included in the image group and is considered to be more easily matched with the target image than the target image from the viewpoint of matching key points or camera information Target 3D point generation means for converting the key points of the high-reliability image into 3D points in the camera coordinate system related to the high-reliability image using the camera information related to the high-reliability image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points, Target key point generation means for generating key points for the target image using the camera information related to the target image from the target 3D points A key point generation device having the above is provided.

[0025] According to the present invention, further, a key point generation method for generating key points of the target image that can be used for matching between the target image and the objective image included in the image group, comprising: Determining a high-reliability image included in the image group, which is considered to be more easily matched with the objective image compared to the target image from the perspective of matching key points or camera information; Converting the key points of the high-reliability image into 3D points in the camera coordinate system related to the high-reliability image using the camera information related to the high-reliability image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points; Generating key points for the target image using the camera information related to the target image from the target 3D points A key point generation method implemented by a computer, characterized by having the above steps, is provided.

Effect of the Invention

[0026] According to the keypoint generation program, apparatus, and method of the present invention, keypoints of a target image that can be used for matching with a target image (query image) can be generated without using a machine learning model such as a transformer that requires a huge amount of calculation time. Further, according to the image matching program of the present invention, it is possible to perform more suitable image matching by performing such keypoint generation processing. Furthermore, according to the camera position and orientation determination program of the present invention, it is possible to determine the position and orientation of a camera related to a more suitable target image by performing such keypoint generation processing.

Brief Description of Drawings

[0027]

Figure 1

Figure 2

Figure 3

Figure 4

Embodiments for Carrying Out the Invention

[0028] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0029] [Keypoint Generation Apparatus, Camera Position and Orientation Determination Apparatus] FIG. 1 is a functional block diagram showing a functional configuration in an embodiment of a keypoint generation apparatus according to the present invention.

[0030] The camera position and orientation determination apparatus 1 of the present embodiment shown in FIG. 1 also serves as an embodiment of a pre-search processing apparatus according to the present invention, and as its pre-search processing, (a) A keypoint generation process that generates keypoints of a target image (hereinafter abbreviated as "target image") that is a target for keypoint generation and is included in an "image group" determined from an image group stored and managed in an image database (DB) 2, and is used for matching with a query image as a target image is an apparatus for performing. Furthermore, (b) A camera position and orientation determination process that determines the position and orientation of a camera related to a query image (target image) using the keypoints of the "target image" generated in (a) above is also an apparatus capable of performing.

[0031] Note that in this embodiment, the "image group" in (a) above is, as will be described later, one of a plurality of images included in the image group of the image database 2 classified based on the position and orientation of the camera related to the image by an image group determination unit 111. Of course, for example, a given image group can also be used. Also, in FIG. 1, this image database 2 is installed outside the camera position and orientation determination apparatus 1, but of course, it may be provided inside the camera position and orientation determination apparatus 1 as a component of the apparatus.

[0032] Here, in order to realize the keypoint generation process in (a) above, the camera position and orientation determination apparatus 1 specifically includes (A) An image reliability determination unit 112 that determines a "high-reliability image" that is a "high-reliability image" included in an "image group" and is considered to be more easily matched with a query image (target image) compared to the "target image" from the perspective of matching keypoints or camera information (B) A target 3D point generation unit 113 that converts the keypoints of the "high-reliability image" into 3D points in the camera coordinate system related to the "high-reliability image" using the camera information related to the "high-reliability image", and generates target 3D points in the camera coordinate system related to the "target image" from the generated 3D points (C) A target keypoint generation unit 114 that generates keypoints for the target image using the camera information related to the target image from the target 3D points has

[0033] In this way, when generating the key points of the "target image", the camera position and orientation determination device 1 uses the 3D points of the "high-reliability image" that are considered to be more easily matched with the query image compared to the "target image". Here, these 3D points are generated based on the more reliable key points in the "high-reliability image", and furthermore, they are information related to the position (in the three-dimensional geometric shape) of the key points in the three-dimensional space. That is, compared with the two-dimensional points on the image, they are information that represents the characteristics of the image content in more detail.

[0034] As a result, the key points of the "target image" generated from these 3D points are suitably usable for matching with the query image. Furthermore, it can be said that when generating such key points, the camera position and orientation determination device 1 does not require a machine learning model that requires a huge amount of calculation time, such as a Transformer. That is, it is possible to generate the key points of the target image that can be used for matching with the query image without relying on a machine learning model such as a Transformer that requires a huge amount of calculation time. Also, thereby, it becomes possible to perform the key point generation process of the "target image" and the matching process with the query image in real time, for example, online, after obtaining the query image.

[0035] Hereinafter, the functional configuration of the camera position and orientation determination device (key point generation device) 1 of the present embodiment will be described in more detail.

[0036] [Device Functional Configuration, Key Point Generation Program / Method, Image Matching / Camera Position and Orientation Determination Program / Method] According to the functional block diagram of FIG. 1, as an embodiment of the present invention, a camera position and orientation determination device (keypoint generation device) 1 includes an input / output interface (IF) unit 101 and a processor / memory (an arithmetic processing system having a memory function). This processor / memory stores an embodiment of a camera position and orientation determination program including a keypoint generation program according to the present invention, and also has a computer function. By executing this camera position and orientation determination program, a camera position and orientation determination process (keypoint generation process) is performed.

[0037] Therefore, the camera position and orientation determination device (keypoint generation device) 1 may be a device dedicated to the camera position and orientation determination process (keypoint generation process), but it can also be a cloud server, a non-cloud server, a personal computer (PC), a notebook or tablet computer, a mobile terminal such as a smartphone, or even a wearable terminal such as an HMD (Head Mounted Display) equipped with the camera position and orientation determination program (keypoint generation program) according to the present invention.

[0038] Furthermore, this processor / memory (a) includes an image group determination unit 111, an image reliability determination unit 112 including a keypoint (KP) - based reliability determination unit 112a and / or a position - based reliability determination unit 112b, a target 3D point generation unit 113 including a high - reliability 3D point generation unit 113a, and a target keypoint generation unit 114, (i) and a camera position and orientation determination unit 121 including a matching unit 121a functions. That is, these functional components can be regarded as functions realized by executing the camera position and orientation determination program (keypoint generation program) stored in the processor / memory. Also, the processing flow shown by connecting the functional components of the camera position and orientation determination device (keypoint generation device) 1 in FIG. 1 with arrows is also understood as an embodiment of the camera position and orientation determination method (keypoint generation method) according to the present invention.

[0039] Incidentally, an apparatus equipped with a key point generation program for embodying the functional configuration unit (a) above can be regarded as a key point generation apparatus according to the present invention even if it does not include the functional configuration (i) above. In this case, this key point generation apparatus, together with an apparatus equipped with a program for embodying the functional configuration unit (i) above, constitutes a camera position and orientation determination system.

[0040] (Image group determination process) Similarly, in the functional block diagram of FIG. 1, in the present embodiment, the image group determination unit 111 extracts the "image group" retrieved from the image database 2 based on the query image Iq via the input / output interface unit 101, and classifies the images included in the retrieved "image group" into a plurality of "image groups" based on the position and orientation of the camera related to the image.

[0041] Here, the search of the "image group" from the image database 2 can be carried out using various image search methods such as a method based on the similarity of images. For example, image search may be performed using a search site provided for searching the image database 2. In any case, the image I included in the retrieved "image group" can be regarded as an image that is somewhat similar or related to the query image Iq.

[0042] Also, as the position and orientation of the camera for each image used when generating a plurality of "image groups", the camera pose information pre-associated with the images stored and managed in the image database 2 can be used. In fact, there are quite a few image databases that associate camera pose information estimated in advance, for example, using SfM (Structure from Motion), with each of the stored and managed images. As a modification, it is also possible for the image group determination unit 111 to determine the camera position and orientation data for each image of the "image group" retrieved, for example, using SfM. In any case, as will be described in detail later, the camera position and orientation determination device 1 finally determines a camera position and orientation with higher accuracy (related to the query image Iq) by using information derived from the position and orientation of the camera associated (or determined) here.

[0043] Next, the image group determination unit 111 generates a plurality of image groups (G1, G2, ···, G = G1 ∪ G2 ∪ ···) based on the acquired camera position and orientation (camera position and camera orientation) for each image of the image group G (retrieved from the image database 2).

[0044] Specifically, in the present embodiment, the image group determination unit 111 calculates the "average value of the camera positions" and the "average value of the camera orientations" for all the images included in the image group G (hereinafter, such an average virtual camera is referred to as a camera centroid). (a) With the orientation of the camera centroid facing forward, the images in which the camera position exists on the right side of the camera centroid (for example, in the range from the orientation tilted 90° forward when viewed from directly to the right side to the orientation tilted 44.9° backward) are regarded as belonging to the image group G1. (b) With the orientation of the camera centroid facing forward, the images in which the camera position exists behind the camera centroid (for example, in the range from the orientation tilted 45° to the right when viewed from directly behind to the orientation tilted 45° to the left) are regarded as belonging to the image group G2. (c) With the direction of the camera centroid as the front, an image in which the position of the camera exists on the left side of the camera centroid (for example, in the range from an orientation tilted 89.9° forward to an orientation tilted 44.9° backward as viewed straight from the left side) shall belong to the image group G3. In the above manner, image groups (G1, G2, G3) are generated.

[0045] Here, for each image I i (i = 1, 2, ···, N) in the N images included in the image group G, if the position of the camera is (X i , Y i , Z i ) and the orientation of the camera is (α i , β i , θ i )(α is the roll angle, β is the pitch angle, θ is the yaw angle), then the following formula

Equation

[0046] Of course, the generation of the image group is not limited to the above. For example, classification into the image group can also be performed according to the orientation of the camera related to each image. Also, for the above (a) to (c), further classification conditions (related to the range) of the orientation of the camera related to the image may be added. Furthermore, without using the camera centroid as a reference, it is also possible to classify each image by performing clustering processing regarding the position, for example, based on the relative relationship of the positions of the cameras related to each image. Incidentally, in FIG. 2 to be described later, a specific example of classifying the images of the image group G into two of the above-described image groups G1 and G2 is shown.

[0047] (Image reliability determination process) Also in the functional block diagram of FIG. 1, in this embodiment, the image reliability determination unit 112, in each image group Gi (i = 1, 2, 3), (a) a highly reliable image I that is considered to be more easily matched with the query image Iq from the perspective of matching keypoints or camera information R (∈Re Gi : highly reliable image group to be described later) and (b) a low-reliability image I that is considered to be more difficult to match with the query image Iq from the perspective of matching keypoints or camera information W (∈W Gi : low-reliability image group to be described later) and are determined.

[0048] Here, the low-reliability image I in the above (b) R corresponds to the "target image" in the above (A) that is the target for keypoint generation. That is, the highly reliable image I in the above (a) R can also be said to be an image that is considered to be more easily matched with the query image Iq compared to the low-reliability image I W (target image).

[0049] Specifically, first, the KP-based reliability determination unit 112a of the image reliability determination unit 112 <KP-based reliability determination> For example, using the nearest neighbor matching method, for each image I included in the image group Gi, after detecting the feature points (keypoints) in the image I, keypoint matching processing with the query image Iq is performed, and based on the ratio Ra of the keypoints that match the query image Iq (to all detected keypoints) or the number of such keypoints, the reliability level RE(I) (= highest, high, medium, low) of the image I is determined and decided.

[0050] Note that this <KP-based reliability determination> is based on the fact that, for example, images with significantly different viewpoints and illuminances usually have fewer matching keypoints, and conversely, images with similar viewpoints and illuminances are likely to have more matching keypoints.

[0051] On the other hand, the position-based reliability determination unit 112b of the image reliability determination unit 112 <Position-based reliability determination> For each image I included in the image group Gi, the distance D between the position of the camera related to the image I and the representative position of the camera related to the image group G (in this embodiment, the position of the camera centroid) is calculated, and based on the calculated distance D, the reliability level RE(I) (= highest, high, medium, low) of the image I is determined and decided.

[0052] Here, this <position-based reliability determination> is based on the high possibility that the position and orientation of the camera centroid of the image group G (searched based on the query image Iq) are generally close to the position and orientation of the camera related to the query image Iq.

[0053] FIG. 2 is a schematic diagram and a table for explaining a specific example of the image group determination process and the image reliability determination process according to the present invention.

[0054] First, as shown in FIG. 2(A), in this specific example, the image group determination unit 111 includes 10 images I i (i = 1, 2, ···, 10) into (a) an image group G1 including images I 1 ~I 4 where the position of the camera exists on the right side of the camera centroid, and (b) an image group G2 including images I 5 ~I 10 where the position of the camera exists behind the camera centroid, and classifies them.

[0055] Next, the image reliability determination unit 112 determines high-reliability images I in each image group (G1, G2)R and low - reliability image I W When determining R and W , it is determined whether to perform <KP - based reliability determination> or <position - based reliability determination> based on the relative relationship between the position of the camera related to the image included in the image group and the representative position (the position of the camera centroid) of the camera related to the image group G. For example, when the variance of the distance between the position of the camera related to the image included in a certain image group and the position of the camera centroid is greater than a predetermined threshold, <position - based reliability determination> is applied to the said image group. On the other hand, it is also possible to apply <KP - based reliability determination> to the image group where the variance of the distance is less than or equal to the predetermined threshold.

[0056] In this specific example, it is considered that the landmark is generally captured from its side. As a result, for image I where there is a high possibility of a large difference in the number of keypoints that match compared to the distance from the camera centroid 1 ~I 4 <KP - based reliability determination> is performed on image group G1 including images. On the other hand, it is considered that the landmark is generally captured from its front. As a result, for image I where there is a high possibility of a large difference in the distance from the camera centroid compared to the number of keypoints that match 5 ~I 10 <position - based reliability determination> is performed on image group G2 including images.

[0057] Incidentally, in the table of Fig. 2(B), a specific example of <KP - based reliability determination> where the higher the ratio Ra of the matching keypoints, the higher the reliability level RE, and a specific example of <position - based reliability determination> where the smaller the distance D from the position of the camera centroid, the higher the reliability level RE are shown. Here, as a modification mode, it is also possible to set to perform only one of <KP - based reliability determination> and <position - based reliability determination> in any image group. However, by using both determinations properly as in this specific example, it is possible to obtain a more accurate reliability level RE.

[0058] The image reliability determination unit 112 performs the determination process as described above for each image group (G1, G2). In this specific example, the images for which the reliability level RE is determined to be 'highest' or 'high' are determined as high-reliability images, and the images for which the reliability level RE is determined to be'medium' or 'low' are determined as low-reliability images (target images). Specifically, (a) In the image group G1, the images I 1 and I 2 for which the reliability level RE is determined to be'medium' are used as low-reliability images. On the other hand, the images I 3 and I 4 for which the reliability level RE is determined to be 'highest' are used as high-reliability images. (b) In the image group G2, the images I 5 ~I 7 for which the reliability level RE is determined to be 'high' are used as high-reliability images. On the other hand, the images I 8 ~I 10 for which the reliability level RE is determined to be 'low' are used as low-reliability images.

[0059] Summarizing the image reliability determination process described above, it is as follows. That is, the image reliability determination unit 112, for each image group Gi (i = 1, 2,...) of the image group G (= G1 ∪ G2 ∪ ···), uses the following expressions (1) Re Gi ={I|RE(I) is 'highest' or 'high', I ∈ Gi} (2) W Gi ={I|RE(I) is'medium' or 'low', I ∈ Gi} to determine the high-reliability image group Re Gi and the low-reliability image group W Gi (Re Gi ∪W Gi = Gi).

[0060] Here, the high-reliability image group Re Gi determined for each image group Gi is thereafter treated as a reliable projection source for key points, while the low-reliability image group W Gi is treated as the projection destination, or in other words, as the 'target image' for generating highly reliable key points. That is, in this embodiment, the high-reliability image group ReGi At least one, preferably a plurality of high - reliability images I included in R A plurality of key points detected from are projected onto each low - reliability image I included in the low - reliability image group W Gi W .

[0061] Note that when the number of low - reliability images I to be projected is large, the calculation time of this projection process will increase. Therefore, in order to reduce the total calculation time, the high - reliability image I that is the projection source of the key points in each low - reliability image I W is limited to those in the vicinity of the low - reliability image I (for example, the high - reliability image I whose camera position is within a predetermined distance from the camera position related to the low - reliability image I W ). Also, it is also possible to use the high - reliability image I whose camera position is closest to the camera position related to the low - reliability image I R as the projection source of the key points in the low - reliability image I W . W . R ) W . R W .

[0062] Also, as a modification, when the image group Gi contains only images with a trust level RE of'maximum' and 'high', it is also possible to set the 'high' images as low - reliability images. Further, when the image group Gi contains only images with a trust level RE of'medium' and 'low', it is also possible to set the'medium' images as high - reliability images. Also, in order to prevent meaningless projection processing from being performed, it is also preferable that an image group G containing a sufficient number of different images is prepared so that each image group Gi contains at least two images with different trust levels RE from each other.

[0063] (Target 3D point generation process) In the functional block diagram of FIG. 1, in this embodiment, in each image group Gi, for each high - reliability image I included in the high - reliability image group Re Gi R , ​​​(a) For example, use the nearest neighbor matching method to detect the key points of the high - reliability image I R Among the detected key points, the key points that match the query image Iq are converted into "high - reliability 3D points", which are 3D points in the camera coordinate system related to the high - reliability image I R using the camera information related to the high - reliability image I R (b) From the "high - reliability 3D points" generated as a result of the conversion, "target 3D points" in the camera coordinate system related to the low - reliability image I W to be projected are generated, and thereby, in each image group Gi, the "target 3D points" related to the projection from the high - reliability image group Re Gi (to the low - reliability image I W ) are determined.

[0064] Specifically, as described in (a) above, the high - reliability 3D point generation unit 113a of the target 3D point generation unit 113 performs keypoint matching between the high - reliability image I Gi included in the high - reliability image group Re and the query image Iq for each high - reliability image I R and determines a set M R of indices of the keypoints that match the query image Iq (={(c Iq→IR , c Iq , c IR )}). Here, c Iq and c IR are respectively the indices of the matching keypoints of the high - reliability image I R and the query image Iq. Furthermore, the following equation (3) K IR ={[x c , y c T |c∈c IR} is used to determine the group of matching keypoints K R in the high - reliability image I. Here, x IR and y c are respectively c the coordinates of the matching keypoints in the high - reliability image I R ​​They are the x - coordinate value and y - coordinate value of the keypoint of index c in the x - y image coordinate system according to

[0065] The high - reliability 3D point generation unit 113a then uses the keypoint group K R included in the high - reliability image I IR of the keypoint k (in the xy image coordinate system) IR (=[x c , y c ) T , k IR ∈K IR ) and converts it into a 3D point (high - reliability 3D point) p in the X - Y - Z camera coordinate system according to the following image - camera coordinate conversion formula

Equation

[0066] Next, as described in (b) above, the target 3D point generation unit 113 specifically generates, for each high-reliability image I Gi included in the high-reliability image group Re R the high-reliability 3D point group P IR in the X-Y-Z camera coordinate system generated for each high-reliability 3D point p IR and converts it into a 3D point (target 3D point) p

Equation

[0067] Incidentally, the target 3D point group P IR→W described above is generated for each keypoint group K Gi related to each high-reliability image I R included in the high-reliability image group Re IR .

[0068] (Target keypoint generation process) Similarly, in the functional block diagram of FIG. 1, the target keypoint generation unit 114 receives, from the target 3D point generation unit 113, each keypoint k Gi related to each high-reliability image I R included in the high-reliability image group Re IR generated for each low-reliability image I WTarget 3D point group P in the U-V-W camera coordinate system according to IR→W is received, and for this target 3D point group P IR→W the target 3D point p included in IR→W using the camera information related to the low-confidence image (target image) I W this low-confidence image (target image) I W keypoints k for IR→W are generated.

[0069] Specifically, the target keypoint generation unit 114 converts the target 3D point P in the U-V-W camera coordinate system IR→W using the following camera-image coordinate conversion formula

Equation

[0070] Next, the target keypoint generation unit 114 combines all the keypoints k R converted and generated from a high-confidence image I IR to generate a keypoint group K IR→W (={k IR→W (={k IR→W}) and further, using the following formula (4) K W =∪ IR K IR→W to determine the keypoint group K W for one low-confidence image (target image) I W Here, ∪ IR is the union of all the high-confidence images I Gi included in the high-confidence image group Re​R (∈Re Gi ) over the high-reliability image I R related to K IR→W is a union operator that forms a set combining them.

[0071] Also, in this embodiment, the target keypoint generation unit 114 determines the keypoint group K Gi for all the low-reliability images (target images) I W included in the low-reliability image group W of the image group Gi as described above, and further performs such target keypoint determination processing for all the image groups Gi (i = 1, 2, ···). W will be carried out.

[0072] FIG. 3 is a schematic diagram for explaining a specific example of the target 3D point generation process and the target keypoint generation process according to the present invention.

[0073] First, according to FIG. 3(A), in the image group G2 (shown in FIG. 2(A)), for each of the high-reliability images I G2 included in the high-reliability image group Re 5 , I 6 and I 7 , using the keypoints that match a given query image Iq, the target 3D points are generated with the low-reliability image I 10 as the target image, and further from these, the keypoint group (K 10 , K I5→I10 , K I6→I10 , K I7→I10 ) for the low-reliability image I I5→I10 is generated and determined. In this specific example, all the keypoints included in these keypoint groups K I6→I10 and K I7→I10 are the keypoints of the low-reliability image I 10 and are suitable keypoints because of the high-reliability images.

[0074] In this regard, according to FIG. 3(B), the low-reliability image I 10It is understood that finally a number of suitable key points are projected (generated). In FIG. 3(B), for ease of viewing the drawing, only the key point(s) related to the high-reliability images I 5 and I 6 are shown. Also, in this specific example, although not shown, for each of the remaining low-reliability images (I 8 , I 9 ), key points are similarly projected (generated) based on each of the high-reliability images I 5 , I 6 and I 7 .

[0075] Thus, in the image group Gi, if a number of high-reliability images I R (∈Re Gi ) are determined, the target 3D point generation unit 113 generates a number of target 3D points related to each of these high-reliability images I R , and further the target key point generation unit 114 can generate a number of suitable key points for each of these low-reliability images (target images) I W (∈W Gi ) from these target 3D points. As a result, it is also possible to perform image matching processing and camera position and orientation determination processing with higher robustness and accuracy later.

[0076] As described above, using the specific example shown in FIG. 3, it has been explained that suitable key points for low-reliability images (target images) can be generated. Such key points can of course be used for the matching processing related to low-reliability images (target images). However, these suitable key points are used to estimate the position and orientation of the camera related to the query image Iq, as will be described in detail later, in the camera position and orientation determination apparatus 1 (FIG. 1) of the present embodiment.

[0077] Here, a brief summary of the target 3D point generation process and the target key point generation process described above will be given. In the present embodiment, these processes are represented by the following equation (5) K IR (2D high - reliability image coordinate system) →P IR (3D X - Y - Z camera coordinate system) → P IR→W (3D U - V - W camera coordinate system) →K IR→W (2D low - reliability image coordinate system) According to the flow shown, the low - reliability image (target image) I W for the suitable keypoint group K IR→W is obtained.

[0078] Here, the intermediate 3D point group P IR is generated based on the highly reliable keypoints in the high - reliability image I R which is considered to be more easily matched with the query image Iq. Furthermore, it is information related to the position (in the three - dimensional geometric shape) of the keypoints in the three - dimensional space. That is, compared with the two - dimensional points on the image, it is information that represents the features of the image content in more detail. As a result, the keypoint group K IR for the low - reliability image (target image) I W generated from this 3D point group P IR→W is more suitably usable for the matching process.

[0079] Also, in this embodiment, when generating such suitable keypoints, for example, a machine - learning model that requires a huge amount of calculation time such as a transformer is not required. That is, the suitable keypoint group K W for the low - reliability image (target image) I IR→W can be generated without relying on a machine - learning model such as a transformer that requires a huge amount of calculation time.

[0080] (Camera position and orientation determination process) In the functional block diagram of FIG. 1, the camera position and orientation determination unit 121 (a) The generated keypoint group K W (= ∪ IR K IR→WLow-trust image (target image) I using W Based on the matching result between the query image (target image) Iq and the low-trust image I in each of the image groups Gi, W Derive the position and orientation of the camera based on the low-trust image I (related to the query image Iq), and (b) High-trust image I R Based on the matching result between the query image (target image) Iq and the high-trust image I in each of the image groups Gi, R Derive the position and orientation of the camera based on the high-trust image I (related to the query image Iq), and (c) Determine the position and orientation of the camera related to the query image (target image) Iq from the positions and orientations of the multiple cameras derived in (a) and (b) above.

[0081] Specifically, in this embodiment, the matching unit 121a of the camera position and orientation determination unit 121 uses the keypoint group K W generated for the low-trust image I W to determine the descriptor group F (6) F W = D W (x, y, :) (x,y)∈KW , D W ∈ R HW×WW×NF of the low-trust image I W by the following formula. Here, D W (x, y, :) W is a dense descriptor corresponding to the keypoint group K (x,y)∈KW generated by, for example, S2DNet. Also, HW and WW are the number of pixels in the x-axis and y-axis directions in the low-trust image I W respectively, and NF is the number of descriptors in the low-trust image I W . W

[0082] ​Note that the above-mentioned dense descriptors and S2DNet are described in detail, for example, in non-patent literature: Hugo Germain, Guillaume Bourmaud, Vincent Lepetit, “S2DNet: Learning Accurate Correspondences for Sparse-to-Dense Feature Matching”, Proceedings of European Conference on Computer Vision (ECCV) 2020, pp.626-643, <https: / / doi.org / 10.1007 / 978-3-030-58580-8_37>, 2020, and non-patent literature: Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, and Marc Pollefeys, “Pixel-Perfect Structure-from-Motion”, International Conference on Computer Vision (ICCV) 2021, <https: / / doi.org / 10.48550 / arXiv.2108.08291>, 2021.

[0083] Next, the matching unit 121a uses the following equation (7) M W =Matching{(K q , F q ) Iq , (K W , F W ) W} to generate the matching result M W , that is, a pair of (indexes of) matched keypoints and descriptors between the query image Iq and the low-confidence image I W . Here, as Matching(·,·) in the above equation (7), various matching methods such as the nearest neighbor matching method and Lowe's Thresholding method can be adopted.

[0084] As a modification, the matching unit 121a determines descriptors corresponding to the individual keypoints included in the generated keypoint group K, for example, without relying on dense descriptors as in the present embodiment, and performs matching processing using these descriptors to generate a matching result M W is also possible. W Finally, the matching unit 121a uses the following formula

[0085] (8) MA =∪ Re =∪ G ∪ IR M Iq→IR (9) MA W =∪ G ∪ W M W to determine the matching results MA Gi for the high-confidence image groups Re Re (i = 1, 2,...) and the low-confidence image groups W Gi (i = 1, 2,...) in all the image groups Gi. Here, M W is the matching result between the high-confidence image I Iq→IR and the query image Iq as described above. Also, ∪ R is a union operator that forms a set by combining the matching results M IR for all the high-confidence images I Gi included in the high-confidence image group Re R (∈Re Gi ). Furthermore, ∪ Iq→IR is a union operator that forms a set by combining the matching results M W for all the low-confidence images I Gi included in the low-confidence image group W W (∈W Gi ). Moreover, ∪ W is a union operator that forms a set by combining the operands over all the image groups Gi (i = 1, 2,...). G

[0086] Incidentally, here, the matching unit 121a described above, together with the image group determination unit 111, the image reliability determination unit 112, the target 3D point generation unit 113, and the target keypoint generation unit 114, which are components of the keypoint generation program according to the present invention already described, can also constitute one embodiment of the "image matching program" according to the present invention. According to this image matching program of the present invention, it is also possible to perform more suitable image matching using the suitable keypoints generated for the target image (low-reliability image I W ).

[0087] Next, in this embodiment, the camera position and orientation determination unit 121 (a) Using the matching result MA (of the above formula (8)) for the high-reliability image group of all the determined image groups, derives the position and orientation of the camera related to the query image Iq based on the high-reliability images in each image group, Re and (b) Using the matching result MA (of the above formula (9)) for the low-reliability image group of all the determined image groups, derives the position and orientation of the camera related to the query image Iq based on the low-reliability images in each image group. W

[0088] Here, the derivation of the camera position and orientation using the matching results in (a) and (b) above can be performed using, for example, a known 8-point algorithm method, a normalized 8-point algorithm method, a cost minimization method, etc. Incidentally, the 8-point algorithm method is described in detail, for example, in the non-patent document: Chahat Deep Singh, ‘Structure from Motion’, [online], [searched on May 11, 2022], Internet <URL: https: / / cmsc426.github.io / sfm / #essential>.

[0089] The camera position and orientation determination unit 121 then uses the following formula (10) (α q , β q , θ​q , X q , Y q , Z q ) =Σ G Σ IRi (Wt IRi / N IR )×(α i , β i , θ i , X i , Y i , Z i )+ Σ G Σ Wi (Wt Wi / N W )×(α i , β i , θ i , X i , Y i , Z i ) Using this, for the camera position and orientation data (α R i , β i , θ i , X i , Y i , Z i ) based on the high - reliability image I i and the camera position and orientation data (α W i , β i , θ i , X i , Y i , Z i ) based on the low - reliability image I i , individual weights Wt IRi (Σ i Wt IRi = 1) and Wt Wi (Σ i Wt Wi = 1) are added, and a weighted average is performed to determine the camera position and orientation data (α q , β q , θ q , X q , Y q , Z q ) of the query image Iq.

[0090] Here, α, β, and θ are the orientations of the camera, which are the roll angle, pitch angle, and yaw angle, respectively. Also, X, Y, and Z are the X - coordinate value, Y - coordinate value, and Z - coordinate value of the position of the camera, respectively. Therefore, (α, β, θ, X, Y, Z) represents the 6 - degree - of - freedom (6 - DoF, six Degrees of Freedom) information of the camera. Furthermore, N IR and N W are the numbers of high - confidence images I R and low - confidence images IW in each image group Gi related to the matching with the query image Iq, respectively. Also, Σ IRi and Σ Wi are the sums for high - confidence images I R i and low - confidence images I W i respectively, and further, Σ G is the sum for the image groups Gi (i = 1, 2, ···).

[0091] As described above, in this embodiment, for target images (low - confidence images) with a small number of originally matching keypoints and inappropriate for deriving camera position and orientation data, suitable keypoints are generated based on high - confidence images (preferably, a large number of suitable keypoints are generated based on a large number of high - confidence images), and the camera position and orientation data related to the query image Iq is also derived using the target images (keypoint - updated low - confidence images) with the generated keypoints added. As a result, (since the information specific to low - confidence images included in the generated keypoints is also considered), a more accurate and (since more keypoints are used) more robust camera position and orientation determination process for the query image Iq can be implemented. Also, the image position of the query image Iq can be specified with higher accuracy (that is, a more suitable image position specification process can also be implemented).

[0092] Note that the weight Wt in the above formula (10) IRi and the weight Wt WiIt may be set in advance to a predetermined value based on experience, or may be appropriately set based on, for example, the confidence level RE. Also, the weight Wt IRi is for the camera position and orientation data related to the high-confidence image. Therefore, the weight Wt Wi is preferably set to a value larger than that.

[0093] FIG. 4 is a schematic diagram for explaining a specific example of the camera position and orientation determination process according to the present invention.

[0094] According to the specific example shown in FIG. 4, first, the camera position and orientation determination unit 121 (a) Camera position and orientation data based on the (KP update) low-confidence images (I G1 , I 1 ) included in the low-confidence image group W of the image group G1, 2 and (b) Camera position and orientation data based on the high-confidence images (I G1 , I 3 ) included in the high-confidence image group Re of the image group G1, 4 and (c) Camera position and orientation data based on the (KP update) low-confidence images (I G2 , I 5 , I 6 , I 7 ) included in the low-confidence image group W of the image group G2, (d) Camera position and orientation data based on the high-confidence images (I G2 , I 8 , I 9 , I 10 ) included in the high-confidence image group Re of the image group G2 are determined.

[0095] Next, the camera position and orientation determination unit 121 performs weighted averaging with the individual weights Wt IRi and Wt Wi (in this specific example, Wt IRi >Wt Wi ) added to the determined camera position and orientation data of the above (a) to (d), and calculates the camera position and orientation data of the final query image Iq.

[0096] Returning to the functional block diagram of FIG. 1, the camera position and orientation determination unit 121 transmits the camera position and orientation data (or image position identification result) of the query image (target image) Iq generated and determined as described above to an external information processing device, for example, the device that sent the request including the query image Iq, via the input / output interface unit 101, and it may be used there. Also, the keypoint information generated by the target keypoint generation unit 114 and the image matching result generated and determined by the matching unit 121a are also transmitted to an external information processing device via the input / output interface unit 101 and can be used there.

[0097] Furthermore, when this camera position and orientation determination device 1 is mounted on an autonomous vehicle, an autonomous mobile robot, etc., this device 1 (camera position and orientation determination unit 121) may output in real time the camera position and orientation data (or image position identification result) related to the camera image (target image) of each moment that photographed the surrounding driving environment to a driving control device also mounted on the autonomous vehicle, the autonomous mobile robot, etc. Thereby, the autonomous vehicle, the autonomous mobile robot, etc. can know their own current positions, and as a result, it is also possible to perform a suitable or safe driving operation.

[0098] As described in detail above, according to the present invention, the keypoints of the target image can be generated using the 3D points of the high-reliability image that is considered to be more easily matched with the target image (query image) compared to the target image. Here, these 3D points are generated based on the more reliable keypoints in the high-reliability image, and further, it is information related to the position on the 3D shape of the keypoints in the three-dimensional space (three-dimensional geometric). That is, compared with the two-dimensional points on the image, it is information that represents the features of the image content in more detail.

[0099] As a result, the keypoints of the target image generated from these 3D points can be suitably used for matching with the target image. Further, in generating such keypoints, a machine learning model that requires a huge amount of computation time, such as a transformer, is not necessary. That is, it is possible to generate keypoints of the target image that can be used for matching with the target image without relying on a machine learning model such as a transformer that requires a huge amount of computation time. Also, more suitable image matching and more suitable determination of the position and orientation of the camera can be performed using the keypoints generated in this way.

[0100] Furthermore, the image matching process and the image position identification process (camera position and orientation determination process) according to the present invention can be utilized for driving control of an autonomous vehicle or the like, and it is also possible to support a huge number of autonomous vehicles or the like to always drive safely, smoothly, and comfortably on the road network in the city. Also, the image matching process and the image position identification process (camera position and orientation determination process) according to the present invention can be utilized for analysis of a huge amount of image data collected by a large number of users taking various landmarks in the city and uploading their camera images, and a huge amount of image data regularly generated by a large number of street cameras, and it is also possible to promote prediction of the flow of people in the city, discovery of installations and suspicious objects and confirmation of their transitions, and further prediction and detection of troubles and crimes. That is, according to the present invention, it is also possible to contribute to Goal 11, "Make cities inclusive, safe, resilient and sustainable," of the Sustainable Development Goals (SDGs) led by the United Nations.

[0101] Furthermore, the image matching process and the image position identification process (camera position and orientation determination process) according to the present invention can be utilized for the analysis of a vast number of on-site image data collected by a large number of users taking pictures of the target area or sea area and uploading their camera images, and it is also possible to conduct investigations on various conditions in such areas and seas, for example, the growth status of crops, the current state of the ecosystem, and the impact of climate change. That is, according to the present invention, it is also possible to contribute to Goal 13, "Take urgent action to combat climate change and its impacts," Goal 14, "Conserve and sustainably use the oceans, seas and marine resources," and Goal 15, "Take urgent action to protect, restore and promote sustainable use of terrestrial ecosystems, sustainably manage forests, combat desertification, and halt and reverse land degradation and halt biodiversity loss" in the Sustainable Development Goals (SDGs) led by the United Nations.

[0102] Regarding the various embodiments of the present invention described above, various changes, modifications, and omissions within the scope of the technical idea and perspective of the present invention can be easily made by those skilled in the art. The explanations given above are merely examples and are not intended to impose any restrictions. The present invention is limited only by the claims and their equivalents.

Explanation of Reference Numerals

[0103] 1 Camera position and orientation determination device (keypoint generation device) 101 Input / output interface (IF) section 111 Image group determination section 112 Image reliability determination section 112a Keypoint (KP)-based reliability determination section 112b Position-based reliability determination section 113 Target 3D point generation section 113a High-reliability 3D point generation section 114 Target keypoint generation section 121 Camera position and orientation determination section 121a Matching section 2 Image database (DB)

Claims

1. A keypoint generation program for generating keypoints of a target image that can be used for matching between a target image included in a target image and an objective image, comprising: Image reliability determination means for determining a highly reliable image included in the image group, which is considered to be more easily matched with the objective image compared to the target image from the viewpoint of matching keypoints or camera information; Target 3D point generation means for converting the keypoints of the highly reliable image into 3D points in the camera coordinate system related to the highly reliable image using the camera information related to the highly reliable image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points; Target keypoint generation means for generating keypoints for the target image using the camera information related to the target image from the target 3D points; A keypoint generation program characterized by causing a computer to function.

2. The image reliability determination means determines a plurality of the highly reliable images; The target 3D point generation means generates target 3D points related to each of the plurality of highly reliable images using the keypoints in each of the plurality of highly reliable images that match the objective image; The target keypoint generation means generates a plurality of the keypoints for the target image from the target 3D points related to each of the plurality of highly reliable images. The keypoint generation program according to claim 1, characterized in that.

3. The image group is one of a plurality of images generated by classifying images included in a certain image group based on the position and orientation of the camera related to the image. The keypoint generation program according to claim 1 or 2, characterized in that.

4. The image reliability determination means determines the highly reliable image based on the ratio or number of keypoints that match the objective image in the images included in the image group, or based on the distance between the position of the camera related to the images included in the image group and the representative position of the camera related to the image group. The keypoint generation program according to claim 3, characterized in that.

5. The image reliability determination means determines, when determining the high-reliability image, whether it is based on the ratio or number of the matching keypoints or based on the distance, according to the relative relationship between the position of the camera related to the image included in the image group and the representative position of the camera related to the image group. The keypoint generation program according to claim 4, characterized in that.

6. An image matching program for performing matching between a target image and a target image included in an image group, The image group is one of a plurality of image groups generated by classifying the images included in a certain image group based on the position and orientation of the camera related to the image, For each image group, there is an image reliability determination means for determining a high-reliability image that is included in the image group and is considered to be more easily matched with the target image than the target image from the perspective of matching keypoints or camera information, For each image group, the keypoints of the high-reliability image are converted into 3D points in the camera coordinate system related to the high-reliability image using the camera information related to the high-reliability image, and a target 3D point in the camera coordinate system related to the target image is generated from the 3D points. Target 3D point generation means, For each image group, target keypoint generation means for generating keypoints for the target image using the camera information related to the target image from the target 3D points, Matching means for generating a dense descriptor for the target image using the generated keypoints of the target image and performing matching between the target image and the target image using the dense descriptor An image matching program characterized by causing a computer to function.

7. A camera position and orientation determination program for determining the position and orientation of the camera related to the target image using the target image included in the image group, The image group is one of a plurality of image groups generated by classifying the images included in a certain image group based on the position and orientation of the camera related to the image, For each image group, there is an image reliability determination means for determining a high-reliability image that is included in the image group and is considered to be more easily matched with the target image than the target image from the perspective of matching keypoints or camera information, For each of the image groups, the key points of the highly reliable image are converted into 3D points in the camera coordinate system related to the highly reliable image using the camera information related to the highly reliable image, and target 3D point generation means for generating target 3D points in the camera coordinate system related to the target image from the 3D points. For each of the image groups, target key point generation means for generating key points for the target image using the camera information related to the target image from the target 3D points. Based on the matching results between the target image and the target image using the generated key points, the position and orientation of the camera based on the target image in each of the image groups are derived, and based on the matching results between the highly reliable image and the target image, the position and orientation of the camera based on the highly reliable image in each of the image groups are derived, and from the derived position and orientation of the plurality of cameras, camera position and orientation determination means for determining the position and orientation of the camera related to the target image. A camera position and orientation determination program characterized by causing a computer to function.

8. The camera position and orientation determination means generates a dense descriptor for the target image using the generated key points of the target image, performs matching between the target image and the target image using the dense descriptor, and generates a matching result between the target image and the target image. The camera position and orientation determination program according to claim 7.

9. The camera position and orientation determination means performs a weighted average with individual weights added to the position and orientation of the camera derived using the matching result between the highly reliable image and the target image and the position and orientation of the camera derived using the matching result between the target image and the target image, and determines the result as the position and orientation of the camera related to the target image. The camera position and orientation determination program according to claim 7 or 8.

10. A key point generation device for generating key points of the target image that can be used for matching between the target image and the target image included in the image group. An image reliability determination means for determining a highly reliable image that is included in the image group and is considered to be more easily matched with the target image than the target image from the viewpoint of matching key points or camera information. Target 3D point generation means for converting the keypoints of the high-reliability image into 3D points in the camera coordinate system related to the high-reliability image using the camera information related to the high-reliability image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points. Target keypoint generation means for generating keypoints for the target image using the camera information related to the target image from the target 3D points A keypoint generation device characterized by comprising the above.

11. A keypoint generation method for generating keypoints of a target image that can be used for matching between a target image included in a target image and an image group, Determining a high-reliability image that is a high-reliability image included in the image group and is considered to be more easily matched with the target image than with the target image from the perspective of matching keypoints or camera information. Converting the keypoints of the high-reliability image into 3D points in the camera coordinate system related to the high-reliability image using the camera information related to the high-reliability image, and generating target 3D points in the camera coordinate system related to the target image from the 3D points. Generating keypoints for the target image using the camera information related to the target image from the target 3D points A keypoint generation method implemented by a computer, characterized by comprising the above.

Citation Information

Patent Citations

  • Scene matching reference data generation system and position measurement system

    JP2011215057A

  • Image processing device and database construction device therefor

    JP2015032256A

  • Unified framework for precise vision-aided navigation

    US20080167814A1

  • Methods and Apparatus for Visual Search

    US20130121600A1

  • Visual localisation

    WO2014044852A2