Three-dimensional space information configuration system and three-dimensional space information configuration method

WO2026181395A1PCT designated stage Publication Date: 2026-09-03HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/036796
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2025-10-20
Publication Date
2026-09-03

Smart Images

  • Figure JP2025036796_03092026_PF_FP_ABST
    Figure JP2025036796_03092026_PF_FP_ABST
Patent Text Reader

Abstract

This three-dimensional space information configuration system for configuring three-dimensional space information is composed of a computer having an arithmetic device for executing predetermined processing and a storage device accessible by the arithmetic device. The three-dimensional space information configuration system comprises: an inter-image feature point association detection unit for detecting a plurality of feature points associated between images from pair image candidates that are pairs of images captured of the same subject; a feature point coordinate pair output unit for evaluating the associated feature points, selecting feature points on the basis of the results of evaluating the feature points, and outputting extracted feature point association information that is a set of the selected feature points; and a configuration unit for configuring three-dimensional space information on the basis of the positional relationship of the extracted feature point association information.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional spatial information construction system and three-dimensional spatial information construction method Incorporation by Reference

[0001] This application claims the priority benefit of Japanese Patent Application No. 2025-28929 filed on February 26, Reiwa 7 (2025), the content of which is incorporated into this application by reference.

[0002] The present invention relates to a three-dimensional spatial information construction system.

[0003] In three-dimensional reconstruction via photogrammetry using RGB images captured by cameras from multiple viewpoints, simultaneous execution of self-localization and environmental map construction via SLAM (Simultaneous Localization and Mapping), photorealistic three-dimensional reconstruction via Gaussian Splatting, and other related processes, high-precision matching processing for common feature points of each image within realistic calculation time and calculation resources is required to implement these processes with high accuracy.

[0004] As background art in the present technical field, there is the following prior art. Patent Document 1 (Japanese Unexamined Patent Application Publication No. 2013-242812) discloses an information processing device that detects the position of a measurement target in a three-dimensional space using captured images, the device comprising: an image acquisition unit that acquires stereo image data obtained by capturing images of the measurement target in parallel with a first camera and a second camera installed at a predetermined distance from each other; a detection surface definition unit that defines a detection surface in the three-dimensional space and defines a detection region obtained by projecting the detection surface onto the image captured by the first camera among the stereo images; a parallax-corrected region derivation unit that derives a parallax-corrected region obtained by shifting a region identical to the detection region in the image captured by the second camera among the stereo images by an amount of parallax determined by the position of the detection surface in the depth direction in a direction that eliminates parallax; a matching unit that performs matching processing between the image of the detection region in the image captured by the first camera and the image of the parallax-corrected region in the image captured by the second camera; and a detection result output unit that outputs a matching result obtained by the matching unit.

[0005] Patent Document 2 (Japanese Patent Publication No. 2013-65247) describes an apparatus for performing stereo matching processing on two images obtained by photographing the same object from different angles, comprising: a first disparity calculation unit that identifies a plurality of pairs of points corresponding to each other in the two images and calculates the disparity for each identified pair; an image division unit that divides the two images so that the divided parts of each image correspond to each other; and for each of the two corresponding divided parts, it identifies the pairs of points present in that part and uses the value of the disparity calculated for the identified pair to perform stereo matching processing. A stereo matching processing device is disclosed, characterized by comprising: a parallax search range setting unit that identifies the maximum and minimum values ​​of the parallax in the relevant portion and sets a parallax search range; and a second parallax calculation unit that identifies a plurality of pairs of points corresponding to each other between the two corresponding divided portions, calculates the parallax for each newly identified pair, and outputs the calculated parallax as the parallax of the newly identified pair in that portion if the calculated value falls within the parallax search range set for that portion.

[0006] Non-patent document 1 discloses a technique for performing sparse feature point matching that is less affected by image rotation and magnification changes.

[0007] Non-patent document 2 discloses a technique for robustly performing dense feature point matching on image pairs captured in diverse environments.

[0008] JP 2013-242812 JP 2013-65247 Lowe, DG Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision 60, 91-110 (2004). https: / / doi.org / 10.1023 / B:VISI.0000029664.99615.94Xuelun Shen, Zhipeng Cai, Wei Yin, Matthias Muller, Zijun Li, Kaixuan Wang, Xiaozhi Chen, and Cheng Wang, "GIM: Learning Generalizable Image Matcher From Internet Videos," The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024.

[0009] In the aforementioned Non-Patent Document 1, the feature point matching process is sparse, resulting in a small number of detected feature points. This presents a problem in that feature points cannot be matched when there are significant changes in conditions or viewpoints.

[0010] The aforementioned Non-Patent Document 2 has a problem in that the feature point matching process requires computation time and resources. Furthermore, it has a problem in that incorrect feature point matching results are output for image pairs that do not truly correspond, based on partial matches or similar parts within the images.

[0011] While Patent Documents 1 and 2 disclose techniques for performing feature point matching at high speed when given pairs of images, if a group of three or more images is input, feature point matching is performed on all image pairs, resulting in a problem where the computation time increases with the square of the number of images.

[0012] This invention was made in view of these circumstances, and aims to achieve high-precision three-dimensional reconstruction from images.

[0013] A typical example of the invention disclosed in this application is as follows: A three-dimensional spatial information configuration system comprising a computer having a computing device that performs predetermined processing and a storage device accessible by the computing device, comprising: an image-to-image feature point correspondence detection unit that detects a plurality of corresponding feature points from pair image candidates which are pairs of images of the same subject; a feature point coordinate pair output unit that evaluates the corresponding feature points, selects feature points based on the evaluation results of the feature points and outputs extracted feature point correspondence information which is a set of the selected feature points; and a configuration unit that configures three-dimensional spatial information based on the positional relationship of the extracted feature point correspondence information.

[0014] According to one aspect of the present invention, high-precision three-dimensional reconstruction from an image can be achieved. Problems, configurations, and effects other than those described above will be clarified by the following description of the embodiments.

[0015] This is a block diagram showing an example of the configuration of the three-dimensional spatial information configuration system of the first embodiment. This is a diagram showing an example of the physical configuration of the three-dimensional spatial information configuration system of the first embodiment. This is a sequence diagram showing an example of the processing of the three-dimensional spatial information configuration system of the first embodiment. This is a sequence diagram showing an example of the processing of the three-dimensional spatial information configuration system of the first embodiment. This is a diagram illustrating an example of matching results using the prior art. This is a diagram illustrating an example of matching results using the prior art. This is a diagram illustrating an example of matching results using the prior art. This is a diagram illustrating an example of matching results using the prior art. This is a diagram showing an example of selecting a pair image candidate in the pair image candidate output unit of the first embodiment. This is a diagram showing an example of extracting feature point correspondences in the feature point coordinate pair output unit of the first embodiment. This is a diagram showing an example of extracting feature point correspondences in the pair image candidate output unit of the first embodiment. This is a diagram showing an example of extracting feature point correspondences in the pair image candidate output unit of the first embodiment. This is a diagram showing an example of calculating the evaluation value of feature point correspondences in the first embodiment. This is a diagram showing an example of calculating the evaluation value of feature point correspondences in the first embodiment. This is a diagram showing an example of calculating the evaluation value of feature point correspondences in the first embodiment. This is a diagram showing an example of the method for calculating evaluation values ​​corresponding to feature points in the first embodiment. This is a diagram showing an example of the method for calculating evaluation values ​​corresponding to feature points in the first embodiment. This is a diagram showing an example of the method for calculating evaluation values ​​corresponding to feature points in the first embodiment. This is a diagram showing the effects of the first embodiment. This is a diagram showing the effects of the first embodiment. This is a block diagram showing an example of the configuration of the three-dimensional spatial information configuration system of the second embodiment. This is a block diagram showing an example of the configuration of the three-dimensional spatial information configuration system of the third embodiment. This is a diagram showing an example of the user interface of the third embodiment. This is a block diagram showing an example of the configuration of the three-dimensional spatial information configuration system of the fourth embodiment. This is a diagram showing an example of the user interface of the fourth embodiment. This is a block diagram showing an example of the configuration of the three-dimensional spatial information configuration system of the fifth embodiment. This is a block diagram showing an example of the configuration of the three-dimensional spatial information configuration system of the fifth embodiment.

[0016] Hereinafter, several embodiments of the present invention will be described with reference to the drawings. In all drawings used to describe each embodiment, the same reference numerals will be used for the same components as a general rule, and repeated descriptions will be omitted. In addition, in the following embodiments, the components (including element steps, etc.) are not necessarily essential unless specifically stated or considered to be clearly essential in principle. Also, when we say "consisting of A," "made of A," "having A," or "including A," we do not exclude other elements unless specifically stated that only that element is included. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of the components, etc., it includes those that are substantially similar or approximate to that shape, etc., unless specifically stated or considered to be clearly not the case in principle.

[0017] [First Embodiment] Figure 1A is a block diagram showing an example of the configuration of a three-dimensional spatial information system 1 according to the first embodiment of the present invention.

[0018] The three-dimensional spatial information configuration system 1 of this embodiment takes three or more image data 2 with different viewpoints as input, calculates the correspondence between feature points between the images with low computational cost, low computational resources, and high accuracy, and uses the correspondence between these feature points to construct three-dimensional spatial information of the subject using techniques such as photogrammetry, Structure from Motion (SfM), Multi-View Stereo (MVS), Neural Radiance Fields (NeRF), and 3D Gaussian Splatting (3DGS). Here, unlike a stereo camera which basically constructs three-dimensional spatial information from two images whose relative shooting positions are known, the relative shooting positions are unknown, and this system assumes the construction of multi-angle three-dimensional spatial information of a subject from three or more images.

[0019] Each component of the three-dimensional spatial information configuration system 1 may be realized by a circuit, or at least a part of it may be realized by a processor such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit) and memory that executes a program.

[0020] The three-dimensional spatial information configuration system 1 includes an input I / F 110, a pair image candidate output unit 12, an image feature point correspondence detection unit 13, a feature point coordinate pair output unit 14, a configuration unit 15, and an output I / F 111.

[0021] The image data 2 input to the three-dimensional spatial information system 1 may be images taken with a general camera, an all-around camera, a smartphone, or images collected via a network. The image data 2 is input to the input I / F 110 of the three-dimensional spatial information system 1.

[0022] The input interface 110 receives the input image data 2, applies various image processing techniques, and outputs it to the paired image candidate output unit 12. For example, the input interface 110 performs image processing such as white balance adjustment, gamma correction, and contrast correction.

[0023] The paired image candidate output unit 12 outputs paired image candidates 31 for three or more image data 2 output from the input I / F 110, which are presumed to contain the same subject 20 in the images. The number of paired image candidates 31 is less than or equal to the total number of combinations of image data 2.

[0024] The image-to-image feature point correspondence detection unit 13 detects the correspondence between feature points in the pair image candidates 31 output from the pair image candidate output unit 12. Since the number of pair image candidates 31 is the same as or less than the total number of combinations of image data 2, the number of times feature point correspondence detection is performed is less than or equal to the number of times it is performed for all combinations. The detected feature point correspondence information is output to the feature point coordinate pair output unit 14. The image-to-image feature point correspondence detection unit 13 may also be input to pre-created pair image candidates 31.

[0025] The feature point coordinate pair output unit 14 calculates an evaluation value for each feature point correspondence relationship based on the feature point correspondence information output from the inter-image feature point correspondence detection unit 13, selects extracted feature point correspondence information based on the evaluation value, and outputs the selected extracted feature point correspondence information to the configuration unit 15.

[0026] The component 15 constructs three-dimensional spatial information using the principle of triangulation from the extracted feature point correspondence information output from the feature point coordinate pair output unit 14.

[0027] The output I / F 111 outputs the three-dimensional spatial information output from the component 15 to the subsequent processing unit. The output three-dimensional spatial information is then processed, saved to a file, or output for other processing.

[0028] Figure 1B shows an example of the physical configuration of the three-dimensional spatial information configuration system 1 according to the first embodiment of the present invention.

[0029] The three-dimensional spatial information configuration system 1 of this embodiment is composed of a computer having a processor (CPU) 501, memory 502, auxiliary storage device 503, and communication interface 504. The three-dimensional spatial information configuration system 1 may also have an input unit 505 and an output unit 508.

[0030] The processor 501 is an arithmetic unit that executes programs stored in the memory 502. The functions of each functional unit of the three-dimensional spatial information configuration system 1 are realized by the execution of various programs by the processor 501. Note that some of the processing performed by the processor 501 when executing programs may be performed by other arithmetic units (for example, hardware such as ASICs or FPGAs).

[0031] Memory 502 is a storage device that includes ROM, a non-volatile memory element, and RAM, a volatile memory element. ROM stores immutable programs (e.g., BIOS). RAM is a high-speed, volatile memory element such as DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor 501 and data used during program execution.

[0032] The auxiliary storage device 503 is, for example, a high-capacity, non-volatile storage device such as a flash memory (SSD) or a magnetic storage device (HDD). The auxiliary storage device 503 also stores data used by the processor 501 when executing a program, and the program that the processor 501 executes. In other words, the program is read from the auxiliary storage device 503, loaded into RAM, and executed by the processor 501, thereby realizing each function of the three-dimensional spatial information configuration system 1.

[0033] The communication interface 504 is a network interface device that controls communication with other devices according to a predetermined protocol.

[0034] The input unit 505 is an interface to which input devices such as a keyboard 506 and a mouse 507 are connected for receiving input from the operator. The output unit 508 is an interface to which output devices such as a display device are connected for outputting the program execution results in a format that can be viewed by the user.

[0035] The program executed by the processor 501 is provided to the three-dimensional spatial information configuration system 1 via removable media (such as DVD-R or flash memory) or a network, and is stored in the non-volatile auxiliary storage device 503, which is a non-temporary storage medium. For this reason, the three-dimensional spatial information configuration system 1 is preferably provided with an interface for reading data from the removable media.

[0036] The three-dimensional spatial information configuration system 1 may be composed of a computer system consisting of a single physical computer, or it may be composed of virtual computers built on multiple physical computer resources. For example, the multiple programs that realize the functions of the three-dimensional spatial information configuration system 1 may each run on separate physical or logical computers, or multiple programs may be combined and run on a single physical or logical computer. In this case, user terminals connected to the three-dimensional spatial information configuration system 1 via a network may provide input and output devices.

[0037] Figures 2A and 2B are sequence diagrams showing an example of processing of the three-dimensional spatial information configuration system 1 according to the first embodiment of the present invention.

[0038] The three-dimensional spatial information configuration system 1 of this embodiment consists of an imaging unit 17 and a server.

[0039] As shown in Figure 2A, the image data 2 input to the three-dimensional spatial information configuration system 1 of this embodiment is, for example, taken by a user using the imaging unit 17. The user may specify shooting parameters such as the exposure time during shooting, in which case the user inputs them to the imaging unit 17. The imaging unit 17 executes the shooting process of the subject according to the specified shooting parameters and acquires shooting data such as the image data 2 of the subject. The captured shooting data is transmitted to the server that constitutes the three-dimensional spatial information along with the shooting parameters. In the case of shooting data acquired from the internet, the shooting data and shooting parameters may be transmitted directly from the user to the server without going through the imaging unit 17. On the server, the paired image candidate output unit 12 executes paired image candidate output processing based on the transmitted shooting data and shooting parameters, the inter-image feature point correspondence detection unit 13 executes inter-image feature point correspondence detection processing, the feature point coordinate pair output unit 14 executes feature point coordinate pair output processing, and the configuration unit 15 executes three-dimensional spatial information configuration processing.

[0040] Furthermore, as shown in Figure 2B, a portion of the pair image candidate output processing in the pair image candidate output unit 12 may be performed in the imaging unit 17. For example, the imaging unit 17 may have a pair image candidate output unit 12, and based on the transmitted shooting data and shooting parameters, it may perform pair image candidate output processing and send the pair image candidates to the server.

[0041] Next, problems of the prior art will be described. In the prior art, the paired image candidate output unit 12 and the feature point coordinate pair output unit 14 shown in FIG. 1 are not provided, and an inter-image feature point correspondence detection unit 13 calculates feature point correspondence information from image data 2 output from an input I / F 110. FIG. 3 shows an example of feature point correspondence information calculated by the inter-image feature point correspondence detection unit 13 of the prior art. A subject 20 is captured in two pieces of image data 2 for which feature point correspondences are to be detected, and a plurality of feature points 21 of the subject 20 are detected. Correspondences between feature points detected between the two images are indicated by solid lines as feature point correspondences 22. On the other hand, incorrectly corresponded feature points are indicated by broken lines as feature point mismatches 23.

[0042] Also, in FIG. 4, for all input image data 2, images for which a predetermined number or more of feature point correspondences 22 have been detected are indicated by solid lines as image correspondences 24, and incorrectly corresponded images are indicated by broken lines as image mismatches 25.

[0043] For example, as a prior art, there is a feature point detection and matching process called SIFT described in Non-Patent Document 1. SIFT is a method of calculating a DoG image D(x, y, σ) represented by equations (1), (2), and (3), and detecting feature points from changes with respect to a scale σ. In equations (1), (2), and (3), x and y are pixel coordinates, and σ is a scale.

[0044]

[0045] SIFT is characterized by being less susceptible to the effects of rotation and scaling of a subject, and is a technique widely used in applications such as three-dimensional reconstruction. On the other hand, SIFT has a small number of detected feature points and is easily affected by changes in shooting direction and shooting conditions. This causes problems such as a small number of correct image correspondences that can be detected. FIGS. 3 and 4 schematically illustrate these problems. When the number of feature points is small, for example, the number of correct image correspondences 24 decreases, the number of incorrect image correspondences 25 increases, and problems arise such as a decrease in the density of three-dimensional reconstruction constructed based on the detected feature points. When the number of image correspondences decreases, the number of images used for constructing three-dimensional spatial information decreases, resulting in omissions in the spatial information. Furthermore, when the number of incorrect image correspondences increases, for example, duplication occurs in the three-dimensional spatial information, which reduces the accuracy of the constructed three-dimensional space. Duplication occurring in three-dimensional spatial information refers to a state where the same subject erroneously exists at a plurality of positions in the three-dimensional space.

[0046] In order to solve these problems, for example, as disclosed in Non-Patent Document 2, there is a feature point detection and matching technique using machine learning. Here, the feature point detection and matching processes may be performed individually in sequence or may be performed simultaneously. By using machine learning, as schematically shown in FIG. 5, the number of detected corresponding feature points increases, making the technique less susceptible to the influence of fluctuations in shooting direction and shooting conditions. As a result, as shown in FIG. 6, the number of detected image correspondences increases.

[0047] Furthermore, the technology described in Non-Patent Document 2, for example, has the problem of requiring time and resources for computation, and increasing the number of image mismatches 25 based only on partial matches within the image. The increase in computation time and resources leads to a decrease in applicable use cases due to reduced usability, and the increase in image mismatches 25 leads to a decrease in the accuracy of the constructed three-dimensional spatial information (e.g., duplication). Also, if feature point correspondence detection between images is performed for all assumed combinations of images for the input image data 2, a computation time of O(N^2) is required. Here, N is the number of image data 2, and if the number of images is increased to make the three-dimensional spatial information more detailed or wider, the computation time increases by the square of the number of images. This effect is particularly large in the case of feature point correspondence detection methods that have relatively long computation times. For example, if it takes 1 second to detect feature point correspondence for one pair of images, and there are 2,000 input images, the computation time for all combinations will be longer than 23 days.

[0048] To summarize the above, if, for example, machine learning-based feature point correspondence detection between images, as disclosed in Non-Patent Document 2, is implemented to improve the accuracy of the constructed three-dimensional spatial information, the number of feature point correspondences and image correspondences will increase, but problems will arise such as increased computation time and resources, and an increase in image mismatches.

[0049] Therefore, in this invention, while utilizing methods that have a large number of detectable feature point correspondences and image correspondences, such as the inter-image feature point correspondence detection using machine learning described in Non-Patent Document 2, the time and resources required for computation are suppressed by pre-selecting candidate pair images 31 for inter-image feature point correspondence detection, and by extracting feature point correspondences with high reliability from the detected feature point correspondences, image mismatches are suppressed.

[0050] Returning to Figure 1, the three-dimensional spatial information configuration system 1 of this embodiment will be described. The pair image candidate output unit 12 selects pair image candidates from the image data 2 input from the input I / F 110 using a selection method described later, and outputs the selected pair image candidates to the inter-image feature point correspondence detection unit 13. For example, the pair image candidate output unit 12 sequentially executes the process of detecting feature points from the image input as image data 2, and the process of associating the detected feature points with each other in different images. By selecting an appropriate algorithm and configuring the pair image candidate output unit 12, the computation time required for the pair image candidate output unit 12 to output pair image candidates can be made shorter than the computation time required for the inter-image feature point correspondence detection unit 13 to detect inter-image feature point correspondences. The inter-image feature point correspondence detection unit 13 detects inter-image feature point correspondences for the input pair image candidates, but since the number of these pair image candidates is the same as or less than the total number of combinations of image data 2, the computation time and computational resources required for inter-image feature point correspondence detection can be suppressed. Furthermore, the feature point coordinate pair output unit 14 extracts highly reliable feature point correspondences from the detected feature points using a selection method described later. This suppresses mismatches between feature points and images, and the component unit 15 can construct highly accurate three-dimensional spatial information using the extracted feature point correspondences.

[0051] Figure 7 shows an example of the selection of pair image candidates in the pair image candidate output unit 12. For example, when 10 image data 2 are input, all possible combinations of two images are shown in the white or gray section in the lower left of the table. Here, the gray section indicates the selected pair image candidate 31. Therefore, for example, in the example shown in Figure 7, image 1 and image 2 are not extracted as pair image candidate 31, while image 1 and image 3 are extracted as pair image candidate 31. In this way, the same number as or fewer than the total number of combinations of pair image candidates 31 are selected. Here, by setting a certain upper limit on the number of times each image in the image data 2 can be selected as a candidate, the number of pair image candidate searches increases linearly with the number of input images, and the computation time in the inter-image feature point correspondence detection unit 13 can be suppressed to a function O(N) that increases linearly with the number N of image data 2.

[0052] Figure 8 shows an example of feature point correspondence extraction by the feature point coordinate pair output unit 14. The feature point coordinate pair output unit 14 calculates an evaluation value for feature point correspondence based on the results of feature point correspondence detection performed by the inter-image feature point correspondence detection unit 13 for the pair image candidate 31 shown in gray, and extracts feature points with high reliability. In Figure 8, for example, a pair image candidate 32 selected for containing a predetermined number or more of the extracted feature point correspondences is shown with a circle. In this way, the feature point coordinate pair output unit 14 suppresses feature point mismatches and image mismatches by extracting feature point correspondences based on evaluation values ​​from the feature point correspondences detected by the inter-image feature point correspondence detection unit 13.

[0053] From here, we will explain the specific processing details of the pair image candidate output unit 12 and the feature point coordinate pair output unit 14 with examples.

[0054] One example of a process executed by the pair image candidate output unit 12 is to detect feature point correspondences between images for all combinations of image data 2 using a method that requires less computation time and resources, and then select a pair image candidate. An example of a processing method that requires less computation time and resources is SIFT and its matching process, as described in Non-Patent Literature 1. One possible method is to perform SIFT and its matching process for all combinations of image data 2, and then, for example, select pair image candidates in order of the number of feature point correspondences detected for each image. In this case, a threshold may be set based on the number of feature point correspondences to select the output pair image candidates. Processing using SIFT requires less computation and can determine the similarity of features between feature points in all image pairs, even if there are many images.

[0055] Figure 9 shows an example of another process performed by the pair image candidate output unit 12, specifically an example of performing graph analysis. Graph analysis is performed on the pair image candidate 31 at a given point in time for the image data 2, and images considered to be incorrect image correspondences are removed from the pair image candidate 31 as image miscorrespondences 25. In the example shown in Figure 9, the correct image correspondences 24 for pair image candidates are shown by solid lines, and the incorrect image miscorrespondences 25 for pair image candidates are shown by dashed lines. For example, graph analysis can be performed using edge between centrality. Edge between centrality is the sum of the proportions of paths that pass through a given edge among all the shortest paths between node pairs, and is expressed by equation (4).

[0056]

[0057] Here, c_B is edge betweenness centrality, e is the edge of interest, V is the set of all edges, σ(s,t) is the shortest path connecting edge s and edge t, and σ(s,t|e) is the shortest path connecting edge s and edge t that passes through edge e. When a network short circuit occurs due to an incorrect pair of image candidates that should not be paired, the edge betweenness centrality of the candidate causing the network short circuit is thought to increase. Therefore, incorrect pair of image candidates can be removed by removing pair of image candidates with high edge betweenness centrality. In addition, incorrect pair of image candidates can be removed by various features such as degree, clustering coefficient, and page rank, or by combinations of these features. For example, pair of image candidates whose degree, clustering coefficient, page rank, or composite features combining them do not fall within a predetermined range can be removed. Furthermore, incorrect pair of image candidates can be removed by anomaly detection using these features, evaluating the image correspondence 24 as an edge, or evaluating the image data 2 as a node. Here, as an example of pair image candidate information held internally by the pair image candidate output unit 12, which is data to be input to graph analysis, the results of inter-image feature point correspondence detection for all combinations using the aforementioned method with low computation time and computational resources may be used. Alternatively, the evaluation value of image similarity calculated for all combinations may be used. Furthermore, the edges of the network may be weighted based on the evaluation value calculated from the inter-image feature point correspondence detection results or image similarity, and used for graph analysis.

[0058] By quantifying the relationships between edges and determining them against a predetermined threshold, highly reliable pairs of images can be extracted.

[0059] Figures 10A and 10B show another example of the processing performed by the pair image candidate output unit 12, specifically when union find is performed. Union find can be applied to the pair image candidate 31 of the image data 2 at a certain point in time to add a pair image that is likely to be correct to the pair image candidate 31. Specifically, considering each image as an element, union find is used to group images together when they become pair image candidate 31. When an image belongs to a certain group, the images within that group are likely to be a group of images taken in close proximity to each other and depicting the same subject, so a pair consisting of that image and other images in the group can be added to the pair image candidate 31. Also, if pairs of images within each group have already been registered as pair image candidate 31, pair image candidate 31 may be added to connect the groups in order to efficiently search for new image pairs. In the example shown in Figure 10A, the group {1, 2, 3, 4, 5} is formed, and in the example shown in Figure 10B, the group {6, 7, 8, 9} is formed. In this example, since images 2 and 4 belong to the same group, they are likely to have been taken in the same vicinity, and it is conceivable to add (2,4) to the candidate pair image 31. Another example of efficiently searching for new image pairs is to add, for example, (5,6) to the candidate pair image 31 as a pair that connects groups.

[0060] Another example of the processing performed by the pair image candidate output unit 12 is that, when a video is input as image data 2, the unit may process the input video to select pair image candidates 31 from frames that are close in time (for example, consecutive frames) taken at similar times.

[0061] Another example of the processing performed by the pair image candidate output unit 12 is that it may refer to location information associated with the image data 2 by the user at the time of shooting and prioritize images with similar locations as pair image candidates 31. The location information associated with the image data 2 may be measured by an inertial measurement unit or a global positioning system.

[0062] Next, an example of the processing performed by the feature point coordinate pair output unit 14 will be described. The feature point coordinate pair output unit 14 calculates an evaluation value for each feature point correspondence relationship for the feature point correspondence information output from the inter-image feature point correspondence detection unit 13, and selects extracted feature point correspondence information based on the calculated evaluation value. For example, the evaluation value of the feature point correspondence relationship may be calculated based on the similarity of the feature quantities of the corresponding feature points (e.g., pixel value, feature point position).

[0063] As an example of the processing of the feature point coordinate pair output unit 14, based on the feature point correspondence information output from the inter-image feature point correspondence detection unit 13, there is a process that outputs all feature point correspondences 22 as extracted feature point correspondence information only for image pairs with a large number of feature point correspondences 22 (see Figure 5). This corresponds to a process in which, for example, the evaluation value of each feature point correspondence relationship is set to 1, and when the sum of the evaluation values ​​of all feature point correspondence relationships in the image pair exceeds a certain threshold, all feature point correspondence information is output as extracted feature point correspondence information. For example, as shown in Figure 11, if a part of the background subject 20B is obscured by a foreground subject 20F that is only visible in one of the image pairs, the number of feature point correspondences 22 will be less than when there is no foreground subject 20F, which may reduce the reliability of the image correspondence or decrease the accuracy of estimating the relative position of the image capture in subsequent processing. By outputting only image pairs in which the number of feature point correspondences 22 is greater than or equal to a predetermined threshold through this processing of the feature point coordinate pair output unit 14, image pairs that cause these accuracy degradations can be eliminated.

[0064] Figure 12 shows another example of the processing of the feature point coordinate pair output unit 14. In the example shown in Figure 12, the feature point coordinate pair output unit 14 extracts corresponding feature points based on the distribution of feature point correspondences 22 on the image. The distribution of feature point correspondences includes, for example, the distribution area, concentration, and shape of feature points in the entire image. For example, an image pair with a large feature point distribution area indicates a global correspondence between the images. Regarding the distribution area, in the example shown in Figure 11, the number of feature point correspondences 22 is small due to the influence of the foreground subject 20F, and the distribution area of ​​feature point correspondences is small. In this way, by outputting only image pairs where the feature point distribution area is above a predetermined threshold, it is possible to eliminate image pairs that correspond only to local feature points in the image, which can be a factor in reducing accuracy. Furthermore, regarding the concentration, as shown in Figure 12, when feature points with a high concentration correspond, the feature point correspondence information is calculated using only some of the features of the subject 20, which may result in an incorrect feature point correspondence or image correspondence. By outputting only feature point pairs where the concentration of feature points is below a predetermined threshold, it is possible to eliminate potentially incorrect feature point correspondences. Furthermore, regarding the shape, by detecting the autocorrelation function of the distribution shape and peaks in the frequency domain, it is possible to determine from the feature point distribution whether a periodic pattern exists in the subject 20. When a periodic pattern, which is common in artificial objects, is detected as a feature point, there is a high possibility that there are many similar subjects 20, and the detected feature point correspondence information may be incorrect feature point correspondence or incorrect image correspondence. For example, by removing feature point pairs whose autocorrelation coefficient of the feature point distribution shape takes a value above a predetermined threshold, feature point pairs that are affected by the periodic pattern of the subject 20 and become a factor in reducing accuracy can be eliminated. Also, as shown in Figure 13, when feature point correspondence is detected for different subjects 20 using only some of their features, the detected feature point distribution will be U-shaped with the center and bottom of the subject missing, and for example, the convexity of the distribution shape will be small. As in this example, the distribution shape can be used to evaluate the absence of feature point correspondences that should be detected. For example, by outputting only image pairs in which the convexity of the feature point distribution shape is above a predetermined threshold, image pairs that are a factor in reducing accuracy because they lack feature point correspondences that should be detected from the image can be eliminated. In this way, analysis using feature point distribution allows us to evaluate the correspondence of each feature point and select appropriate extracted feature point correspondence information.Furthermore, by assuming camera internal parameters, it is possible to select appropriate extracted feature point correspondence information by considering the three-dimensional shape information obtainable from the distribution of feature point correspondences. For example, depending on the subject, the evaluation value of extracted feature point correspondence information in which only planes are selected as feature points may be lowered.

[0065] Figure 14 shows another example of a method for calculating evaluation values ​​for feature point correspondence. In the example shown in Figure 14, the feature point coordinate pair output unit 14 extracts the corresponding feature point using the pixel information on the image at the location of the feature point 21. For example, in Figure 14, the front of the subject 20 is bright with high brightness, but the top and sides are dark with low brightness. Generally, pixels with low brightness are relatively more affected by noise, so there is a possibility that the error in the detected feature point position and feature point correspondence will be large. Also, pixels with excessively high brightness may have a large error in the detected feature point position and feature point correspondence due to the effect of overexposure. Therefore, for example, by calculating an evaluation value based on the brightness of each pixel and selecting feature points with an evaluation value higher than a predetermined threshold as extracted feature point correspondence information, it is possible to eliminate feature point correspondences 22 with low accuracy. The evaluation value here is high when the brightness of a pixel falls within a predetermined range, and low when it falls outside that range. Furthermore, for example, in image data 2 taken outdoors, feature point correspondences related to time-varying subjects 20, such as clouds and shade, which cause a decrease in three-dimensional reconstruction accuracy, can be eliminated using pixel information such as color information.

[0066] Figure 15 shows another example of a method for calculating evaluation values ​​for feature point correspondence. In the example shown in Figure 15, the feature point coordinate pair output unit 14 extracts corresponding feature points based on the image quality evaluated for areas exceeding a predetermined area in each of the two input images. For example, in Figure 15, the subject 20 is blurred in the right-hand image. Generally, feature point detection in blurred images may have a large error in detection position. Therefore, the blur of the image can be calculated by, for example, a Laplacian filter or analysis of high-frequency components, and the feature point correspondence within the image can be evaluated. In addition to the amount of blur indicating the amount of blur, one or more of the following may be used as image quality for evaluation: amount of noise, brightness, resolution, contrast (amount of blown-out highlights, amount of crushed blacks), artifacts associated with compression, moiré, exposure, flare, chromatic aberration, etc. By eliminating feature point pairs exceeding a predetermined threshold for blur, noise, highlight clipping, black clipping, artifacts associated with compression, moiré, flare, and chromatic aberration, and by eliminating feature point pairs below a predetermined threshold for resolution, and by eliminating feature point pairs that do not fall within a predetermined range for brightness and exposure, the accuracy of the feature point correspondence 22 can be reduced.

[0067] Figure 16 shows another example of a method for calculating the evaluation value of feature point correspondence. In the example shown in Figure 16, the feature point coordinate pair output unit 14 performs graph analysis of a network where image data 2 are nodes and image pairs with feature point correspondences are edges, based on the feature point correspondence information output from the inter-image feature point correspondence detection unit 13. As a result of the graph analysis, for image pairs evaluated as correct image pairs (image correspondences 24), the feature point correspondences contained in those images are output as extracted feature point correspondence information. In Figure 16, image correspondences 24 evaluated as correct as a result of the graph analysis are shown with solid lines, and image miscorresponds 25 evaluated as incorrect are shown with dashed lines. Graph analysis can be performed, for example, by selecting using edge-between centrality as described in the processing details of the pair image candidate output unit 12. Furthermore, incorrect pair image candidates may be removed by various features such as short circuit, order, clustering coefficient, and page rank, or combinations thereof. Also, anomaly detection using these features may be performed, and if the amount of anomalies exceeds a predetermined threshold, the image correspondence 24 as an edge may be evaluated as abnormal. The image data 2 as a node may be evaluated, and the feature point correspondences that are considered correct may be extracted and output as feature point correspondence information.

[0068] While examples of processing performed by the pair image candidate output unit 12 and the feature point coordinate pair output unit 14 have been described, the processing performed by each of the pair image candidate output unit 12 and the feature point coordinate pair output unit 14 does not have to be of only one type. Multiple types of processing may be combined as needed through user interface operations. Alternatively, multiple types of processing may be combined using a machine learning model. This machine learning model learns the relationship between the combination of image and processing and a predetermined index indicating the accuracy of the image pair. When an image and processing combination is input, it outputs the accuracy of the image pair. Then, it can select a combination of processing that produces an image pair with good accuracy to generate a high-accuracy image pair.

[0069] Figures 17A, 17B, and 17C show the effects of the first embodiment of the present invention.

[0070] In Figures 17A, 17B, and 17C, solid lines indicate the case where the present invention is applied, and dashed lines indicate the case where the prior art is applied, resulting in feature point mismatches 23 and image mismatches 25 in the inter-image feature point correspondence detection unit 13, for example, duplication where a single object appears multiple times in the three-dimensional spatial information, thus reducing the accuracy of the three-dimensional spatial information. For comparison, a dashed line shows the case where the prior art is applied and feature point mismatches 23 and image mismatches 25 do not occur in the inter-image feature point correspondence detection unit 13. In each figure, lines are shifted as necessary to avoid overlapping lines.

[0071] Figure 17A schematically shows the relationship between the number of images 2 and the computation time when constructing three-dimensional spatial information of the same subject. In the conventional technology shown by the dashed and dotted lines, the inter-image feature point correspondence detection unit 13 performs processing for all combinations of image data 2, so the computation time increases in proportion to the square of the number of images N. On the other hand, in the first embodiment of the present invention, as shown by the solid line, for example as explained in Figure 7, the pair image candidate output unit 12 outputs pair image candidate 31 that are likely to correspond, and the inter-image feature point correspondence detection unit 13 performs efficient processing, thereby suppressing the computation time to a function O(N) that increases in proportion to the number of images N.

[0072] Figure 17B schematically shows the relationship between the number of images 2 in the image data and the accuracy of the three-dimensional spatial information, expressed, for example, in decibels as PSNR (Peak signal-to-noise ratio), when three-dimensional spatial information of the same subject is constructed. In the conventional technology, as shown by the dashed line, as the number of images N increases, feature point mismatches 23 and image mismatches 25 occur in the inter-image feature point correspondence detection unit 13, resulting in, for example, duplication of the three-dimensional spatial information and a decrease in the accuracy of the three-dimensional spatial information. On the other hand, in the first embodiment of the present invention, as shown by the solid line, a pair image candidate 31 with a high probability of correspondence is output from the pair image candidate output unit 12, and the feature point coordinate pair output unit 14 selects the extracted feature point correspondence information based on the calculated evaluation value, thereby suppressing the occurrence of feature point mismatches 23 and image mismatches 25. Therefore, the decrease in accuracy due to, for example, duplication of three-dimensional spatial information can be reduced. Here, we are evaluating the accuracy of the three-dimensional spatial information using PSNR, assuming, for example, image generation from a free viewpoint using Gaussian slatting. However, other values ​​may be used as the accuracy evaluation value.

[0073] Figure 17C schematically shows the relationship between the computation time shown in Figure 17A and the accuracy of the three-dimensional spatial information shown in Figure 17B, with the number of images as a parameter. In the conventional technology, as shown by the dashed line, when the same computation time is applied, the accuracy of the three-dimensional spatial information decreases because fewer images can be used compared to the first embodiment of the present invention shown by the solid line. Furthermore, if duplication of three-dimensional spatial information occurs due to feature point mismatching 23 or image mismatching 25, the accuracy decreases even further than in the first embodiment of the present invention, as shown by the dashed line.

[0074] When conventional technology is applied, the reduction in the accuracy of three-dimensional spatial information due to feature point mismatch 23 and image mismatch 25 occurs with a frequency that cannot be ignored. When a pair of candidate images 31 is input to the inter-image feature point correspondence detection unit 13, for example, let's assume that image mismatch 25 occurs with a probability of 0.1%. If 1000 image data 2 are input, image mismatch 25 will occur in 0.1% of all combinations, or 500 pairs.

[0075] The first embodiment of the present invention has been described above. According to the first embodiment, robust dense feature point matching can be performed while reducing computation time and resources, and feature point mismatching can be suppressed. The reduction in computation time and resources can be achieved by appropriately selecting candidate pair images for feature point matching, and the suppression of mismatching can be achieved based on the selection of candidate pair images and evaluation values ​​calculated from the feature point matching results. By using these feature point matching results, high-precision three-dimensional reconstruction from images becomes possible. Furthermore, high-precision three-dimensional reconstruction becomes possible from images without location information or distance data, such as those taken with an inexpensive camera (for example, a camera that does not acquire location information or distance data).

[0076] [Second Embodiment] Figure 18 is a block diagram showing an example of the configuration of a three-dimensional spatial information system 1 according to the second embodiment of the present invention.

[0077] The three-dimensional spatial information configuration system 1 of the second embodiment includes an input I / F 110, a pair image candidate output unit 12, an image inter-feature point correspondence detection unit 13, a feature point coordinate pair output unit 14, a configuration unit 15, and an output I / F 111. In the second embodiment, the same reference numerals are used for the same components and processes as in the first embodiment described above, and their descriptions are omitted.

[0078] The extracted feature point correspondence information output by the feature point coordinate pair output unit 14 is input to the configuration unit 15 as well as the pair image candidate output unit 12. As a result, the pair image candidate output unit 12 receives dense and robust feature point correspondence information calculated by the inter-image feature point correspondence detection unit 13 and extracted by the feature point coordinate pair output unit 14. Therefore, the pair image candidate output unit 12 can adaptively and sequentially select the next pair image candidate to be output, taking into account the dense and robust feature point correspondence information.

[0079] In the first embodiment, the pair image candidate output unit 12, the inter-image feature point correspondence detection unit 13, and the feature point coordinate pair output unit 14 are basically executed sequentially for a series of image data 2. On the other hand, in the second embodiment, images are extracted and processed step by step for a series of image data 2, and the series of processes can be executed iteratively. In other words, the pair image candidate output unit 12, the inter-image feature point correspondence detection unit 13, and the feature point coordinate pair output unit 14 execute processing on a subset of the image data 2, and considering the dense and robust feature point correspondence information output by the feature point coordinate pair output unit 14, the series of processes can be executed on the next subset of the image data 2 in an iterative manner.

[0080] As an example, let's describe the case where the paired image candidate output unit 12 performs graph analysis. In this example, in the network used for graph analysis, the image correspondence 24 that forms an edge can be used as the image correspondence 24 based on the extracted feature point correspondence information output from the feature point coordinate pair output unit 14. For example, the number of feature point pairs in an image pair can be selected using a predetermined threshold, and an image pair with a large number of feature point pairs can be selected. This allows for high-precision graph analysis, improving the accuracy of the paired image candidates output by the paired image candidate output unit 12. Specifically, consider the case where the paired image candidate output unit 12 determines that an image correspondence 24 estimated to be a paired image candidate 32 by, for example, SIFT is an image mismatch 25 based on the extracted feature point correspondence information output from the feature point coordinate pair output unit 14. In this case, the paired image candidate output unit 12 can remove the edge corresponding to this image correspondence 24 from the graph used for analysis. By excluding the coordinate pairs excluded by the feature point coordinate pair output unit 14 in the pair image candidate output unit 12, the graph from which the edges corresponding to the image correspondence 24 determined to be an image mismatch 25 have been removed can be used for analysis, optimizing the processing performed by the pair image candidate output unit 12 and improving the accuracy of the pair image candidates 32 output from the pair image candidate output unit 12.

[0081] As another example, we will describe the case where union find is performed in the paired image candidate output unit 12. In this example, in grouping by union find, the image correspondence 24 based on the extracted feature point correspondence information output from the feature point coordinate pair output unit 14 is used. This allows union find to be performed with high accuracy, and the accuracy of the paired image candidates output by the paired image candidate output unit 12 is improved.

[0082] Furthermore, the paired image candidate output unit 12 and the feature point coordinate pair output unit 14 may adaptively select the processing content in each output unit or adaptively change the parameters of each processing unit based on the output results of the other. For example, if the feature point coordinate pair output unit 14 determines that a periodic pattern exists from the distribution shape corresponding to the feature points, the paired image candidate output unit 12 may choose to perform graph analysis processing rather than SIFT processing, which is susceptible to the influence of periodic patterns.

[0083] The second embodiment of the present invention has been described above. According to the second embodiment, the accuracy of the pair image candidates output by the pair image candidate output unit 12 can be improved by iteratively performing processing using the extracted feature point correspondence information output by the feature point coordinate pair output unit 14.

[0084] [Third Embodiment] Figure 19 is a block diagram showing an example configuration of the three-dimensional spatial information configuration system 1 according to the third embodiment of the present invention.

[0085] The three-dimensional spatial information configuration system 1 of the third embodiment includes an input I / F 110, a pair image candidate output unit 12, an image feature point correspondence detection unit 13, a feature point coordinate pair output unit 14, a configuration unit 15, an output I / F 111, and a user selection processing input unit 161. In the third embodiment, the same reference numerals are used for the same configurations and processes as in the first embodiment described above, and their descriptions are omitted.

[0086] In the third embodiment, the user selection processing input unit 161 receives the user's selection as input and outputs it to the pair image candidate output unit 12 and the feature point coordinate pair output unit 14. The pair image candidate output unit 12 and the feature point coordinate pair output unit 14 execute processing according to the user's selection. The input received by the user selection processing input unit 161 is the analysis content to be executed by the pair image candidate output unit 12 and the feature point coordinate pair output unit 14, respectively, such as graph analysis. The user selection processing input unit 161 may present candidate analysis content to the user. Alternatively, it may receive input that combines multiple analysis content.

[0087] Figure 20 shows an example of the user interface of this embodiment. The example of the user interface shown in Figure 20 includes an image capture data information area 210, a processing content selection area 220 in the paired image candidate output unit 12, and a processing content selection area 230 in the feature point coordinate pair output unit 14.

[0088] The shooting data information area 210 includes, for example, an image enlargement area 211 for checking image data 2, a list display area 212, a shooting parameter area 213 for checking camera settings at the time of shooting, a shooting status area 214 for checking the situation at the time of shooting, and a memo area 215 for which supplementary information is written. The shooting data information area 210 may also be made available to allow the user to modify or add information as needed.

[0089] The processing content selection area 220 in the paired image candidate output unit 12 includes, for example, a processing candidate area 221 where a list of processes that can be executed in the paired image candidate output unit 12 can be viewed, a processing flow area 222 where the user can select a process from the processing candidates to construct a processing flow, a processing parameter area 223 that displays the parameters of each process and allows the user to set them, and a processing result example area 224 that presents an example of the processing result obtained as a result of the constructed processing flow. For example, the processing flow area 222 may be an interface in which processing blocks are dragged from the processing candidate area 221 to construct a processing flow based on visual programming. In the processing parameter area 223, recommended parameters may be displayed. Recommended parameters are, for example, parameters calculated in advance to accommodate a wide range of shooting conditions for each process. In the processing result example area 224, a graph format display is exemplified in Figure 20, but the user or system may appropriately select a display format, such as the tabular format shown in Figure 7.

[0090] The processing content selection area 230 in the feature point coordinate pair output unit 14 includes, for example, a processing candidate area 231 where a list of processes that can be executed by the feature point coordinate pair output unit 14 can be viewed, a processing flow area 232 where the user can select a process from the processing candidates to construct a processing flow, a processing parameter area 233 that displays the parameters of each process and allows the user to set them, and a processing result example area 234 that presents an example of a processing result obtained as a result of the constructed processing flow. Recommended parameters may be displayed in the processing parameter area 233. Since the output of the feature point coordinate pair output unit 14 is extracted feature point correspondence information, the processing result example area may, for example, display feature point correspondence information for a pair of images.

[0091] The third embodiment of the present invention has been described above. According to the third embodiment, the user can select the processing to be performed in the pair image candidate output unit 12 and the feature point coordinate pair output unit 14, thereby enabling processing suitable for the input image data 2, improving the accuracy of the pair image candidates 31 output by the pair image candidate output unit 12, and improving the accuracy of the extracted feature point correspondence information output by the feature point coordinate pair output unit 14.

[0092] [Fourth Embodiment] Figure 21 is a block diagram showing an example of the configuration of the three-dimensional spatial information system 1 according to the fourth embodiment of the present invention.

[0093] The three-dimensional spatial information configuration system 1 of the fourth embodiment includes an input I / F 110, a pair image candidate output unit 12, an image inter-feature point correspondence detection unit 13, a feature point coordinate pair output unit 14, a configuration unit 15, an output I / F 111, an image capture information input unit 162, and a processing optimization unit 163. In the fourth embodiment, the same reference numerals are used for the same configurations and processes as in the first embodiment described above, and their descriptions are omitted.

[0094] In the fourth embodiment, the shooting information input unit 162 receives image data 2 and its shooting information as input from the user or another system, and outputs the received image data 2 and shooting information to the processing optimization unit 163. Based on the image data 2 and its shooting information output from the shooting information input unit 162, the processing optimization unit 163 estimates the optimal processing to be performed in the paired image candidate output unit 12 and the feature point coordinate pair output unit 14. For example, the shooting information may include information about the indoor or outdoor shooting location, information about the subject such as a factory or dam, camera settings at the time of shooting such as exposure time, whether or not the lighting was moved, and whether or not moving objects were captured in the image.

[0095] For example, blurring is expected to occur when a long exposure time is set. In this case, the processing optimization unit 163 may select image quality-based extraction as one of the processes to be performed in the feature point coordinate pair output unit 14.

[0096] As another example of selecting the optimal processing method, if the input information for the image is an indoor location or a factory, it is assumed that the subject 20 will contain many man-made objects and that periodic patterns will exist. In this case, the processing optimization unit 163 may select graph analysis as the processing for the pair image candidate output unit 12, and select extraction based on the distribution concentration and distribution shape of the feature point correspondence 22 as the processing for the feature point coordinate pair output unit 14.

[0097] Furthermore, by presenting the selected optimal processing method and its rationale to the user, it is possible to show the user areas for improvement during image capture for high-precision 3D reconstruction.

[0098] Figure 22 shows an example of the user interface of this embodiment. The example of the user interface shown in Figure 22 includes an image capture data information area 260, a processing content optimization result area 270 in the paired image candidate output unit 12, and a processing content optimization result area 280 in the feature point coordinate pair output unit 14.

[0099] The shooting data information area 260 includes, for example, an image enlargement area 261 for checking the image data 2, a list display area 262, a shooting parameter area 263 for checking the camera settings at the time of shooting, a shooting status area 264 for checking the conditions at the time of shooting, and a memo area 265 containing supplementary information.

[0100] The processing content optimization result area 270 in the paired image candidate output unit 12 includes, for example, a processing candidate area 271 where a list of processes that can be executed in the paired image candidate output unit 12 can be confirmed, a processing flow area 272 where the optimal processing flow selected by the processing optimization unit 163 is displayed, a processing parameter area 273 where the parameters of each process are displayed and can be set by the user, a processing result example area 274 which presents an example of the processing result obtained as a result of the constructed processing flow, a selection basis area 275 which presents the basis for the optimal processing selection in the processing optimization unit 163, and a shooting advice area 276 which presents a suggestion for a shooting method for the next shooting. In the processing parameter area 2733, recommended parameters may be displayed. In the processing result example area 274, although a graph format display is exemplified in Figure 22, the user or system may appropriately select a display format, such as the tabular format shown in Figure 7.

[0101] The processing content optimization result area 280 in the feature point coordinate pair output unit 14 includes, for example, a processing candidate area 281 where a list of processes that can be executed in the feature point coordinate pair output unit 14 can be viewed, a processing flow area 282 where the optimal processing flow selected by the processing optimization unit 163 is displayed, a processing parameter area 283 where the parameters of each process are displayed and can be set by the user, a processing result example area 284 which presents an example of a processing result obtained as a result of the constructed processing flow, a selection basis area 285 which presents the basis for the optimal processing selection in the processing optimization unit 163, and a shooting advice area 286 which presents a suggestion for a shooting method for the next shooting. Recommended parameters may be displayed in the processing parameter area 283. Since the output of the feature point coordinate pair output unit 14 is extracted feature point correspondence information, the processing result example area may, for example, display feature point correspondence information for a pair of images.

[0102] In the shooting data information area 260 and the processing content optimization result areas 270 and 280, the user may be allowed to modify or add information as needed. In response to the user's modification or addition of information, the processing content optimization results are updated, and the optimization results displayed in the processing content optimization result areas 270 and 280 are also updated.

[0103] The fourth embodiment of the present invention has been described above. According to the fourth embodiment, the processing to be performed in the paired image candidate output unit 12 and the feature point coordinate pair output unit 14 can be optimized using the image data 2 and shooting information received as input from the user. As a result, the accuracy of the paired image candidates 31 output by the paired image candidate output unit 12 is improved, and the accuracy of the extracted feature point correspondence information output by the feature point coordinate pair output unit 14 is improved.

[0104] [Fifth Embodiment] Figures 23A and 23B are block diagrams showing an example configuration of the three-dimensional spatial information system 1 according to the fifth embodiment of the present invention.

[0105] The third-dimensional spatial information configuration system 1 of the fifth embodiment, as shown in Figure 23A, includes an imaging unit 17, an input I / F 110, a paired image candidate output unit 12, an inter-image feature point correspondence detection unit 13, a feature point coordinate pair output unit 14, a configuration unit 15, and an output I / F 111. In the fifth embodiment, the same components and processes as in the first embodiment described above are denoted by the same reference numerals, and their descriptions are omitted. In the fifth embodiment, as shown in the figure, the third-dimensional spatial information configuration system 1 includes an imaging unit 17.

[0106] The imaging unit 17 captures the subject 20 to be reconstructed in three dimensions based on various camera settings, generates image data 2, and outputs it to the input I / F 110. For example, the imaging unit 17 can be any device capable of acquiring still images or videos, such as a camera, smartphone, 360° camera, stereo camera, or RGBD camera.

[0107] As shown in Figure 23B, the component 15 includes a camera position and orientation estimation unit 151, a triangulation unit 152, a bundle adjustment unit 153, a depth map calculation unit 154, a depth map integration unit 155, and a three-dimensional spatial information configuration unit 156.

[0108] The camera position and orientation estimation unit 151 calculates the position and orientation of the camera from which each image data 2 was taken, based on the extracted feature point correspondence information output from the feature point coordinate pair output unit 14, and outputs this to the triangulation unit 152. The camera position and orientation are calculated using, for example, a basic matrix or foundation matrix and an outlier removal algorithm. At this time, the image correspondences 24 used as extracted feature point correspondence information may be improved in accuracy by setting a threshold for the number of feature point correspondences and limiting them to only images from which a predetermined number or more feature point correspondences have been extracted. Furthermore, the relative orientation of each image correspondence 24 calculated from the extracted feature point correspondence information may also be improved in accuracy by setting a threshold for accuracy and limiting it to only those with an accuracy of a predetermined level or higher.

[0109] The triangulation unit 152 uses the camera position and orientation output from the camera position and orientation estimation unit 151 to calculate a three-dimensional point cloud of the subject 20 using the principle of triangulation, and outputs it to the bundle adjustment unit 153.

[0110] The bundle adjustment unit 153 optimizes the three-dimensional point cloud and camera position / orientation output from the triangulation unit 152, and outputs the optimized result to the depth map calculation unit 154. Specifically, based on the extracted feature point correspondence information between multiple images, it takes this extracted feature point correspondence information, the three-dimensional point cloud, and the camera position / orientation as input and performs a nonlinear optimization process to minimize the reprojection error. Furthermore, a robust error function may be used to mitigate the effects of outliers. This improves the accuracy of the three-dimensional point cloud and makes the camera position / orientation estimation more accurate. The optimized three-dimensional point cloud and camera position / orientation are used as input data for the depth map calculation unit 154 to generate accurate depth information.

[0111] The depth map calculation unit 154 uses the optimized three-dimensional point cloud and camera position / orientation output from the bundle adjustment unit 153 to calculate a depth map for each image and outputs it to the depth map integration unit 155.

[0112] The depth map integration unit 155 integrates the depth maps output from the depth map calculation unit 154 to calculate a dense three-dimensional point cloud and outputs it to the three-dimensional spatial information constructing unit 156.

[0113] The three-dimensional spatial information constructor 156 calculates three-dimensional spatial information using the dense three-dimensional point cloud output from the depth map integration unit 155 and outputs it to the output I / F 111. For example, the calculation of three-dimensional spatial information may involve generating a polygon mesh from the dense three-dimensional point cloud, applying a texture, and generating three-dimensional spatial information with surface information. As another example of calculating three-dimensional spatial information, an image from an arbitrary viewpoint may be generated by NeRF using image data 2 and camera position and orientation as input, and learning the radiance and density of the entire scene. Alternatively, a realistic image from an arbitrary viewpoint may be generated at high speed using Gaussian Splatting, which uses a model trained to represent the three-dimensional space as a set of three-dimensional Gaussians, with image data 2, camera position and orientation, and the dense point cloud as input. Furthermore, the output of the three-dimensional spatial information constructor 156 may be the dense three-dimensional point cloud itself, or a dense three-dimensional point cloud with noise removed.

[0114] The input I / F 110 to output I / F 111 of the three-dimensional spatial information configuration system 1 is, for example, a computing system located outside the imaging unit 17, and can be configured as, for example, a general-purpose PC. The imaging unit 17 and the input I / F 110 are connected directly or via a communication network to input and output image data 2. Alternatively, image data 2 may be input and output using an information storage medium.

[0115] The paired image candidate output unit 12 and the feature point coordinate pair output unit 14 may appropriately select the processing to be performed depending on the type of imaging unit 17 and the type of image data 2 to be acquired. For example, when using a camera with a small sensor size, such as a smartphone, as the imaging unit 17, a lot of noise is generated in dark places, so it is advisable to select feature point correspondence extraction using pixel information as the processing for the feature point coordinate pair output unit 14.

[0116] The fifth embodiment of the present invention has been described above. According to the fifth embodiment, a specific series of processing pipelines can be implemented, from the acquisition of image data 2 of the subject 20 by the imaging unit 17 to the output of various types of three-dimensional spatial information.

[0117] It should be noted that the present invention is not limited to the embodiments described above, but includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are described in detail for the purpose of clearly illustrating the present invention, and the present invention is not necessarily limited to having all the configurations described. Furthermore, some of the configurations of one embodiment may be replaced with those of another embodiment. Furthermore, configurations of other embodiments may be added to the configuration of one embodiment. Furthermore, some of the configurations of each embodiment may be added, deleted, or replaced with those of other embodiments.

[0118] Furthermore, each of the aforementioned configurations, functions, processing units, and processing means may be implemented in hardware, for example, by designing them as integrated circuits, or they may be implemented in software by having a processor interpret and execute programs that realize each function.

[0119] Information such as programs, tables, and files that implement each function can be stored in memory, hard disks, SSDs (Solid State Drives), or recording media such as IC cards, SD cards, and DVDs.

[0120] Furthermore, the control lines and information lines shown are those deemed necessary for explanation purposes and do not necessarily represent all control lines and information lines required for implementation. In reality, it can be assumed that almost all components are interconnected.

Claims

1. A three-dimensional spatial information configuration system comprising a computer having a computing device that performs predetermined processing and a storage device accessible by the computing device, the system comprising: an image-to-image feature point correspondence detection unit that detects a plurality of corresponding feature points from pair image candidates that form pairs of images of the same subject; a feature point coordinate pair output unit that evaluates the corresponding feature points, selects feature points based on the evaluation results of the feature points, and outputs extracted feature point correspondence information which is a set of the selected feature points; and a configuration unit that configures three-dimensional spatial information based on the positional relationship of the extracted feature point correspondence information.

2. A three-dimensional spatial information configuration system according to claim 1, characterized in that it comprises a pair image candidate output unit that outputs the pair image candidates from three or more images.

3. A three-dimensional spatial information configuration system according to claim 2, wherein the pair image candidate output unit sequentially performs the processes of detecting feature points from the image and associating the detected feature points with each other across different images.

4. A three-dimensional spatial information configuration system according to claim 3, characterized in that the number of pair image candidate searches in the pair image candidate output unit increases linearly with the number of input images by limiting the number of search target images to a predetermined upper limit.

5. A three-dimensional spatial information configuration system according to claim 2, wherein the pair image candidate output unit selects pair image candidates based on the graph structure of the image pair.

6. A three-dimensional spatial information configuration system according to claim 2, wherein the pair image candidate output unit selects a pair image candidate based on the result of grouping by union-finding of the image pair.

7. A three-dimensional spatial information configuration system according to claim 2, wherein the pair image candidate output unit selects image pairs that are close in position as pair image candidates, based on position information associated with the images.

8. A three-dimensional spatial information configuration system according to claim 1, wherein the feature point coordinate pair output unit selects the feature points based on the feature point information output from the inter-image feature point correspondence detection unit.

9. A three-dimensional spatial information configuration system according to claim 8, wherein the information of the feature points is characterized in that at least one of the number of feature points, the degree of concentration of the feature point distribution, the shape of the feature point distribution, and the pixel information of the feature point location.

10. A three-dimensional spatial information configuration system according to claim 1, wherein the feature point coordinate pair output unit evaluates the feature points based on the image quality feature quantities and selects the feature points based on the evaluation results of the feature points.

11. A three-dimensional spatial information configuration system according to claim 10, characterized in that the image quality feature quantity is at least one of blur amount, noise amount, brightness, highlight clipping amount, and black clipping amount.

12. A three-dimensional spatial information configuration system according to claim 1, wherein the feature point coordinate pair output unit evaluates the feature points based on the graph structure of the image pair or the result of grouping by unionfinding.

13. A three-dimensional spatial information configuration system according to any one of claims 8 to 12, wherein the feature point coordinate pair output unit removes erroneous feature point coordinate pairs and image pairs by detecting anomalies in nodes or edges based on at least one of the information of the feature point, the image quality features of the image, the graph structure of the image pair, and the result of grouping by unionfind.

14. A three-dimensional spatial information configuration system according to claim 2, wherein the pair image candidate output unit receives the result of the evaluation of the feature points output from the feature point coordinate pair output unit as input.

15. A three-dimensional spatial information configuration system according to claim 2, characterized in that it provides an interface that allows the user to adjust the parameters of the pair image candidate output unit and the feature point coordinate pair output unit.

16. A three-dimensional spatial information configuration system according to claim 2, characterized in that it presents recommended parameters for the pair image candidate output unit and the feature point coordinate pair output unit.

17. A three-dimensional spatial information configuration system according to claim 1, wherein the configuration unit comprises: a camera position and orientation estimation unit that calculates the position and orientation of a camera from which the plurality of images were taken based on the extracted feature point correspondence information and the plurality of images; a triangulation unit that calculates a three-dimensional point cloud from the position and orientation of the camera using the principle of triangulation; a bundle adjustment unit that optimizes the three-dimensional point cloud and the position and orientation of the camera; a depth map calculation unit that calculates depth maps corresponding to the plurality of images from the optimized three-dimensional point cloud and the position and orientation of the camera; a depth map integration unit that integrates the depth maps to calculate a dense three-dimensional point cloud; and a three-dimensional spatial information configuration unit that calculates three-dimensional spatial information from the dense three-dimensional point cloud.

18. A three-dimensional spatial information configuration system comprising: a camera that outputs multiple images of a space; a pair image candidate output unit that outputs pair image candidates from three or more images output from the camera, which are pairs of images of the same subject; an inter-image feature point correspondence detection unit that detects multiple feature points corresponding to each other from the pair image candidates; a feature point coordinate pair output unit that evaluates the corresponding feature points, selects feature points based on the evaluation results of the feature points, and outputs extracted feature point correspondence information which is a set of the selected feature points; and a component that constructs three-dimensional spatial information based on the positional relationship of the extracted feature point correspondence information.

19. A three-dimensional spatial information configuration system according to claim 18, wherein the camera captures a moving image, and the pair image candidate output unit outputs the pair image candidates from frame images with close shooting times included in the moving image captured by the camera.

20. A three-dimensional spatial information configuration system according to claim 18, wherein the camera has the pair image candidate output unit, and the inter-image feature point correspondence detection unit receives the pair image candidates output from the pair image candidate output unit of the camera as input.

21. A method for constructing three-dimensional spatial information by a computer, wherein the computer comprises an arithmetic unit that performs predetermined processing and a storage device accessible by the arithmetic unit, and the method for constructing three-dimensional spatial information comprises: an inter-image feature point correspondence detection procedure for detecting a plurality of feature points corresponding between images from pair image candidates that form pairs of images of the same subject; a feature point coordinate pair output procedure for calculating evaluation values ​​of the corresponding feature points, selecting feature points based on the calculated evaluation values, and outputting extracted feature point correspondence information which is a set of the selected feature points; and a configuration procedure for constructing three-dimensional spatial information based on the positional relationship of the extracted feature point correspondence information.