Image processing device, program, and image processing method
The image processing device addresses the challenge of tilting in satellite images by generating corresponding point pairs and using looser criteria for axis determination, ensuring accurate alignment despite tilted features.
Patent Information
- Application Number
- PCT/JP2024/023382
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-02
AI Technical Summary
Conventional image alignment methods struggle with accurately aligning satellite images due to the tilting of tall objects like buildings, leading to insufficient feature points and reduced alignment accuracy when many tall buildings are present.
An image processing device that generates corresponding point pairs by extracting points from images captured from above, calculates difference vectors based on image capturing direction and posture, and determines inappropriate pairs using looser criteria for one axis to ensure a sufficient number of feature points, thereby enhancing alignment accuracy.
Ensures a sufficient number of feature points while suppressing the influence of tilting, enabling highly accurate alignment of satellite images.
Smart Images

Figure JP2024023382_02012026_PF_FP_ABST
Abstract
Description
Image processing device, program, and image processing method
[0001] The present disclosure relates to an image processing device, a program, and an image processing method.
[0002] Satellite images are used in a variety of fields. However, since there is some error in the location information indicating the coverage area of satellite images, it is necessary to accurately identify the coverage area of the satellite image by overlaying it with a previous image or map. To perform such overlay, positioning is required based on the features visible on the image.
[0003] One method for automatically aligning images involves extracting feature points from the images, detecting corresponding points from the extracted feature points using feature point matching, and then aligning the images using the detected corresponding points. However, tall objects such as buildings can be "tilted," meaning that they appear differently depending on the angle at which they are captured. For example, when a tall object is captured from an angle, its position on the image will differ from when it is captured from directly above.
[0004] The image processing device described in Patent Document 1 determines whether a feature point is affected by collapse based on edge information around the feature point, and removes feature points affected by collapse from a feature point list. Images are then aligned using the feature points included in the feature point list.
[0005] JP 2014-126893 A
[0006] However, as with conventional technology, when there are many tall buildings, if all tilted feature points are removed, a sufficient number of corresponding points cannot be obtained, resulting in alignment failure or reduced accuracy.
[0007] Therefore, one or more aspects of the present disclosure aim to ensure a sufficient number of feature points while suppressing the influence of the collapse, thereby enabling highly accurate alignment.
[0008] An image processing device according to one aspect of the present disclosure includes a corresponding point extraction unit that generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of a first image and a second image captured from above so that at least a predetermined space is included; a difference vector calculation unit that calculates, in accordance with the image capturing direction and posture, a difference vector that is the difference between a first vector indicating the orientation and size of a line segment that stands upright in the space and has a known length when projected onto the first image, and a second vector indicating the orientation and size of the line segment when projected onto the second image; and a difference vector calculation unit that calculates positions of the plurality of corresponding point pairs by: The image processing system includes a coordinate determination unit that determines the difference vector using coordinate values in a coordinate system having a first axis parallel to the difference vector and a second axis perpendicular to the first axis, and a registration unit that determines inappropriate corresponding point pairs from the plurality of corresponding point pairs using the coordinate values of the plurality of corresponding point pairs, and aligns the first image and the second image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs, wherein the registration unit is characterized in that, in determining that a corresponding point pair is not inappropriate, the standard for the coordinate values of the first axis is looser than the standard for the coordinate values of the second axis.
[0009] A program according to one aspect of the present disclosure includes a computer including: a corresponding point extraction unit that generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of a first image and a second image captured from above so that at least a predetermined space is included; a difference vector calculation unit that calculates, in accordance with the image capturing direction and posture, a difference vector that is the difference between a first vector indicating the orientation and size of a line segment that stands upright in the space and has a known length when projected onto the first image, and a second vector indicating the orientation and size of the line segment when projected onto the second image; and a difference vector calculation unit that calculates positions of the plurality of corresponding point pairs by The image processing device functions as a coordinate determination unit that determines the corresponding points by coordinate values in a coordinate system having a first axis parallel to the difference vector and a second axis perpendicular to the first axis, and an alignment unit that determines inappropriate corresponding point pairs from the plurality of corresponding point pairs using the coordinate values of the plurality of corresponding point pairs, and aligns the first image and the second image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs, and the alignment unit is characterized in that, in determining that a corresponding point pair is not inappropriate, the standard for the coordinate values of the first axis is looser than the standard for the coordinate values of the second axis.
[0010] An image processing method according to one aspect of the present disclosure generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of a first image and a second image captured from above so that the images include at least a predetermined space, calculates a difference vector which is the difference between a first vector indicating the direction and size of a line segment that stands upright in the space and has a known length projected onto the first image, and a second vector indicating the direction and size of the line segment projected onto the second image, depending on the imaging direction and posture, specifies the positions of the plurality of corresponding point pairs by coordinate values in a coordinate system having a first axis parallel to the difference vector and a second axis perpendicular to the first axis, determines inappropriate corresponding point pairs from the plurality of corresponding point pairs using the coordinate values of the plurality of corresponding point pairs, and aligns the first image and the second image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs, wherein the image processing method is characterized in that the criteria for determining that a corresponding point pair is not inappropriate are looser for the coordinate values of the first axis than for the coordinate values of the second axis.
[0011] According to one or more aspects of the present disclosure, it is possible to ensure a sufficient number of feature points while suppressing the influence of tilting, thereby enabling highly accurate alignment.
[0012] 1 is a block diagram schematically illustrating the configuration of an image processing device according to Embodiments 1 and 2. (A) to (D) are schematic diagrams showing examples of a reference image, a target image, and a collapse vector. (A) and (B) are schematic diagrams showing an example of calculating a collapse difference vector from a collapse vector. (A) and (B) are schematic diagrams for explaining corresponding points in a reference image and a target image. (A) and (B) are schematic diagrams for explaining a position error vector of a corresponding point pair identified from a reference image and a target image. (A) and (B) are schematic diagrams showing an example of converting an image coordinate system into a transformation coordinate system. (A) and (B) are schematic diagrams showing a first example of calculating registration parameters. (A) and (B) are schematic diagrams showing a first example of calculating registration parameters. (A) and (B) are block diagrams showing an example of a hardware configuration.
[0013] 1 is a block diagram showing a schematic configuration of an image processing device 100 according to embodiment 1. The image processing device 100 includes a reference image acquisition unit 101, a target image acquisition unit 102, a corresponding point extraction unit 103, a collapse estimation unit 104, a corresponding point conversion unit 105, and a registration unit 106.
[0014] The reference image acquisition unit 101 acquires a reference image that serves as a reference for alignment. For example, the reference image acquisition unit 101 may acquire the reference image from a network such as the Internet via a communication unit (not shown), or, if the reference image is already stored in a storage unit (not shown), the reference image may be acquired from the storage unit. The acquired reference image is provided to the corresponding point extraction unit 103 and the collapse estimation unit 104.
[0015] In addition, the reference image acquisition unit 101 acquires, along with the reference image, reference image supplementary information, which is information indicating the position and orientation of the imaging device when the reference image was captured, and provides the reference image supplementary information to the collapse estimation unit 104.
[0016] The target image acquisition unit 102 acquires a target image, which is an image to be aligned. For example, the target image acquisition unit 102 may acquire the target image from a network such as the Internet via a communication unit (not shown), or, if the target image is already stored in a storage unit (not shown), may acquire the target image from the storage unit. The acquired target image is provided to the corresponding point extraction unit 103 and the collapse estimation unit 104.
[0017] In addition, the target image acquisition unit 102 acquires, along with the target image, target image accompanying information, which is information indicating the position and orientation of the imaging device when the target image was captured, and provides the target image accompanying information to the collapse estimation unit 104.
[0018] Here, both the reference image and the target image may be optical images or other types of images, such as images generated using a Synthetic Aperture Radar (SAR).
[0019] Generally, the reference image is an image previously captured of the same area as the target image, and an image that has been accurately map-projected is used, but an image that has not been map-projected may also be used. When performing map projection, for example, GCPs (Ground Control Points), which are control points acquired on the ground by GNSS (Global Navigation Satellite System) surveying or the like, are separately prepared, and the GCPs are associated with pixel coordinates in the satellite image, thereby enabling accurate map projection.
[0020] From the above, it is sufficient that the reference image and the target image are images captured from above so as to include at least a predetermined space. The reference image is also referred to as a first image, and the reference image acquisition unit 101 is also referred to as a first image acquisition unit. The target image is also referred to as a second image, and the target image acquisition unit 102 is also referred to as a second image acquisition unit.
[0021] The corresponding point extraction unit 103 generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of the reference image and the target image. For example, the corresponding point extraction unit 103 identifies corresponding point pairs by extracting points that appear in both the reference image and the target image as corresponding points from each of the reference image and the target image, and identifies the coordinates of the corresponding point pairs on each image. Corresponding point information indicating the coordinates of the identified corresponding point pairs is provided to the corresponding point conversion unit 105.
[0022] The corresponding point extraction unit 103 can extract corresponding points by using feature-based matching or region-based matching. In the case of feature-based matching, feature points and feature amounts of the feature points are extracted from each image, and feature points in the two images with similar feature amounts are considered to be corresponding point pairs. For the feature amounts, known methods such as ORB (Oriented-Brief), SIFT (Scale-Invariant Feature Transform), or AKAZE (Accelerated KAZE) can be used. In this case, the corresponding point extraction unit 103 considers a feature point extracted from the reference image and a feature point extracted from the target image that matches the extracted feature point as a corresponding point pair.
[0023] In the case of region-based matching, template matching is performed for each small region cut out from one image with the other image, and if the similarity is equal to or greater than a certain value, the position from which the template was cut out and the matching position are regarded as a corresponding point pair. In this case, the corresponding point extraction unit 103 identifies the region in the reference image that best matches the region identified in the target image, and each position indicating the identified region is regarded as a corresponding point pair.
[0024] The collapse estimation unit 104 functions as a difference vector calculation unit that calculates a collapse difference vector as a difference vector between a collapse vector, which is a first vector that indicates the direction and size of a line segment of a known length that stands upright in the same space contained in the reference image and the target image and is projected onto the reference image, and a collapse vector, which is a second vector that indicates the direction and size of the line segment that is projected onto the target image, depending on the imaging direction and posture.
[0025] For example, the collapse estimation unit 104 uses the reference image attendant information and the target image attendant information to calculate, as a collapse vector, the direction and size of collapse of an object of unit height for each of the reference image and the target image, based on the position and orientation of the imaging device at the time of imaging. Here, it is assumed that the unit height is a known length.
[0026] For example, a collapse vector 121 as shown in Fig. 2(B) can be calculated from a reference image 120 shown in Fig. 2(A). Also, a collapse vector 123 as shown in Fig. 2(D) can be calculated from a target image 122 shown in Fig. 2(C).
[0027] The collapse estimation unit 104 can then calculate the difference between the collapse vectors calculated from the reference image and the target image, thereby obtaining a collapse difference vector. For example, as shown in Fig. 3, a collapse difference vector 124 can be calculated from the difference between a collapse vector 121 of the reference image 120 and a collapse vector 123 of the target image 122.
[0028] The collapse estimation unit 104 then provides the collapse information indicating the calculated collapse difference vector to the corresponding point conversion unit 105 .
[0029] The corresponding point transformation unit 105 functions as a coordinate identification unit that identifies the positions of multiple corresponding point pairs using coordinate values in a coordinate system having a first axis parallel to the tilt difference vector and a second axis perpendicular to the first axis.
[0030] For example, the corresponding point transformation unit 105 transforms the coordinates of the corresponding point pair indicated by the corresponding point information into a transformation coordinate system, which is a coordinate system in which one axis is a direction parallel to the collapse difference vector indicated by the collapse information, and the other axis is a direction perpendicular to the collapse difference vector.
[0031] For example, tall buildings are captured in both the reference image 120 shown in Fig. 4(A) and the target image 122 shown in Fig. 4(B). Corresponding points P1 to P4 in the reference image 120 shown in Fig. 4(A) are paired with corresponding points P1# to P4# in the target image 122 shown in Fig. 4(B).
[0032] The corresponding point pairs P1, P1# and P3, P3# are not affected by the collapse because they have no height. The corresponding point pairs P2, P2# and P4, P4# are affected by the collapse because they have height.
[0033] When the reference image 120 and the target image 122 are overlaid in the correct positions, the corresponding points pairs P1, P1# and P3, P3#, which are not affected by the collapse, match, as shown in Figure 5. On the other hand, the corresponding points pairs P2, P2# and P4, P4#, which are affected by the collapse, do not match.
[0034] Here, the position error of the corresponding point pair affected by the collapse is represented by a vector, which is referred to as a position error vector. In the example shown in FIG. 5, the position error vectors are represented by reference numerals 125 and 126.
[0035] The directions of the position error vector and the tilt difference vector are the same, and the magnitude of the position error vector is proportional to the height of the corresponding point pair.
[0036] Returning to Fig. 1 , the corresponding point transformation unit 105 transforms the coordinates of the corresponding point pair extracted by the corresponding point extraction unit 103 using the collapse difference vector calculated by the collapse estimation unit 104. For example, as shown in Fig. 6 , the corresponding point transformation unit 105 transforms the image coordinate system (x, y) into a transformed coordinate system (v, u). The transformed coordinate system (v, u) is a coordinate system specified by an axis v parallel to the collapse difference vector and an axis u perpendicular to the axis v. Here, the v axis is also referred to as the first axis, and the u axis is also referred to as the second axis.
[0037] In this case, an error occurs in the v-axis direction for the corresponding points due to the influence of the tilt, but no error occurs in the u-axis direction. In other words, the u-axis is not affected by the tilt, and only the v-axis is affected by the tilt.
[0038] The registration unit 106 uses the coordinate values of the plurality of corresponding point pairs to determine an inappropriate corresponding point pair from the plurality of corresponding point pairs.
[0039] For example, the alignment unit 106 determines, as an inappropriate corresponding points pair, a corresponding points pair that has been determined to be an outlier by an algorithm that determines an outlier using the distance between each of a plurality of corresponding points pairs. Here, in this embodiment, since the coordinates of a corresponding points pair can be expressed by coordinate values on the first axis and coordinate values on the second axis as described above, the algorithm can use a looser threshold for the coordinate values on the first axis as a criterion for determining that a point is not an outlier than the threshold for the coordinate values on the second axis.
[0040] Furthermore, when the absolute value of the difference between the coordinate values of one of the corresponding points pairs along the first axis is equal to or greater than a first threshold value, or when the absolute value of the difference between the coordinate values of the one of the corresponding points pairs along the second axis is equal to or greater than a second threshold value, the alignment unit 106 determines that the one of the corresponding points pairs is an inappropriate corresponding points pair. Again, the first threshold value can be greater than the second threshold value.
[0041] Alternatively, the alignment unit 106 determines that a corresponding point pair whose coordinates of the corresponding point extracted from the target image are not within an ellipse whose center is the coordinates of the corresponding point extracted from the reference image, whose major axis is parallel to the first axis, and whose minor axis is parallel to the second axis, is an inappropriate corresponding point pair.
[0042] A more detailed description will be given below. The registration unit 106 performs outlier removal, which removes inappropriate corresponding point pairs as outliers from the corresponding point pairs of the target image whose coordinates have been transformed by the corresponding point transformation unit 105. The corresponding point pairs of the target image include pairs that are falsely detected or pairs with large errors due to the influence of collapsing. For this reason, outlier removal of inappropriate corresponding point pairs is performed.
[0043] For the outlier removal, a known algorithm such as RANSAC (RANdom Sample Consensus) or M-estimation may be used.
[0044] Here, we will use the example of finding an affine matrix using RANSAC to explain how to utilize corresponding point pairs that have been decomposed into axes that are not affected by collapsing and axes that are affected by collapsing, in other words, that have undergone coordinate transformation.
[0045] First, the registration unit 106 randomly extracts a small number of corresponding points from the corresponding points in the target image whose coordinates have been transformed by the corresponding point transformation unit 105. Then, the registration unit 106 calculates a provisional affine matrix from the extracted corresponding points for the corresponding points in the reference image whose coordinates have been transformed by the corresponding point transformation unit 105. Furthermore, the registration unit 106 determines whether the remaining corresponding points in the target image whose coordinates have been transformed by the corresponding point transformation unit 105 also fit the affine matrix. In this manner, the registration unit 106 removes, as outliers, corresponding point pairs whose reprojection error is equal to or exceeds a predetermined threshold value from among the corresponding points in the target image whose coordinates have been transformed by the corresponding point transformation unit 105. Here, the reprojection error is the distance between the position coordinates of corresponding point Pi# (i is an identification number for identifying the corresponding point pair and is an integer equal to or greater than 1) in the target image and the position coordinates of corresponding point Pi in the reference image when the position coordinates of the corresponding point Pi# are affine transformed.
[0046] Furthermore, although vertical and horizontal errors are usually evaluated on the same level, in this embodiment, the coordinates are transformed into the u axis, which is not affected by the tilt, and the v axis, which is affected by the tilt, so that these can be evaluated independently. For example, the coordinates of the corresponding points in the reference image, which have been transformed by the corresponding point transformation unit 105, in the corresponding point pair i are expressed as (u r,i , v r,i ), and the coordinates of the corresponding point in the target image are (u t,i , v t,i ), if the following formulas (1) and (2) are not satisfied, it is an outlier. t,i -u r,i |<t 1 (1) |v t,i -v r,i |<t 2 (2)
[0047] For example, if the majority of the surface is covered by features that are affected by collapse, and a sufficient number of corresponding points cannot be obtained using conventional thresholds, the threshold value of the v axis, t 2 In this case, the accuracy decreases because the influence of outliers or collapse remains in the v axis, but the outliers can be removed from the u axis, so it is possible to calculate the alignment parameters with higher accuracy compared to when the threshold is relaxed (larged) in the conventional method. 1 For , setting a value close to "0" makes it possible to reliably remove outliers.
[0048] As another example, when the height of a three-dimensional object in the image is known, such as when height data of a building within the shooting range is available, the threshold value t of the v axis is set to a value smaller than the tilt caused by the three-dimensional object. 2 By setting the above, it is possible to more accurately remove corresponding point pairs affected by the collapse.
[0049] Although the method of determining whether to remove an outlier for each corresponding point pair has been described above, the registration unit 106 may perform outlier determination for each of the u-axis and v-axis. In other words, the registration unit 106 may use corresponding point pairs affected by collapse for registration on the u-axis but not on the v-axis. In other words, for corresponding point pairs whose absolute value of the difference in coordinate values on the first axis is equal to or greater than a predetermined threshold, the registration unit 106 may also align the reference image and target image using only the coordinate values on the second axis. This allows more corresponding point pairs to be used on the u-axis, enabling highly accurate registration.
[0050] Alternatively, the alignment unit 106 may determine an outlier by comprehensively evaluating both the u axis and the v axis. For example, the alignment unit 106 may determine that a corresponding point pair is an outlier if it does not satisfy the following formula (3):
[0051]
[0052] In equation (3), for the corresponding point pair i, an ellipse (radius = t÷a) is defined with the corresponding point of the reference image as the center. 1 and t÷a 2 ), it is determined whether there is a corresponding point of the target image within the ellipse. If the corresponding point of the target image is outside the ellipse, the corresponding point pair i is determined to be an outlier.
[0053] a 1 and a 2 is a predetermined coefficient, and t is a threshold value. For example, if the features affected by the collapse occupy most of the surface of the earth and a sufficient number of corresponding points cannot be obtained using the conventional threshold value, the threshold value of the v axis can be loosened (here, a 2 In this case, the accuracy decreases because the influence of outliers and collapse remains large on the v-axis, but the outliers can be removed from the u-axis, so it is possible to calculate the alignment parameters with higher accuracy than when the threshold is relaxed in the conventional method.
[0054] As another example, when the height of a three-dimensional object in an image is known, such as when height data of a building within the shooting range is available, the distance between the building and the object can be calculated by dividing the distance by t / a.2 Set the threshold so that a becomes small (here, a 2 By increasing the value of ( ), it is possible to more accurately remove corresponding point pairs that are affected by the collapse. As described above, in this embodiment, the reference for the axis that is affected by the collapse can be determined independently from the reference for the axis that is not affected by the collapse.
[0055] The registration unit 106 then aligns the reference image and the target image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs. Here, the registration unit 106 uses a looser criterion for determining that a corresponding point pair is not an inappropriate corresponding point pair with respect to the coordinate values of the second axis than with respect to the coordinate values of the first axis.
[0056] For example, after removing outliers, the registration unit 106 calculates registration parameters using as input the coordinates of the corresponding points pair transformed by the corresponding point transformation unit 105. When the registration is performed by affine transformation, the registration parameters are six numerical values of an affine matrix.
[0057] The model used for the alignment may be a parallel translation, an affine transformation, a homography transformation, or any of these with a higher-order term added. The alignment unit 106 calculates the necessary parameters according to the model used.
[0058] Although the description here has been given mainly assuming satellite images or aerial photographs, the image processing device 100 of this embodiment can also be used for other purposes such as inspection images on a factory line. In such cases, it is necessary to measure or estimate the inclination vector for each camera or the inclination difference vector between cameras by, for example, mounting an inclination sensor on the camera or performing calibration before or after the fact.
[0059] Embodiment 2 As shown in FIG. 1 , an image processing apparatus 200 according to embodiment 2 includes a reference image acquisition unit 101, a target image acquisition unit 102, a corresponding point extraction unit 103, a collapse estimation unit 104, a corresponding point conversion unit 105, and a registration unit 206.
[0060] The reference image acquisition unit 101, the target image acquisition unit 102, the corresponding point extraction unit 103, the collapse estimation unit 104, and the corresponding point conversion unit 105 of the image processing device 200 according to embodiment 2 are similar to the reference image acquisition unit 101, the target image acquisition unit 102, the corresponding point extraction unit 103, the collapse estimation unit 104, and the corresponding point conversion unit 105 of the image processing device 100 according to embodiment 1.
[0061] When deriving alignment parameters from corresponding point pairs, the alignment unit 206 in the second embodiment calculates the maximum likelihood value of the v axis based on the error distribution of the u axis. For example, a case will be described in which the alignment model is a translation and RANSAC is used for outlier removal. As shown in FIG. 7A, outlier removal by RANSAC involves selecting a range of width d on the error histogram of the corresponding point pairs, which is the outlier determination threshold t × 2, and calculating the u t -t and u t + t so that the area selected within the range enclosed by u is the largest. t Then, calculating the alignment parameters by the least squares method corresponds to calculating the centroid c of the selected region.
[0062] As shown in Figure 7(A), when there is no tilt, the error distribution is generally nearly symmetrical, making it possible to calculate alignment parameters with high accuracy using the method described above. On the other hand, as shown in Figure 7(B), when there is tilt, the error distribution is biased in one direction. In the example shown in Figure 7(B), the error distribution is shifted to the right.
[0063] Generally, if there is no tilt, the error distribution is considered to be isotropic, so if there is no tilt, the v-axis should have the same error distribution as the u-axis. Utilizing this, a method for calculating the alignment parameters for the v-axis with high accuracy will be described.
[0064] First, the alignment unit 206 finds the center of gravity c on the u-axis using the method described above, as shown in FIG. 8A. Next, the alignment unit 206 fits the error distribution on the u-axis to the v-axis, as shown in FIG. 8B. In the example shown in FIG. 8B, the error distribution is shifted in the positive direction of the v-axis in proportion to the height of the corresponding point, so fitting is performed using only the negative side of the error distribution on the u-axis, in other words, the left side (i.e., the left half) of the center of gravity c in FIG. 8A. The alignment unit 206 sets the value c# corresponding to the center of gravity c of the error distribution on the u-axis after fitting as the alignment parameter for the v-axis.
[0065] Specifically, if the error distribution is assumed to be a normal distribution, the error distribution is expressed by the following equation (4).
[0066]
[0067] In other words, a large number of corresponding point pairs (u r,i , v r,i ), (u t,i , v t,i ) to find the maximum likelihood value of the deviation amount by fitting the error distribution to the above equation μ u , μ v This corresponds to the request for
[0068] Here, we assume that the error distribution is isotropic, i.e., σ u = σ v = σ, then from the distribution on the u axis, a u , μ u and σ can be obtained, and a v and μ v is the only unknown value.
[0069] In addition, the influence of the collapse on the v-axis is thought to be more likely to appear on the positive side of the error distribution. Therefore, it is thought to be effective to use only values below the mode, exclude samples above the mode as outliers, and use only samples on the negative side.
[0070] From the above, v t,i -v r,iBy fitting the graph shown in the following equation (5) using the samples below the mode value and σ calculated on the u axis, we can obtain a highly accurate μ v is required.
[0071]
[0072] Here, the alignment unit 206 t,i -v r,i Without distinguishing between the positive and negative sides of v and μ v Alternatively, the alignment unit 206 may perform fitting to a graph expressed by the following equation (6) without using σ.
[0073]
[0074] In addition, the alignment unit 206 t,i -v r,i When extracting samples on the positive and negative sides of the curve, a value other than the mode may be used, such as a value obtained by adding a predetermined constant to the mode. Note that the fitting described above is a process of aligning curves so that the distance between them is minimized.
[0075] As described above, in the second embodiment, the registration unit 206 can align the reference image and the target image by fitting a first graph showing the distribution of differences in coordinate values along the first axis between a plurality of corresponding point pairs excluding inappropriate corresponding point pairs to a second graph showing the distribution of differences in coordinate values along the second axis between a plurality of corresponding point pairs excluding inappropriate corresponding point pairs. Here, the registration unit 206 can align the reference image and the target image by fitting the second graph to a portion of the first graph where the differences in coordinate values along the first axis between a plurality of corresponding point pairs excluding inappropriate corresponding point pairs are smaller than a specific value.
[0076] As described above, in the second embodiment, the influence of the bias in the error distribution due to the collapse can be suppressed, and the registration parameters can be calculated with high accuracy. Here, the translation model and RANSA have been described as examples, but other registration models and outlier removal methods can also be used.
[0077] 9 is a block diagram showing a schematic configuration of an image processing device 300 according to embodiment 3. The image processing device 300 includes a reference image acquisition unit 101, a target image acquisition unit 102, a corresponding point extraction unit 103, a collapse estimation unit 104, a corresponding point conversion unit 105, a registration unit 106, and a height estimation unit 307.
[0078] The reference image acquisition unit 101, the target image acquisition unit 102, the corresponding point extraction unit 103, the collapse estimation unit 104, the corresponding point conversion unit 105, and the registration unit 106 of the image processing device 300 according to the third embodiment are the same as the reference image acquisition unit 101, the target image acquisition unit 102, the corresponding point extraction unit 103, the collapse estimation unit 104, the corresponding point conversion unit 105, and the registration unit 106 of the image processing device 100 according to the first embodiment. However, the collapse estimation unit 104 in the third embodiment also provides collapse information to the height estimation unit 307. Furthermore, the registration unit 106 provides information indicating the coordinates of the corresponding point pair whose coordinates have been converted by the corresponding point conversion unit 105 to the height estimation unit 307. Note that the processing of the registration unit 206 in the second embodiment may also be performed in the third embodiment.
[0079] The height estimation unit 307 compares the distance in the transformed coordinate system of a selected corresponding point pair, which is one corresponding point pair selected from multiple corresponding point pairs, with the magnitude of the collapse difference vector, thereby estimating the height of the selected corresponding point pair from the length of the line segment used to calculate the collapse difference vector.
[0080] For example, the height estimation unit 307 estimates the height of each corresponding point pair from a collapse difference vector indicated by the collapse information. As described above, the magnitude of the collapse vector indicates the magnitude of collapse per unit height, and therefore the collapse difference vector also indicates the magnitude of the position error per unit height. Therefore, the height estimation unit 307 can estimate the height of the corresponding point pair by comparing the difference in the v-axis values of the corresponding point pair with the magnitude of the collapse difference vector. Then, the height estimation unit 307 generates height information indicating the height of the corresponding point pair.
[0081] As shown in FIG. 10A , some or all of the reference image acquisition unit 101, target image acquisition unit 102, corresponding point extraction unit 103, collapse estimation unit 104, corresponding point conversion unit 105, alignment units 106 and 206, and height estimation unit 307 described above can be configured by a memory 10 and a processor 11 such as a CPU (Central Processing Unit) that executes a program stored in the memory 10. Such a program may be provided via a network or may be provided by being recorded on a recording medium. That is, such a program may be provided, for example, as a computer program product.
[0082] 10(B), some or all of the reference image acquisition unit 101, target image acquisition unit 102, corresponding point extraction unit 103, collapse estimation unit 104, corresponding point conversion unit 105, alignment units 106 and 206, and height estimation unit 307 can be configured by a processing circuit 12 such as a single circuit, a composite circuit, a processor operated by a program, a parallel processor operated by a program, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array). As described above, the reference image acquisition unit 101, target image acquisition unit 102, corresponding point extraction unit 103, collapse estimation unit 104, corresponding point conversion unit 105, alignment units 106 and 206, and height estimation unit 307 can be realized by a processing circuit network.
[0083] 100, 200, 300 Image processing device, 101 Reference image acquisition unit, 102 Target image acquisition unit, 103 Corresponding point extraction unit, 104 Collapse estimation unit, 105 Corresponding point conversion unit, 106, 206 Alignment unit, 307 Height estimation unit.
Claims
1. A corresponding point extraction unit generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of a first image and a second image captured from above so that at least a predetermined space is included; a difference vector calculation unit calculates a difference vector between a first vector indicating the direction and size of a line segment that stands upright in the space and has a known length when projected onto the first image, and a second vector indicating the direction and size of the line segment when projected onto the second image, according to the imaging direction and posture; a coordinate identification unit identifies the positions of the plurality of corresponding point pairs using coordinate values of a coordinate system having a first axis parallel to the difference vector and a second axis perpendicular to the first axis; and a registration unit uses the coordinate values of the plurality of corresponding point pairs to determine inappropriate corresponding point pairs from the plurality of corresponding point pairs, and performs registration of the first image and the second image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs, the alignment unit, in determining that the corresponding points are not an inappropriate pair, sets looser criteria for coordinate values along the first axis than criteria for coordinate values along the second axis.
2. The image processing device according to claim 1, wherein the alignment unit determines that a corresponding point pair determined to be an outlier by an algorithm that determines an outlier using the respective distances in the plurality of corresponding point pairs is an inappropriate corresponding point pair, and wherein the algorithm uses a looser threshold value for the coordinate values of the first axis as a criterion for determining that a point is not an outlier than the threshold value for the coordinate values of the second axis.
3. The image processing device of claim 1, wherein the alignment unit determines one of the plurality of corresponding point pairs to be an inappropriate corresponding point pair when the absolute value of the difference between the coordinate values of the first axis of the corresponding point pair included in the plurality of corresponding point pairs is equal to or greater than a first threshold value, or when the absolute value of the difference between the coordinate values of the second axis of the corresponding point pair is equal to or greater than a second threshold value, and the first threshold value is greater than the second threshold value.
4. The image processing device according to claim 1, characterized in that the alignment unit determines that a corresponding point pair whose coordinates of the corresponding point extracted from the second image do not lie within an ellipse whose center is the coordinates of the corresponding point extracted from the first image, whose major axis is parallel to the first axis, and whose minor axis is parallel to the second axis, is an inappropriate corresponding point pair.
5. An image processing device as claimed in any one of claims 1 to 4, characterized in that the alignment unit aligns the first image and the second image by fitting a first graph showing the distribution of differences in coordinate values of the first axis of the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs to a second graph showing the distribution of differences in coordinate values of the second axis of the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs.
6. The image processing device according to claim 5, wherein the alignment unit aligns the first image and the second image by fitting the second graph to a portion of the first graph where the difference in coordinate values on the first axis of the plurality of corresponding point pairs, excluding the inappropriate corresponding point pairs, is smaller than a specific value.
7. The image processing device according to any one of claims 1 to 4, wherein the alignment unit aligns the first image and the second image using only the coordinate values of the second axis for corresponding point pairs where the difference in coordinate values of the first axis is equal to or greater than a predetermined threshold.
8. An image processing device according to any one of claims 1 to 7, further comprising a height estimation unit that estimates the height of the selected corresponding points pair from the length of the line segment by comparing the distance in the coordinate system of a selected corresponding points pair, which is one corresponding points pair selected from the plurality of corresponding points pairs, with the magnitude of the difference vector.
9. An image processing device according to any one of claims 1 to 8, characterized in that the corresponding point extraction unit sets the feature points extracted from the first image and the feature points extracted from the second image that match the feature points as the corresponding point pairs.
10. An image processing device according to any one of claims 1 to 8, characterized in that the corresponding point extraction unit identifies an area of the first image that best matches the area identified in the second image, and sets each position indicating the identified area as the corresponding point pair.
11. A computer is caused to function as: a corresponding point extraction unit that generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of a first image and a second image captured from above so that at least a predetermined space is included; a difference vector calculation unit that calculates a difference vector that is the difference between a first vector that indicates the direction and size of a line segment that stands upright in the space and has a known length when projected onto the first image, and a second vector that indicates the direction and size of the line segment when projected onto the second image, according to the imaging direction and posture; a coordinate identification unit that identifies the positions of the plurality of corresponding point pairs using coordinate values of a coordinate system having a first axis parallel to the difference vector and a second axis orthogonal to the first axis; and a registration unit that uses the coordinate values of the plurality of corresponding point pairs to determine inappropriate corresponding point pairs from the plurality of corresponding point pairs, and aligns the first image and the second image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs. the alignment unit determines whether the corresponding points are not an inappropriate pair by setting looser criteria for the coordinate values of the first axis than criteria for the coordinate values of the second axis.
12. An image processing method that generates a plurality of corresponding point pairs by extracting a plurality of corresponding points from each of a first image and a second image captured from above so that at least a predetermined space is included; calculates a difference vector, which is the difference between a first vector indicating the direction and size of a line segment that stands upright in the space and has a known length projected onto the first image, and a second vector indicating the direction and size of the line segment projected onto the second image, depending on the imaging direction and posture; specifies the positions of the plurality of corresponding point pairs by coordinate values in a coordinate system having a first axis parallel to the difference vector and a second axis perpendicular to the first axis; determines inappropriate corresponding point pairs from the plurality of corresponding point pairs using the coordinate values of the plurality of corresponding point pairs; and aligns the first image and the second image using the plurality of corresponding point pairs excluding the inappropriate corresponding point pairs, wherein the standard for determining that a corresponding point pair is not inappropriate is set looser for the coordinate values of the first axis than for the coordinate values of the second axis.
Citation Information
Patent Citations
Image processing method, image processing apparatus, and image processing program
JP2014126893A
Earth surface condition grasping method, earth surface condition grasping device, and earth surface condition grasping program
JP2022167343A
Image processing device and image processing method
WO2022259451A1