An automatic precise positioning and correction method based on keyhole decryption of historical images

By constructing an image pyramid and a panoramic camera mathematical model, and combining it with a multi-threshold matching enhancement algorithm, the complex distortion problem of keyhole images was solved, achieving high-precision image positioning and correction, and generating images without panoramic distortion.

CN116823671BActive Publication Date: 2025-11-04CHINESE ACAD OF SURVEYING & MAPPING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310907843.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-11-04
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle the complex distortions in keyhole images, especially nonlinear radiation distortion, rotational distortion, scale distortion, and large temporal phase differences, which makes it impossible to accurately extract corresponding points. Furthermore, there is a lack of mathematical models suitable for panoramic cameras and high-precision image correction methods.

Method used

An image pyramid is constructed based on keyhole images and reference images. Affine transformation relationships are obtained through feature matching. Combined with a panoramic camera mathematical model and a multi-threshold matching enhancement algorithm, high-precision connection points are generated through precise matching layer by layer. Orthorectification is performed using a digital elevation model.

Benefits of technology

It achieves high-precision positioning and correction of keyhole images, can adapt to changes in focal length, overcome complex distortions, improves the robustness of feature matching and the reliability of control points, and generates images without panoramic distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116823671B_ABST
    Figure CN116823671B_ABST
Patent Text Reader

Abstract

The application discloses an automatic precise positioning and correction method based on lock eye decryption historical images, and constructs an image pyramid based on lock eye images and reference images; an affine transformation relationship between the lock eye images and the reference images is obtained at the top layer of the image pyramid through a feature matching step; template matching and multiple threshold matching enhancement are performed on the lock eye images and the reference images based on the affine transformation relationship, and the image pyramid is circulated downward and transmitted until the original resolution to obtain high-precision connection points; a panoramic camera mathematical model is constructed, image posture parameters are generated through the panoramic camera mathematical model and the high-precision connection points; model guided matching is performed through the image posture parameters to generate connection points without the influence of panoramic distortion, and the image posture parameters are regenerated; precise positioning is realized by using the regenerated image posture parameters in combination with a digital elevation model, and precise orthographic correction is performed to generate lock eye orthographic images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of satellite image data processing, and particularly relates to an automatic accurate positioning and correction method based on lock eye decryption historical image. BACKGROUND

[0002] The lock eye image provides strong data support for fields in urgent need of high-resolution historical images, such as historical research, urban planning, and environmental monitoring research. However, due to the lack of geographic reference information and the existence of complex distortion, the lock eye image cannot be directly used for scientific research, and the recovery of geographic reference information and orthorectification work have great difficulties: the current attitude recovery and correction method of the lock eye historical image adopts manual extraction of control point information between the lock eye image and the reference image providing geographic information, but the mathematical model for fitting the panoramic camera is different. Therefore, according to its mathematical model, it can be divided into two categories: non-panoramic camera model and panoramic camera model.

[0003] 1) Panoramic camera model: fully consider the imaging principle of panoramic camera, and propose a mathematical model with strict derivation. Therefore, this kind of model has high precision, but is complex and difficult to implement.

[0004] 2) Non-panoramic camera model: traditional frame model or improved frame model, and rational function polynomial model (RFM) are used to fit the panoramic camera. This kind of model is relatively simple and easy to implement, but its precision is low and can only be used for small range of image blocks.

[0005] Therefore, the current lock eye image technology has the following problems:

[0006] 1. In order to obtain high-resolution images with large coverage, KH-4 series satellites use panoramic cameras, so a strict mathematical model is needed to fit the panoramic camera model, and there is a lack of a panoramic camera mathematical model suitable for lock eye images.

[0007] 2. The current feature matching algorithm cannot resist the nonlinear radiometric distortion, rotation distortion, scale distortion and panoramic distortion between the lock eye image and the reference image providing geographic information, so there is a lack of an algorithm that can extract homonymic points between large temporal difference remote sensing historical images with complex distortion (nonlinear radiometric distortion, rotation distortion, scale distortion, temporal difference and ground object change).

[0008] 3. It is impossible to remove low-reliability control points in ground object change area and improve the number of control points in weak texture and repeated texture area. SUMMARY

[0009] In view of the deficiencies of the prior art, the present application provides an automatic accurate positioning and correction method based on lock eye decryption historical image.

[0010] An automatic precise positioning and correction method based on keyhole decryption history image, comprising the following steps:

[0011] Building an image pyramid based on the keyhole image and the reference image;

[0012] Obtaining the affine transformation relationship between the keyhole image and the reference image at the top layer of the image pyramid through the feature matching step;

[0013] Based on the affine transformation relationship, template matching and multiple threshold matching enhancement are performed on the keyhole image and the reference image, and the image pyramid is continuously circulated downward until the original resolution to obtain high-precision connection points;

[0014] Building a panoramic camera mathematical model, generating image pose parameters through the panoramic camera mathematical model and the high-precision connection points;

[0015] Performing model-guided matching through the image pose parameters to generate high-precision connection points without panoramic distortion images, and regenerating image pose parameters;

[0016] Using the regenerated image pose parameters in combination with a digital elevation model to achieve precise positioning and orthographic correction.

[0017] In some embodiments, the image pyramid is built based on the keyhole image and the reference image, including building a first mean pyramid and a Gaussian scale space;

[0018] The first mean pyramid building step is to use mean filtering to eliminate high-frequency noise in the keyhole image and the reference image, and then sample the filtered image to half of the unfiltered image. The process of continuous filtering and sampling is repeated to obtain images of different resolutions at a predetermined number of layers, thereby building the first mean pyramid;

[0019] The Gaussian scale space building step is to use a predetermined Gaussian multiple and a downsampling multiple to perform Gaussian and downsampling on the keyhole image and the reference image, thereby obtaining a multi-layer resolution image set under the corresponding Gaussian downsampling multiple, and building the Gaussian scale space according to the multi-layer resolution image set.

[0020] In some embodiments, the feature matching step obtains the affine transformation relationship between the keyhole image and the reference image at the top layer of the image pyramid, and the feature matching step includes feature detection, feature description, and feature point matching;

[0021] The feature detection is:

[0022] An image edge in the image pyramid is detected by using a sobel operator, a low-parameter threshold FAST algorithm is used to extract N corner points from the image edge, a feature point set is generated after sorting according to a Harris score, and a KD tree is used to search for a neighbor point in a range of the feature point set , wherein w and h represent the width and height of the image respectively, and the first M points of the feature point set are output as feature points according to the elimination result;

[0023] The feature description includes constructing a multi-direction channel feature, estimating a main direction, and constructing a feature descriptor;

[0024] The multi-direction channel feature is constructed as follows:

[0025] A multi-scale multi-direction Log-Gabor filter is used to extract an image texture feature in the image pyramid, and a multi-direction filter feature is obtained by summing all the scale filter features in each direction, and a formula of the Log-Gabor filter is as follows:

[0026]

[0027] The size of the LG feature is W*H*O, wherein W is the width of the image, H is the height of the image, and O is the number of directions of the Log-Gabor filter, represents an amplitude value calculated from a convolution result of the odd-symmetry and even-symmetry log-Gabor filters on the image;

[0028] The main direction is estimated as follows:

[0029] A multi-scale filter feature in a circular neighborhood of the feature point is extracted, and a norm feature is calculated, the norm feature is Gaussian weighted, the number of fan-shaped regions with repeated areas in the circular neighborhood is determined according to the weighted result, the weighted norm sum of each fan-shaped region with repeated areas is calculated, and the direction of the central axis corresponding to the fan-shaped region with the maximum weighted norm sum is selected as the main direction;

[0030] The feature descriptor is constructed as follows:

[0031] n3 concentric circles with different radii are constructed in the neighborhood of the feature point, each concentric circle is divided into n4 directions from the selected main direction, and a sampling point is determined in each direction, thereby obtaining n3*n4+1 sampling points;

[0032] The sampling points on the concentric circles with different sizes are weighted and summed on the multi-direction filter features in the neighborhood of each concentric circle by using corresponding Gaussian kernels, and an o-dimensional sampling vector is obtained at each sampling point;

[0033] Starting from the selected main direction, the sampling vectors of each sampling point are spliced in a clockwise order to obtain a complete feature descriptor;

[0034] The feature point matching is:

[0035] The feature point matching is performed based on the feature descriptor, the reference image and the lock eye image at each layer of the image pyramid are matched in pairs by a nearest neighbor matching manner and mismatched points are removed, then all matching results are added to the same name point set and mismatched points are removed again to obtain a final matching result and an affine transformation relationship between the reference image and the lock eye image.

[0036] In some embodiments, the lock eye image and the reference image are subjected to template matching and multiple threshold value matching enhancement based on the affine transformation relationship, and the image pyramid is continuously circulated downward until the original resolution to obtain high-precision connection points, and the template matching comprises:

[0037] The reference image is orthophoto resampled based on the affine transformation relationship provided by the feature matching;

[0038] A second mean value pyramid of the resampled orthophoto is reconstructed;

[0039] The second mean value pyramid is subjected to layer-by-layer accurate matching by using a template matching algorithm to obtain high-precision connection points.

[0040] In some embodiments, the second mean value pyramid is subjected to layer-by-layer accurate matching by using a template matching algorithm to obtain high-precision connection points, comprising:

[0041] Step 1: in the top layer of the second mean value pyramid, the corner points of the lock eye image gradient graph are extracted by using a FAST feature detector;

[0042] Step 2: mapping to the orthophoto based on the transformation matrix, and performing local matching by using a multi-directional gradient template algorithm;

[0043] The multi-directional gradient template algorithm is:

[0044] Firstly, the Sobel algorithm is used to obtain the horizontal and vertical gradients, and each direction gradient is interpolated:

[0045]

[0046] wherein, is the direction gradient; is the direction, is the horizontal gradient, is the vertical gradient;

[0047] Step 3: After obtaining the feature points of the top layer, the matching points are mapped to the next layer as connection points according to the resolution difference;

[0048] Step 4: Steps 1-3 are repeatedly executed until the original resolution, thereby obtaining high-precision connection points step by step.

[0049] In some embodiments, the multiple threshold matching enhancement includes: extracting scale difference and direction difference of a local area based on the final matching result, thereby determining reliability;

[0050] Using an affine transformation model, the scale and direction difference in the affine matrix are extracted by the following formula:

[0051]

[0052] Where M represents the affine transformation matrix, T, R, S, and H represent translation, rotation, scale, and skew components, respectively;

[0053] According to the affine transformation matrix, translation, rotation, scale, and skew components, an equation set is obtained:

[0054]

[0055]

[0056]

[0057]

[0058] Where t=0, the scale difference of the local feature points based on the affine transformation relationship is obtained according to the calculation result of the equation set 、 and the direction difference .

[0059] In some embodiments, the panoramic camera mathematical model is constructed, and image pose parameters are generated through the panoramic camera mathematical model and the high-precision connection points, and the construction of the panoramic camera mathematical model includes:

[0060] Rotation angle calculation:

[0061]

[0062] Where, is the horizontal coordinate of the panoramic camera coordinate system, and f is the focal length of the panoramic camera;

[0063] Fitting dynamic panoramic deformation:

[0064]

[0065] Where Represents the satellite's operating speed. Represents satellite altitude. Represents the camera's rotational angular velocity. It is the interior orientation element at time t. Represents dynamic panoramic deformation;

[0066] Eliminate the scale factor based on the rotation and pose matrices:

[0067]

[0068]

[0069]

[0070] in, The rotation matrix representing the gap. Represents the scale factor. The pose matrix of the gap at time t. , These represent the initial exterior orientation elements of the camera. These represent the coefficients that represent the changes in the camera's exterior orientation elements over time. These represent the camera exterior orientation elements at time t, which are obtained by left-multiplying the formula. ,get:

[0071]

[0072] make

[0073] in, Let S represent the element in the i-th row and j-th column of the rotation matrix R. Dividing the first and second rows of the equation by the third row eliminates the scale factor S, resulting in the following equation:

[0074]

[0075]

[0076] Substituting the rotation angle and dynamic panoramic deformation term into the equations yields the panoramic camera coordinates. ;

[0077]

[0078]

[0079] Where X, Y, and Z are ground coordinates. The longitudinal coordinate of the lock eye image is the coordinate, and the mathematical model of the panoramic camera has 14 parameters: 6 linear elements for describing the position of the photographic center P relative to the object space coordinate system 6 angular elements for describing the aerial attitude of the image plane at the moment of photography and dynamic panoramic deformation parameters and focal length f.

[0080] In some embodiments, the image pose parameters are generated by the panoramic camera mathematical model and the high-precision connection points, comprising:

[0081] Obtaining the image plane conjugate point and the object plane conjugate point through the panoramic camera mathematical model;

[0082] Constructing a collinear equation according to the image plane conjugate point and the object plane conjugate point;

[0083] The collinear equation is:

[0084]

[0085]

[0086] After linearizing the collinear equation and calculating the initial value, the image pose parameters are obtained.

[0087] In some embodiments, the high-precision connection points of the panoramic distortion-free image are generated by model-guided matching using the image pose parameters, and the image pose parameters are regenerated, comprising:

[0088] Using the image pose parameters to obtain the corresponding lock eye image points through the Google image points;

[0089] Using the pose parameters to correct the neighborhood image block of the lock eye image point into an orthographic image block, so that the image blocks to be matched are in the same model, thereby eliminating the geometric distortion of the panoramic model and the orthographic model;

[0090] Wherein, the inverse calculation formula of orthographic correction is:

[0091]

[0092]

[0093] Wherein, and correspond to the horizontal coordinate and the longitudinal coordinate of the lock eye image point, respectively.

[0094] Re-performing template matching to obtain the image connection points of the panoramic distortion-free image, and then regenerating the image pose parameters.

[0095] In some embodiments, the regenerated image pose parameters are used in combination with a digital elevation model to achieve accurate positioning and orthorectification, the rectification steps comprising:

[0096] Step one: set a time initial value;

[0097] Step two: obtain exterior orientation elements according to the time initial value, the exterior orientation elements including 、 and ;

[0098] Step three: calculate the image coordinates corresponding to the lock eye image according to the regenerated image pose parameters and the ground coordinates;

[0099] Step four: calculate the time result value according to the image coordinates corresponding to the lock eye image;

[0100] Step five: cyclically execute steps one to four until the correction value of the image coordinates is unchanged, thereby completing the rectification.

[0101] The beneficial effects of the present application are:

[0102] 1. A panoramic camera mathematical model suitable for focal length changes is proposed.

[0103] 2. A robust feature matching algorithm is proposed, which realizes scale, rotation, and radiation invariance. By introducing scale space technology to detect feature scale, extract maximum sector norm feature to detect feature direction, and propose multi-scale and multi-direction filter information to overcome radiation invariance, it can identify the position, rotation, and scale differences between lock eye images with large temporal differences and Google orthographic images.

[0104] 3. A two-stage generalized control feature point extraction algorithm is proposed. The first stage adopts a pyramid multi-level matching strategy. Specifically, a feature matching algorithm is used to extract the scale and rotation differences between the lock eye image and the orthographic reference image, and then a rotation and scale invariant template matching algorithm is used for further accurate matching.

[0105] 4. A multi-threshold matching enhancement technique is proposed to eliminate low-reliability control points in building change areas and improve the number of control points in weak texture and repetitive texture areas. BRIEF DESCRIPTION OF DRAWINGS

[0106] Figure 1 The flowchart of the present application.

[0107] Figure 2 The image pyramid construction schematic diagram.

[0108] Figure 3 The main direction estimation method schematic diagram.

[0109] Figure 4 A schematic diagram for feature descriptor construction method.

[0110] Figure 5 A schematic diagram for feature point matching.

[0111] Figure 6 A schematic diagram for the whole process. DETAILED DESCRIPTION

[0112] Exemplary embodiments of the present application will be described herein below with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be limited by the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the application to those skilled in the art.

[0113] An automatic precise positioning and correction method based on lock eye decryption historical image, as shown in Figure 1 , comprises the following steps:

[0114] S100: constructing an image pyramid based on the lock eye image and the reference image;

[0115] In some embodiments, the step of constructing the image pyramid based on the lock eye image and the reference image comprises constructing a first mean pyramid and a Gaussian scale space.

[0116] The step of constructing the first mean pyramid comprises using mean filtering to eliminate high-frequency noise in the lock eye image and the reference image, then sampling the filtered image to one-half of the unfiltered image, and repeatedly performing filtering and sampling to obtain images of different resolutions in a preset number of layers, thereby constructing the first mean pyramid.

[0117] The step of constructing the Gaussian scale space comprises using a preset Gaussian multiple and a down-sampling multiple to perform Gaussian and down-sampling on the lock eye image and the reference image, thereby obtaining a multi-layer resolution image set corresponding to the Gaussian down-sampling multiple, and constructing the Gaussian scale space based on the multi-layer resolution image set.

[0118] As shown in Figure 2 , the image pyramid is divided into two parts: 1, the mean pyramid, and 2, the Gaussian scale space. Both of them are composed of multiple layers of images of different resolutions, but the mean pyramid is a means to speed up image matching and improve accuracy, while the scale space is a means to resist scale distortion in feature matching.

[0119] The specific construction process is as follows:

[0120] (1) The mean pyramid is constructed using mean down-sampling. First, the original image is down-sampled to one-half of the original image, and then the down-sampled image is repeatedly filtered and sampled to obtain images of different resolutions in a preset number of layers. The mean filter is used to eliminate high frequency noise in the image, and then the filtered result is down-sampled to an image with 1 / 2 size , and the process is repeatedly performed to obtain images with different resolutions of the layer , .

[0121] (2) A Gaussian down-sampling method is used to construct a scale space: the scale space is composed of two multi-resolution image sets: , , n is generally set to 4, and it can be seen that there are n images in each group, and there are 2*n images in the scale space, represent the original image used for feature matching , is obtained by 1.5 times Gaussian down-sampling of , in addition, and are obtained by 2 times Gaussian down-sampling of and , and the relationship between other images can be obtained in the same way.

[0122] S200: Obtain the affine transformation relationship between the lock eye image and the reference image at the top layer of the image pyramid through a feature matching step;

[0123] Wherein, the FAST feature detection algorithm has a serious key point clustering phenomenon, which will reduce the accuracy of matching and increase redundant information, in order to effectively resist nonlinear radiation distortion and obtain good repeatability and unique feature points, a uniform distribution FAST corner detection algorithm for extracting image gradient (referred to as: GU-FAST algorithm) is proposed.

[0124] In some embodiments, the affine transformation relationship between the lock eye image and the reference image at the top layer of the image pyramid is obtained through a feature matching step, and the feature matching step includes feature detection, feature description and feature point matching;

[0125] The feature detection is:

[0126] The sobel operator is used to detect the image edge in the image pyramid, the FAST algorithm with low parameter threshold is used to extract N corner points from the image edge, and the feature point set is generated after sorting according to the Harris score. The KD tree is used to search for the nearest neighbor points in the range of , where w and h represent the width and height of the image respectively, and the first M points of the feature point set are output as feature points according to the elimination result;

[0127] The feature description includes constructing a multi-directional channel feature, a principal direction estimation, and constructing a feature descriptor;

[0128] The constructing a multi-directional channel feature is:

[0129] A multi-scale multi-directional Log-Gabor filter is used to extract image texture features in an image pyramid, and the filtered features of all scales in each direction are summed to obtain multi-directional filtered features, and the formula of the Log-Gabor filter is:

[0130]

[0131] The size of the LG feature is W*H*O, where W is the width of the image, H is the height of the image, and O is the number of directions of the Log-Gabor filter, represents the amplitude calculated from the convolution results of the odd-symmetry and even-symmetry log-Gabor filters on the image;

[0132] Affected by nonlinear radiation distortion, the principal direction estimation algorithm based on traditional image features (image gray scale, image gradient) fails. Therefore, a principal direction estimation algorithm based on maximum fan norm features is proposed.

[0133] The principal direction estimation is:

[0134] Multi-scale filtered features in the circular neighborhood of the feature point are extracted, and the norm features are calculated. The norm features are Gaussian weighted, the number of fan-shaped regions with repeated areas in the circular neighborhood is determined according to the weighted results, the weighted norm sum of each fan-shaped region with repeated areas is calculated, and the direction of the central axis corresponding to the fan-shaped region with the maximum weighted norm sum is selected as the principal direction.

[0135] As shown in Figure 3 , the specific algorithm process is as follows:

[0136] (1) Extract multi-scale filtered features in the circular neighborhood of the feature point, and calculate the norm features thereof;

[0137] (2) Gaussian weight the norm features;

[0138] (3) Determine a plurality of fan-shaped regions with repeated areas in the feature neighborhood: start from 0 degrees, determine a fan-shaped region every degrees, and the angle of each fan-shaped region is degrees;

[0139] (4) Calculate the weighted norm sum of each fan-shaped region;

[0140] (5) Take the direction of the central axis corresponding to the fan-shaped region with the maximum weighted norm sum as the principal direction;

[0141] In addition, in order to avoid repeated calculation, the circular neighborhood is divided into small fan-shaped blocks by degrees and the norm sum of each block is calculated, and then the sum of the large fan-shaped block is synthesized quickly by using the sum result of the small blocks. In order to achieve more robust rotation invariance, the second largest fan-shaped region larger than 70% of the maximum value is called the second principal direction.

[0142] The feature descriptor is constructed: the multi-directional filtered features in the neighborhood of the feature point are obtained, and the initial direction of the filter is set to the principal direction by adjusting the order. In addition, in order to extract a robust feature descriptor, multi-scale and multi-directional sampling points are set in the neighborhood of the feature point to extract the local structural features at different positions in the neighborhood of the feature point.

[0143] The feature descriptor is constructed as follows:

[0144] n3 concentric circles of different radii are constructed in the neighborhood of the feature point, each concentric circle is divided into n4 directions starting from the selected principal direction, and a sampling point is determined in each direction, thereby obtaining n3*n4+1 sampling points;

[0145] The multi-directional filtered features in the neighborhood of each concentric circle are weighted and summed by using the corresponding Gaussian kernel of the sampling point on the concentric circle of different sizes, and an o-dimensional sampling vector is obtained at each sampling point;

[0146] Starting from the selected principal direction, the sampling vectors of each sampling point are spliced in a clockwise order, thereby obtaining a complete feature descriptor;

[0147] The feature point matching is as follows:

[0148] The feature point matching is performed based on the feature descriptor, the reference image and the lock eye image at each layer of the image pyramid are matched two by two by using the nearest neighbor matching method and the mismatched points are removed, then all the matching results are added to the same name point set and the mismatched points are removed again, thereby obtaining the final matching result and the affine transformation relationship between the reference image and the lock eye image.

[0149] In the feature point matching stage, firstly, the nearest neighbor matching method is used to match each layer of the reference image and the pyramid of the target image two by two and remove the mismatched points, then all the matching results are added to the final matching point set and the mismatched points are removed again (random sample consensus algorithm: RANSAC), thereby obtaining the final matching points and the transformation relationship between the images.

[0150] As Figure 5 ​As shown, the first column of circles represents the feature point set of each layer of the image pyramid of the lock eye image, the second column of circles represents the feature point set of each layer of the image pyramid of the reference image, and the connection of the circles in the first two columns represents the matching (nearest neighbor matching) of the point sets. Then, the matching result is added to the same point set, and the mismatched points are removed again to obtain the final matching result.

[0151] S300: performing template matching and multi-threshold matching enhancement on the lock eye image and the reference image based on the affine transformation relationship, and continuously circulating downward through the image pyramid until the original resolution to obtain high-precision connection points;

[0152] In some embodiments, the template matching and multi-threshold matching enhancement on the lock eye image and the reference image based on the affine transformation relationship, and continuously circulating downward through the image pyramid until the original resolution to obtain high-precision connection points, the template matching includes:

[0153] Resampling the orthographic image of the reference image based on the affine transformation relationship provided by the feature matching;

[0154] Reconstructing a second mean pyramid of the resampled orthographic image;

[0155] Using a template matching algorithm to perform layer-by-layer accurate matching on the second mean pyramid to obtain high-precision connection points.

[0156] In some embodiments, the template matching algorithm used to perform layer-by-layer accurate matching on the second mean pyramid to obtain high-precision connection points includes:

[0157] Step 1: using a FAST feature detector to extract the corner points of the lock eye image gradient map at the top layer of the second mean pyramid;

[0158] Step 2: mapping to the orthographic image based on the transformation matrix, and performing local matching using a multi-directional gradient template algorithm;

[0159] The multi-directional gradient template algorithm is:

[0160] First, use the Sobel algorithm to obtain the horizontal and vertical gradients, and interpolate to obtain the gradient in each direction:

[0161]

[0162] wherein, is the directional gradient; is the direction, is the horizontal gradient, is the gradient in vertical direction; where we use the default parameter, i.e. 0 to 360 degrees are divided into 9 directions, so the size of the template feature of each point is W*W*9, W is the neighborhood size of the point template. In order to further improve the robustness, the template feature is further processed by 3D Gaussian kernel convolution and normalization in the direction gradient dimension.

[0163] Step 3: After obtaining the feature points of the top layer, the matching points are mapped to the next layer as connection points according to the resolution difference;

[0164] Step 4: Repeat steps 1-3 until the original resolution, so as to gradually obtain high-precision connection points.

[0165] Based on this, a multi-threshold matching enhancement strategy based on matching diffusion and local consistency is proposed; the strategy can effectively eliminate low-reliability homonym points in local feature change areas, and improve the number of homonym points in repeated and weak texture areas. The basic principle is: based on the assumption that the image eliminates large angle and scale difference, the texture, angle and scale information of the local area are extracted through joint large-scale local homonym point pairs to judge the reliability of the matching result.

[0166] In some embodiments, the multi-threshold matching enhancement includes: based on the final matching result, the scale difference and direction difference of the local area are extracted to determine the reliability;

[0167] Wherein, the detailed process is described as follows:

[0168] First, in order to alleviate the problem of few matching points in weak texture areas, search for potential matching points in the feature neighborhood and add them to the matching point set. Then, jointly use the point set as a whole, which can use larger range of texture structure information.

[0169] Second, in order to alleviate the influence of large-scale geometric distortion on matching accuracy, the matching result of the point set is allowed to change within a specific transformation model.

[0170] Finally, based on the matching result, the scale difference and direction difference of the local area are extracted to determine the reliability.

[0171] Using affine transformation model, the scale and direction difference in affine matrix are extracted by the following formula:

[0172]

[0173] Where M represents the affine transformation matrix, T, R, S, H represent translation, rotation, scale, and shear components, respectively;

[0174] According to the affine transformation matrix, translation, rotation, scale and shear components, the equation group is obtained:

[0175]

[0176]

[0177]

[0178]

[0179] where t=0, the scale difference of the local feature points based on the affine transformation relationship is obtained according to the calculation results of the equation group , and the direction difference .

[0180] S400: Construct a panoramic camera mathematical model, and generate image pose parameters through the panoramic camera mathematical model and the high-precision connection point;

[0181] In some embodiments, the construction of the panoramic camera mathematical model and the generation of the image pose parameters through the panoramic camera mathematical model and the high-precision connection point include:

[0182] Rotation angle calculation:

[0183]

[0184] wherein, is the horizontal coordinate of the panoramic camera coordinate system, and f is the focal length of the panoramic camera;

[0185] Fitting dynamic panoramic deformation:

[0186]

[0187] wherein represents the running speed of the satellite, represents the satellite height, represents the rotation angular velocity of the camera, is the internal orientation element at time t, represents the dynamic panoramic deformation term;

[0188] Eliminate the scale factor according to the rotation matrix and the attitude matrix:

[0189]

[0190]

[0191]

[0192] wherein, represents the rotation matrix of the gap, represents the scale factor, the pose matrix representing the pose of the slit at time t, the initial exterior orientation elements of the camera, respectively, the coefficients of the time-varying camera exterior orientation elements, respectively, the camera exterior orientation elements at time t, obtained by multiplying the left side of the equation by

[0193]

[0194] Let

[0195] wherein, the element in the i-th row and j-th column of the rotation matrix R, dividing the first and second rows of the equation by the third row can eliminate the scale factor S, and the following equation is obtained:

[0196]

[0197]

[0198] the rotation angle and the dynamic panoramic deformation term are brought into to obtain the panoramic camera coordinates ;

[0199]

[0200]

[0201] wherein, X, Y, and Z are ground coordinates, the longitudinal coordinate of the keyhole image coordinates, the constructed panoramic camera mathematical model has 14 parameters: of which 6 are linear elements, used to describe the position of the photographic center P relative to the object space coordinate system , and 6 are angular elements, used to describe the aerial pose of the image plane at the photographic moment , and the dynamic panoramic deformation parameter , and the focal length f.

[0202] S500: performing model-guided matching through the image pose parameters to generate high-precision connection points of images without panoramic distortion, and re-generating the image pose parameters;

[0203] In some embodiments, the generating the image pose parameters through the panoramic camera mathematical model and the high-precision connection points comprises:

[0204] obtaining the image-side conjugate points and the object-side conjugate points through the panoramic camera mathematical model;

[0205] constructing a collinear equation according to the image-side conjugate points and the object-side conjugate points; ​​

[0206] The collinear equation is:

[0207]

[0208]

[0209] After linearization of the collinear equation and calculation of initial values, the image pose parameters are obtained.

[0210] When the mathematical model of the panoramic camera is determined, a certain number of image points and object points in the image coverage can be used. In this paper, image matching determines the homonymic points between a large number of lock eye images and reference images (which can be Google Earth, Landsat, etc. images with geographic information): Since the reference image has geographic information, the connection point corresponding to the ground coordinates (X, Y) can be obtained by the geographic information of the reference image. At this time, we get the image conjugate point and the object conjugate point (X, Y, Z). At this time, according to the collinear condition equation, the exterior orientation elements of the image are solved, which is called the space resection of a single image. The collinear equation is as shown above. Since the mathematical model has 14 unknowns, and each pair of image and object conjugate points can list 2 equations, seven or more control points with known ground coordinates are needed to solve the pose parameters. The more the number of control points and the more uniform the distribution, the higher the accuracy and reliability of the solved pose parameters. When the number of observation equations is greater than the number of unknowns, the least square adjustment method is needed to solve it. In addition, since the collinear equation is a nonlinear equation, it needs to be linearized and provide the initial value of the solving parameter.

[0211] S600: Using the regenerated image pose parameters to realize accurate positioning and orthographic correction in combination with the digital elevation model.

[0212] In some embodiments, the model-guided matching through the image pose parameters to generate high-precision connection points without panoramic distortion images, and the regeneration of image pose parameters, include:

[0213] Using the image pose parameters to obtain the corresponding lock eye image points from the Google image points;

[0214] Using the pose parameters to correct the neighborhood image block of the lock eye image point into an orthographic image block, so that the image blocks to be matched are in the same model, thereby eliminating the geometric distortion of the panoramic model and the orthographic model;

[0215] The inverse calculation formula of orthographic correction is:

[0216]

[0217]

[0218] wherein, and correspond to the horizontal and vertical coordinates of the lock eye image point respectively;

[0219] re-performing template matching to obtain an image connection point of the image without panoramic distortion, and then re-generating the image pose parameter.

[0220] In some embodiments, the precise positioning and orthographic correction are achieved by using the re-generated image pose parameter in combination with a digital elevation model, and the correction step comprises:

[0221] Step one: setting a time initial value;

[0222] Step two: obtaining exterior orientation elements according to the time initial value, the exterior orientation elements including , and ;

[0223] Step three: calculating the image coordinate corresponding to the lock eye image according to the re-generated image pose parameter and the ground coordinate;

[0224] Step four: calculating a time result value according to the image coordinate corresponding to the lock eye image;

[0225] Step five: cyclically executing steps one to four until the correction value of the image coordinate is unchanged, thereby completing the correction.

[0226] wherein, after obtaining the pose parameter of the panoramic camera of the image, an inverse algorithm is used to generate an orthographic image of any resolution, that is, since the resolution of the orthographic image is consistent, each pixel has a corresponding ground coordinate, therefore, by determining the image coordinate (x, y) on the lock eye image corresponding to the ground coordinate (X, Y, Z) of the pixel on the orthographic image to be solved, the elevation value Z is also obtained by interpolation in the digital elevation model, in addition, if the coordinate is not on the integer pixel, the pixel value is interpolated by the bilinear interpolation method, and finally the value is assigned to the orthographic image. The relationship between the ground point coordinate (X, Y, Z) and the panoramic coordinate is determined by the formula of the inverse algorithm, thereby completing the precise positioning. In addition, since the exterior orientation elements such as and are related to the time t, and t is related to x, therefore, iteration is also required, and the process is as follows:

[0227] (1) the initial value t0 is set to 0.5,

[0228] (2) then the exterior orientation elements are obtained: the parameters with t subscript,

[0229] (3) According to the attitude parameter and X, Y, Z, calculate x, y,

[0230] (4) Then according to x, y, calculate time t1, re-perform (1)-(3) until the correction value of x, y is basically unchanged, thereby completing the correction.

[0231] In summary, as shown in the figure, the overall process of the scheme is as follows: Figure 6

[0232] Construct the image pyramid of the lock eye and the reference image;

[0233] Use feature matching at the top layer of the image pyramid to obtain the rough affine transformation relationship between the images, and use it for subsequent template matching layer by layer to obtain the connection point; (that is, use a feature matching algorithm at the top layer to obtain the rough transformation relationship between the images; then based on the transformation relationship, template matching is performed on the two images, and is continuously passed down to the original resolution through the image pyramid, so as to obtain a high-precision connection point;)

[0234] Finally, use the image attitude parameter to correct the panoramic distortion to re-perform accurate matching to improve the accuracy of the control points, and generate an orthographic image.

[0235] Among them, according to the attitude parameter, we can directly combine the digital elevation model to obtain the ground coordinates corresponding to each point on the image, so as to realize accurate positioning, and at the same time generate an orthographic image through orthographic correction.

[0236] The above is only the preferred embodiment of the present application, it should be pointed out that for the skilled in the art without departing from the technical solutions of the present application, several modified and improved technical solutions should also be considered to fall within the scope of the present application.​

Claims

1. A method for automatic and precise positioning and correction based on the decryption of historical images of keyholes, characterized by: The method comprises the following steps: constructing an image pyramid based on the keyhole image and the reference image; obtaining an affine transformation relationship between the keyhole image and the reference image at the top layer of the image pyramid through a feature matching step, the feature matching step comprising feature detection, feature description and feature point matching; the feature detection is: An image edge in an image pyramid is detected by using a sobel operator, a low-parameter threshold FAST algorithm is used to extract N corner points from the image edge, a feature point set is generated after sorting according to Harris scores, and a KD tree is used to search for a near neighbor point in a range of the feature point set , wherein w and h respectively represent a width and a height of the image, and the first m points of the feature point set are output as feature points according to a result of the elimination; and the feature point set is used as an input of a feature matching algorithm to search for a feature point set of a target image. the feature description comprises constructing a multi-directional channel feature, main direction estimation and constructing a feature descriptor; the constructing a multi-directional channel feature is: a multi-scale multi-directional Log-Gabor filter is used to extract image texture features in the image pyramid, and the filtered features of all scales in each direction are summed to obtain a multi-directional filtered feature, and the formula of the Log-Gabor filter is: , The size of the LG feature is where w is the width of the image, h is the height of the image, o1 is the number of directions of the Log-Gabor filter, denotes the amplitude calculated from the convolution result of the image by the odd-symmetric and even-symmetric log-Gabor filter pair. the main direction estimation is: multi-scale filtered features in a circular neighborhood of the feature point are extracted, and a norm feature is calculated, the norm feature is Gaussian weighted, the number of fan-shaped regions with repeated areas in the circular neighborhood is determined according to the weighted result, the weighted norm sum of each fan-shaped region with repeated areas is calculated, and the direction of the central axis corresponding to the fan-shaped region with the maximum weighted norm sum is selected as the main direction; the constructing a feature descriptor is: n3 different radius concentric circles are constructed in the neighborhood of the feature point, and each of the concentric circles is divided into n4 directions from the selected main direction, and a sampling point is determined in each direction, thereby obtaining sampling points; sampling points on concentric circles of different sizes are used to weight and sum the multi-directional filtered features of each concentric circular neighborhood using corresponding Gaussian kernels, and an o2-dimensional sampling vector is obtained at each sampling point; starting from the selected main direction, the sampling vectors of each sampling point are spliced in a clockwise order, thereby obtaining a complete feature descriptor; the feature point matching is: feature point matching is performed based on the feature descriptor, each reference image and keyhole image at each layer of the image pyramid is matched in pairs through a nearest neighbor matching method and mismatched points are removed, then all matching results are added to the same point set and mismatched points are removed again to obtain a final matching result and an affine transformation relationship between the reference image and the keyhole image; template matching and multi-threshold matching enhancement are performed on the keyhole image and the reference image based on the affine transformation relationship, and the image pyramid is continuously circulated downward until the original resolution is obtained, and a high-precision connection point is obtained; a panoramic camera mathematical model is constructed, and image pose parameters are generated through the panoramic camera mathematical model and the high-precision connection point; model-guided matching is performed through the image pose parameters to generate a high-precision connection point without panoramic distortion, and image pose parameters are regenerated; precise positioning and orthographic correction are realized by using the regenerated image pose parameters in combination with a digital elevation model.

2. The method of claim 1, wherein: The image pyramid is constructed based on the keyhole image and the reference image, comprising constructing a first mean pyramid and a Gaussian scale space; the first mean pyramid construction step is: using mean filtering to eliminate high-frequency noise in the keyhole image and the reference image, then sampling the filtered image to one-half of the unfiltered image, continuously circulating the filtering and sampling process, thereby obtaining images of different resolutions of a preset number of layers, thereby constructing the first mean pyramid; The step of constructing the Gaussian scale space is: adopting a preset Gaussian multiple and a down-sampling multiple to perform Gaussian and down-sampling on the lock eye image and the reference image, so as to obtain a multi-layer resolution image set under a corresponding Gaussian down-sampling multiple, and constructing the Gaussian scale space according to the multi-layer resolution image set.

3. The method of claim 2, wherein: The template matching and the multiple threshold matching enhancement are performed on the lock eye image and the reference image based on the affine transformation relationship, and the image pyramid is continuously circulated downward until the original resolution, so as to obtain a high-precision connection point. The template matching comprises: Resampling the orthographic image of the reference image based on the affine transformation relationship provided by the feature matching; Reconstructing a second mean value pyramid of the resampled orthographic image; Performing layer-by-layer accurate matching on the second mean value pyramid by using a template matching algorithm to obtain the high-precision connection point.

4. The method of claim 3, wherein: The step of performing layer-by-layer accurate matching on the second mean value pyramid by using a template matching algorithm to obtain the high-precision connection point comprises: Step 1: using a FAST feature detector to extract corner points of a gradient graph of the lock eye image at a top layer of the second mean value pyramid; Step 2: mapping to the orthographic image based on a transformation matrix, and performing local matching by using a multi-directional gradient template algorithm; The multi-directional gradient template algorithm is: Firstly, using a Sobel algorithm to obtain horizontal and vertical gradients, and interpolating to obtain a gradient in each direction: wherein, is the directional gradient; is the direction, is the gradient in the horizontal direction, is the gradient in the vertical direction; Step 3: after obtaining the feature points of the top layer, mapping the matching points to a next layer as the connection points according to a resolution difference; Step 4: repeating steps 1-3 until the original resolution, so as to gradually obtain the high-precision connection point.

5. The method of claim 4, wherein: The multiple threshold matching enhancement comprises: extracting scale difference and direction difference of a local region based on the final matching result, so as to determine reliability; Using an affine transformation model, extracting scale difference and direction difference in the affine matrix by the following formula: Wherein M1 represents the affine transformation matrix, a, b, c, d, o, p represent elements of the M1 matrix respectively, T1, R1, S1, H1 represent translation, rotation, scale, and skew components respectively; Obtaining an equation group according to the affine transformation matrix, the translation, the rotation, the scale and the skew component: , , , , where t1=0, the scale difference of the local feature points based on the affine transformation relationship is obtained according to the calculation result of the equation set , and the direction difference .

6. The method of claim 5, wherein: The step of constructing the mathematical model of the panoramic camera, generating image pose parameters by the mathematical model of the panoramic camera and the high-precision connection point, comprises: Rotation angle calculation: wherein, is the horizontal coordinate of the panoramic camera coordinate system, and f is the focal length of the panoramic camera. Fitting dynamic panoramic deformation: wherein represents a running speed of the satellite, H1 represents a satellite height, represents a rotational angular velocity of the camera, is an exterior orientation element at time t, represents a dynamic panoramic deformation term; Eliminating scale factors according to the rotation matrix and the pose matrix: , , , where, a rotation matrix representing the slit, a scale factor, a pose matrix representing the slit at time t, , respectively represent the initial extrinsic elements of the cameras, respectively represent the coefficients of the time variation of the extrinsic elements of the cameras, respectively represent the extrinsic elements of the cameras at time t, obtained by left-multiplying the equation by , we obtain: Let wherein Representing the element in the i-th row and j-th column of the camera rotation matrix, dividing the first and second rows of the equation by the third row, respectively, eliminates the scale factor s, resulting in the following: , Bringing the rotation angle and the dynamic panoramic deformation term into the resulting panoramic camera coordinates ; , where X, Y, Z are ground coordinates, The longitudinal coordinate of the lock eye image, the mathematical model of the panoramic camera constructed has 14 parameters: of which 6 are linear elements, used to describe the coefficients of the change of the position of the photographic center P relative to the object space coordinate system with time , and 6 are angular elements, used to describe the aerial attitude of the image plane at the moment of photography , and dynamic panoramic deformation parameters , and focal length f.

7. The method of claim 6, wherein: The step of generating image pose parameters by the mathematical model of the panoramic camera and the high-precision connection point comprises: Obtaining image-side conjugate points and object-side conjugate points by the mathematical model of the panoramic camera; Constructing a collinear equation according to the image-side conjugate points and the object-side conjugate points; The collinear equation is: , , After linearizing the collinear equation and setting an initial value, the image pose parameters are obtained.

8. The method of claim 7, wherein: The step of performing model-guided matching by the image pose parameters to generate a high-precision connection point of a panoramic distortion-free image, and regenerating the image pose parameters comprises: Obtaining corresponding lock eye image points by the image pose parameters through Google image points; The neighborhood image block of the lock eye image point is rectified into an orthographic image block by using the attitude parameter, so that the image blocks to be matched are in the same model, thereby eliminating the geometric distortion of the panoramic model and the orthographic model; The inverse calculation formula of the orthographic rectification is: wherein, and correspond to the abscissa and ordinate of the image point of the geometric distortion-eliminated keyhole, respectively; The template matching is performed again to obtain the image connection point of the image without panoramic distortion, and then the image attitude parameter is regenerated.

9. The method of claim 8, wherein: The regenerated image attitude parameter is combined with the digital elevation model to realize accurate positioning and orthographic rectification. The rectification steps include: Step one: setting a time initial value; Step two: obtaining exterior orientation elements based on the time initial value, the exterior orientation elements including , and ; Step three: calculating the image coordinate corresponding to the lock eye image according to the regenerated image attitude parameter and the ground coordinate; Step four: calculating the time result value according to the image coordinate corresponding to the lock eye image; Step five: cyclically executing steps one to four until the correction value of the image coordinate is unchanged, thereby completing the rectification.

Citation Information

Patent Citations

  • Three-dimensional reconstruction-based unmanned aerial vehicle image stitching method and system

    CN108765298A

  • Multi-source remote sensing image feature matching method based on directional phase consistency

    CN109523585A