Multi-scene adaptive cross-correlation peak ratio registration verification and bicentric error evaluation method

By adopting a multi-scenario adaptive cross-correlation peak ratio registration verification method, and using homography matrix and dual center points to evaluate image registration error, the usability and accuracy evaluation problem of image registration under true value conditions is solved, and efficient and accurate registration result judgment is achieved. It is applicable to fields such as autonomous driving, UAV navigation and remote sensing mapping.

CN121921348BActive Publication Date: 2026-05-29NORTHWEST A & F UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWEST A & F UNIV
Filing Date
2026-03-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies cannot effectively evaluate the accuracy of image registration without true values. In particular, in multimodal and complex scenarios, traditional methods cannot guarantee the usability and accuracy of registration results and lack engineering practicality.

Method used

A multi-scenario adaptive cross-correlation peak ratio registration verification method is adopted. By calculating the homography matrix and performing geometric transformation, structural consistency features are extracted, cross-correlation is performed, the amplitude ratio relationship between the main peak and the secondary peak is determined, and the registration error is evaluated using dual center points.

Benefits of technology

It enables automated verification and accuracy evaluation of image registration results under true value-free conditions, reduces false alarm rate and missed alarm rate, improves the accuracy of judgment and computational efficiency, and is applicable to fields such as autonomous driving, UAV navigation and remote sensing mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921348B_ABST
    Figure CN121921348B_ABST
Patent Text Reader

Abstract

The technical scheme belongs to the technical field of remote sensing image registration, and discloses a multi-scene adaptive cross-correlation peak ratio registration verification and double-center error evaluation method.The innovative core of the technical scheme is to construct a registration result judgment and precision evaluation method based on cross-correlation operation.Under any data set (whether a multi-modal image data set or a view angle change image data set) or in any scene, a homography matrix is calculated by using the registration result, an affine transformation is performed on a real-time image, cross-correlation operation is performed on the affine transformed real-time image and a reference image, whether the registration is successful or not is judged by using a correlation coefficient matrix, and the registration error is evaluated by using a method of two center points.The method has the characteristics of low complexity, high effectiveness, short operation time and the like, and can be successfully implemented in engineering application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technical solution belongs to the field of remote sensing image registration technology, and in particular relates to a method and system for multi-scene adaptation cross-correlation peak ratio registration verification and dual-center error evaluation. Background Technology

[0002] With the continuous development of low-altitude economy and autonomous driving, platforms such as drones and automobiles are using visual navigation technology as a key technology for composite guidance, thus enabling the application and implementation of image registration technology. However, existing research on image registration focuses more on improving the robustness and accuracy of registration methods, lacking a judgment on whether the registration results are usable. Traditional image registration methods, such as POSGIFT ("POS-GIFT: Ageometric and intensity-invariant feature transformation for multimodal images" (Information Fusion, 2023.10-2027.)) and HOWP ("Histogram of the orientation of the weighted phase descriptor for multi-modal remote sensing image matching" (ISPRS Journal of Photogrammetry and Remote Sensing, 2022.12.018"), emphasize the development of algorithms towards multimodal approaches, achieving excellent results on multimodal remote sensing image datasets. These non-neural network methods are also easier to deploy on devices such as drones and vehicle platforms. However, when faced with image pairs exhibiting irregular distortions, such as geometric transformations or perspective transformations, the registration success rate and accuracy drop sharply. These traditional methods all overlook one crucial point: there is no perfect registration method. Rather than focusing on improving registration success rate and accuracy, the primary concern in autonomous navigation and positioning is whether the registration results can be used for subsequent data processing. Registration failure can introduce significant deviations into the positioning results.

[0003] In recent years, a number of excellent deep learning-based image registration algorithms have emerged, such as SuperGlue ("SuperGlue: Learning Feature Matching with Graph Neural Networks" (arXiv: 1911.11763v2"), DUST3R ("DUSt3R: Geometric 3D Vision Made Easy" (arXiv: 2312.14132) and VDFT ("VDFT: Robust feature matching of a continent and ground images using viewpoint-invariant defor..."). While methods like "Mable Feature Transformation" (ISPRS Journal of Photogrammetry and Remote Sensing, 2024.09.016) utilize large amounts of irregularly distorted data for training, maintaining good registration results and accuracy even with severely distorted image pairs, the training dataset is always finite. These methods exhibit severe registration failures when dealing with multimodal datasets such as far-infrared or SAR images. This demonstrates that both traditional feature registration algorithms and data-driven registration algorithms have limitations and cannot guarantee successful registration in all datasets and scenarios. Currently, most papers focus on improving registration robustness, success rate, and accuracy within specific datasets, and statistical metrics such as matching success rate and matching accuracy are presented on a limited number of datasets. Whether in airborne or vehicle-mounted autonomous navigation and positioning technologies, or in technologies that combine registration results for image fusion or change detection, it is desirable for registration technology to not only register image pairs but also provide a digital result indicating whether the registration was successful and the accuracy of the registration. Only such a result can be "understood" by the computer.

[0004] Furthermore, aside from fields like autonomous driving and visual navigation, when algorithms face large-scale image registration tasks, there is no ground truth to measure the registration results. Relying solely on manual evaluation of the matching results cannot guarantee the correctness of the judgment, nor can it provide a quantifiable registration accuracy. Currently, no articles provide a method for evaluating image registration accuracy without ground truth; all reports provide the accuracy of the registration results under known ground truth conditions, which lacks practical engineering applicability.

[0005] Finally, considering the practicality required for engineering applications, the judgment method cannot be too complex, and the computation time cannot be too long. These all place stringent requirements on the registration result judgment method.

[0006] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:

[0007] When algorithms face large-scale image registration tasks, there is no ground truth to measure the registration results. Relying solely on manual evaluation of the matching results cannot guarantee the correctness of the judgment, nor can it provide a quantifiable registration accuracy. Currently, no articles provide a method for evaluating image registration accuracy without ground truth; all articles provide the accuracy of the registration results under known ground truth conditions, which is not practical for engineering applications. Summary of the Invention

[0008] To address the problems existing in the current technology, this technical solution provides a method for multi-scenario adapted cross-correlation peak ratio registration verification and dual-center error evaluation.

[0009] This technical solution is implemented as follows: A multi-scenario adapted cross-correlation peak ratio registration verification and dual-center error evaluation method includes:

[0010] Step 1: Define the image pair to be processed as a real-time image and a reference image, and calculate the homography matrix between them based on the image feature matching results;

[0011] Step 2: Use the homography matrix to perform a geometric transformation on the real-time image to make the real-time image and the reference image consistent in geometric structure, so as to eliminate the effects of rotation, scale changes and geometric distortion.

[0012] Step 3: Extract structural consistency features from the geometrically consistent real-time graph and the reference graph, and perform cross-correlation operation based on the structural consistency features to obtain the cross-correlation matrix;

[0013] Step 4: Determine the position of the main peak in the cross-correlation coefficient matrix, and divide the cross-correlation coefficient matrix into eight non-overlapping regions extending along the horizontal, vertical and diagonal directions with the main peak position as the center. Select the local maximum value with the largest amplitude in each region as the secondary peak.

[0014] Step 5: Verify the consistency of the registration results based on the proportional relationship between the amplitude of the main peak and the amplitudes of each secondary peak.

[0015] Step 6: When the verification result meets the preset conditions, determine the coordinates of the first center point based on the position of the main peak of cross-correlation, and at the same time, map the center position of the real-time graph to the reference graph based on the homography matrix to determine the coordinates of the second center point.

[0016] Step 7: Calculate the Euclidean distance between the coordinates of the first center point and the coordinates of the second center point, and use this Euclidean distance as the evaluation result of the registration error.

[0017] Furthermore, the homography matrix is ​​calculated from the pixel correspondence between the real-time image and the reference image, and the geometric transformation is an affine transformation performed on the real-time image based on the homography matrix, so that the two images maintain consistency in spatial structure.

[0018] Furthermore, the structural consistency features are obtained through phase consistency feature extraction to suppress the influence of multimodal imaging differences on the cross-correlation results and enhance the significance of the main cross-correlation peak relative to the side peaks.

[0019] Another objective of this technical solution is to provide an image registration and verification method based on cross-correlation primary and secondary peak spatial constraints, characterized by comprising:

[0020] Perform cross-correlation on the two images after geometric consistency processing to obtain the cross-correlation matrix;

[0021] The main peak with the largest amplitude is determined in the cross-correlation matrix, and the cross-correlation matrix is ​​divided into multiple regions in eight directions with the location of the main peak as the center.

[0022] In each of the aforementioned regions, a local maximum with the largest amplitude is determined as the secondary peak;

[0023] Based on the proportional relationship between the amplitude of the main peak and the amplitudes of each secondary peak, a peak ratio criterion is constructed, and the validity of the image registration result is determined based on the peak ratio criterion.

[0024] Furthermore, the region division method is to divide the cross-correlation matrix into eight triangular regions along the horizontal, vertical and two diagonal directions with the main peak position as the origin, and the regions do not overlap with each other.

[0025] Furthermore, the peak ratio criterion is implemented by comparing whether the ratio of the amplitude of the main peak to the amplitude of the secondary peak exceeds a preset threshold, so as to reduce the probability of misjudgment.

[0026] Another objective of this technical solution is to provide a multi-scenario adaptable cross-correlation peak ratio registration verification and dual-center error evaluation system, including:

[0027] The registration unit is used to perform feature matching between the real-time image and the reference image, and to calculate the homography matrix;

[0028] A geometric transformation unit is used to perform geometric transformation on the real-time graph based on the homography matrix to eliminate the difference in geometric distortion between the graph and the reference graph.

[0029] The feature processing unit is used to extract structural consistency features between the real-time graph and the reference graph, and to perform cross-correlation operations to obtain the cross-correlation matrix;

[0030] The peak analysis unit is used to determine the position of the main peak in the cross-correlation matrix and to determine multiple secondary peaks according to a predetermined spatial rule;

[0031] The verification unit is used to determine whether the registration was successful based on the peak ratio relationship between the main peak and the secondary peak.

[0032] The error assessment unit is used to determine the coordinates of the first center point and the second center point when registration is successful, and to calculate the Euclidean distance between them as the registration error.

[0033] Furthermore, the coordinates of the first center point are determined by the position of the main cross-correlation peak in the reference image, representing the center of the region in the reference image that is most similar to the structure of the real-time image.

[0034] Furthermore, the coordinates of the second center point are obtained by mapping the center position of the real-time graph to the reference graph using a homography matrix.

[0035] Furthermore, when the coordinates of the first center point and the second center point completely coincide, it indicates that the registration result is correct. When there is an offset between the two, their Euclidean distance is used to reflect the magnitude of the registration error.

[0036] Based on the above technical solutions and the technical problems they solve, please analyze the advantages and positive effects of the technical solution to be protected from the following aspects:

[0037] This technical solution proposes a multi-scenario adapted cross-correlation peak ratio registration verification and dual-center error evaluation method. This method determines whether the registration result is a successful digital result and returns an evaluation error that is close to the true registration error.

[0038] The core innovation of this technical solution lies in the construction of a registration result judgment and accuracy evaluation method based on cross-correlation operation. Under any dataset (whether it is a multimodal image dataset or a viewpoint change image dataset) or any scene, the homography matrix is ​​calculated using the registration result, an affine transformation is performed on the real-time image, and a cross-correlation operation is performed between the affine transformation real-time image and the reference image. The registration success or failure is judged by the correlation coefficient matrix, and the registration error is evaluated by combining the two center points.

[0039] 1. Cross-correlation matching based on feature extraction using phase consistency algorithm

[0040] Compared to traditional Gaussian filtering-based feature extraction methods, using phase consistency algorithms for feature extraction significantly improves the sensitivity of the judgment method to slight geometric distortions such as rotation and scaling caused by registration results. This is because phase consistency methods retain only image texture information and simultaneously sharpen the texture. When correct registration aligns the texture structures of two images, the energy generated by resonance reaches its maximum, resulting in a higher correlation peak than side peaks, thus achieving side peak suppression and making the main peak "sharper." Furthermore, in terms of computational efficiency, image enhancement using Gaussian filtering often requires filtering in 4-8 directions, corresponding to 4-8 filtering results, which need to be convolved separately and then superimposed. The method proposed in this paper generates only one feature extraction result and performs only one convolution, improving computational efficiency.

[0041] 2. Peak ratio determination method based on eight directions

[0042] Compared to the method based on Gaussian filtering enhancement followed by cross-correlation calculation to determine the global primary and secondary peak ratio, this technical solution, using the same judgment threshold, reduces the false alarm rate from 44.3% to 10%, the missed alarm rate from 12% to 8%, while increasing the correct judgment rate from 43.7% to 82%. It can also correctly judge banded peaks caused by repetitive characteristics. This demonstrates the improvement brought by the proposed method in reducing the false alarm rate and increasing the probability of correct judgment.

[0043] 3. Registration accuracy evaluation method based on two center points

[0044] In practical applications of image registration methods, it is necessary to register unknown image pairs. In such cases, there is no ground truth value to evaluate registration accuracy, and existing techniques do not address how to provide an effective answer. To solve this problem, this technical solution proposes a registration accuracy evaluation method based on two center points. The difference between the estimated error and the actual registration error is 0.13-1.33 pixels (results from different datasets), and the standard deviation of the difference is 0.36-1.21, indicating that the error closely surrounds the mean. This demonstrates that the Euclidean distance between the two center points can be used to reflect the magnitude of the registration error.

[0045] 4. Calculation time and efficiency

[0046] The proposed method in this technical solution is not highly complex. Both the feature extraction based on phase consistency and the peak-to-peak ratio (PPR) determination method in eight directions are innovative advancements built upon existing technologies, effectively improving the accuracy of registration results. Furthermore, the integration of Fourier transform to accelerate image convolution significantly reduces computation time. The registration accuracy evaluation method based on two center points utilizes the product of cross-correlation matching, which not only avoids adding extra computation time but also provides highly accurate assessments of registration accuracy.

[0047] 5. Project Applicability

[0048] The proposed method, characterized by low complexity, high effectiveness, and short computation time, can be successfully implemented in engineering applications. It provides a novel approach for UAV positioning and navigation, radar guidance, SLAM, and large-scale image registration tasks, enabling automated evaluation of registration results and determining whether to utilize them based on assessment accuracy. This also opens up a new field of registration accuracy evaluation.

[0049] Second, as supporting evidence of the inventive step of the claims of this technical solution, it is also reflected in the following important aspects:

[0050] (1) The expected benefits and commercial value of this technical solution after its transformation are as follows:

[0051] The expected benefits and commercial value of this technical solution are mainly reflected in three core areas: autonomous driving, drone navigation, and remote sensing mapping, and it has significant industrialization potential.

[0052] autonomous driving field

[0053] This method can verify the image registration results of in-vehicle vision navigation in real time and output a quantized accuracy evaluation value. This effectively avoids positioning deviations caused by registration failures, improving the safety and reliability of autonomous driving systems. Related technologies can be integrated into in-vehicle vision modules, providing core algorithm support for autonomous driving solution providers.

[0054] Unmanned aerial vehicle (UAV) navigation field

[0055] For low-altitude drone operations, this method is adaptable to multimodal image registration and verification requirements. It can quickly determine registration validity without relying on ground truth data, significantly reducing drone positioning errors in complex environments. The technology can be transformed into a drone navigation algorithm plugin, empowering drone applications in plant protection, surveying, and inspection.

[0056] Remote sensing mapping field

[0057] For large-scale, multi-source remote sensing image registration tasks, this method enables automated registration result verification and accuracy assessment. It replaces the traditional manual interpretation method, significantly improving remote sensing image processing efficiency and reducing operating costs. The technology can be applied to image data processing platforms in industries such as satellite remote sensing and aerial mapping.

[0058] (2) The technical solution of this invention fills a technical gap in the industry both domestically and internationally:

[0059] This technical solution fills the technical gap in the automated verification and accuracy evaluation of multi-scene image registration results under the condition of no truth value.

[0060] Existing technologies focus on improving the robustness and accuracy of registration algorithms, but do not address the issue of determining whether the registration results are "usable." Both domestically and internationally, there is a lack of registration result verification methods that can be directly applied in engineering.

[0061] Existing methods struggle to meet the registration and verification needs of complex scenarios involving multimodalities and geometric distortions. This technical solution, by combining phase consistency feature extraction, overcomes the technical bottleneck of multi-scenario adaptation. Furthermore, the method of determining the peak ratio in eight directions enables highly reliable registration result evaluation with a high accuracy rate.

[0062] Traditional accuracy assessment relies on ground truth data, which is not applicable to large-scale, multimodal real-world engineering scenarios. The dual-center error assessment method proposed in this technical solution achieves quantitative assessment of registration accuracy for the first time without ground truth data, human intervention, or dependence on the features themselves.

[0063] (3) Whether the technical solution of this technical solution solves the technical problem that people have long wanted to solve but have never been able to solve successfully:

[0064] This technical solution solves two major technical challenges that have long existed in the fields of autonomous navigation and remote sensing mapping: "judging the validity of registration results" and "assessing accuracy without true values".

[0065] In the field of autonomous navigation, the industry has long desired a technology that can provide real-time feedback on the validity of registration results. If registration failures are not identified in a timely manner, they can lead to positioning errors and even safety incidents. Previous technologies have been unable to achieve high-precision, low-latency registration verification in complex scenarios.

[0066] In the field of remote sensing and mapping, large-scale image registration tasks lack ground truth data. The industry has long needed an automated accuracy assessment method to replace inefficient manual interpretation. Previous technical solutions have failed to meet the requirements of engineering practicality.

[0067] The core challenge of the aforementioned problem is "balancing adaptability to multiple scenarios, engineering practicality, and evaluation accuracy." This technical solution achieves a unification of these three aspects for the first time through a combination of cross-correlation peak ratio verification and dual-center error assessment.

[0068] (4) Whether the technical solution of this invention overcomes technical bias:

[0069] This technical solution overcomes the industry's bias of "emphasizing registration algorithm optimization while neglecting result validity verification".

[0070] Traditional technical biases hold that improving the robustness and accuracy of registration algorithms can solve problems in practical applications. This overlooks the fact that the core prerequisite for engineering is whether the registration result is usable.

[0071] It is widely believed in the industry that registration accuracy assessment must rely on true value data. This has led to a technical bias that "accuracy cannot be assessed without true values." This technical solution demonstrates that high-precision assessment can be achieved without true values ​​by calculating the Euclidean distance between two centers.

[0072] Previous technical solutions relied on complex computational processes to verify registration results. This new solution overcomes this bias by simplifying phase consistency feature extraction and peak-to-peak ratio (PTR) determination, thereby reducing computational complexity while maintaining accuracy and meeting real-time requirements in engineering applications. Attached Figure Description

[0073] Figure 1 This is a flowchart of the multi-scenario adaptation cross-correlation peak ratio registration verification and dual-center error evaluation method provided in this embodiment of the technical solution.

[0074] Figure 2 This is a block diagram of the cross-correlation peak ratio registration verification and dual-center error evaluation system provided in this embodiment of the technical solution.

[0075] Figure 3 This is a detailed flowchart of the multi-scenario adaptation cross-correlation peak ratio registration verification and dual-center error evaluation method provided in this embodiment of the technical solution.

[0076] Figure 4 This is a registration image of a multimodal image provided in this embodiment of the technical solution.

[0077] Figure 5 This is a feature extraction image of the real-time image and the reference image provided in this embodiment of the technical solution.

[0078] Figure 6 This is an Euclidean distance diagram for calculating the coordinates of two center points, provided in this embodiment of the technical solution.

[0079] Figure 7The embodiment of this technical solution solves for the optimal threshold and uses ROC curves to exhaustively enumerate the corresponding FTR and TPR graphs under different thresholds.

[0080] Figure 8 This is the route planning diagram provided in this embodiment of the technical solution.

[0081] Figure 9 These are reference diagrams and real-time diagrams provided in the embodiments of this technical solution.

[0082] Figure 10 This is a demonstration diagram of the registration result provided in this embodiment of the technical solution.

[0083] Figure 11 This is a visualization of the registration accuracy provided in this embodiment of the technical solution. Detailed Implementation

[0084] To make the objectives, technical solutions, and advantages of this technical solution clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the technical solution and are not intended to limit the scope of the technical solution.

[0085] like Figure 1 As shown in the figure, the multi-scenario adaptation cross-correlation peak ratio registration verification and dual-center error evaluation method provided in this embodiment of the technical solution includes the following steps:

[0086] S101, the image pair to be processed is defined as a real-time image and a reference image, and the homography matrix between the two is calculated based on the image feature matching results;

[0087] S102, the real-time image is geometrically transformed using the homography matrix to ensure that the real-time image and the reference image are consistent in geometric structure, thereby eliminating the effects of rotation, scale changes and geometric distortion.

[0088] S103, extract structural consistency features from the geometrically consistent real-time graph and the reference graph, and perform cross-correlation operation based on the structural consistency features to obtain the cross-correlation matrix;

[0089] S104, determine the position of the main peak in the cross-correlation matrix, and divide the cross-correlation matrix into eight non-overlapping regions extending along the horizontal, vertical and diagonal directions with the main peak position as the center. Select the local maximum value with the largest amplitude in each region as the secondary peak.

[0090] S105, based on the proportional relationship between the amplitude of the main peak and the amplitudes of each secondary peak, the consistency of the registration results is verified;

[0091] S106, when the verification result meets the preset conditions, the coordinates of the first center point are determined based on the position of the main peak of cross-correlation, and the center position of the real-time graph is mapped to the reference graph based on the homography matrix to determine the coordinates of the second center point.

[0092] S107, calculate the Euclidean distance between the coordinates of the first center point and the coordinates of the second center point, and use the Euclidean distance as the evaluation result of the registration error.

[0093] The homography matrix provided in this embodiment of the technical solution is calculated from the pixel correspondence between the real-time image and the reference image. The geometric transformation is an affine transformation performed on the real-time image based on the homography matrix, so that the two images maintain consistency in spatial structure.

[0094] The structural consistency features provided in this embodiment are obtained through phase consistency feature extraction to suppress the influence of multimodal imaging differences on cross-correlation results and enhance the significance of the main cross-correlation peak relative to the side peaks.

[0095] This technical solution embodiment provides an image registration and verification method based on cross-correlation primary and secondary peak spatial constraints, including:

[0096] Perform cross-correlation on the two images after geometric consistency processing to obtain the cross-correlation matrix;

[0097] The main peak with the largest amplitude is determined in the cross-correlation matrix, and the cross-correlation matrix is ​​divided into multiple regions in eight directions with the location of the main peak as the center.

[0098] In each of the aforementioned regions, a local maximum with the largest amplitude is determined as the secondary peak;

[0099] Based on the proportional relationship between the amplitude of the main peak and the amplitudes of each secondary peak, a peak ratio criterion is constructed, and the validity of the image registration result is determined based on the peak ratio criterion.

[0100] The region division method provided in this embodiment of the technical solution is to divide the cross-correlation matrix into eight triangular regions with the main peak position as the origin, along the horizontal, vertical and two diagonal directions, and the regions do not overlap with each other.

[0101] The peak ratio criterion provided in this technical solution embodiment is achieved by comparing whether the ratio of the amplitude of the main peak to the amplitude of the secondary peak exceeds a preset threshold, so as to reduce the probability of misjudgment.

[0102] The multi-scene adaptive cross-correlation peak ratio registration verification and dual-center error evaluation method proposed in this technical solution is not a simple superposition of several isolated techniques. Instead, it constructs a collaborative working mechanism with an inherent logical closed loop, centered around the core technical goal of achieving highly reliable image registration effectiveness judgment and interpretable error quantification under complex imaging conditions. This mechanism takes geometric consistency correction, structural consistency constraints, spatially constrained peak discrimination, and dual-center error mapping as its core components. Each component is coupled and mutually restrictive in terms of functional positioning, input-output relationship, and constraints, forming an inseparable overall technical solution.

[0103] This technical solution first obtains the pixel correspondence between the real-time image and the reference image through feature matching, and then calculates the homography matrix accordingly. This homography matrix is ​​not merely output as a conventional registration result, but is used as a unified geometric benchmark for all subsequent criteria and error assessments. By performing a geometric transformation on the real-time image based on this homography matrix, the two images achieve spatial structural consistency, eliminating the interference of geometric factors such as rotation, scale changes, and perspective distortion on the cross-correlation response distribution from the source, thus creating a stable premise for subsequent correlation analysis. This step ensures that the formation of the main cross-correlation peak is primarily determined by the true spatial correspondence, rather than a spurious peak caused by geometric mismatch.

[0104] After completing geometric consistency correction, this technical solution does not directly perform pixel-level cross-correlation, but further introduces structural consistency features as an intermediate representation layer. These structural consistency features are extracted using a phase consistency method, highlighting modality-independent structural information in the image and effectively suppressing the influence of illumination differences, sensor response differences, and imaging noise on the correlation results. This results in a response pattern of "concentrated main peak and discrete side peaks" in the cross-correlation matrix. This structural feature constraint ensures that the significance of the main cross-correlation peak originates from the alignment of the true structure, rather than simple gray-level similarity.

[0105] Based on this, this technical solution proposes a joint primary and secondary peak discrimination mechanism based on spatial orientation constraints. By centering on the primary cross-correlation peak, the correlation matrix is ​​divided into eight non-overlapping regions along the horizontal, vertical, and diagonal directions. Within each region, only the local maximum with the largest amplitude is selected as the secondary peak, ensuring that the selection of the secondary peak simultaneously satisfies the requirements of spatial dispersion and amplitude representativeness. This regional constraint avoids the misjudgment risk caused by the concentration of secondary peaks in the neighborhood of the primary peak in traditional methods, ensuring that the amplitude relationship between the primary and secondary peaks has stable statistical significance. The peak ratio criterion constructed in this way is not an isolated threshold judgment, but rather works in conjunction with the aforementioned geometric consistency and structural consistency constraints to form a reliability verification mechanism with multiple superimposed conditions.

[0106] After the registration results pass the consistency verification described above, this technical solution further introduces a dual-center error evaluation strategy: on the one hand, the position of the main peak of cross-correlation is used to determine the first center point based on the correlation response; on the other hand, the homography matrix is ​​used to map the geometric center of the real-time graph to the reference graph to form a second center point. The two types of center points reflect the "optimal position of the correlation response" and the "theoretical position of the geometric mapping," respectively, and the Euclidean distance between them is the quantified result of the registration error. This error definition inherits information from both the correlation domain and the geometric domain, avoiding the one-sidedness of a single evaluation scale, and giving the evaluation results clear physical meaning and engineering interpretability.

[0107] This technical solution organically integrates geometric correction, structural features, spatially constrained peak discrimination, and dual-center error mapping within a unified homography framework, forming a collaborative working mechanism that is dependent on each other and complementary in function. No single technical means can achieve the same technical effect without this overall structure, thus demonstrating significant overall creativity and irreplaceability.

[0108] like Figure 2 As shown in the figure, this embodiment of the technical solution provides a multi-scenario adaptable cross-correlation peak ratio registration verification and dual-center error evaluation system, including:

[0109] The registration unit is used to perform feature matching between the real-time image and the reference image, and to calculate the homography matrix;

[0110] A geometric transformation unit is used to perform geometric transformation on the real-time graph based on the homography matrix to eliminate the difference in geometric distortion between the graph and the reference graph.

[0111] The feature processing unit is used to extract structural consistency features between the real-time graph and the reference graph, and to perform cross-correlation operations to obtain the cross-correlation matrix;

[0112] The peak analysis unit is used to determine the position of the main peak in the cross-correlation matrix and to determine multiple secondary peaks according to a predetermined spatial rule;

[0113] The verification unit is used to determine whether the registration was successful based on the peak ratio relationship between the main peak and the secondary peak.

[0114] The error assessment unit is used to determine the coordinates of the first center point and the second center point when registration is successful, and to calculate the Euclidean distance between them as the registration error.

[0115] The coordinates of the first center point provided in this embodiment of the technical solution are determined by the position of the main peak of the cross-correlation in the reference figure, representing the center of the region in the reference figure that is most similar to the structure of the real-time figure.

[0116] The second center point coordinates provided in this embodiment of the technical solution are obtained by mapping the center position of the real-time graph to the reference graph through a homography matrix.

[0117] The embodiment of this technical solution indicates that the registration result is correct when the coordinates of the first center point and the second center point are completely coincident. When there is an offset between the two, their Euclidean distance is used to reflect the magnitude of the registration error.

[0118] The multi-scenario adaptable cross-correlation peak ratio registration verification and dual-center error assessment system proposed in this technical solution is not a simple splicing of existing image registration, cross-correlation analysis, or error calculation modules. Instead, it constructs a system-level collaborative mechanism that is highly coupled in terms of functional flow, data dependence, and criterion logic, centered around the core objective of achieving reliable verification of registration results and interpretable error assessment under complex imaging conditions. The functional units in the system do not work independently, but rather form a progressive and mutually supportive overall structure under a unified geometric mapping framework and related analysis constraints.

[0119] First, the registration unit establishes the pixel correspondence between the real-time image and the reference image by performing feature matching, and calculates the homography matrix accordingly. This homography matrix is ​​not merely an output of the registration result, but serves as a data link and the basis for geometric constraints throughout the entire system. Based on this, the geometric transformation unit performs a geometric transformation on the real-time image, ensuring that the real-time image maintains spatial structural consistency with the reference image. This eliminates interference from factors such as rotation, scale changes, and perspective distortion at the system level, thus preventing distortion caused by different modules operating independently under different geometric assumptions. This geometric consistency processing provides a unified spatial reference frame for the subsequent feature processing unit and peak analysis unit, avoiding result distortion caused by different modules operating independently under different geometric assumptions.

[0120] After completing geometric consistency correction, the feature processing unit does not directly perform cross-correlation calculations based on the original pixels. Instead, it first extracts structural consistency features between the real-time image and the reference image. This feature extraction process emphasizes the preservation of structural information and the suppression of imaging modal differences, ensuring that the similarity reflected by the cross-correlation calculation originates from the real spatial structure rather than accidental consistency in grayscale or noise. The resulting cross-correlation matrix exhibits an overall distribution characteristic of prominent main peaks and dispersed side peaks, providing a stable foundation for subsequent peak analysis.

[0121] The peak analysis unit determines the location of the primary peak with the largest amplitude in the aforementioned cross-correlation matrix. Centered on this primary peak, it divides regions in multiple directions according to predetermined spatial rules, selecting only one representative local maximum from each region as the secondary peak. This spatially constrained peak extraction method keeps the primary and secondary peaks spatially separated, avoiding the multi-peak confusion problem caused by local perturbations or repetitive textures in traditional correlation analysis. Therefore, the amplitude relationship between the primary and secondary peaks is no longer an isolated numerical comparison, but a stable criterion formed under spatial structural constraints.

[0122] The verification unit makes a comprehensive judgment based on the peak ratio relationship between the above-mentioned main peak and multiple sub-peaks. Only when the main peak is significantly superior to the sub-peaks in both amplitude and spatial distribution, the registration result is determined to be valid. This verification result directly determines whether the error evaluation unit is triggered, making the error evaluation based on the premise of "credible registration result", rather than treating all registration results equally for numerical calculation.

[0123] After the registration verification passes, the error evaluation unit simultaneously introduces two types of center definitions: one is the first center point determined by the position of the cross-correlation main peak, which is used to represent the best structure matching position in the sense of correlation; the other is the second center point obtained by mapping the center of the real-time image to the reference image through the homography matrix, which is used to represent the theoretical alignment position in the sense of geometric mapping. The two types of center points have different sources and complementary physical meanings. Their Euclidean distance not only reflects the geometric mapping deviation but also reflects the correlation response offset, thus constituting a registration error quantification index with clear engineering significance.

[0124] The system of this technical solution organically integrates registration, feature processing, peak analysis, verification criterion, and error evaluation under the homography geometry framework, forming an overall collaborative mechanism with closely linked links. Any unit cannot achieve the same technical effect when separated from this system structure, which reflects significant overall creativity at the system level and makes the refutation based on simple combination of existing technologies lack a basis for establishment.

[0125] A multi-scenario adaptation cross-correlation peak ratio registration verification and dual-center error evaluation method provided by an embodiment of this technical solution, the specific method:

[0126] S1. Name the image pair as the real-time image and the reference image respectively, use the SuperGlue point feature registration algorithm to implement the registration of the image pair, obtain the pixel correspondence of the image pair, and calculate the homography matrix;

[0127] S2. Use the homography matrix to perform an affine transformation on the real-time image to eliminate the geometric distortion difference of the corresponding image content in the real-time image and the reference image;

[0128] S3. Use the phase consistency algorithm to extract features from the real-time image and the reference image to eliminate the multi-modal difference between the two, and then perform a convolution operation on the extracted real-time image and the reference image to obtain the correlation coefficient matrix;

[0129] S4. Obtain the highest peak and the corresponding row and column numbers from the correlation coefficient matrix, and then solve all the local maxima and the positions of the local maxima points of the matrix; taking the position of the highest peak point as the origin, divide the matrix into eight triangular regions in the shape of a "plus" sign, sort the local maxima of each region, and take out the largest local maximum and the position of the largest local maximum point (excluding the origin). These eight local maxima are defined as sub-peaks;

[0130] S5. Use the threshold method to evaluate the registration results; the peak ratio is significantly positively correlated with the quality of the registration results, and the threshold method can quickly judge the registration results; use ROC curves to quickly find the optimal threshold and verify it on the test set;

[0131] S6. Only when the image registration is determined to be successful can the registration accuracy be further evaluated;

[0132] Once registration is successful, the real-time image after affine transformation is convolved with the reference image. The location of the highest peak indicates that the structure of that region in the reference image is most similar to that of the real-time image after affine transformation. This location is called the first center point coordinate. At the same time, based on the homography matrix calculated from the registration result, the center point of the real-time image before affine transformation can be projected onto a certain row and column coordinate in the reference image. This is called the second center point coordinate. Theoretically, if registration is completely successful, the two center point coordinates will completely coincide.

[0133] S7. Calculate the Euclidean distance between the coordinates of the two center points and use the Euclidean distance as the prediction error.

[0134] In the image pair described in S1 of this technical solution embodiment, the real-time image can be an image captured by a drone or an image obtained by other sensors, and the reference image can be an image captured by a satellite or an image obtained by other means, depending on different task requirements; the two can be a multimodal image pair or a viewpoint change image pair.

[0135] In the embodiment of this technical solution, the homography matrix calculated from the registration result in S1 is a matrix that reflects the mathematical relationship between pixels in an image. The image can be regarded as a two-dimensional matrix. Multiplying the real-time image by the homography matrix can transform the real-time image to the geometric state of the reference image, eliminating the rotation, scaling, geometric distortion, and other relationships between the two, and maintaining the same structural information. Therefore, an affine transformation is performed on the real-time image in S2. The purpose of doing this is to determine whether the registration is successful in the following steps by judging the peak ratio of the cross-correlation coefficient matrix.

[0136] The phase consistency feature extraction of the image in S3 provided in this embodiment can eliminate the influence of multimodal differences between image pairs, allowing this technical solution to be applied in multimodal datasets. Furthermore, this method can achieve side peak suppression and obtain more accurate cross-correlation matching results. The convolution operation in S3 can essentially be regarded as a resonance phenomenon. When there is structural texture information similar to the real-time image in the reference image, the correlation coefficient matrix will obtain a main peak with a relatively higher intensity than the side peaks due to resonance.

[0137] In the S4 step provided in this technical solution embodiment, the peak value of the main peak is obtained and its ratio is calculated with the peak values ​​of the eight surrounding secondary peaks. The matching success is determined by judging the number of times the ratio of the main peak to the secondary peak exceeds a threshold. This can effectively reduce the false alarm rate and the missed alarm rate, and achieve correct judgment of the registration result.

[0138] The judgment threshold in S5 provided in this technical solution embodiment is obtained by registering the image using an SG network and then performing convolution operations to obtain the peak ratio. The optimal threshold is found through the ROC curve. This technical solution aims to minimize the false alarm rate and improve the correct judgment rate while minimizing the false alarm rate. Therefore, the threshold corresponding to the maximum Youden index was selected.

[0139] In step S6, when a successful match is determined, there will be a region in the reference image that is very similar to the structure of the real-time image. The position of the center point of this region in the reference image is taken out and called the first center point coordinates. The homography matrix calculated based on the registration result can also project the center point of the real-time image onto the reference image, and the position of the center point in the reference image is called the second center point coordinates.

[0140] In S7, when the registration is theoretically perfect, the two center points coincide. However, after testing with a multimodal image dataset, when registration errors occur, the difference between the Euclidean distance between the two center points and the actual registration error is 0.13-1.33 pixels, with a standard deviation of 0.36-1.21, closely surrounding the mean. This demonstrates that the Euclidean distance between the two center points can be used to reflect the magnitude of the registration error.

[0141] like Figure 3 A method for judging registration results and evaluating accuracy based on cross-correlation calculation includes the following steps:

[0142] Step 1: To register multimodal images, this study proposes a parallel computation approach using multiple methods, selecting the SuperGlue method for computation. The registration results are presented, and the homography matrix is ​​calculated, as shown in the appendix. Figure 4 As shown.

[0143] Step 2: Use the homography matrix to perform an affine transformation on the real-time image to eliminate the geometric distortion differences (such as rotation, scaling, geometric transformation, etc.) between the corresponding image content in the real-time image and the reference image.

[0144] Step 3: Use the phase consistency algorithm to extract features from the real-time image and the reference image to eliminate multimodal differences between them. Then, perform a convolution operation between the extracted real-time image and the reference image to obtain the correlation coefficient matrix, as shown in the appendix. Figure 5 As shown.

[0145] Step 4: Obtain the highest peak and the corresponding row and column numbers from the correlation coefficient matrix, and then solve for all local maxima and their positions in the matrix. Taking the position of the highest peak point as the origin, divide the matrix into eight triangular regions in a "cross" shape. Sort the local maxima in each region and pick out the largest local maximum and its position (excluding the origin). These eight local maxima are named sub-peaks.

[0146] Step 5: Dataset description: This technical solution constructs a dataset containing multiple modalities and multiple perspectives. The dataset contains 1000 pairs of multi-modal and multi-perspective remote sensing image pairs, covering multiple cities around the world, including Beijing, Shanghai, Suzhou, Wuhan, Sanhe, Yancheng, Dengfeng, Zhongshan, Zhuhai, etc. in China, Rennes in France, Tucson, Omaha, Guam, and Jacksonville in the United States, and Dwarka and Agra in India. The types include optical-SAR image pairs, optical-infrared image pairs, satellite optical-UAV optical image pairs. The training and test sets are divided in a ratio of 6:4, and the training set is used to solve the threshold.

[0147] Step 6: To solve for the optimal threshold, use the ROC curve to exhaustively list the corresponding FTR and TPR at different thresholds. When a certain threshold maximizes the Youden index, it means that when the peak ratio exceeds this threshold, the missed alarm rate can be reduced to the minimum, that is, the correct true value is likely to be correct at this time. When a certain threshold is a point on the ROC curve closest to the upper left corner point, it means that when the peak ratio is less than this value, the false alarm rate reaches the minimum, that is, the incorrect true value is likely to be incorrect at this time. As Figure 7 shown is the optimal threshold obtained by this study using the training set, where the training and test sets are divided in a ratio of 6:4, and the data comes from the registration results of optical and SAR, infrared, and optical image pairs. It can be observed from the figure that when the Youden index reaches the maximum, the missed alarm rate reaches the lowest, and at the same time, the false alarm rate period is controlled to be as small as possible. When using 1.35 as the threshold, after calculation, the missed alarm rate in the test set is 0.08, the false alarm rate is 0.10, and the correct judgment rate reaches 0.82. As shown in the appendix Figure 7 shown.

[0148] Step 7: The judgment rule of this technical solution is: judge one by one whether these 8 ratios exceed the threshold. When 6 or more ratios exceed the threshold, it is considered that the image pair registration is successful; otherwise, it is judged as failed.

[0149] Step 8: Only when image registration is determined to be successful can the registration accuracy be further evaluated. When registration is successful, the real-time image after affine transformation is convolved with the reference image. The location of the highest peak indicates that the region in the reference image is most structurally similar to the real-time image after affine transformation; this location is called the first center point coordinate. Simultaneously, based on the homography matrix calculated from the registration results, the center point of the real-time image before affine transformation can be projected onto a certain row and column coordinate in the reference image; this is called the second center point coordinate. Theoretically, with complete registration, the two center point coordinates will completely coincide.

[0150] Step 9: Calculate the Euclidean distance between the coordinates of the two center points. The estimated error is maintained at 0.13-1.33 pixels, and the standard deviation of the difference is 0.36-1.21.

[0151] Step 10, Phase Consistency Algorithm Parameter Description: Frequency decomposition layer number is 3, octave band amplification factor of filter bank is 1.6, frequency domain standard deviation ratio of Gaussian filter is 0.75, gradient calculation smoothing parameter is 3, and noise suppression ratio coefficient is 1.

[0152] In the image pairs described in step 1, the two images may or may not share the same image structure information. For example, in large-scale image registration tasks, image pairs may become disordered, resulting in completely different image pairs being registered. Alternatively, during UAV positioning, the captured image content may not be present in the pre-stored image information. In such cases, registration disorder can occur. However, the peak ratio (BPRR) method described in this paper can also determine registration failure, so there is no need to worry about the lack of duplicate information. The registration algorithm mentioned only needs to be a point feature registration algorithm. This type of algorithm can provide the pixel correspondence between image pairs and can be used to calculate the homography matrix.

[0153] The homography matrix calculated from the registration results in step 2 is a matrix that reflects the mathematical relationship between pixels in an image; the image can be viewed as a two-dimensional matrix. Multiplying the real-time image by the homography matrix transforms the real-time image to the geometric state of the reference image, eliminating rotation, scaling, geometric distortion, and other relationships between the two while maintaining identical structural information.

[0154] Step 3 involves extracting phase consistency features from the image, which eliminates the influence of multimodal differences between image pairs. This allows the proposed solution to be applied to multimodal datasets. Furthermore, this method can achieve sidepeak suppression and obtain more accurate cross-correlation matching results, as shown in the attached figure. Figure 5 As shown. The convolution operation in step 3 can essentially be seen as a resonance phenomenon. When there is structural texture information in the reference image that is similar to that in the real-time image, the correlation coefficient matrix will have a main peak with a relatively higher intensity than the side peaks due to resonance.

[0155] In step 7, if at least 6 out of 8 directions exceed the threshold, the region is considered to be most similar to the real-time image. Conversely, if the highest peak is nearly identical to the eight secondary peaks, the similarity is considered to be no significant compared to other regions, and registration is not considered successful, or the reference image does not contain a structure similar to the real-time image. (See attached image.) Figure 5 As shown.

[0156] In step 7, the peak value of the main peak is obtained and its ratio is calculated with the peak values ​​of the eight surrounding secondary peaks. The success of the match is determined by the number of times the ratio exceeds a threshold, which effectively reduces the false alarm rate and the missed alarm rate. Tested on multimodal and viewpoint variation datasets, this method achieved a false alarm rate of 10%, a missed alarm rate of 8%, and a correct identification rate of 82%. It can also correctly identify banded peak clusters caused by repetitive features.

[0157] In step 8, when a successful match is determined, a region in the reference image will appear that is extremely similar to the structure of the real-time image. The position of the center point of this region in the reference image is extracted and referred to as the coordinates of the first center point, as shown in the attached figure. Figure 3 The location of the highest peak is shown in the figure. The homography matrix calculated based on the registration results can also project the center point of the real-time image onto the reference image, and the position of this center point in the reference image is called the coordinates of the second center point.

[0158] In step 9, when the registration is theoretically perfect, the two center points coincide. However, testing on a multimodal image dataset shows that when registration errors occur, the difference between the Euclidean distance between the two center points and the actual registration error is 0.13-1.33 pixels, with a standard deviation of 0.36-1.21, closely surrounding the mean. This demonstrates that the Euclidean distance between the two center points can be used to reflect the magnitude of the registration error. (See attached image) Figure 6 As shown, the image on the right is the mosaicking result of the maximum similarity position obtained after cross-correlation operation, which is the first center point position; while the image on the left is the fusion result of the real-time image mosaicked into the reference image after affine transformation. In this case, the second center point position can be obtained. The Euclidean distance between these two center points is used to evaluate the registration error. The evaluation accuracy is 0.13-1.33 pixels, and the standard deviation of the accuracy is 0.36-1.21.

[0159] Example 1: Registration Verification and Error Assessment of Visible Light Remote Sensing Images under Large-Scale Viewpoint Variations

[0160] Visible light images acquired by UAVs at different altitudes and flight paths were selected as real-time images, while high-resolution remote sensing orthophotos of the same region were used as reference images. First, the homography matrix was calculated based on image feature matching results, and a geometric transformation was performed on the real-time images to ensure overall structural consistency between the two images. Then, structural consistency features were extracted from the transformed real-time and reference images, and cross-correlation was performed to obtain a cross-correlation matrix. The dominant peak was identified in the cross-correlation matrix, and regions were divided in eight directions centered on the dominant peak. Secondary peaks were selected from each region to construct the ratio of dominant to secondary peaks. Experiments showed that in cases of successful registration, the dominant peak was significantly higher than the secondary peaks in each region. However, in scenarios of misregistration or localized texture repetition, the ratio of dominant to secondary peaks decreased significantly, thus achieving reliable registration verification. After successful verification, the Euclidean distance between the first center point obtained based on the cross-correlation dominant peak and the second center point obtained based on geometric mapping was further calculated. This distance stably reflects the actual registration error.

[0161] (1) Route planning

[0162] like Figure 8 As shown, to achieve the highest possible accuracy in image registration, a town park area with rich features was selected as the application scenario for the image registration technology. The baseline optical satellite image was taken on October 11, 2024, and provided by Century Space Technology Co., Ltd., with a resolution of 3m / pixel.

[0163] (2) Reference graph and real-time graph

[0164] To reduce memory requirements for image storage and improve registration speed, after flight path planning, a 1000*1200 pixel image is cropped from the optical satellite image, centered on the planned point; this is called the reference image. The specific pixel coordinates of the reference image within the optical satellite image are then cropped and recorded. The real-time image was captured on December 17, 2025, a day with clear skies and ample light. The reference image and the real-time image are shown below. Figure 9 As shown.

[0165] (3) UAV registration and pixel coordinate calculation

[0166] First, we show the registration results between the real-time image and the reference image. Below are the registration results obtained using the DUST3R algorithm. Figure 10 As can be seen, except for point 2, the other three points have excellent registration results. Even though the real-time image may have geometric distortion due to the influence of inertial navigation, robust registration can still be achieved.

[0167] To visualize the accuracy of the registration results, the registered pixel coordinates are compared with the actual pixel coordinates of the UAV, and the results are then visualized. Figure 11As shown in the figure. The red coordinates represent the actual pixel coordinates of the drone, and the green coordinates represent the registered pixel coordinates.

[0168] (4) Statistical analysis of image registration results

[0169] In addition to comparing the actual pixel coordinates with the registered pixel coordinates, this paper also statistically analyzes the judgment results and the actual results to quantify the prediction accuracy, as shown in Table 1. Accuracy refers to the pixel Euclidean distance between the two coordinates. In terms of accuracy, using SuperGlue achieves an optimal registration accuracy of 1 pixel, while the worst is around 6 pixels. This is due to the nonlinear geometric distortion between the real-time image and the reference image of aerial point 2, and the scarcity of feature points around the lake. If the image resolution is higher and the features are richer, the horizontal error can reach approximately 3 meters. Combined with an infrared altimeter, this method can serve as an important supplementary means for satellite navigation and positioning. Furthermore, in optical image registration tasks, it can accurately identify whether registration is successful, and the estimated accuracy is very close to the actual registration accuracy. This demonstrates that the registration result evaluation method has extremely high reliability and accuracy in predicting registration results without human intervention, and can promote the practical application and deployment of registration technology. Further analysis of the double-center error reveals that the predicted accuracy increases with the actual accuracy, indicating that the distance between the two center points changes monotonically with the artificially introduced registration deviation, proving that this error assessment method has good physical consistency.

[0170] Table 1. Statistics of Actual Flight Judgment Results and Prediction Accuracy

[0171]

[0172] Example 2: Robust Registration Verification of Infrared and Visible Multimodal Images

[0173] Nighttime infrared images were selected as the real-time image, while daytime visible light images were used as the reference image. Significant differences exist between the two in imaging mechanisms and grayscale distribution. After geometrically homogenizing the real-time image using a homography matrix, a structural consistency feature extraction method was employed to reduce multimodal differences, retaining only edge and structural information. In the cross-correlation matrix, the dominant peak exhibits a clear spatial concentration upon correct registration, while the secondary peaks are dispersed across eight directional regions. A dominant-secondary peak ratio determination mechanism effectively avoids misjudgments caused by locally similar structures.

[0174] When dealing with optical and infrared image datasets, which have high image resolution, and where phase consistency algorithms can effectively overcome modal differences between the two, the accuracy of the classification depends on the precision of the registration results. Therefore, the statistical results in Table 2 show that RIFT, HAPCG, and POS-GIFT, three registration methods adaptable to multimodal images, have high accuracy rates.

[0175] Table 2. Statistics of Optical-Infrared Image Judgment Results

[0176] Index RIFT HAPCG CMM-Net SP+SG CombGlue DUST3R POS-GIFT False alarm rate (%) 10.88 8.34 9.72 5.98 8.67 2.20 3.60 False alarm rate (%) 2.45 1.59 3.77 3.19 2.26 1.83 1.30 Correct judgment rate (%) 86.67 90.07 86.51 90.83 89.07 95.97 95.1

[0177] Because the image quality is good and the registration accuracy assessment is based on a successful judgment, it often means that the registration accuracy is high at this point, so the estimated accuracy is closer to the actual accuracy. Looking at the two indicators in Table 3—the error between the estimated accuracy and the actual accuracy, and the dispersion of the error—RIFT, HAPCG, and POS-GIFT have relatively small errors in most matching points in their registration results, resulting in more accurate homography matrices. Therefore, the difference and dispersion between the estimated accuracy and the actual accuracy are smaller, meaning the results are more accurate and stable.

[0178] Table 3. Statistics on Prediction Accuracy of Optical-Infrared Images

[0179] Index RIFT HAPCG CMM-Net SP+SG CombGlue DUST3R POS-GIFT Prediction accuracy (pixels) 1.83 4.50 4.22 1.65 1.68 4.11 1.18 Actual accuracy (pixels) 2.36 2.69 2.25 2.64 2.11 1.89 1.37 Prediction error (pixel) 0.53 0.99 1.97 1.81 1.43 2.22 0.19 Dispersion 0.69 0.53 2.56 2.41 1.84 3.96 0.86

[0180] Example 3: Verification of Mismatch Suppression in Complex Urban Texture Scenes

[0181] In urban areas with dense high-rise buildings and repetitive road structures, two aerial images with partial occlusion and repetitive textures were selected for processing. Traditional correlation methods based on a single peak are prone to false alarms. In this embodiment, the cross-correlation matrix is ​​divided into spatial regions centered on the main peak, and only the largest secondary peak is retained in each region, preventing the side peaks generated by repetitive textures from concentrating in the same direction. Experimental results show that even with multiple highly correlated regions, the main-secondary peak ratio mechanism can still effectively distinguish between true registration and false matching, demonstrating the innovative advantages of this spatial constraint mechanism in complex scenes.

[0182] When faced with satellite-UAV perspective optical image datasets, although only DUST3R achieved excellent registration results, the judgment method proposed in this paper can still accurately identify the registration results. As shown in Table 4, the reason is that, except for DUST3R, the proportion of correct matching points for all other methods is too small, and most registration results are incorrect registrations. Incorrect registration leads to incorrect homography matrices, thus the evaluation results are consistent with the true results.

[0183] Table 4. Statistics of Satellite-UAV Perspective Image Judgment Results

[0184] Index RIFT HAPCG CMM-Net SP+SG CombGlue DUST3R POS-GIFT False alarm rate (%) 13.22 8.94 6.79 13.87 10.18 6.70 11.44 False alarm rate (%) 1.82 3.35 3.61 2.43 1.16 1.67 3.41 Correct judgment rate (%) 84.96 87.71 89.60 83.70 88.66 91.63 85.15

[0185] Because the matching success rates of the other six methods are too low, this paper only analyzes the error between the estimated accuracy and the true accuracy of DUST3R. Since the method proposed in this paper extracts the corresponding image region based on the distribution of the matching point set before convolution, the estimated accuracy obtained is closer to the true accuracy. However, it is worth noting that since the ground truth values ​​of this dataset are all manually labeled, there will be a certain standard error, approximately 1-2 pixels.

[0186] Table 5. Statistics on Prediction Accuracy of Satellite-UAV Perspective.

[0187] Index DUST3R Prediction accuracy (pixels) 1.97 Actual accuracy (pixels) 1.09 Prediction error (pixel) 0.88 Dispersion 2.15

[0188] Table 5 shows the success rate and estimated accuracy of registration results under various registration methods. However, in practical applications, the optimal solution is selected as the output based on three aspects: registration success flag, estimated accuracy, and registration time.

[0189] Example 4: Verification of Consistency of Two-Center Errors under Significant Scale Variation

[0190] Image pairs of the same region acquired at different resolutions were selected, and the real-time image showed a significant scale change relative to the reference image. After unifying scale and rotation using a homography matrix, cross-correlation analysis was performed. Different degrees of scale perturbation were gradually introduced into the experiment, and the offset between the first and second center points was recorded. The results show that as the scale error increases, the Euclidean distance between the two center points increases synchronously, and the trend is smooth and continuous. This indicates that the dual-center error evaluation mechanism can not only reflect whether registration was successful but also quantify the registration accuracy.

[0191] Example 5: Application Verification of Cross-Sensor Data in Weak Texture Regions

[0192] We selected weakly textured areas such as surface water and desert as test scenarios, where traditional registration methods are easily affected by noise. After geometric uniformity and structural consistency feature extraction, the main cross-correlation peak still forms at the true registration location, while the secondary peaks are dispersed in amplitude across the eight-directional region. The main-to-secondary peak ratio determination effectively avoids misjudgments caused by noise peaks. Furthermore, the bicenter distance remains stable under different noise intensities, indicating that this method remains reliable in information-sparse scenarios.

[0193] Example 6: Adaptability Verification under a Unified Framework for Multiple Scenarios

[0194] The above method is uniformly applied to various types of image data, including visible light, infrared, multi-scale, and multi-view images, without changing the core processing flow; only different types of image pairs are input. Experimental results show that regardless of changes in image source, imaging modality, or scene complexity, the cross-correlation principal and secondary peak spatial constraint mechanism and the dual-center error evaluation mechanism can work together to achieve unified registration verification and error evaluation. This uniformity demonstrates that this technical solution is not an empirical method for a single scene, but an innovative solution with a universal technical mechanism.

[0195] The above embodiments systematically verify the significant technical effects of this technical solution in terms of registration verification reliability and error assessment consistency from the perspectives of multimodal, multi-scale, and multi-texture complexity, fully supporting its protection scope and demonstrating inventiveness.

[0196] It should be noted that the implementation of this technical solution can be achieved through hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of this technical solution can be implemented using hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or programmable hardware devices such as field-programmable gate arrays and programmable logic devices. They can also be implemented using software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0197] The above description is merely a specific implementation of this technical solution, but the scope of protection of this technical solution is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in this technical solution, and within the spirit and principles of this technical solution, should be covered within the scope of protection of this technical solution.

Claims

1. A method for multi-scenario adapted cross-correlation peak ratio registration verification and dual-center error evaluation, characterized in that, Includes the following steps: Step 1: Define the image pair to be processed as a real-time image and a reference image, and calculate the homography matrix between them based on the image feature matching results; Step 2: Use the homography matrix to perform a geometric transformation on the real-time image to make the real-time image and the reference image consistent in geometric structure, so as to eliminate the effects of rotation, scale changes and geometric distortion. Step 3: Extract structural consistency features from the geometrically consistent real-time graph and the reference graph, and perform cross-correlation operation based on the structural consistency features to obtain the cross-correlation matrix; Step 4: Determine the position of the main peak in the cross-correlation coefficient matrix, and divide the cross-correlation coefficient matrix into eight non-overlapping regions extending along the horizontal, vertical and diagonal directions with the main peak position as the center. Select the local maximum value with the largest amplitude in each region as the secondary peak. Step 5: Verify the consistency of the registration results based on the proportional relationship between the amplitude of the main peak and the amplitudes of each secondary peak. Step 6: When the verification result meets the preset conditions, determine the coordinates of the first center point based on the position of the main peak of cross-correlation, and at the same time, map the center position of the real-time graph to the reference graph based on the homography matrix to determine the coordinates of the second center point. Step 7: Calculate the Euclidean distance between the coordinates of the first center point and the coordinates of the second center point, and use this Euclidean distance as the evaluation result of the registration error; The structural consistency features are obtained through phase consistency feature extraction to suppress the influence of multimodal imaging differences on cross-correlation results and enhance the significance of the main cross-correlation peak relative to the side peaks; The peak-to-peak ratio criterion is achieved by comparing whether the ratio of the amplitude of the main peak to the amplitude of the secondary peak exceeds a preset threshold, thereby reducing the probability of misjudgment.

2. The method as described in claim 1, characterized in that, The homography matrix is ​​calculated from the pixel correspondence between the real-time image and the reference image. The geometric transformation is an affine transformation performed on the real-time image based on the homography matrix, so that the two images maintain consistency in spatial structure.

3. An image registration and verification method based on cross-correlation primary and secondary peak spatial constraints, implementing the multi-scene adaptation cross-correlation peak ratio registration verification and dual-center error evaluation method as described in any one of claims 1-2, characterized in that, include: Perform cross-correlation on the two images after geometric consistency processing to obtain the cross-correlation matrix; The main peak with the largest amplitude is determined in the cross-correlation matrix, and the cross-correlation matrix is ​​divided into multiple regions in eight directions with the location of the main peak as the center. In each of the aforementioned regions, a local maximum with the largest amplitude is determined as the secondary peak; Based on the proportional relationship between the amplitude of the main peak and the amplitudes of each secondary peak, a peak ratio criterion is constructed, and the validity of the image registration result is determined based on the peak ratio criterion.

4. The method as described in claim 3, characterized in that, The region division method is to divide the cross-correlation matrix into eight triangular regions with the main peak position as the origin, along the horizontal, vertical and two diagonal directions, with each region not overlapping with the others.

5. A system for multi-scenario adaptation cross-correlation peak ratio registration verification and dual-center error assessment implementing the multi-scenario adaptation cross-correlation peak ratio registration verification and dual-center error assessment method as described in any one of claims 1-2, characterized in that, include: The registration unit is used to perform feature matching between the real-time image and the reference image, and to calculate the homography matrix; A geometric transformation unit is used to perform geometric transformation on the real-time graph based on the homography matrix to eliminate the difference in geometric distortion between the graph and the reference graph. The feature processing unit is used to extract structural consistency features between the real-time graph and the reference graph, and to perform cross-correlation operations to obtain the cross-correlation matrix; The peak analysis unit is used to determine the position of the main peak in the cross-correlation matrix and to determine multiple secondary peaks according to a predetermined spatial rule; The verification unit is used to determine whether the registration was successful based on the peak ratio relationship between the main peak and the secondary peak. The error assessment unit is used to determine the coordinates of the first center point and the second center point when registration is successful, and to calculate the Euclidean distance between them as the registration error.

6. The system as described in claim 5, characterized in that, The coordinates of the first center point are determined by the position of the main cross-correlation peak in the reference image, representing the center of the region in the reference image that is most similar to the structure of the real-time image.

7. The system as described in claim 5, characterized in that, The second center point coordinates are obtained by mapping the center position of the real-time graph to the reference graph using a homography matrix.

8. The system as described in claim 5, characterized in that, When the coordinates of the first center point and the second center point completely coincide, the registration result is correct. When there is a misalignment between the two, their Euclidean distance is used to reflect the magnitude of the registration error.