Land space large scene image matching positioning method and device based on multi-level clustering balance, electronic equipment and storage medium

By optimizing the distribution of matching points through multi-level clustering equilibrium processing and robust estimation algorithms, the problems of cumulative error, matching point imbalance and noise interference in large-scale land space scenarios are solved, achieving high-precision and robust visual positioning capabilities. It is applicable to land space monitoring in fields such as natural resources, agriculture and rural areas, water conservancy, transportation, forestry and emergency management.

CN121746748BActive Publication Date: 2026-05-15CHONGQING FUPEIHE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING FUPEIHE TECH CO LTD
Filing Date
2026-02-24
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as large cumulative errors, unbalanced distribution of matching points, severe noise interference, and difficulty in cross-scale matching in large-scale territorial spatial scenarios, resulting in insufficient positioning accuracy and robustness.

Method used

A multi-level clustering equalization method is adopted, which replaces one or more preset position images with panoramic images to construct a three-level processing architecture of global distance equalization, local density equalization and global re-equalization. This optimizes the spatial distribution of matching points and introduces a robust estimation algorithm to filter outliers, thereby improving the quality and distribution uniformity of matching points.

Benefits of technology

It significantly improves the solution accuracy and positioning reliability of the spatial transformation model, and realizes high-precision, consistent visual positioning capability in complex national spatial scenarios, meeting the monitoring needs of natural resources, forestry, emergency response and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746748B_ABST
    Figure CN121746748B_ABST
Patent Text Reader

Abstract

The application discloses a land space large-scene image matching positioning method and device based on multi-stage clustering balancing, electronic equipment and a storage medium. The method first acquires a panoramic image collected by a pan-tilt camera and a remote sensing image of a corresponding region, and performs multi-scale cross-modal feature matching on the acquired panoramic image and remote sensing image to obtain an initial matching point set. Then, multi-stage clustering balancing processing including global distance balancing, local density balancing and global rebalancing in sequence is performed, and a final matching point set with uniform distribution and noise suppression is output. Finally, a spatial conversion model is solved based on the set to realize accurate positioning of a camera picture to a geographic coordinate. The application effectively solves the problems of cross-scale image correlation difficulty, uneven distribution of matching points and cumulative error, and significantly improves the positioning accuracy and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision and geographic information technology, specifically to a method, device, electronic device and storage medium for spatial mapping of cameras and maps for precise two-way positioning in large-scale land space scenes such as natural resources, agriculture and rural areas, water conservancy, ecological environment, transportation, forestry, and emergency response. Background Technology

[0002] In large-scale automated monitoring of national land space in fields such as natural resources, agriculture and rural areas, water conservancy, ecological environment, transportation, forestry, and emergency management, PTZ cameras deployed on towers or poles acquire images through 360-degree rotation, providing visual support for the identification and handling of various events. These tasks generally require accurately mapping the pixel coordinates of targets to geospatial coordinates while identifying them, thereby achieving the operational goal of "seeing and locating accurately." The core of this process lies in establishing a high-precision spatial mapping relationship between remote sensing maps (covering a wide national territory) and regional 1x1 maps (focusing on local details), and image matching technology is the key link supporting the construction of this mapping relationship.

[0003] Currently, the mainstream positioning matching schemes in the industry are mainly divided into two categories: single preset position matching strategy and multi-preset position matching strategy. The former fits unified positioning model parameters to images and remote sensing base maps under a fixed viewpoint; the latter models multiple preset views separately in order to improve coverage. Although these two methods have been widely used, in real-world application environments with large-scale, multi-scale, and highly complex land use scenarios, they have both exposed a series of insurmountable technical defects, which seriously restrict positioning accuracy and system robustness.

[0004] Specifically, existing technologies have the following five core problems:

[0005] First, single-preset position matching suffers from significant cumulative error. This method relies on a global positioning model fitted from a single viewpoint. As the positioning range expands or the task continues to run, the model parameter errors accumulate, causing the positioning deviation to increase continuously with distance or time. This fails to meet the requirement of consistent high-precision positioning across the entire territory in large-scale spatial scenarios.

[0006] Second, multi-preset matching is prone to model overfitting. Although the multi-preset strategy alleviates the accumulated error to some extent, the spatial distribution of matching points is extremely uneven because each preset is matched independently. Matching points are over-accumulated in textured areas (such as urban building clusters and dense farmland), while matching points are severely scarce in textured areas (such as deserts, water bodies, and open tidal flats). This imbalance leads to the localization model overfitting to high-density regional features during training, resulting in poor generalization ability and difficulty in adapting to complex scenes such as terrain undulations, vegetation occlusion, and building occlusion.

[0007] Third, defects in the quality of matching points directly reduce the accuracy of the model solution. Even with a multi-preset position scheme, the original matching results still generally contain a large number of mismatched noise points, and the feature distribution is highly uneven. When these low-quality matching points are used as input to solve spatial transformation models such as homography matrices or affine transformations, they will significantly amplify the solution error, ultimately leading to large deviations in the positioning results, making it impossible to provide reliable data support for downstream analysis, early warning, or disposal tasks.

[0008] Fourth, cross-scale image matching and association are difficult. Remote sensing images and regional 1x1 maps differ significantly in resolution, viewpoint, illumination, and semantic representation, representing typical cross-modal and cross-scale image pairs. Traditional matching methods lack effective global constraint mechanisms, easily leading to matching ambiguities or missed matches, making it difficult to establish stable and reliable correspondences. Although some studies have attempted to introduce panoramic images for stitching and matching, they have not fully exploited the global geometric consistency advantages inherent in panoramic images, nor have they designed multi-scale adaptive matching mechanisms for national land space scenarios, thus failing to fundamentally solve the cross-scale association problem.

[0009] Fifth, there is a lack of effective balancing mechanisms for imbalanced matching point distribution. Existing balancing strategies are mostly limited to single clustering (such as K-Means densification) or simple point expansion, failing to construct a multi-level balancing architecture that coordinates "global-local" factors. This makes it impossible to simultaneously consider the quantity balance across large areas and the density consistency within micro-regions. More importantly, existing methods have not established a quantitative correlation between the balancing of matching points and the final positioning accuracy, resulting in the balancing process becoming merely a formality and failing to truly improve positioning performance.

[0010] In summary, the large cumulative error, imbalanced distribution of matching points, noise interference, difficulties in cross-scale correlation, and the resulting model overfitting and low solution accuracy are intertwined and mutually reinforcing, constituting the main bottlenecks of current large-scale land spatial positioning and matching technology. The industry urgently needs a comprehensive solution that can simultaneously eliminate the sources of cumulative error, effectively fuse cross-scale images, and systematically optimize the spatial distribution quality of matching points. Summary of the Invention

[0011] To overcome the problems of large cumulative error, unbalanced distribution of matching points, severe noise interference, and difficulty in cross-scale matching in existing technologies, this application proposes a method, device, electronic device, and storage medium for matching and locating large-scale images of national land space based on multi-level clustering equilibrium. The aim is to cut off the error accumulation path through panoramic image matching and construct a three-level processing architecture of global distance equilibrium, local density equilibrium, and global re-equilibrium to achieve a uniform and robust distribution of matching points across the entire scene, thereby significantly improving the solution accuracy and positioning reliability of the spatial transformation model.

[0012] The first objective of this application is to provide a method for matching and locating large-scale spatial images of national territory based on multi-level clustering equilibrium.

[0013] The aforementioned objective of this application is achieved through the following technical solution:

[0014] A method for matching and locating large-scale spatial images of national territory based on multi-level clustering equilibrium, the method comprising:

[0015] Acquire panoramic images captured by a gimbal camera, and acquire remote sensing images that at least partially cover the geographical area corresponding to the panoramic images;

[0016] Multi-scale cross-modal feature matching is performed on the panoramic image and the remote sensing image to obtain an initial set of matching points;

[0017] A multi-level clustering equalization process is performed on the initial set of matching points to output a final set of matching points with uniform distribution and suppressed noise. The multi-level clustering equalization process includes global distance equalization, local density equalization, and global re-equalization performed sequentially.

[0018] The spatial transformation model is solved based on the final set of matching points, and the spatial transformation model is used to achieve accurate positioning from camera view to geographic coordinates.

[0019] Preferably, acquiring the panoramic image captured by the gimbal camera and acquiring a remote sensing image that at least partially covers the geographical area corresponding to the panoramic image includes:

[0020] Multiple 1x magnification images of a region captured by a gimbal camera within a 360-degree horizontal range are stitched together to generate a panoramic image with an isometric rectangular projection.

[0021] Based on the geographical location information of the gimbal camera, a remote sensing image within a preset radius centered on that location is captured from the digital map service as the target remote sensing image.

[0022] Preferably, the step of performing multi-scale cross-modal feature matching between the panoramic image and the remote sensing image to obtain an initial set of matching points includes:

[0023] The panoramic image and the remote sensing image are scaled to multiple preset resolution levels respectively;

[0024] At each resolution level, feature detection and description algorithms are used to extract key points and their descriptors from the image;

[0025] Cross-image matching is performed based on descriptor similarity to obtain matching point pairs at various resolutions;

[0026] The initial set of matching points is formed by merging matching point pairs from all resolution levels.

[0027] Preferably, the global distance equalization step includes:

[0028] Using the pixel coordinates of the matching points as features, the K-Means clustering algorithm is used to divide the matching points into K global sub-regions, where K is an integer greater than or equal to 8;

[0029] Calculate the maximum number of matching points in K global sub-regions;

[0030] For sub-regions where the number of matching points is less than the maximum number of matching points, random resampling is performed, combined with a small Gaussian perturbation, to expand the number of matching points in each sub-region so that the number of matching points in each sub-region is consistent.

[0031] Preferably, the step of local density equalization includes:

[0032] For each global sub-region after global distance equalization, the DBSCAN density clustering algorithm is used to further divide it into several local clusters based on the local point cloud density;

[0033] For each local cluster, if the number of matching points is lower than the average number of points in the global sub-region, then neighborhood interpolation or random perturbation is used to supplement the local cluster to balance the local density.

[0034] Preferably, the global rebalancing step includes:

[0035] All matching points after local density equalization are treated as a whole, and the DBSCAN density clustering algorithm is used again for global clustering to form multiple global density clusters.

[0036] Calculate the maximum number of matching points in all global density clusters;

[0037] For global density clusters with insufficient matching points, random sampling is performed within the neighborhood of existing matching points in the cluster, and Gaussian noise perturbation conforming to a preset standard deviation is applied to the sampled points to generate new virtual matching points. This process continues until the number of matching points in the cluster reaches the maximum number of matching points, thus obtaining the final set of matching points.

[0038] Preferably, before performing multi-level clustering equalization processing on the initial set of matching points, the method further includes:

[0039] A robust estimation algorithm is used to perform regional geometric consistency checks on the initial set of matching points, and abnormal matching points with reprojection errors exceeding the threshold are removed. The resulting set of pre-filtered matching points is then used as input for multi-level clustering equilibrium processing.

[0040] The second objective of this application is to provide a large-scale spatial image matching and positioning device based on multi-level clustering equilibrium.

[0041] The second objective of this application is achieved through the following technical solution:

[0042] A spatial image matching and localization device based on multi-level clustering equilibrium, the device comprising:

[0043] The image acquisition module is used to acquire panoramic images captured by the gimbal camera and to acquire remote sensing images that at least partially cover the geographical area corresponding to the panoramic images.

[0044] The image matching module is used to perform multi-scale cross-modal feature matching between the panoramic image and the remote sensing image to obtain an initial set of matching points;

[0045] The equalization processing module is used to perform multi-level cluster equalization processing on the initial matching point set to output a final matching point set with uniform distribution and suppressed noise. The multi-level cluster equalization processing includes global distance equalization, local density equalization and global re-equalization performed sequentially.

[0046] The positioning execution module is used to solve the spatial transformation model based on the final matching point set, and to use the spatial transformation model to achieve accurate positioning from the camera image to geographic coordinates.

[0047] Preferably, the image acquisition module is specifically used for:

[0048] Multiple 1x magnification images of a region captured by a gimbal camera within a 360-degree horizontal range are stitched together to generate a panoramic image with an isometric rectangular projection.

[0049] Based on the geographical location information of the gimbal camera, a remote sensing image within a preset radius centered on that location is captured from the digital map service as the target remote sensing image.

[0050] Preferably, the image matching module is specifically used for:

[0051] The panoramic image and the remote sensing image are scaled to multiple preset resolution levels respectively;

[0052] At each resolution level, feature detection and description algorithms are used to extract key points and their descriptors from the image;

[0053] Cross-image matching is performed based on descriptor similarity to obtain matching point pairs at various resolutions;

[0054] The initial set of matching points is formed by merging matching point pairs from all resolution levels.

[0055] Preferably, the global distance equalization step includes:

[0056] Using the pixel coordinates of the matching points as features, the K-Means clustering algorithm is used to divide the matching points into K global sub-regions, where K is an integer greater than or equal to 8;

[0057] Calculate the maximum number of matching points in K global sub-regions;

[0058] For sub-regions where the number of matching points is less than the maximum number of matching points, random resampling is performed, combined with a small Gaussian perturbation, to expand the number of matching points in each sub-region so that the number of matching points in each sub-region is consistent.

[0059] Preferably, the step of local density equalization includes:

[0060] For each global sub-region after global distance equalization, the DBSCAN density clustering algorithm is used to further divide it into several local clusters based on the local point cloud density;

[0061] For each local cluster, if the number of matching points is lower than the average number of points in the global sub-region, then neighborhood interpolation or random perturbation is used to supplement the local cluster to balance the local density.

[0062] Preferably, the global rebalancing step includes:

[0063] All matching points after local density equalization are treated as a whole, and the DBSCAN density clustering algorithm is used again for global clustering to form multiple global density clusters.

[0064] Calculate the maximum number of matching points in all global density clusters;

[0065] For global density clusters with insufficient matching points, random sampling is performed within the neighborhood of existing matching points in the cluster, and Gaussian noise perturbation conforming to a preset standard deviation is applied to the sampled points to generate new virtual matching points. This process continues until the number of matching points in the cluster reaches the maximum number of matching points, thus obtaining the final set of matching points.

[0066] Preferably, the device further includes:

[0067] The outlier filtering module is used to perform regional geometric consistency checks on the initial set of matching points using a robust estimation algorithm, remove outlier matching points whose reprojection errors exceed a threshold, and obtain a pre-filtered set of matching points as input for multi-level clustering equilibrium processing.

[0068] The third objective of this application is to provide an electronic device.

[0069] The aforementioned objective three of this application is achieved through the following technical solution:

[0070] An electronic device, comprising:

[0071] It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the large-scale spatial image matching and localization method based on multi-level clustering equilibrium as described in any one of the first objectives of this application.

[0072] The fourth objective of this application is to provide a computer-readable storage medium.

[0073] The fourth objective of this application is achieved through the following technical solution:

[0074] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the large-scale spatial image matching and localization method based on multi-level clustering equilibrium as described in any of the first objectives of this application.

[0075] Compared with the prior art, this application has the following beneficial effects:

[0076] 1. By using panoramic images instead of single or multiple preset position images as the matching basis, a unified spatial mapping relationship covering a 360-degree field of view is established, avoiding the problem of cumulative error in model parameters caused by independent modeling of multiple preset positions, thereby improving the consistency of localization results in large scenes.

[0077] 2. By constructing a multi-level clustering equalization processing architecture—global distance equalization, local density equalization, and global re-equalization—the spatial distribution of matching points can be optimized hierarchically. The innovation of this architecture lies in its targeted modification of traditional clustering algorithms, upgrading them from simple classification tools to dedicated processing engines for matching point equalization: At the macro scale, a globally uniform sampling initialization strategy is introduced into the K-Means algorithm, effectively avoiding initial cluster centers biased towards densely populated matching point regions. Combined with a point expansion mechanism based on the maximum cluster size, the number of matching points across different regions is balanced. At the micro scale, an adaptive neighborhood radius is adopted for the DBSCAN algorithm, enabling it to flexibly adapt to local density differences between dense and sparse texture regions. Neighborhood interpolation or perturbation point supplementation strategies alleviate local clustering or scarcity problems. Finally, the modified DBSCAN is used again for global density clustering, and a unified equalization operation is applied to complete the final verification of the consistency of the distribution across the entire scene. This three-level architecture, through the above algorithmic innovations, systematically improves the uniformity of the overall distribution of matching points and effectively suppresses the overfitting phenomenon in spatial transformation models caused by data imbalance.

[0078] 3. Before the multi-level clustering equalization process, a robust estimation algorithm is introduced to filter out the outliers of the initial matching points, effectively eliminating mismatched points with large reprojection errors. Combined with the subsequent equalization process, the set of matching points input to the solution stage of the spatial transformation model (such as homography matrix) has higher geometric consistency and reasonable distribution, thereby improving the stability and reliability of the model solution.

[0079] 4. By integrating multi-scale cross-modal feature matching and multi-level equalization mechanism, it can still effectively establish a reliable correspondence between remote sensing images and regional 1x maps even when there are significant differences in scale, viewpoint and texture. It also maintains a certain ability to supplement matching points in areas with sparse texture (such as water areas and deserts) or severely occluded areas, thus improving the applicability and robustness of the method in diverse and complex scenarios of national land space. Attached Figure Description

[0080] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0081] Figure 1 This is a flowchart illustrating a method for matching and locating large-scale spatial images based on multi-level clustering equilibrium in one embodiment of this application.

[0082] Figure 2 This is a schematic diagram of the structure of a large-scene image matching and positioning device for national land space based on multi-level clustering equilibrium in one embodiment of this application;

[0083] Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation

[0084] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0085] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described below are merely illustrative. For example, the division of units and modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or modules can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, and can be electrical, mechanical, or other forms.

[0086] In addition, each functional unit in the various embodiments of this application can be integrated into a single processor, or each unit can be a separate device, or two or more units can be integrated into a single device; each functional unit in the various embodiments of this application can be implemented in hardware or in the form of hardware plus software functional units.

[0087] Those skilled in the art will understand that all or part of the steps of the following method embodiments can be implemented by program instructions and related hardware. The aforementioned program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, they perform the steps of the following method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0088] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "a plurality of" or "several" means two or more, unless otherwise explicitly specified.

[0089] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. It should be noted that the core innovative ideas of this application can be implemented through a variety of specific technical means.

[0090] On the one hand, the core of this application lies in using stitched equilateral rectangular projection panoramic images and remote sensing images to perform multi-scale resolution matching, in order to construct a global spatial positioning matching relationship covering a 360-degree field of view. The implementation of this idea does not depend on a specific matching algorithm; any feature matching method that can effectively associate heterogeneous, cross-scale images can be applied to this framework.

[0091] On the other hand, the core of this application lies in proposing a general matching point preprocessing architecture, namely, a four-stage pipeline of "original matching point filtering → distance clustering global balancing → density clustering local balancing → density clustering global rebalancing", combined with adaptive balancing strategies (such as downsampling in dense matching point regions and controllably expanding sparse regions) to systematically solve the problem of unbalanced matching point distribution. The implementation of this architecture is not limited to specific clustering or balancing algorithms; the K-Means and DBSCAN algorithms used in subsequent embodiments of this application are merely preferred implementations of this idea.

[0092] The technical solution of this application will now be described in detail with reference to the accompanying drawings, using a specific preferred embodiment.

[0093] like Figure 1 As shown in the figure, this application provides a method for matching and locating large-scale spatial images of national territory based on multi-level clustering equilibrium. The method may include the following steps:

[0094] S1, acquire a panoramic image captured by a gimbal camera, and acquire a remote sensing image that at least partially covers the geographic area corresponding to the panoramic image;

[0095] In this embodiment, a forestry resource monitoring and forest fire early warning scenario is used as an example. The pan-tilt camera to be located is deployed on top of a forest fire lookout tower, at a height of approximately 30 meters. This camera has 360° horizontal rotation and ±90° pitch adjustment capabilities, and can periodically (e.g., every 5 minutes) or triggered by a fire warning to acquire a panoramic image covering its entire field of view. Simultaneously, the system acquires satellite remote sensing images with a resolution of 0.5 meters from the National Geographic Information Public Service Platform (e.g., Tianditu), covering a radius of 10 kilometers centered on the lookout tower, ensuring as complete a coverage of the geographical area observable by the panoramic image as possible. By acquiring the panoramic image captured by the pan-tilt camera and the corresponding remote sensing image, a basic data source is provided for subsequent cross-scale image matching.

[0096] S2, perform multi-scale cross-modal feature matching on panoramic images and remote sensing images to obtain an initial set of matching points;

[0097] This embodiment employs deep learning feature matching algorithms (such as SuperPoint / SuperGlue) for cross-viewpoint and cross-modal feature point matching. After successful matching, for each keypoint in the panoramic image, its pixel coordinates are converted into observation angle information using the camera intrinsic parameter matrix; for the corresponding keypoint in the remote sensing image, its pixel coordinates are used to query the remote sensing image metadata to obtain the geographic coordinate information in the WGS84 coordinate system. This forms an initial set of matching points that associates the observation angle with geographic coordinates. This set serves as the original basis for constructing spatial mapping relationships.

[0098] S3, perform multi-level cluster equalization on the initial set of matching points to output a final set of matching points with uniform distribution and suppressed noise. The multi-level cluster equalization process includes global distance equalization, local density equalization and global re-equalization performed in sequence.

[0099] While multi-scale cross-modal matching yields an initial set of matching points representing the correspondence between panoramic and remote sensing images, this initial set often suffers from two fundamental flaws due to the complexity of the national spatial scene (e.g., extremely uneven distribution of ground textures, numerous occlusions, or homogeneous areas). First, the spatial distribution is severely unbalanced, with matching points excessively dense in texture-rich areas (e.g., cities, building clusters) and extremely scarce or even absent in sparsely textured areas (e.g., water bodies, deserts, farmland). Second, it contains significant noise or mismatched points. If such low-quality matching points are directly used to solve subsequent spatial transformation models (e.g., homography matrices), the model will overfit high-density areas due to data bias, leading to significant deviations in localization results in sparse areas and poor overall robustness.

[0100] Therefore, in this embodiment, a multi-level clustering equalization process is introduced to systematically optimize the spatial distribution quality of matching points. The basic principle is that it does not rely on manually set rules or simple downsampling, but rather uses a hierarchical, adaptive clustering equalization mechanism to structurally reorganize the matching point set. This mechanism can identify and quantify the uneven distribution of matching points at different spatial scales, and through controllable data augmentation or adjustment strategies, actively improve the global uniformity of the point set while preserving the original geometric consistency, and simultaneously suppress the influence of noise.

[0101] The final set of matching points output after multi-level clustering and equalization processing is more evenly and reasonably distributed in space, and noise is effectively suppressed. This high-quality point set serves as the input for subsequent spatial transformation model solving, significantly improving the stability, accuracy, and generalization ability of the spatial transformation model. This ensures consistent and reliable geographic coordinate positioning results regardless of where the target appears in the image (whether it's a densely built-up urban area or an open beach), ultimately achieving truly accurate positioning in large-scale national spatial scenarios.

[0102] S4 solves the spatial transformation model based on the final matching point set, and uses the spatial transformation model to achieve accurate positioning from camera view to geographic coordinates.

[0103] In this embodiment, the homography matrix is ​​solved using a high-quality set of matching points after multi-level equalization processing, employing algorithms such as Direct Linear Transform (DLT). In practical applications, when the system detects a suspected fire target in the monitoring screen, the homography matrix can be used to convert the target's pixel coordinates into its corresponding geographic coordinates (longitude and latitude) in real time and accurately, thereby supporting the rapid response and precise dispatch of fire-fighting forces.

[0104] As described above, this embodiment provides a method for matching and locating large-scale spatial images based on multi-level clustering equilibrium. First, a panoramic image captured by a gimbal camera is acquired, along with a remote sensing image that at least partially covers the corresponding geographic area of ​​the panoramic image. Then, multi-scale cross-modal feature matching is performed between the panoramic image and the remote sensing image to obtain an initial set of matching points. Next, a multi-level clustering equilibrium process, including sequential global distance equilibrium, local density equilibrium, and global re-equilibrium, is performed on the initial set of matching points to output a final set of matching points with uniform distribution and suppressed noise. Finally, a spatial transformation model is solved based on the final set of matching points, and the spatial transformation model is used to achieve accurate positioning from the camera image to geographic coordinates.

[0105] The large-scale spatial image matching and localization method based on multi-level clustering equilibrium in this embodiment has the following advantages:

[0106] 1. By using panoramic images as the matching benchmark, the traditional scheme that relies on one or more discrete preset positions is replaced. The panoramic image covers a 360° continuous field of view and establishes a globally unified spatial mapping relationship with the remote sensing image. This avoids the problem of model parameter errors accumulating as the spatial range expands due to independent modeling of multiple preset positions, thereby achieving consistent and reliable positioning accuracy across the entire large scene.

[0107] 2. A multi-scale, cross-modal feature matching strategy is adopted, which can adaptively associate ground features at different scales and from different perspectives in remote sensing images (macroscopic, top-down view) and panoramic images (microscopic, side view). This mechanism effectively alleviates the problems of matching failure or insufficient points caused by large differences in image modalities and uneven texture distribution (such as dense urban areas and desert water areas), and significantly improves the matching success rate and feature coverage in complex land space scenarios.

[0108] 3. A multi-level clustering balancing process is introduced to structurally optimize the initial matching point set output from the previous step. This process addresses the two core issues of imbalance in the number of matching points between regions and uneven density within regions through macro-level, micro-level, and global verification, while simultaneously suppressing noise interference. The resulting final matching point set is uniformly distributed and geometrically consistent, fundamentally preventing the spatial transformation model from overfitting to local high-density regions due to data bias, and significantly improving the model's generalization ability.

[0109] 4. The spatial transformation model is solved based on high-quality matching points optimized through multi-level clustering and equalization. Because the input data possesses uniform distribution, noise robustness, and geometric consistency, the solved model more realistically reflects the mapping relationship between the camera's field of view and geographic space. Therefore, in practical applications, regardless of whether the target appears in a texture-rich or sparse area of ​​the image, stable and accurate geographic coordinates can be obtained, meeting the core requirements of "visible and accurate location" for land and space monitoring operations such as natural resources, forestry, and emergency response.

[0110] In summary, this embodiment, through a combination of panoramic image matching, multi-scale feature fusion, and multi-level clustering equilibrium, collaboratively solves multiple bottlenecks in existing technologies, such as cumulative error, imbalance of matching points, noise sensitivity, and difficulties in cross-scale association, thereby achieving high-precision, robust, and consistent visual positioning capabilities across large-scale national land space scenarios.

[0111] In one embodiment, before performing multi-level cluster balancing on the initial set of matching points, the following steps may also be included:

[0112] A robust estimation algorithm is used to perform regional geometric consistency checks on the initial set of matching points, and abnormal matching points with reprojection errors exceeding the threshold are removed. The resulting set of matching points after preliminary filtering is used as the input for multi-level clustering equilibrium processing.

[0113] In this embodiment, the initial set of matched points typically contains a certain proportion of outliers. To prevent these noise points from interfering with subsequent equalization processing, a robust estimation algorithm (such as RANSAC (Random Sample Consensus)) is first used to filter them. Specifically, the camera field of view is divided into several regions, and RANSAC is run independently in each region to fit a local homography model. The reprojection error threshold is set to 3 pixels, and the confidence level is 0.9. All matched points with a reprojection error exceeding 3 pixels are identified as outliers and removed. The remaining inliers constitute the initially filtered set of matched points, which is then used as input for multi-level clustering equalization processing.

[0114] This embodiment effectively removes significant noise and mismatches in the initial matching results by introducing robust estimation for outlier filtering before the equalization process. This ensures the quality of the data input to the multi-level clustering equalization module and prevents noisy points from being incorrectly included in the equalization calculation, thereby improving the effectiveness of the entire equalization process and the reliability of the final positioning results.

[0115] In one embodiment, step S1, acquiring a panoramic image captured by a gimbal camera and acquiring a remote sensing image that at least partially covers the geographic area corresponding to the panoramic image, includes:

[0116] Multiple 1x magnification images of a region captured by a gimbal camera within a 360-degree horizontal range are stitched together to generate a panoramic image with an isometric rectangular projection.

[0117] Based on the geographical location information of the gimbal camera, remote sensing images within a preset radius centered on that location are captured from the digital map service as the target remote sensing image.

[0118] In this embodiment, the gimbal camera acquires 36 images of the region with a resolution of 360×360 pixels during a complete 360-degree horizontal scan. These images cover a range with a horizontal viewing angle of 360 degrees and a vertical viewing angle of approximately 90 degrees. Using a mature image stitching algorithm (such as OpenCV's Stitcher module), these 36 images are seamlessly stitched into a 3600×1800 pixel equirectangular projection panoramic image. This projection method maintains the orthogonality of meridians and parallels, facilitating subsequent processing. Simultaneously, based on the GPS coordinates of the camera's installation location (e.g., latitude and longitude of 120.045°, 30.098°), the system requests and downloads a high-resolution remote sensing image centered at that point with a radius of 5 kilometers from an online map service API as the target remote sensing map. This remote sensing map covers a land area of ​​approximately 5 km², including various typical land features such as farmland, mountains, ponds, buildings, and forests, meeting the complex needs of actual land space monitoring.

[0119] This embodiment stitches together multiple 1x1 area images into a panoramic image, which not only obtains a unified image benchmark covering the entire field of view, but also makes full use of the global geometric consistency advantage of panoramic images, creating conditions for establishing a stable and unambiguous cross-scale correspondence with remote sensing images. It fundamentally solves the problem that traditional preset position matching is difficult to associate feature distortion and scale differences between heterogeneous images.

[0120] In one embodiment, step S2, performing multi-scale cross-modal feature matching between the panoramic image and the remote sensing image to obtain an initial set of matching points includes:

[0121] Scaling panoramic and remote sensing images to multiple preset resolution levels;

[0122] At each resolution level, feature detection and description algorithms are used to extract key points and their descriptors from the image;

[0123] Cross-image matching is performed based on descriptor similarity to obtain matching point pairs at various resolutions;

[0124] The matching point pairs from all resolution levels are merged to form the initial matching point set.

[0125] In large-scale spatial scenarios, there is a significant scale difference between remote sensing images (overhead, macroscopic) and pan-tilt-zoom (PTZ) panoramic images (side view, local): the same feature may appear as a single pixel in a remote sensing image, but occupy tens or even hundreds of pixels in a close-up panoramic image. If feature matching is performed only at a single scale, the algorithm is prone to failure due to scale mismatch. At high resolution, it cannot associate distant small targets, while at low resolution, it will lose close-up details.

[0126] To this end, this embodiment simultaneously scales the panoramic image and the remote sensing image to multiple preset resolution levels (such as 600, 900, 1200, 1500, etc.), and independently executes a complete feature detection, description, and matching process at each scale. Thus, at higher resolutions (such as 1500), the algorithm can capture more subtle structures and effectively match distant small-scale features (such as ridgelines and field boundaries); at lower resolutions (such as 600), the image is smoothed and denoised, making it easier to extract stable large-scale structural features (such as building outlines, road intersections, and the edges of large bodies of water). Finally, the matching point pairs obtained at each scale are fused to form an initial matching point set covering multiple scales and distances. This set naturally possesses stronger scene representativeness and geometric diversity.

[0127] In this embodiment, fusing the matching point pairs obtained at different scales is a process of deduplication, verification, and integration, aiming to construct an initial set of matching points that is both free of redundancy and comprehensive. Specifically, in this embodiment, the specific fusion process of the multi-scale matching point pair fusion mechanism is as follows:

[0128] 1. Collect matching results at each scale: First, obtain the set of matching point pairs output from the matching process at each preset resolution level (e.g., 600, 900, 1200, 1500). Each point pair consists of the coordinates of a keypoint in a panoramic image and the coordinates of a corresponding keypoint in a remote sensing image, along with a confidence score for the matching point pair. The confidence score of the matching point pair is calculated based on the similarity of the feature descriptors.

[0129] 2. Establish a unified matching point index: Since images at different scales originate from scaling of the same original panoramic and remote sensing images, the coordinates of matching points at all scales are uniformly back-projected back into the coordinate system of the original high-resolution image using the coordinate mapping relationship between scales. This step ensures that matching points from different scales can be compared and processed under the same spatial reference.

[0130] 3. Perform non-maximum suppression (NMS) or neighborhood-based deduplication:

[0131] For point pairs that are repeatedly detected and matched at multiple scales in the same physical location (e.g., in the original image coordinate system, the Euclidean distance is less than a preset threshold, such as 3-5 pixels), a deduplication operation will be performed.

[0132] Deduplication strategies typically employ confidence-based selection, specifically retaining the matching pair with the highest confidence score and eliminating other duplicates with lower confidence scores. This effectively avoids redundancy in matching points caused by multi-scale detection, ensuring the simplicity of the final set.

[0133] 4. Integration to Form the Initial Matching Point Set: After deduplication, all remaining unique matching point pairs are integrated to form the final initial matching point set. Since different resolution levels have complementary scene perception capabilities—high resolution excels at capturing small-scale details in the distance, while low resolution is better at extracting large-scale stable structures in the foreground—effective matching points from each scale naturally possess differences and complementarity in spatial location and land cover type. The fusion process retains these non-redundant matching results from different scales, ensuring that the final initial matching point set simultaneously covers large targets in the foreground and small targets in the distance, as well as high-texture and low-texture areas. This results in a comprehensive matching point set covering multiple scales, distances, and rich geometric diversity. Each point pair in this set represents a highly reliable correspondence of land cover features successfully associated at a specific scale.

[0134] Through the above fusion mechanism, the final initial set of matching points has the following characteristics:

[0135] Multi-scale coverage: It simultaneously includes fine structures captured at high resolution (such as distant field ridges and paths) and stable large structures extracted at low resolution (such as houses and main roads), achieving effective coverage of features at different scales.

[0136] Multi-distance coverage: Large targets nearby and small targets far away can be matched at their respective most suitable scales, thereby expanding the geographical range of effective matching.

[0137] It has stronger scene representativeness and geometric diversity: the set is no longer a skewed result under a single scale, but a comprehensive product that integrates information from the whole scene and multiple scales. It includes both dense points in texture-rich areas and key anchor points in texture-sparse areas, providing a high-quality and diverse data foundation for subsequent equalization processing.

[0138] The multi-scale cross-modal feature matching method implemented in this paper has the following advantages:

[0139] 1. Traditional single-scale matching methods struggle to simultaneously consider both distant and near-field features, often resulting in large-area matching gaps in complex scenes. This embodiment significantly expands the effective matching geographic range and feature types through multi-scale parallel matching, particularly improving the matching capability for distant small targets and near-field large structures.

[0140] 2. After integrating multi-scale results, the initial set of matching points not only increases in total quantity, but also covers more comprehensively in space. It includes dense points in texture-rich areas such as urban building clusters, as well as key anchor points in texture-sparse areas such as deserts, water areas, and forest areas, providing a more sufficient data foundation for subsequent equalization processing.

[0141] 3. In scenarios with shading, changes in illumination, or seasonal changes in land cover (such as the tillage and harvest seasons in farmland), single-scale features may become completely ineffective. Multi-scale strategies provide redundant observation paths, so even if a matching at one scale fails, other scales may still successfully establish a correspondence, thereby improving the fault tolerance and stability of the overall matching process.

[0142] 4. The quality of the initial matching point set directly determines the final positioning accuracy. The multi-scale fused matching point set generated in this embodiment has higher completeness, representativeness, and geometric consistency, providing a better data foundation for subsequent multi-level clustering and equalization processing, thereby ensuring the accuracy and reliability of the spatial transformation model solution.

[0143] In summary, this embodiment effectively solves the problem of feature association difficulties caused by huge scale differences in large-scale territorial spatial scenarios by introducing a multi-scale cross-modal matching mechanism, which helps to achieve more accurate positioning with high robustness across all scenarios.

[0144] This embodiment adopts a multi-scale resolution matching strategy, which enables the matching process to take into account the features of land features at different distances and scales, significantly improving the matching success rate and the number of matching points in complex land space scenarios, and providing richer and more representative initial data for subsequent equilibrium processing and high-precision model solving.

[0145] In one embodiment, the global distance equalization step in step S3 includes:

[0146] Using the pixel coordinates of the matching points as features, the K-Means clustering algorithm is used to divide the matching points into K global sub-regions, where K is an integer greater than or equal to 8;

[0147] Calculate the maximum number of matching points in K global sub-regions;

[0148] For sub-regions with fewer matching points than the maximum number of matching points, random resampling combined with a small Gaussian perturbation is performed to expand the number of matching points in each sub-region, so that the number of matching points in each sub-region is consistent.

[0149] The global distance equalization in this embodiment belongs to the first level of multi-level cluster equalization processing. Its core objective is to solve the problem of severely uneven distribution of matching points on a macro scale in large-scale territorial spatial scenarios. That is, some geographical areas (such as urban building clusters) have excessively dense matching points, while other areas (such as water areas, deserts, and farmland) have extremely sparse or even missing matching points. If this phenomenon of excessively large differences in the number of matching points between large areas is not addressed, it will directly lead to overfitting of high-density areas when solving subsequent spatial transformation models (such as homography matrices), resulting in an imbalance in global positioning accuracy.

[0150] To achieve coarse-grained equilibrium at the macro level, this embodiment proposes a task-optimized K-Means clustering equilibrium mechanism, which specifically includes the following two stages:

[0151] Phase 1: Global distance partitioning based on a modified K-Means algorithm.

[0152] Input data: The set of matching points after preliminary filtering. This set has eliminated obvious false matches, but still retains the original uneven distribution characteristics.

[0153] Parameter configuration:

[0154] Cluster size K: Based on the panoramic image resolution (e.g., 3600×1800 pixels) and scene complexity (e.g., diversity of terrain features), K is dynamically set to an integer greater than or equal to 8, with an optimal range of 8~16 and a default value of 12. This range ensures that the sub-regions are neither too coarse (K too small) nor too fragmented (K too large).

[0155] Random seed: Set to a fixed integer between 30 and 50 (e.g., seed=42) to ensure that the clustering results are completely reproducible under the same input, meeting the stability requirements of industrial systems.

[0156] Clustering features: The pixel coordinates (u,v) of each matching point are used as the unique clustering feature, which directly reflects its spatial position in the camera's field of view.

[0157] Key points of algorithm modification:

[0158] Initialization strategy: Abandoning the standard K-Means random or k-means++ initialization, a global uniform sampling strategy is adopted, that is, K initial centers are selected from the set of matching points after preliminary filtering in a grid or equidistant manner on the image plane, and the initial centers are forced to be uniformly distributed in the entire field of view, which effectively avoids the clustering centers from all clustering in high-density areas due to the skewness of the original point set.

[0159] Convergence criteria: Set dual termination criteria: cluster center offset ≤ 1 pixel (aligned with positioning accuracy requirements) or maximum number of iterations = 100, thus balancing efficiency and accuracy.

[0160] Output: After the algorithm converges, the set of matching points after initial filtering is divided into K mutually exclusive subsets. (1≤i≤K), each Corresponding to a global sub-region partitioned based on Euclidean distance, containing There are 1 matching point.

[0161] This stage proactively constructs a set of macro-level partitions that cover the entire field of view and have a reasonable spatial distribution, providing a structured foundation for subsequent equilibration operations.

[0162] Phase 2: Global equilibrium processing based on the maximum number benchmark.

[0163] Calculate the baseline number: Traverse all K sub-regions, find the sub-region with the most matching points, and record its number of points as . Adaptive point augmentation:

[0164] For each sub-region, the corresponding subset ,like Then calculate the required expansion quantity. This is achieved using a random resampling + small Gaussian perturbation strategy, as detailed below: From Random selection with replacement inside A number of points were selected, and the standard deviation was applied to each selected point. Gaussian noise perturbation of pixels generates new virtual matching points;

[0165] Add the new point to the currently processed subset. To form an equilibrium subset Its points are exactly .

[0166] Merge Output: Merge all balanced sub-regions to obtain a globally spatially balanced set of matching points. ;

[0167] Simultaneously, the corresponding matching points on the remote sensing image are updated to form a complete set of point pairs. This serves as the input for the next level (local density equilibrium).

[0168] This stage uses controlled data augmentation to force consistency in the number of matching points in each macroscopic region without introducing external information or disrupting local geometric relationships.

[0169] This embodiment solves the fundamental problem of excessive differences in the number of matching points between large regions through the above two-stage processing, enabling different terrain regions such as cities, villages, water areas, and mountains to have a balanced sample contribution ability in model training. This avoids the spatial transformation model from being overly biased towards large regions with dense matching points during the solution process; it also avoids regional overfitting caused by data skewness in the spatial transformation model, significantly improving the model's generalization ability and positioning stability across the entire 360° field of view. The modified K-Means used is not a simple call to library functions, but rather an engineering innovation in initialization, convergence, and reproducibility tailored to task requirements, ensuring the rationality of the partitioning results and the effectiveness of the balancing operation.

[0170] In one embodiment, the local density equalization step in step S3 includes:

[0171] For each global sub-region after global distance equalization, the DBSCAN density clustering algorithm is used to further divide it into several local clusters based on the local point cloud density;

[0172] For each local cluster, if the number of matching points is lower than the average number of points in the global sub-region, then neighborhood interpolation or random perturbation is used to supplement the local cluster to balance the local density.

[0173] The local density equalization in this embodiment is the second level of multi-level cluster equalization, aiming to solve the problem of uneven distribution of matching points at the microscale that still exists after global distance equalization. Although global distance equalization ensures that the total number of matching points is consistent across major regions (such as cities and water bodies), within a single region, due to significant differences in landform textures (for example, the same sub-region may contain both high-texture building walls and low-texture water surfaces), matching points may still exhibit localized dense clustering or sparse missingness. Without processing, the spatial transformation model may still be overly sensitive to high-density local structures (such as the corner of a building) during the solution process, leading to a decrease in positioning accuracy in adjacent sparse areas (such as water surfaces and open spaces).

[0174] Therefore, this embodiment proposes a local refinement and equalization mechanism based on density clustering, which specifically includes the following two stages:

[0175] Phase 1: Local density partitioning based on adaptive DBSCAN.

[0176] Input data: Each sub-region after global distance equalization For each processing unit (1≤i≤K), a set is formed by extracting matching points from its source image (panoramic image). .

[0177] Clustering algorithm selection: DBSCAN (Density-Based Spatial Clustering of Applications with Noise) density clustering algorithm was adopted. Unlike K-Means equidistant clustering, DBSCAN can automatically identify clusters of arbitrary shapes and effectively distinguish between high-density and low-density areas, making it particularly suitable for handling local texture heterogeneity.

[0178] Adaptive setting of key parameters:

[0179] Neighborhood radius The value is dynamically adjusted based on local texture density, ranging from 1 to 3 pixels. In areas with dense texture (such as building facades and road intersections), the spacing between feature points is small, therefore... Use a smaller value (e.g., 1 pixel) to avoid excessive merging; in areas with sparse texture (such as water surfaces or wastelands), the spacing between feature points is large, therefore... Take a larger value (such as 3 pixels) to ensure that sparse points can be effectively clustered.

[0180] Minimum number of core points This setting allows even isolated points to form independent clusters, preventing valid matching points in sparse areas from being misjudged as noise and discarded because they do not meet the density threshold, thus ensuring the preservation of key anchor points.

[0181] Clustering execution: Using the pixel coordinates (u,v) of the matching points as features, ... Perform DBSCAN clustering and output several local density clusters. (1≤j≤t, where t is the number of clusters generated adaptively). For example, in a sub-region containing buildings and water surfaces, the algorithm can automatically divide the area into a high-density "building wall cluster" and a low-density "water surface cluster".

[0182] The essence of this stage is not noise reduction, but to identify and separate local structural units with different density characteristics within the region, providing operational objects for fine-grained equalization.

[0183] Second stage: Local equilibrium processing based on regional average benchmark.

[0184] Calculate the equilibrium benchmark: for the current global sub-region Calculate the average number of points among all matching points within it. .

[0185] Adaptive point-filling strategy:

[0186] For each local cluster If the number of matching points is lower than the average number of points mentioned above, a point replacement operation is initiated, which generates new points using neighborhood linear interpolation or random perturbation.

[0187] Neighborhood interpolation: Performs linear interpolation between existing matching points to generate geometrically reasonable virtual points, which is suitable for structurally continuous regions (such as wall edges).

[0188] Random perturbation: Adding small Gaussian noise to the neighborhood of a point, suitable for sparse regions with no clear structure.

[0189] Merged Output: All equalized local clusters are merged to obtain a set of matching points with consistent local density. This set is then integrated into a full-scene matching point set E, and the corresponding points on the remote sensing image are updated synchronously to form a set of equalized point pairs. This serves as the input for the next level (global rebalancing).

[0190] This stage, while preserving the original local structural semantics, bridges the local density gap through controllable data augmentation, enabling high-texture areas and low-texture areas to have similar matching point densities at the local scale.

[0191] This embodiment's local density equalization, through the aforementioned two-stage processing, delves into the microscopic level, resolving the problem of local accumulation or scarcity of matching points within the same region, and compensating for the fine-grained unevenness that global equalization cannot address; through adaptive DBSCAN parameter settings (especially... and This approach balances the fine segmentation of high-density structures with the preservation of key points in sparse regions, avoiding the shortcomings of traditional clustering methods that miss or lose points in sparse areas. It employs a point-filling strategy that combines neighborhood interpolation with random perturbation, ensuring the geometric rationality of newly added points while enhancing adaptability to unstructured regions. It significantly suppresses the model's tendency to overfit on local high-density features, enabling the localization results to maintain stable accuracy in mixed terrain scenarios such as building clusters, water bodies, and farmland, thereby improving the overall robustness of the system.

[0192] In one embodiment, the global rebalancing step in step S3 includes:

[0193] All matching points after local density equalization are treated as a whole, and the DBSCAN density clustering algorithm is used again for global clustering to form multiple global density clusters.

[0194] Calculate the maximum number of matching points in all global density clusters;

[0195] For global density clusters with insufficient matching points, random sampling is performed within the neighborhood of existing matching points in the cluster, and Gaussian noise perturbation with a preset standard deviation is applied to the sampled points to generate new virtual matching points. This process continues until the number of matching points in the cluster reaches the maximum number of matching points, resulting in the final set of matching points.

[0196] The global rebalancing in this embodiment is the third level of the multi-level clustering balancing process, and it is the final verification and fine-tuning stage of the entire balancing process. Although the first two levels of processing (global distance balancing and local density balancing) have achieved preliminary optimization of the number and density of matching points at the macro-region partitioning and micro-structure internal levels, respectively, the results may still have residual density inconsistencies at the full scene scale because the two levels of processing use different clustering logics (the former is based on distance partitioning, and the latter is based on local density refinement). For example, some sparse clusters generated by local balancing may have a low overall number of points, while dense clusters formed by high-texture regions may have a high number of points, resulting in the final matching point set not achieving an ideal uniform distribution state from a global perspective.

[0197] To completely solve this problem, this embodiment proposes a global consistency rebalancing mechanism based on density clustering. The core idea is to treat the matching point set after the first two rounds of processing as a whole, re-cluster it globally based on density, and force the balance of the number of points between each density cluster, thereby outputting a highly uniform and noise-suppressed final matching point set.

[0198] Phase 1: Global density clustering based on DBSCAN.

[0199] Input data: Receive the set of matching points E of the entire scene after local density equalization processing. This set has completed preliminary equalization at the macro-region and local structure levels.

[0200] Clustering algorithm selection: The DBSCAN density clustering algorithm was used again, but this time it was applied to the entire map area, rather than a single sub-region. DBSCAN can adaptively identify clusters of arbitrary shapes based on the local density of points, making it naturally suitable for capturing complex density patterns formed by terrain differences (such as urban blocks, farmland grids, water edges, and mountain treelines) in the entire scene.

[0201] Parameter adaptive configuration:

[0202] Neighborhood radius The value is dynamically adjusted based on the local texture density, ranging from 1 to 3 pixels. A smaller value (e.g., 1 pixel) is used in areas with dense texture (such as building clusters) to finely distinguish adjacent structures; a larger value (e.g., 3 pixels) is used in areas with sparse texture (such as water bodies or deserts) to ensure that sparse points can effectively aggregate into clusters.

[0203] Minimum number of core points This setting ensures that even isolated valid matching points (such as distant power line tower corners) are preserved as independent clusters, preventing them from being lost as noise during the final equalization stage and guaranteeing the integrity of critical positioning anchors.

[0204] Random seed: A fixed setting to ensure that the clustering results are reproducible and meet the stability requirements of industrial deployment.

[0205] Clustering Execution and Output: Using the pixel coordinates (u,v) of matching points as features, perform DBSCAN clustering on set E, and output t global density clusters. (1≤j≤t), each It corresponds to a geographical region or land cover type unit with similar local density characteristics (such as "dense building cluster", "sparse farmland cluster", "linear road cluster", etc.).

[0206] This stage is not about repeating local operations, but about re-evaluating the density distribution pattern of matching points from a holistic perspective, providing a unified and self-consistent basis for the final equilibrium.

[0207] Phase 2: Global rebalancing based on the maximum number of cluster points.

[0208] Calculate the equilibrium baseline: Traverse all global density clusters Find the cluster with the most matching points, and denote its number of points as . ;

[0209] Adaptive point augmentation: for each global density cluster :

[0210] like Then calculate the required expansion quantity. ;

[0211] Random sampling is performed within the neighborhood of existing matching points in the cluster (e.g., within 2 pixels), and a preset standard deviation is applied to each sampling point (e.g., ...). Gaussian noise perturbation (pixels) is used to generate new virtual matching points;

[0212] Add new points To form an equilibrium cluster So that its points are exactly equal to .

[0213] Merge Output: Merge all equalized density clusters to obtain the final set of matching points. ;

[0214] Synchronously update the corresponding matching points on the remote sensing image to form the final set of equilibrium point pairs. , which serves as the input for solving the space transformation model.

[0215] This stage eliminates density deviations that may have been left over from the previous two stages of processing by performing a final global-scale point alignment, ensuring a highly consistent and uniform distribution of matching points across the entire scene.

[0216] The global rebalancing in this embodiment, as the final step in the multi-level balancing architecture, achieves final consistency verification and correction of the distribution of matching points across the entire scene, effectively bridging the global density residuals that may be caused by differences in hierarchical processing logic. Through full-map DBSCAN clustering, it naturally integrates terrain semantics and geometric density, making the balancing operation more in line with the actual distribution patterns of land features in the national land space. It uses a small Gaussian perturbation (σ=0.3 pixels) to expand the points, increasing the number of points while maintaining the original geometric relationships to the maximum extent, avoiding the introduction of significant positioning deviations. It ensures that the final set of matching points F has both global uniformity and local rationality, providing the optimal and most robust data foundation for solving subsequent spatial transformation models (such as homography matrices), significantly improving positioning accuracy and reliability.

[0217] like Figure 2 As shown, this application provides a large-scale spatial image matching and localization device based on multi-level clustering equilibrium. The device may include:

[0218] The image acquisition module 201 is used to acquire panoramic images captured by the pan-tilt camera and to acquire remote sensing images that at least partially cover the geographical area corresponding to the panoramic images.

[0219] Image matching module 202 is used to perform multi-scale cross-modal feature matching between panoramic images and remote sensing images to obtain an initial set of matching points;

[0220] The equalization processing module 203 is used to perform multi-level cluster equalization processing on the initial set of matching points to output a final set of matching points with uniform distribution and suppressed noise. The multi-level cluster equalization processing includes global distance equalization, local density equalization and global re-equalization performed sequentially.

[0221] The positioning execution module 204 is used to solve the spatial transformation model based on the final matching point set, and to use the spatial transformation model to achieve accurate positioning from the camera image to geographic coordinates.

[0222] In one embodiment, the image acquisition module 201 is specifically used for:

[0223] Multiple 1x magnification images of a region captured by a gimbal camera within a 360-degree horizontal range are stitched together to generate a panoramic image with an isometric rectangular projection.

[0224] Based on the geographical location information of the gimbal camera, remote sensing images within a preset radius centered on that location are captured from the digital map service as the target remote sensing image.

[0225] In one embodiment, the image matching module 202 is specifically used for:

[0226] Scaling panoramic and remote sensing images to multiple preset resolution levels;

[0227] At each resolution level, feature detection and description algorithms are used to extract key points and their descriptors from the image;

[0228] Cross-image matching is performed based on descriptor similarity to obtain matching point pairs at various resolutions;

[0229] The matching point pairs from all resolution levels are merged to form the initial matching point set.

[0230] In one embodiment, the global distance equalization steps include:

[0231] Using the pixel coordinates of the matching points as features, the K-Means clustering algorithm is used to divide the matching points into K global sub-regions, where K is an integer greater than or equal to 8;

[0232] Calculate the maximum number of matching points in K global sub-regions;

[0233] For sub-regions with fewer matching points than the maximum number of matching points, random resampling combined with a small Gaussian perturbation is performed to expand the number of matching points in each sub-region, so that the number of matching points in each sub-region is consistent.

[0234] In one embodiment, the local density equalization step includes:

[0235] For each global sub-region after global distance equalization, the DBSCAN density clustering algorithm is used to further divide it into several local clusters based on the local point cloud density;

[0236] For each local cluster, if the number of matching points is lower than the average number of points in the global sub-region, then neighborhood interpolation or random perturbation is used to supplement the local cluster to balance the local density.

[0237] In one embodiment, the global rebalancing steps include:

[0238] All matching points after local density equalization are treated as a whole, and the DBSCAN density clustering algorithm is used again for global clustering to form multiple global density clusters.

[0239] Calculate the maximum number of matching points in all global density clusters;

[0240] For global density clusters with insufficient matching points, random sampling is performed within the neighborhood of existing matching points in the cluster, and Gaussian noise perturbation with a preset standard deviation is applied to the sampled points to generate new virtual matching points. This process continues until the number of matching points in the cluster reaches the maximum number of matching points, resulting in the final set of matching points.

[0241] In one embodiment, the device may further include:

[0242] The outlier filtering module is used to perform regional geometric consistency checks on the initial set of matching points using a robust estimation algorithm, and remove outlier matching points whose reprojection errors exceed a threshold. The resulting set of matching points after preliminary filtering is used as input for multi-level clustering equilibrium processing.

[0243] It should be noted that the working principles and technical effects of the above embodiments of the large-scale spatial image matching and positioning device based on multi-level clustering equilibrium are the same as those of the corresponding embodiments of the above-mentioned large-scale spatial image matching and positioning method based on multi-level clustering equilibrium, and will not be repeated here.

[0244] like Figure 3 As shown, this application provides an electronic device 3, which includes a memory 301, a processor 302, and a computer program 303 stored in the memory 301 and executable on the processor 302. The memory 301 and the processor 302 are connected via a bus 304. When the processor 302 executes the computer program 303, it implements the large-scale spatial image matching and positioning method based on multi-level clustering equilibrium as described in the above method embodiment of this application.

[0245] Specifically, the electronic device 3 can be an intelligent device with memory and processor, such as an industrial control computer, PC, or smart mobile terminal, or a computer component with memory and processor, such as a CPU or GPU.

[0246] In this embodiment, electronic device 3 is a remote monitoring terminal such as a PC or smartphone.

[0247] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the large-scene image matching and positioning method for territorial space based on multi-level clustering equilibrium as described in the above-described method embodiments of this application.

[0248] It should be noted that the large-scale spatial image matching and positioning device, electronic device and computer-readable storage medium based on multi-level clustering equilibrium in the above embodiments have the same working principle and technical effect as the large-scale spatial image matching and positioning method based on multi-level clustering equilibrium in the above embodiments, and will not be repeated here.

[0249] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0250] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0251] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0252] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for matching and locating large-scale spatial images of national territory based on multi-level clustering equilibrium, characterized in that, The method includes: Acquire panoramic images captured by a gimbal camera, and acquire remote sensing images that at least partially cover the geographical area corresponding to the panoramic images; Multi-scale cross-modal feature matching is performed on the panoramic image and the remote sensing image to obtain an initial set of matching points; A multi-level clustering equalization process is performed on the initial set of matching points to output a final set of matching points with uniform distribution and suppressed noise. The multi-level clustering equalization process includes global distance equalization, local density equalization, and global re-equalization performed sequentially. The spatial transformation model is solved based on the final set of matching points, and the spatial transformation model is used to achieve accurate positioning from camera view to geographic coordinates; in, The step of performing multi-scale cross-modal feature matching between the panoramic image and the remote sensing image to obtain an initial set of matching points includes: The panoramic image and the remote sensing image are scaled to multiple preset resolution levels respectively; At each resolution level, feature detection and description algorithms are used to extract key points and their descriptors from the image; Cross-image matching is performed based on descriptor similarity to obtain matching point pairs at various resolutions; The initial set of matching points is formed by merging matching point pairs from all resolution levels. The process of solving the spatial transformation model based on the final set of matching points and using the spatial transformation model to achieve accurate positioning from camera view to geographic coordinates includes: Based on the final set of matching points, the solution space transformation model is obtained using the direct linear transformation algorithm, and the spatial transformation model is used to achieve accurate positioning from camera image to geographic coordinates. The spatial transformation model is a homography matrix.

2. The method for matching and locating large-scale spatial images based on multi-level clustering equilibrium according to claim 1, characterized in that, The acquisition of panoramic images captured by a gimbal camera, and the acquisition of remote sensing images that at least partially cover the geographical area corresponding to the panoramic images, include: Multiple 1x magnification images of a region captured by a gimbal camera within a 360-degree horizontal range are stitched together to generate a panoramic image with an isometric rectangular projection. Based on the geographical location information of the gimbal camera, a remote sensing image within a preset radius centered on that location is captured from the digital map service as the target remote sensing image.

3. The method for matching and locating large-scale spatial images based on multi-level clustering equilibrium according to claim 1, characterized in that, The steps of the global distance equalization include: Using the pixel coordinates of the matching points as features, the K-Means clustering algorithm is used to divide the matching points into K global sub-regions, where K is an integer greater than or equal to 8; Calculate the maximum number of matching points in K global sub-regions; For sub-regions where the number of matching points is less than the maximum number of matching points, random resampling is performed, combined with a small Gaussian perturbation, to expand the number of matching points in each sub-region so that the number of matching points in each sub-region is consistent.

4. The method for matching and locating large-scale spatial images based on multi-level clustering equilibrium according to claim 1, characterized in that, The steps of local density equalization include: For each global sub-region after global distance equalization, the DBSCAN density clustering algorithm is used to further divide it into several local clusters based on the local point cloud density; For each local cluster, if the number of matching points is lower than the average number of points in the global sub-region, then neighborhood interpolation or random perturbation is used to supplement the local cluster to balance the local density.

5. The method for matching and locating large-scale spatial images of a territory based on multi-level clustering equilibrium as described in claim 1, characterized in that, The steps of the global rebalancing include: All matching points after local density equalization are treated as a whole, and the DBSCAN density clustering algorithm is used again for global clustering to form multiple global density clusters. Calculate the maximum number of matching points in all global density clusters; For global density clusters with insufficient matching points, random sampling is performed within the neighborhood of existing matching points in the cluster, and Gaussian noise perturbation conforming to a preset standard deviation is applied to the sampled points to generate new virtual matching points. This process continues until the number of matching points in the cluster reaches the maximum number of matching points, thus obtaining the final set of matching points.

6. The method for matching and locating large-scale spatial images based on multi-level clustering equilibrium according to any one of claims 1-5, characterized in that, Before performing multi-level clustering equalization on the initial set of matching points, the method further includes: A robust estimation algorithm is used to perform regional geometric consistency checks on the initial set of matching points, and abnormal matching points with reprojection errors exceeding the threshold are removed. The resulting set of pre-filtered matching points is then used as input for multi-level clustering equilibrium processing.

7. A spatial image matching and positioning device for large-scale land use based on multi-level clustering equilibrium, characterized in that, The device includes: The image acquisition module is used to acquire panoramic images captured by the gimbal camera and to acquire remote sensing images that at least partially cover the geographical area corresponding to the panoramic images. The image matching module is used to perform multi-scale cross-modal feature matching between the panoramic image and the remote sensing image to obtain an initial set of matching points; The equalization processing module is used to perform multi-level cluster equalization processing on the initial matching point set to output a final matching point set with uniform distribution and suppressed noise. The multi-level cluster equalization processing includes global distance equalization, local density equalization and global re-equalization performed sequentially. The positioning execution module is used to solve the spatial transformation model based on the final matching point set, and to use the spatial transformation model to achieve accurate positioning from the camera image to geographic coordinates; in, The image matching module is specifically used for: The panoramic image and the remote sensing image are scaled to multiple preset resolution levels respectively; At each resolution level, feature detection and description algorithms are used to extract key points and their descriptors from the image; Cross-image matching is performed based on descriptor similarity to obtain matching point pairs at various resolutions; The initial set of matching points is formed by merging matching point pairs from all resolution levels. The positioning execution module is specifically used for: Based on the final set of matching points, the solution space transformation model is obtained using the direct linear transformation algorithm, and the spatial transformation model is used to achieve accurate positioning from camera image to geographic coordinates. The spatial transformation model is a homography matrix.

8. An electronic device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the large-scale spatial image matching and localization method based on multi-level clustering equilibrium as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the large-scene image matching and localization method for national land space based on multi-level clustering equilibrium as described in any one of claims 1-6.