Large-scene unmanned aerial vehicle positioning method and system based on multi-source cross-modal image hierarchical matching

Through the multi-source cross-modal image hierarchical matching method, the problem of insufficient accuracy of drone positioning in complex environments is solved, and high-precision drone positioning under low GNSS is achieved.

CN120293157AActive Publication Date: 2025-07-11HUNAN UNIV

Patent Information

Application Number
CN202510781864.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing drone positioning technology is difficult to provide accurate navigation information in complex environments, especially when GNSS is affected by electromagnetic interference and INS errors are divergent, visual positioning technology has the problem of insufficient positioning accuracy.

Method used

The multi-source cross-modal image hierarchical matching method is adopted to generate tile map pyramids with different ground resolutions, correct the real-time image direction, perform image matching and feature point correction, and combine PnP pose solution to achieve accurate positioning of the drone.

Benefits of technology

When GNSS provides low-precision positions, visual assistance enables high-precision positioning of the drone in large scenarios, reducing electromagnetic interference and environmental impact, and improving positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120293157A_ABST
    Figure CN120293157A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scene unmanned aerial vehicle positioning method and system based on multi-source cross-modal image hierarchical matching, and the method comprises the steps: generating image pyramids which are different in ground resolution and are represented by a tile map according to a satellite optical map of an unmanned aerial vehicle airport region; acquiring a real-shot image of the unmanned aerial vehicle airport area and correcting the real-shot image to be in the same direction as the satellite optical map; determining a tile map with the same ground target according to the ground resolution of the real shot image; splicing the found tile map with a neighborhood tile map to obtain a reference map; and performing image matching on the actually-shot image and the reference image by using a hierarchical matching method, correcting the geographic coordinates of the feature points in the actually-shot image according to the corresponding feature points, and performing pose calculation to solve the pose of the unmanned aerial vehicle. The invention aims to realize the vision-assisted positioning of the unmanned aerial vehicle, so that the accurate positioning of the unmanned aerial vehicle in a large scene is realized through vision assistance under the condition that a positioning system can only provide a low-precision position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a large-scene UAV positioning method and system for multi-source cross-modal image hierarchical matching. Background Art

[0002] Accurate navigation and positioning of unmanned aerial vehicles (UAVs) is crucial for UAV operations.

[0003] In traditional UAV positioning, the Global Navigation Satellite System (GNSS) and the Inertial Navigation System (INS) are commonly used positioning systems. GNSS relies on the communication information between artificial satellites and ground receivers to calculate the position and speed of the carrier. It has the advantages of being lightweight and inexpensive, and can provide real-time position. It is currently the most widely used navigation system. However, in a complex environment, UAVs are subject to electromagnetic interference and shielding during flight, and the external environment also has a great impact on GNSS, which makes it difficult to provide accurate positions for UAVs, reducing the survival rate and mission completion rate of UAVs. INS relies on the angular acceleration and linear acceleration provided by gyroscopes and accelerometers and their integration over time to calculate speed and displacement. It has the advantages of being immune to electromagnetic interference, small in size, and low in cost. However, due to its positioning characteristics, its error will diverge over time, resulting in the inability to provide long-term and high-precision positioning information.

[0004] Therefore, how to provide accurate navigation information for UAVs through auxiliary positioning means and overcome electromagnetic interference, shielding, and environmental impacts has become a key technical problem to be solved urgently in UAV positioning technology. With the development of computer vision technology, visual positioning technology has developed vigorously. Based on obtaining the accurate geographical location of the pixels in the captured real image, existing methods have provided a method using Perspective-n-Points (PnP) pose estimation to calculate the pose of the carrier. Since visual positioning technology is a fully autonomous positioning technology, it is not affected by electromagnetic interference, and the error does not diverge over time, having the advantage of working all-weather and all-day. However, there are errors in obtaining the geographical location of the pixels in the captured real image in the existing technology, resulting in the problem of insufficient positioning accuracy in visual positioning technology. Summary of the Invention

[0005] The technical problem to be solved by the present invention: Aiming at the above problems of the existing technology, a large-scene UAV positioning method and system for multi-source cross-modal image hierarchical matching are provided. The present invention aims to realize the visual-aided positioning of UAVs, so as to achieve the precise positioning of UAVs in a large scene through visual assistance when the positioning system can only provide low-precision positions.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A large-scene UAV positioning method for multi-source cross-modal image hierarchical matching, including: S1. Generate an image pyramid presented in a tile map with different ground resolutions based on the satellite optical map of the UAV airport area, obtain the actual shot image of the UAV airport area and correct it to the same direction as the satellite optical map; S2. Determine the corresponding tile map level according to the ground resolution of the actual shot image, and retrieve in the tile map corresponding to the tile map level according to the geographical coordinates of the central pixel point of the actual shot image to find the tile map with the same ground target as the actual shot image; S3. Stitch the found tile map with its neighboring tile maps to obtain a reference map; S4. Perform image matching on the actual shot image and the reference map using a hierarchical matching method to obtain the corresponding feature points and pixel transformation matrix between the actual shot image and the reference map; S5. According to the accurate geographical coordinates of the reference map and its corresponding feature points in the actual shot image, correct the geographical coordinates of the feature points with position errors in the actual shot image to obtain geographical coordinates with higher accuracy; S6. Solve the pose of the UAV according to the geographical coordinates of the feature points after correction in the actual shot image.

[0007] Optionally, determining the corresponding tile map level according to the ground resolution of the actual shot image in step S2 includes: S2.1. Calculate the distance between the actual shot image and each tile map according to the following formula d ; , In the above formula, and are the ground resolutions of the two axes of the actual shot image respectively, and are the ground resolutions of the two axes of the i-th tile map respectively; S2.2. Find the tile map level corresponding to the tile map with the smallest distance d as the tile map level corresponding to the ground resolution of the actual shot image.

[0008] Optionally, when retrieving in the tile map corresponding to the tile map level according to the geographical coordinates of the central pixel point of the actual shot image in step S2 to find the tile map with the same ground target as the actual shot image, the functional expression of the constraint condition satisfied by the tile map with the same ground target as the actual shot image is: , In the above formula, is the geographical coordinate of the central pixel point of the real-shot image, is the upper left coordinate of the i-th tile map, is the lower right coordinate of the i-th tile map.

[0009] Optionally, in step S4, using the hierarchical matching method to match the real-shot image and the reference image includes: S4.1, downsample the real-shot image and the reference image respectively, and then use a preset matching algorithm to match the downsampled reference image and the infrared map to obtain a set of feature point pairs , where is the k-th feature point of the reference image, is the k-th feature point of the real-shot image, is the total number of feature point pairs; S4.2, cluster the feature points of the real-shot image and the reference image respectively to obtain the cluster center coordinates; S4.3, restore the cluster center coordinates to the original ground resolution image coordinates respectively, and crop out a specified-size neighborhood as the region of interest according to the original ground resolution image coordinates to obtain the region-of-interest matching image; S4.4, use a preset matching algorithm to match the region-of-interest matching images of the real-shot image and the reference image to obtain a set of feature point pairs , where is the k-th feature point of the region-of-interest matching image of the reference image, is the k-th feature point of the region-of-interest matching image of the real-shot image, is the total number of feature point pairs.

[0010] Optionally, when downsampling the real-shot image and the reference image respectively in step S4.1, it includes performing times downsampling on the reference image, and performing downsampling on the real-shot image according to the following formula: , In the above formula, and are the downsampling ratios in the x and y axis directions of the real-shot image, and are the ground resolutions of the two axes of the reference image, and are the ground resolutions of the two axes of the real-shot image, is the downsampling multiple performed on the reference image.

[0011] Optionally, the functional expressions for clustering the feature points of the real-shot image and the reference image respectively in step S4.2 to obtain the cluster center coordinates are: , , In the above formula, is the coordinate of the clustering center of the real-shot image, and are respectively the x coordinate and y coordinate of the i-th feature point of the real-shot image, is the coordinate of the clustering center of the reference image, and are respectively the x coordinate and y coordinate of the i-th feature point of the reference image, is the total number of feature point pairs.

[0012] Optionally, the function expressions for respectively restoring the clustering center coordinates to the original ground resolution image coordinates in step S4.3 are: , , In the above formula, is the coordinate of the clustering center of the real-shot image after being restored to the original ground resolution image coordinates, and are the downsampling ratios in the x and y axis directions of the real-shot image, is the coordinate of the clustering center of the reference image after being restored to the original ground resolution image coordinates, is the downsampling multiple performed on the reference image, is the coordinate of the clustering center of the real-shot image, is the coordinate of the clustering center of the reference image.

[0013] In addition, the present invention also provides a large-scale UAV positioning system for multi-source cross-modal image hierarchical matching, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the large-scale UAV positioning method for multi-source cross-modal image hierarchical matching.

[0014] In addition, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is used to be programmed or configured by a microprocessor to execute the large-scale UAV positioning method for multi-source cross-modal image hierarchical matching.

[0015] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the large-scale UAV positioning method for multi-source cross-modal image hierarchical matching through a processor.

[0016] Compared with the prior art, the present invention mainly has the following advantages: In the large-scene UAV positioning method of multi-source cross-modal image hierarchical matching of the present invention, when the main positioning system of the UAV (such as GNSS) can only provide low-precision positions, precise positioning of the UAV in a large scene is achieved through visual assistance. The satellite optical map reference map area is determined based on the low-precision position provided by the main positioning system, and then the reference map and the actual captured image are subjected to hierarchical image matching to obtain the geographical location coordinates of the feature points of the actual captured image accurately. Finally, the UAV pose is solved. It is less affected by electromagnetic interference and the external environment, and has high positioning accuracy, which can provide navigation guarantee for the UAV to perform tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic diagram of the basic process of the method in the embodiment of the present invention.

[0019] Figure 2 It is a schematic diagram of different-level tile maps in the embodiment of the present invention, where (a) and (b) are a group of tile maps with different resolutions at the 12th layer, and (c) and (d) are a group of tile maps with different resolutions at the 14th layer.

[0020] Figure 3 It is a schematic diagram of a reference map and an infrared image pair in the embodiment of the present invention, where (a) and (b) are the first group of reference map and infrared image pair, and (a) is the reference map and (b) is the actual captured image; (c) and (d) are the second group of reference map and infrared image pair, and (c) is the reference map and (d) is the actual captured image.

[0021] Figure 4 It is a schematic diagram of the result of using hierarchical matching for the reference map and the infrared image in the embodiment of the present invention, where (a) is the result of rough matching, (b) is the result of fine matching, and (c) is the final result. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0023] As Figure 1As shown in the figure, the large - scene UAV positioning method for multi - source cross - modal image hierarchical matching in this embodiment includes: S1. Generate an image pyramid presented in tile maps with different ground resolutions based on the satellite optical map of the UAV airport area, obtain the actual shot image of the UAV airport area, and correct it to the same direction as the satellite optical map; S2. Determine the corresponding tile map level according to the ground resolution of the actual shot image, and retrieve in the tile map corresponding to the tile map level according to the geographical coordinates of the central pixel point of the actual shot image to find the tile map with the same ground target as the actual shot image; S3. Stitch the found tile map with its neighboring tile maps to obtain a reference map; S4. Use the hierarchical matching method to match the actual shot image and the reference map to obtain the corresponding feature points and pixel transformation matrix between the actual shot image and the reference map; S5. According to the accurate geographical coordinates of the reference map and its corresponding feature points in the actual shot image, correct the geographical coordinates of the feature points with position errors in the actual shot image to obtain higher - precision geographical coordinates; S6. Solve the pose of the UAV according to the geographical coordinates of the feature points in the actual shot image after correction.

[0024] Step S1 of this embodiment is a step of image pre - processing. In step S1, the UAV airport area can be divided according to the flight speed and duration of the UAV, generally within plus or minus 20 longitude and latitude. After generating an image pyramid presented in tile maps with different ground resolutions based on the satellite optical map (including geographical location information) of the UAV airport area, the size of each tile map is, and the geographical coordinates (WGS84 coordinate system) of the upper - left pixel of the tile map are known. The image pyramid is generally 18 layers, that is, it includes 18 kinds of ground resolutions, and the highest is about 0.5m. For example Figure 2 , shows examples of tile maps at different levels, where (a) and (b) are a group of tile maps with different resolutions at the 12th layer, and (c) and (d) are a group of tile maps with different resolutions at the 14th layer. The actual shot image used in this embodiment is an actual shot infrared image, and its purpose is to realize UAV positioning at night. If only considering the daytime, a common visible - light actual shot image can be used. Based on the flight direction of the UAV, correct the actual shot infrared image to the due - north direction (the satellite optical maps are all in the due - north direction) to reduce the difficulty of image matching. Given the flight direction of the UAV, which is the direction of the infrared image, rotate the actual shot image to the due - north direction.

[0025] In step S2 of this embodiment, determining the corresponding tile map level according to the ground resolution of the actual shot image includes: S2.1. Calculate the distance between the actual shot image and each tile map according to the following formula d ; , In the above formula, and are the two-axis ground resolutions of the real-shot image respectively, and are the two-axis ground resolutions of the i-th tile map respectively; S2.2. Find the tile map level corresponding to the tile map with the smallest distance d as the tile map level corresponding to the ground resolution of the real-shot image.

[0026] In step S2 of this embodiment, determining the corresponding tile map level according to the ground resolution of the real-shot image is achieved by first determining, based on the ground resolution of the infrared map, the tile map level with the smallest resolution difference for geographical matching, which can reduce the scale gap between the reference map and the real-shot map and make the subsequent image matching less difficult. When retrieving in the tile map corresponding to the tile map level according to the geographical coordinates of the central pixel point of the real-shot image in step S2 of this embodiment to find the tile map with the same ground target as the real-shot image, the functional expression of the constraint conditions satisfied by the tile map with the same ground target as the real-shot image is: , In the above formula, is the geographical coordinate of the central pixel point of the real-shot image, is the upper left coordinate of the i-th tile map, is the lower right coordinate of the i-th tile map. Given the upper left geographical coordinate of the infrared map, the two-axis ground resolutions and , and the pixel resolution , then the center point coordinate is: , , Similarly, given that the upper left coordinate of each tile map is , the two-axis ground resolutions and , and the pixel resolution , where represents the tile map serial number, the lower right coordinate is: , .

[0027] In step S3 of this embodiment, when splicing the found tile map with its neighboring tile maps to obtain the reference map, the found tile map has a similar ground resolution size, but its pixel resolution is unified to , while the pixel resolution of the actual infrared image is generally , which results in a small overlap rate between the tile map and the infrared image, making it difficult to achieve successful matching. In response to this situation, in step S3 of this embodiment, the tile map is expanded by splicing the tile maps in its neighborhood. The number of spliced maps can be specified as needed. For example, as an optional implementation, in this embodiment, the tile maps in the neighborhood are spliced to finally obtain sized reference map (with accurate geographic coordinates) for subsequent image matching. As Figure 3 shows an example of the reference map and the infrared image pair after geographic matching and tile splicing, where (a) and (b) are the first group of reference map and infrared image pair, and (a) is the reference map and (b) is the actual image; (c) and (d) are the second group of reference map and infrared image pair, and (c) is the reference map and (d) is the actual image.

[0028] As Figure 1 shown, in step S4, the actual image and the reference map are matched using a hierarchical matching method, which includes: S4.1, Coarse matching: The actual image and the reference map are respectively downsampled, and then a preset matching algorithm is used to match the downsampled reference map and the infrared map to obtain a set of feature point pairs , where is the k-th feature point of the reference map, is the k-th feature point of the actual image, is the total number of feature point pairs; S4.2, Feature point clustering: The feature points of the actual image and the reference map are respectively clustered to obtain the cluster center coordinates; S4.3, Region of interest cropping: The cluster center coordinates are respectively restored to the original ground resolution image coordinates, and a neighborhood of a specified size is cropped according to the original ground resolution image coordinates as the region of interest to obtain the region of interest matching image; S4.4, Fine matching: The region of interest matching images of the actual image and the reference map are matched using a preset matching algorithm to obtain a set of feature point pairs , where is the k-th feature point of the region of interest matching image of the reference map, is the k-th feature point of the region of interest matching image of the actual image, is the total number of feature point pairs.

[0029] It should be noted that the preset matching algorithm can adopt various known matching algorithms according to needs. In this embodiment, cross-modal remote sensing image matching is involved. The method based on deep learning is limited by the dataset and difficult to solve cross-modal remote sensing image matching. The traditional mathematics-based matching method is difficult to match successfully due to the too large image resolution of the reference image and the infrared map. Therefore, the MSFT algorithm (see: Maoqing Hu, Bin Sun, Xudong Kang, Shutao Li, Multiscale structural feature transform for multi-modal image matching, Information Fusion, Vol 95, 2023, pp 341-354.) is adopted as the preset matching algorithm.

[0030] When the captured image and the reference image are respectively downsampled in step S4.1 of this embodiment, it includes performing times downsampling on the reference image, and performing downsampling on the captured image according to the following formula: , In the above formula, and are the downsampling ratios in the x and y axis directions of the captured image, and are the ground resolutions of the two axes of the reference image, and are the ground resolutions of the two axes of the captured image, is the downsampling multiple performed on the reference image. The downsampling multiple performed on the reference image can be taken as needed. For example, as an optional implementation manner, the downsampling multiple performed on the reference image in this embodiment is taken as 4, then there is: , Then use the MSFT algorithm to match the downsampled reference image and the infrared map to obtain the set of feature point pairs.

[0031] The function expression for clustering the feature points of the captured image and the reference image respectively to obtain the clustering center coordinates in step S4.2 of this embodiment is: , , In the above formula, is the clustering center coordinate of the captured image, and are the x coordinate and y coordinate of the i-th feature point of the captured image respectively, is the clustering center coordinate of the reference image, and are the x - coordinate and y - coordinate of the i - th feature point of the reference image respectively, is the total number of feature point pairs.

[0032] In step S4.3 of this embodiment, the function expressions for restoring the clustering center coordinates to the original ground - resolution image coordinates respectively are as follows: , , In the above formula, is the clustering center coordinate of the real - shot image after being restored to the original ground - resolution image coordinates, and are the down - sampling magnification factors in the x - and y - axis directions of the real - shot image, is the clustering center coordinate of the reference image after being restored to the original ground - resolution image coordinates, is the down - sampling multiple executed on the reference image, is the clustering center coordinate of the real - shot image, is the clustering center coordinate of the reference image. In this embodiment, based on the clustering center coordinates obtained from the above formula, the corresponding neighborhoods are respectively cropped to obtain the regions of interest, and the cropping size is , and these are the matching image pairs that have the same ground targets and are easy to match.

[0033] In step S4.4 of this embodiment, during the fine - matching, the matching images of the regions of interest of the real - shot image and the reference image are continued to be matched using the MSFT algorithm to obtain the set of feature point pairs , where is the k - th feature point of the matching image of the region of interest of the reference image, is the k - th feature point of the matching image of the region of interest of the real - shot image, is the total number of feature point pairs, Figure 4 shows the comparison diagrams of the rough - matching, fine - matching and the final result, where (a) is the result of the rough - matching, (b) is the result of the fine - matching, and (c) is the final result.

[0034] Then, in step S5, according to the accurate geographical coordinates of the reference image and its corresponding feature points in the real - shot image, the geographical coordinates of the feature points with position errors in the real - shot image are corrected to obtain higher - precision geographical coordinates; in step S6, the pose of the UAV is solved according to the corrected geographical coordinates of the feature points in the real - shot image. In this embodiment, specifically, the PnP pose - solving method is used to solve the high - precision pose of the UAV. The PnP pose - solving method is a well - known pose - solving method, including according to the accurate geographical coordinates of the feature points and the corresponding pixel coordinates First, use P3P to solve the initial pose, and then use Bundle Adjustment (BA) to iteratively optimize the pose to obtain the final accurate pose. The principle of P3P to solve the initial pose is as follows: given three pairs of matching pairs of known geographic coordinates and pixel coordinates, that is, the pixel coordinates and geographic coordinates of the infrared image feature points obtained in step five. Then, based on similar triangles, a set of quadratic equations is constructed jointly, and then the Wu elimination method is used to solve it to obtain the 3D coordinates of the three pairs of feature points in the camera coordinate system. Finally, based on the 3D-3D point pair (camera coordinate system-geographic coordinate system), the pose of the camera in the geographic coordinate system is estimated. For specific calculation methods, please refer to the literature (Gao Xiang, Zhang Tao, "Fourteen Lectures on Visual SLAM", Electronic Industry Press, 2017.). The principle of bundle adjustment (BA) includes: based on the initial pose provided by the PnP algorithm, construct an optimization function with reprojection error as the objective function and drone pose as the optimization term, and then use Gauss-Newton iteration and other methods to iteratively optimize the drone pose to obtain the final accurate pose. For specific calculation methods, please refer to the literature: (Gao Xiang, Zhang Tao, "Fourteen Lectures on Visual SLAM", Electronic Industry Press, 2017.).

[0035] In summary, the large-scene UAV positioning method of multi-source cross-modal image hierarchical matching in this embodiment includes geographic matching based on the UAV aerial infrared image and the satellite optical map to obtain a reference map with precise geographic coordinates near the UAV; by performing hierarchical image matching on the infrared image and the reference map to obtain the pixel coordinate correspondence between the two, wherein the hierarchical image matching algorithm mainly includes the steps of coarse matching, clustering, cropping, and fine matching; finally, based on the pixel coordinate correspondence between the infrared image and the reference map, the geographic coordinates of the infrared image are corrected, and finally the accurate positioning of the UAV is achieved through posture solution. The present invention can achieve accurate positioning of the UAV in a large scene through visual assistance when the global positioning system of the UAV is interfered with and can only provide a low-precision position.

[0036] In addition, this embodiment also provides a large-scene UAV positioning system for multi-source cross-modal image hierarchical matching, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the large-scene UAV positioning method for multi-source cross-modal image hierarchical matching. This embodiment also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is used to be programmed or configured by a microprocessor to execute the large-scene UAV positioning method for multi-source cross-modal image hierarchical matching. This embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the large-scene UAV positioning method for multi-source cross-modal image hierarchical matching through a processor.

[0037] Those skilled in the art should understand that the technical solutions provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0038] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A large-scale UAV positioning method for multi-source cross-modal image hierarchical matching, characterized in that Including: S1. Generate an image pyramid presented in a tile map with different ground resolutions based on the satellite optical map of the UAV airport area, obtain the real-shot image of the UAV airport area and correct it to the same orientation as the satellite optical map. S2. Determine the corresponding tile map level according to the ground resolution of the real-shot image, and retrieve in the tile map corresponding to the tile map level according to the geographical coordinates of the central pixel point of the real-shot image to find the tile map with the same ground target as the real-shot image. S3. Stitch the found tile map with its neighboring tile maps to obtain a reference map. S4. Perform image matching on the real-shot image and the reference map using a hierarchical matching method to obtain the corresponding feature points and pixel transformation matrix between the real-shot image and the reference map. S5. According to the accurate geographical coordinates of the reference map and its corresponding feature points in the real-shot image, correct the geographical coordinates of the feature points with position errors in the real-shot image to obtain geographical coordinates with higher accuracy. S6. Solve the UAV pose according to the geographical coordinates of the feature points after correction in the real-shot image.

2. The large-scene UAV positioning method for multi-source cross-modal image hierarchical matching according to claim 1, wherein In step S2, determining the corresponding tile map level according to the ground resolution of the real-shot image includes: S2.1, calculate the distance between the real-shot image and each tile map according to the following formula d ; , In the above formula, and are respectively the two-axis ground resolutions of the actual captured image, and are respectively the two-axis ground resolutions of the i-th tile map; S2.2, find the d tile map level corresponding to the tile map with the smallest distance as the tile map level corresponding to the ground resolution of the real-shot image.

3. The large-scale UAV positioning method for multi-source cross-modal image hierarchical matching according to claim 1, wherein When retrieving in the tile map corresponding to the tile map level according to the geographical coordinates of the central pixel point of the real-shot image in step S2 to find the tile map with the same ground target as the real-shot image, the function expression of the constraint condition satisfied by the tile map with the same ground target as the real-shot image is: , In the above formula, is the geographical coordinate of the central pixel point of the real-shot image, is the upper-left coordinate of the i-th tile map, is the lower-right coordinate of the i-th tile map.

4. The large-scene UAV positioning method for multi-source cross-modal image hierarchical matching according to claim 1, wherein In step S4, performing image matching on the real-shot image and the reference map using a hierarchical matching method includes: S4.1, downsample the real-shot image and the reference image respectively, and then use a preset matching algorithm to match the downsampled reference image and the infrared map to obtain a set of feature point pairs , where is the k-th feature point of the reference image, is the k-th feature point of the real-shot image, is the total number of feature point pairs; S4.

2. Cluster the feature points of the real-shot image and the reference map respectively to obtain the cluster center coordinates. S4.

3. Restore the cluster center coordinates to the original ground resolution image coordinates respectively, and crop out a specified-sized neighborhood as the region of interest according to the original ground resolution image coordinates to obtain the region of interest matching image. S4.4, match the region of interest matching images of the real-shot image and the reference image using a preset matching algorithm to obtain a set of feature point pairs , where is the k-th feature point of the region of interest matching image of the reference image, is the k-th feature point of the region of interest matching image of the real-shot image, is the total number of feature point pairs.

5. The large-scale UAV positioning method for multi-source cross-modal image hierarchical matching according to claim 4, wherein When performing downsampling on the actual captured image and the reference image respectively in step S4.1, it includes performing multiple downsampling on the reference image, and performing downsampling on the actual captured image according to the following formula: , In the above formula, and are the downsampling ratios in the x and y axis directions of the actual captured image, and are the ground resolutions of the two axes of the reference image, and are the ground resolutions of the two axes of the actual captured image, is the downsampling multiple performed on the reference image.

6. The large-scale UAV positioning method for multi-source cross-modal image hierarchical matching according to claim 4, wherein The function expression for clustering the feature points of the real-shot image and the reference map respectively to obtain the cluster center coordinates in step S4.2 is: , , In the above formula, is the coordinate of the clustering center of the real-shot image, and are the x coordinate and y coordinate of the i-th feature point of the real-shot image respectively, is the coordinate of the clustering center of the reference image, and are the x coordinate and y coordinate of the i-th feature point of the reference image respectively, is the total number of feature point pairs.

7. The large-scale UAV positioning method for multi-source cross-modal image hierarchical matching according to claim 4, wherein The function expression for restoring the cluster center coordinates to the original ground resolution image coordinates respectively in step S4.3 is: , , In the above formula, is the clustering center coordinate after the real-shot image is restored to the image coordinates of the original ground resolution, and are the downsampling ratios in the x and y axis directions of the real-shot image, is the clustering center coordinate after the reference image is restored to the image coordinates of the original ground resolution, is the downsampling multiple performed on the reference image, is the clustering center coordinate of the real-shot image, is the clustering center coordinate of the reference image.

8. A large-scale UAV positioning system for multi-source cross-modal image hierarchical matching, comprising a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the large-scale UAV positioning method of multi-source cross-modal image hierarchical matching described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program therein, characterized in that, The computer program is used to be programmed or configured by the microprocessor to execute the large-scale UAV positioning method of multi-source cross-modal image hierarchical matching described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the large-scale UAV positioning method of multi-source cross-modal image hierarchical matching described in any one of claims 1 to 7 through the processor.

Citation Information

Patent Citations

  • Feature extraction method based on color feature

    CN106650755A

  • Self-adaptive satellite image generation method for visual positioning of unmanned aerial vehicle

    CN114201633A

  • Unmanned aerial vehicle visual positioning method based on multi-source image matching

    CN114842220A

  • Unmanned aerial vehicle multi-mode visual positioning method and system in GNSS denial environment

    CN119511332A

Cited By

  • Unmanned aerial vehicle image and multi-modal map area cutting method and system based on attitude information

    CN121527101A