A fast scene matching positioning method for UAVs in large-scale scenes
By generating a feature descriptor database offline and introducing a secondary scene adaptation area, combined with semantic segmentation and visual place recognition algorithms, the problems of slow computing speed and low robustness of drones in large-scale scenes are solved, and fast and accurate drone positioning is achieved.
Patent Information
- Application Number
- CN202411047988.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-07-31
AI Technical Summary
In large-scale scenarios, the drone scene matching method has slow calculation speed and low robustness, making it difficult to achieve real-time and high-precision positioning and navigation.
By generating a feature descriptor database offline and introducing a secondary scene adaptation area, combined with semantic segmentation and visual place recognition algorithms, the computational complexity is reduced and the system robustness is improved.
It greatly reduces the amount of real-time calculations, lowers the probability of mismatching, improves the robustness and positioning accuracy of the system, and achieves fast and accurate drone positioning.
Smart Images

Figure CN119251455B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to scene matching technology and unmanned aerial vehicle navigation technology, and in particular to a scene matching navigation method based on fast feature matching. Background Art
[0002] With the rapid development of drone technology and chips, the use of drones equipped with embedded computing platforms in military, civilian, and scientific research fields is increasing. Rapid drone navigation is a fundamental prerequisite for achieving other tasks. However, as applications become increasingly complex, they often need to operate in environments where GNSS signals are unavailable or unreliable, requiring the use of other sensors for drone positioning and navigation. Image matching technology is a viable solution for achieving positioning and navigation, but large-scale image data navigation tasks consume a significant amount of time in data preprocessing and repeated feature extraction. Furthermore, traditional feature extraction algorithms are sensitive to image variations and suffer from low robustness, limiting their real-time performance and accuracy. Therefore, improving the computational speed, accuracy, and robustness of scene matching methods is a key area of research.
[0003] Aiming at the problems of low calculation speed and low robustness of scene matching methods in large-scale scenes, the present invention proposes a fast scene matching positioning method for unmanned aerial vehicles in large-scale scenes. Summary of the Invention
[0004] In order to address the deficiencies of the above-mentioned prior art, the present invention aims to provide a method for rapid scene matching and positioning of UAVs in large-scale scenarios, which reduces the amount of computation by generating a feature descriptor database offline and introducing a secondary scene adaptation area; in the event of a matching failure, a visual place recognition algorithm is used to quickly retrieve the correct secondary scene adaptation area to improve the robustness of the system.
[0005] This application achieves the above effects through the following technical solutions: a method for rapid scene matching and positioning of drones in a large-scale scene, the method comprising the following steps:
[0006] S1 uses drones to obtain real aerial images of the target area, projects and corrects them, synthesizes them into a large scene range base map, and constructs a dataset;
[0007] S2 builds a semantic segmentation model and fine-tunes the pre-trained semantic segmentation model according to the data set; uses the visual place recognition (VPR) algorithm to process the large scene base map in blocks to generate a block recognition vector database; uses the feature extraction algorithm to process the large scene base map in blocks to generate a block matching feature database;
[0008] The S3 airborne computing platform obtains the real-time image from the airborne camera and uses the secondary scene adaptation area generation algorithm to generate the secondary scene adaptation area based on the real-time image.
[0009] S4 uses a precise matching algorithm to perform pixel-level precise matching on the real-time image of the onboard camera and outputs the precise matching pixel coordinates;
[0010] S5 uses a pixel coordinate to geographic coordinate mapping algorithm to calculate the geographic coordinates of the center pixel coordinates of the real-time image of the drone's onboard camera, and the air pressure information to generate NMEA format GNSS information.
[0011] Furthermore, the S1 is specifically:
[0012] S11 defines the flight area and uses commercial drones to perform overlapping traversal flights over the target area;
[0013] S12 inputs the obtained overlapping UAV images into ArcGIS software for projection correction, generates a large scene range base map in TIFF format with geographic coordinate information, and generates a TFW file based on the geographic information of the image;
[0014] S13 crops and labels the obtained base map to construct a segmentation dataset.
[0015] Furthermore, the S2 is specifically:
[0016] S21 fine-tunes the pre-trained semantic segmentation network according to the segmentation dataset to obtain fine-tuned training weights;
[0017] S22 uses the fine-tuned semantic segmentation network to infer the cropped base image to obtain a semantic segmentation mask;
[0018] S23 inputs the obtained segmentation mask into the VPR algorithm to generate a block retrieval vector library;
[0019] S24 feeds the cropped large scene range base map into the feature extraction algorithm to obtain the matching features of each block and generate a matching feature library.
[0020] Furthermore, the S3 is specifically:
[0021] S31 determines whether the matching result of the previous frame is reliable;
[0022] If the matching result of the previous frame is reliable in S32, jump to S33; if the matching result of the previous frame is unreliable, jump to S34;
[0023] S33 takes the block area where the coordinates of the previous frame matching result are located as the central block of the secondary scene adaptation area, and expands it into an N×N block as the secondary scene adaptation area with this block as the center, and directly enters S4;
[0024] S34 sends the current frame image to the semantic segmentation model to obtain a segmentation mask;
[0025] S35 inputs the segmentation mask into the VPR algorithm to obtain the searched vector;
[0026] S36 calculates the cosine similarity between the searched vector and each vector in the search vector library. The cosine similarity formula is as follows:
[0027]
[0028] Where, vector A is the searched vector, vector B is the vector in the search vector library, A·B is the dot product of vector A and vector B, ‖A‖ and ‖B‖ are the Euclidean norms of vector A and vector B respectively;
[0029] S37: The block with the highest similarity is used as the central block of the secondary scene adaptation area, and the block is expanded into N×N blocks with the central block of the secondary scene adaptation area as the center, as the secondary scene adaptation area, and then the process goes to S4.
[0030] Furthermore, the S4 is specifically:
[0031] S41 uses a feature extraction algorithm to extract features from the current frame image;
[0032] S42 obtains the features of each block from the matching feature database according to the range of the secondary scene adaptation area, and merges the feature vector queues of each block into a complete feature vector queue;
[0033] S43 uses a precise matching algorithm to match the feature vector queue of the current frame image with the feature vector queue of the secondary scene adaptation area;
[0034] S44 uses the RANSAC algorithm to remove mismatched points after matching;
[0035] S45 calculates the homography transformation matrix H to obtain the transformation matrix between the current frame image and the base map, and obtains the position of the center coordinates of the current frame image in the base map. The relationship between the homography transformation matrix and the mapping coordinate point pair is shown in the following formula:
[0036]
[0037] Where (x′, y′) is the mapping point coordinate of the feature point (x, y), and H is the homography matrix;
[0038] S46 calculates the pixel coordinates of the center coordinates of the current frame in the base map using the homography matrix.
[0039] Furthermore, the S5 is specifically as follows:
[0040] S51 reads the mapping relationship between the image space and the geographic space in the TFW file generated in S12;
[0041] S52 calculates the geographic coordinates of the center pixel coordinates according to the mapping relationship. The calculation formula is as follows:
[0042] x′=Ax+By+C
[0043] y′=Dx+Ey+F
[0044] Where x' is the geographic X coordinate corresponding to the pixel, y' is the geographic Y coordinate corresponding to the pixel, x is the pixel coordinate, y is the pixel coordinate, A is the pixel resolution in the X direction, D and B are the translation and rotation coefficients, E is the pixel resolution in the Y direction, C is the X coordinate of the pixel center in the upper left corner of the grid map, and F is the Y coordinate of the pixel center in the upper left corner of the grid map.
[0045] The S53 calculates the barometric pressure and generates GNSS sentence information in NMEA format. The barometric pressure uses the hypsometric formula, which is as follows:
[0046]
[0047] Where P0 is the standard atmospheric pressure, which is 101.325 kPa; P is the actual measured atmospheric pressure, in kPa; and T is the actual measured temperature, in °C. This formula considers both temperature and pressure to calculate altitude.
[0048] Compared with the prior art, the present invention has the following advantages:
[0049] (1) The present invention uses an offline method to extract matching features of a large base map, which greatly reduces the amount of real-time calculations. At the same time, there is no need to store the original base map on the computing platform, saving storage space.
[0050] (2) The present invention uses the secondary scene adaptation area as a buffer zone to match the real-time image, which reduces the amount of calculation and the probability of mismatching.
[0051] (3) The present invention uses the VPR algorithm to correct the matching algorithm when a mismatch occurs, thereby improving the robustness of the system.
[0052] (4) The feature extraction algorithm of the present invention can also be replaced, and deep learning methods can also be used for feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Flowchart of the method for rapid scene matching and positioning of UAVs in large-scale scenes;
[0054] Figure 2 Schematic diagram of SIFT feature storage data structure;
[0055] Figure 3 This is a schematic diagram of the secondary scene adaptation area;
[0056] Figure 4 Flowchart for generating secondary scene adaptation area for applying VPR algorithm;
[0057] Figure 5 This is a schematic diagram of the client and server structures in the precise matching algorithm;
[0058] Figure 6 Schematic diagram of image matching. DETAILED DESCRIPTION
[0059] In order to better illustrate the technical details of the present invention, the present invention is further described below with reference to the accompanying drawings.
[0060] The present invention proposes a method for rapid scene matching and positioning of UAVs in large-scale scenarios. The method reduces the computational complexity by generating a feature descriptor database offline and introducing a secondary scene adaptation region. In the event of matching failure, the semantic segmentation and VPR algorithm are used to quickly retrieve the correct secondary scene adaptation region to improve the system robustness.
[0061] Specifically, the present invention provides a method for fast scene matching and positioning of UAVs in a large-scale scene, such as Figure 1 As shown, the following steps are included:
[0062] S1 first uses a drone to obtain real aerial images of the target area, corrects and synthesizes them into a large scene range base map, and constructs a dataset;
[0063] S11, demarcate the flight area and use commercial drones to perform traversal flights over the target area with overlapping areas;
[0064] S12, inputting the obtained overlapping UAV images into ArcGIS software for projection correction, synthesizing and generating a base map with geographic coordinate information;
[0065] S13: Crop the obtained base image without overlapping to a size of 1024×1024; randomly select 100 cropped images (which can be adjusted according to the difficulty of the task) for annotation, and divide them into training set and test set according to the ratio of 0.85:0.15;
[0066] S2 builds a semantic segmentation model, fine-tunes the pre-trained semantic segmentation model according to the dataset, and uses the Visual Place Recognition (VPR) algorithm and feature extraction algorithm to process the base map to generate a base map database;
[0067] S21: Build the ISDNet semantic segmentation network and select pre-trained weights based on the scene content. Fine-tune the semantic segmentation network using the dataset created in the previous step and obtain the fine-tuned training weights.
[0068] S22, use the semantic segmentation network with fine-tuned parameters to infer all the cropped images to obtain the semantic segmentation mask;
[0069] S23, build the VPR algorithm: AnyLoc, select the corresponding pre-trained weights according to the scene content, input the segmentation mask obtained in the previous step into AnyLoc, infer the vector and save it according to the image number, and generate a block retrieval vector library. The input image here can be larger than 1024×1024, but if the image is too large, it is easy to have insufficient computing resources. When the input image is larger than 1024×1024, the model will automatically maintain the proportion and downsample the input image to a maximum side size of 1024. Therefore, when subsequently calculating and generating real-time image retrieval vectors in real time, the spatial resolution of the input image should be as close as possible to the spatial resolution when constructing the block retrieval vector library, or when constructing the block retrieval vector library, the spatial resolution should be close to the spatial resolution of the task;
[0070] In step S24, the cropped image is fed into a feature extraction algorithm to obtain matching features for each block and generate a matching feature library. Since the subsequent generation of the secondary scene adaptation area requires a balance between reliability and computational speed, the cropped image size must satisfy the following equation to ensure that the secondary scene adaptation area completely encompasses the camera's framing range.
[0071]
[0072] Where s represents the pixel size of the cropped image, GSD S It represents the ground truth distance corresponding to each pixel of the base map when extracting features from the base map. a represents the ground truth distance corresponding to one pixel of the camera image when extracting features from the camera image. The units are both cm.
[0073] This step uses the GPU version of SIFT as an example. This example uses PopSIFT as the feature extraction and matching algorithm library. The library's API is used to process images. Because feature descriptors need to be saved intermediately, code for saving them from the GPU is added.
[0074] In this step, the feature descriptor needs to be stored in a specific format. In this example, SIFT features are used as an example. Figure 2 As shown, each block feature is saved in txt text format, and its name should be the pixel coordinates of the block (0,0) coordinates in the map. The first line of text contains the total number of feature descriptors of the current block, and each subsequent line represents a feature descriptor. As shown in the figure, the first line is 916, which means there are 916 lines of feature descriptors. The first two floating-point numbers in each line of feature descriptors are the coordinates of the feature point in the base map, the third floating-point number is the sigma scale information, and the following 128 floating-point numbers are the 128-dimensional feature vectors of the SIFT features. In this example, a space is used as the separator for each number;
[0075] S3 uses the secondary scene adaptation area generation algorithm to generate the secondary scene adaptation area based on the previous result and the current frame content for better matching. It is worth noting that the previous part is all offline database preparation, and the subsequent parts are all real-time calculation parts.
[0076] S31 determines whether the matching result of the previous frame is reliable based on whether the number of matching pairs in the previous frame exceeds the set threshold. This threshold is an empirical value. After a lot of experiments, this example sets it to 20 to achieve a balance between accuracy and speed. Note that the number of matching pairs should be the number after eliminating the wrong matching points;
[0077] If the matching result of the previous frame is reliable in S32, since the UAV will not perform large-scale instantaneous maneuvers, it can be considered that the last matching coordinates can be directly used as the reference point of the secondary scene adaptation area without executing the VPR algorithm to find a new reference point, so jump to S33. If the matching result of the previous frame is unreliable, the matching coordinates may have drifted over a large range and need to find a new reference point, so jump to S34.
[0078] S33 compares the previous frame matching result with the previous frame's secondary scene adaptation area update threshold, and decides whether to update the secondary scene adaptation area based on whether the threshold is exceeded. Figure 3 Where Δl is the updating threshold of the secondary scene adaptation area, which should satisfy the following formula.
[0079] △l×GSD s <V max ×T max
[0080] In the formula, GSD S Indicates the ground truth distance corresponding to each pixel of the base map when extracting features from the base map, in cm, V max T is the maximum flight speed of the UAV, in cm / s. max The time period for the system to process one frame, in seconds.
[0081] When Δy is less than Δl, the update mechanism is triggered, and the block area where the coordinates of the previous frame matching result are located is used as the central block of the secondary scene adaptation area, and the block is expanded as the center to the following: Figure 3 The 3×3 block shown is used as a secondary scene adaptation area and enters S4;
[0082] S34 downsamples or upsamples the current frame image so that the current frame GSD is close to the base image GSD, and sends it to the semantic segmentation model for inference to generate a segmentation mask. The specific process is as follows: Figure 4 As shown, the upper part is the offline calculation part, which is used to form the retrieval vector library, and the lower part is the real-time calculation part, that is, the process from S34 to S37;
[0083] S35 inputs the segmentation mask into the AnyLoc algorithm to generate a search vector;
[0084] S36 calculates the cosine similarity between the searched vector and each vector in the search vector library. The cosine similarity formula is as follows:
[0085]
[0086] Where A·B is the dot product of vector A and vector B, ‖A‖ and ‖B‖ are the Euclidean norms (i.e., the lengths of the vectors) of vector A and vector B, respectively.
[0087] S37 is similar to S33. The center coordinates of the block with the highest similarity are used as reference coordinates to determine whether it exceeds the update threshold. If it exceeds the update threshold, the block is used as the center block of the secondary scene adaptation area, and the block is expanded as the center to form the following image: Figure 3 The 3×3 block shown is used as a secondary scene adaptation area and enters S4;
[0088] S4 uses a precise matching algorithm to perform pixel-level precise matching on optical images and output precise matching pixel coordinates. This example implements the algorithm in two processes: client and server. Each process uses multi-threading to implement tasks. The client is mainly responsible for feature extraction and feature matching, while the server is mainly responsible for reading and combining features in the secondary adaptation area, coordinate system conversion, and output of geographic coordinates. The specific process structure is as follows: Figure 5 As shown;
[0089] In the S41 example, the client obtains the current frame image from the camera stream, and then uses the Popsift algorithm to call CUDA to extract features from the current image output to obtain a feature descriptor. To balance speed and accuracy, this example sets the upper limit of feature point extraction to 2048.
[0090] The S42 server uses a thread pool to read the features of each block from the matching feature database based on the range of the secondary scene adaptation area, and merges the feature vector queues of each block into a complete feature vector queue, and transmits it to the client via TCP / IP. In this example, you can choose to deploy the client and server on the same embedded development board or on different development boards. The client program must run on a development board that supports CUDA, but there is no requirement for the server. It should be noted that if different development boards are used, the two development boards need to be connected via a network cable and a gateway needs to be set up. At this time, the network transmission speed may become a bottleneck. In addition, in this example, the client uses the SIFT algorithm to extract features. It is also possible to use a deep learning-based method to extract and match features. You only need to replace the client program and modify the server code slightly. This example mainly uses SIFT features as the main illustrative example;
[0091] S43 The client receives the feature vector queue generated by the server and calls the GPU-accelerated BF matching algorithm in Popsift to match the current frame features with the features of the secondary scene adaptation area to obtain a set of matching point pairs without eliminating false matching points;
[0092] The S44 client uses the RANSAC algorithm to remove mismatched points from the matched point pairs;
[0093] The S45 client calculates the homography transformation matrix H by eliminating the mismatched points and obtains the transformation matrix between the current frame image and the base image. The relationship between the homography transformation matrix and the mapping coordinate point pairs is shown in the following formula:
[0094]
[0095] Where (x′, y′) is the coordinate of the mapping point of the feature point (x, y), and H is the homography matrix, which can be calculated using 4 or more sets of feature matching point pairs. This example uses the findHomography() function in the OpenCV library to calculate the matrix. The more point pairs are input, the higher the accuracy. The matching diagram is shown below. Figure 6 As shown, the green lines are matching point pairs, and the blue boxes are the specific positions of the left image in the right image;
[0096] The S46 client uses the pixel coordinates of the current frame center point in the current frame image to multiply the homography matrix to obtain the pixel coordinates of the current frame image center coordinates in the base map. The client packages the coordinates and sends them to the server to process the next frame image.
[0097] S5 uses a pixel coordinate to geographic coordinate mapping algorithm to calculate the geographic coordinates of the center pixel coordinates of the drone image and the air pressure information to generate NMEA format GNSS information.
[0098] S51 reads the mapping relationship between image space and geographic space in the TFW file and obtains the geographic information of the base map, which includes the following six data:
[0099] 1 The resolution scale of one pixel in the map unit in the X direction.
[0100] 2 translation amount.
[0101] 3. Rotation amount. (angle)
[0102] The negative value of the Y resolution scale in the Y direction is one pixel in 4 map units.
[0103] 5 The X coordinate of pixel 0,0 (upper left).
[0104] 6 The Y coordinate of pixel 0,0 (upper left).
[0105] S52 calculates the geographic coordinates of the center pixel coordinates according to the mapping relationship. The calculation formula is as follows:
[0106] x′=Ax+By+C
[0107] y′=Dx+Ey+F
[0108] Where x' is the geographic X coordinate corresponding to the pixel, y' is the geographic Y coordinate corresponding to the pixel, x is the pixel coordinate, y is the pixel coordinate, A is the pixel resolution in the X direction, D and B are the translation and rotation coefficients, E is the pixel resolution in the Y direction, C is the X coordinate of the pixel center in the upper left corner of the grid map, and F is the Y coordinate of the pixel center in the upper left corner of the grid map.
[0109] The S53 calculates the barometric pressure and generates GNSS sentence information in NMEA format. The barometric pressure uses the hypsometric formula, which is as follows:
[0110]
[0111] Where P0 is the standard atmospheric pressure, which is 101.325 kPa; P is the actual measured atmospheric pressure in kPa; and T is the actual measured temperature in °C. T+273.15 is the conversion from Celsius to Kelvin. This formula considers both temperature and pressure to calculate altitude.
[0112] Finally, the server transmits the generated GNSS sentence to the flight controller through the serial port. The specific sentence is as follows:
[0113] $GNRMC,095025.654,A,3200.15746,N,11836.81634,E,15.143,236.50,030624,,,A*7 A$GNGGA,095025.654,3200.1575,N,11836.8163,E,1,12,1.50,115.1,M,81.4,M,,*76
[0114] This statement is output through logging during the experiment and conforms to the NMEA format. The flight controller used in this example only requires RMC and GGA statements to recognize it, so this statement can be recognized and received by the flight controller normally.
[0115] The above provides a detailed introduction to the rapid scene matching positioning method for drones in large-scale scenarios provided by the present invention. The principles and implementation of this technology are illustrated through specific examples. These examples are intended only to facilitate understanding of the core concepts of this technology. It should be emphasized that professionals skilled in this field are free to make appropriate adjustments and optimizations to this technology without violating its fundamental principles. Such adjustments and optimizations fall within the scope of protection of the claims of this invention.
Claims
1. A method for rapid scene matching positioning of UAVs in a large-scale scene, characterized by: The method comprises the following steps: S1 uses drones to obtain real aerial images of the target area, projects and corrects them, synthesizes them into a large scene range base map, and constructs a dataset; S2 builds a semantic segmentation model and fine-tunes the pre-trained semantic segmentation model according to the data set; uses the visual place recognition (VPR) algorithm to process the large scene base map in blocks to generate a block recognition vector database; uses the feature extraction algorithm to process the large scene base map in blocks to generate a block matching feature database; The S3 airborne computing platform obtains the real-time image from the airborne camera and uses the secondary scene adaptation area generation algorithm to generate the secondary scene adaptation area based on the real-time image. The S3 is specifically as follows: S31 determines whether the matching result of the previous frame is reliable; If the matching result of the previous frame is reliable in S32, jump to S33; if the matching result of the previous frame is unreliable, jump to S34; S33 takes the block area where the coordinates of the previous frame matching result are located as the central block of the secondary scene adaptation area, and expands it into an N×N block as the secondary scene adaptation area with this block as the center, and directly enters S4; S34 sends the current frame image to the semantic segmentation model to obtain a segmentation mask; S35 inputs the segmentation mask into the VPR algorithm to obtain the searched vector; S36 calculates the cosine similarity between the searched vector and each vector in the search vector library. The cosine similarity formula is as follows: Where, vector A is the searched vector, vector B is the vector in the search vector library, A·B is the dot product of vector A and vector B, ‖A‖ and ‖B‖ are the Euclidean norms of vector A and vector B respectively; S37: The block with the highest similarity is used as the central block of the secondary scene adaptation area, and the central block of the secondary scene adaptation area is expanded into N×N blocks as the secondary scene adaptation area, and the process goes to S4; S4 uses a precise matching algorithm to perform pixel-level precise matching on the real-time image of the onboard camera and outputs the precise matching pixel coordinates; S5 uses a pixel coordinate to geographic coordinate mapping algorithm to calculate the geographic coordinates of the center pixel coordinates of the real-time image of the drone's onboard camera, and the air pressure information to generate NMEA format GNSS information.
2. The method for rapid scene matching positioning of UAVs in a large-scale scenario according to claim 1 is characterized in that: The S1 is specifically: S11 defines the flight area and uses commercial drones to perform overlapping traversal flights over the target area; S12 inputs the obtained overlapping UAV images into ArcGIS software for projection correction, generates a large scene range base map in TIFF format with geographic coordinate information, and generates a TFW file based on the geographic information of the image; S13 crops and labels the obtained base map to construct a segmentation dataset.
3. The method for rapid scene matching positioning of a UAV in a large-scale scene according to claim 2 is characterized in that: The S2 is specifically: S21 fine-tunes the pre-trained semantic segmentation network according to the segmentation dataset to obtain fine-tuned training weights; S22 uses the fine-tuned semantic segmentation network to infer the cropped base image to obtain a semantic segmentation mask; S23 inputs the obtained segmentation mask into the VPR algorithm to generate a block retrieval vector library; S24 feeds the cropped large scene range base map into the feature extraction algorithm to obtain the matching features of each block and generate a matching feature library.
4. The method for rapid scene matching positioning of unmanned aerial vehicles in a large-scale scene according to claim 1 is characterized in that: The S4 is specifically: S41 uses a feature extraction algorithm to extract features from the current frame image; S42 obtains the features of each block from the matching feature database according to the range of the secondary scene adaptation area, and merges the feature vector queues of each block into a complete feature vector queue; S43 uses a precise matching algorithm to match the feature vector queue of the current frame image with the feature vector queue of the secondary scene adaptation area; S44 uses the RANSAC algorithm to remove mismatched points from the matched point pairs; S45 calculates the homography transformation matrix H to obtain the transformation matrix between the current frame image and the base map, and obtains the position of the center coordinates of the current frame image in the base map. The relationship between the homography transformation matrix and the mapping coordinate point pair is shown in the following formula: Where (x′, y′) is the mapping point coordinate of the feature point (x, y), and H is the homography matrix; S46 calculates the pixel coordinates of the center coordinates of the current frame in the base map using the homography matrix.
5. The method for rapid scene matching positioning of unmanned aerial vehicles in a large-scale scene according to claim 1 is characterized in that: The S5 is specifically: S51 reads the mapping relationship between the image space and the geographic space in the TFW file generated in S12; S52 calculates the geographic coordinates of the center pixel coordinates according to the mapping relationship. The calculation formula is as follows: x′=Ax+By+C y′=Dx+Ey+F Where x' is the geographic X coordinate corresponding to the pixel, y' is the geographic Y coordinate corresponding to the pixel, x is the pixel coordinate, y is the pixel coordinate, A is the pixel resolution in the X direction, D and B are the translation and rotation coefficients, E is the pixel resolution in the Y direction, C is the X coordinate of the pixel center in the upper left corner of the grid map, and F is the Y coordinate of the pixel center in the upper left corner of the grid map. The S53 calculates the barometric pressure and generates GNSS sentence information in NMEA format. The barometric pressure uses the hypsometric formula, which is as follows: Where P0 is the standard atmospheric pressure, which is 101.325 kPa; P is the actual measured atmospheric pressure, in kPa; T is the actual measured temperature, in °C.
Citation Information
Patent Citations
Navigation and positioning method for unmanned aerial vehicle
CN114509070A
Scene matching method based on unmanned aerial vehicle image and satellite map in denial environment
CN118053010A