Feature extraction and matching method applied to navigation and positioning of unmanned aerial vehicle in underground space

By improving the SURF point feature extraction algorithm and combined visual inertial sensor, combined with IMU pre-integration to eliminate mismatch points, the drone navigation and positioning accuracy and real-time problems in underground space are solved, and more efficient feature point matching and navigation positioning are achieved.

CN120219489APending Publication Date: 2025-06-27SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272757.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In complex underground spaces, drone navigation positioning has factors such as poor lighting and weak texture that affect the navigation accuracy, and the mismatch problem in large maneuver states is serious, affecting the accuracy and real-timeness of navigation.

Method used

The improved SURF point feature extraction algorithm is adopted, combined with vision and inertial sensors, and false matching points are eliminated through IMU pre-integration, and feature extraction and matching algorithm are optimized, reducing the calculation amount and improving the accuracy and robustness of matching.

Benefits of technology

It improves the navigation and positioning accuracy and real-time performance of the drone in underground space, enhances the adaptability to complex and variable environments, and ensures the accuracy and reliability of feature point matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219489A_ABST
    Figure CN120219489A_ABST
Patent Text Reader

Abstract

A feature extraction and matching method applied to navigation and positioning of an unmanned aerial vehicle in an underground space comprises a data preprocessing step of processing underground space image data collected by a camera to generate an underground space image data set, a point feature extraction step of adopting an improved SURF technology to generate underground space point features, and a point feature extraction step of extracting the underground space point features. In the point feature matching step, a point feature matching method for eliminating mismatching based on IMU pre-integration is adopted to realize matching of underground space point features; a clear image data set with high contrast is obtained in the data preprocessing step, an improved extreme value detection algorithm is adopted in the point feature extraction step, direction distribution is carried out by combining the fast feature positioning characteristic of BOX filtering and fusing texture information of image gradient, and stable point features for rotation, noise, scale change and brightness change can be obtained. In the point feature matching step, the accuracy and robustness of a matching result can be realized by optimizing a data structure and utilizing an RANSAC (Random Sample Consensus) algorithm and IMU (Inertial Measurement Unit) pre-integration mismatching screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous navigation and positioning of unmanned aerial vehicles, and specifically relates to a method for feature extraction and matching applied to the navigation and positioning of unmanned aerial vehicles in underground spaces. Background Art

[0002] Facing the complexity and unknownness of the underground environment, unmanned aerial vehicles have the characteristics of small size, flexibility, wide activity space, and low cost. In the complex underground space, unmanned aerial vehicles do not rely on GNSS navigation information and can perform tasks such as rapid target search and rescue and unknown environment exploration in underground mines, tunnels, and urban underground areas. In underground spaces lacking GNSS signals, unmanned aerial vehicles need other auxiliary or alternative methods to achieve precise navigation and positioning.

[0003] The Simultaneous Localization and Mapping algorithm (SLAM) refers to a main body equipped with specific sensors that, in the absence of prior information about the environment, builds a model during movement while estimating its own movement. Feature extraction and matching is a key technology at the front end of the SLAM system. Unmanned aerial vehicles face problems such as insufficient single-body performance, limited environmental situation perception, and autonomous navigation ability in complex underground space environments. There are factors such as poor lighting and weak texture in underground spaces that directly affect the navigation and positioning accuracy of unmanned aerial vehicles. At the same time, due to large attitude and speed changes of unmanned aerial vehicles themselves, false matching problems will also occur.

[0004] The inertial navigation system can provide high-frequency measurements, and its measurements are not affected by external environments such as lighting, texture, and weather. However, its measurements have time-varying biases and require sensors that sense the external environment to provide effective constraints to reduce measurement errors. Cameras can sense rich texture information in the environment. By combining visual / inertial sensors, the environmental perception ability, autonomous navigation ability, and positioning accuracy of unmanned aerial vehicles can be improved to adapt to complex and changeable underground environments. At the same time, considering the limited computing power of unmanned aerial vehicles, the corresponding feature extraction and matching algorithms are optimized and improved to reduce the calculation time while ensuring accurate navigation and positioning, thereby improving the real-time performance of unmanned aerial vehicles. Summary of the Invention

[0005] In view of the requirements for accuracy and real-time performance of drones in the complex and changeable underground space operation environment, it is necessary to ensure that the extracted point features remain stable against rotation, noise, scale transformation, and brightness change, while improving the calculation efficiency. The present invention proposes a point feature extraction technology that improves SURF with accelerated robust characteristics. In view of the problem of false matching caused by the large maneuvering state of drones, the present invention proposes a visual SURF point feature matching method based on IMU pre-integration to eliminate false matching points, so as to improve the accuracy and robustness of matching. By reducing the calculation amount, optimizing the accelerated data structure, and using the RANSAC algorithm and IMU pre-integration for screening, the accuracy and robustness of matching are further improved, thus ensuring reliable feature point matching.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A feature extraction and matching method applied to the navigation and positioning of drones in underground spaces, comprising the following steps:

[0008] Step 1: Collect the original image I1(x) of the underground space, and perform image preprocessing on the obtained original image to obtain I2(x);

[0009] Step 2: Establish an integral image according to the preprocessed image I2(x), and obtain the sum I of the pixels of the input image above and to the left of the pixel at the position (x, y) Σ(x) , and adopt an improved SURF visual point feature extraction algorithm to calculate the set of pixel coordinates of the point features in the image

[0010] Step 3: Assign a feature main direction to the detected point features for feature description, construct a feature descriptor, perform dimensionality reduction processing on the descriptor by optimizing the data structure, and calculate the Euclidean distance between feature points for the reference image and the image to be matched to obtain a feature point matching feature set;

[0011] Step 4: Preliminary elimination of false matching points based on the RANSAC algorithm;

[0012] Step 5: Screen for false matching of point features based on IMU pre-integration.

[0013] As a further improvement of the present invention, in step 1, the image preprocessing of the collected underground space image data specifically includes:

[0014] Perform denoising and enhance the contrast.

[0015] As a further improvement of the present invention, the specific steps of step 2 are as follows:

[0016] Establish an integral image, I Σ(x)Represents the sum of the pixels of the input image above and to the left of the pixel at position (x, y), and its formula is:

[0017]

[0018] Then, replace the second-order Gaussian filter with a Box filter to start detecting and locating feature points, aiming to reduce the computational cost. The principle of the Box filter is to approximate the second-order Gaussian partial derivative using a template. Through the above two steps, the scale-space function is obtained. The core of SURF feature point location is the Hessian matrix, and the formula is as follows:

[0019]

[0020] Where L xx (X, σ) is the second-order Gaussian partial derivative Convolved with each point X in the image, L xy (X, σ) is Convolved with point X, L yy (X, σ) is Convolved with point X, and the Gaussian function formula therein:

[0021]

[0022] Where x and y represent the coordinates of point X;

[0023] Next, perform convolution with a multi-directional BOX filter to obtain D xx , D xy , D yy , and replace L xx , L xy , L yy in the Hessian matrix, so as to obtain the SURF fast Hessian matrix, and its formula is as follows:

[0024] det(H) = D xx D yy - ω(D xy ) 2

[0025] Then, apply non-maximum suppression in the 3×3×3 scale space. The extreme points among the 26 neighbors of the selected points in the scale space are regarded as SURF feature points, which are invariant to the scale of the source image, and the set of pixel coordinates of the point features is obtained

[0026] As a further improvement of the present invention, the specific steps of step 3 are as follows:

[0027] The direction of each SURF feature point is represented by the convolution of the Gaussian distribution-weighted Haar wavelet in the x and y directions, with the current feature point as the center and a sliding π / 3 window with a radius of 6σ, where σ is the size of the filter, to scan the wavelet response of the circular area here, and the sector with the largest wavelet response is determined as the main feature direction;

[0028] After determining the main direction of the current feature point, a square sampling area is constructed. Then, the sampling area is divided into 16 sub-areas, and each sub-area is further divided into 25 sample grids. Four descriptors in the positive and negative directions of the x-axis and y-axis of each sub-area are obtained from the Haar wavelet convolution. Subsequently, the 64-dimensional descriptor of the SURF feature point is obtained.

[0029] For the reference image and the image to be matched, the corresponding descriptors are used to calculate the Euclidean distance d between all SURF feature points in the two images. The minimum and second minimum values ​​of d are defined as the shortest distance d1 and the second shortest distance d2, respectively. If the ratio d1 / d2 is lower than the specified threshold, a pair of matching points can be obtained. All SURF feature points are obtained in this way, and all feature points form a matching feature set.

[0030] As a further improvement of the present invention, the specific steps of step 4 are as follows:

[0031] (a) Assume that the set of feature point matching pairs between the image to be tested and the template image is P, extract a certain number of feature point matching pairs n from it, and use n to initialize the transformation matrix M;

[0032] (b) If the error between the feature points corresponding to the remaining matching pairs and M is less than the agreed threshold, then they form a consistent set with the extracted n matching pairs, and the error E i The calculation formula is as follows:

[0033]

[0034] (c) Use M to traverse the feature points corresponding to the matching pairs in P, calculate the proportion S of consistent sets in P that satisfy M under the error threshold, repeat the above steps, and select the transformation matrix with the largest proportion, denoted as *M;

[0035] (d) Eliminate the elements of the set P whose errors exceed the threshold in the transformation matrix *M, and then use the remaining feature point matching pairs to obtain the transformation matrix, which is the final transformation matrix.

[0036] As a further improvement of the present invention, the specific steps of step 5 are as follows:

[0037] The pose relationship between two adjacent frames can be obtained by the IMU pre-integration method, so a certain landmark point P = [x, y, z, 1] TThe conversion relationships between the camera coordinate system and the world coordinate system for two frames can be respectively as follows;

[0038]

[0039] Where c1 and c2 are the camera coordinate systems of two adjacent frames, and b1 and b2 are the carrier coordinate systems of two adjacent frames. Since the carrier coordinate system is fixedly connected to the IMU, the installation method of the camera and the IMU determines the conversion matrix between the two And the conversion matrix between the w system and the b system is;

[0040]

[0041] Where the rotation and translation matrices can be obtained from the IMU pre-integration formula. Integrate the IMU measurements between the k-th frame and the (k + 1)-th frame, only considering the IMU biases ba, bg and the random noises na, ng, and the time interval Δt, and derive the translation vector from the carrier coordinate system to the world coordinate system The moving speed vector The rotation quaternion The formulas are as follows;

[0042]

[0043] The conversion between the camera coordinate system and the pixel coordinate system can be obtained by the intrinsic matrix K. The pixel coordinates of the landmark point P in two adjacent frames The conversion relationship with the camera coordinates is;

[0044]

[0045] Where the intrinsic matrix K can be obtained by off-line calibration of the camera;

[0046]

[0047] The matrix product is

[0048]

[0049] Ignoring the change in the depth of the feature point between two adjacent frames, it can be simplified to

[0050]

[0051] The above formula is the conversion relationship between the pixel coordinates of the landmark point P between two frames. Among them, the matrix E uses the IMU pre-integration. The pixel coordinates of the landmark point in the first frame are used to predict the pixel coordinates of the landmark point in the second frame according to the IMU pre-integration result between the two frames. The predicted value is compared with the pixel coordinates of the second frame obtained by initially removing the false matches. If the distance is greater than the threshold, the tracked point is considered a false match point to be removed, otherwise the tracking is successful, and finally the accurate matching of the point features is achieved.

[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0053] (1) The present invention uses vision and inertial sensors. The inertial navigation system can provide high-frequency measurements, and its measurements are not affected by external environments such as light, texture, and weather. The camera can perceive rich texture information in the environment. By combining multiple sensors, the environmental perception ability, autonomous navigation ability, and positioning accuracy of the drone are improved, thus adapting to the complex and changeable underground environment.

[0054] (2) The improved SURF point feature extraction algorithm proposed by the present invention. For the influence of light on the SURF algorithm, stable features can be extracted through the calculation of local gradient information and integral images, reducing the influence of light changes on feature extraction. For weak texture scenes, the SURF algorithm can capture feature information under weak textures by using Haar wavelet responses and integral image techniques, improving the stability and reliability of feature extraction. For structures with different depths and distances in the underground space, the scale differences of feature points are large. The SURF algorithm has scale invariance and can extract stable feature points at different scales in the underground space. The SURF algorithm has high computational efficiency by using integral images and fast feature extraction techniques.

[0055] (3) A visual SURF point feature matching method based on IMU pre-integration to eliminate false matches proposed by the present invention is used to eliminate false matching points, thereby improving the accuracy and robustness of matching. For the feature points extracted from the image by the SURF algorithm, the feature descriptors are dimensionally reduced to retain important information, reducing the computational amount of matching. To further improve the matching speed, an accelerated data structure is introduced to effectively reduce the comparison times between matching point pairs, thereby greatly improving the matching efficiency. Screening is carried out using the RANSAC algorithm and IMU pre-integration to further improve the accuracy and robustness of matching, thereby achieving reliable feature point matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the flowchart of the improved SURF point feature extraction algorithm in the present invention;

[0057] Figure 2 is the flowchart of the point feature matching based on IMU pre-integration elimination in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0059] A feature extraction and matching method for underground UAV navigation and positioning proposed by the present invention, the method comprises the following steps, and the flow chart of the improved SURF point feature extraction algorithm is as Figure 1 shown, and the flow chart of point feature matching based on IMU pre-integration rejection is as Figure 2 shown:

[0060] Step 1: Perform image preprocessing on the original underground space image I1(x). Denoise the image, and Gaussian filtering is intended to be used to reduce the noise in the image. Enhance the contrast of the image to make the scale space extreme points more obvious for detection, and the histogram equalization method is intended to be used to enhance the contrast of the image. The original image obtained is subjected to image preprocessing to obtain I2(x).

[0061] Step 2: Establish an integral image for the preprocessed image I2(x), where I ∑(x) represents the sum of the pixels of the input image above and to the left of the pixel at position (x, y), and its formula is:

[0062]

[0063] Then, replace the second-order Gaussian filter with a Box filter to start detecting and positioning feature points, aiming to reduce the computational cost. The principle of the Box filter is to approximate the second-order Gaussian partial derivative using a template. Through the above two steps, the scale space function can be obtained. Different from the Gaussian difference pyramid of SIFT, SURF does not need to perform Gaussian blur to establish the scale space. The core of SURF feature point positioning is the Hessian matrix, and the formula is as follows:

[0064]

[0065] where L xx (X, σ) is the convolution of the second-order Gaussian partial derivative with each point X in the image, L xy (X, σ) is the convolution with point X, and L yy (X, σ) is the convolution with point X. The Gaussian function formula among them is:

[0066]

[0067] where x and y represent the coordinates of point X.

[0068] Then, perform convolution with a multi-directional BOX filter to obtain D xx , D xy , D yy , and replace L in the Hessian matrix xx , L xy , Lyy , thus obtaining the SURF fast Hessian matrix, and its formula is as follows:

[0069] det(H) = D xx D yy -ω(D xy ) 2

[0070] Then, non-maximum suppression is applied in the 3×3×3 scale space. The extreme points among the 26 neighbors of the selected points in the scale space are regarded as SURF feature points, which are invariant to the scale of the source image, and a set of point feature pixel coordinates is obtained.

[0071] Step 3: Feature description is performed for the detected point features. The main direction of the feature is assigned, and a feature descriptor is constructed. By optimizing the data structure, dimensionality reduction processing is performed on the descriptor. For the reference image and the image to be matched, the Euclidean distance between feature points is calculated through the descriptor, and a feature point matching feature set is obtained.

[0072] The main direction of each SURF feature point is represented by the Haar wavelet convolution weighted by the Gaussian distribution in the x and y directions. With the current feature point as the center, a sliding π / 3 sector window with a radius of 6σ (σ is the size of the filter) is used to scan the wavelet response of the circular area here. The sector with the largest wavelet response is determined as the main direction of the feature. After determining the main direction of the current feature point, a square sampling area is constructed. Then, this sampling area is divided into 16 sub-areas, and each sub-area is further divided into 25 sample grids, and four descriptors in the positive and negative directions of the x-axis and y-axis of each sub-area are obtained from the Haar wavelet convolution. Subsequently, a 64-dimensional descriptor of the SURF feature point is obtained.

[0073] For the reference image and the image to be matched, the corresponding descriptors are used to calculate the Euclidean distance d between all SURF feature points in the two images. The minimum value and the second minimum value of d are defined as the shortest distance d1 and the second shortest distance d2 respectively. If the ratio d1 / d2 is lower than the specified threshold, a pair of matching points can be obtained. All SURF feature points are obtained in this way, and all feature points form a matching feature set.

[0074] Step 4: Preliminary elimination of mis-matched points based on the RANSAC algorithm. Assume that the set of feature point matching pairs between the image to be measured and the template image is P, and a certain number of feature point matching pairs n are extracted from it. The transformation matrix M is initialized using n.

[0075] If the error between the feature points corresponding to the remaining matching pairs and M is less than the agreed threshold, they form a consensus set with the n groups of matching pairs extracted. The error E i The calculation formula of is as follows:

[0076]

[0077] Use M to traverse the feature points corresponding to the matching pairs in P, calculate the proportion S of the consistent set in P that satisfies M under the error threshold, repeat the above steps, and select the transformation matrix with the largest proportion and denote it as *M;

[0078] Remove the elements in the set P whose errors exceed the threshold in the most transformation matrix *M, and then use the remaining feature point matching pairs to obtain the transformation matrix, which is the final transformation matrix.

[0079] Step 5: Screen out the incorrect point feature matches based on IMU pre-integration. To solve the problem of incorrect feature point matches based on SURF under large motion, introduce the IMU pre-integration results between frames as a constraint to remove the incorrect matching points. The pose relationship between two adjacent frames can be obtained by the IMU pre-integration method. Then a certain landmark point P = [x, y, z, 1] T The conversion relationships between the camera coordinate system and the world coordinate system of two frames can be respectively

[0080]

[0081] In the formula, c1 and c2 are the camera coordinate systems of two adjacent frames, b1 and b2 are the carrier coordinate systems of two adjacent frames. Since the carrier coordinate system is fixedly connected to the IMU, the installation method of the camera and the IMU determines the conversion matrix between the two And the conversion matrix between the w system and the b system is

[0082]

[0083] In the formula, the rotation and translation matrices can be obtained from the IMU pre-integration formula. Integrate the IMU measurements between the k-th frame and the (k + 1)-th frame, only considering the IMU biases ba, bg and the random noises na, ng, and the time interval Δt, and deduce the translation vector from the carrier coordinate system to the world coordinate system The moving speed vector The rotation quaternion The formula is as follows

[0084]

[0085] The conversion between the camera coordinate system and the pixel coordinate system can be obtained by the internal parameter matrix K. The pixel coordinates of the landmark point P in two adjacent frames The conversion relationship with the camera coordinates is

[0086]

[0087] In the formula, the internal parameter K can be obtained by offline calibration of the camera

[0088]

[0089] The matrix product is

[0090]

[0091] Ignoring the change in the depth of feature points between two adjacent frames, it can be simplified to

[0092]

[0093] The conversion relationship between the pixel coordinates of the landmark point P between two frames. The matrix E utilizes IMU pre-integration. The pixel coordinates of the landmark point in the first frame are used to predict the pixel coordinates of the landmark point in the second frame according to the IMU pre-integration result between two frames. The predicted value is compared with the pixel coordinates of the second frame obtained by initially eliminating false matches. If the distance is greater than the threshold, the tracked point is considered a false match point to be eliminated; otherwise, the tracking is successful, and finally, the accurate matching of point features is achieved.

[0094] As described above, it is only a preferred embodiment of the present invention, and it is not a limitation to the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still falls within the scope of protection required by the present invention.

Claims

1. A feature extraction and matching method for unmanned aerial vehicle navigation and positioning in underground space, characterized in that: The following steps are involved: Step 1: Collect the original image I1(x) of the underground space, and perform image preprocessing on the acquired original image to obtain I2(x); Step 2: Create an integral image based on the preprocessed image I2(x) to obtain the sum of the pixels I of the input image on the left side above the pixel at position (x, y). Σ(x) , using the improved SURF visual point feature extraction algorithm, calculate the point feature pixel coordinate set in the image Step 3: Describe the detected point features, assign the main direction of the features, construct feature descriptors, optimize the data structure, reduce the dimension of the descriptors, calculate the Euclidean distance between the feature points of the reference image and the image to be matched through the descriptors, and obtain the feature point matching feature set; Step 4: Preliminary elimination of mismatched points based on the RANSAC algorithm; Step 5: Perform point feature mismatch screening based on IMU pre-integration.

2. The feature extraction and matching method for unmanned aerial vehicle navigation and positioning in underground space according to claim 1 is characterized in that: The step 1 of performing image preprocessing on the collected underground space image data specifically includes: De-noising and contrast enhancement.

3. The feature extraction and matching method for unmanned aerial vehicle navigation and positioning in underground space according to claim 1 is characterized in that: The specific steps of step 2 are as follows: Create an integral image, I Σ(x) represents the sum of the pixels of the input image located to the left above the pixel at position (x, y), and its formula is: Then, the Box filter is used to replace the second-order Gaussian filter to detect and locate feature points in order to reduce the computational cost. The principle of the Box filter is to use the template to approximate the second-order partial derivative of Gaussian. The scale space function is obtained through the above two steps. The core of SURF feature point positioning is the Hessian matrix, and the formula is as follows: Where L xx (X, σ) is the second-order Gaussian partial derivative Convolution with each point X in the image, L xy (X, σ) is Convolution with point X, L yy (X, σ) is Convolution with point X, where the Gaussian function formula is: Among them, x and y represent the coordinates of point X; Next, convolution is performed with a multi-directional BOX filter to obtain D xx , D xy , D yy , replace L in the Hessian matrix xx , L xy , L yy , thus obtaining the SURF fast Hessian matrix, whose formula is as follows: it(H)=D xx D yy -w(D xy ) 2 Then, non-maximum suppression is applied in the 3×3×3 scale space, and the extreme points among the 26 neighbors of the selected point in the scale space are regarded as SURF feature points, which are invariant to the scale of the source image, and the point feature pixel coordinate set is obtained. .

4. The feature extraction and matching method for unmanned aerial vehicle navigation and positioning in underground space according to claim 1, characterized in that: The specific steps of step 3 are as follows: The direction of each SURF feature point is represented by the convolution of the Gaussian distribution-weighted Haar wavelet in the x and y directions, with the current feature point as the center and a sliding π / 3 window with a radius of 6σ, where σ is the size of the filter, to scan the wavelet response of the circular area here, and the sector with the largest wavelet response is determined as the main feature direction; After determining the main direction of the current feature point, a square sampling area is constructed. Then, the sampling area is divided into 16 sub-areas, and each sub-area is further divided into 25 sample grids. Four descriptors in the positive and negative directions of the x-axis and y-axis of each sub-area are obtained from the Haar wavelet convolution. Subsequently, the 64-dimensional descriptor of the SURF feature point is obtained. For the reference image and the image to be matched, the corresponding descriptors are used to calculate the Euclidean distance d between all SURF feature points in the two images. The minimum and second minimum values ​​of d are defined as the shortest distance d1 and the second shortest distance d2, respectively. If the ratio d1 / d2 is lower than the specified threshold, a pair of matching points can be obtained. All SURF feature points are obtained in this way, and all feature points form a matching feature set.

5. The feature extraction and matching method for unmanned aerial vehicle navigation and positioning in underground space according to claim 1 is characterized in that: The specific steps of step 4 are as follows: (a) Assume that the set of feature point matching pairs between the image to be tested and the template image is P, extract a certain number of feature point matching pairs n from it, and use n to initialize the transformation matrix M; (b) If the error between the feature points corresponding to the remaining matching pairs and M is less than the agreed threshold, then they form a consistent set with the extracted n matching pairs, and the error E i The calculation formula is as follows: (c) Use M to traverse the feature points corresponding to the matching pairs in P, calculate the proportion S of consistent sets in P that satisfy M under the error threshold, repeat the above steps, and select the transformation matrix with the largest proportion, denoted as *M; (d) Eliminate the elements of the set P whose errors exceed the threshold in the transformation matrix *M, and then use the remaining feature point matching pairs to obtain the transformation matrix, which is the final transformation matrix.

6. The feature extraction and matching method for unmanned aerial vehicle navigation and positioning in underground space according to claim 1, characterized in that: The specific steps of step 5 are as follows: The pose relationship between two adjacent frames can be obtained by the IMU pre-integration method, so a certain landmark point P = [x, y, z, 1] T The transformation relationship between the camera coordinate system and the world coordinate system in the two frames can be respectively: Where c1 and c2 are the camera coordinate systems of two adjacent frames, b1 and b2 are the carrier coordinate systems of two adjacent frames. Since the carrier coordinate system is fixedly connected to the IMU, the installation method of the camera and IMU determines the conversion matrix of the two. The transformation matrix of w system and b system is; The rotation and translation matrices can be obtained by the IMU pre-integration formula. The IMU measurements between the kth frame and the (k+1)th frame are integrated, and only the IMU zero bias ba, bg and random noise na, ng are considered, separated by time Δt, to derive the translation vector from the carrier coordinate system to the world coordinate system: Moving speed vector Rotation Quaternion The formula is as follows; The conversion between the camera coordinate system and the pixel coordinate system can be obtained by the intrinsic parameter matrix K. The pixel coordinates of the landmark point P in two adjacent frames are The conversion relationship between the camera coordinates is: The internal parameter K can be obtained by offline camera calibration; The matrix product is Ignoring the change in the depth of feature points between two adjacent frames, the above formula can be simplified to the conversion relationship between the pixel coordinates of the landmark point P between the two frames, where the matrix E utilizes the IMU pre-integration. The pixel coordinates of the landmark point in the first frame are predicted based on the IMU pre-integration results between the two frames. The predicted value is compared with the pixel coordinates of the second frame obtained by preliminarily eliminating false matches. If the distance is greater than the threshold, the tracked point is considered to be a false match point that needs to be eliminated. Otherwise, the tracking is successful, and accurate matching of point features is ultimately achieved.