Sliding window factor graph optimization method applied to underground space unmanned aerial vehicle
By applying the sliding window factor graph optimization method on the drone and combining multi-sensor data, the problems of high computing complexity and poor real-time performance in the underground space environment are solved, and efficient multi-sensor data fusion and drone real-time navigation are achieved.
Patent Information
- Application Number
- CN202510272755.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
In the underground space environment, the existing multi-sensor fusion technology is large in large-scale and dynamic environments and has poor real-time performance, making it difficult to meet the needs of real-time navigation and map updates of drones.
The sliding window factor graph optimization method is adopted, combined with IMU, lidar and camera sensors, and the position of the drone is optimized in real time through data preprocessing, feature matching and sliding window factor graph optimization.
It effectively reduces the computational complexity, improves real-timeness and positioning accuracy, and adapts to the complex and dynamic environmental changes in underground space.
Smart Images

Figure CN120219488A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous navigation and positioning of unmanned aerial vehicles, and more specifically, it is a sliding window factor graph optimization method applied to unmanned aerial vehicles in underground spaces. Background Art
[0002] With the rapid development of human society, the available surface space is becoming increasingly tense. As the third major exploitable field after cosmic space and ocean resources, the importance of underground space is becoming increasingly prominent. In underground spaces, due to the lack of GPS signals and complex environmental conditions, the autonomous flight, positioning, and mapping of unmanned aerial vehicles face huge challenges. To achieve precise positioning and environmental perception of unmanned aerial vehicles in underground spaces, multiple sensors (such as IMUs, lidars, and cameras) are usually used for data fusion. In existing multi-sensor fusion technologies, factor graph optimization methods are widely used in pose estimation and map construction. However, traditional graph optimization methods have large computational amounts and poor real-time performance in large-scale and dynamic environments, and it is difficult to meet the requirements of real-time navigation and map updating of unmanned aerial vehicles.
[0003] Currently, the sliding window factor graph optimization method has been proposed and used to reduce computational complexity, but how to efficiently fuse multi-sensor data and process dynamically changing environmental data in an underground space environment is still a technical problem. Therefore, how to reduce the computational burden while ensuring real-time performance and accuracy has become an urgent problem to be solved. Summary of the Invention
[0004] Aiming at how to perform efficient multi-sensor data fusion based on the sliding window factor graph optimization method in complex underground environments, improve the positioning accuracy of unmanned aerial vehicles, while reducing computational complexity and ensuring the real-time performance of the optimization process, the present invention proposes a sliding window factor graph optimization method applied to unmanned aerial vehicles in underground spaces. This method combines sensors such as IMUs, lidars, and cameras, and optimizes the pose of the unmanned aerial vehicle in real time through the sliding window factor graph optimization method.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A sliding window factor graph optimization method applied to unmanned aerial vehicles in underground spaces, characterized by comprising the following steps:
[0007] Step 1: Data preprocessing;
[0008] Preprocess the data collected from the IMU, lidar, and camera sensors. The IMU provides high-frequency acceleration and angular velocity data. The state of the drone can be obtained through IMU pre-integration. The lidar acquires the depth information of the environment and generates point cloud data. Preprocess the generated lidar point cloud data. The camera obtains the image information of the environment and provides visual constraints. Preprocess the camera images;
[0009] Step 2: Feature matching;
[0010] After preprocessing the data collected by the IMU, lidar, and camera, perform surface feature matching of the point cloud data line by line and visual point-line feature tracking respectively, and calculate the corresponding point cloud matching residuals and visual point-line residuals;
[0011] Step 3: Construct a sliding window factor graph for backend optimization;
[0012] After data preprocessing and feature matching, use the sliding window factor graph optimization method to estimate the optimal state of the system.
[0013] As a further improvement of the present invention, the data preprocessing in step 1 specifically includes:
[0014] 1) IMU pre-integration;
[0015] The state quantity of the drone at time k is;
[0016] I k =[R k ,p k ,v k ,b k
[0017] Among them, R k is the rotation matrix at time k, p k , v k , b k are the position, velocity, and bias obtained by IMU pre-integration respectively;
[0018] Add the state of the key frame to the sliding window, perform bundle adjustment optimization, realize the transformation from the IMU coordinate system to the world coordinate system, and obtain the relative motion measurement value between two consecutive IMU key frames m i and m i+1 :
[0019]
[0020] 2) Lidar point cloud data preprocessing;
[0021] First, the radar point cloud data is segmented using a fast point cloud segmentation algorithm based on random sample consensus. Then, the curvature of the segmented radar point cloud data is calculated using the depth map. The non-ground points with larger curvature are marked as edge feature points, and those with smaller curvature are marked as plane feature points. Finally, the pre-integrated data of the IMU is used to perform motion compensation on the radar point cloud data;
[0022] 3) Camera image preprocessing;
[0023] Before visual feature extraction, the image quality is enhanced to highlight the feature information in the image. First, the image is transformed from the RGB color space to the HSV color space to avoid image distortion caused by processing in the RGB color space. Second, only the luminance component is subjected to adaptive Gamma correction processing. Then, the CLAHE algorithm is applied to the Gamma correction result to transform the unevenly distributed histogram in the image into a uniformly distributed histogram, effectively stretching the dynamic range of the gray values and improving the image contrast. Finally, the unprocessed hue component, saturation component, and the processed luminance component are fused and inversely transformed back to the RGB color space to obtain the final enhanced image;
[0024] Visual point and line features are extracted from the enhanced image. First, the Kanade-Lucas-Tomasi feature tracking sparse optical flow algorithm is used to extract and track the initial feature points in each new image frame, keeping the number of point features in each frame of the image between 200 and 400. Then, the RANSAC and fundamental matrix model are used for outlier suppression to remove the feature points with large outliers. Finally, the Line Segment Detector is used to extract line segments, and the Line BandDescriptor is used to describe the line features.
[0025] As a further improvement of the present invention, the feature extraction in step 2 specifically includes:
[0026] 1) Matching of surface and line features of the point cloud data;
[0027] Considering that there are far more plane features than line features in the underground space environment, the plane features are used for the minimum distance constraint from point to plane to estimate the change amounts of the Z direction, roll angle, and pitch angle. On this basis, the minimum distance constraint from point to line is performed to estimate the change amounts of the X direction, Y direction, and yaw angle. The relative pose of consecutive frames is optimized by jointly considering the ground points and corner points. The Levenberg-Marquardt algorithm of Marquardt is used to iteratively solve the optimal pose, and the initial value of the iterative calculation is the pose of the current frame's IMU pre-integration, which can effectively reduce the number of iterations and eliminate the mis-matching caused by the zero initial value or the initial value assumption of uniform motion;
[0028]
[0029] 2) Visual point-line feature tracking;
[0030] First, select similar line segments for tracking, and add the set of feature points on image frames with similar line lengths to ensure that there are at least two feature points on the line segment in the new image. Then, track the line segments in two consecutive image frames. Finally, to improve the accuracy of the visual pose, construct local residual constraints using the point feature projection residuals, line feature projection residuals, and the measurement results of the IMU. When inserting a new key frame, perform a bundle adjustment to adjust the current mapping. In the sliding window, optimize all state variables by minimizing the prior terms and cost terms of all measurement residuals.
[0031]
[0032] Among them, is the set of adjacent key frames; and are the point feature projection residual and the line feature projection residual respectively; Γ p and Γ ζ are the set of point features and the set of line features that are observed at least twice in the sliding window respectively; and are the j-th visual point feature residual value and the k-th visual line feature residual value in the i-th image frame c i respectively; and are the set of point features and the set of line features in the i-th image frame c i respectively; is the prior information of the marginalized residual; r t is the marginalized residual information; J t is the marginalized Jacobian matrix; ∑ t is the marginalized set;
[0033] For point features, the reprojection error can be defined as the distance between the observed point and the reprojected point on the same pixel plane. Assume that the j-th visual point feature is first observed in the h-th image frame c h , then the reprojection error in the i-th image frame c i is;
[0034]
[0035] For line features, the reprojection error can be defined as the distance from the endpoints of the observed line to the reprojected line in the same pixel plane. Project a three-dimensional line onto the two-dimensional image plane, and the two-dimensional projected line segment i of the image frame c
[0036]
[0037] Among them, is the k-th visual line feature in image frame c; K is the internal parameter matrix of the three-dimensional line; N i is the rotation matrix of image frame c; f c , f v , f u are respectively the pixel values in the two-dimensional pixel coordinate system; c u , c v are respectively the pixel conversion values of image frame c in the two-dimensional pixel coordinate system; L1, L2, and l3 are respectively the line segment components in the three-axis directions; is the reprojection error of the visual line feature; is the distance between the k-th visual line feature in image frame c i and the k-th visual line feature projected onto the two-dimensional pixel plane; s k and l k are the end points observed in the two-dimensional pixel plane.
[0038] As a further improvement of the present invention, the construction of the sliding window factor graph in step 3 for backend optimization specifically includes:
[0039] Define a sliding window with a fixed size. When the window is full, the state nodes in the window need to be slid to achieve local bundle adjustment;
[0040] Adopt a non-linear optimization method to jointly optimize the observation values of lidar, camera, and IMU sensors. First, construct the measurement equations of different types of observations with respect to the state to be estimated, then construct the residual terms, and further add up the residuals of different types to obtain the cost function of the entire optimization problem;
[0041]
[0042] Among them, is the IMU pre-integration factor; is the weighted lidar line-plane residual factor; is the pose residual term between lidar key frame ψ and key frame γ; is the covariance matrix weight; is the weighted visual point-line residual factor; the pose residual term between visual key frame ψ and key frame γ; E m is the edge residual term.
[0043] Beneficial effects:
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] 1. Enhanced real-time performance: By adopting the sliding window optimization method, the scale of the factor graph is restricted, reducing the computational complexity and ensuring that the optimization process can be executed in real time to meet the positioning and navigation requirements of drones in complex underground environments.
[0046] 2. Reduced computational overhead: By using the marginalization technique to remove unnecessary historical data, the computational burden and memory consumption during the optimization process are significantly reduced, improving the computational efficiency of the system.
[0047] 3. Adaptation to dynamic environments: The sliding window optimization can quickly adjust the optimization results as sensor data is updated, adapting to the complex and dynamic environmental changes in underground spaces. Description of the Drawings
[0048] Figure 1 is the overall system flowchart of the present invention;
[0049] Figure 2 is the sliding window factor graph of the present invention. Detailed Embodiment
[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0051] As Figure 1 shown, the present invention provides a sliding window factor graph optimization method for underground space drones, which includes the following steps:
[0052] Step 1: Data preprocessing. Preprocess the data collected from sensors such as IMU, lidar, and camera. The IMU provides high-frequency acceleration and angular velocity data, and the drone state can be obtained through IMU pre-integration. The lidar obtains the depth information of the environment to generate point cloud data, and preprocess the generated lidar point cloud data. The camera obtains the image information of the environment to provide visual constraints, and preprocess the camera images.
[0053] Step 2: Feature matching. After preprocessing the data collected by the IMU, lidar, and camera, perform surface feature matching of the point cloud data line by line and visual point-line feature tracking respectively, and calculate the corresponding point cloud matching residuals and visual point-line residuals.
[0054] Step 3: Construct a sliding window factor graph for backend optimization. After data preprocessing and feature matching, use the sliding window factor graph optimization method to estimate the optimal state of the system. Define a sliding window with a fixed size.
[0055] Among them, step 1: Data preprocessing. Preprocess the data collected from sensors such as IMU, lidar, and camera. The IMU provides high-frequency acceleration and angular velocity data, and the UAV state can be obtained through IMU pre-integration. The lidar obtains the depth information of the environment, generates point cloud data, and preprocesses the generated lidar point cloud data. The camera obtains the image information of the environment, provides visual constraints, and preprocesses the camera images. Specifically as follows:
[0056] 1) IMU pre-integration. To avoid adding new state variables during each IMU measurement, a re-parameterization process is added between two frames to achieve motion constraints and avoid repeated integration.
[0057] The state of the UAV at time k is
[0058] I k =[R k , p k , v k , b k
[0059] where R k is the rotation matrix at time k, p k , v k , b k are the position, velocity, and bias obtained from IMU pre-integration respectively.
[0060] Add the state of the key frame to the sliding window and perform bundle adjustment optimization to achieve the transformation from the IMU coordinate system to the world coordinate system. The relative motion measurement value between two consecutive IMU key frames m i and m i+1 can be obtained:
[0061]
[0062] where: is the IMU pre-integration residual; is the set of IMU measurements pre-integrated in the sliding window; is the residual value between the i-th key frame and the (i + 1)-th key frame; I is the state of the UAV; δ is the pre-integration; is the IMU observation value between two adjacent key frames; b a and b g are the acceleration a bias and the gravity acceleration g bias respectively.
[0063] 2) Preprocessing of lidar point cloud data. The preprocessing of lidar point cloud data mainly includes point cloud data segmentation, line and plane feature extraction, and motion distortion correction. First, the lidar point cloud data is segmented using a fast point cloud segmentation algorithm based on Random Sample Consensus (RANSAC). Then, the curvature of the segmented lidar point cloud data is calculated using the depth map. Non-ground points with larger curvature are marked as edge feature points, and those with smaller curvature are marked as plane feature points, thereby extracting the line features and plane features of the underground space. Finally, to eliminate the problem of point cloud data distortion caused by sensor movement, the pre-integrated data of the IMU is used to perform motion compensation on the lidar point cloud data.
[0064] 3) Preprocessing of camera images. The weak texture and weak illumination in the underground space result in uneven illumination of the whole or part of the image, seriously affecting the extraction and tracking of visual features and easily causing drift or even failure of visual SLAM. The image quality is enhanced before visual feature extraction to highlight the feature information in the image. First, the image is transformed from the RGB space to the HSV space to avoid image distortion caused by processing in the RGB color space. Second, only the brightness component is subjected to adaptive Gamma correction processing. Then, the CLAHE algorithm is applied to the Gamma correction result to transform the unevenly distributed histogram in the image into a uniformly distributed histogram, effectively stretching the dynamic range of the gray values to improve the image contrast. Finally, the unprocessed hue component, saturation component, and the processed brightness component are fused and inversely transformed back to the RGB color space to obtain the final enhanced image.
[0065] Visual point and line feature extraction is performed on the enhanced image. First, the Kanade-Lucas-Tomasi (KLT) feature tracking sparse optical flow algorithm is used to extract and track the initial feature points in each new image frame, keeping the number of point features in each frame of the image between 200 and 400. Then, the RANSAC and fundamental matrix model are used for outlier suppression to remove feature points with large outliers. Finally, the Line Segment Detector (LSD) is used to extract line segments, and the Line Band Descriptor (LBD) is used to describe the line features.
[0066] Among them, step 2: Feature matching. After preprocessing the data collected by the IMU, lidar, and camera, point cloud data line and plane feature matching and visual point and line feature tracking are performed respectively, and the corresponding point cloud matching residuals and visual point and line residuals are calculated. Specifically as follows:
[0067] 1) Point cloud data line-plane feature matching. Considering that there are far more plane features than line features in the underground space environment, plane features are used to perform the minimum constraint of the distance from a point to a plane, and the change amounts of the Z direction, roll angle, and pitch angle are estimated. On this basis, the minimum constraint of the distance from a point to a line is performed to estimate the change amounts of the X direction, Y direction, and yaw angle, and the relative pose of consecutive frames is optimized by combining ground points and corner points. The Levenberg-Marquardt (LM) algorithm is used to iteratively solve the optimal pose, and the initial value of the iterative calculation is the pose of the current frame IMU pre-integration, which can effectively reduce the number of iterations and eliminate the mis-matching caused by the zero initial value or the initial value assumption of uniform motion.
[0068]
[0069] Among them, and are the distances from the A-th point to the line and from the B-th line to the plane respectively; Γ A and Λ B are the main direction and normal direction corresponding to the feature respectively; Q A and P A are the line feature points in the online feature point set set; R is the rotation matrix, t is the time; Q B and P B are the plane feature points in the plane feature point set set; is the residual of the distance from a point to a line and from a point to a plane; and are the prediction pose prior values provided by the IMU pre-integration factor and the IMU odometer respectively.
[0070] 2) Visual point-line feature tracking. First, similar line segments are selected for tracking, and the feature point set on the image frames with similar line segment lengths is added to ensure that there are at least 2 feature points on the line segment in the new image. Then, the lines in two consecutive image frames are tracked. Finally, to improve the accuracy of the visual pose, local residual constraints are constructed using the point feature projection residual, line feature projection residual, and the measurement results of the IMU. When inserting a new key frame, a bundle adjustment is performed on the current mapping. In the sliding window, all state variables are optimized by minimizing the prior terms and cost terms of all measurement residuals.
[0071]
[0072] Among them, is the set of adjacent key frames; and are the point feature projection residual and line feature projection residual respectively; Γ p and Γ ζThey are respectively the point feature set and the line feature set that are observed at least twice in the sliding window; and They are respectively the residual value of the j-th visual point feature and the residual value of the k-th visual line feature in the i-th image frame c i ; and They are respectively the point feature set and the line feature set in the i-th image frame c i ; is the prior information of the marginalized residual; r t is the marginalized residual information; J t is the marginalized Jacobian matrix; ∑ t is the marginalized set.
[0073] For point features, the reprojection error can be defined as the distance between the observed point and the reprojected point on the same pixel plane. Assume that the j-th visual point feature is first observed in the h-th image frame c h , then the reprojection error in the i-th image frame c i is
[0074]
[0075] For line features, the reprojection error can be defined as the distance from the endpoints of the observed line to the reprojected line in the same pixel plane. Project a three-dimensional line onto the two-dimensional image plane, and the two-dimensional projected line segment of the image frame c i can be obtained
[0076]
[0077] where is the k-th visual line feature in the image frame c i ; K is the internal parameter matrix of the three-dimensional line; N c is the rotation matrix of the image frame c; f v , f u are respectively the pixel values in the two-dimensional pixel coordinate system; c u , c v are respectively the pixel conversion values of the image frame c in the two-dimensional pixel coordinate system; L1, L2, and L3 are respectively the line segment components in the three-axis directions; is the reprojection error of the visual line feature; is the distance between the k-th visual line feature in the image frame c i and the k-th visual line feature projected onto the two-dimensional pixel plane; s k and l k are the endpoints observed in the two-dimensional pixel plane.
[0078] Among them, Step 3: Construct a sliding window factor graph for backend optimization. After data preprocessing and feature matching, the optimal state of the system is estimated using the sliding window factor graph optimization method. Define a sliding window of a fixed size. Specifically as follows:
[0079] To ensure high precision and real-time performance, this paper selects the sliding window size to be 8. When the window is full, the state nodes within the window need to be slid to achieve local bundle adjustment.
[0080] The observation values of lidar, camera, and IMU sensors are jointly optimized using a nonlinear optimization method. First, construct the measurement equations of different types of observations relative to the state to be estimated, then construct the residual terms, and further add up the different types of residuals to obtain the cost function of the entire optimization problem.
[0081]
[0082] Among them, is the IMU pre-integration factor; is the weighted lidar line-plane residual factor; is the pose residual term between lidar keyframe ψ and keyframe γ; is the covariance matrix weight; is the weighted visual point-line residual factor; The pose residual term between visual keyframe ψ and keyframe γ; E m is the edge residual term.
[0083] After the residual terms are constructed, the sliding window factor graph of the system is as Figure 2 shown. Using the GTSAM library, a fixed-lag smoothing framework based on the efficient incremental optimizer iSAM2 is used to solve the factor graph. Add all visual and lidar factors to the graph using a dynamically weighted covariance scaling (DCS) robust cost function to reduce the influence of outliers. Finally, the optimized pose of the drone is output.
[0084] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.
Claims
1. A sliding window factor graph optimization method applied to underground space drones, characterized by: The following steps are involved: Step 1: Data preprocessing; Preprocess the data collected from IMU, LiDAR and camera sensors. IMU provides high-frequency acceleration and angular velocity data. The drone status can be obtained through IMU pre-integration. LiDAR obtains the depth information of the environment and generates point cloud data. The generated LiDAR point cloud data is preprocessed. The camera obtains the image information of the environment, provides visual constraints, and performs data preprocessing on the camera image. Step 2: Feature matching; After preprocessing the data collected by IMU, LiDAR and phase, point cloud data line and surface feature matching and visual point and line feature tracking are performed respectively to calculate the corresponding point cloud matching residuals and visual point and line residuals; Step 3: Build a sliding window factor graph for backend optimization; After data preprocessing and feature matching, the optimal state of the system is estimated using a sliding window factor graph optimization method.
2. The sliding window factor graph optimization method for underground space drones according to claim 1 is characterized in that: The data preprocessing in step 1 specifically includes: 1) IMU pre-integration; The state quantity of the drone at time k is; I k =[R k ,p k ,v k ,b k ] Among them, R k is the rotation matrix at time k, p k ,v k ,b k They are the position, velocity and bias obtained by IMU pre-integration respectively; The key frame state is added to the sliding window, and the bundle adjustment optimization is performed to realize the transformation from the IMU coordinate system to the world coordinate system, and the continuous IMU key frame m is obtained. i and m i+1 Relative motion measurements between: 2) LiDAR point cloud data preprocessing; Firstly, the radar point cloud data is segmented by a fast point cloud segmentation algorithm based on random sample consistency. Then, the curvature of the segmented radar point cloud data is calculated using the depth map. Non-ground points with larger curvature are marked as edge feature points, and those with smaller curvature are marked as plane feature points. Finally, the pre-integrated data of the IMU is used to perform motion compensation on the radar point cloud data. 3) Camera image preprocessing; Before visual feature extraction, the image quality is enhanced to highlight the feature information in the image. First, the image is transformed from RGB space to HSV space to avoid image distortion caused by processing in RGB color space. Secondly, only the brightness component is subjected to adaptive Gamma correction. Then, the Gamma correction result is processed by CLAHE algorithm to transform the unevenly distributed histogram in the image into a uniformly distributed histogram, effectively stretching the dynamic range of grayscale values and improving the image contrast. Finally, the unprocessed hue component, saturation component and processed brightness component are fused and inversely transformed back to RGB color space to obtain the final enhanced image. Visual point and line feature extraction is performed on the enhanced image. First, the Kanade-Lucas-Tomasi feature tracking sparse optical flow algorithm is used to extract and track the initial feature points in each new image frame, keeping the number of point features in each frame image at 200-400. Then, RANSAC and the basic matrix model are used to suppress outliers and remove feature points with large outliers. Finally, the LineSegment Detector is used to extract line segments, and the Line Band Descriptor is used to describe the line features.
3. The sliding window factor graph optimization method for underground space drones according to claim 1 is characterized in that: The feature extraction in step 2 specifically includes: 1) Line and surface feature matching of point cloud data; Considering that there are far more plane features than line features in underground space environments, plane features are used to constrain the minimum distance from point to surface, and the changes in the Z direction, roll angle, and pitch angle are estimated; on this basis, the minimum distance from point to line is constrained, and the changes in the X direction, Y direction, and yaw angle are estimated. The relative pose of continuous frames is estimated by combining ground points and corner points for optimization; the Levenberg-Marquardt algorithm is used to iteratively solve the optimal pose, and the initial value of the iterative calculation is the pose of the IMU pre-integration of the current frame, which can effectively reduce the number of iterations and eliminate the mismatch caused by the zero initial value or the uniform motion assumption initial value; 2) Visual point and line feature tracking; First, similar line segments are selected for tracking, and feature point sets on image frames with similar line segment lengths are added to ensure that there are at least two feature point line segments in the new image. Then, two consecutive image frame lines are tracked. Finally, in order to improve the accuracy of visual pose, local residual constraints are constructed using point feature projection residuals, line feature projection residuals and IMU measurement results. When a new key frame is inserted, a bundle adjustment is performed on the current map. In the sliding window, all state variables are optimized by minimizing the prior terms and cost terms of all measurement residuals. in, is a set of adjacent key frames; and are the point feature projection residual and line feature projection residual respectively; Γ p and Γ ζ They are the point feature set and line feature set observed at least twice in the sliding window; and are the i-th image frame c i The j-th visual point feature residual value and the k-th visual line feature residual value in ; and are respectively in the i-th frame image frame c i Midpoint feature set, line feature set; is the prior information of marginalized residual; r t is the marginalized residual information; J t is the marginalized Jacobian matrix; ∑ t for the marginalized collection; For point features, the reprojection error can be defined as the distance between the observation point and the reprojection point on the same pixel plane. Assuming that the jth visual point feature is in the hth image frame c h is first observed in the i-th image frame c i The reprojection error in is: For line features, the reprojection error can be defined as the distance from the endpoint of the observed line to the reprojected line in the same pixel plane. Projected onto the two-dimensional image plane, we can get the image frame c i The two-dimensional projection line segment of in, For image frame c i The kth visual line feature in; K is the intrinsic parameter matrix of the three-dimensional line; N c is the rotation matrix of image frame c; f v , f u are the pixel values in the two-dimensional pixel coordinate system; c u , c v are the pixel conversion values of the image frame c in the two-dimensional pixel coordinate system; L1, L2, L3 are the line segment components in the three-axis directions; is the reprojection error of the visual line feature; is the image frame c i The distance between the kth visual line feature in and the kth visual line feature projected onto the two-dimensional pixel plane; s k and l k is the endpoint observed on the two-dimensional pixel plane.
4. The sliding window factor graph optimization method for underground space drones according to claim 1 is characterized in that: The step 3 of constructing a sliding window factor graph for backend optimization specifically includes: Define a sliding window of fixed size. When the window is full, the state nodes in the window need to be slid to achieve local bundle adjustment. The nonlinear optimization method is used to jointly optimize the observation values of the lidar, camera, and IMU sensors. First, the measurement equations of different types of observation values relative to the state to be estimated are constructed, and then the residual terms are constructed. The different types of residuals are further added together to obtain the cost function of the entire optimization problem. in, is the IMU pre-integration factor; is the weighted radar line-surface residual factor; is the pose residual term between the radar key frame ψ and the key frame γ; is the covariance matrix weight; is the weighted visual point-line residual factor; The pose residual term between the visual keyframe ψ and the keyframe γ; E m is the marginal residual term.