Multimodal robust positioning method for UAVs in underground space based on confidence factor
Through the multimodal robust positioning method based on confidence factor, the problem of insufficient underground space positioning accuracy and robustness of the UAV is solved. Through the front-end preprocessing and back-end optimization stage, dot-line features are extracted and adaptive confidence factors are constructed, sensor weights are dynamically adjusted, and positioning is optimized, so that high-precision and robust underground space positioning are achieved.
Patent Information
- Application Number
- CN202411810800.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-12-10
AI Technical Summary
The existing underground space positioning method of UAVs lacks positioning accuracy and robustness in scenarios such as poor lighting conditions, weak texture characteristics and sensor degradation, and lacks effective evaluation of the degree of sensor degradation and dynamic weight allocation.
Using a multimodal robust positioning method based on confidence factor, through front-end preprocessing and back-end optimization stages, dotted and line features are extracted and adaptive confidence factors are constructed, asymptotic non-convex factor graph optimization function is designed, sensor weights are dynamically adjusted, and position poses are optimized to improve positioning accuracy and robustness.
In low illumination and complex underground environments, the positioning accuracy and system stability of the drone are significantly improved, the influence of sensor outliers is reduced, and the adaptive sensor weight allocation is achieved.
Smart Images

Figure CN119758352B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and multi-sensor fusion, and in particular to the technology of visual-inertial-lidar fusion, and mainly relates to a multi-modal robust positioning method for unmanned aerial vehicles in underground space based on confidence factors. Background Art
[0002] With the development of simultaneous localization and mapping (SLAM) technology, multi-sensor fusion for unmanned systems has become one of the most active research areas in recent years. While IMUs provide high-frequency measurements unaffected by external factors such as illumination, texture, and weather, they exhibit time-varying biases, requiring other sensors to provide effective constraints to reduce measurement errors. Cameras perceive rich texture information in the environment, but vision-based SLAM systems struggle to track visual features in areas with weak textures or under poor lighting conditions, leading to system degradation or even failure. LiDARs provide accurate depth measurements and perceptual structural information, and are less susceptible to illumination effects. However, their point cloud resolution is relatively low. In scenarios with degraded structure, performance degrades due to insufficient constraints. Consequently, in recent years, many researchers have focused on fusion visual-inertial-liDAR positioning systems.
[0003] Existing methods only fuse these three sensors and fail to consider extreme conditions such as underground spaces. The main challenges encountered in such situations include: 1) Underground spaces have poor lighting conditions and weak texture features. The insufficient features extracted by the camera make the visual sensor constraints unreliable during the fusion process; 2) Due to the complexity of underground spaces, periodic outliers appear in sensor measurements. Existing fusion frameworks fail to consider the degree of degradation of different sensors, resulting in low robustness of back-end optimization; 3) During the pose graph optimization process, existing methods lack quantitative evaluation of degraded scenes and cannot adaptively assign weights to different factors in the visual-inertial-lidar fusion framework. Summary of the Invention
[0004] The present invention is aimed at the problems existing in the prior art and proposes a multimodal robust positioning method for unmanned aerial vehicles in underground space based on confidence factors. The method includes two stages, front-end and back-end. In the front-end stage, the measurement data of the camera, lidar and IMU are preprocessed, and point and line features are extracted for adaptive visual / inertial / lidar initialization; in the back-end stage, IMU measurement residuals, lidar edge-plane residuals, visual point-line residuals and Manhattan structure constraint residuals are constructed according to the features extracted and processed in the front-end stage; an adaptive confidence factor is designed to evaluate the degree of degradation of the camera and lidar by the number of visual feature tracking and the difference between the reference feature and the transformed feature in the lidar point cloud. The adaptive confidence factor includes at least a visual confidence factor and a lidar confidence factor; according to the weight of each sensor, an asymptotically non-convex factor graph optimization function is constructed, and dynamic adjustment of the posture optimization is achieved according to the degradation degree weight of each sensor to obtain a positioning result, thereby improving the positioning accuracy and robustness of the unmanned aerial vehicle in underground space.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is: a multimodal robust positioning method for underground space of unmanned aerial vehicles based on confidence factors, including two stages: front-end and back-end.
[0006] In the front-end stage: pre-processing the measurement data of the camera, lidar and IMU, extracting point and line features for adaptive vision / inertial / lidar initialization;
[0007] In the back-end stage: IMU measurement residuals, lidar edge-plane residuals, visual point-line residuals, and Manhattan structure constraint residuals are constructed based on the features extracted and processed in the front-end stage; then, an adaptive confidence factor is constructed to evaluate the degree of degradation of the camera and lidar by the number of visual feature tracking and the difference between the reference feature and the transformed feature in the lidar point cloud. The adaptive confidence factor includes at least a visual confidence factor and a lidar confidence factor; according to the weight of each sensor, an asymptotically non-convex factor graph optimization function is constructed, and dynamic adjustment of the pose optimization is achieved according to the degradation degree weight of each sensor to obtain a positioning result; wherein,
[0008] The laser radar edge-plane residual is used to minimize the geometric error between the feature points and the target points in the point cloud;
[0009] The visual point-line residual: uses the geometric relationship between point features and line features in the image to constrain the camera's position;
[0010] The Manhattan structure constrained residual uses the Manhattan world assumption in the scene to constrain pose estimation and map optimization.
[0011] As an improvement of the present invention, the front-end stage includes a lightweight deep learning network DCE-Net for image enhancement. If the lidar is not degraded, lidar / inertial initialization is performed to obtain the bias of the accelerometer and gyroscope, and the visual scale is restored through the depth value of the point cloud.
[0012] As an improvement of the present invention, in the back-end stage, the lidar edge-plane residual is constructed by the point-to-edge and point-to-plane matching residuals, specifically:
[0013]
[0014] in, represents the marginal residual, represents the plane residual, T i j represents the transition matrix between time i and time j, Indicates the mth edge point and nth surface point at time i, Represents the coordinates of points A, B, and C.
[0015] As an improvement of the present invention, the construction of the visual point-line residual in the back-end stage specifically includes the following steps:
[0016] S21: Detecting ORB features in the camera image, including FAST key points and BRI EF descriptors, extracting line features, and enhancing the line features. The enhancement method at least includes introducing hidden layer parameter adjustment, short line removal, and broken line merging.
[0017] S22: After receiving the new visual features obtained in step S21, the visual point and line residuals are constructed through the point and line reprojection residuals:
[0018]
[0019] in, represents the residual of the point feature set a, Represents the residual of the line feature set b, l b,1 ,l b,2 Represents line features l b, The components in the x and y planes, represents the camera projection model, Represents the coordinates of the tracking point features at time i and j in set a, are the projection coordinates of the two endpoints of the tracking line feature in the i-th frame on the image.
[0020] As another improvement of the present invention, in the back-end stage, the Manhattan structure constraint residual constructed is specifically:
[0021]
[0022] in, is the normalized 3D line feature, They are vertical direction error and parallel direction error respectively.
[0023] As another improvement of the present invention, in the back-end stage, the visual confidence factor and the lidar confidence factor in the adaptive confidence factor are respectively:
[0024] s visual =u / W
[0025] s lidar =exp(-d mse / λ 2 )
[0026] Among them, u represents the number of visual feature tracking, W represents the size of the sliding window, and d mse represents the average error between the reference features and the transformed features in the lidar point cloud, and λ is the distance scale adjustment threshold. Based on these two confidence factors, the penalty function of the visual and lidar weights is designed:
[0027]
[0028] Among them, ω visual ,ω lidar represents the weight of the residual of visual and lidar features, and θ represents the control parameter that controls the non-convexity of the penalty function.
[0029] As another improvement of the present invention, the asymptotically non-convex factor graph optimization function in the back-end stage is specifically:
[0030]
[0031] Among them, x * represents the optimized state variable, N1, N2, N3, N4, and N5 represent the number of IMU pre-integration residuals, the number of lidar edge residuals, the number of lidar plane residuals, the number of visual point feature residuals, and the number of Manhattan structure constraint residuals in the sliding window, respectively; r i imu , They represent the i-th IMU pre-integration residual, the m-th lidar edge residual, the n-th lidar plane residual, the a-th visual point feature residual, the b-th visual line feature and the b-th Manhattan structure constraint residual respectively; Φ(ω visual ) and Φ(ω lidar ) represent the penalty functions of vision and lidar weights respectively.
[0032] As a further improvement of the present invention, in the back-end stage, dynamic adjustment of the pose optimization is achieved according to the degradation degree weight of each sensor, specifically by solving the asymptotically non-convex factor graph optimization function. The solution process includes inner iteration and outer iteration:
[0033] In each outer iteration, the control parameter θ is fixed;
[0034] In each inner iteration, first fix the weight of t-1 To optimize the state variable x at time t t , and then by fixing the state variable x at time t t To optimize the weight at time t Get the residual weights of visual and lidar features at time t
[0035] By changing the value of θ to repeat the internal and external iteration process, θ is adjusted to increase the non-convexity of the optimization function, and the optimal UAV posture is obtained to achieve positioning.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) The present invention introduces a lightweight DCE-Net in the preprocessing process of the front-end stage, and extracts robust point and line features at the same time, which improves the image quality under low illumination, makes the visual feature constraints reliable, and reduces the time cost.
[0038] (2) The visual / inertial / lidar asymptotically non-convex factor graph optimization function constructed in the present invention reduces the impact of abnormal values of each sensor on state estimation in underground space degradation scenarios.
[0039] (3) The present invention designs an adaptive confidence factor to evaluate the degree of degradation of the camera and lidar by the number of visual feature tracking and the difference between the reference features and the transformed features in the lidar point cloud, thereby realizing dynamic adjustment of graph optimization and improving the robustness of underground space state estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the working principle of the method of the present invention;
[0041] Figure 2 Schematic diagram of the structure of the robust factor graph optimization model in the back-end stage of the method of the present invention;
[0042] Figure 3 Schematic diagram of the UAV platform built in the test example of the present invention;
[0043] Figure 4 The scene diagram of the dataset selected for the test example in this invention;
[0044] Figure 5 This is a comparison chart of the trajectory results of different algorithms on the UG_03 sequence in the test example of the present invention;
[0045] Figure 6 This is a comparison chart of the trajectory results of different algorithms on the UC_02 sequence in the test example of the present invention;
[0046] Figure 7 This is a comparison chart of the trajectory results of different algorithms on the EB_01 sequence in the test example of the present invention. DETAILED DESCRIPTION
[0047] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0048] Example 1
[0049] The multimodal robust positioning method for UAV underground space based on confidence factor works as follows Figure 1 As shown, the following steps are included:
[0050] Step S1, front-end: First, the measurement data of the camera, lidar and IMU are input into the preprocessing module. In the preprocessing module, IMU pre-integration assists in removing the distortion of the lidar point cloud, and these dedistorted point clouds are used to perform depth association with visual features. Then, a lightweight deep learning network DCE-Net is introduced to enhance low-light images in underground spaces. Next, a feature extraction method based on LSD is adopted to improve the LSD algorithm from three aspects: hidden layer parameter adjustment, short line removal and broken line fusion, to extract robust point and line features. Finally, adaptive visual / inertial / lidar initialization is performed to obtain reliable initial values. If the lidar is not degraded, lidar / inertial initialization is performed to obtain the bias of the accelerometer and gyroscope, and the visual scale is restored through the depth value of the point cloud. Otherwise, visual / inertial initialization is performed using the point and line features extracted in the previous step.
[0051] In step S2, in the backend, first, the features processed by the frontend are used to construct the IMU measurement residuals, lidar edge-plane residuals, visual point-line residuals, and Manhattan structure constraint residuals; then, an adaptive confidence factor is designed to evaluate the degree of degradation of the camera and lidar through the number of visual feature tracking and the difference between the reference features and the transformed features in the lidar point cloud; finally, according to the weight of each sensor, an asymptotically non-convex factor graph optimization function is constructed to achieve dynamic adjustment of the pose optimization according to the degradation degree weight of each sensor, thereby obtaining a high-precision and robust positioning result.
[0052] Among them, the LiDAR edge-plane residual is used to minimize the geometric error between the feature points (such as edge points and plane points) in the point cloud and the target point, thereby improving the accuracy of pose estimation. During the operation of the UAV, a new LiDAR feature is received at time i. After, among them represents the lidar edge feature at time i, Represents the LiDAR plane feature at time i, and constructs the LiDAR edge-plane residual through point-to-edge and point-to-plane matching residuals:
[0053]
[0054] in, represents the marginal residual, represents the plane residual, T i j represents the transition matrix between time i and time j, Indicates the mth edge point and nth surface point at time i, Represents the coordinates of points A, B, and C.
[0055] Visual point-line residual, utilizes the geometric relationship between point features and line features in the image to constrain the camera pose, thereby improving the robustness and accuracy of the system.
[0056] When the visual sensor is working normally, point feature extraction is first performed to detect ORB features in the image, including FAST key points and BRI EF descriptors. At the same time, line feature extraction is performed, and three strategies are introduced to enhance line feature extraction: hidden layer parameter adjustment, short line removal, and broken line merging. This not only ensures the system's effective feature information but also enhances the system's real-time performance.
[0057] After receiving the new visual features, the visual point and line residuals are constructed through the point and line reprojection residuals:
[0058]
[0059] in, represents the residual of the point feature set a, Represents the residual of the line feature set b. b,1 ,l b,2 Represents line features l b, Components in the x, y plane. represents the camera projection model, Represents the coordinates of the tracking point features at time i and j in set a, are the projection coordinates of the two endpoints of the tracking line feature in the i-th frame on the image.
[0060] Manhattan structure constraint residuals are mainly used to constrain pose estimation and map optimization by using the Manhattan world assumption in the scene. It reduces the degrees of freedom and improves robustness and accuracy by assuming that the geometric structure of the scene follows a set of global orthogonal directions.
[0061] During the operation of the drone, after receiving the line features extracted from each frame, the three Manhattan axes are used to parallel judge each line feature. If there is a parallel relationship, the line feature is associated with the corresponding Manhattan axis. Before building the nonlinear optimization model, it is necessary to define the parallel and perpendicular structural residuals between line features. For the normalized 3D line features The vertical direction error and the parallel direction error can be defined as Further construct the structural constraint between the Manhattan axis and the associated line features, which is defined as the Manhattan structural constraint residual:
[0062]
[0063] Due to the abnormality of the sensor and the complex interference of the underground space, the measurement information of the sensor is easily interfered with seriously. In order to improve the positioning accuracy and system stability in the feature degradation environment, the visual confidence factor s is designed. visual and the lidar confidence factor s lidar , which can quantitatively reflect the quality of feature matching and evaluate the degree of degradation of the camera and lidar. The degree of degradation of the camera and lidar is evaluated by the number of visual feature tracking and the difference between the reference feature and the transformed feature in the lidar point cloud, and the visual confidence factor and lidar confidence factor are designed respectively:
[0064] s visual =u / W
[0065] s lidar =exp(-d mse / λ 2 )
[0066] Among them, u represents the number of visual feature tracking, W represents the size of the sliding window, and d mse represents the average error between the reference features and the transformed features in the lidar point cloud, and λ is the distance scale adjustment threshold, which ensures that the distance error is within a certain scale range.
[0067] The more visual features tracked in the sliding window, the more accurate the visual information. Under these two confidence factors, the penalty function of visual and lidar weights is designed:
[0068]
[0069] Among them, ω visual ,ω lidarrepresents the weight of the residual of visual and lidar features, and θ is a control parameter that controls the non-convexity of the penalty function.
[0070] While factor graphs can effectively handle heterogeneous and nonlinear measurement information from visual-inertial-lidar systems, they cannot correct for sensor anomalies. To mitigate the impact of factor graph degradation in underground spatial environments, a visual / inertial / lidar asymptotically non-convex factor graph optimization function is constructed using designed IMU measurement residuals, lidar edge-plane residuals, visual point-line residuals, and Manhattan structure constraint residuals.
[0071] First, define the state variables to be optimized for the integrated navigation system:
[0072]
[0073] in, They represent the position, velocity and rotation of the carrier in the world coordinate system w of the i-th frame, respectively, and b a ,b g represents the bias of the accelerometer and gyroscope, λ m Represents the inverse depth of the mth tracked feature. The state x can be obtained by nonlinear least squares optimization problem i Considering the sensor degradation, an asymptotically non-convex optimization function is designed.
[0074]
[0075] where r(y i ,x) represents the measurement residual of each sensor, ω i represents the weight of each sensor, Φ ρ (ω i ) represents the penalty function of each sensor weight. Design visual confidence factor s visual and the lidar confidence factor s lidar , construct a robust factor graph optimization model, as follows Figure 2 As shown in the figure, circles represent nodes in the factor graph, which represent sensors; lines between nodes represent edges in the factor graph, which represent the constraint relationships between sensors.
[0076] The asymptotically non-convex factor graph optimization function for visual / inertial / lidar is:
[0077]
[0078] Among them, x *represents the optimized state variable, N1, N2, N3, N4, and N5 represent the number of IMU pre-integration residuals, the number of lidar edge residuals, the number of lidar plane residuals, the number of visual point feature residuals, and the number of Manhattan structure constraint residuals in the sliding window, respectively; r i imu , They represent the i-th IMU pre-integration residual, the m-th lidar edge residual, the n-th lidar plane residual, the a-th visual point feature residual, the b-th visual line feature and the b-th Manhattan structure constraint residual respectively; Φ(ω visual ) and Φ(ω lidar ) represent the penalty functions of vision and lidar weights respectively.
[0079] Solve the asymptotically non-convex factor graph optimization function of visual / inertial / lidar through iterative optimization, which mainly includes inner iteration and outer iteration.
[0080] In each outer iteration, the control parameter θ is fixed.
[0081] In the inner iteration, we first fix the weight of t-1 To optimize the state variable x at time t t :
[0082]
[0083] Among them, the penalty function Φ that is irrelevant to x is removed ρ (ω i ), The weight ω contains the residuals of visual and lidar features visual ,ω lidar Then by fixing the state variable x at time t t To optimize the weight at time t
[0084]
[0085] Among them, the residual r(y i ,x t ) is a constant in this step. Combining the above formula and the designed confidence factor, the residual weights of the visual and lidar features at time t can be obtained.
[0086]
[0087] The larger the residual factor of different sensors, the smaller the confidence, and the smaller the assigned weight. By changing the value of θ to repeat the internal and external iteration process, θ is adjusted to increase the non-convexity of the optimization function, and finally the optimized drone posture is obtained.
[0088] Test Case
[0089] Existing public datasets lack challenging scenarios such as underground spaces. In order to verify the robustness of the proposed algorithm, a drone experimental platform was built to simulate challenging scenarios such as poor lighting conditions, sparse features, and structural degradation in underground spaces for real data collection.
[0090] The drone experimental platform built in this test case is based on the PX4 autopilot, with a 50A four-in-one electronic speed controller and a 2810-900kV brushless motor. The flight control system uses a CUAV V5 nano, and the power supply is a 6s 5500mAh 75C aircraft model battery. Environmental perception uses an Intel D435i camera, an Xsens MT i-7MEMS-I MU, a Livox mid360 lidar, and an MTF-01 optical flow module. An Intel onboard computer is also used for data processing. The completed drone platform is shown below. Figure 3 shown.
[0091] In order to fully verify the proposed method in the complex environment of underground space, multiple sets of data were collected in the underground garage (UG), experimental building (EB) and underground patio (UC) on the campus of Southeast University. The scene pictures are as follows: Figure 4 As shown in the figure, UG has dim lighting, simulating the visual degradation of underground space; EB has a parallel corridor structure and transparent glass, simulating the degradation of underground space lidar; UC has variable lighting and complex characteristics, simulating the complex scene of underground space.
[0092] In order to evaluate the robustness of the proposed method, experiments were conducted on a self-collected UAV flight dataset. In order to improve the positioning accuracy and system stability in feature degradation environments, two adaptive confidence factors s are designed. visual and s lidar To reflect the quality of feature matching and evaluate the degree of degradation of the camera and lidar. In this test case, an ablation experiment was set up by removing these two factors to verify their impact on positioning accuracy. Three typical sequences were selected for the ablation experiment, and the complete algorithm was compared with the algorithm without the visual confidence factor, the algorithm without the lidar confidence factor, and the fast l ivo algorithm. "proposed-sv" means removing the visual confidence factor s. visual "proposed-s l" means deleting the lidar confidence factor s lidar , “proposed-sv-s l” means deleting two confidence factors, “proposed” means the complete method, and the comparison of trajectory results is as follows Figures 5 to 7 As shown, Figure 5 represents the UG_03 sequence, Figure 6 represents the UC_02 sequence, Figure 7 Represents the EB_01 sequence. From the visualization trajectory results, it can be seen that the complete method has stronger robustness.
[0093] Furthermore, the positioning accuracy is quantitatively evaluated by comparing the ground distance between the starting point and the end point of the positioning results of different algorithms. The real translation distance is calculated from the recorded start and end point coordinates and defined as the "truth". The results are shown in Table 1, where "proposed-sv" means deleting the visual confidence factor s. visual The proposed-s l method means deleting the lidar confidence factor s lidar “proposed-sv-s l” represents the proposed method with two confidence factors removed, and “proposed” represents the complete method.
[0094] Table 1 Network model training parameters
[0095]
[0096] As can be seen from the table above, the positioning results of the proposed method are closer to the real data. The confidence factor in the proposed method is crucial; the lack of any one of these factors will lead to a decline in method performance. Furthermore, a comparison of the proposed method with Fast-LIVO using a self-collected dataset shows that the proposed method achieves much higher positioning accuracy than Fast-LIVO.
[0097] In summary, the method of the present invention preprocesses the input data at the front end, uses a lightweight deep network DCE-Net to improve the quality of the image under low illumination conditions, and then extracts point and line features and initializes the system; in the back end, first constructs the IMU measurement residual, the lidar edge-plane residual, the visual point and line residual and the Manhattan structure constraint residual, and then designs an adaptive confidence factor, and uses the tracking number of visual feature points and the matching difference between the reference feature and the target feature in the lidar point cloud to evaluate the degradation degree of the camera and lidar respectively; finally, constructs an asymptotically non-convex factor graph optimization function, and realizes dynamic adjustment of the pose optimization according to the degradation degree weight of each sensor, thereby improving the positioning accuracy and robustness of the underground space drone, with better results.
[0098] It should be noted that the above content merely illustrates the technical idea of the present invention and cannot be used to limit the scope of protection of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications all fall within the scope of protection of the claims of the present invention.
Claims
1. A multimodal robust positioning method for UAVs in underground spaces based on confidence factors, characterized by: It includes two stages: front-end and back-end. In the front-end stage, the camera, lidar, and IMU measurement data are preprocessed, and point and line features are extracted for adaptive visual / inertial / lidar initialization. The front-end stage includes a lightweight deep learning network DCE-Net for image enhancement. If the lidar is not degraded, the lidar / inertial initialization is performed to obtain the bias of the accelerometer and gyroscope, and the visual scale is restored using the depth value of the point cloud. If the lidar is degraded, the extracted visual point and line features are used for visual / inertial initialization; In the back-end stage: IMU measurement residuals, lidar edge-plane residuals, visual point-line residuals, and Manhattan structure constraint residuals are constructed based on the features extracted and processed in the front-end stage; then, an adaptive confidence factor is designed to evaluate the degree of degradation of the camera and lidar by the number of visual feature tracking and the difference between the reference feature and the transformed feature in the lidar point cloud, and the adaptive confidence factor includes at least a visual confidence factor and a lidar confidence factor; according to the weight of each sensor, an asymptotically non-convex factor graph optimization function is constructed, and dynamic adjustment of the pose optimization is achieved according to the degradation degree weight of each sensor to obtain a positioning result; wherein, The laser radar edge-plane residual is used to minimize the geometric error between the feature points and the target points in the point cloud; The visual point-line residual: uses the geometric relationship between point features and line features in the image to constrain the camera's position; The Manhattan structure constrained residual: uses the Manhattan world assumption in the scene to constrain pose estimation and map optimization; The visual confidence factor and lidar confidence factor in the adaptive confidence factor are: s visual =u / W s lidar =exp(-d mse / λ 2 ) Among them, u represents the number of visual feature tracking, W represents the size of the sliding window, and d mse represents the average error between the reference features and the transformed features in the lidar point cloud, and λ is the distance scale adjustment threshold. Based on the two confidence factors, the penalty function of the visual and lidar weights is designed: Among them, ω visual ,ω lidar represents the weight of the residual of visual and lidar features, and θ represents the control parameter that controls the non-convexity of the penalty function.
2. The confidence factor-based multimodal robust positioning method for underground UAVs according to claim 1, characterized in that: In the back-end stage, the lidar edge-plane residual is constructed through the point-to-edge and point-to-plane matching residuals, specifically: in, represents the marginal residual, represents the plane residual, T i j represents the transition matrix between time i and time j, Indicates the mth edge point and nth surface point at time i, Represents the coordinates of points A, B, and C.
3. The confidence factor-based multimodal robust positioning method for underground UAVs according to claim 2, characterized in that: The construction of the visual point-line residual in the back-end stage specifically includes the following steps: S21: Detect ORB features in the camera image, including FAST key points and BRIEF descriptors, perform line feature extraction, and enhance the line features. The enhancement method at least includes introducing hidden layer parameter adjustment, short line removal, and broken line merging. S22: After receiving the new visual features obtained in step S21, the visual point and line residuals are constructed through the point and line reprojection residuals: in, represents the residual of the point feature set a, Represents the residual of the line feature set b, l b,1 ,l b,2 Represents the line features l b Components in the x, y plane, π: represents the camera projection model, Represents the coordinates of the tracking point features at time i and j in set a, are the projection coordinates of the two endpoints of the tracking line feature in the i-th frame on the image.
4. The confidence factor-based multimodal robust positioning method for underground UAVs according to claim 3, characterized in that: In the back-end stage, the Manhattan structure constraint residual constructed is specifically: in, is the normalized 3D line feature, They are vertical direction error and parallel direction error respectively.
5. The confidence factor-based multimodal robust positioning method for underground UAVs according to claim 4, characterized in that: The asymptotically non-convex factor graph optimization function in the back-end stage is specifically: Among them, x * represents the optimized state variable, N1, N2, N3, N4, and N5 represent the number of IMU pre-integration residuals, the number of lidar edge residuals, the number of lidar plane residuals, the number of visual point feature residuals, and the number of Manhattan structure constraint residuals in the sliding window, respectively; They represent the i-th IMU pre-integration residual, the m-th lidar edge residual, the n-th lidar plane residual, the a-th visual point feature residual, the b-th visual line feature and the b-th Manhattan structure constraint residual respectively; Φ(ω visual ) and Φ(ω lidar ) represent the penalty functions of vision and lidar weights respectively.
6. The confidence factor-based multimodal robust positioning method for underground UAVs according to claim 5, characterized in that: In the back-end stage, dynamic adjustment of the pose optimization is achieved according to the degradation degree weight of each sensor. Specifically, the asymptotically non-convex factor graph optimization function is solved. The solution process includes inner iteration and outer iteration: In each outer iteration, the control parameter θ is fixed; In each inner iteration, first fix the weight of t-1 To optimize the state variable x at time t t , and then by fixing the state variable x at time t t To optimize the weight at time t Get the residual weights of visual and lidar features at time t By changing the value of θ to repeat the internal and external iteration process, θ is adjusted to increase the non-convexity of the optimization function, and the optimal UAV posture is obtained to achieve positioning.
Citation Information
Patent Citations
Multi-source fusion SLAM system based on visual point-line feature optimization
CN113837277A
Multi-sensor pose estimation method and device considering perceptual degradation
CN117745821A