A SLAM positioning and mapping method suitable for multi-modal information fusion of robots in a hazard scene

By constructing a hazardous scene degradation vector and a modal credibility correction factor map, the problem of sensor degradation in hazardous scenes of multimodal SLAM is solved, and continuous localization and robust mapping in scenarios such as smoke, dust and high temperature are achieved, which improves the robot's localization accuracy and map consistency in hazardous environments.

CN122631155APending Publication Date: 2026-08-25SUZHOU MINGTAI INTELLIGENT EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610759466.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing multimodal SLAM methods cannot perceive sensor degradation in real time in hazardous scenarios, resulting in fixed fusion weights, low-quality observation of pollution pose optimization, inability of the hazard map to participate in loop closure and relocalization, and a lack of multimodal degradation recovery mechanisms.

Method used

By constructing a degradation vector for hazardous scenarios, dynamically adjusting the multimodal SLAM constraint strength, and utilizing the modal credibility correction factor graph covariance matrix, cross-modal correlation is enhanced, which participates in keyframe selection and loop closure detection, and enters the degradation recovery mode for relocation.

Benefits of technology

It achieves continuous localization and robust mapping in scenarios such as smoke, dust, high temperature, and electromagnetic interference, improving the robot's localization accuracy and map consistency in hazardous scenarios, and enhancing the real-time perception of hazard information and the closed-loop reliability of repetitive structural scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122631155A_ABST
    Figure CN122631155A_ABST
Patent Text Reader

Abstract

The application discloses a kind of SLAM positioning and mapping method suitable for robot multimodal information fusion in hazardous scene, it is related to robot simultaneous localization and mapping, multi-sensor fusion, hazardous environment perception and dangerous area mapping technical field.The laser radar point cloud, visible light image, thermal infrared image, inertial measurement, odometry and environmental hazard data are obtained;Degeneration vector containing visible light, point cloud, thermal infrared, mileage and environmental hazard degradation is constructed by synchronization and preprocessing;According to observation quality, normalized residual and degeneration vector, the credibility of mode is calculated, and the covariance of corresponding factor in multimodal factor graph is corrected to optimize pose;Based on the optimized pose, composite map containing geometric occupancy layer, semantic layer and hazard risk layer is constructed, and risk layer is used for key frame insertion, loop detection and degradation recovery repositioning, to improve the reliability of robot positioning and mapping in hazardous scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of robot synchronous localization and mapping, multi-sensor fusion, hazardous environment perception, and dangerous area mapping, and particularly to a SLAM localization and mapping method suitable for robot multimodal information fusion in hazardous scenarios. Background Technology

[0002] Mobile robots are required to perform autonomous localization, environmental mapping, target search, and risk marking in hazardous scenarios such as fire rescue, chemical plant inspection, mine exploration, underground space surveying, nuclear facility inspection, and post-disaster search and rescue. Simultaneous localization and mapping (SLAM) is the fundamental technology for robots to perform these tasks.

[0003] Existing SLAM methods typically employ lidar, visible light cameras, inertial measurement units (IMUs), or odometry for pose estimation. LiDAR SLAM provides relatively stable geometric constraints, visual SLAM utilizes texture and semantic information, IMUs provide short-term motion prediction, and odometry provides continuous relative displacement constraints. To improve robustness, existing technologies also include multimodal SLAM methods such as lidar-inertial fusion, vision-inertial fusion, and lidar-vision-inertial fusion.

[0004] However, in hazardous scenarios, sensor degradation is sudden and multi-source. For example, smoke and dust can reduce the quality of visible light images and may also cause sparse point clouds or abnormal echoes in lidar; high-temperature fire sources, reflective metal surfaces, or strong thermal radiation can cause local saturation of thermal infrared images; slippery surfaces, gravel surfaces, and track idling can increase odometer errors; toxic gas leaks, high-temperature areas, and electromagnetic interference can affect robot safe passage and map reliability.

[0005] Existing multimodal SLAM typically focuses on geometric fusion between sensors, often employing fixed weights, empirical thresholds, or single residual evaluations to process sensor data. These methods suffer from the following problems: they fail to convert inherent hazards such as smoke, gas, high temperatures, and electromagnetic interference into degradation factors that directly impact SLAM optimization; when a sensor mode degrades, fixed fusion weights may continue to introduce low-quality constraints, causing pose drift or map ghosting; hazard information is often displayed as a post-processing result, not participating in keyframe selection, loop closure detection, or relocalization; and when multiple external sensing modes degrade simultaneously, there is a lack of recoverable relocalization mechanisms.

[0006] Therefore, a method is needed that can perceive the sensor degradation status in real time in hazardous scenarios, dynamically adjust the multimodal SLAM constraint strength, and use the hazard risk map in reverse to participate in the localization and mapping process. Summary of the Invention

[0007] Technical problems to be solved

[0008] The technical problem to be solved by this invention is to provide a SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios, so as to solve the problems of existing multimodal SLAM in scenarios such as smoke, dust, low light, high temperature, toxic gas, electromagnetic interference and slippery ground, which are difficult to quantify sensor degradation, fixed fusion weights, low-quality observation pollution pose optimization, and the inability of hazard maps to participate in loop closure and relocalization.

[0009] Technical solution

[0010] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0011] A SLAM localization and mapping method for robots in hazardous scenarios, applicable to multimodal information fusion, includes the following steps:

[0012] S1. Acquire multimodal sensing data collected by the robot in hazardous scenarios. The multimodal sensing data includes lidar point cloud data, visible light image data, thermal infrared image data, inertial measurement data, odometry data, and at least one type of environmental hazard sensor data. The environmental hazard sensor data includes at least one of gas concentration data, ambient temperature data, smoke and dust concentration data, and electromagnetic interference intensity data.

[0013] S2. The multimodal sensing data is synchronized in time, unified in coordinates and preprocessed, and a hazard scene degradation vector is constructed based on the preprocessed multimodal sensing data. The hazard scene degradation vector includes visible light degradation, point cloud degradation, thermal infrared degradation, mileage degradation and environmental hazard.

[0014] Specifically, the visible light degradation is determined based on the average brightness of the visible light image, image entropy, image blur, and the number of effective visual feature points; the point cloud degradation is determined based on the unit voxel point cloud density, laser echo intensity stability, number of effective edge features, and number of effective planar features; the thermal infrared degradation is determined based on the saturated pixel ratio, number of temperature gradients, and number of effective thermal contours of the thermal infrared image; the odometer degradation is determined based on the displacement and angle deviations between the odometer prediction and the inertial measurement prediction; and the environmental hazard is obtained by normalizing at least one of the following: gas concentration data, ambient temperature data, smoke and dust concentration data, and electromagnetic interference intensity data.

[0015] S3. Extract geometric features from lidar point cloud data, visual features from visible light image data, and thermal infrared contour features from thermal infrared image data. Generate short-term motion predictions between adjacent keyframes of the robot based on inertial measurement data and odometry data.

[0016] Specifically, the thermal infrared contour features are obtained by sequentially performing temperature normalization, saturation region removal, temperature gradient calculation, connected component filtering, and contour description on the thermal infrared image; and the thermal infrared contour features are projected onto the lidar point cloud coordinate system or the visible light image coordinate system to establish a cross-modal association between the thermal infrared contour features and the point cloud boundary features or visual edge features.

[0017] S4. Based on the observation quality of each sensing mode, the normalized residual between the sensing mode and the short-term motion prediction, and the degradation vector of the hazardous scene, calculate the mode confidence level corresponding to the lidar mode, visible light mode, thermal infrared mode, and odometry mode.

[0018] Specifically, the modal confidence R_m(t) of any sensing mode m at time t is calculated as follows:

[0019] R_m(t)=clip{Q_m(t)·exp[-αD_m(t)-βE_m(t)],R_min,1};

[0020] Where Q_m(t) is the observation quality of sensing mode m, D_m(t) is the degradation amount corresponding to sensing mode m in the degradation vector of the hazardous scene, E_m(t) is the normalized residual between sensing mode m and the short-time motion prediction, α and β are weight coefficients greater than 0, R_min is the minimum confidence level, and clip represents the clipping function.

[0021] S5. Construct a multimodal factor map, which includes inertial pre-integration factor, odometry factor, lidar matching factor, visual reprojection factor, thermal infrared profile matching factor, and closed-loop constraint factor. Then, modify the covariance matrix of the corresponding factor according to the modal confidence level so that the factor constraint corresponding to the sensing mode with reduced confidence level is weakened.

[0022] Specifically, the covariance matrix of the corresponding factor is corrected based on the modality confidence level, including:

[0023] Σ'_m(t)=Σ_m / [R_m(t)+ε];

[0024] Where Σ_m is the initial covariance matrix of the factor corresponding to sensing mode m, Σ'_m(t) is the corrected covariance matrix, R_m(t) is the modal confidence of sensing mode m, and ε is a positive number to prevent division by zero.

[0025] S6. Optimize the multimodal factor map to obtain the robot pose sequence, and fuse the lidar point cloud data, visible light image data, thermal infrared image data and environmental hazard sensor data according to the robot pose sequence to construct a composite map, which includes a geometric occupancy layer, a semantic layer and a hazard risk layer.

[0026] Specifically, the hazard risk layer stores risk values ​​and risk confidence in the form of a two-dimensional grid or a three-dimensional voxel. The risk value of any grid or voxel is obtained by weighting at least three of the following: temperature anomaly value, gas concentration value, smoke and dust concentration value, electromagnetic interference intensity, accessibility risk, and dynamic obstacle density. After optimization of the multimodal factor map, the risk value is reprojected and updated based on the optimized robot pose.

[0027] S7. Determine whether to insert a keyframe based on the risk gradient of the hazard risk layer and the robot pose uncertainty, and generate a cross-modal closed-loop descriptor based on the lidar scanning descriptor, visual descriptor, thermal infrared contour descriptor and hazard risk distribution descriptor for closed-loop detection.

[0028] Specifically, the conditions for inserting keyframes include at least two of the following: robot displacement exceeding a displacement threshold, robot posture change exceeding an angle threshold, robot pose covariance exceeding an uncertainty threshold, risk gradient of the hazard risk layer exceeding a risk change threshold, and a switch in the dominant localization mode. The cross-modal closed-loop descriptor includes a lidar scanning context descriptor, a visual bag-of-words descriptor, a thermal infrared contour orientation histogram, and a hazard risk distribution histogram. When a candidate historical keyframe simultaneously satisfies the cross-modal closed-loop descriptor similarity threshold, geometric consistency threshold, and risk distribution consistency threshold, a closed-loop constraint factor is added to the multimodal factor graph.

[0029] S8. When the confidence level of at least two of the lidar mode, visible light mode, and thermal infrared mode is lower than the preset confidence level threshold, the degradation recovery mode is entered. In the degradation recovery mode, short-term pose estimation is performed based on inertial measurement data and odometry data, and when the confidence level of any external sensing mode recovers to above the preset confidence level threshold, relocation is performed using the cross-modal closed-loop descriptor.

[0030] Specifically, when the external perception modality confidence level is lower than a preset confidence level threshold, the robot's current pose covariance is increased, and high uncertainty regions are marked in the composite map. After the external perception modality is recovered, historical keyframes are retrieved using cross-modal closed-loop descriptors, and relocation constraint factors are generated after geometric consistency verification.

[0031] The present invention also provides a robot SLAM localization and mapping system, including a robot body, a lidar, a visible light camera, a thermal infrared camera, an inertial measurement unit, an odometry, at least one environmental hazard sensor, a processor, and a memory. The memory stores a program that can be executed by the processor. When the program is executed by the processor, it implements the above-mentioned SLAM localization and mapping method applicable to robot multimodal information fusion in hazardous scenarios.

[0032] Beneficial effects

[0033] Compared with existing technologies, this invention provides a SLAM localization and mapping method for robots in hazardous scenarios, which integrates multimodal information and has the following advantages:

[0034] 1. This invention transforms environmental hazards into degradation quantities in the SLAM optimization process. It not only records data on gases, temperatures, smoke, or electromagnetic interference, but also combines these data with images, point clouds, thermal infrared data, and odometer degradation to form a degradation vector for hazardous scenes, thereby providing a basis for adjusting the weights of multimodal fusion.

[0035] 2. This invention automatically weakens the constraints of degenerate modes by adjusting the covariance matrix of the corresponding factors in the modality confidence correction factor graph, thus avoiding low-quality observation contamination of pose optimization caused by fixed-weight fusion.

[0036] 3. This invention utilizes thermal infrared contour features and point cloud boundary features or visual edge features to establish cross-modal associations, which can enhance continuous positioning capabilities in smoke, darkness and high temperature scenes.

[0037] 4. The hazard risk layer of this invention is not only used to display dangerous areas, but also participates in keyframe insertion, cross-modal closed-loop detection, and degradation recovery relocation, improving the closed-loop reliability and risk level in repetitive structure scenes. Figure 1 To the point of being responsive.

[0038] 5. This invention enters a degradation recovery mode when multiple external sensing modes degrade, uses inertial measurement data and odometry data to perform short-term pose estimation, and performs relocalization through a cross-modal closed-loop descriptor after external sensing recovery, thereby improving the continuity and robustness of robot localization in hazardous scenarios. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the overall process of the method of the present invention;

[0040] Figure 2 This is a schematic diagram of the system composition and multimodal sensor arrangement of the present invention;

[0041] Figure 3 This is a schematic diagram of the process for constructing the degradation vector for hazardous scenarios in this invention;

[0042] Figure 4 A schematic diagram illustrating the process of establishing cross-modal associations between thermal infrared contour features and point cloud boundary features or visual edge features in this invention;

[0043] Figure 5 This is a schematic diagram illustrating the modal reliability calculation and factor graph covariance correction relationship of the present invention;

[0044] Figure 6 This is a schematic diagram of the multimodal factor graph structure of the present invention;

[0045] Figure 7 This is a schematic diagram of the composite map structure of the present invention;

[0046] Figure 8 This is a schematic diagram of the cross-modal closed-loop detection process of the present invention;

[0047] Figure 9 This is a schematic diagram of the degradation recovery and relocation process of the present invention. Detailed Implementation

[0048] The present invention will be further described below with reference to the accompanying drawings and embodiments. It should be understood that the following detailed description is only for explaining the present invention and is not intended to limit the scope of protection of the present invention.

[0049] Example 1: System Composition and Data Acquisition

[0050] Reference Appendix Figure 1 and attached Figure 2 In one embodiment, the robot is a tracked fire reconnaissance robot. The robot body is equipped with a 3D lidar, a visible light camera, a thermal infrared camera, an inertial measurement unit, a track odometer, a temperature sensor, a carbon monoxide sensor, and a smoke concentration sensor.

[0051] 3D LiDAR is used to acquire point cloud data of building interior walls, passageways, door frames, stairs, and obstacles. Visible light cameras are used to acquire RGB or grayscale images. Thermal infrared cameras are used to acquire thermal infrared images of heat sources, humans, and high-temperature objects. Inertial measurement units are used to acquire angular velocity and linear acceleration. Tracked odometers are used to estimate the robot's travel distance. Temperature sensors, carbon monoxide sensors, and smoke concentration sensors are used to acquire environmental hazard information.

[0052] Reference Appendix Figure 2 Each sensor synchronizes its time using a unified clock or software timestamp. Using the robot's body coordinate system B as a reference, the extrinsic parameters of the LiDAR coordinate system L, visible light camera coordinate system C, thermal infrared camera coordinate system T, and inertial measurement unit coordinate system I relative to the robot's body coordinate system B are obtained through offline calibration. While the robot moves in hazardous environments, it continuously collects multimodal sensor data and inputs the collected data into the processor for subsequent SLAM localization and mapping processing.

[0053] Example 2: Time Synchronization, Coordinate Unification, and Preprocessing

[0054] Reference Appendix Figure 1 and attached Figure 3 The processor first performs time synchronization, coordinate unification, and preprocessing on the multimodal sensing data.

[0055] For sensor data with different sampling frequencies, nearest neighbor time matching, linear interpolation, or buffer queuing methods can be used to obtain data frames at a unified time. For LiDAR point cloud data, outlier removal, voxel filtering, and motion distortion compensation are performed. For visible light image data, distortion correction, brightness normalization, and noise reduction are performed. For thermal infrared image data, temperature normalization, bad pixel repair, and non-uniformity correction are performed. For inertial measurement data, zero-bias compensation is performed. For odometer data, abnormal jump detection is performed. Through the above processing, data from different modalities can participate in subsequent fusion calculations under a unified time and coordinate reference.

[0056] Example 3: Construction of Degradation Vectors for Hazardous Scenarios

[0057] Reference Appendix Figure 3 Let the current time be t, and the degradation vector of the hazardous scene be:

[0058] D(t)=[d_v(t),d_l(t),d_ir(t),d_o(t),d_h(t)];

[0059] Wherein, d_v(t) is the visible light degradation, d_l(t) is the point cloud degradation, d_ir(t) is the thermal infrared degradation, d_o(t) is the odometer degradation, and d_h(t) is the environmental hazard.

[0060] The observation quality Q_v(t) of a visible light image can be obtained by the following formula:

[0061] Q_v(t)=clip{a1·N_v(t) / N_v0+a2·H_v(t) / H_v0+a3·B_v(t) / B_v0,0,1};

[0062] Where N_v(t) is the number of effective visual feature points in the current frame, N_v0 is the reference value for the number of effective visual feature points under normal conditions, H_v(t) is the current image entropy, H_v0 is the image entropy reference value, B_v(t) is the image sharpness, B_v0 is the sharpness reference value, and a1, a2, and a3 are weights, with a1+a2+a3=1. The visible light degradation is d_v(t)=1-Q_v(t). When the robot enters areas of dense smoke, darkness, or flashing bright light, the number of effective visual feature points decreases, the image entropy decreases, or the image blurriness increases, resulting in an increase in the visible light degradation d_v(t).

[0063] The point cloud observation quality Q_l(t) can be calculated from the point cloud density per unit voxel, the number of effective edge features, the number of effective planar features, and the stability of echo intensity:

[0064] Q_l(t)=clip{b1·ρ_l(t) / ρ_l0+b2·N_e(t) / N_e0+b3·N_p(t) / N_p0+b4·I_l(t) / I_l0,0,1};

[0065] Where ρ_l(t) is the point cloud density per unit voxel, N_e(t) is the number of effective edge features, N_p(t) is the number of effective planar features, I_l(t) is the echo intensity stability, ρ_l0, N_e0, N_p0, and I_l0 are reference values, and b1 to b4 are weights that sum to 1. The point cloud degradation is d_l(t) = 1 - Q_l(t).

[0066] The thermal infrared observation quality Q_ir(t) can be determined by the number of effective thermal profiles, the number of temperature gradients, and the proportion of saturated pixels: Q_ir(t) = clip{c1·N_ir(t) / N_ir0+c2·G_ir(t) / G_ir0+c3·[1-S_ir(t)],0,1};

[0067] Where N_ir(t) is the number of effective thermal contours, G_ir(t) is the number of temperature gradients, S_ir(t) is the saturation pixel ratio, N_ir0 and G_ir0 are reference values, and c1 to c3 are weights and their sum is 1.

[0068] The thermal infrared degradation is d_ir(t) = 1 - Q_ir(t). When a fire source, strong thermal radiation, or high-temperature metal causes large-area saturation in the thermal infrared image, S_ir(t) increases, and the thermal infrared degradation d_ir(t) increases accordingly.

[0069] The motion prediction ΔT_imu(t) for adjacent time moments is obtained through pre-integration of inertial measurement data, and the motion prediction ΔT_o(t) is obtained through odometry. The displacement deviation between the two is δp(t), and the angular deviation is δθ(t). The odometer degradation is:

[0070] d_o(t)=clip{k1·δp(t) / δp0+k2·δθ(t) / δθ0,0,1};

[0071] Where δp0 and δθ0 are deviation reference thresholds, and k1+k2=1. When the tracks or wheels slip, spin freely, or the robot is impacted, the deviation between the odometer prediction and the inertial measurement prediction increases, and d_o(t) increases.

[0072] The environmental hazard quantity d_h(t) can be calculated using the following formula:

[0073] d_h(t)=clip{r1·G(t) / G_max+r2·T(t) / T_max+r3·S(t) / S_max+r4·M(t) / M_max,0,1};

[0074] Wherein, G(t) is the gas concentration, T(t) is the temperature anomaly, S(t) is the smoke and dust concentration, M(t) is the electromagnetic interference intensity, G_max, T_max, S_max and M_max are the corresponding normalized reference values, and r1 to r4 are the weights.

[0075] Example 4: Multimodal Feature Extraction and Cross-modal Association

[0076] Reference Appendix Figure 1 and attached Figure 4 Voxel filtering, outlier removal, and motion distortion compensation are performed on lidar point cloud data.

[0077] The process involves compensation, followed by calculation of point cloud curvature, extraction of point cloud edge features, point cloud planar features, and local geometric descriptors. Edge features are used for point-to-line matching, and planar features are used for point-to-surface matching.

[0078] Reference Appendix Figure 4 The visible light image data undergoes distortion correction, brightness normalization, and denoising, and then corner points, edges, and local visual descriptors are extracted. Optionally, semantic targets such as doors, walls, stairs, pipelines, people, equipment, or vehicles are identified using a semantic recognition model.

[0079] Reference Appendix Figure 4 The thermal infrared image data is temperature normalized to map the original temperature values ​​to a uniform scale; regions exceeding the saturation threshold are masked and removed; then, a temperature gradient map is calculated, and connected component filtering is performed on continuous gradient regions to retain thermal infrared contours with areas greater than a set threshold and stable shapes; finally, the orientation histogram, area, principal direction, and center position of the thermal infrared contours are calculated to obtain the thermal infrared contour descriptor.

[0080] Reference Appendix Figure 4 By utilizing the extrinsic parameters between the thermal infrared camera and the lidar / visible light camera, the thermal infrared contour is projected onto the lidar point cloud coordinate system or the visible light image coordinate system. When the thermal infrared contour boundary and the point cloud boundary or visual edge meet preset thresholds in projection position, direction, and scale, a cross-modal association is established between the thermal infrared contour features and the point cloud boundary features or visual edge features.

[0081] Example 5: Short-time motion prediction and modal reliability calculation

[0082] Reference Appendix Figure 5The inertial short-time motion prediction between adjacent keyframes of the robot is obtained by pre-integrating inertial measurement data, and the mileage short-time motion prediction between adjacent keyframes is obtained by odometry data. The two can be weighted and fused to obtain the short-time motion prediction, or prediction constraints can be established separately based on the inertial measurement data and the odometry data.

[0083] Reference Appendix Figure 5 The modal reliability is calculated for lidar mode, visible light mode, thermal infrared mode, and odometer mode. Taking lidar mode as an example, the lidar observation quality is Q_l(t), the corresponding degradation is d_l(t), and the normalized residual between the relative motion obtained from lidar point cloud matching and the short-time motion prediction is E_l(t). Then, the lidar mode reliability is:

[0084] R_l(t)=clip{Q_l(t)·exp[-αd_l(t)-βE_l(t)],R_min,1};

[0085] R_v(t)=clip{Q_v(t)·exp[-αd_v(t)-βE_v(t)],R_min,1}; R_ir(t)=clip{Q_ir(t)·exp[-α d_ir(t)-βE_ir(t)],R_min,1}; R_o(t)=clip{Q_o(t)·exp[-αd_o(t)-βE_o(t)],R_min,1};

[0086] Wherein, R_min can be taken from 0.05 to 0.20, α from 0.5 to 3.0, and β from 0.5 to 3.0. These parameters can be adjusted according to the robot type, sensor performance, and application scenario. When visible light images degrade in dense smoke or dark environments, R_v(t) decreases; when the lidar has sparse point clouds in dusty environments, R_l(t) decreases; when thermal infrared images are saturated over a large area due to high-temperature fire sources, R_ir(t) decreases; and when the robot slips, R_o(t) decreases.

[0087] Example 6: Construction of Multimodal Factor Plots and Covariance Correction

[0088] Reference Appendix Figure 5 and attached Figure 6 The state of the robot in the i-th keyframe can be represented as:

[0089] X_i={R_i,p_i,v_i,b_gi,b_ai};

[0090] Where R_i is attitude, p_i is position, v_i is velocity, b_gi is gyroscope bias, and b_ai is accelerometer bias.

[0091] Reference Appendix Figure 6 The multimodal factor map includes inertial pre-integration factor, odometry factor, lidar matching factor, visual reprojection factor, thermal infrared profile matching factor, and closed-loop constraint factor. The optimization objective function of the multimodal factor map can be expressed as:

[0092]

[0093] ²_Σ'ir)+Σρ(||e_loop||²_Σloop);

[0094] Where e_imu is the inertial pre-integration residual, e_o is the odometer residual, e_l is the lidar matching residual, e_v is the visual reprojection residual, e_ir is the thermal infrared profile matching residual, e_loop is the closed-loop residual, and ρ is the robust kernel function.

[0095] Reference Appendix Figure 5 For any sensing mode m, the covariance matrix of the corresponding factor is corrected as follows:

[0096] Σ'_m(t)=Σ_m / [R_m(t)+ε];

[0097] When the confidence level of a certain mode decreases, the corresponding covariance matrix increases, weakening the constraint effect of the mode residual in factor graph optimization. Conversely, when the confidence level of a certain mode increases, the corresponding covariance matrix decreases or remains at a normal level, strengthening or stabilizing the constraint effect of that mode on pose optimization. Through this approach, the system can automatically adjust the influence of different sensing modes in scenarios involving smoke, dust, low light, high temperature, electromagnetic interference, or slippage, reducing the contamination of robot pose optimization by degraded observations.

[0098] Example 7: Composite Map Construction

[0099] Reference Appendix Figure 7 A composite map is constructed by fusing LiDAR point cloud data, visible light image data, thermal infrared image data, and environmental hazard sensor data based on the optimized robot pose sequence. The composite map includes a geometric occupancy layer, a semantic layer, and a hazard risk layer.

[0100] The geometric occupancy layer can be implemented using a 2D raster map, a 3D voxel map, or an octree map to represent obstacles, free space, and unknown space. The semantic layer stores semantic target categories, target locations, and target confidence levels. Semantic targets include doors, walls, stairs, passageways, pipelines, people, equipment, and vehicles.

[0101] The hazard risk layer stores the risk value H(v) and risk confidence level C(v) in the form of a two-dimensional raster or a three-dimensional voxel. For any raster or voxel v, the risk value H(v) can be calculated using the following formula:

[0102] H(v)=w1T(v)+w2G(v)+w3S(v)+w4M(v)+w5P(v)+w6D(v);

[0103] Where T(v) is the temperature anomaly, G(v) is the gas concentration, S(v) is the smoke and dust concentration, M(v) is the electromagnetic interference intensity, P(v) is the passability risk, D(v) is the dynamic obstacle density, and w1 to w6 are weights. (See attached reference.) Figure 7 Whenever the multimodal factor map is optimized and the robot's historical pose is updated, the historical hazard observations are reprojected onto the hazard risk layer according to the optimized robot pose, so as to reduce the risk position deviation caused by early positioning drift.

[0104] For example, in a firefighting scenario, if the robot's positioning is slightly drifted when it first enters the corridor, and the trajectory is subsequently corrected through closed-loop optimization, then areas with abnormal temperatures, areas with excessive toxic gases, and areas obscured by smoke and dust will be remapped onto the composite map based on the corrected pose, thereby improving the consistency between the hazard risk layer and the geometric occupancy layer.

[0105] Example 8: Risk Perception Keyframe Insertion

[0106] Reference Appendix Figure 1 Appendix Figure 7 and attached Figure 8 This invention introduces a hazard risk layer and a dominant localization mode switching condition on the basis of traditional SLAM keyframe insertion conditions.

[0107] A keyframe is inserted when at least two of the following conditions are met: the robot displacement exceeds the displacement threshold; the robot posture change exceeds the angle threshold; the robot pose covariance exceeds the uncertainty threshold; the risk gradient of the hazard risk layer exceeds the risk change threshold; or the dominant localization mode switches.

[0108] Reference Appendix Figure 7 When the robot passes through gas leak boundaries, high-temperature zone boundaries, or dense smoke boundaries, even with small robot displacement, keyframes can be inserted due to significant changes in the risk gradient of the hazard risk layer, thereby improving the mapping accuracy of hazardous area boundaries. (See attached reference.) Figure 8 When the dominant positioning mode switches from visible light mode to lidar mode or thermal infrared mode, the system inserts key frames to record the observation status before and after the mode switch, reducing the uncertainty of subsequent closed-loop detection and map optimization.

[0109] Example 9: Cross-modal closed-loop detection

[0110] Reference Appendix Figure 8 Each keyframe generates a cross-modal closed-loop descriptor Ψ_k:

[0111] Ψ_k=[S_l(k),S_v(k),S_ir(k),S_h(k)];

[0112] Wherein, S_l(k) is the lidar scanning context descriptor, S_v(k) is the visual bag-of-words descriptor, S_ir(k) is the thermal infrared profile orientation histogram, and S_h(k) is the hazard risk distribution histogram.

[0113] Reference Appendix Figure 8 During loop closure detection, candidate historical keyframes are first retrieved based on cross-modal loop closure descriptors. Then, geometric consistency verification, thermal infrared consistency verification, and risk distribution consistency verification are performed on these candidate historical keyframes. Geometric consistency verification compares the point cloud registration residuals between the current keyframe and the candidate historical keyframes; thermal infrared consistency verification compares the principal direction, area ratio, and spatial distribution of the thermal infrared contours; and risk distribution consistency verification compares the hazard risk distribution histogram and the topological relationship of high-risk areas. When a candidate historical keyframe simultaneously meets the cross-modal loop closure descriptor similarity threshold, geometric consistency threshold, and risk distribution consistency threshold, a loop closure constraint factor is added to the multimodal factor graph. Through this method, the risk of erroneous loop closures in scenarios such as repeated corridors, repeated pipe galleries, and mine roadways can be reduced, and the reliability of loop closure judgment can be improved by utilizing hazard risk distribution.

[0114] Example 10: Degradation Recovery Relocation

[0115] Reference Appendix Figure 9 When the confidence level of at least two of the lidar mode, visible light mode, and thermal infrared mode falls below a preset confidence level threshold, the system enters a degradation recovery mode. In this invention, the external sensing mode refers to the lidar mode, visible light mode, and thermal infrared mode; the dominant positioning mode refers to the external sensing mode with the highest confidence level at the current moment and the greatest contribution to pose optimization constraints.

[0116] Reference Appendix Figure 9 In degradation recovery mode, the robot performs short-term pose estimation based on inertial measurement and odometry data, and improves the current pose covariance. Simultaneously, this area is marked as a high-uncertainty region in the composite map. At this point, the system can either use low-confidence external observations as weak constraints in the optimization process, or temporarily store low-confidence external observations for later verification.

[0117] Reference Appendix Figure 9When the modal credibility of any external sensing modality recovers to above a preset credibility threshold, the system retrieves historical keyframes using cross-modal closed-loop descriptors. If a candidate historical keyframe passes geometric consistency verification, a relocation constraint factor is generated and added to a multimodal factor graph for local or global optimization. After optimization, the system simultaneously corrects the robot's current pose, historical trajectory, and composite map.

[0118] For example, when a robot traverses a dense smoke area, the visible light image and the laser point cloud may degrade simultaneously, and the system enters a degradation recovery mode; after the robot leaves the dense smoke area, the laser radar point cloud is restored, and the system retrieves historical key frames through the laser radar scanning descriptor, thermal infrared contour descriptor, and hazard risk distribution descriptor, and completes relocation.

[0119] Example 11: Parameter Setting Example

[0120] In one specific implementation, the following parameter ranges can be used: the reference value for the number of effective visible light feature points N_v0 is 100 to 300; the reference value for the unit voxel point cloud density ρ_l0 is 500 to 3000 points per cubic meter; when the thermal infrared saturated pixel ratio exceeds 0.25, the thermal infrared degradation increases significantly; when the displacement deviation predicted by the odometer and inertial measurement exceeds 0.10 meters to 0.30 meters, the odometer degradation increases; when the angular deviation predicted by the odometer and inertial measurement exceeds 3 degrees to 8 degrees, the odometer degradation increases.

[0121] The preset confidence threshold can be set to 0.25 to 0.40; R_min can be set to 0.05 to 0.20; ε can be set to 10^-6 to 10^-3; the keyframe displacement threshold can be set to 0.3 meters to 1.0 meter; and the keyframe angle threshold can be set to 5 degrees to 15 degrees. These parameters are merely examples and do not constitute a limitation on the scope of protection of this invention. Those skilled in the art can adjust them according to the robot platform, sensor performance, and actual application scenarios.

[0122] This invention can be applied to fire rescue robots, chemical inspection robots, mine exploration robots, underground space inspection robots, nuclear facility inspection robots, post-disaster search and rescue robots, and unmanned emergency reconnaissance platforms. Based on existing available sensors and computing platforms, this method can output continuous pose and composite maps containing hazardous area information under hazardous conditions such as smoke, dust, low light, high temperature, toxic gases, electromagnetic interference, and slippery surfaces, demonstrating industrial applicability.

Claims

1. A SLAM localization and mapping method for robots in hazardous scenarios, characterized in that, The steps include the following: S1. Acquire multimodal sensing data collected by the robot in hazardous scenarios. The multimodal sensing data includes lidar point cloud data, visible light image data, thermal infrared image data, inertial measurement data, odometry data, and at least one environmental hazard sensor data. The environmental hazard sensor data includes at least one of gas concentration data, ambient temperature data, smoke and dust concentration data, and electromagnetic interference intensity data. S2. The multimodal sensing data is synchronized in time, unified in coordinates and preprocessed, and a hazard scene degradation vector is constructed based on the preprocessed multimodal sensing data. The hazard scene degradation vector includes visible light degradation, point cloud degradation, thermal infrared degradation, mileage degradation and environmental hazard. S3. Extract geometric features from lidar point cloud data, visual features from visible light image data, and thermal infrared contour features from thermal infrared image data. Generate short-term motion predictions between adjacent keyframes of the robot based on inertial measurement data and odometry data. S4. Based on the observation quality of each sensing mode, the normalized residual between the sensing mode and the short-term motion prediction, and the degradation vector of the hazardous scene, calculate the mode confidence level corresponding to the lidar mode, visible light mode, thermal infrared mode and odometry mode respectively. S5. Construct a multimodal factor map, which includes an inertial pre-integration factor, an odometry factor, a lidar matching factor, a visual reprojection factor, a thermal infrared profile matching factor, and a closed-loop constraint factor. Then, modify the covariance matrix of the corresponding factor according to the modal confidence level so that the factor constraint corresponding to the sensing mode with reduced confidence level is weakened. S6. Optimize the multimodal factor map to obtain the robot pose sequence, and fuse the lidar point cloud data, visible light image data, thermal infrared image data and environmental hazard sensor data according to the robot pose sequence to construct a composite map, which includes a geometric occupancy layer, a semantic layer and a hazard risk layer. S7. Determine whether to insert a key frame based on the risk gradient of the hazard risk layer and the robot pose uncertainty, and generate a cross-modal closed-loop descriptor based on the lidar scanning descriptor, visual descriptor, thermal infrared contour descriptor and hazard risk distribution descriptor for closed-loop detection. S8. When the confidence level of at least two of the lidar mode, visible light mode, and thermal infrared mode is lower than the preset confidence level threshold, the degradation recovery mode is entered. In the degradation recovery mode, short-term pose estimation is performed based on inertial measurement data and odometry data, and when the confidence level of any external sensing mode recovers to above the preset confidence level threshold, relocation is performed using the cross-modal closed-loop descriptor.

2. The SLAM localization and mapping method for robots in hazardous scenarios based on multimodal information fusion as described in claim 1, characterized in that, In step S2, the visible light degradation amount is determined based on the average brightness of the visible light image, image entropy, image blur, and the number of effective visual feature points; The point cloud degradation amount is determined based on the point cloud density per unit voxel, the stability of laser echo intensity, the number of effective edge features, and the number of effective planar features; The thermal infrared degradation amount is determined based on the saturation pixel ratio, the number of temperature gradients, and the number of effective thermal contours in the thermal infrared image. The mileage degradation is determined based on the displacement and angular deviations between the odometer prediction and the inertial measurement prediction. The environmental hazard level is obtained by normalizing at least one of the following: gas concentration data, ambient temperature data, smoke and dust concentration data, and electromagnetic interference intensity data.

3. The SLAM localization and mapping method for robots in hazardous scenarios based on multimodal information fusion as described in claim 1, characterized in that, In step S4, the modal confidence R_m(t) of any sensing mode m at time t is calculated as follows: R_m(t)=clip{Q_m(t)·exp[-αD_m(t)-βE_m(t)],R_min,1}; Where Q_m(t) is the observation quality of sensing mode m, D_m(t) is the degradation amount corresponding to sensing mode m in the degradation vector of the hazardous scene, E_m(t) is the normalized residual between sensing mode m and the short-time motion prediction, α and β are weight coefficients greater than 0, R_min is the minimum confidence level, and clip represents the clipping function.

4. The SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios according to claim 1, characterized in that, In step S5, correcting the covariance matrix of the corresponding factor based on the modal confidence level includes: Σ'_m(t)=Σ_m / [R_m(t)+ε]; Where Σ_m is the initial covariance matrix of the factor corresponding to sensing mode m, Σ'_m(t) is the corrected covariance matrix, R_m(t) is the modal confidence of sensing mode m, and ε is a positive number to prevent division by zero.

5. The SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios according to claim 1, characterized in that, In step S3, the thermal infrared contour features are obtained by sequentially performing temperature normalization, saturation region removal, temperature gradient calculation, connected component filtering, and contour description on the thermal infrared image; and the thermal infrared contour features are projected onto the lidar point cloud coordinate system or the visible light image coordinate system to establish a cross-modal association between the thermal infrared contour features and the point cloud boundary features or visual edge features.

6. The SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios according to claim 1, characterized in that, In step S6, the hazard risk layer stores risk values ​​and risk confidence in the form of a two-dimensional grid or a three-dimensional voxel. The risk value of any grid or voxel is obtained by weighting at least three of the following: temperature anomaly value, gas concentration value, smoke and dust concentration value, electromagnetic interference intensity, accessibility risk and dynamic obstacle density. After optimization of the multimodal factor map, the risk value is reprojected and updated according to the optimized robot pose.

7. The SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios according to claim 1, characterized in that, In step S7, the conditions for inserting a keyframe include at least two of the following: robot displacement exceeds a displacement threshold, robot posture change exceeds an angle threshold, robot pose covariance exceeds an uncertainty threshold, risk gradient of the hazard risk layer exceeds a risk change threshold, and the dominant localization mode switches.

8. The SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios according to claim 1, characterized in that, In step S7, the cross-modal closed-loop descriptor includes a lidar scanning context descriptor, a visual bag-of-words descriptor, a thermal infrared profile orientation histogram, and a hazard risk distribution histogram; when a candidate historical keyframe simultaneously satisfies the cross-modal closed-loop descriptor similarity threshold, geometric consistency threshold, and risk distribution consistency threshold, a closed-loop constraint factor is added to the multimodal factor graph.

9. The SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios according to claim 1, characterized in that, In step S8, the degradation recovery mode includes: increasing the robot's current pose covariance and marking high uncertainty areas in the composite map when the external perception modality confidence is lower than a preset confidence threshold; after the external perception modality is recovered, re-localization constraint factors are generated by retrieving historical keyframes using cross-modal closed-loop descriptors and after geometric consistency verification.

10. A robot SLAM localization and mapping system, characterized in that, The system includes a robot body, a lidar, a visible light camera, a thermal infrared camera, an inertial measurement unit, an odometry system, at least one environmental hazard sensor, a processor, and a memory. The memory stores a program that can be executed by the processor. When the program is executed by the processor, it implements the SLAM localization and mapping method for robot multimodal information fusion in hazardous scenarios as described in any one of claims 1 to 9.