Multi-Modal SLAM Sensor Fusion With Adaptive Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current SLAM solutions, such as VLOAM, do not effectively utilize LIDAR data jointly with camera data, leading to potential errors due to environmental conditions like poor lighting, and lack redundancy for ensuring functional safety and reliability in robot navigation.
Innovation Solution
The MM-SLAM system dynamically adjusts the weighting factors of LIDAR and camera sensor data based on environmental conditions, using a probability density function and historical data to combine information from both sensors, ensuring reliable localization and mapping even in adverse conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If only cameras or only LIDAR systems are used as sensors independently, then the system complexity is reduced, but the reliability and functional safety are insufficient due to lack of redundancy
Solution Approach 1:
The patent combines LIDAR and camera sensors into a unified multi-modal SLAM system, where both sensors operate simultaneously and their data are integrated through a common optimization framework. This merging provides redundancy and improves reliability while maintaining manageable system complexity through shared processing architecture.
Solution Approach 2:
The patent creates a multi-functional sensor system where both LIDAR and camera sensors serve multiple purposes: LIDAR provides depth information and structural data, while cameras provide texture, color, and semantic information. The system can adaptively switch between or combine these functions based on environmental conditions, improving reliability without proportionally increasing complexity.
2Measurement precision
If visual odometry runs at high frequency (60 Hz) and LIDAR odometry runs at low frequency (1 Hz), then the measurement precision is improved, but error accumulation reaches arbitrarily high values due to independent estimation
Solution Approach 1:
The patent merges visual odometry and LIDAR odometry into a unified optimization framework where both data streams are processed together rather than independently. This joint estimation eliminates error accumulation by consistently fusing measurements from both sensors at their respective frequencies, maintaining high measurement precision while ensuring reliability.
Solution Approach 2:
The patent implements a feedback mechanism where the optimization framework continuously adjusts motion estimates by incorporating new measurements from both visual and LIDAR sensors. The system uses the high-frequency visual data for real-time correction and the low-frequency LIDAR data for long-term consistency, with each sensor type providing feedback to correct errors from the other.
3Adaptability or versatility
If LIDAR and camera data are combined with fixed weighting factors, then the device complexity is reduced, but the adaptability to environmental conditions deteriorates
Solution Approach 1:
The patent implements dynamic weighting factors that automatically adjust based on environmental conditions, sensor quality, and data consistency. The system monitors measurement uncertainty and adaptively changes the relative contribution of LIDAR and camera data, enabling adaptability to varying conditions without requiring complex manual configuration or additional hardware.
Data Source
AI summary
A method for motion tracking is provided including receive first data, receive second data, transform the second data to generate transformed second data corresponding to the first frame; determine a first weighting factor for the first data and a second weighting factor for the transformed second data; weight the first data using the first weighting factor to generate first weighted data; weight the transformed second data using the second weighting factor to generate second weighted data; and combine the weighted first data and the weighted second data to generate combined image data. The first data include a first frame of a first scene of an environment detected by a camera or image sensor. The second data include a second frame of a second scene of an environment detected by a light detection and ranging (LIDAR) sensor. At least a subset of the second scene corresponds to the first scene.


