Automatic driving environment intelligent sensing and positioning method based on multi-modal data fusion
By using a multimodal data fusion method, the characteristics of vehicle motion state are quantified, a sensor reliability model is constructed, and dynamic weight adjustments are made. This solves the perception delay and accuracy problems of autonomous driving systems in complex environments, achieving high-precision, low-latency environmental perception and improving the safety and reliability of autonomous driving.
Patent Information
- Application Number
- CN202511639332.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-01-16
AI Technical Summary
Existing autonomous driving systems struggle to achieve high-precision, low-latency environmental perception in complex traffic environments, especially in high-speed driving or complex curve scenarios. Traditional methods fail to adaptively adjust sensor weights based on the vehicle's motion state, resulting in lag in path planning and vehicle control, which affects safety and reliability.
A multimodal data fusion method is adopted to receive GNSS, IMU and visual sensor data, quantify the vehicle motion state characteristics, construct a dynamic reliability model of the sensors, perform multi-level fault detection and hierarchical fault isolation, dynamically update sensor weights, realize weight smoothing constraints and normalization, and perform weighted fusion to solve the vehicle state vector.
It achieves high-precision, low-latency environmental perception in high-speed, high-curvature, and complex road environments, improving the safety and reliability of autonomous vehicles and providing a stable and robust foundation for environmental perception.
Smart Images

Figure CN121346782A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of navigation and positioning for unmanned systems, and in particular to an intelligent perception and positioning method for autonomous driving environments using multimodal data fusion. Background Technology
[0002] With the rapid development of autonomous driving technology, the demand for high-precision positioning and environmental perception in complex traffic environments is becoming increasingly urgent. Autonomous driving systems must maintain stable, safe, and real-time driving capabilities under different road conditions, which relies on the vehicle's accurate perception of its own state and the surrounding environment. Multi-source sensor collaboration has become a core approach: GNSS provides global position reference, IMU captures high-frequency dynamic information, and cameras acquire lane line and surrounding obstacle data. Traditional systems often struggle to balance low latency and high accuracy, especially in high-speed driving or complex curve scenarios, where perception latency can lead to lags in path planning and vehicle control, affecting safety and reliability. Therefore, there is an urgent need for a multi-source fusion method that can simultaneously utilize the fast response of IMU, visual accuracy, and global constraints of GNSS to provide autonomous vehicles with stable, continuous, and low-latency environmental perception capabilities.
[0003] Current autonomous driving environmental perception methods primarily rely on fixed-weight sensor fusion or simple filtering algorithms. While visual and GNSS information can guarantee high accuracy in straight-line driving or low-speed scenarios, camera processing suffers computational delays during sharp turns, rapid acceleration / deceleration, or complex road curvature changes. GNSS is significantly affected by occlusion or multipath effects, and although IMUs respond quickly, long-term integration suffers from drift, leading to unreliable short-term predictions. Traditional methods fail to adaptively adjust sensor weights based on vehicle motion, thus failing to fully leverage the low-latency advantage of IMUs in dynamic scenarios and neglecting to utilize lane line information for joint correction of GNSS and IMU data. These issues limit the accurate positioning, lane keeping, and collision avoidance capabilities of autonomous vehicles in high-speed dynamic environments, necessitating improvements to existing fusion strategies to achieve high-precision, low-latency environmental perception in dynamic scenarios. Summary of the Invention
[0004] To address the aforementioned technical issues, this invention provides a multimodal data fusion-based intelligent perception and localization method for autonomous driving environments. This method significantly improves the safety and reliability of autonomous vehicles in high-speed, high-curvature, and complex road environments, providing a stable, robust, and high-precision environmental perception foundation for intelligent driving systems.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows: A multimodal data fusion method for intelligent perception and localization in autonomous driving environments includes the following steps: Step S1: Receive data from GNSS, IMU and vision sensors mounted on the autonomous vehicle, and establish measurement models for each sensor respectively; Step S2: Based on the measurement model, quantify the vehicle motion state characteristics, which include at least steering angle characteristics, longitudinal acceleration characteristics, lateral acceleration characteristics, and angular velocity characteristics; Step S3: Based on the measurement model in Step S1 and the feature quantization results in Step S2, construct the dynamic reliability model of each sensor, perform multi-level fault detection on the sensors, perform graded fault isolation according to the test results, and select the retained sensor data. Step S4: Based on the retained sensor data, determine the vehicle driving scenario; set the optimization target and initialize the weight values of each sensor; dynamically update the sensor weights according to the vehicle driving scenario, and perform weight smoothing constraints and normalization; solve the vehicle state vector through weighted fusion and output the vehicle positioning information.
[0006] In the above scheme, step 2, quantifying the vehicle motion state characteristics, includes: The specific method for quantifying the steering angle feature is as follows: ; in, for Quantized value of steering angle feature at any time; for The steering angle of the vehicle's front wheels at any given moment is positive when turning left; This refers to the sensor sampling time step. The length of the time window; , These are the steering angle feature weight coefficients, all of which are positive and satisfy the following conditions: ; The specific method for quantizing the longitudinal acceleration feature is as follows: ; in, for Quantized value of longitudinal acceleration feature at any given time; for Real-time longitudinal acceleration of the vehicle; for The actual longitudinal acceleration at any given moment; For longitudinal acceleration threshold, when The system was judged to be due to road surface interference rather than active acceleration / deceleration. , These are the longitudinal acceleration characteristic coefficients, all of which are positive numbers. ; The specific method for quantifying the lateral acceleration features is as follows: ; in, for Quantized value of lateral acceleration characteristic at any given time; for Real-time longitudinal speed of the vehicle; for Lane curvature at any time, straight road time When turning The absolute value increases with the steepness of the turn; The specific method for quantifying the vehicle angular velocity is as follows: ; in, for Vehicle angular velocity characteristics at any given moment; , , They are respectively The actual angular velocities of the vehicle along its x, y, and z axes at any given moment; , , They are respectively The true angular velocities of the vehicle along its x, y, and z axes at any given time.
[0007] In the above scheme, step 3 involves constructing the dynamic reliability model for each sensor as follows: The specific formula for the GNSS reliability model is as follows: ; in, for The reliability of GNSS at any given time; for GNSS horizontal accuracy factor at any time; , These are the minimum and maximum thresholds for the GNSS horizontal accuracy factor, respectively. DOP impact coefficient; Number of frames; for Time before Number of valid intra-frame GNSS observations; Indicates the signal continuity factor; The specific method for IMU reliability modeling is as follows: ; in, for The reliability of the IMU at any given moment; for Time before IMU Average drift residual of the frame; This represents the maximum permissible drift residual for the IMU. This is the drift influence coefficient; The maximum threshold for motion features; for Quantized value of steering angle feature at any time; , They are respectively Quantitative values of longitudinal and lateral acceleration characteristics at any given time; The specific formula for the reliability model of vision sensors is as follows: ; in, for Reliability of time-lapse vision sensors for The number of feature point matches detected by vision at any given moment; The maximum number of matches; The coefficient representing the influence of the number of matches; for The degree of visual lane line occlusion at any given time; for The cumulative duration of occlusion at any given moment; This is the occlusion duration coefficient.
[0008] In the above scheme, step 3, the method for multi-level fault detection of the sensor is as follows: The first-level instantaneous residual test calculates the reliability-weighted instantaneous residual and compares it with a threshold. The specific method is as follows: ; ; in, for Time sensor Reliability-weighted instantaneous residuals; for Time sensor Observed values; sensor The observation matrix; for The system predicts the state at any given time. For sensors The normal observation variance; For sensors The variance of fault observations; It is a sensor Reliability; The residual threshold; This serves as the first level of verification criteria; a reliability-weighted residual denominator is introduced: reliable sensors use their own observation variance. Suspected unreliable sensors use fault variance ; The second layer of multi-sensor cross-validation involves verifying the consistency of observations from suspected faulty sensors with similar observations from other normal sensors. The specific method is as follows: ; ; in, , Sensors and Similar observations; This is the consistency threshold; This is the second layer of inspection mark; Suspected faulty sensor The observed variance; For sensors The observed variance; The third layer of long-term fault confirmation uses the fault probability of historical frames to make the final fault determination. The specific method is as follows: ; ; in, for Time sensor Historical failure probability; This is the failure probability threshold; It is the number of frames; This is the second layer of inspection mark; As the third layer of inspection mark, only when If this happens, the sensor is determined to be faulty.
[0009] In the above scheme, step 3, the method for graded fault isolation based on the test results is as follows: The specific formula for calculating the sensor's failure impact factor is as follows: ; in, for Time sensor Fault influencing factors for System state at time 1 dimension, for Time sensor No. 3D observation, For state dimension weights, for Time sensor Reliability model; Based on the magnitude of the fault impact factor and the sensitivity of the state dimension, different levels of isolation strategies are implemented, including: isolating only the high-sensitivity dimension, isolating the medium-to-high-sensitivity dimension, or completely isolating all observation dimensions of the sensor.
[0010] In the above scheme, step 4, the method for determining the vehicle driving scenario is as follows: The specific formula for scene determination is as follows: ; in, The result of scene determination; For motion feature threshold, , , , These are the threshold values for steering angle, longitudinal acceleration, lateral acceleration, and angular velocity, respectively. for Quantized value of steering angle feature at any time; for Quantized value of longitudinal acceleration feature at any given time; for Quantized value of lateral acceleration characteristic at any given time; for The angular velocity characteristics of the vehicle at any given moment.
[0011] In the above scheme, step 4 sets the optimization objective as the vehicle state vector, as follows: ; in, for The vehicle motion state vector at any given moment; , , They represent in The position of the vehicle in the x, y, z directions in the geographic coordinate system at any given time; , , Indicates in The velocity components of the vehicle in the x, y, and z directions in the vehicle coordinate system at any given moment; , , They represent in The vehicle's roll angle, pitch angle, and yaw angle at all times; The method for initializing the weight values of each sensor is as follows: During the vehicle startup phase, initial weights are assigned to the IMU, GNSS, and vision sensors respectively. The specific formula for the initial weight of the IMU is as follows: ; in, for Initial weights of the IMU at time step; for Time-based IMU baseline weighting coefficients; for Time-based IMU calibration; for GNSS readiness at any given time indicates the accuracy of the GNSS sensor; for The real-time visual readiness rating indicates the accuracy of the visual sensors; the clearer the lane lines, the higher this rating. The specific formula for the initial weights of GNSS is as follows: ; in, for Initial GNSS weights at time; for GNSS base weight coefficient at any given time; The specific formula for the initial visual weights is as follows: ; in, for Initial visual weights at each moment; for Moment-based visual weighting coefficients; The initial weights are normalized so that the sum of the weights is 1, which satisfies the following condition. The initial weight normalization formula is as follows: ; in, , , These are the initial weights of each sensor after normalization; for Initial weights of the IMU at time step; for Initial GNSS weights at time; for Initial visual weights at each moment.
[0012] In the above scheme, step 4, the method for dynamically updating the sensor weights based on the driving scenario is as follows: In stable scenarios, increase the weight of GNSS and vision, and decrease the weight of IMU; in dynamic scenarios, increase the weight of IMU, and decrease the weight of GNSS and vision. The specific formula for updating IMU weights is as follows: ; in, for Time-based IMU weights; These are the IMU baseline coefficients; The Sigmoid normalization function, For motion feature vectors; for Time-based IMU calibration; The threshold for motion features; The specific formula for GNSS weight update is as follows: ; in, for Time-based GNSS weights; The basic coefficients for stable GNSS scenes; For motion feature vectors; GNSS readiness; The specific formula for updating the visual sensor weights is as follows: ; in, for Momentary visual weight; For visual baseline coefficients; Visual readiness.
[0013] In the above scheme, the method for applying weight smoothing constraints in step 4 is as follows: ; in, , , They are respectively Final weights of each sensor after time-smoothing; For smoothing coefficients; , , These are the smoothed weights from the previous time step; Then, the contributions of each sensor are balanced using upper and lower limit constraints, as follows: ; in, This serves as the lower limit for weights, ensuring that the weight of any sensor is not zero. The upper limit of the weight is set to ensure that the weight of any sensor does not exceed 0.8.
[0014] In the above scheme, step 4, the weight normalization method is as follows: ; ; ; in, , , They are respectively The dynamic weighted information matrix of each sensor at any given time. The diagonal elements represent the local covariance of the IMU, with the position, velocity, and attitude variances respectively. For GNSS local covariance, only the position dimension has non-zero variance; For visual local covariance, only the relative position and heading dimensions have non-zero variance; The overall weights are normalized to ensure that the sum of the weights for each state dimension is an identity matrix. The specific method is as follows: ; ; ; ; in, for The total information matrix at any given time is the sum of three dynamically weighted information matrices; , , They are respectively A dynamic weighted information matrix of IMU, GNSS, and visual sensors at any given time; , , They are respectively A comprehensive fusion weight matrix for real-time IMU, GNSS, and visual sensors; The vehicle state vector is obtained by normalizing the weights and performing weighted fusion. ; in, for The optimal estimate of the vehicle's motion state vector at any given time; , , They are respectively A comprehensive fusion weight matrix for real-time IMU, GNSS, and visual sensors; for The local state of the IMU at any given time includes the position, velocity, and attitude updated by the IMU pre-integration and sub-filter. for Local GNSS state at any given time, containing only global position; for The real-time visual local state, including the relative position offset of lane lines and heading deviation, is obtained by solving the vehicle state vector to obtain the real-time position information of the vehicle.
[0015] Through the above technical solution, the intelligent perception and localization method for autonomous driving environment provided by the present invention, which involves multimodal data fusion, has the following beneficial effects: 1. This invention breaks through the limitations of traditional single-feature quantification in resisting interference by quantifying key features such as steering angle, acceleration and angular velocity. It effectively suppresses the negative impact of external interference such as road bumps on the judgment of vehicle status, realizes a comprehensive and accurate characterization of vehicle motion status, and provides highly reliable data for intelligent driving decision-making.
[0016] 2. This invention designs a novel fault-tolerant filtering framework, constructs a dynamic model of sensor reliability, and implements a multi-level fault detection mechanism. This enables rapid location and isolation of sensor faults and resistance to interference from complex environments, ensuring the reliability and stable output of sensor data and significantly improving the robustness of the navigation system.
[0017] 3. This invention uses a threshold-based judgment to perceive the vehicle's driving scenario and employs an adaptive weight allocation strategy driven by dynamic scenario matching. This reduces misjudgment of the vehicle's driving scenario, solves the problem that fixed weight allocation cannot adapt to different driving scenarios, and achieves high-precision vehicle positioning when the driving scenario changes dynamically. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0019] Figure 1 This is a schematic diagram of a multimodal data fusion-based intelligent perception and localization method for autonomous driving environments disclosed in an embodiment of the present invention.
[0020] Figure 2 The structure diagram of the multi-sensor fusion fault-tolerant filtering algorithm for reliability dynamic modeling and multi-level verification. Detailed Implementation
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0022] This invention provides a multimodal data fusion-based intelligent perception and localization method for autonomous driving environments, such as... Figure 1 As shown, it includes the following steps: Step S1: Sensor data modeling: It receives data from GNSS, IMU, and vision sensors mounted on autonomous vehicles, and establishes measurement models for each sensor, as detailed below: The IMU sensor uses the output data from the accelerometer and gyroscope to calculate and obtain the position, velocity, and attitude information of the autonomous vehicle. The specific formula for the IMU sensor measurement model is as follows: ; ; in, The system acceleration vector output by the accelerometer. This is the system angular velocity vector output by the gyroscope. , These are the random walk biases for acceleration and angular velocity as a function of time, respectively. These are zero-mean Gaussian white noise during the acceleration and angular velocity measurements, respectively. , These are the vehicle's actual acceleration and angular velocity, respectively.
[0023] GNSS provides global position observation information for vehicles. The specific formula for the GNSS measurement model is as follows: ; in, , , These represent the vehicle's position in the x, y, and z directions of the geographic coordinate system, respectively. Indicates the vehicle's position in the geographic coordinate system; A visual sensor measurement model is established, utilizing the visual sensor to capture key information between two consecutive frames, including feature extraction, keypoint matching, and camera pose estimation. The specific formulas are as follows: ; ; ; in, Represents the three-dimensional coordinates of a feature point in the world coordinate system. The intrinsic parameter matrix representing the camera calibration. This represents the extrinsic parameter matrix of the camera calibration. Represents the pixel coordinates of the output feature points. Represents the actual pixel coordinates of the feature points. Represents Gaussian noise; This represents the relative rotation matrix calculated from the two image frames. Represents the actual rotation matrix. Represents disturbance noise, obeys ; This represents the relative translation vector calculated from the two image frames. For the actual translation vector, This represents Gaussian noise.
[0024] The goal of system modeling is to map the outputs of different sensors to a unified state space, providing a unified framework for subsequent fusion, while also facilitating error analysis and weight allocation.
[0025] Step S2: Quantification of vehicle motion state features: Based on the measurement model, the characteristics of vehicle motion state are quantified, including at least steering angle characteristics, longitudinal acceleration characteristics, lateral acceleration characteristics, and angular velocity characteristics.
[0026] First, the front wheel steering angle It is a key parameter affecting the rate of change of heading angle. The specific method for quantifying the characteristics of the steering angle is as follows: ; in, for Quantized value of steering angle feature at any time; for The steering angle of the vehicle's front wheels at any given moment is positive when turning left; This refers to the sensor sampling time step. The length of the time window; , These are the steering angle feature weight coefficients, all of which are positive and satisfy the following conditions: ; Secondly, longitudinal acceleration directly reflects the severity of a vehicle's active acceleration and deceleration; minor acceleration fluctuations caused by uneven road surfaces are not considered active driving behavior and need to be filtered through a threshold. The specific method for quantifying longitudinal acceleration features is as follows: ; in, for Quantized value of longitudinal acceleration feature at any given time; for Real-time longitudinal acceleration of the vehicle; for The actual longitudinal acceleration at any given moment; For longitudinal acceleration threshold, when The system was judged to be due to road surface interference rather than active acceleration / deceleration. , These are the longitudinal acceleration characteristic coefficients, all of which are positive numbers. ; Lateral acceleration is generated by the centrifugal force when a vehicle turns, and is directly related to the dynamic risks of turning scenarios. Its magnitude is positively correlated with vehicle speed and lane curvature. The specific method for quantifying lateral acceleration features is as follows: ; in, for Quantized value of lateral acceleration characteristic at any given time; for Real-time longitudinal speed of the vehicle; for Lane curvature at any time, straight road time When turning The absolute value increases with the steepness of the turn; Finally, during vehicle movement, road bumps (such as gravel roads or speed bumps) or sharp turns can cause drastic changes in vehicle attitude (roll, pitch), affecting the accuracy of sensor observations. The specific method for quantifying vehicle angular velocity is as follows: ; in, for Vehicle angular velocity characteristics at any given moment; , , They are respectively The actual angular velocities of the vehicle along its x, y, and z axes at any given moment; , , They are respectively The true angular velocities of the vehicle along its x, y, and z axes at any given time.
[0027] By calculating the Euclidean distance of the changes in angular velocity across the three axes, the changes in roll angular velocity caused by unilateral bumps and the changes in heading and angular velocity during sharp turns can be effectively captured, avoiding local misjudgments caused by single-axis fluctuations and comprehensively reflecting the vehicle's attitude stability.
[0028] Step S3: Reliability dynamic modeling and multi-level verification of multi-sensor fusion fault-tolerant filtering: Based on the measurement model in step S1 and the feature quantization results in step S2, a dynamic reliability model for each sensor is constructed. Multi-level fault detection is performed on the sensors, and graded fault isolation is conducted according to the inspection results. Only retained sensor data is then selected. The process is as follows: Figure 2 As shown.
[0029] (1) Dynamic modeling of sensor reliability Traditional reliability modeling relies solely on sensor observations. This design utilizes the correlation between historical error trends and motion states to construct a three-dimensional reliability index.
[0030] The specific formula for the GNSS reliability model is as follows: ; in, for The reliability of GNSS at any given time; for GNSS horizontal accuracy factor at any time; , These are the minimum and maximum thresholds for the GNSS horizontal accuracy factor, respectively. DOP impact coefficient; Number of frames; for Time before Number of valid intra-frame GNSS observations; Indicates the signal continuity factor; The specific method for IMU reliability modeling is as follows: ; in, for The reliability of the IMU at any given moment; for Time before IMU Average drift residual of the frame; This represents the maximum permissible drift residual for the IMU. This is the drift influence coefficient; The maximum threshold for motion features; for Quantized value of steering angle feature at any time; , They are respectively Quantitative values of longitudinal and lateral acceleration characteristics at any given time; The specific formula for the reliability model of vision sensors is as follows: ; in, for Reliability of time-lapse vision sensors for The number of feature point matches detected by vision at any given moment; The maximum number of matches; The coefficient representing the influence of the number of matches; for The degree of visual lane line occlusion at any given time; for The cumulative duration of occlusion at any given moment; This is the occlusion duration coefficient.
[0031] (2) Multi-level fault detection Based on the sensor reliability model described above, the method for multi-level fault detection of sensors is as follows: First-level instantaneous residual check: First, for each sensor, calculate the reliability-weighted instantaneous residual and compare it with a threshold to check whether the current observation deviates from the predicted state. The specific method is as follows: ; ; in, for Time sensor Reliability-weighted instantaneous residuals; for Time sensor Observed values; sensor The observation matrix; for The system predicts the state at any given time. For sensors The normal observation variance; For sensors The variance of fault observations; It is a sensor Reliability; The residual threshold; This serves as the first level of verification criteria; a reliability-weighted residual denominator is introduced: reliable sensors use their own observation variance. Suspected unreliable sensors use fault variance .
[0032] Second-layer multi-sensor cross-validation: Then, consistency checks are performed on the observations of the first-layer suspected faulty sensors and other normal sensors (e.g., comparing GNSS position with visually estimated position). The specific method is as follows: ; ; in, , Sensors and Similar observations; This is the consistency threshold; This is the second layer of inspection mark; Suspected faulty sensor The observed variance; For sensors The observed variance; Cross-validate suspected faulty sensors by observing normal sensors to avoid misjudging a single sensor due to transient noise (e.g., large residuals caused by GNSS transient multipath, but consistent visual position, which can rule out fault).
[0033] Third-level long-term fault confirmation: Finally, for the second layer of suspected sensors, long-term fault confirmation and analysis were conducted. To avoid misjudgment due to short-term fluctuations, the failure probability of a frame is used to make the final failure determination based on the failure probability of historical frames. The specific method is as follows: ; ; in, for Time sensor Historical failure probability; This is the failure probability threshold; It is the number of frames; This is the second layer of inspection mark; As the third layer of inspection mark, only when If this happens, the sensor is determined to be faulty.
[0034] By introducing a reliability-weighted historical failure probability, reliable sensors are given a higher weight for suspected failures, while unreliable sensors are given a lower weight for suspected failures, ensuring that long-term failures are confirmed and short-term fluctuations are filtered out.
[0035] (3) Hierarchical fault isolation Traditional isolation means complete discarding upon failure. This invention achieves hierarchical isolation based on the failure impact factor, retaining the effective observation dimensions of the faulty sensor and discarding only the failed dimensions. The failure impact factor quantifies the influence of each observation dimension of the sensor on the system state. Dimensions with small impact can be retained, while dimensions with large impact need to be isolated to avoid complete discarding.
[0036] The method for grading and isolating faults based on test results is as follows: First, calculate the sensor's failure impact factor using the following formula: ; in, for Time sensor Fault influencing factors for System state at time 1 dimension, for Time sensor No. 3D observation, For state dimension weights, for Time sensor Reliability model; Depending on the magnitude of the fault impact factor and the sensitivity of the state dimension, different levels of isolation strategies are implemented, including: isolating only the high-sensitivity dimension, isolating the medium-to-high-sensitivity dimension, or completely isolating all observation dimensions of the sensor. For example, when (Low-impact fault): Isolate sensitivity only Dimensions (e.g., visual heading observation failure, retain lateral position observation); when (Medium-term impact of faults): Isolation sensitivity Dimensions (e.g., IMU attitude observation failure, retaining velocity observation); when (High-impact fault): Completely isolate the sensor All observations (such as complete GNSS disconnection, no valid position observations).
[0037] By implementing a full-process fault-tolerance mechanism that includes dynamic modeling of sensor reliability, multi-level fault detection, and graded fault isolation, the system achieves robustness in fault identification and system performance, further enhancing filtering robustness and fault tolerance in complex scenarios, and providing continuous, reliable, and high-precision positioning results for autonomous vehicles.
[0038] Step S4: Dynamic weight allocation based on vehicle motion state: Based on the retained sensor data, the vehicle driving scenario is determined; optimization objectives are set and the weight values of each sensor are initialized; according to the vehicle driving scenario, the sensor weights are dynamically updated, and weight smoothing constraints and normalization are performed. The vehicle state vector is solved through weighted fusion, and the vehicle positioning information is output.
[0039] The specific method is as follows: (1) Vehicle driving scenario judgment First, it is necessary to determine the vehicle driving scenario, distinguishing between stable and dynamic scenarios, to avoid misjudgments caused by relying solely on the motion state. Stable scenarios require gentle motion and reliable sensors, while dynamic scenarios require violent motion or unreliable sensors.
[0040] The specific formula for scene determination is as follows: ; in, The result of scene determination; For motion feature threshold, , , , These are the threshold values for steering angle, longitudinal acceleration, lateral acceleration, and angular velocity, respectively. for Quantized value of steering angle feature at any time; for Quantized value of longitudinal acceleration feature at any given time; for Quantized value of lateral acceleration characteristic at any given time; for The angular velocity characteristics of the vehicle at any given moment.
[0041] (2) Optimization objective and weight initialization The optimization objective is set as the vehicle state vector, as follows: ; in, for The vehicle motion state vector at any given moment; , , They represent in The position of the vehicle in the x, y, z directions in the geographic coordinate system at any given time; , , Indicates in The velocity components of the vehicle in the x, y, and z directions in the vehicle coordinate system at any given moment; , , They represent in The vehicle's roll angle, pitch angle, and yaw angle at all times; (3) Weight initialization The method for initializing the weight values of each sensor is as follows: During the vehicle startup phase, initial weights are assigned to the IMU, GNSS, and vision sensors. Initially, the high-frequency output of the IMU is prioritized during startup, with its weight gradually decreasing as the startup process progresses.
[0042] The specific formula for the initial weights of the IMU is as follows: ; in, for Initial weights of the IMU at time step; for Time-based IMU baseline weighting coefficients; for Time-based IMU calibration; for GNSS readiness at any given time indicates the accuracy of the GNSS sensor; for The real-time visual readiness rating indicates the accuracy of the visual sensors; the clearer the lane lines, the higher this rating. Initially, the GNSS weight is reduced, and it is increased as the startup progresses and GNSS readiness improves, gradually enhancing its global positioning capabilities. The specific formula for the initial GNSS weight is as follows: ; in, for Initial GNSS weights at time; for GNSS base weight coefficient at any given time; Initially, the visual weights are reduced, and then increased as the startup process progresses and visual readiness improves, providing local accuracy constraints for subsequent joint calibration. The specific formula for the initial visual weights is as follows: ; in, for Initial visual weights at each moment; for Moment-based visual weighting coefficients; During the startup phase, the initial weights of each sensor may deviate from 1 due to variable fluctuations. Therefore, it is necessary to normalize the initial weights to ensure the sum of the weights equals 1, thus satisfying the condition... To avoid numerical deviations in subsequent fusion calculations, the initial weight normalization formula is as follows: ; in, , , These are the initial weights of each sensor after normalization; for Initial weights of the IMU at time step; for Initial GNSS weights at time; for Initial visual weights at each moment.
[0043] (4) Dynamic update of weights The method for dynamically updating sensor weights based on the driving scenario is as follows: During the driving phase, the weights need to be adjusted according to the intensity of vehicle movement (dynamic / stable scenario) and sensor reliability: GNSS and vision are prioritized in stable scenarios, and IMU is prioritized in dynamic scenarios.
[0044] In stable scenarios, where vehicle movement is gentle (e.g., constant speed in a straight line) and sensors are reliable, the weight of GNSS and vision should be increased first, while the weight of IMU should be decreased to suppress drift. In dynamic scenarios, where vehicle movement is violent (e.g., sharp turns, rapid acceleration) or a sensor is unreliable (e.g., GNSS multipath, visual occlusion), the weight of IMU should be increased first, while the weight of GNSS and vision should be decreased.
[0045] In stable scenarios, the IMU provides only basic high-frequency output. Weights are suppressed by a low fundamental coefficient and smooth motion to prevent IMU drift from affecting fusion accuracy. In dynamic scenarios, the IMU's low latency advantage is significant. Weights are increased by a high fundamental coefficient and intense motion to maximize the effect of low-latency output. The specific formula for IMU weight update is as follows: ; in, for Time-based IMU weights; These are the IMU baseline coefficients; The Sigmoid normalization function, For motion feature vectors; for Time-based IMU calibration; This is the threshold for motion characteristics.
[0046] In stable scenarios, GNSS global positioning accuracy is high. By increasing the base coefficient and the smoothness of motion, the weights are maximized to maximize the global positioning constraint effect and ensure reliable GNSS participation in fusion. In dynamic scenarios, GNSS is susceptible to multipath interference. By decreasing the base coefficient and the intensity of motion, the weights are reduced to ensure reliable GNSS participation in fusion. The specific formula for GNSS weight update is as follows: ; in, for Time-based GNSS weights; The basic coefficients for stable GNSS scenes; For motion feature vectors; For GNSS readiness.
[0047] In stable scenarios, lane lines are clearly observed visually. By increasing the weights with high base coefficients and smooth motion, the local accuracy constraint is maximized, ensuring reliable visual participation in fusion. In dynamic scenarios, visual processing is delayed. By reducing the weights with low base coefficients and intense motion, reliable visual participation in fusion is ensured. The specific formula for updating visual sensor weights is as follows: ; in, for Momentary visual weight; For visual baseline coefficients; Visual readiness.
[0048] (5) Weight smoothing and constraints When switching scenes (such as from smooth to dynamic), unsmoothed weights may change abruptly. The smoothing coefficient is adjusted according to the scene change rate to ensure a balance between adaptability and continuity. The method for implementing weight smoothing constraints is as follows: ; in, , , They are respectively Final weights of each sensor after time-smoothing; For smoothing coefficients; , , These are the smoothed weights from the previous time step.
[0049] To avoid over-reliance on a single sensor (e.g., in dynamic scenes, the IMU weight increases to 0.9, accumulating drift risk; in stable scenes, the visual weight increases to 0.8, increasing the risk of failure during occlusion), the contributions of each sensor are balanced through upper and lower bound constraints. The specific method is as follows: ; in, This serves as the lower limit for weights, ensuring that the weight of any sensor is not zero. The upper limit of the weight is set to ensure that the weight of any sensor does not exceed 0.8.
[0050] (6) Weight normalization This section is based on the final dynamic weights. , , The global optimal state is solved through weighted information fusion. The final dynamic weights are used as scene adaptation coefficients and multiplied by the information matrix to ensure scene and reliability matching. The weight normalization method is as follows: ; ; ; in, , , They are respectively The dynamic weighted information matrix of each sensor at any given time. The diagonal elements represent the local covariance of the IMU, with the position, velocity, and attitude variances respectively. For GNSS local covariance, only the position dimension has non-zero variance; For visual local covariance, only the relative position and heading dimensions have non-zero variance; The overall weights are normalized to ensure that the sum of the weights for each state dimension is an identity matrix. The specific method is as follows: ; ; ; ; in, for The total information matrix at any given time is the sum of three dynamically weighted information matrices; , , They are respectively A dynamic weighted information matrix of IMU, GNSS, and visual sensors at any given time; , , They are respectively A comprehensive fusion weight matrix of real-time IMU, GNSS, and visual sensors.
[0051] The vehicle state vector is obtained by normalizing the weights and performing weighted fusion. ; in, for The optimal estimate of the vehicle's motion state vector at any given time; , , They are respectively A comprehensive fusion weight matrix for real-time IMU, GNSS, and visual sensors; for The local state of the IMU at any given time includes the position, velocity, and attitude updated by the IMU pre-integration and sub-filter. for Local GNSS state at any given time, containing only global position; for The real-time visual local state, including the relative position offset of lane lines and heading deviation, is obtained by solving the vehicle state vector to obtain the real-time position information of the vehicle.
[0052] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal data fusion method for intelligent perception and localization in autonomous driving environments, characterized in that, Includes the following steps: Step S1: Receive data from GNSS, IMU and vision sensors mounted on the autonomous vehicle, and establish measurement models for each sensor respectively; Step S2: Based on the measurement model, quantify the vehicle motion state characteristics, which include at least steering angle characteristics, longitudinal acceleration characteristics, lateral acceleration characteristics, and angular velocity characteristics; Step S3: Based on the measurement model in Step S1 and the feature quantization results in Step S2, construct the dynamic reliability model of each sensor, perform multi-level fault detection on the sensors, perform graded fault isolation according to the test results, and select the retained sensor data. Step S4: Determine the vehicle driving scenario based on the retained sensor data; Set the optimization target and initialize the weight values for each sensor; Based on the vehicle driving scenario, the sensor weights are dynamically updated, and weight smoothing constraints and normalization are performed. The vehicle state vector is solved by weighted fusion, and the vehicle positioning information is output.
2. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, Step 2, quantifying the vehicle motion state characteristics includes: The specific method for quantifying the steering angle feature is as follows: ; in, for Quantized value of steering angle feature at any time; for The steering angle of the vehicle's front wheels at any given moment is positive when turning left; This refers to the sensor sampling time step. The length of the time window; , These are the steering angle feature weight coefficients, all of which are positive and satisfy the following conditions: ; The specific method for quantizing the longitudinal acceleration features is as follows: ; in, for Quantized value of longitudinal acceleration feature at any given time; for Real-time longitudinal acceleration of the vehicle; for The actual longitudinal acceleration at any given moment; For longitudinal acceleration threshold, when The system was initially determined to be road interference rather than active acceleration / deceleration. , These are the longitudinal acceleration characteristic coefficients, all of which are positive numbers. ; The specific method for quantifying the lateral acceleration features is as follows: ; in, for Quantized value of lateral acceleration characteristic at any given time; for Real-time longitudinal speed of the vehicle; for Lane curvature at any time, straight road time When turning The absolute value increases with the steepness of the turn; The specific method for quantifying the vehicle angular velocity is as follows: ; in, for Vehicle angular velocity characteristics at any given moment; , , They are respectively The actual angular velocities of the vehicle along its x, y, and z axes at any given moment; , , They are respectively The true angular velocities of the vehicle along its x, y, and z axes at any given time.
3. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 3, the dynamic reliability models for each sensor are constructed as follows: The specific formula for the GNSS reliability model is as follows: ; in, for The reliability of GNSS at any given time; for GNSS horizontal accuracy factor at any time; , These are the minimum and maximum thresholds for the GNSS horizontal accuracy factor, respectively. DOP impact coefficient; Number of frames; for Time before Number of valid intra-frame GNSS observations; Indicates the signal continuity factor; The specific method for IMU reliability modeling is as follows: ; in, for The reliability of the IMU at any given moment; for Time before IMU Average drift residual of the frame; This represents the maximum permissible drift residual for the IMU. This is the drift influence coefficient; The maximum threshold for motion features; for Quantized value of steering angle feature at any time; , They are respectively Quantitative values of longitudinal and lateral acceleration characteristics at any given time; The specific formula for the reliability model of vision sensors is as follows: ; in, for Reliability of time-lapse vision sensors for The number of feature point matches detected by vision at any given moment; The maximum number of matches; The coefficient representing the influence of the number of matches; for The degree of visual lane line occlusion at any given time; for The cumulative duration of occlusion at any given moment; This is the occlusion duration coefficient.
4. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 3, the method for multi-level fault detection of the sensor is as follows: The first-level instantaneous residual test calculates the reliability-weighted instantaneous residual and compares it with a threshold. The specific method is as follows: ; ; in, for Time sensor Reliability-weighted instantaneous residuals; for Time sensor Observed values; sensor The observation matrix; for The system predicts the state at any given time. For sensors The normal observation variance; For sensors The variance of fault observations; It is a sensor Reliability; The residual threshold; This serves as the first level of verification criteria; a reliability-weighted residual denominator is introduced: reliable sensors use their own observation variance. Suspected unreliable sensors use fault variance ; The second layer of multi-sensor cross-validation involves verifying the consistency of observations from suspected faulty sensors with similar observations from other normal sensors. The specific method is as follows: ; ; in, , Sensors and Similar observations; This is the consistency threshold; This is the second layer of inspection mark; Suspected faulty sensor The observed variance; For sensors The observed variance; The third layer of long-term fault confirmation uses the fault probability of historical frames to make the final fault determination. The specific method is as follows: ; ; in, for Time sensor Historical failure probability; This is the failure probability threshold; It is the number of frames; This is the second layer of inspection mark; As the third layer of inspection mark, only when If this happens, the sensor is determined to be faulty.
5. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 3, the method for grading and isolating faults based on the test results is as follows: The specific formula for calculating the sensor's failure impact factor is as follows: ; in, for Time sensor Fault influencing factors for System state at time 1 dimension, for Time sensor No. 3D observation, For state dimension weights, for Time sensor Reliability model; Based on the magnitude of the fault impact factor and the sensitivity of the state dimension, different levels of isolation strategies are implemented, including: isolating only the high-sensitivity dimension, isolating the medium-to-high-sensitivity dimension, or completely isolating all observation dimensions of the sensor.
6. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 4, the method for determining the vehicle's driving scenario is as follows: The specific formula for scene determination is as follows: ; in, The result of scene determination; For motion feature threshold, , , , These are the threshold values for steering angle, longitudinal acceleration, lateral acceleration, and angular velocity, respectively. for Quantized value of steering angle feature at any time; for Quantized value of longitudinal acceleration feature at any given time; for Quantized value of lateral acceleration characteristic at any given time; for The angular velocity characteristics of the vehicle at any given moment.
7. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 4, the optimization objective is set as the vehicle state vector, as follows: ; in, for The vehicle motion state vector at any given moment; , , They represent in The position of the vehicle in the x, y, z directions in the geographic coordinate system at any given time; , , Indicates in The velocity components of the vehicle in the x, y, and z directions in the vehicle coordinate system at any given moment; , , They represent in The vehicle's roll angle, pitch angle, and yaw angle at all times; The method for initializing the weight values of each sensor is as follows: During the vehicle startup phase, initial weights are assigned to the IMU, GNSS, and vision sensors respectively. The specific formula for the initial weight of the IMU is as follows: ; in, for Initial weights of the IMU at time step; for Time-based IMU baseline weighting coefficients; for Time-based IMU calibration; for GNSS readiness at any given time indicates the accuracy of the GNSS sensor; for The real-time visual readiness rating indicates the accuracy of the visual sensors; the clearer the lane lines, the higher this rating. The specific formula for the initial weights of GNSS is as follows: ; in, for Initial GNSS weights at time; for GNSS base weight coefficient at any given time; The specific formula for the initial visual weights is as follows: ; in, for Initial visual weights at each moment; for Moment-based visual weighting coefficients; The initial weights are normalized so that the sum of the weights is 1, which satisfies the following condition. The initial weight normalization formula is as follows: ; in, , , These are the initial weights of each sensor after normalization; for Initial weights of the IMU at time step; for Initial GNSS weights at time; for Initial visual weights at each moment.
8. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 4, the method for dynamically updating the sensor weights based on the driving scenario is as follows: In stable scenarios, increase the weight of GNSS and vision, and decrease the weight of IMU; in dynamic scenarios, increase the weight of IMU, and decrease the weight of GNSS and vision. The specific formula for updating IMU weights is as follows: ; in, for Time-based IMU weights; These are the IMU baseline coefficients; The Sigmoid normalization function, For motion feature vectors; for Time-based IMU calibration; The threshold for motion features; The specific formula for GNSS weight update is as follows: ; in, for Time-based GNSS weights; The basic coefficients for stable GNSS scenes; For motion feature vectors; GNSS readiness; The specific formula for updating the visual sensor weights is as follows: ; in, for Momentary visual weight; For visual baseline coefficients; Visual readiness.
9. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 4, the method for applying weight smoothing constraints is as follows: ; in, , , They are respectively Final weights of each sensor after time-smoothing; For smoothing coefficients; , , These are the smoothed weights from the previous time step; Then, the contributions of each sensor are balanced using upper and lower limit constraints, as follows: ; in, This serves as the lower limit for weights, ensuring that the weight of any sensor is not zero. The upper limit of the weight is set to ensure that the weight of any sensor does not exceed 0.
8.
10. The method for intelligent perception and localization of autonomous driving environment based on multimodal data fusion according to claim 1, characterized in that, In step 4, the weight normalization method is as follows: ; ; ; in, , , They are respectively The dynamic weighted information matrix of each sensor at any given time. The diagonal elements represent the local covariance of the IMU, with the position, velocity, and attitude variances respectively. For GNSS local covariance, only the position dimension has non-zero variance; For visual local covariance, only the relative position and heading dimensions have non-zero variance; The overall weights are normalized to ensure that the sum of the weights for each state dimension is an identity matrix. The specific method is as follows: ; ; ; ; in, for The total information matrix at any given time is the sum of three dynamically weighted information matrices; , , They are respectively A dynamic weighted information matrix of IMU, GNSS, and visual sensors at any given time; , , They are respectively A comprehensive fusion weight matrix for real-time IMU, GNSS, and visual sensors; The vehicle state vector is obtained by normalizing the weights and performing weighted fusion. ; in, for The optimal estimate of the vehicle's motion state vector at any given time; , , They are respectively A comprehensive fusion weight matrix for real-time IMU, GNSS, and visual sensors; for The local state of the IMU at any given time includes the position, velocity, and attitude updated by the IMU pre-integration and sub-filter. for Local GNSS state at any given time, containing only global position; for The real-time visual local state, including the relative position offset of lane lines and heading deviation, is obtained by solving the vehicle state vector to obtain the real-time position information of the vehicle.