Multi-sensor fusion method and system for intelligent driving and vehicle
By dynamically assessing and adaptively fusing the states of multiple sensors, virtual perception data is generated, which solves the problems of low perception accuracy and missing perception dimensions under adverse weather conditions, thereby improving the perception robustness and reliability of the intelligent driving system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN JIANGXIA CHUNENG AUTOMOBILE TECHNOLOGY R&D CO LTD
- Filing Date
- 2026-02-02
- Publication Date
- 2026-04-28
AI Technical Summary
Existing multi-sensor fusion technologies suffer from low perception accuracy and missing perception dimensions under adverse weather conditions. They are unable to effectively handle complex scenarios with multiple interference factors coupled together, and lack a fine quantitative assessment of sensor data quality, resulting in the system being unable to maintain basic perception capabilities when some sensors fail.
By dynamically assessing the raw sensing data from multiple sensors, environmental state parameters and sensor quality assessment indicators are generated, dynamic confidence weights are calculated, and virtual sensing data is generated when sensors fail to assist the fusion process, thus achieving adaptive fusion.
It improves the perception robustness of the intelligent driving system in adverse weather conditions, optimizes resource utilization efficiency, and achieves safe performance degradation through algorithm redundancy in extreme situations, thereby enhancing the overall reliability of the system.
Smart Images

Figure CN121929178A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent driving technology, and specifically to a multi-sensor fusion method, system, and vehicle for intelligent driving. Background Technology
[0002] In the field of intelligent driving, multi-sensor fusion technology is typically used to achieve comprehensive and reliable perception of the vehicle's surrounding environment. This involves comprehensively utilizing data from sensors based on different principles, such as cameras, lidar, and millimeter-wave radar. By integrating the advantages of each sensor, the aim is to overcome the limitations of a single sensor, thereby improving the overall robustness and reliability of the environmental perception system under various operating conditions (such as heavy rain, heavy snow, dense fog, and complex and adverse weather conditions like strong light at night).
[0003] However, existing multi-sensor fusion technologies still have significant shortcomings when dealing with the aforementioned complex and severe weather conditions. First, existing fusion strategies switch between a limited number of weather modes, which cannot effectively handle complex scenarios where multiple interfering factors (such as reflections from wet roads combined with thin fog) are coupled. Second, existing solutions generally lack a fine-grained quantitative assessment of the instantaneous data quality of the sensors themselves, failing to diagnose in real time specific data quality degradation issues such as partial obstruction of camera lenses, noise from dense raindrops in LiDAR, and multipath reflections in millimeter-wave radar, resulting in low-quality or failed data continuing to participate in fusion calculations. Finally, when a sensor completely fails or its performance is severely degraded due to extreme weather, existing architectures often simply discard the sensor's data, leading to a lack of perception dimensions in the system and an inability to provide the vehicle with the minimum environmental information needed for intelligent decision-making to perform safety-keeping operations such as parking.
[0004] Therefore, how to achieve higher perception accuracy under adverse weather conditions and maintain basic perception capabilities when some sensors fail is a technical problem that urgently needs to be solved in the field of intelligent driving environmental perception. Summary of the Invention
[0005] In view of this, it is necessary to provide a multi-sensor fusion method, system and vehicle for intelligent driving to solve the technical problems of low perception accuracy and missing perception dimensions in adverse weather conditions caused by the rigidity of fusion strategies and the coarseness of sensor state assessment in existing methods.
[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides a multi-sensor fusion method for intelligent driving, comprising:
[0007] Acquire raw sensing data from at least two types of sensors, and perform dynamic state assessment on the raw sensing data to obtain environmental state parameters and sensor quality assessment indicators. Based on the environmental state parameters and sensor quality evaluation indicators, dynamic confidence weights for each sensor are generated. Based on the dynamic confidence weight, data from different sensors are adaptively fused to output fused perception results. During the fusion process, when the preset sensor failure conditions are met, virtual perception data is generated based on multi-source information, and the virtual perception data is incorporated into the adaptive fusion process with auxiliary weights.
[0008] In one possible implementation, the dynamic state assessment of the raw sensing data to obtain environmental state parameters and sensor quality assessment indicators includes: For camera data, evaluate at least one of image sharpness, occlusion area, and glare intensity; For lidar data, evaluate at least one of the following: point cloud noise ratio, effective point cloud density, and echo intensity. For millimeter-wave radar data, evaluate at least one of the following: signal-to-noise ratio and multipath interference intensity.
[0009] In one possible implementation, generating the dynamic confidence weights for each sensor based on the environmental state parameters and sensor quality assessment indicators includes: The environmental state parameters and the sensor quality assessment indicators are input into the confidence calculation function corresponding to the sensor type to obtain the initial confidence weights; The initial confidence weights are subjected to time-smoothing filtering to obtain the dynamic confidence weights.
[0010] In one possible implementation, the preset sensor failure condition is: the dynamic confidence weight of at least one sensor remains below a preset threshold for a predetermined duration.
[0011] In one possible implementation, generating virtual sensing data based on multi-source information includes: Sensors whose dynamic confidence weight remains below a preset threshold for a predetermined duration are considered to be faulty sensors. Based on the type of the failed sensor, call the corresponding physical or geometric model; The data from other high-confidence sensors and the environmental state parameters are input into the called model, and virtual sensing data is generated through cross-modal inference.
[0012] In one possible implementation, the step of inputting data from other high-confidence sensors and the environmental state parameters into the invoked model to generate virtual sensing data through cross-modal inference includes: When the failed sensor is a lidar, virtual point cloud data is generated based on the millimeter-wave radar target information and geometric projection model; When the failed sensor is a camera, virtual image data is generated based on the 3D scene information and image rendering model from other sensors.
[0013] In one possible implementation, the adaptive fusion of data from different sensors based on the dynamic confidence weights includes at least one of the following processing levels: At the data preprocessing level, data is enhanced or removed based on the dynamic confidence weights. At the feature fusion level, features extracted from different sensors are weighted and fused according to the dynamic confidence weights. At the target fusion level, targets detected by different sensors are associated and tracked based on the dynamic confidence weights.
[0014] In one possible implementation, the auxiliary weight is preset to a fixed value lower than the normal sensor dynamic confidence weight range, or is preset to a low value range.
[0015] On the other hand, the present invention also provides a multi-sensor fusion system for intelligent driving, comprising: The data acquisition module is used to acquire raw sensing data from at least two types of sensors and to perform dynamic state evaluation on the raw sensing data to obtain environmental state parameters and sensor quality evaluation indicators. The weight generation module is used to generate dynamic confidence weights for each sensor based on the environmental state parameters and sensor quality evaluation indicators. The adaptive fusion module is used to adaptively fuse data from different sensors based on the dynamic confidence weights and output fused perception results. During the fusion process, when the preset sensor failure conditions are met, virtual perception data is generated based on multi-source information, and the virtual perception data is integrated into the adaptive fusion process with auxiliary weights.
[0016] Thirdly, the present invention also provides a vehicle including the aforementioned multi-sensor fusion system for intelligent driving.
[0017] The beneficial effects of this invention are as follows: The multi-sensor fusion method for intelligent driving provided by this invention dynamically evaluates the raw data of each sensor, outputting environmental parameters and refined quality indicators, which helps to accurately identify the degree of sensor interference. Then, based on these evaluation results, dynamic confidence weights are generated, thereby assigning a quantitative indicator reflecting the instantaneous credibility of each data source. Subsequently, multi-source data is adaptively fused according to this weight, so that the fusion process can automatically focus on more reliable information, improving the accuracy of the output results. Finally, when continuous sensor failure is detected, virtual perception data is generated based on high-confidence data and integrated into the fusion process with low weight, so as to maintain the baseline perception capability of the system when critical sensors fail, avoid complete functional interruption, thereby significantly improving the perception robustness of the intelligent driving system in adverse weather conditions, optimizing resource utilization efficiency, and achieving safe performance degradation through algorithm redundancy in extreme cases, thus enhancing the overall reliability of the system. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic flowchart of an embodiment of the multi-sensor fusion method for intelligent driving provided by the present invention; Figure 2 For the present invention Figure 1 A schematic diagram of an embodiment of S102; Figure 3 For the present invention Figure 1 A schematic diagram of an embodiment of S103; Figure 4 For the present invention Figure 3 A schematic diagram of an embodiment of S303; Figure 5 This is a schematic diagram of an embodiment of the multi-sensor fusion system for intelligent driving provided by the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0021] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0022] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.
[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0024] This invention provides a multi-sensor fusion method, system, and vehicle for intelligent driving. The technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] Figure 1 This is a schematic flowchart of an embodiment of the multi-sensor fusion method for intelligent driving provided by the present invention, as shown below. Figure 1 As shown, the multi-sensor fusion method for intelligent driving includes: S101. Acquire raw sensing data from at least two types of sensors, and perform dynamic state evaluation on the raw sensing data to obtain environmental state parameters and sensor quality evaluation indicators. S102. Based on environmental state parameters and sensor quality evaluation indicators, generate dynamic confidence weights for each sensor. S103. Based on the dynamic confidence weight, adaptively fuse the data from different sensors and output the fused perception result. During the fusion process, when the preset sensor failure condition is met, virtual perception data is generated based on multi-source information, and the virtual perception data is integrated into the adaptive fusion process with auxiliary weights.
[0026] Dynamic state assessment refers to the real-time analysis of raw sensor data using dedicated algorithm units to extract quantitative features reflecting data quality and environmental conditions. Specific implementation methods may include, but are not limited to: using lightweight neural network models for image sharpness scoring and occlusion region segmentation; employing point cloud statistical analysis algorithms to calculate noise ratio and density; and utilizing signal processing techniques to evaluate the signal-to-noise ratio and interference intensity of radar signals.
[0027] Adaptive fusion refers to a data processing architecture that dynamically adjusts the fusion strategy and weights based on the real-time confidence level of the input data. Its implementation can cover multiple levels, such as: at the data preprocessing level, denoising or enhancing low-confidence data; at the feature extraction level, performing confidence-weighted concatenation of feature maps from different sources; and at the target decision level, performing confidence-based association and trajectory fusion on detection results from different sensors.
[0028] Virtual sensing data refers to synthetic data generated by algorithms that simulate the format and characteristics of real sensor data when real sensor data fails or is of low quality. Its generation depends on specific physical or geometric models. For example, based on a priori models of millimeter-wave radar target trajectories and vehicle contours, sparse point clouds simulating lidar are generated through geometric projection; or based on multi-view visual and map information, image patches of occluded areas are generated through 3D reconstruction and rendering techniques.
[0029] Step S101 involves simultaneously acquiring raw perception data streams from at least two types of sensors on the vehicle (e.g., cameras, LiDAR, millimeter-wave radar). Subsequently, dynamic state assessments are performed on these raw data in parallel.
[0030] This evaluation process aims to output two key types of information: macroscopic environmental state parameters (such as weather type and overall visibility level) and fine-grained diagnostic results on the data quality of each sensor, i.e., sensor quality evaluation indicators. This embodiment not only goes beyond the traditional method of coarsely classifying weather but also achieves a simultaneous and refined measurement of the intensity of environmental interference and the instantaneous health status of the sensors themselves.
[0031] Next, in step S102, the system comprehensively processes the environmental state parameters obtained from the above evaluation with the sensor quality evaluation indicators to generate a dynamic confidence weight for each sensor. This weight is a value between 0 and 1, which quantitatively represents the reliability of the corresponding sensor data at the current moment. Step S102 transforms the multi-dimensional heterogeneous evaluation indicators into a unified confidence metric through a configurable mapping mechanism. Therefore, for the data collected by the sensors, a continuously changing confidence level can be assigned based on its actual interference situation, rather than simply trusting or discarding it, thus providing an objective data quantification source for achieving smooth and adaptive fusion.
[0032] Finally, the adaptive fusion stage begins in step S103. In this stage, the data from all sensors do not participate in the fusion equally, but are weighted according to the dynamic confidence weights obtained in step S102. The entire fusion process can be carried out in a multi-level architecture, ensuring that each step from the raw data to the final target is guided by confidence.
[0033] It should be noted that the adaptive fusion process in this embodiment incorporates a virtual compensation mechanism.
[0034] Specifically, at the start of the fusion process, the confidence weights of each sensor are continuously monitored. When at least one sensor is determined to meet a preset failure condition (e.g., its confidence weight remains below a threshold for a certain period of time), this mechanism is triggered. This mechanism proactively uses data from other sensors that still maintain high confidence, combined with environmental state parameters, to generate virtual sensing data that matches the data type of the failed sensor through model inference. This virtual data is then re-injected into the fusion process with a preset, lower auxiliary weight.
[0035] This virtual compensation mechanism constructs a set of fault-softening safety redundancy paths. When a critical sensor temporarily fails due to extreme weather, it no longer simply loses its sensing capability in that dimension, but can make reasonable inferences and supplements based on existing information and physical laws, thereby maintaining a minimum level of environmental awareness of the system. This achieves a smooth transition from complete failure to controllable performance degradation, greatly improving the functional safety and robustness of the entire sensing system under harsh conditions.
[0036] This embodiment dynamically assesses the state of raw data from each sensor, outputting environmental parameters and refined quality indicators to accurately identify the degree of sensor interference. Then, based on these assessment results, dynamic confidence weights are generated, assigning a quantitative indicator reflecting the instantaneous reliability of each data source. Subsequently, multi-source data is adaptively fused according to these weights, allowing the fusion process to automatically prioritize more reliable information and improve the accuracy of the output results. Finally, when continuous sensor failure is detected, virtual perception data is generated based on high-confidence data and integrated into the fusion process with low weight. This maintains the system's baseline perception capability when critical sensors fail, preventing complete functional interruption. This significantly improves the perception robustness of the intelligent driving system in adverse weather conditions, optimizes resource utilization efficiency, and achieves safe performance degradation through algorithmic redundancy in extreme situations, enhancing the overall reliability of the system.
[0037] In some embodiments of the present invention, step S101 performs dynamic state assessment on the raw sensing data to obtain environmental state parameters and sensor quality assessment indicators, including: For camera data, evaluate at least one of image sharpness, occlusion area, and glare intensity; For lidar data, evaluate at least one of the following: point cloud noise ratio, effective point cloud density, and echo intensity. For millimeter-wave radar data, evaluate at least one of the following: signal-to-noise ratio and multipath interference intensity.
[0038] Image sharpness assessment refers to the technique of quantitatively analyzing the degree of image blur through algorithms. Specific implementation methods include, but are not limited to: image gradient-based methods (such as the Tenengrad function), frequency domain analysis-based methods (such as calculating the high-frequency energy after the Fourier transform of the image), and quality scoring using deep learning models.
[0039] Point cloud noise refers to discrete, outlier points in lidar point cloud data that do not represent real physical objects. Specific types include random scattering points caused by airborne particles (rain, snow, fog), and anomalous points caused by the sensor itself or multiple reflections. Detection methods typically include statistical distance filtering and density-based clustering analysis.
[0040] Multipath interference, in millimeter-wave radar detection, refers to the phenomenon where electromagnetic waves, besides their direct path, are reflected by other objects (such as the ground or walls) before reaching the receiving antenna, thus creating false or misaligned targets. Detection methods typically require combining target trajectory analysis, historical correlation, and consistency verification based on geometric models.
[0041] In some embodiments of the present invention, step S101 performs dynamic state assessment on the raw sensing data to obtain specific implementations of environmental state parameters and sensor quality assessment indicators, which vary depending on the type of sensor, thereby achieving fine-grained diagnosis of data quality.
[0042] For camera sensors, the evaluation primarily focuses on the quality and integrity of the image data. Image sharpness evaluation aims to quantify the degree of blur in the image.
[0043] As a preferred approach, this embodiment achieves this by calculating the gradient magnitude or frequency domain features of the image. For example, when the lens is covered with water droplets due to heavy rain, the evaluation module will output a lower sharpness score. The evaluation of occlusion areas is used to identify areas of the lens partially covered by mud, water, or stains. This can be achieved through semantic segmentation networks or background subtraction algorithms, thereby generating a mask map that identifies invalid pixel areas. The evaluation of glare intensity is used to quantify the impact of overexposure areas caused by oncoming headlights at night or water reflections. This can be achieved by detecting changes in the area and saturation of high-brightness areas in the image. The beneficial effect of this evaluation is that it can accurately identify whether the quality degradation of camera data is caused by global weather (such as heavy fog) or localized dirt (such as mud spots), providing a basis for subsequent differentiated processing.
[0044] For lidar sensors, the evaluation focuses on the reliability and validity of point cloud data. The assessment of point cloud noise ratio is used to count invalid scattering points caused by reflections from particles such as raindrops and snowflakes. A common calculation method is to analyze statistical outliers in the distance to neighboring points. The assessment of effective point cloud density reflects the number of points usable to represent real objects per unit volume. This value decreases significantly in dense fog, and its calculation typically relies on voxelization statistics. Analysis of echo intensity and its attenuation rate reflects the laser's penetrating power in the medium; fog and haze cause echo intensity to decrease exponentially with distance. The beneficial effect of this part of the evaluation is that it quantifies the real impact of weather on lidar detection capabilities at a physical level, rather than simply judging it as a failure, allowing the system to know at what detection distance the data remains reliable.
[0045] For millimeter-wave radar sensors, the evaluation focuses on signal-level reliability and interference. Signal-to-noise ratio (SNR) is a core indicator of radar target echo signal quality, calculated through spectral analysis of the received signal. A low SNR means the target may be overwhelmed by severe weather noise. Multipath interference intensity assessment is used to detect false targets caused by ground or guardrail reflections. One approach is to identify these false targets by analyzing the discrepancies between the target's dynamic trajectory and its physical motion. This effectively distinguishes between genuine targets and false alarms caused by weather noise or complex reflections, improving data reliability at the signal source.
[0046] Through the parallel evaluation process described above, this embodiment can output a comprehensive diagnostic report that includes not only macro-environmental qualitative judgments (such as heavy rain) but also a series of fine-grained quality indicators. This transforms the black-box sensor data in traditional solutions into feature information with clear quality labels and credibility annotations, providing accurate data support for subsequent dynamic intelligent fusion decision-making.
[0047] In some embodiments of the present invention, such as Figure 2As shown, step 102 generates the dynamic confidence weights for each sensor based on environmental state parameters and sensor quality assessment indicators, including: S201. Input the environmental state parameters and sensor quality assessment indicators into the confidence calculation function corresponding to the sensor type to obtain the initial confidence weight; S202. Perform time-smoothing filtering on the initial confidence weights to obtain dynamic confidence weights.
[0048] The confidence calculation function refers to a mathematical model or algorithm rule that maps multiple input variables to a single output value, used to comprehensively evaluate the reliability of sensor data. Its specific implementation can be a linear weighted function, a rule-based conditional decision tree, or a trained small neural network, etc.
[0049] Time smoothing filtering refers to a signal processing method that uses signal processing techniques to smooth time series data in order to suppress high-frequency noise or outliers. Specific implementation methods in this embodiment include, but are not limited to: moving average filtering, exponential smoothing filtering, and low-pass digital filtering (such as a first-order IIR filter).
[0050] In some embodiments of the present invention, the specific implementation of generating dynamic confidence weights for each sensor includes two ordered processing stages.
[0051] The first stage involves the application of a confidence calculation function. This function is a mathematical mapping model specifically designed for different sensor types, and its inputs are the environmental state parameters and sensor quality evaluation indicators obtained in the preceding steps.
[0052] Specifically, for camera sensors, the function input typically includes image sharpness score, pixel-level occlusion rate, ambient visibility estimate, and glare intensity; for LiDAR sensors, the input focuses on point cloud noise ratio, effective point cloud density, and echo intensity attenuation rate related to detection distance; and for millimeter-wave radar sensors, the input mainly relies on signal-to-noise ratio, multipath interference detection intensity, and target trajectory continuity. The function output is an initial confidence weight between 0 and 1, whose value directly reflects the instantaneous reliability of the sensor data at the current moment without considering historical frames.
[0053] By customizing computational models for each sensor, multi-dimensional heterogeneous physical indicators are accurately fused and transformed into a unified reliability quantification value, providing a direct numerical basis for achieving refined fusion weight allocation.
[0054] In one example, taking a camera as an example, its dynamic confidence level It can be represented as:
[0055] in, For clarity, For occlusion rate, For visibility, Glare intensity, sharpness, occlusion rate, and visibility are derived from environmental condition parameters and sensor quality assessment indicators, and are outputs of real-time evaluation. , , , This is an adjustable coefficient. f This is the normalization function.
[0056] It should be noted that different sensors have different failure mechanisms, so the evaluation metrics for their confidence function inputs are different.
[0057] The second stage involves time-smoothing filtering of the initial confidence weights. This process aims to suppress random fluctuations that may arise from single-frame data evaluation, ensuring the temporal stability and continuity of the final dynamic confidence weights. One possible implementation is to use a first-order infinite impulse response filter, such as a weighted average of the initial weights of the current frame and the dynamic weights output from the previous frame. This method effectively filters out abnormal jumps in confidence caused by transient interference (such as a bird flying by or an isolated splash of mud in a single frame of camera footage), preventing the fusion strategy from making unnecessary and frequent drastic adjustments due to such transient noise, thereby ensuring the smoothness of the perception system's decision output and the consistency of the driving experience.
[0058] This embodiment achieves the dynamic confidence weights by combining instantaneous calculation based on physical indicators with time smoothing incorporating historical information. This results in a dynamic confidence weight that is both sensitive to changes in the real environment and maintains the stability required for decision-making.
[0059] In some embodiments of the present invention, the preset sensor failure condition is: the dynamic confidence weight of at least one sensor is continuously lower than a preset threshold for a predetermined duration.
[0060] The preset threshold and the predetermined duration can be set according to the performance of different sensors and the actual application needs, and there are no restrictions here.
[0061] In some embodiments of the present invention, such as Figure 3 As shown, step 103, generating virtual sensing data based on multi-source information, includes: S301. Sensors whose dynamic confidence weight is continuously lower than a preset threshold for a predetermined time are considered as failed sensors. S302. Based on the type of the failed sensor, call the corresponding physical or geometric model; S303. Input the data from other high-confidence sensors and environmental state parameters into the called model, and generate virtual sensing data through cross-modal reasoning.
[0062] In this context, a physical model or geometric model refers to a mathematical and computational model used to describe the physical processes of a specific sensor's perception or the geometric relationships of data generation. In this embodiment, its specific form depends on the sensor being simulated. For example, a model simulating a camera might involve a 3D-to-2D perspective projection model and a light scattering model; a model simulating a lidar might include a laser beam propagation attenuation model, a target surface reflectivity model, and a point cloud generation geometric model; and a model simulating millimeter-wave radar might involve an electromagnetic wave propagation model and a radar cross-section model.
[0063] Cross-modal reasoning refers to the technical process of using information from one or more perception modalities (such as vision, point clouds, and radio frequency signals) to infer or generate information from another perception modality after processing and transformation. Its core lies in establishing semantic or geometric relationships between data from different modalities. For example, by combining the target distance and orientation information detected by millimeter-wave radar with prior knowledge of the target category (such as a vehicle), geometric reasoning can be used to generate the point cloud distribution that the target might appear in in the lidar coordinate system.
[0064] In this embodiment, the virtual sensing data is generated through algorithmic simulation, mimicking the format and attributes of real physical sensor output data. Its generation is not based on the acquisition of real physical signals, but rather on computational synthesis from other information sources and models. To accurately characterize its uncertainty as inferred information, expected noise can be actively injected or confidence downsampling can be applied during or after generation.
[0065] In some embodiments of the present invention, the preset sensor failure condition is explicitly defined as follows: the dynamic confidence weight of at least one sensor remains below its corresponding preset threshold for a predetermined period of time. By setting this condition, the duration is introduced as a judgment dimension, effectively distinguishing between transient performance fluctuations caused by instantaneous, sporadic interference (such as a bird flying by in a single frame or an isolated splash of water) and substantial functional failures caused by persistent severe weather or sensor hardware problems.
[0066] For example, the predetermined duration can be set to 3 to 5 processing frame cycles (approximately 100-200 milliseconds). This duration is usually shorter than the upper limit of the safety response time of the downstream vehicle control module, while being significantly longer than the duration of typical transient interference. This ensures that the system can respond to real faults in a timely manner, while minimizing the risk of falsely triggering the compensation mechanism due to transient interference, thus guaranteeing the stability of the system operation.
[0067] After determining that the above failure conditions are met, the virtual perception data generation process is initiated.
[0068] Specifically, the system first marks sensors that meet the failure criteria as failed sensors. Then, based on the specific type of the failed sensor (e.g., camera, LiDAR, or millimeter-wave radar), it retrieves the corresponding pre-built physical or geometric model from the model library. These models encapsulate the physical operating principles and data generation rules of the specific sensor. Finally, the system inputs data from other sensors that still maintain high confidence at the current moment (e.g., a list of available high-confidence millimeter-wave radar targets when LiDAR fails) and real-time environmental state parameters (e.g., estimated visibility and attenuation coefficients) into the retrieved model.
[0069] Through cross-modal inference computation, the model outputs virtual perception data that simulates the failed sensor in terms of data structure and features. For example, if the failed camera is a forward-facing camera, the system may call an image rendering model to fuse information from side cameras, 3D scene reconstruction data from LiDAR, and map semantics to generate virtual image patches for the occluded areas; if the failed camera is LiDAR, it may call a geometric projection and scattering model to generate sparse virtual point clouds at the corresponding spatial locations based on the stable target trajectory and vehicle outline priors from millimeter-wave radar.
[0070] In addition to traditional hardware redundancy, this embodiment constructs an information redundancy path based on algorithms and knowledge. When critical sensors temporarily fail due to extreme weather, information is proactively reconstructed and supplemented based on interpretable physical laws. This provides the intelligent driving system with valuable reaction time to maintain basic environmental awareness under extreme conditions, achieving a smooth transition from functional interruption to moderate performance degradation, which is beneficial to improving driving safety.
[0071] In some embodiments of the present invention, such as Figure 4 As shown, data from other high-confidence sensors and environmental state parameters are input into the called model, and virtual sensing data is generated through cross-modal inference, including: S401. When the failed sensor is a lidar, generate virtual point cloud data based on the millimeter-wave radar target information and geometric projection model. S402. When the failed sensor is a camera, virtual image data is generated based on the 3D scene information and image rendering model of other sensors.
[0072] In this context, the geometric projection model refers to a mathematical model used to describe how the geometric shape of an object in three-dimensional space is observed by a sensor (specifically, LiDAR) and converted into a specific data format (point cloud). In the scenario of generating a virtual LiDAR point cloud, the core is to simulate the intersection calculation between the laser beam and the three-dimensional geometric model (such as a standard vehicle body box), and to determine the spatial density and distribution of the point cloud based on prior knowledge such as the radar cross-section.
[0073] In computer graphics, the image rendering model refers to a series of algorithms and processes that convert a 3D scene description into a 2D image. In the context of generating virtual camera images in this embodiment, the process typically includes 3D scene reconstruction (or retrieval), viewpoint setting, geometric projection (transforming 3D coordinates to a 2D image plane), and simple shading or texture mapping. The goal is to generate an image with correct geometric and semantic layout, rather than highly realistic textures.
[0074] Among them, 3D scene information refers to a 3D digital description of the vehicle's surrounding environment obtained through sensor perception and calculation. Its specific forms can include: a 3D point set directly composed of LiDAR point clouds, a textured 3D mesh model generated by point cloud or visual SLAM algorithms, or an environment occupancy grid map formed by multi-sensor fusion in units of cubes or grids.
[0075] In some embodiments of the present invention, the generation logic of virtual sensing data is executed differently according to the specific type of the failed sensor, thereby achieving targeted compensation for the lack of a specific sensing dimension.
[0076] When the system determines that the failed sensor is a lidar, the virtual data generation process mainly relies on the geometric projection model.
[0077] Specifically, the process uses current high-confidence millimeter-wave radar sensor data as its core input. Millimeter-wave radar typically provides a stable list of targets, including status information such as target range, azimuth, and radial velocity. The generation module combines this target information with a prior geometric model of the contours of common traffic participants such as vehicles and pedestrians, and reconstructs them in three-dimensional space using a geometric projection model.
[0078] For example, for a vehicle target being stably tracked by millimeter-wave radar, the generation module assumes it has a standard 3D bounding box model of a car or truck and places the model in the world coordinate system based on the distance and orientation measured by the radar. Subsequently, the model simulates the scanning line mechanism of the lidar, calculates the laser reflection points that may be generated on the surface of these virtual 3D models, and finally synthesizes a sparse but reasonably spatially structured virtual point cloud data.
[0079] This method can restore the basic geometric outline and position of key obstacles ahead by taking advantage of the strong penetrating power of millimeter-wave radar when lidar fails completely due to dense fog or heavy rain. Although the point cloud is sparse and based on speculation, it is sufficient to support collision warning and obstacle avoidance decision-making in emergency situations.
[0080] When the system determines that the failed sensor is the camera, the virtual data generation process focuses on the image rendering model. The input to this process is more diverse, typically incorporating 3D scene information constructed from other compliant sensors.
[0081] For example, real-time 3D scene reconstruction can be performed using point clouds from still-functioning LiDAR cameras, or multi-view geometric reasoning can be performed by combining images from other perspectives (such as side-view and rear-view cameras). Simultaneously, static scene semantics provided by high-precision maps can be incorporated. Based on this information, the generation module synthesizes corresponding virtual image data from the perspectives that the failed cameras should have observed, using an image rendering model.
[0082] In practice, the 3D scene information (such as a mesh reconstructed from point clouds or a simplified geometry) can be rasterized first, and basic texture or semantic tags (such as road surface, vehicles, guardrails) can be attached. Then, a 2D image patch can be generated through the projection transformation of a virtual camera.
[0083] It should be understood that when the main forward-facing camera is completely covered by mud or blinded by strong light, this method can predict the general outline of the scene that the camera should see based on the understanding of the environment by other sensors. Although it is lacking in texture details, it can provide the system with crucial semantic information such as lane lines, drivable areas and the location of large obstacles, preventing the vehicle from being paralyzed due to complete visual interruption.
[0084] This embodiment achieves both accuracy and rationality in redundancy compensation by customizing differentiated virtual data generation paths for different types of sensors. This allows the perception system to recover specific modal environmental information when some sensors fail, thereby maximizing the integrity of the system's understanding of the environment and improving the safety of autonomous driving.
[0085] In some embodiments of the present invention, adaptive fusion of data from different sensors based on dynamic confidence weights includes at least one of the following processing levels: At the data preprocessing level, data is enhanced or removed based on dynamic confidence weights; At the feature fusion level, features extracted from different sensors are weighted and fused based on dynamic confidence weights. At the target fusion level, targets detected by different sensors are associated and tracked based on dynamic confidence weights.
[0086] In some embodiments of the present invention, the auxiliary weight is preset to a fixed value that is lower than the normal sensor dynamic confidence weight range, or is preset to a low value range.
[0087] The data preprocessing level refers to the processing stage that operates on the raw sensor data or primary processed data. Specific technical methods include, but are not limited to: image processing algorithms for dehazing, deraining, and contrast enhancement; point cloud processing for statistical outlier removal and voxel-based downsampling; and signal processing for radar noise filtering.
[0088] The feature fusion layer refers to the stage in machine learning or deep learning models where feature representations extracted from different data sources or different network layers are integrated. Specific implementation mechanisms include feature map concatenation, weighted summation, and feature reweighting based on attention mechanisms (such as channel attention and spatial attention).
[0089] The target fusion layer refers to the stage after target detection, where results from multiple independent detectors are processed uniformly. This includes data association (matching detection boxes from different sources to the same real object) and state estimation (fusion of multi-source observations to update the target's position, velocity, and other states). Commonly used methods include the Hungarian algorithm, joint probabilistic data association, and trajectory fusion based on Kalman filtering or particle filtering.
[0090] At the data preprocessing level, targeted enhancement or removal operations are performed based on the dynamic confidence weights of each sensor data point or its specific region. For example, for camera image regions with low confidence, image dehazing or deraining algorithms based on physical models may be triggered for enhancement; for local point sets in LiDAR point clouds that are evaluated as having a high noise ratio (corresponding to low confidence), they may be directly filtered and removed. This approach allows for the repair or cleaning of low-quality data at the beginning of the fusion process, effectively reducing the propagation of noise and interference to subsequent advanced processing modules and improving the quality of input data.
[0091] At the feature fusion level, adaptive fusion is applied to abstract features extracted from different sensors. Specifically, in neural networks for object detection, when multi-sensor feature concatenation or attention calculation is required, feature maps from high-confidence sensors are given higher fusion weights.
[0092] For example, before feature stitching, a channel attention weight vector is generated based on the real-time confidence of each sensor, and the feature channels from different sensors are recalibrated so that the network automatically focuses on more reliable information sources during training and inference.
[0093] This approach enables intelligent and refined fusion decision-making, allowing the perception model to dynamically adjust its internal information utilization strategy based on real-time changes in data quality, thereby optimizing the accuracy and robustness of the fusion results at the feature level.
[0094] At the target fusion level, adaptive fusion is reflected in the process of associating and unifying independent detection results (i.e., bounding boxes, categories, and trajectories) from different sensors. When performing target association matching or trajectory updates, the confidence level of each sensor's detection results is used as a key reference.
[0095] For example, when updating the target state using Kalman filtering, observations from high-confidence sensors will be given a larger innovative covariance, thus having a greater impact on the final fused trajectory.
[0096] This approach ensures that the final target list output to the downstream planning module comprehensively reflects the most reliable observation information from all sensors, improves the reliability of system-level sensing output, and provides a more accurate environmental situation assessment for planning and control.
[0097] To better implement the multi-sensor fusion method for intelligent driving in this embodiment of the invention, based on the multi-sensor fusion method for intelligent driving, correspondingly, as follows: Figure 5 As shown, this embodiment of the invention also provides a multi-sensor fusion system for intelligent driving. The multi-sensor fusion system 500 for intelligent driving includes: The data acquisition module 501 is used to acquire raw sensing data from at least two types of sensors and to perform dynamic state evaluation on the raw sensing data to obtain environmental state parameters and sensor quality evaluation indicators. The weight generation module 502 is used to generate dynamic confidence weights for each sensor based on the environmental state parameters and sensor quality evaluation indicators. The adaptive fusion module 503 is used to adaptively fuse data from different sensors according to the dynamic confidence weight, and output the fused perception result. During the fusion process, when the preset sensor failure condition is met, virtual perception data is generated based on multi-source information, and the virtual perception data is integrated into the adaptive fusion process with auxiliary weights.
[0098] The multi-sensor fusion system 500 for intelligent driving provided in the above embodiments can realize the technical solutions described in the above embodiments of the multi-sensor fusion method for intelligent driving. The specific implementation principles of each module or unit can be found in the corresponding content in the above embodiments of the multi-sensor fusion method for intelligent driving, and will not be repeated here.
[0099] This application also provides a vehicle including the above-described multi-sensor fusion system for intelligent driving.
[0100] Accordingly, this application also provides a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, they can implement the steps or functions of the multi-sensor fusion method for intelligent driving provided in the above-described method embodiments.
[0101] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0102] The multi-sensor fusion method, system, and vehicle for intelligent driving provided by the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A multi-sensor fusion method for intelligent driving, characterized in that, include: Acquire raw sensing data from at least two types of sensors, and perform dynamic state assessment on the raw sensing data to obtain environmental state parameters and sensor quality assessment indicators. Based on the environmental state parameters and sensor quality evaluation indicators, dynamic confidence weights for each sensor are generated. Based on the dynamic confidence weight, data from different sensors are adaptively fused to output fused perception results. During the fusion process, when the preset sensor failure conditions are met, virtual perception data is generated based on multi-source information, and the virtual perception data is incorporated into the adaptive fusion process with auxiliary weights.
2. The method according to claim 1, characterized in that, The dynamic state assessment of the raw sensing data to obtain environmental state parameters and sensor quality assessment indicators includes: For camera data, evaluate at least one of image sharpness, occlusion area, and glare intensity; For lidar data, evaluate at least one of the following: point cloud noise ratio, effective point cloud density, and echo intensity. For millimeter-wave radar data, evaluate at least one of the following: signal-to-noise ratio and multipath interference intensity.
3. The method according to claim 1 or 2, characterized in that, The process of generating dynamic confidence weights for each sensor based on the environmental state parameters and sensor quality assessment indicators includes: The environmental state parameters and the sensor quality assessment indicators are input into the confidence calculation function corresponding to the sensor type to obtain the initial confidence weights; The initial confidence weights are subjected to time-smoothing filtering to obtain the dynamic confidence weights.
4. The method according to claim 1, characterized in that, The preset sensor failure condition is: the dynamic confidence weight of at least one sensor is continuously lower than a preset threshold for a predetermined period of time.
5. The method according to claim 4, characterized in that, The virtual sensing data generated based on multi-source information includes: Sensors whose dynamic confidence weight remains below a preset threshold for a predetermined duration are considered to be faulty sensors. Based on the type of the failed sensor, call the corresponding physical or geometric model; The data from other high-confidence sensors and the environmental state parameters are input into the called model, and virtual sensing data is generated through cross-modal inference.
6. The method according to claim 5, characterized in that, The step of inputting data from other high-confidence sensors and the environmental state parameters into the invoked model to generate virtual perception data through cross-modal inference includes: When the failed sensor is a lidar, virtual point cloud data is generated based on the millimeter-wave radar target information and geometric projection model; When the failed sensor is a camera, virtual image data is generated based on the 3D scene information and image rendering model from other sensors.
7. The method according to claim 1, characterized in that, The adaptive fusion of data from different sensors based on the dynamic confidence weights includes at least one of the following processing levels: At the data preprocessing level, data is enhanced or removed based on the dynamic confidence weights. At the feature fusion level, features extracted from different sensors are weighted and fused according to the dynamic confidence weights. At the target fusion level, targets detected by different sensors are associated and tracked based on the dynamic confidence weights.
8. The method according to claim 1, characterized in that, The auxiliary weight is preset to a fixed value that is lower than the normal sensor dynamic confidence weight range, or is preset to a low value range.
9. A multi-sensor fusion system for intelligent driving, characterized in that, include: The data acquisition module is used to acquire raw sensing data from at least two types of sensors and to perform dynamic state evaluation on the raw sensing data to obtain environmental state parameters and sensor quality evaluation indicators. The weight generation module is used to generate dynamic confidence weights for each sensor based on the environmental state parameters and sensor quality evaluation indicators. The adaptive fusion module is used to adaptively fuse data from different sensors based on the dynamic confidence weights and output fused perception results. During the fusion process, when a preset sensor failure condition is met, virtual perception data is generated based on multi-source information, and the virtual perception data is integrated into the adaptive fusion process with auxiliary weights.
10. A vehicle, characterized in that, Including the multi-sensor fusion system for intelligent driving as described in claim 9.