Self-repairing method and system for multi-modal sensor fusion, electronic equipment and medium
By performing health scores and generative adversarial network compensation on the sensors, dynamic allocation of fusion weights, the problem of perceived performance degradation caused by sensor failure in autonomous driving systems is solved, and higher perceived reliability and stability are achieved.
Patent Information
- Application Number
- CN202510636056.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
AI Technical Summary
The sensor perception performance of existing autonomous driving systems in complex environments is degraded, making it difficult to dynamically deal with sensor failures, resulting in a high perception error rate, affecting system stability and safety.
By obtaining the physical and logical layer health scores of multiple heterogeneous sensors, the sensor failure is determined, and a generative adversarial network is used to compensate cross-modal data, dynamically allocate fusion weights, and output perceptual results in combination with the data fusion algorithm.
It improves the perceived reliability and stability of the autonomous driving system in complex environments, significantly reduces the perceived error rate, and enhances the system's adaptability in extreme scenarios.
Smart Images

Figure CN120541772A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a self-repair method, system, electronic device and medium for multimodal sensor fusion. Background Art
[0002] The development of autonomous driving technology relies on multimodal sensor fusion for accurate environmental perception. Currently, autonomous driving systems typically employ a variety of sensors (such as lidar, cameras, and millimeter-wave radar) to work together. The collected data is fused and processed using algorithms such as weighted averaging, Kalman filtering, or Bayesian networks to improve the accuracy and robustness of environmental perception. In hardware design, a multi-sensor redundant deployment strategy is often adopted, such as installing multiple sensors of the same type to prevent single points of failure. Furthermore, the system combines real-time signal strength analysis and multi-sensor data consistency verification to dynamically monitor sensor status. To address sensor failure scenarios, relevant technologies have explored single-mode repair methods based on historical data interpolation, generative adversarial networks, and dynamic weighting of confidence levels. These technologies, combined with edge computing platforms, enable lightweight model deployment to meet the real-time processing requirements of onboard environments. Furthermore, some relevant solutions utilize reinforcement learning to optimize fusion strategies, enhancing the system's adaptability in complex scenarios.
[0003] However, related technologies still have many limitations. Traditional data fusion methods mainly rely on static redundant designs and fixed weight distribution patterns, showing obvious lack of adaptability in complex environments. When encountering extreme weather or special scenarios, the perception performance of multi-sensor systems will be significantly reduced, making it difficult to dynamically respond to sudden sensor failures. Data compensation methods in related technologies are mostly focused on interpolation and repair within a single mode, with high related compensation errors, especially in dynamic environments. The performance is insufficient, resulting in a dynamic scene compensation error rate exceeding 30%, posing a challenge to the stability and safety of autonomous driving systems in complex environments such as severe weather and temporary occlusion. Summary of the Invention
[0004] Based on this, it is necessary to provide a self-repair method, system, electronic device and medium for multimodal sensor fusion that can significantly improve the reliability of multimodal sensor dynamic fusion perception in response to the above technical problems.
[0005] To achieve the above objectives, a first embodiment of the present invention provides a self-repair method for multimodal sensor fusion, comprising:
[0006] Acquire raw data from a plurality of heterogeneous sensors, and generate a physical layer health score and a logical layer health score for the heterogeneous sensors based on the raw data;
[0007] Determining whether the heterogeneous sensor is failed by combining the physical layer health score and the logical layer health score;
[0008] When it is determined that a failed heterogeneous sensor exists, inputting valid raw data into a generative adversarial network for cross-modal generation, and outputting first virtual data of the failed heterogeneous sensor; the valid raw data represents raw data from the valid heterogeneous sensor;
[0009] Dynamically allocating fusion weight coefficients of the heterogeneous sensors according to the comprehensive health scores of the heterogeneous sensors; the comprehensive health score is calculated by weighting the physical layer health score and the logical layer health score;
[0010] Based on the fusion weight coefficient, a data fusion algorithm is used to fuse the valid original data and the first virtual data, and a fusion perception result is output.
[0011] The second embodiment of the present invention proposes a multi-modal sensor fusion self-repair system.
[0012] Sensor modules for configuring heterogeneous sensor arrays;
[0013] a sensor health monitoring module configured to acquire raw data from a plurality of heterogeneous sensors and generate a physical layer health score and a logical layer health score for the heterogeneous sensors based on the raw data; and determine whether the heterogeneous sensors have failed based on the physical layer health score and the logical layer health score;
[0014] a cross-modal data compensation module for inputting valid raw data into a generative adversarial network to perform cross-modal generation upon determining the presence of a failed heterogeneous sensor, and outputting first virtual data of the failed heterogeneous sensor; the valid raw data representing raw data from the valid heterogeneous sensor;
[0015] A dynamic weighted data fusion module is configured to dynamically assign fusion weight coefficients to the heterogeneous sensors based on their comprehensive health scores, where the comprehensive health scores are calculated by weighting the physical layer health scores and the logical layer health scores. Based on the fusion weight coefficients, a data fusion algorithm is used to fuse the valid original data and the first virtual data, and output a fused perception result.
[0016] The third aspect of the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the self-repair method of multimodal sensor fusion when executing the computer program.
[0017] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, the steps of the self-repair method of multimodal sensor fusion are implemented.
[0018] The self-repairing method, system, electronic device, and medium for multimodal sensor fusion described above utilize a dual-path health scoring mechanism at the physical and logical layers to comprehensively assess sensor status, making judgments more comprehensive and reliable. For failed sensors, a generative adversarial network is used to achieve cross-modal data compensation, generating corresponding virtual data based on multiple valid modalities to improve compensation accuracy. Dynamically assigning fusion weights based on health scores enhances fusion stability. Finally, a fusion algorithm fuses valid raw data and virtual data to produce reliable perception results, significantly improving the autonomous driving system's environmental perception capabilities in complex or abnormal scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 1 is a flow chart of a self-repair method for multimodal sensor fusion in one embodiment;
[0020] Figure 2 1 is a flow chart illustrating the reasoning process of a generative adversarial network in one embodiment;
[0021] Figure 3 1 is a flow chart of data fusion based on the Kalman filter algorithm in one embodiment;
[0022] Figure 4 1 is a flow chart of data fusion based on the Kalman filter algorithm in one embodiment;
[0023] Figure 5 A schematic diagram of a process flow for optimizing a generative adversarial network in one embodiment;
[0024] Figure 6 A multimodal sensor fault self-repair and dynamic fusion process in one embodiment;
[0025] Figure 7 The figure is a structural block diagram of a self-repairing system for multimodal sensor fusion of a water dispenser in one embodiment. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0027] The following describes in detail the implementation details of the technical solutions of the embodiments of the present application.
[0028] In one embodiment, Figure 1 As shown, a flow chart of a self-repair method for multimodal sensor fusion is provided, which may include the following steps:
[0029] Step S101 : acquiring raw data from a plurality of heterogeneous sensors, and generating a physical layer health score and a logical layer health score of the heterogeneous sensors based on the raw data.
[0030] In order to achieve efficient collection and redundant coverage of multimodal data, it is necessary to configure a heterogeneous sensor array consisting of multiple different types of sensor groups to simultaneously collect multiple types of environment or object information. For example, the heterogeneous sensor array may include lidar, camera, millimeter-wave radar, ultrasonic sensor, and inertial measurement unit (IMU), where the lidar obtains three-dimensional point cloud, the camera collects image information, the millimeter-wave radar is used to detect the distance and speed of obstacles, the ultrasonic sensor is used for close-range distance perception, and the IMU is mainly used to measure acceleration and angular velocity. These different types of sensors differ in perception mechanism, data type, observation accuracy or environmental adaptability, and the data they collect have different physical properties and perception dimensions, such as images, point clouds, distance information, thermal imaging, acceleration data, etc. These data are collectively referred to as multimodal data. Through multi-source complementarity, the sensor array can improve the overall perception robustness of the system in complex environments.
[0031] In practical applications, heterogeneous sensors are prone to performance degradation or failure in complex environments due to their physical characteristics. To evaluate the operating status of each heterogeneous sensor, this embodiment obtains raw data from multiple heterogeneous sensors and calculates health scores for the heterogeneous sensors in two dimensions based on the raw data.
[0032] The physical layer health score is used to measure the quality stability of the sensor at the physical signal level. It can be measured using physically measurable indicators such as image clarity, point cloud density, and noise level. In actual applications, different sensors use different indicator systems, specifically:
[0033] For LiDAR, the point cloud density and signal-to-noise ratio are comprehensively evaluated to obtain the physical layer health score of the LiDAR. The point cloud density is calculated by counting the valid points within the unit solid angle, and the effect of field of view occlusion is taken into account. The corresponding scoring formula is:
[0034]
[0035] In the above formula, N valid Indicates the number of points with effective reflection intensity (such as ≥30) per unit time; N totalIndicates the theoretical maximum number of points of the lidar, determined based on the frame rate and scanning mode; FOV current Indicates the effective field of view of the current non-blocked area; FOV max Indicates the maximum field of view of the lidar under ideal conditions.
[0036] The signal-to-noise ratio of the lidar is calculated by the mean and standard deviation of the echo intensity. The corresponding scoring formula is:
[0037]
[0038] In the above formula, μ signal represents the average intensity of the reflected signal; σ noise Represents the standard deviation of the noise signal.
[0039] Finally, by weighted fusion of the point cloud density score and the signal-to-noise ratio score, the physical layer health score of the lidar is obtained, specifically:
[0040] HS lidar =α1·S lidar +(1-α1)·SNR lidar
[0041] Among them, α1 is the weight coefficient of point cloud density score, which ranges from 0 to 1.
[0042] For cameras, an image clarity assessment method is used to obtain the physical layer health score of the camera. The Brenner gradient function (Brenner) can be used to calculate the blurriness of the image captured by the camera, so the image clarity assessment score formula is:
[0043]
[0044] In the above formula, I(x,y) represents the grayscale value of the image at the pixel position (x,y); W and H represent the width and height of the image.
[0045] The image clarity assessment score is normalized to obtain the physical layer health score of the camera, which is:
[0046]
[0047] In the above formula, T blur =5×10 6 , when G Brenner A larger value indicates a blurrier image and a corresponding lower physical layer health score.
[0048] For millimeter-wave radars, the target stability of the millimeter-wave radar is evaluated to obtain the millimeter-wave radar health score. The target stability evaluation of the millimeter-wave radar can be calculated using target stability matching, which counts the target matching rate of the millimeter-wave radar within 10 consecutive frames. Specifically,
[0049]
[0050] Among them, N total_targets Indicates the total number of detected targets at present; N matched Indicates the number of successfully matched targets in multiple consecutive frames.
[0051] Based on this, the calculation of the physical layer health scores of different heterogeneous sensors is completed.
[0052] The logical health score determines the credibility of sensor data from the information output level. It can be comprehensively evaluated based on the matching relationship between sensor data outputs, multi-sensor collaborative reasoning capabilities, and system-level algorithm performance. Specific indicators may include:
[0053] (1) Cross-modal data consistency, which is used to check the degree of difference in the detection results (such as position, speed, and category) of the same target by different sensors. It is specifically expressed as:
[0054]
[0055] (2) Algorithm reliability. This is calculated by combining anomaly detection rate and classification confidence. The anomaly detection rate is used to measure the ability of the fusion algorithm to identify erroneous input data, such as the proportion of abnormal points eliminated by the filtering algorithm; the classification confidence represents the average confidence score of the target classification (such as vehicles and pedestrians). Based on this, the algorithm reliability is specifically expressed as:
[0056] Reliability = 0.5 classification confidence + 0.5 anomaly detection rate
[0057] (3) Time synchronization, which reflects the accuracy of time alignment between sensors, can be measured using timestamp deviation. Timestamp deviation refers to the time alignment error of multi-sensor data. For example, a delay between the lidar and camera frames of less than 10ms is considered normal. Based on this, time synchronization is specifically expressed as:
[0058]
[0059] (4) Logical conflict detection is used to measure whether multiple sensors have consistent judgments on the same scene. This can be measured using the conflicting decision ratio, which represents the conflicting rate of decisions made by multiple sensors on the same scene. For example, a camera detects a pedestrian, but a lidar does not. The conflicting decision ratio is specifically expressed as:
[0060]
[0061] Based on this, the logical layer health score (LLHS) of heterogeneous sensors is obtained by weighted fusion of the above different logical layer indicators, specifically:
[0062] LLHS=w′1·Consistency+w′2·Reliability+w′3·SyncScore+w′4·ConflictScore
[0063] In practical applications, the weight w′ i It can be dynamically adjusted according to the task stage or environment. For example, in highly dynamic scenarios (such as high-speed driving), the weights corresponding to time synchronization and data consistency can be increased; while in complex environments (such as dense obstacles), the weights corresponding to algorithm reliability and logical conflict detection can be increased. In practical applications, the weight w′ i The specific calculation can be:
[0064]
[0065] Based on this, the physical layer health score and logical layer health score of each heterogeneous sensor can be output, thereby more comprehensively and dynamically perceiving the operating status of the sensor.
[0066] Step S102 : Determine whether the heterogeneous sensor is failed by combining the physical layer health score and the logical layer health score.
[0067] Combining the physical and logical layer health scores, a comprehensive assessment is made to determine whether each sensor is in a failed state. This failure not only refers to a complete cessation of sensor hardware operation (e.g., power outage or damage), but also includes situations such as significant performance degradation, severe distortion of observed information, or significant inconsistency with other sensor outputs.
[0068] In the application scenario of autonomous driving, when a sensor is determined to be failed by the system, the raw data collected by the sensor will no longer be directly used in the subsequent perception fusion process, and its data path will be marked as unavailable.
[0069] It should be noted that a sensor that is judged to be failed may still be physically powered on and continue to output data, but due to the insufficient health score of the failed sensor, the output of the failed sensor is regarded as unusable data to avoid misleading information interfering with the perception decision-making process.
[0070] In one embodiment, after obtaining the physical layer health score and the logical layer health score of each heterogeneous sensor, the health status of each heterogeneous sensor is evaluated using a preset determination rule, thereby identifying a sensor that is currently in a failed state.
[0071] In practical applications, sensors may experience malfunctions due to a variety of reasons. The physical layer health score primarily reflects the stability of the sensor's signal quality. An abnormal physical layer health score typically indicates a problem with the sensor's hardware. For example, a significantly decreased physical layer health score (e.g., below 0.4) often indicates a hard fault such as signal distortion, significantly increased observation noise, or unstable power supply.
[0072] The logical health score measures the reliability of a sensor at the multimodal fusion semantic level, primarily assessing its observation consistency, reasoning accuracy, and temporal alignment with other sensors. A low logical health score for a sensor (e.g., below 0.5) may indicate an anomaly in collaborative perception, such as significant target positioning deviation, inconsistent data, or reduced confidence in algorithm output.
[0073] Based on this, a hierarchical determination strategy can be used to identify sensor failure states. When the sensor's physical layer health score is less than θ1 (e.g., θ1 = 0.4), the sensor is marked as "suspected hardware failure." When the sensor's logical layer health score is less than θ2 (e.g., θ2 = 0.5), the sensor is marked as "suspected logical failure."
[0074] In actual applications, if the physical layer health score does not fall below the judgment threshold but shows a continuous downward trend, it may also indicate that the sensor is experiencing aging or performance degradation due to environmental influences. For example, the point cloud density of a lidar gradually decreases in a high temperature or dusty environment, which is a common phenomenon of physical performance degradation.
[0075] In one determination method, based on a hierarchical determination strategy, to more accurately determine whether a sensor has failed, the physical layer health score and the logical layer health score can be combined to determine whether the sensor is in a failed state. When the physical layer health score is less than θ1 and the logical layer health score is less than θ2, the sensor is considered failed. This situation typically corresponds to the coexistence of hardware-level anomalies and fusion performance anomalies, which are highly characteristic of failure. Furthermore, when there is a significant separation between the physical layer health score and the logical layer health score, this is also valuable for determination. For example, if the physical layer health score is low but the logical layer health score is high, the sensor may have experienced a brief interference (such as a momentary signal loss), requiring continuous monitoring to determine whether the sensor has failed. On the other hand, when the physical layer health score is high but the logical layer health score is low, it usually indicates that the sensor itself is not faulty, but its observed data conflicts with the environment or other sensors. This is often caused by interference scenarios such as algorithm misjudgment, occlusion, and backlighting. Therefore, priority should be given to troubleshooting consistency issues with the algorithm model or data source.
[0076] In practical applications, to further improve judgment accuracy and anti-interference capabilities, a mechanism combining threshold triggering and time series analysis has been introduced. If the physical and logical layer health scores approach a boundary, or one is high and the other is low, their subsequent change trends are continuously observed. A single low score will trigger an alert, but it will not immediately determine that the sensor is invalid. If one of the scores remains low for multiple consecutive periods (e.g., below the threshold for five consecutive frames), it will also trigger a failure determination, effectively filtering out transient score fluctuations caused by occasional interference.
[0077] In another determination method, for different types of sensors, it also supports adaptive adjustment of the weights of the physical layer health score and the logical layer health score in determining sensor failure. The physical layer health score and the logical layer health score are further weighted to form a more comprehensive and sensitive comprehensive failure score of the sensor's current operating status, and based on this, it is determined whether the sensor is in a failure state. Among them, the comprehensive failure score of heterogeneous sensors is specifically expressed as:
[0078] FS=α2·(1-PLHS)+β1·(1-LLHS)
[0079] In the above formula, FS represents the comprehensive failure score of heterogeneous sensors, which is used to measure the overall failure degree of the sensor. A larger value indicates a higher failure risk. α2 represents the weight coefficient of the physical layer health score, which reflects the impact of physical faults on failure. β1 represents the weight coefficient of the logical layer health score, which reflects the impact of logical layer anomalies such as fusion anomalies on failure.
[0080] In practical applications, different weights are assigned to the physical layer health score and the logical layer health score depending on the type of sensor. For example, for hardware-sensitive sensors (such as lidar), priority is given to the proportion of the physical layer score in the overall judgment, that is, a higher weight is assigned to the physical layer health score. α2 = 0.7 and β1 = 0.3 can be set to enhance sensitivity to hardware failures. For sensors that rely more on algorithms and data fusion (such as cameras), the weight of the logical layer score is increased. α2 = 0.4 and β1 = 0.6 can be set to better identify problems such as fusion performance degradation. The configuration of the weight coefficients α2 and β1 can be preset in the system parameters, or dynamically and adaptively adjusted according to the environment to enhance the accuracy and adaptability of failure judgment.
[0081] After obtaining the comprehensive failure score, it is compared with a set third preset threshold (such as 0.6). When the comprehensive failure score of a sensor exceeds the threshold, the sensor is determined to be in a failed state.
[0082] In this embodiment, the failure judgment mechanism integrates threshold judgment, trend analysis and weighted scoring mechanism, and can perform strategy adaptation according to the actual operating environment, thereby accurately identifying different types of sensor failures.
[0083] Step S103 : When it is determined that there is a failed heterogeneous sensor, the valid original data is input into the generative adversarial network to perform cross-modal generation and output first virtual data of the failed heterogeneous sensor.
[0084] After one or more sensors are determined to be failed, the raw data collected from the heterogeneous sensor will no longer participate in data fusion. In order to make up for the perception blind spot caused by the lack of modal data, it is necessary to restore the modal data corresponding to the failed sensor through compensation to avoid deviations in the overall perception results. In this embodiment, a generative adversarial network is introduced to generate cross-modal data. The raw data from other valid heterogeneous sensors (i.e., non-failed heterogeneous sensors) is used as an effective information source and input into the generative adversarial network. The generative adversarial network performs a forward reasoning process to generate virtual data corresponding to the failed mode. The generated virtual data has structural consistency and semantic integrity and can be regarded as the compensation result of the failed sensor data. It will then participate in subsequent data fusion as one of the perception inputs, avoiding data interruption or local perception gaps caused by sensor failure.
[0085] Among them, generative adversarial networks (GANs) possess cross-modal generation capabilities and can support multiple modal combinations of input and output, such as generating point clouds from images and radar, generating infrared images from images, and generating images from radar, enabling complementary restoration of multimodal perception data. GANs automatically learn the semantic and geometric mapping relationships between modalities, enabling them to predict and reconstruct virtual data of completely different modalities based on valid data from other modalities. For example, raw data collected by cameras and radars is input into a GAN, which then generates point cloud data for lidar, achieving multimodal data complementarity. GANs provide a cross-modal mapping mechanism, breaking the limitations of intra-modal restoration and fully leveraging the complementary relationships between different perception sources, significantly improving the expressiveness and adaptability of the compensated data. For example, in extreme environments such as heavy rain, dust, and strong sunlight, multimodal data complementarity generation (e.g., generating virtual lidar point cloud data from millimeter-wave radar data) can compensate for optical sensor failures and reduce target detection error rates.
[0086] In order to achieve the above-mentioned cross-modal generation effect, it is necessary to systematically train and optimize the generative adversarial network. The training process of the generative adversarial network will be described below.
[0087] First, a multimodal perception dataset with a calibration relationship is used for training. The training data covers modalities such as camera images, millimeter-wave radar, and lidar point clouds. The number of samples is no less than 100,000 groups, covering a variety of typical scenarios, including sunny days, rainy days, nighttime, foggy days, and complex traffic environments. In the preprocessing stage, time synchronization between different modalities is achieved through hardware timestamp alignment (error is less than 1ms), and the coordinate system is converted through calibration parameters. The radar data and image data are projected into a unified pixel space to ensure that the input modal data are aligned in time and space. Among them, the projection of radar polar coordinate data into the image pixel coordinate system is specifically expressed as:
[0088]
[0089] In the above formula, x pixel and y pixel represents the pixel coordinates after the radar data is projected onto the image plane, f represents the focal length of the camera, r represents the distance of the radar detection target point (polar coordinate radius), θ represents the radar scanning angle (polar coordinate angle), z represents the depth of the target point in the camera coordinate system (i.e., the z-axis distance from the camera), c x and c y represents the optical center coordinates.
[0090] Generative adversarial networks employ a typical adversarial architecture, consisting of a generator and a discriminator. The generator receives multimodal input (e.g., 256×256×4, consisting of a 3-channel image and a 1-channel radar image) that has been spatially and temporally aligned and concatenated into a unified tensor. It then performs feature encoding and decoding via a convolution-deconvolution architecture, outputting a single-channel pseudo point cloud image. The discriminator determines the similarity between the generated data and the real data, outputting a local perception probability map matrix.
[0091] In order to balance the authenticity and numerical accuracy of the generated results, the total loss function used in the training process is:
[0092]
[0093] In the above formula, To combat losses; is the mean square error loss; λ is the balancing factor, which is set to 100 in this example.
[0094] The adversarial loss is used to encourage the generator to generate a distribution close to the real point cloud, which is specifically expressed as:
[0095]
[0096] In the above formula, represents the mathematical expectation; x real represents the real data sample; D(x real ) indicates that the discriminator determines x real is the probability of true data; G(x input ) represents the generator's response to the input modal data x input Generated pseudo data; D(G(x input )) indicates that the discriminator determines x input is the probability of true data.
[0097] The mean square error loss is suitable for constraining the numerical consistency between the generated point cloud and the corresponding real point cloud, which is specifically expressed as:
[0098]
[0099] In the above formula, Represents the multimodal input data of the i-th training sample; Represents the real data corresponding to the i-th training sample; Represents the i-th pseudo data generated by the generator.
[0100] Based on the total loss function, an adaptive moment estimation optimizer (β1 = 0.5, β2 = 0.999) was used for iterative optimization with a learning rate of 2×10 -4The batch size is 8 and the training period is 50 training cycles. The training process is completed on a high-performance computing platform, and each training cycle takes about 2 hours, ensuring that the generative adversarial network has good generalization ability in various typical scenarios.
[0101] It should be noted that after the generative adversarial network is trained in the cloud, it is lightweighted through structural pruning, depthwise separable convolution, knowledge distillation technology, etc., adapted to the edge computing platform, and achieves end-to-end low-latency reasoning, which not only ensures real-time performance, but also improves the generalization ability of the model in multiple scenarios and reduces the system's dependence on high-computing power hardware. Moreover, the generative adversarial network only performs forward reasoning operations and does not update parameters to ensure stable operation on resource-constrained devices. In a typical test scenario, the root mean square error (RMSE) between the virtual data corresponding to the missing mode generated by the generative adversarial network and the real point cloud data is less than 0.2 meters, and the total delay of end-to-end reasoning and post-processing is controlled within 50 milliseconds, which can meet the real-time processing requirements of the edge computing platform.
[0102] In practical applications, to further enhance the generalization and compensation accuracy of generative adversarial networks in multiple environments and scenarios, a federated learning collaborative optimization mechanism is introduced. In this mechanism, after completing local inference and some online self-learning optimization, the edge device uploads the updated local generation model parameters to the cloud server within a set period (for example, every 24 hours).
[0103] The cloud server will collect local model parameters from K vehicles And use weighted average method to aggregate the model to obtain the global model parameters, which are specifically expressed as:
[0104]
[0105] Global model parameters are periodically distributed to each edge device via encrypted transmission, replacing the parameters of its local generation network. This federated learning strategy enables cross-device model collaborative training without transmitting raw data, significantly improving the generation model's adaptability and long-term stability. Furthermore, actual test results demonstrate that the federated learning collaborative mechanism can reduce cross-modal generation error by an average of 40% and improve the model's generalization ability to diverse environmental scenarios by approximately 30%, effectively enhancing the ability to compensate for failed data and system robustness.
[0106] In one embodiment, Figure 2 As shown in Figure 2, the reasoning process of the generative adversarial network may include the following steps:
[0107] Step S201 : normalize and perform spatiotemporal registration processing on the valid original data to construct input data in a unified format.
[0108] In order to ensure the consistency of multimodal input data in space and time, it is necessary to perform preprocessing operations on the valid original data, including spatiotemporal registration processing and normalization processing.
[0109] Specifically, in the temporal dimension, the camera and millimeter-wave radar sensor data are aligned through a hardware timestamp synchronization mechanism to ensure that their acquisition time error is controlled within 1 millisecond. In the spatial dimension, the radar data is transformed from polar coordinates to pixel coordinates using calibration parameters and projected onto the image plane. The coordinate transformation relationship can be obtained by referring to Equation (1).
[0110] After completing spatiotemporal alignment, the input data is processed into a unified tensor format. For example, the camera image is downsampled to a 256×256 resolution, and the radar data is mapped to the same resolution through interpolation. The resulting data is then concatenated to form a 4-channel input tensor (i.e., 3 channels for RGB images and 1 channel for radar). The image data is then normalized, scaling it to the range [-1, 1] and mapping the radar data to the range [0, 1] to meet the input requirements of the generator network.
[0111] Step S202: Input the input data into the generator of the generative adversarial network to generate an output result corresponding to the mode of the failed sensor.
[0112] The uniformly formatted input data is then fed into the generator of the generative adversarial network for inference. The generator employs a symmetrical encoder-decoder architecture, consisting of a five-layer convolutional encoder and a five-layer deconvolutional decoder. The encoder's convolution kernel size is 4×4, with a stride of 2. The number of channels increases from 64 to 512, and the activation function used is the leaky rectified linear unit (LeakyReLU) activation function (with a leakage coefficient of 0.2). The decoder's channel count decreases layer by layer from 512 to 1, ultimately outputting a single-channel pseudo point cloud image. The output layer uses the hyperbolic tangent activation function (Tanh). Skip connections are introduced in layers 2, 3, and 4 to concatenate the encoder's intermediate features with the decoder's alignment layer to preserve local detail information.
[0113] During the actual inference process, the inference phase does not involve model parameter updates and only performs forward propagation. When the generative adversarial network is deployed on an edge computing platform, it can maintain a latency of no more than 20 milliseconds during a single inference, meeting the real-time requirements of scenarios such as autonomous driving.
[0114] Step S203 , filtering the output result using a statistical outlier removal algorithm, and performing coordinate transformation on the filtered output result to obtain first virtual data.
[0115] After the generator completes inference, its output enters the post-processing stage. First, the output is filtered using a statistical outlier removal algorithm to automatically remove isolated noise points and improve the structural stability of the generated point cloud. Subsequently, an inverse coordinate transformation is performed to remap the generated point cloud from the image coordinate system back to the vehicle coordinate system, aligning its spatial positioning with the other modal data.
[0116] Step S104 : dynamically allocating fusion weight coefficients of the heterogeneous sensors based on the comprehensive health scores of the heterogeneous sensors.
[0117] After completing the identification of the sensor status, the weight coefficient of each heterogeneous sensor participating in multimodal data fusion is dynamically assigned based on its comprehensive health score.
[0118] To comprehensively assess the stability of sensors at both the physical signal and semantic data levels, a weighted fusion of the sensor's physical and logical health scores is performed to generate a comprehensive health score representing its overall health status. This score reflects the sensor's measurement capabilities and data reliability. The comprehensive health score is specifically expressed as:
[0119] S i =α3·PLHS i +(1-α3)·LLHS i
[0120] Among them, S i represents the comprehensive health score of the i-th sensor; PLHS i Represents the physical layer health score of sensor i; LLHS i represents the logical health score of the i-th sensor; α3 is a weight coefficient (ranging from 0 to 1) that can be dynamically set based on the system scenario. For example, the weight of the logical health score can be increased in a strong interference environment. The comprehensive health score reflects the stability and credibility of the sensor at both the physical signal level and the semantic reasoning level.
[0121] It should be noted that the comprehensive health score S i The definition and purpose of the FS are fundamentally different from the previously mentioned comprehensive failure score (FS) used to determine sensor failure status. The FS is a positive indicator; higher values indicate a healthier overall sensor state and should be given a higher weight in subsequent data fusion. The FS, on the other hand, is a negative metric; higher values indicate a higher risk of sensor failure and are primarily used to trigger data compensation mechanisms.
[0122] By calculating this comprehensive score, a refined state quantification basis can be provided for sensors in different health states, thereby dynamically allocating a fusion weight coefficient in the fusion calculation to each sensor.
[0123] In one embodiment, a detailed description is provided of how to dynamically allocate fusion weight coefficients. Here, based on the comprehensive health score of each heterogeneous sensor, its fusion weight coefficient in the multimodal fusion process is further determined. To maintain a reasonable relative proportional relationship between the fusion weights, the system adopts a normalization strategy, normalizing the comprehensive health scores of all sensors to calculate their contribution to the fusion. The specific formula for allocating fusion weight coefficients is expressed as:
[0124]
[0125] Among them, W i is the fusion weight coefficient of the i-th sensor; S i is the comprehensive health score of the i-th sensor; N is the total number of sensors currently participating in the fusion. This distribution ensures that the weight of each sensor in the fusion process is not only related to its own state but also to the relative states of other sensors, which helps to achieve a more robust data fusion control strategy.
[0126] By dynamically allocating fusion weight coefficients, sensors with high health scores will obtain higher fusion weight systems, thus having a greater impact on the final perception results; conversely, sensors with low health scores will have lower corresponding fusion weight systems, thus having higher environmental adaptability and perception stability.
[0127] Step S105 : Based on the fusion weight coefficient, a data fusion algorithm is used to fuse the valid original data and the first virtual data, and a fusion perception result is output.
[0128] The valid original data is the original observation data from the valid sensor, which can be directly used for fusion; the first virtual data is the virtual observation data obtained by the cross-modal generative adversarial network compensation, which is used to replace the missing information of the failed sensor. These two types of data are uniformly organized into a data set with consistent structure and participate in the fusion calculation according to their corresponding fusion weight coefficients. In the data fusion process,
[0129] Fusion algorithms can be based on state estimation models (such as Kalman filtering) or probabilistic inference models (such as Bayesian networks). They use fusion weights for different modalities to perform a weighted fusion of the data from each modality, adaptively controlling the impact of various observational data on the fusion result. This effectively suppresses the interference of ineffective modalities while fully utilizing data from valid modalities to achieve a robust estimation of environmental targets or states, ultimately outputting a fused perception result. In practical applications, the fused perception result includes not only the physical states of each target object, such as position and velocity, but also high-level semantic information, such as category, confidence level, and behavior prediction.
[0130] Since the data fusion process takes into account the sensor state differences and weight adjustments, the final output results have high reliability and robustness, and can maintain stable perception capabilities in the event of sensor failure or complex environment.
[0131] In one embodiment, Figure 3 As shown in FIG, a specific process of data fusion based on the Kalman filter algorithm is given, which may include the following steps:
[0132] Step S301: Dynamically adjust the observation noise covariance matrix according to the fusion weight coefficient.
[0133] Before executing the Kalman filter, the observation noise covariance matrix is dynamically adjusted based on the fusion weight coefficient of each sensor. This adjustment strategy is used to reduce the dependence on sensor observations with low health scores during the state estimation process, thereby improving the robustness of the fusion process.
[0134] Here, an original observation noise covariance matrix R is preset for each heterogeneous sensor i , is used to characterize the measurement error level of the sensor under normal working conditions. In actual operation, according to the fusion weight coefficient W obtained in the previous step i , adjust the covariance matrix to obtain the current effective observation noise covariance It is calculated as follows:
[0135]
[0136] in, represents the effective observation noise covariance of the i-th sensor; W i is the corresponding fusion weight coefficient, the smaller the value, the lower the health score; ∈ is a very small constant (such as 10 -6 ), which is used to prevent the denominator from approaching zero and causing resin instability.
[0137] Through this dynamic adjustment mechanism, the interference effect of failed or abnormal sensors on state estimation can be automatically reduced during the Kalman filtering process. When the health score of a sensor is low, its corresponding W i Approaching 0, leading to The data is amplified exponentially, thereby significantly reducing the weight of the sensor's observation data in the filter gain calculation, effectively improving the stability and accuracy of fusion perception.
[0138] Step S302: Use the Kalman filter algorithm to predict and update the adjusted observation noise covariance matrix, the valid original data, and the first virtual data to obtain a target state estimation result of the perception object.
[0139] After completing the dynamic adjustment of the observation noise covariance matrix, the target state is predicted and the observation update operation is performed based on the Kalman filter algorithm, thereby realizing the time series fusion of multi-source observation data and the estimation of the target state.
[0140] First, in the prediction phase, the target state at the current moment is estimated a priori based on the state estimation results at the previous moment and the preset state transition model. The prediction calculation includes the following two core formulas:
[0141]
[0142] in, is the prior state estimate at time k, F k is the state transfer matrix (for example, using a constant velocity model), B k u k Represents control input (such as vehicle acceleration, etc.); is the prediction error covariance matrix, Q k represents the process noise covariance.
[0143] Then, in the update phase, the observation data of multiple sensors (i.e., valid original data and first virtual data) are weighted and fused with the prior state to output the target state estimation result at the current moment. First, the weighted observation value is calculated:
[0144]
[0145] Among them, z i,k represents the observation value of the i-th sensor at time k, W i is the corresponding fusion weight coefficient. Then, according to the adjusted observation noise covariance Compute the total observation covariance and construct the Kalman gain:
[0146]
[0147] Among them, H k is the observation matrix, which maps the state space to the observation space; K k is the Kalman gain, which is used to adjust the response of the predicted state to the observed value.
[0148] Finally, the state estimate and covariance matrix are updated:
[0149]
[0150] Among them, z i,k represents the observation value of the i-th sensor at time k; H k Represents the observation matrix, which maps the state to the observation space; K krepresents the Kalman gain, which determines the degree of correction of the observation value to the state estimate; w i represents the weight of the i-th sensor.
[0151] In autonomous driving scenarios, perceived objects can be key entities in the traffic environment, including vehicles ahead, pedestrians, bicycles, obstacles, curbs, traffic signs, and so on. Sensors can provide information such as their position, velocity, and acceleration at a given moment. Different types of sensors perceive the same object with a certain degree of redundancy and complementarity. Kalman filtering can fuse observations from different modalities to form a consistent target state estimate. The target state estimate is the system's optimal estimate of the motion state of a specific perceived object at the current moment, after fusing the currently valid raw data with the compensated virtual data. This estimate typically includes the object's position coordinates, velocity vector, and, where necessary, other state quantities such as heading angle and dimensions. For example, for a detected moving vehicle ahead, the state estimate might include two-dimensional position, velocity, and dimensions such as length, width, and height.
[0152] Through the above-mentioned Kalman filter prediction and update process, the valid original data and the virtual data of the failed sensor are integrated, and the degree of participation is adaptively adjusted based on their respective weights to achieve dynamic estimation of the target state. It has good real-time and fault tolerance capabilities, and is particularly suitable for target perception in scenarios with data distortion or perception loss.
[0153] In one embodiment, the data fusion algorithm also includes a Bayesian algorithm, such as Figure 4 As shown, Figure 4 The specific process of data fusion based on the Kalman filter algorithm is shown, which may include the following steps:
[0154] Step S401 : constructing a conditional dependency relationship among the sensor node, the environment state node, and the target state node, and adjusting the conditional probability table of the sensor node according to the fusion weight coefficient.
[0155] To achieve multimodal data fusion processing, a Bayesian network model is first constructed to express the conditional dependencies between sensor observations, environmental conditions, and the target state. As a directed acyclic graph (DAG), a Bayesian network can depict the causal relationships and uncertainty transmission paths between variables through probabilistic modeling, making it suitable for state estimation in multi-source information fusion tasks.
[0156] In a Bayesian network, there are three types of nodes: sensor nodes, environment state nodes, and target state nodes. Sensor nodes represent the observation data output by various heterogeneous sensors. The observation data of valid heterogeneous sensors is the raw data collected, while the observation data of invalid heterogeneous sensors is virtual data generated using a generative adversarial network. Environment state nodes describe external environmental conditions, such as weather, light intensity, and occlusion rate. Target state nodes represent the actual state of the observed object, including key motion parameters such as position, velocity, and acceleration.
[0157] These nodes are connected by edges, indicating conditional dependencies. For example, the probability distribution of a target node is influenced by its parent nodes (including multiple sensor nodes and environmental nodes). Based on this, a conditional probability table (CPT) can be set through training or knowledge rules to quantify the probability of a node under different parent node states.
[0158] To enhance the model's ability to distinguish between sensors of varying quality, the conditional probability table for each sensor node is dynamically adjusted based on each sensor's fusion weight coefficient. Specifically, higher sensor health scores correspond to more concentrated conditional probability values and higher confidence levels. As the weight of a sensor decreases, the distribution in its conditional probability table flattens, indicating increased uncertainty in the sensor's observation data, thereby reducing its impact on the target state node's inference.
[0159] Based on this, in this embodiment, an adaptive Bayesian inference structure based on fusion weights is implemented, which improves the ability to respond to changes in observation reliability.
[0160] Step S402 , based on the Bayesian algorithm, the posterior probability distribution of the target state is calculated by combining the target state estimation result, the valid original data and the first virtual data, and the fusion perception result is output.
[0161] After completing the Bayesian network modeling and setting the conditional probability table, the posterior probability inference of the target state is performed based on the Bayesian network structure. This process uses multi-source observation data as evidence and combines it with the prior probability distribution modeled in the network to achieve a comprehensive estimate of the target state.
[0162] During the inference process, the target state estimate obtained by the previous Kalman filter is used as a priori input. The original observation data provided by the current valid sensors and the virtual data generated by compensating for the failed sensors are used as observation evidence and input into the corresponding sensor nodes in the Bayesian network. By calculating the posterior probability distribution of the target state node, the combined influence of all observation data and environmental conditions is integrated to output a fused perception result.
[0163] The posterior probability distribution refers to the probabilistic inference result of the target state after the observation data is known, reflecting the possibility distribution of the target state under the given evidence conditions. In this step, the posterior probability distribution is established for the target state of the perceived object. Its input includes data from multiple sensors (valid original data and virtual data), fusion weight coefficients, prior state prediction values, etc. In the Bayesian network, the posterior probability distribution is composed of the probability output of the target state node under the given observation evidence conditions, and is formally expressed as:
[0164]
[0165] Among them, P (target state) represents the prior distribution of the target state, which can be provided by Kalman filtering or historical priors; P (observation i |Target state, w i ) represents the given target state and sensor weight w i The conditional probability distribution of the i-th sensor observation value is the weighted likelihood function. This function reflects the regulatory effect of the current fusion weight on the sensor credibility. The higher the weight, the greater the contribution of the observation value to the posterior probability.
[0166] Ultimately, the position parameter with the highest probability in the posterior distribution is selected as the fused perception output of the perceived object, and its confidence interval or uncertainty range is optionally output. This data fusion process ensures that the fused perception result has both observation consistency and weight adjustment capabilities, effectively suppressing estimation bias caused by invalid data and significantly improving the system's robustness in target perception under multimodal mismatch and complex environments.
[0167] In one embodiment, Figure 5 As shown, Figure 5 The optimization process of the generative adversarial network is shown, which may include the following steps:
[0168] Step S501 : After the failed heterogeneous sensor is restored, real data collected by the restored heterogeneous sensor within a set time and second virtual data of the restored heterogeneous sensor outputted by a generative adversarial network are obtained.
[0169] In order to improve the adaptability and generation accuracy of the generative adversarial network in long-term operation, after detecting that the failed heterogeneous sensor has resumed normal operation, the online optimization process is started and the data collection operation is performed.
[0170] Specifically, after a previously failed sensor completes self-tests and recovers, real data from that sensor is continuously collected within a set time window (e.g., 5 consecutive minutes) to serve as a reference for subsequent optimization training. To ensure data comparability, the multimodal input corresponding to the real data is fed back into the currently deployed generative adversarial network to generate a second set of virtual data matching that time period.
[0171] The real data and the second virtual data constitute a training sample pair. The real data here reflects actual observations of the sensor under normal operation, while the second virtual data reflects the generative adversarial network's ability to generate results under the current model parameters. By collecting this set of data pairs, an automatic evaluation and adjustment mechanism for the recovery sensor can be established, providing the necessary input for subsequent model updates.
[0172] Step S502: Calculate the loss parameters of the generative adversarial network based on the real data and the second virtual data, and update the parameters of the generative adversarial network based on the loss parameters.
[0173] Based on the collected real data and the corresponding second virtual data, an online training loss function is constructed, and the parameter update process of the generative adversarial network is performed to optimize its generation performance and cross-modal compensation accuracy.
[0174] In order to measure the output effect of the generator in the current environment, an unsupervised self-learning loss function is defined The loss consists of two parts: one is the mean square error between the generated data and the real data, which is used to measure the numerical difference between the generated result and the target output; the other is the generator parameter Regularization term is used to constrain network complexity and prevent overfitting. The loss function is expressed as follows:
[0175]
[0176] in, represents the i-th real data sample after recovery, represents the second virtual data sample corresponding thereto; θ G is the parameter vector of the generator; λ is the regularization coefficient, which can be set to 0.01; N is the number of samples.
[0177] After the loss function is calculated, the generator parameters are optimized using the stochastic gradient descent algorithm with momentum (SGD with Momentum). The optimizer parameters can be set to the learning rate η = 1×10 -5 , with a momentum factor of 0.9. The update rule is as follows:
[0178]
[0179] Based on this, without relying on manual labeling, the generative model is incrementally trained based on the real feedback of the restored sensors. This allows it to continuously improve the generalization ability of the generative adversarial network through online learning optimization when facing rare scenarios (such as road collapses and temporary obstacles), reduce the compensation error in long-tail scenarios, and further improve the reliability of cross-modal compensation.
[0180] It should be noted that in order to ensure that the update process of the generative model has sufficient computing resources and training stability, the online optimization process here is usually completed on the cloud server. Specifically, after the edge device collects and processes the real data of the restored sensor and the corresponding second virtual data, it uploads the data pair to the cloud training platform through a secure channel. The cloud uses its stronger computing power and storage resources to perform the complete loss function calculation and network parameter optimization operations, and after the optimization is completed, the updated generator model parameters are packaged and sent to the edge device. This cloud training-edge inference collaborative mechanism not only significantly improves the training efficiency and model convergence speed, but also effectively reduces the computing burden of the edge device, while ensuring the real-time response capability of the system in low-latency scenarios. By sending it periodically or on demand, the edge side always runs the optimal or most adaptable generative model, thereby maintaining the high-precision compensation capability and continuous adaptability of the multimodal perception system in a dynamic environment.
[0181] Combined with Figure 6 The multimodal sensor fault self-repair and dynamic fusion process shown in the figure takes multimodal sensor data as input and combines sensor status monitoring, anomaly detection, fault self-repair, data fusion and online optimization to achieve high stability and fusion processing of diversified sensor information. Figure 6 The complete process of self-repair and dynamic fusion is explained.
[0182] First, it receives data input from multiple heterogeneous sensors, covering multi-source information with different physical quantities and sampling characteristics, as the basis for subsequent processing.
[0183] Next, we conduct sensor health monitoring, evaluating sensor status through both physical and logical layer testing. Physical layer testing primarily determines sensor health based on the sensor's hardware status and performance parameters, while logical layer testing utilizes consistency analysis across multimodal sensor data to comprehensively assess the rationality and accuracy of sensor data. Combining the results of these two testing methods, we assign a comprehensive sensor health score, determining whether the sensor is experiencing failure or anomalies.
[0184] When a sensor is judged to be normal, its raw data is used as the source for subsequent data fusion. If a sensor is judged to be faulty, a cross-modal generative adversarial network (GAN) data compensation mechanism is triggered. This mechanism, based on a trained GAN, processes the valid raw input data through spatiotemporal alignment and normalization, enabling the generator to generate virtual data that is highly consistent with the sensor's failure modal characteristics.
[0185] The generated virtual data is then merged with the valid original data, achieving effective multimodal data fusion through a dynamic weighted allocation algorithm. During the fusion process, a Kalman filter algorithm and a Bayesian belief network are combined to comprehensively optimize the uncertainty and filtering of various sensor data, improving the accuracy and robustness of the fusion results. After fusion is complete, the fused perception results are output.
[0186] After a failed sensor returns to normal operation, a comparative calculation is performed based on the difference between the sensor's real data and the generated virtual data. This difference feedback is used to dynamically update the model parameters of the generative adversarial network, realizing the adaptive self-repair capability of the generative model.
[0187] The above-mentioned self-repair and dynamic fusion process not only realizes the intelligent fusion of multimodal sensor data, but also has the ability to automatically identify anomalies and compensate for repairs, thereby greatly enhancing stability and environmental adaptability.
[0188] In the above-mentioned embodiment, the self-repair method for multimodal sensor fusion can perform real-time, comprehensive health assessments of the operating status of heterogeneous sensors at both the physical and logical levels, ensuring high fault detection accuracy. In the event of sensor failure, cross-modal compensation is generated based on valid raw data, effectively restoring perception capabilities under the failed data dimension and ensuring compensation accuracy in dynamic scenarios. At the same time, the combination of a health score-driven weight distribution mechanism and an adaptive data fusion strategy ensures the dynamic responsiveness of the fusion process to differences in data quality, improves the accuracy and robustness of perception results, and can adapt to perception fusion scenarios in different environments, thereby enhancing perception reliability.
[0189] In one embodiment, a self-repairing system for modal sensor fusion is provided, referring to Figure 7 As shown, the self-repairing system 700 for automatic modal sensor fusion may include: a sensor module 701, a sensor health monitoring module 702, a cross-modal data compensation module 703, a dynamic weighted data fusion module 704 and a cloud collaboration module 705.
[0190] The sensor module 701 is used to configure a heterogeneous sensor array. The sensor module 701 integrates a multi-type heterogeneous sensor array, covering various sensor types such as lidar, camera, millimeter wave radar, infrared perception, ultrasonic sensor and IMU. Various types of sensors complement each other in terms of spatial layout, performance parameters and information dimensions, building a highly reliable and comprehensive perception system. Among them,
[0191] High-line-count LiDAR is used as the core modality for 3D spatial modeling. The example configuration uses a 128-line LiDAR with a 360° horizontal field of view and a 40° vertical field of view. It offers a maximum detection range of 150 meters, a point cloud density of 12,800 points per degree, and a frame rate of 10Hz. The LiDAR is mounted in the center of the vehicle's roof, ensuring unobstructed coverage of the primary 180° forward sensing area.
[0192] The camera includes a front-facing main camera, side cameras, and infrared cameras. The front-facing main camera uses a 2-megapixel RGB image sensor with a 120° field of view, a 30Hz frame rate, and HDR imaging capabilities (dynamic range ≥ 120dB) for high-precision visual perception. The side cameras are deployed on the left and right sides of the vehicle body, each with a 2-megapixel wide-angle camera and a 180° field of view. They are mainly used to monitor the vehicle's side blind spots. The infrared camera has a resolution of 640×512 and a frame rate of 30Hz, and can provide redundant perception support in low light, at night, or in foggy weather.
[0193] Millimeter-wave radars include forward-facing long-range radar and corner-facing radars. The forward-facing long-range radar operates at 77 GHz and has a detection capability of 200 meters, a ranging accuracy of ±0.1 meters, an angular resolution of 1°, and a frame rate of 20 Hz. It is primarily used for long-range obstacle detection in high-speed scenarios. The corner-facing radars, deployed at the four corners of the vehicle, consist of four 24 GHz short-range millimeter-wave radars with a detection range of 80 meters, suitable for close-range static and dynamic target recognition and tracking.
[0194] Up to 12 ultrasonic sensors can be installed, supporting distance sensing from 0.1 to 5 meters, with an accuracy of ±0.02 meters and a frame rate of 10Hz. The sensors are evenly spaced on the front and rear bumpers, with spacing within 0.5 meters, ensuring near-field obstacle recognition and parking assistance.
[0195] The IMU is six-axis, including a 3-axis accelerometer (±18g) and a 3-axis gyroscope (±2000° / s), with a data sampling frequency of 100Hz, which can provide stable vehicle posture and motion estimation compensation.
[0196] In the sensor module, after each sensor is installed, it is necessary to undergo high-precision mechanical calibration to ensure that the field of view overlap rate within the main sensing area is not less than 30%, in order to meet the spatial alignment requirements of cross-modal data. The sensor module enables 360° full-scene perception without blind spots, and the forward 180° key area is triple-covered by lidar, forward-looking camera, and forward-facing millimeter-wave radar. Moreover, optical sensors (camera, lidar) and radio frequency sensors (millimeter-wave radar, ultrasonic) form a complementary perception network that can effectively cope with complex environments such as rain, snow, backlight, and nighttime, thereby improving the perception robustness and adaptability of the overall system.
[0197] The sensor health monitoring module 702 is configured to obtain raw data from multiple heterogeneous sensors and generate physical layer health scores and logical layer health scores for the heterogeneous sensors based on the raw data; and determine whether the heterogeneous sensors have failed by combining the physical layer health scores and the logical layer health scores.
[0198] The cross-modal data compensation module 703 is configured to input valid raw data into a generative adversarial network to perform cross-modal generation upon determining that a failed heterogeneous sensor exists, and output first virtual data of the failed heterogeneous sensor; the valid raw data represents raw data from a valid heterogeneous sensor;
[0199] Dynamic weighted data fusion module 704 is used to dynamically assign fusion weight coefficients to heterogeneous sensors based on their comprehensive health scores; the comprehensive health score is calculated by weighting the physical layer health score and the logical layer health score; based on the fusion weight coefficients, a data fusion algorithm is used to fuse the valid original data and the first virtual data, and output a fused perception result.
[0200] In one embodiment, the sensor health monitoring module 702 is specifically configured to determine that the heterogeneous sensor is failed if the physical layer health score is less than a first preset threshold and the logical layer health score is less than a second preset threshold;
[0201] or,
[0202] Based on the physical layer health score and the logical layer health score, a comprehensive failure score of the heterogeneous sensor is determined by a weighted calculation method. If the comprehensive failure score is greater than a third preset threshold, the heterogeneous sensor is determined to be failed.
[0203] In one embodiment, the dynamic weighted data fusion module 704 is specifically configured to determine a fusion weight coefficient based on the comprehensive health score of the heterogeneous sensor and the comprehensive health scores of all heterogeneous sensors; wherein, the greater the comprehensive health score of the heterogeneous sensor, the greater the fusion weight coefficient.
[0204] In one embodiment, the data fusion algorithm includes a Kalman filter algorithm, and the dynamic weighted data fusion module 704 is specifically used to dynamically adjust the observation noise covariance matrix according to the fusion weight coefficient; use the Kalman filter algorithm to predict and observe the adjusted observation noise covariance matrix, the valid original data and the first virtual data to obtain the target state estimation result of the perceived object.
[0205] In one embodiment, the data fusion algorithm includes a Bayesian algorithm, and the dynamic weighted data fusion module 704 is specifically used to adjust the conditional probability table of the sensor node using a weight coefficient; based on the Bayesian algorithm, the posterior probability distribution of the target state is calculated by combining the target state estimation result, the valid original data and the first virtual data, and the fused perception result is output; the posterior probability distribution is determined by the conditional probability table of the sensor node.
[0206] In one embodiment, the cross-modal data compensation module 703 is specifically used to normalize and spatiotemporally align the valid original data to construct input data in a unified format; input the input data into the generator of the generative adversarial network to generate the output results of the corresponding mode of the failed sensor; use a statistical outlier removal algorithm to filter the output results, and perform coordinate transformation on the filtered output results to obtain first virtual data.
[0207] In one embodiment, the self-repairing system 700 for automatic modal sensor fusion further includes a cloud-based collaborative module 705 for obtaining, after the failed heterogeneous sensor is restored, real data collected by the restored heterogeneous sensor within a set time and second virtual data about the restored heterogeneous sensor based on the output of a generative adversarial network; calculating the loss parameters of the generative adversarial network based on the real data and the second virtual data, and updating the parameters of the generative adversarial network based on the loss parameters.
[0208] In actual applications, the cloud system module 705 is used to perform functions such as model aggregation, parameter optimization and model distribution in the cloud, thereby achieving cross-device joint optimization of the model among multiple terminals and enhancing the global generalization capability.
[0209] For the specific definition of the automatic modal sensor fusion self-repair system 700, please refer to the definition of the automatic modal sensor fusion self-repair method above, which will not be repeated here. The various modules in the above-mentioned automatic modal sensor fusion self-repair system 700 can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0210] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements a self-repair method for automatic modal sensor fusion when executing the computer program.
[0211] In one embodiment, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a self-repair method for automatic modal sensor fusion is implemented.
[0212] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0213] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0214] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0215] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0216] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A self-repair method for multimodal sensor fusion, characterized in that: include: Acquire raw data from a plurality of heterogeneous sensors, and generate a physical layer health score and a logical layer health score for the heterogeneous sensors based on the raw data; Determining whether the heterogeneous sensor is failed by combining the physical layer health score and the logical layer health score; When it is determined that a failed heterogeneous sensor exists, inputting valid raw data into a generative adversarial network for cross-modal generation, and outputting first virtual data of the failed heterogeneous sensor; the valid raw data represents raw data from the valid heterogeneous sensor; dynamically allocating fusion weight coefficients of the heterogeneous sensors according to the comprehensive health scores of the heterogeneous sensors; The comprehensive health score is calculated by weighting the physical layer health score and the logical layer health score; Based on the fusion weight coefficient, a data fusion algorithm is used to fuse the valid original data and the first virtual data, and a fusion perception result is output.
2. The self-repair method of multimodal sensor fusion according to claim 1, characterized in that: The determining whether the heterogeneous sensor is failed by combining the physical layer health score and the logical layer health score includes: If the physical layer health score is less than a first preset threshold and the logical layer health score is less than a second preset threshold, determining that the heterogeneous sensor is failed; or, Based on the physical layer health score and the logical layer health score, a comprehensive failure score of the heterogeneous sensor is determined by weighted calculation. If the comprehensive failure score is greater than a third preset threshold, the heterogeneous sensor is determined to be failed.
3. The self-repair method of multimodal sensor fusion according to claim 1, characterized in that: The allocating fusion weight coefficients of the heterogeneous sensors according to the comprehensive health scores of the heterogeneous sensors includes: The fusion weight coefficient is determined according to the comprehensive health score of the heterogeneous sensor and the comprehensive health scores of all the heterogeneous sensors; wherein, the greater the comprehensive health score of the heterogeneous sensor, the greater the fusion weight coefficient.
4. The self-repair method of multimodal sensor fusion according to claim 1, characterized in that: The data fusion algorithm includes a Kalman filter algorithm. Based on the fusion weight coefficient, the data fusion algorithm is used to fuse the valid original data and the first virtual data, and output a fusion perception result, including: Dynamically adjust the observation noise covariance matrix according to the fusion weight coefficient; The Kalman filter algorithm is used to predict and update the adjusted observation noise covariance matrix, the effective original data and the first virtual data to obtain a target state estimation result of the perception object.
5. The self-repair method of multimodal sensor fusion according to claim 4, characterized in that: The data fusion algorithm includes a Bayesian algorithm, and based on the fusion weight coefficient, the data fusion algorithm is used to fuse the valid original data and the first virtual data, and output a fusion perception result, further comprising: Constructing a conditional dependency relationship among the sensor node, the environment state node, and the target state node, and adjusting the conditional probability table of the sensor node according to the fusion weight coefficient; Based on the Bayesian algorithm, the posterior probability distribution of the target state is calculated by combining the target state estimation result, the valid original data and the first virtual data, and the fusion perception result is output; the posterior probability distribution is determined by the conditional probability table of the sensor node.
6. The self-repair method of multimodal sensor fusion according to claim 1, characterized in that: The method further comprises: After the failed heterogeneous sensor is restored, obtaining real data collected by the restored heterogeneous sensor within a set time and second virtual data of the restored heterogeneous sensor based on the output of the generative adversarial network; Based on the real data and the second virtual data, a loss parameter of the generative adversarial network is calculated, and the parameters of the generative adversarial network are updated based on the loss parameter.
7. The self-repair method of multimodal sensor fusion according to claim 1, characterized in that: When it is determined that there is a failed heterogeneous sensor, valid original data is input into a generative adversarial network for cross-modal generation, and first virtual data of the failed heterogeneous sensor is output, including: Normalizing and spatiotemporally registering the valid raw data to construct input data in a unified format; Inputting the input data into the generator of the generative adversarial network to generate an output result corresponding to the mode of the failed sensor; A statistical outlier removal algorithm is used to filter the output result, and coordinate transformation is performed on the filtered output result to obtain the first virtual data.
8. A multimodal sensor fusion self-repair system, characterized in that: include: Sensor modules for configuring heterogeneous sensor arrays; a sensor health monitoring module, configured to acquire raw data from a plurality of heterogeneous sensors and generate a physical layer health score and a logical layer health score for the heterogeneous sensors based on the raw data; Determining whether the heterogeneous sensor is failed by combining the physical layer health score and the logical layer health score; a cross-modal data compensation module for, when determining that a failed heterogeneous sensor exists, inputting valid raw data into a generative adversarial network to perform cross-modal generation and outputting first virtual data of the failed heterogeneous sensor; the valid raw data representing raw data from the valid heterogeneous sensor; a dynamic weighted data fusion module, configured to dynamically assign fusion weight coefficients of the heterogeneous sensors according to the comprehensive health scores of the heterogeneous sensors; The comprehensive health score is calculated by weighting the physical layer health score and the logical layer health score; Based on the fusion weight coefficient, a data fusion algorithm is used to fuse the valid original data and the first virtual data, and a fusion perception result is output.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the self-repair method for multimodal sensor fusion according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the self-repair method for multimodal sensor fusion according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-source sensing data processing method, medium and equipment
CN121117731A
Sensing analysis self-healing control method and device for integrated circuit power supply
CN122092646A