Multi-modal sensor fusion algorithm for multi-signal processing and system thereof
Through methods such as dynamic feature mapping, mutual feature clustering, scenario priority graph and modal bridging matrix, the problems of information loss, sensor failure and high computational complexity in multimodal sensor fusion are solved, and high-precision, reliable and real-time multimodal sensor fusion is achieved to adapt to complex dynamic scenarios.
Patent Information
- Application Number
- CN202510690392.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
The existing multimodal sensor fusion method has problems such as information loss and insufficient consistency when processing heterogeneous data, sensor failure leading to data loss and reduced system reliability, and high computational complexity leading to insufficient real-time performance.
The method of dynamic feature mapping, mutual feature clustering, scenario priority graph analysis, modal bridging matrix and edge-cloud collaborative computing is adopted. The heterogeneous data is converted into a unified feature space through dynamic feature mapping, and mutual feature clustering is performed to generate a fusion feature set with high consistency and low redundancy. The scenario priority graph is constructed to dynamically optimize the fusion parameters. When the sensor fails, the missing data is reconstructed through the modal bridging matrix, and the computing tasks are dynamically allocated to the edge or cloud platform through the task priority prediction model.
It significantly improves the data fusion accuracy and robustness of the multimodal sensor system, improves the reliability and real-time performance of the system in sensor failure scenarios, adapts to complex dynamic scenarios, reduces communication delays, and enhances the applicability of the system in application scenarios with high security and real-time requirements.
Smart Images

Figure CN120632764A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sensor data processing technology, and in particular to a multi-modal sensor fusion algorithm for multi-signal processing and a system thereof. Background Art
[0002] According to the data fusion method and system based on multimodal sensors disclosed in China Publication No. "CN118097352A", which relates to the field of data processing technology, the method includes: performing multimodal data sensing on a predetermined traffic location through a multimodal data acquisition device to obtain a multimodal traffic data set, wherein the multimodal data acquisition device includes a camera and a lidar; performing data matching and alignment on the multimodal traffic data set to generate a pre-processed multimodal data set; inputting the pre-processed multimodal data set into a fusion module for bidirectional feature alignment to obtain two-dimensional features and three-dimensional features under the same dimension; and adaptively fusing the two-dimensional features and three-dimensional features to obtain a data fusion result. This solves the technical problem in the prior art that the fusion of image and point cloud data is one-way and does not fully form the complementary advantages of the two modal data. It realizes the feature mining and fusion of image features and point cloud features, forming the technical effect of feature fusion with complementary advantages.
[0003] The above patent documents and prior art have the following technical problems when used:
[0004] Problem 1: Traditional multimodal sensor fusion methods in the aforementioned documents and existing technologies often suffer from information loss and insufficient fusion feature consistency when processing heterogeneous data due to differences in data format, dimensionality, and noise characteristics, thus affecting perception accuracy.
[0005] The second issue is that in multimodal sensor systems, the aforementioned document states that failure of one or more sensors, such as lidar or camera failures, often leads to data loss, severely impacting the integrity of the fusion results and system reliability. This is particularly true in high-security scenarios like autonomous driving and drones, where traditional methods lack effective data reconstruction mechanisms and struggle to address such failures.
[0006] Question three: In the above-mentioned documents and existing technologies, traditional multimodal sensor fusion algorithms usually rely on cloud processing due to their high computational complexity, resulting in high communication latency and insufficient real-time performance, making it difficult to meet the high low-latency requirements of scenarios such as intelligent transportation and industrial automation. Summary of the Invention
[0007] Technical problems solved
[0008] In view of the shortcomings of the existing technology, the present invention provides a multi-signal processing multi-modal sensor fusion algorithm and system thereof, which solves the following problems:
[0009] 1. Address the issues of information loss and lack of consistency in heterogeneous data fusion;
[0010] 2. Address the issues of data loss and reduced system reliability caused by sensor failure;
[0011] 3. Address the problems of high computational complexity and lack of real-time performance.
[0012] Technical Solution
[0013] To achieve the above objectives, the present invention is implemented through the following technical solutions: a multi-modal sensor fusion algorithm for multi-signal processing and a system thereof, wherein the multi-modal sensor fusion algorithm comprises the following steps:
[0014] Sp1: Collect raw data from multimodal sensors and convert heterogeneous data into a unified feature space through dynamic feature mapping, which adaptively adjusts the mapping function based on the feature retention coefficient;
[0015] Sp2: Perform mutual feature clustering based on the dynamic correlation between multimodal sensors, quantify data consistency and conflict through feature mutual construction scoring, and generate a highly consistent and low-redundancy fusion feature set;
[0016] Sp3: Builds a real-time scenario vector, analyzes environmental characteristics through a scenario priority graph, and dynamically optimizes fusion parameters to adapt to different scenarios;
[0017] Sp4: When multimodal sensors fail, the modal bridge matrix is used to reconstruct missing data using cross-modal feature correlation to generate high-fidelity compensation features.
[0018] Sp5: Evaluates the spatiotemporal consistency of the fusion results through fusion consistency metrics, and dynamically adjusts the verification threshold based on the context to output highly reliable perception information.
[0019] Sp6: During the execution of Sp1 to Sp5, the computational complexity of each step and the real-time load of the edge device are evaluated through the task priority prediction model, and the fusion tasks are dynamically allocated to the edge computing nodes and cloud platform.
[0020] Preferably, the feature retention coefficient of the dynamic feature mapping of Sp1 generates a weighted retention factor based on the real-time signal-to-noise ratio and feature dimension distribution of the multimodal sensor data, and adjusts the parameters of the mapping function through a nonlinear optimization algorithm to minimize the information loss of heterogeneous data in a unified feature space.
[0021] Preferably, the feature mutual construction score of the mutual construction feature cluster in Sp2 is calculated by constructing a dynamic correlation matrix between sensors, and the dynamic correlation matrix is based on the mutual information entropy and time series correlation of each sensor data, dynamically updates the scoring threshold, eliminates conflicting features and retains high consistency features.
[0022] Preferably, the mutual construction feature clustering in Sp2 further includes feature cluster optimization, which performs multiple rounds of screening on the fused feature set through an iterative clustering algorithm, and adjusts the cluster boundary according to the increment of the feature mutual construction score in each iteration to increase the simplicity and consistency of the feature set.
[0023] Preferably, when constructing the scenario priority graph in Sp3, the environmental features of the real-time scenario vector are mapped to nodes of a weighted directed graph, the edge weights between nodes are dynamically calculated based on the rate of change of the environmental features and the reliability of the sensor, and the selection of fusion parameters is optimized by the shortest path algorithm.
[0024] Preferably, the dynamic optimization fusion parameters in Sp3 include the adjustment of sensor weights and feature priorities, and the adjustment of feature priorities is based on the node priority sorting of the scenario priority graph, and a multi-level fusion strategy is generated through real-time updating of environmental features to adapt to complex dynamic scenarios.
[0025] Preferably, the construction of the modal bridging matrix of Sp4 is based on sample training of multimodal sensor data to generate a nonlinear mapping relationship of cross-modal features, and the mapping relationship is dynamically updated through real-time sensor performance evaluation to increase the fidelity of the reconstructed features.
[0026] Preferably, the Sp4 modal bridging matrix includes sensor failure detection when reconstructing missing data, determining the failed sensor by calculating the missing rate and abnormal value ratio of each sensor data, and preferentially using other sensor data with high feature correlation with the failed sensor for reconstruction.
[0027] Preferably, the fusion consistency measure of Sp5 is achieved by constructing a spatiotemporal deviation matrix, which quantifies the continuity of the fusion results in time series and the consistency of spatial distribution, and dynamically adjusts the verification threshold according to the environmental characteristics of the real-time scenario vector to correct abnormal fusion results.
[0028] Preferably, the system of the multimodal sensor fusion algorithm includes the following contents:
[0029] A multimodal sensor array collects the raw data in Sp1. The multimodal sensor array includes a lidar, a visible light camera, a millimeter-wave radar, an ultrasonic sensor, and an inertial measurement unit, and adjusts the data collection frequency in real time based on sensor performance through a dynamic calibration mechanism.
[0030] An edge computing unit that performs dynamic feature mapping, mutual feature clustering, and scenario priority graphs in Sp1 to Sp3. The edge computing unit integrates an adaptive load balancing module that dynamically allocates processing resources based on the computational complexity of Sp2 and Sp3. The edge computing unit includes a sensor collaborative management module that monitors the data quality of the multimodal sensor array in real time and dynamically adjusts the acquisition priority and data weight of the remaining sensors when a sensor failure is detected in Sp4 to optimize the input data quality for modal bridging reconstruction.
[0031] A cloud computing platform that performs reconstruction of the modal bridging matrix in Sp4, updates the modal bridging matrix through periodic training, and stores historical data of the scenario priority graph to optimize Sp3. The cloud computing platform includes a scenario context database that stores real-time scenario vectors in Sp3 and regularly updates node weights of the scenario priority graph through an incremental learning algorithm to improve the accuracy of dynamically optimized fusion parameters.
[0032] Communication module, supporting 5G and dedicated short-range communication protocols, ensuring low-latency data interaction between edge computing units and cloud computing platforms, and used to support real-time consistency verification in Sp5;
[0033] Visualization interface, outputs high-reliability perception information in Sp5, and provides dynamic adjustment function of fusion parameters.
[0034] Beneficial effects
[0035] The present invention provides a multi-modal sensor fusion algorithm and system for multi-signal processing. It has the following beneficial effects:
[0036] 1. The present invention adopts dynamic feature mapping and adaptive adjustment of feature retention coefficients to convert heterogeneous multimodal sensor data into a unified feature space, significantly improving the accuracy and robustness of data fusion. Traditional sensor fusion methods often lead to information loss due to inconsistent heterogeneous data formats. The present invention generates weighted retention factors based on real-time signal-to-noise ratio and feature dimension distribution, and dynamically adjusts mapping function parameters through nonlinear optimization algorithms to effectively minimize information loss. The adaptive mechanism can dynamically optimize feature extraction according to sensor data quality and environmental changes, and adapt to complex dynamic scenarios. Compared with traditional static mapping methods, this solution has made breakthroughs in feature retention and data consistency, so that the fusion results can still maintain high accuracy in high-noise or complex scenarios, providing a more reliable data foundation for multimodal sensor systems and greatly improving the stability and adaptability of the system in practical applications.
[0037] 2. The present invention uses a modal bridging matrix to achieve cross-modal feature reconstruction when the sensor fails, which significantly improves the robustness and reliability of the system in scenarios where some sensors fail. Traditional fusion systems often experience performance degradation due to data loss when the sensor fails. The present invention generates a nonlinear mapping relationship of cross-modal features through sample training, and dynamically updates the mapping matrix in combination with real-time sensor performance evaluation, thereby achieving high-fidelity reconstruction of missing data. In particular, through failure detection and feature correlation priority sorting, it uses other highly correlated sensor data for reconstruction, greatly improving the accuracy of the reconstructed features. This is particularly important in high-reliability scenarios such as autonomous driving and drone navigation. It can maintain normal system operation when the sensor fails, significantly reducing the risk of mission failure due to hardware failure, improving the system's fault tolerance and practical application value, ensuring that high-reliability perception output can be maintained in complex environments, and providing key technical guarantees for high-security application scenarios.
[0038] 3. The present invention uses a task priority prediction model and edge-cloud collaborative computing to achieve dynamic allocation of fusion tasks, significantly improving the real-time performance and computing efficiency of the system. Traditional multimodal sensor fusion systems often rely on the cloud due to their high computational complexity, resulting in increased delays and communication costs. The present invention uses an edge computing unit to perform dynamic feature mapping, mutual feature clustering, and scenario priority graph analysis, and combines it with an adaptive load balancing module to dynamically allocate tasks to the edge or cloud based on computational complexity and real-time load. It fully utilizes the high real-time performance of edge devices and the powerful computing power of the cloud, optimizes resource utilization, and reduces communication delays. The communication module supports 5G and dedicated short-range protocols to ensure low-latency data interaction, further improving the efficiency of real-time consistency verification, and significantly improving the applicability of the system in scenarios with high real-time requirements, providing an innovative solution for efficient and real-time data fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A diagram showing the steps of the multimodal sensor fusion algorithm of the present invention;
[0040] Figure 2 Output transmission relationship diagram for each step of the multimodal sensor fusion algorithm of the present invention;
[0041] Figure 3 This is a system architecture diagram of the multimodal sensor fusion algorithm of the present invention. DETAILED DESCRIPTION
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Specific embodiment one:
[0044] like Figures 1 to 3 As shown, a multi-modal sensor fusion algorithm and system for multi-signal processing, the multi-modal sensor fusion algorithm includes the following steps:
[0045] Sp1: Data Acquisition and Dynamic Feature Mapping. Raw data is collected from multimodal sensors. Dynamic feature mapping is used to transform heterogeneous data into a unified feature space. Dynamic feature mapping adaptively adjusts the mapping function based on feature retention coefficients. The feature retention coefficients generate weighted retention factors based on the real-time signal-to-noise ratio (SNR) and feature dimensionality distribution of the multimodal sensor data. The mapping function parameters are adjusted using a nonlinear optimization algorithm to minimize information loss in the unified feature space. Data heterogeneity is eliminated through normalization, providing consistent input for subsequent fusion. Data processing consists of data acquisition, preprocessing, feature extraction, and dynamic mapping. First, raw data is collected from a multimodal sensor array, including point cloud data from lidar, RGB images from cameras, range and velocity from millimeter-wave radar, proximity information from ultrasonic sensors, and acceleration and angular velocity from an inertial measurement unit (IMU). The preprocessing phase uses high-precision timestamps and motion compensation to align the data time series. Gaussian filtering is used for noise reduction, ensuring a signal-to-noise ratio greater than 10dB. During the feature extraction phase, point cloud data is voxelized to generate 128-dimensional feature vectors, image data is edge detected to generate 256-dimensional feature vectors, and distance data is generated to generate 32-dimensional vectors. Dynamic feature mapping is implemented using a nonlinear mapping network. The network structure consists of an input layer, a feature extraction layer, and a mapping layer. The mapping layer uses a multi-layer perceptron with a ReLU activation function. The feature retention coefficient is a parameter that controls the degree of information retention during the mapping process. It generates a weighted retention factor based on the real-time signal-to-noise ratio and the feature dimension distribution. The feature retention coefficient is calculated using the formula:
[0046] FRC=w1·SNR+w2·DimDist
[0047] in:
[0048] FRC is the feature retention coefficient, ranging from [0, 1]. The larger the value, the more information is retained.
[0049] SNR is the real-time signal-to-noise ratio, which is calculated by the ratio of signal power to noise power. The formula is SNR = 10*log 10 (Psignal / P noise ), P signal is the signal power, calculated by the mean square of the sensor output, P noise The noise power is estimated by the background sample variance. For example, the SNR of a lidar point cloud is calculated by dividing the point cloud intensity mean by the noise variance, with a typical value of 15-30 dB.
[0050] DimDist is the entropy value of the feature dimension distribution, reflecting the uniformity of the feature vector dimension. It is calculated by Shannon entropy. The formula is DimDist = -∑(pi*log2(pi)). Pi is the probability of the i-th dimension feature, reflecting the dispersion of the feature distribution. The typical value is 0.2-0.5.
[0051] w1 is the weighting coefficient of the signal-to-noise ratio, with an initial value of 0.6. It is dynamically adjusted according to the change of the signal-to-noise ratio. If the SNR drops by 10%, w1 increases to 0.7.
[0052] w2 is the weighted coefficient of feature dimension distribution, with an initial value of 0.4, which is adjusted according to the change of entropy value.
[0053] The mapping function parameters are adjusted by nonlinear optimization algorithm. The optimization goal is to minimize the information loss. The loss function is the mean square error, that is, MSE = ∑(X maped -X target ) 2 , where X mapped is the mapping feature, X target For the target feature, the system determines the solution to monitor the signal-to-noise ratio and entropy value. If the signal-to-noise ratio is less than 10dB or the entropy value is greater than 0.5, the noise reduction or interpolation process is triggered and the mapping is re-executed. The weight of the feature retention coefficient is dynamically adjusted according to the real-time signal-to-noise ratio. If the signal-to-noise ratio drops by 10%, w1 is increased to 0.6. The optimization formula is Where θ is the mapping function parameter, η is the learning rate, the initial value is 0.01, It is the gradient of the loss function with respect to the parameters, calculated by back propagation. The optimization process iterates 10 times per second, and the convergence condition is that the MSE is less than 0.01 or 100 iterations. If the system determines that the SNR is less than 10dB, the noise reduction process is triggered; if DimDist is greater than 0.5, the interpolation samples are added. The weighted retention factor balances the influence of the signal-to-noise ratio and the dimensionality distribution by dynamically adjusting w1 and w2, improving the mapping accuracy by 15%, which is significantly better than the linear normalization with fixed weights. Data processing includes format conversion and dimensionality unification, and outputs a 128-dimensional standardized feature vector. The input is multimodal sensor data and environmental data, and the output is a feature vector in a unified feature space for use by Sp2. Through the adaptive adjustment and nonlinear optimization of the feature retention coefficient, the information loss is significantly reduced, which is better than the traditional linear normalization.
[0054] Sp2: Mutual Feature Clustering. This method performs mutual feature clustering based on the dynamic correlation between multimodal sensors. It quantifies data consistency and conflict through feature mutual construction scores to generate a highly consistent, low-redundancy fused feature set. Dynamic correlation refers to the time-varying correlation and dependency between multimodal sensor data. It is quantified through mutual information entropy and time series correlation, reflecting the degree of coordination between sensor features. For example, the consistency between LiDAR point clouds and millimeter-wave radar distances in target detection. Mutual feature clustering generates a highly consistent, low-redundancy fused feature set through an iterative clustering algorithm based on dynamic correlation. The process includes correlation matrix construction, feature mutual construction score calculation, iterative clustering, and feature screening.
[0055] The feature mutual construction score is calculated by constructing a dynamic correlation matrix between sensors. The dynamic correlation matrix is based on the mutual information entropy and time series correlation of each sensor data, dynamically updates the scoring threshold, eliminates conflicting features and retains highly consistent features. The mutual construction feature clustering further includes feature clustering optimization, and the fusion feature set is screened multiple times through an iterative clustering algorithm. Each iteration adjusts the clustering boundary according to the increment of the feature mutual construction score to increase the simplicity and consistency of the feature set. It is used to eliminate data redundancy and conflict and generate an efficient fusion feature set. The data processing includes correlation matrix construction, score calculation, iterative clustering and feature screening. The system takes the 128-dimensional standardized feature vector of Sp1 as input and constructs an n×n dynamic correlation matrix, where n is the number of sensors. The matrix elements are calculated by mutual information entropy and time series correlation. Mutual information entropy measures the dependency between features, and the correlation is calculated by the Pearson coefficient. The feature mutual construction score is calculated by the formula:
[0056] MIS=α*MI+β*Corr
[0057] in:
[0058] MIS is the feature mutual construction score, ranging from [0, 1], where higher values indicate stronger feature consistency;
[0059] α is the weight of mutual information entropy, with an initial value of 0.6. It is adjusted through historical data training. If the consistency decreases by 10%, α is increased to 0.7.
[0060] MI is the mutual information entropy, which measures the dependency between features. The formula is MI(X,Y) = ∑(p(x,y)*log(p(x,y) / (p(x)*p(y)), where p(x,y) is the joint probability of features X and Y, and p(x) and p(y) are marginal probabilities.
[0061] Corr is the time series correlation, calculated by Pearson coefficient, Corr = ∑((X i -μ X )(Y i -μY )) / (σ X *σ Y ), μ is the mean, σ is the standard deviation;
[0062] β is the weight of the correlation, with an initial value of 0.4. If the stability of the time series decreases, β increases to 0.5. The initial values are determined by historical data training, and are 0.6 and 0.4 respectively. The scoring threshold is dynamically updated according to the MIS distribution. new =Thresh old old *(1-0.05*ΔMIS), where ΔMIS is the percentage decrease in the mean MIS. If the mean MIS decreases by 10%, the threshold is lowered by 5%. Iterative clustering uses an improved K-means algorithm. The initial number of clusters, k = √N, where N is the total number of features. Each iteration adjusts the cluster boundaries based on the MIS increment, removing features with MIS values below the threshold. After three iterations, a 64-dimensional fused feature set is generated. The mean MIS and the proportion of conflicting features are monitored. If the mean MIS is less than 0.8 or the proportion of conflicting features is greater than 10%, the number of iterations is increased to five. The clustering threshold is adjusted based on the MIS distribution. If the distribution variance is greater than 0.1, the threshold is raised by 10%. Data processing includes feature compression and conflict elimination, outputting a 64-dimensional fused feature set. The input is the normalized feature vector of Sp1, and the output is a highly consistent feature set for use by Sp3. Through a dynamic correlation matrix and iterative clustering optimization, feature consistency is significantly improved, outperforming traditional feature selection methods.
[0063] Sp3: Scenario-driven control constructs a real-time scenario vector, analyzes environmental characteristics through a scenario priority graph, and dynamically optimizes fusion parameters to adapt to different scenarios. When constructing the scenario priority graph, the environmental features of the real-time scenario vector are mapped to nodes in a weighted directed graph. Edge weights between nodes are dynamically calculated based on the rate of change of environmental features and sensor reliability. Fusion parameter selection is optimized using a shortest path algorithm. Dynamic optimization of fusion parameters includes adjusting sensor weights and feature priorities. Feature priority adjustment is based on the node priority ranking of the scenario priority graph. A multi-level fusion strategy is generated through real-time updates of environmental features to adapt to complex and dynamic scenarios. The fusion strategy is dynamically adjusted according to environmental changes to improve scenario adaptability. The processing steps include scenario vector construction, priority graph generation, and parameter optimization. The system takes the 64-dimensional fusion feature set of Sp2 and environmental sensor data as input. It extracts environmental features, including light intensity, rainfall, and target density, through a convolutional neural network to generate a 256-dimensional real-time scenario vector. The scenario priority graph is represented as a weighted directed graph, with nodes corresponding to environmental features and edge weights calculated using the formula:
[0064] Weight=γ*ChangeRate+δ*Reliability
[0065] in:
[0066] is the edge weight, ranging from [0, 1], indicating the strength of association between environmental features. The larger the value, the higher the priority.
[0067] ChangeRate is the rate of environmental change, calculated by the time difference of the characteristic value, ChangeRate=|V t -V (t-1) | / V (t-1) , V t is the current eigenvalue;
[0068] Reliability is the reliability of the sensor, ranging from [0 to 1], and is calculated by the signal-to-noise ratio and the historical failure rate. Reliability = 0.7*SNR norm +0.3*(1-FailureRate), SNR norm is the normalized signal-to-noise ratio, FailureRate is the failure rate in the last 100 frames;
[0069] δ is the weight of sensor reliability, with an initial value of 0.5. If reliability decreases, δ is reduced to 0.3
[0070] γ is the weight of the environmental change rate, with an initial value of 0.5. If the change rate increases by 50%, γ increases to 0.7;
[0071] The corresponding parameters for different scenarios include sensor weight vectors and feature priority lists. For example: in rainy scenes (rainfall greater than 10mm / h), the camera weight is reduced to 0.2, the millimeter-wave radar weight is increased to 0.5, and image edge features are prioritized; in sunny scenes (light greater than 500lux), the camera weight is increased to 0.4, and texture features are prioritized. The graph optimization uses the shortest path algorithm to calculate the optimal path from the initial node to the target node to generate a sensor weight vector and a feature priority list. Feature priority is based on node priority sorting. If the light node has the highest weight, the image feature priority is increased by 20%. Sensor weight allocation is based on the node weights and edge weights of the scenario priority graph in Sp3. The weight vector is generated by the shortest path algorithm. For example, the sum of the weights of the five sensors (lidar, camera, millimeter-wave radar, ultrasonic, inertial measurement unit) is 1, and the initial value is [0.2, 0.2, 0.2, 0.2, 0.2]. The weight is adjusted according to the environmental characteristics, and the formula is Weight i =NodeWeight i / ∑NodeWeigh t i ,NodeWeogh t iis the weight of the scene priority graph node corresponding to the i-th sensor (based on lighting, rainfall, etc.). For example, on rainy days (rainfall greater than 10 mm / h), the camera weight is reduced to 0.2, and the millimeter-wave radar weight is increased to 0.5. Feature priority adjustment is based on the priority sorting of the scene priority graph nodes. If the lighting node has the highest weight (0.4), the priority of image features (such as edges and textures) is increased by 20% and recorded in the priority list (such as [edge: 0.5, texture: 0.3, point cloud: 0.2]). The multi-level fusion strategy includes a highly robust mode and a real-time mode, which are switched based on the environment feature update. The multi-level fusion strategy includes a highly robust mode (for inclement weather, weights favor millimeter-wave radar and lidar, such as [0.4, 0.1, 0.3, 0.1, 0.1]) and a real-time mode (for normal scenarios, weights are balanced, such as [0.2, 0.3, 0.2, 0.2, 0.1]). These modes correspond to environmental characteristics: the highly robust mode corresponds to rainfall greater than 10 mm / h or illumination less than 200 lux, while the real-time mode corresponds to illumination greater than 500 lux. If the fusion error exceeds 5%, the system re-adjusts the weights. This mechanism improves scenario adaptability by 25% through dynamic allocation and priority sorting. The system monitors the edge weight update frequency and parameter error. If the frequency is less than 1 Hz or the error exceeds 5%, the scenario vector is reconstructed. The weights of the scenario priority graph are adjusted based on the rate of environmental change. If the rate of change increases by 50%, γ is increased to 0.7. Data processing includes feature mapping and weight calculation, outputting optimized fusion parameters. The input is the feature set and environmental data of Sp2, and the output is the sensor weights and priority list for Sp4. Through dynamic weight updates and multi-level strategies of the scenario priority graph, scenario adaptability is significantly improved.
[0072] Sp4: Modal Bridging Reconstruction. When a multimodal sensor fails, the modal bridging matrix is used to reconstruct missing data using cross-modal feature correlations to generate high-fidelity compensatory features. The modal bridging matrix is constructed based on sample training of multimodal sensor data to generate a nonlinear mapping relationship between cross-modal features. This mapping relationship is dynamically updated through real-time sensor performance evaluation, increasing the fidelity of the reconstructed features. The modal bridging matrix reconstruction includes sensor failure detection. By calculating the missing rate and outlier ratio of each sensor data point, the failed sensor is identified. Data from other sensors with high feature correlation with the failed sensor is prioritized for reconstruction, enhancing the system's fault tolerance and ensuring reliable feature generation even in the event of sensor failure. The operational process includes failure detection, matrix construction, and data reconstruction. The system uses the fusion parameters of Sp3, the 64-dimensional feature set of Sp2, and real-time sensor data as input. It monitors the quality of each sensor data in real time, calculates the missing rate and outlier ratio, and deems a sensor failure if the missing rate exceeds 30% or the outlier ratio exceeds 20%. The modal bridging matrix is constructed through a deep neural network. The network input is a 64-dimensional feature vector and the output is the compensation feature. The structure includes 3 fully connected layers and the activation function is ReLU. The modal bridging matrix generates a nonlinear mapping relationship through sample training. The training data is historical sensor data and the loss function is the mean square error. The mapping relationship is updated through real-time sensor performance. If the signal-to-noise ratio drops by 20%, the modal bridging matrix parameters are retrained. The reconstruction process gives priority to high-correlation sensor data and generates 64-dimensional compensation features through the modal bridging matrix. Monitor the reconstruction error and fidelity. If the mean square error is greater than 0.1 or the fidelity is less than 90%, select sensor data with higher correlation. Sensor failure detection is achieved by real-time monitoring of data missing rate and outlier ratio. If the missing rate is greater than 30% or the outlier ratio is greater than 20%, it is judged to be a failure. The missing rate is the proportion of missing samples, and the formula is:
[0073] MissingRate=N missing / N total
[0074] Among them, N missing is the number of missing samples (no output or invalid value), N total is the total number of samples, 100 frames / second;
[0075] The outlier ratio is calculated using the 3σ criterion:
[0076] OutlierRate=N outlier / N total
[0077] Among them, N outlierThe number of samples exceeding ±3 standard deviations of the mean is based on the most recent 100 frames of data. For example, a camera may experience a 60% loss rate due to occlusion on rainy days. The system uses an embedded microcontroller to perform calculations every second, and the judgment result triggers the modal bridging reconstruction of Sp4, prioritizing the use of high-correlation sensor data. Detection accuracy reaches 95%, ensuring timely failure identification. Reconstruction priority is adjusted based on the number of failed sensors. If there are more than two failed sensors, lidar data is prioritized. Data processing includes feature mapping and compensation generation, outputting compensation features consistent with the Sp2 format. The input is the Sp3 parameters and the Sp2 feature set, and the output is the compensation features used by Sp5. Dynamic updating of the modal bridging matrix and high-correlation priority reconstruction improves fault tolerance and feature fidelity.
[0078] Sp5: Fusion consistency measurement, which evaluates the spatiotemporal consistency of the fusion results through fusion consistency measurement, and dynamically adjusts the verification threshold based on the scenario context to output highly reliable perception information. The fusion consistency measurement is achieved by constructing a spatiotemporal deviation matrix. The spatiotemporal deviation matrix quantifies the continuity of the fusion results in the time series and the consistency of the spatial distribution. It also dynamically adjusts the verification threshold according to the environmental characteristics of the real-time scenario vector, corrects abnormal fusion results, verifies the reliability of the fusion results, and eliminates spatiotemporal deviations. The operation process includes deviation matrix construction, deviation calculation, threshold adjustment, and result correction. The system takes the compensation features of Sp4, the fusion parameters of Sp3, and the feature set of Sp2 as input to construct an m×n spatiotemporal deviation matrix, where m is the time step and n is the spatial dimension;
[0079] Time continuity is calculated by the autocorrelation coefficient, the formula is:
[0080]
[0081] Spatial consistency is calculated by the Euclidean distance between eigenvectors:
[0082]
[0083] The deviation is calculated using the formula:
[0084] Deviation=λ*TemporalDev+μ*SpatialDev
[0085] in:
[0086] Deviation is the total deviation, ranging from [0, ∞), where smaller values indicate higher consistency;
[0087] λ is the time deviation weight, with an initial value of 0.5. If the time continuity requirement is reduced (such as rainy days), λ is reduced to 0.3;
[0088] TemporalDev is the time deviation, calculated by the autocorrelation coefficient, TemporalDev = 1-R(τ), τ is the time lag;
[0089] μ is the spatial bias weight, with an initial value of 0.5. If the spatial consistency requirement is reduced, μ is reduced to 0.3;
[0090] SpatialDev is the spatial deviation, calculated by Euclidean distance, , X i and Y i is the vector component of the feature;
[0091] The verification threshold is adjusted according to the scenario vector of Sp3, such as lowering the time continuity threshold by 10% on rainy days. If the deviation is greater than the threshold, the feedback mechanism is triggered, and Sp4 reconstruction or Sp3 parameter optimization is re-executed. The final fusion result is generated by weighted fusion, and the weight comes from the sensor weight vector of Sp3, outputting the target position (x, y, z), speed (vx, vy, vz) and category (vehicle, pedestrian, obstacle). The system judgment scheme monitors the deviation mean and threshold coverage. If the mean is greater than 0.05 or the coverage is less than 95%, the number of feedback iterations is increased. Adjust the threshold according to the scenario vector, Threshold new =Thresh old *(1-0.1*EnvFactor) is the normalized value of environmental features, with an initial threshold of 0.05. If the deviation is greater than the threshold, Sp4 reconstruction is triggered. If the system determines that the coverage is less than 95%, feedback iterations are added. The output perception information is target data in JSON format, which is used for path planning, obstacle detection, etc., to improve system reliability by 20%. If the rainfall increases by 50%, the spatial consistency threshold is reduced by 15%. Data processing includes deviation quantification, result correction and format conversion, and outputs perception information in JSON format. The input is Sp4 features and Sp3 parameters, and the output is high-reliability perception information for use by external systems. The innovation of this step lies in the adaptive threshold and spatiotemporal deviation matrix, which improves the robustness of the fusion results.
[0092] Sp6: Adaptive computing load sharing. During the execution of Sp1 to Sp5, the computational complexity of each step and the real-time load of the edge device are evaluated through the task priority prediction model, and the fusion tasks are dynamically allocated to the edge computing nodes and cloud platforms to optimize the allocation of computing resources and ensure real-time performance and efficiency. It includes task evaluation, load monitoring, priority prediction and task allocation. The system takes the intermediate data and computing tasks of Sp1 to Sp5 as input to evaluate the computational complexity of each step, such as the number of clustering iterations of Sp2 and the graph optimization path length of Sp3. Load monitoring collects the CPU occupancy, memory usage and network bandwidth of the edge device. Task priority prediction is achieved through a regression model, and the priority is calculated by the formula:
[0093] Priority=η*Complexity+θ*Latency
[0094] in:
[0095] Priority is the task priority, ranging from [0, 1]. The higher the value, the higher the priority.
[0096] η is the weight of computational complexity, with an initial value of 0.6. If the complexity increases by 50%, η increases to 0.7.
[0097] Complexity is the computational complexity of the task, estimated by the number of iterations or path length, such as the number of clustering iterations of Sp2, and is normalized to the range [0, 1];
[0098] θ is the weight of the predicted delay, with an initial value of 0.4. If the delay increases, θ increases to 0.5;
[0099] Latency is the predicted task delay, in milliseconds, estimated by a regression model (linear regression, based on historical data), ranging from [0, 100];
[0100] Priority is used for task allocation of Sp6. High-priority tasks (such as Sp1 and Sp2) are allocated to edge nodes to optimize real-time performance by 20%, and low-priority tasks are allocated to the cloud platform. The allocation plan is generated by priority sorting to ensure that the edge device load is less than 80% and the task delay is less than 50ms. The system judgment plan monitors the load and delay. If the load is greater than 80% or the delay is greater than 50ms, the task priority prediction weight is adjusted and η is increased to 0.7. Tasks are dynamically allocated according to the real-time load. If the edge CPU occupancy rate exceeds 85%, part of the calculation of Sp3 is transferred to the cloud. Data processing includes task segmentation and allocation optimization, and outputs a task allocation plan. The input is task data from Sp1 to Sp5, and the output is an allocation plan for system execution. The dynamic prediction of task priority prediction and load balancing significantly improve real-time performance and resource utilization.
[0101] When the system is initialized, the sensor configuration, dynamic feature mapping model parameters, initial weights of the scenario-driven control model, and the initial mapping relationship of the modal bridging matrix are loaded. Then, the execution is cyclically executed in the order of Sp1 to Sp5, and Sp6 monitors and optimizes the task allocation in real time. Each step exchanges intermediate data through shared memory. Sp1 generates standardized features, Sp2 generates a fusion feature set, Sp3 optimizes fusion parameters, Sp4 generates compensation features, and Sp5 outputs perception information. Abnormal situations such as sensor failure and computational overload are handled through a feedback mechanism, and the termination condition is the output of perception information or the triggering of a new task. The system operates in a collaborative manner on edge computing nodes and cloud platforms. The edge nodes handle real-time tasks, and the cloud handles complex training to ensure that the delay is less than 100ms. Specific embodiment two:
[0103] like Figures 1 to 3 As shown, based on the content in the above specific embodiments, the following contents are further disclosed:
[0104] The system of multimodal sensor fusion algorithm includes the following:
[0105] The multimodal sensor array, responsible for raw data acquisition in Sp1, includes a lidar, visible light camera, millimeter-wave radar, ultrasonic sensor, and inertial measurement unit (IMU). A dynamic calibration mechanism adjusts data acquisition frequency in real time based on sensor performance to ensure data quality and system adaptability. The lidar uses 16 or 32 lines, 100Hz, and 1024 points per frame to collect 3D point cloud data, covering a 360° field of view and generating spatial position information. The visible light camera uses a CMOS sensor with a resolution of 1080p, 30fps, and a 120° field of view to capture RGB images and provide visual information. The millimeter-wave radar has a ranging accuracy of ±0.1m and a velocity accuracy of ±0.05m / s, collecting target distance and velocity. The ultrasonic sensor has a ranging range of 0.2-5m and an accuracy of ±0.01m, collecting information about close-range obstacles. The IMU includes a 3-axis accelerometer and a 3-axis gyroscope with an accuracy of ±0.01g, collecting acceleration and angular velocity, and providing attitude and motion data. Data acquisition is coordinated by an embedded microcontroller, using high-precision timestamps with an accuracy of ±1ms to ensure time synchronization. Motion compensation is used to align data sequences based on inertial measurement unit data. The acquisition frequency is adjusted based on sensor performance through a dynamic calibration mechanism, for example, reducing the camera frame rate to 20fps on rainy days. Based on sensor collaborative sampling and real-time calibration, the microcontroller collects signal-to-noise ratio and environmental data from each sensor every second. A pre-trained calibration model is used to calculate the frequency adjustment coefficient. If the signal-to-noise ratio falls below 10dB or environmental interference exceeds a threshold, the system triggers frequency adjustment or data noise reduction to ensure time synchronization. High-precision timestamps and inertial measurement unit-assisted calibration eliminate data latency. Data processing includes format conversion and preprocessing, outputting a standardized data stream for processing by the edge computing unit. Dynamic calibration improves data quality by 30%, maintaining over 90% data integrity even in harsh environments. It supports dynamic feature mapping in Sp1, providing high-quality input for subsequent steps.
[0106] The edge computing unit is responsible for dynamic feature mapping, co-constructed feature clustering, and scenario-driven control in Sp1 through Sp3. An integrated adaptive load balancing module dynamically allocates processing resources based on the computational complexity of Sp2 and Sp3. It also includes a sensor coordination management module for real-time monitoring of data quality across the multimodal sensor array. Upon detecting a sensor failure in Sp4, it dynamically adjusts the acquisition priority and data weighting of the remaining sensors to optimize the input data quality for modal bridging reconstruction. The unit's hardware architecture is based on a high-performance embedded platform with integrated sensor interfaces and storage modules. Task scheduling is managed by a real-time operating system. Dynamic feature mapping in Sp1 utilizes a GPU-accelerated nonlinear mapping network. Co-constructed feature clustering in Sp2 utilizes iterative clustering performed by the CPU. Scenario-driven control in Sp3 combines GPU and CPU processing of the scenario priority graph. The adaptive load balancing module estimates CPU / GPU utilization and task complexity per second based on the number of Sp2 iterations and the path length of the Sp3 graph. If the complexity of Sp2 causes the CPU utilization to exceed 80%, some computation in Sp3 is offloaded to the GPU. The sensor collaborative management module analyzes the sensor data missing rate and outlier ratio in real time. If the missing rate exceeds 30%, it is deemed failed, the millimeter-wave radar weight is increased to 0.5, and the acquisition priority is adjusted. Based on parallel computing and dynamic scheduling, the unit stores the 128-dimensional feature vector of Sp1, the 64-dimensional feature set of Sp2, and the fusion parameters of Sp3 in shared memory, with data exchange latency less than 10ms. If the system determines that task latency exceeds 50ms, load balancing is triggered. If the data quality of the failed sensor falls below 90%, weights are reallocated. Resource allocation is dynamically adjusted based on complexity, prioritizing the real-time performance of Sp1 and Sp2. Data processing includes feature compression, matrix calculation, and parameter optimization, outputting sensor weights and a priority list for Sp3. Load balancing reduces computational latency by 30%, optimizes the quality of reconstructed input by 20% in sensor failure scenarios, and supports efficient execution of Sp1 through Sp3.
[0107] The cloud computing platform is responsible for performing modal bridging reconstruction in Sp4, updating the modal bridging matrix through periodic training, and storing historical data of the scenario priority graph to optimize Sp3. It includes a scenario context database and regularly updates the node weights of the scenario priority graph through an incremental learning algorithm to improve the accuracy of dynamic optimization fusion parameters. The hardware structure of the platform includes a high-performance server (equipped with 2 Intel Xeon CPUs, 4 NVIDIA A100 GPUs, 1TB RAM, 10TB SSD storage) and a distributed database (based on MongoDB). Computing tasks are managed through containerization technology. The modal bridging matrix training uses a GPU-accelerated deep neural network. The modal bridging matrix parameters are updated every 24 hours based on 100,000 frames of historical data. The loss function is the mean square error, and the updated data is synchronized to the edge computing unit. The scenario context database stores Sp3's 256-dimensional real-time scenario vectors. An incremental learning algorithm is used to update the node weights of the scenario priority graph every 12 hours. Based on distributed computing and data persistence, the platform receives the Sp2 feature set and Sp3 scenario vectors from edge units via a high-bandwidth network, performs Sp4 reconstruction, generates 64-dimensional compensation features, and feeds the optimized scenario priority graph weights back to the edge units. If the system determines that the modal bridging matrix reconstruction error exceeds 0.1, retraining is triggered. If the fusion error of the updated scenario priority graph weights exceeds 5%, the learning rate is adjusted to 0.02. The control logic prioritizes GPU resources for modal bridging matrix training, ensuring training latency of less than one hour. Data processing includes feature mapping, model training, and data storage, outputting compensation features and updated scenario priority graph weights. As a result, the platform's incremental learning improves the accuracy of the scenario priority graph by 15% and the fidelity of the modal bridging matrix reconstruction by 92%, significantly enhancing the fault tolerance of Sp4 and the scenario adaptability of Sp3.
[0108] The communication module supports 5G or dedicated short-range communication protocols to ensure low-latency data interaction between the edge computing unit and the cloud computing platform, and is used to support the real-time consistency check in claim 1Sp5, ensuring timely transmission and verification of the fusion results. The hardware structure of this module includes a 5G modem, a dedicated short-range communication unit and a network interface controller. Data transmission is managed through a protocol stack. 5G is used for edge-cloud interaction, and the dedicated short-range communication unit is used for communication between local devices. Based on low-latency and high-reliability transmission, the module transmits the 64-dimensional feature set of Sp2, the fusion parameters of Sp3 and the compensation features of Sp4 to the cloud every second, and simultaneously receives the modal bridge matrix parameters and the weights of the scenario priority map from the cloud for the calculation of the spatiotemporal deviation matrix of Sp5. If the system determines that the transmission delay is greater than 30ms or the packet loss rate is greater than 1%, it switches to the dedicated short-range communication unit or adjusts the 5G frequency band, giving priority to the transmission of Sp5 verification data, using the TCP protocol to ensure data integrity, and compressing the data packets if the bandwidth occupancy rate exceeds 90%. Data processing includes data packaging, encryption (AES-256) and decompression, and outputs a real-time interactive data stream. Effectively, the module ensures Sp5 verification delay of less than 50ms, data integrity of 99.9%, supports real-time fusion result output, and maintains stable communication, especially in highly dynamic scenarios.
[0109] The visualization interface is responsible for outputting high-reliability perception information from Sp5 and providing dynamic adjustment of fusion parameters, allowing users to monitor fusion results in real time and optimize system performance. The hardware structure of this interface includes a high-resolution display, an embedded processor, and a touch screen controller. Perception information is presented through a graphical user interface (GUI), displaying target location (x, y, z), velocity (vx, vy, vz), and category (vehicle, pedestrian, obstacle). The information is formatted as JSON and rendered as a 3D scene or 2D heat map. Users can adjust Sp3 fusion parameters, such as sensor weights and feature priorities, via the touchscreen. Adjustment instructions are transmitted in real time to the edge computing unit. Based on real-time rendering and interaction, the interface receives Sp5 perception information every second and communicates with the edge unit via the WebSocket protocol, with latency less than 10ms. If the rendering frame rate falls below 30fps, the system reduces the resolution to 1080p. If the fusion error exceeds 5% after user parameter adjustment, a prompt is issued to readjust the parameters. High-priority targets (such as the vehicle ahead) are displayed first, and user adjustment history is recorded to optimize Sp3's SPG weights. Data processing includes JSON parsing, 3D rendering, and parameter encoding, with output of visualization results and adjustment instructions. This improves user interaction efficiency by 40% and fusion accuracy by 10% after parameter adjustment, significantly enhancing system operability and perception transparency in autonomous driving and industrial monitoring scenarios.
[0110] The system collects high-quality data through a multimodal sensor array, with edge computing units performing real-time processing and cloud computing platforms handling complex tasks. The communication module ensures low-latency interaction and provides intuitive output and interaction. Based on modular collaboration, the system loads sensor configurations, model parameters, and communication protocols during initialization, then loops through steps Sp1 to Sp6. Abnormal conditions such as sensor failure and computational overload are handled through a feedback mechanism. System judgment and control logic ensure that all components work together, with an overall latency of less than 100ms and a fusion accuracy exceeding 95%. The system maintains 90% perception reliability in harsh environments and improves computing efficiency by 35%. It is suitable for autonomous driving, industrial automation, and smart cities, supporting high-precision perception in complex scenarios. Specific embodiment three:
[0112] like Figures 1 to 3 As shown, based on the content in the above specific embodiments, the following contents are further disclosed:
[0113] In order to further verify the feasibility and core technical effects of the technical solution of this application, the following actual application cases are further disclosed:
[0114] The present invention is applied to target detection and tracking of autonomous driving in complex urban environments: In complex traffic environments, autonomous driving vehicles need to detect and track pedestrians, vehicles and other obstacles in real time, facing challenges such as illumination variation range of 200-1000 lux, rainfall rate of 0-20 mm / h and sensor noise. This solution converts lidar, camera and millimeter wave radar data into a 128-dimensional unified feature space through dynamic feature mapping of Sp1. Gaussian filtering is used for preprocessing, and the signal-to-noise ratio is improved to 20 dB. Feature extraction generates 128-dimensional vectors of point clouds, 256-dimensional vectors of images and 32-dimensional vectors of radars. Dynamic feature mapping uses a multi-layer perceptron, and the feature retention coefficient is calculated based on the signal-to-noise ratio and feature dimension entropy (0.25-0.4), w1=0.6, w2=0.4, and nonlinear optimization (learning rate 0.01) reduces the MSE to 0.007. Sp2 generates a 64-dimensional fused feature set through mutual feature clustering (K-means, k = √N, N = 128, 3 iterations), with a mutual information entropy weight wMIS = 0.6, and removes conflicting features (ratio < 8%). Sp3's scenario priority map adjusts sensor weights based on illumination and rainfall, reducing the camera weight to 0.2 on rainy days and increasing the millimeter-wave radar weight to 0.5, optimizing target detection accuracy. Sp5's spatiotemporal deviation matrix has a temporal continuity autocorrelation coefficient of 0.9 and a spatial consistency Euclidean distance of <0.05, ensuring the reliability of the fusion results. It outputs the target position (x, y, z), velocity (vx, vy, vz), and category (pedestrian, vehicle). Compared with traditional linear fusion methods, this solution improves target detection accuracy by 18% in rainy scenes, significantly improving robustness in complex environments. Dynamic feature mapping significantly reduces information loss through adaptive feature retention coefficients. The optimized MSE (0.007) is 53% lower than the traditional method (0.015), and it still maintains high detection accuracy (93%) in low signal-to-noise ratio (12dB) scenarios, significantly improving the accuracy and robustness of data fusion. It is particularly suitable for autonomous driving target detection in complex urban environments.
[0115] In mountain rescue missions, drones must avoid obstacles in foggy conditions with visibility less than 50 meters and rainfall of 15 mm / h. Camera occlusion can cause a 60% loss rate. This solution reconstructs missing data using the Sp4 modal bridging matrix to ensure navigation reliability. The system uses lidar, ultrasonic sensors, and an inertial measurement unit (IMU). Sp4 detects camera failures with a loss rate greater than 30%. Using a three-layer fully connected neural network with Reluctant Unit (ReLU) activation and 2000 frames of training data with an average significance error (MSE) of 0.009, it constructs the modal bridging matrix, prioritizing lidar with a correlation of 0.85 to generate 64-dimensional compensatory features with 94% fidelity. Sp1 preprocesses and aligns time series, providing high-precision timestamps with an error of less than 1 millisecond, generating 128-dimensional feature vectors. The inter-construction feature clustering in Sp2 achieved an MIS mean of 0.82 and a conflicting feature ratio of 6%. The feature set was optimized. Sp3 adjusted sensor weights based on rainfall and visibility: 0.5 for lidar, 0.3 for IMU, and 0.2 for ultrasonic. The spatiotemporal deviation matrix in Sp5 achieved temporal continuity of 0.88 and spatial consistency <0.04. Verification results show that obstacle positions and velocities are output. Compared to traditional methods (82% fidelity and 3m positioning error), this solution reduced positioning error to 0.8m and increased obstacle avoidance success rate to 97%, significantly enhancing fault tolerance. The modal bridging matrix reconstructs missing features from highly correlated sensor data, with a correlation of 0.85 for lidar. This improved the system's fault tolerance and robustness, ensuring highly reliable perception output even in complex environments. This solution ensures high-precision navigation even in camera failure scenarios, making it suitable for drone rescue missions in inclement weather.
[0116] In the real-time signal optimization in intelligent transportation systems, the timing of traffic lights needs to be optimized in real time to alleviate congestion during peak hours, with a vehicle density of 15 vehicles / 100m 2), this application realizes efficient task allocation and optimizes real-time performance through the edge-cloud collaborative computing of Sp6. The system processes lidar, camera and millimeter-wave radar data. Sp1 generates a 128-dimensional feature vector with a processing time of 10ms. Sp2 generates a 64-dimensional feature set (15ms) through mutual feature clustering (3 rounds of iterations, MIS0.86). Sp3 adjusts the sensor weight according to traffic density and light 500lux (0.4 for camera and 0.3 for lidar), generates fusion parameters in 5ms, and Sp4 handles occasional failures of millimeter-wave radar (missing rate 20%) and reconstructs features (MSE0.01, 10ms). Sp6's task priority prediction model (η = 0.6, θ = 0.4) allocates tasks based on Sp2's three iterations and Sp3's 80-node path length. Sp1 and Sp2 are processed on edge nodes (CPU 2.5GHz, 4GB RAM, 75% load), while Sp4 is processed in the cloud with a 15ms delay. 5G communication delay of 8ms supports data interaction, reducing total latency to 38ms and signal adjustment time from 1.8s to 0.9s, improving traffic flow by 22%. Edge-cloud collaboration, through dynamic task allocation (Sp1 and Sp2 edge processing, Sp4 cloud processing), reduces latency by 58% (38ms vs. 90ms) and signal adjustment time by 50% (0.9s vs. 1.8s). This demonstrates an effective balance between computational efficiency and real-time performance, breaking through the bottleneck of traditional fusion systems for high-complexity tasks and significantly improving the system's applicability in scenarios with high real-time requirements. This solution efficiently addresses the high computational demands of peak hours and significantly optimizes the performance of intelligent transportation systems.
[0117] In summary, the above application cases verify the feasibility of the technical solution of this application in autonomous driving, drone navigation and intelligent transportation, respectively reflecting the core technical effects of dynamic feature mapping, modal bridging matrix and edge-cloud collaboration, and verifying its high precision, fault tolerance and real-time advantages in complex environments.
[0118] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the statement "comprising a reference structure" does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0119] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A multimodal sensor fusion algorithm for multi-signal processing, characterized by: The multimodal sensor fusion algorithm includes the following steps: Sp1: Collect raw data from multimodal sensors and convert heterogeneous data into a unified feature space through dynamic feature mapping, which adaptively adjusts the mapping function based on the feature retention coefficient; Sp2: Perform mutual feature clustering based on the dynamic correlation between multimodal sensors, quantify data consistency and conflict through feature mutual construction scoring, and generate a highly consistent and low-redundancy fusion feature set; Sp3: Builds a real-time scenario vector, analyzes environmental characteristics through a scenario priority graph, and dynamically optimizes fusion parameters to adapt to different scenarios; Sp4: When multimodal sensors fail, the modal bridge matrix is used to reconstruct missing data using cross-modal feature correlation to generate high-fidelity compensation features. Sp5: Evaluates the spatiotemporal consistency of the fusion results through fusion consistency metrics, and dynamically adjusts the verification threshold based on the context to output highly reliable perception information. Sp6: During the execution of Sp1 to Sp5, the computational complexity of each step and the real-time load of the edge device are evaluated through the task priority prediction model, and the fusion tasks are dynamically allocated to the edge computing nodes and cloud platform.
2. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The feature retention coefficient of the dynamic feature mapping of Sp1 generates a weighted retention factor according to the real-time signal-to-noise ratio and feature dimension distribution of the multimodal sensor data, and adjusts the parameters of the mapping function through a nonlinear optimization algorithm to minimize the information loss of heterogeneous data in a unified feature space.
3. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The feature mutual construction score of the mutual construction feature cluster in Sp2 is calculated by constructing a dynamic correlation matrix between sensors, and the dynamic correlation matrix is based on the mutual information entropy and time series correlation of each sensor data, dynamically updates the scoring threshold, eliminates conflicting features and retains highly consistent features.
4. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The mutual construction feature clustering in Sp2 further includes feature clustering optimization, which performs multiple rounds of screening on the fused feature set through an iterative clustering algorithm. In each iteration, the cluster boundaries are adjusted according to the increment of the feature mutual construction score to increase the simplicity and consistency of the feature set.
5. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: When constructing the scenario priority graph in Sp3, the environmental features of the real-time scenario vector are mapped to nodes of a weighted directed graph. The edge weights between nodes are dynamically calculated based on the rate of change of environmental features and sensor reliability, and the selection of fusion parameters is optimized through the shortest path algorithm.
6. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The dynamic optimization fusion parameters in Sp3 include the adjustment of sensor weights and feature priorities, and the adjustment of feature priorities is based on the node priority sorting of the scenario priority graph, and a multi-level fusion strategy is generated through real-time updating of environmental features to adapt to complex dynamic scenarios.
7. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The construction of the modal bridging matrix of Sp4 is based on sample training of multimodal sensor data to generate a nonlinear mapping relationship between cross-modal features, and the mapping relationship is dynamically updated through real-time sensor performance evaluation to increase the fidelity of the reconstructed features.
8. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The Sp4 modal bridging matrix includes sensor failure detection when reconstructing missing data. By calculating the missing rate and outlier ratio of each sensor data, the failed sensor is determined, and the other sensor data with high feature correlation with the failed sensor is preferentially used for reconstruction.
9. The multimodal sensor fusion algorithm for multi-signal processing according to claim 1, characterized in that: The Sp5 fusion consistency measure is achieved by constructing a spatiotemporal deviation matrix, which quantifies the continuity of the fusion results in time series and the consistency of spatial distribution. The verification threshold is dynamically adjusted according to the environmental characteristics of the real-time scenario vector to correct abnormal fusion results.
10. A multimodal sensor fusion algorithm for multi-signal processing according to any one of claims 1 to 9, characterized in that: The system of the multimodal sensor fusion algorithm includes the following: A multimodal sensor array collects the raw data in Sp1. The multimodal sensor array includes a lidar, a visible light camera, a millimeter-wave radar, an ultrasonic sensor, and an inertial measurement unit, and adjusts the data collection frequency in real time based on sensor performance through a dynamic calibration mechanism. An edge computing unit that performs dynamic feature mapping, mutual feature clustering, and scenario priority graphs in Sp1 to Sp3. The edge computing unit integrates an adaptive load balancing module that dynamically allocates processing resources based on the computational complexity of Sp2 and Sp3. The edge computing unit includes a sensor collaborative management module that monitors the data quality of the multimodal sensor array in real time and dynamically adjusts the acquisition priority and data weight of the remaining sensors when a sensor failure is detected in Sp4 to optimize the input data quality for modal bridging reconstruction. A cloud computing platform that performs reconstruction of the modal bridging matrix in Sp4, updates the modal bridging matrix through periodic training, and stores historical data of the scenario priority graph to optimize Sp3. The cloud computing platform includes a scenario context database that stores real-time scenario vectors in Sp3 and regularly updates node weights of the scenario priority graph through an incremental learning algorithm to improve the accuracy of dynamically optimized fusion parameters. Communication module, supporting 5G and dedicated short-range communication protocols, ensuring low-latency data interaction between edge computing units and cloud computing platforms, and used to support real-time consistency verification in Sp5; Visualization interface, outputs high-reliability perception information in Sp5, and provides dynamic adjustment function of fusion parameters.
Citation Information
Cited By
Unmanned aerial vehicle obstacle avoidance method based on sensing fusion and reinforcement learning
CN120871979A
Intelligent enterprise operation data analysis system based on artificial intelligence
CN121167508A
Positioning method of robot in complex scene
CN121323645A
Traffic engineering premises mountain torrent risk early warning analysis method based on multi-source perception
CN121352526A
Method and device for acquiring CRP gather of seismic data acquired by multidirectional towrope
CN121432531A