A data fusion method and system based on deep learning

By adopting deep learning technology in unmanned systems, the data fusion model is built, and the problem of difficult fusion of multiple sensor data in complex environments is solved, more efficient and accurate data fusion is achieved, and the decision-making capabilities of unmanned systems are improved.

CN119625478BActive Publication Date: 2025-05-13NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162058.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Unmanned systems have difficulty making full use of multiple sensor data in complex environments, which affects the accuracy and effectiveness of decision-making.

Method used

Using a deep learning-based data fusion method, a data fusion model is constructed through deep recursive neural network DRNN, bidirectional long and short-term memory unit Bi-LSTM and sparse attention mechanism, multi-step preprocessing and data-level fusion are performed, and the model is optimized to improve computing efficiency and generalization capabilities.

Benefits of technology

It significantly improves the accuracy and reliability of data fusion, fully taps the potential of multi-source information, improves the accuracy and effectiveness of decision-making in complex environments, and is suitable for a variety of sensor data processing tasks in unmanned systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625478B_ABST
    Figure CN119625478B_ABST
Patent Text Reader

Abstract

The present invention is a data fusion method and system based on deep learning, which relates to the field of deep learning technology, including classifying multiple sensor data obtained in an unmanned system according to set classification rules to obtain classified data; preprocessing the classified data based on the basic attributes of the classified data to obtain preprocessed classified data; fusing the preprocessed classified data at the data level to obtain multiple classified fusion data; building a data fusion model based on a deep recurrent neural network DRNN, combining a bidirectional long short-term memory unit Bi-LSTM and a sparse attention mechanism, optimizing the data fusion model through model pruning, and obtaining an optimized data fusion model; inputting multiple classified fusion data into the optimized data fusion model for secondary fusion to obtain the final fusion feature. Fully tap the potential of multi-source information and improve the accuracy and effectiveness of decision-making in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a data fusion method and system based on deep learning. Background Art

[0002] With the rapid development of science and technology, unmanned systems are increasingly widely used. Their efficient operation depends on the accurate processing and rapid decision-making of multi-source heterogeneous data. Unmanned systems such as drones, unmanned vehicles and unmanned ships play an important role in traffic monitoring, disaster relief, environmental monitoring, etc. These application environments are often unpredictable, requiring unmanned systems to be able to adapt to different weather and geographical conditions. In the commercial field, unmanned systems are used in logistics distribution, agricultural monitoring, infrastructure inspection, etc., and need to complete tasks efficiently and accurately in a wide geographical area.

[0003] At present, unmanned systems face a series of technical challenges in data processing and decision-making. In particular, the problem of making full use of sensor data is particularly prominent. The differences in the design principles and physical properties of various sensors have great differences in the description and representation of the environment. Faced with many types of sensors and their heterogeneous data structures, due to the large amount of data and high requirements for data timeliness, existing technical means often find it difficult to fully tap the potential of multi-source information, resulting in serious impacts on the accuracy and effectiveness of decision-making in complex environments. How to efficiently extract and effectively utilize key information from these data has become a technical problem that needs to be solved urgently in the field of unmanned systems. Summary of the invention

[0004] In order to overcome the above-mentioned shortcomings in the prior art that it is difficult to fully tap the potential of multi-source information, resulting in serious impact on the accuracy and effectiveness of decision-making in complex environments, the main purpose of the present invention is to provide a data fusion method and system based on deep learning.

[0005] To achieve the above object, the present invention adopts the following technical solution, a data fusion method based on deep learning, comprising the following steps:

[0006] According to the set classification rules, the various sensor data obtained in the unmanned system are classified to obtain classified data; based on the basic attributes of the classified data, pre-processing is performed to obtain the pre-processed classified data;

[0007] The preprocessed classification data is fused at the data level to obtain multiple classification fusion data;

[0008] Based on the deep recurrent neural network DRNN (Deep Recurrent Neural Network), combined with the bidirectional long short-term memory unit Bi-LSTM (Bidirectional Long Short-Term Memory) and the sparse attention mechanism, a data fusion model is constructed, and the data fusion model is optimized through model pruning to obtain an optimized data fusion model;

[0009] The multiple classified fusion data are input into the optimized data fusion model for secondary fusion to obtain the final fusion features.

[0010] The construction of the data fusion model includes the following steps:

[0011] Select the deep recurrent neural network DRNN as the basic architecture, replace its neurons with bidirectional long short-term memory units Bi-LSTM, and set the parameters of the bidirectional long short-term memory unit Bi-LSTM, including the weights and biases of the forget gate, input gate, cell state, and output gate;

[0012] Add residual connections between each Bidirectional Long Short-Term Memory Unit Bi-LSTM layer formed,

[0013] A skip connection is added between every four bidirectional long short-term memory unit Bi-LSTM layers, a sparse self-attention mechanism is inserted into each skip connection, and PReLU is added as the activation function after the skip connection to build a data fusion model.

[0014] The step of optimizing the data fusion model by model pruning to obtain an optimized data fusion model comprises the following steps:

[0015] Acquire a training set to train the data fusion model until convergence, and obtain a trained data fusion model;

[0016] Based on the trained data fusion model, select the pruning strategy as gradient pruning, and obtain the gradient and cumulative gradient of each weight;

[0017] Based on the accumulated gradient, set the threshold and identify the parameters that can be pruned;

[0018] Perform pruning according to the accumulated gradient and the set threshold to obtain a pruned data fusion model;

[0019] The pruned data fusion model is fine-tuned and optimized through SGD to obtain the optimized data fusion model.

[0020] The step of obtaining the gradient of each weight comprises the following steps:

[0021] Obtain the loss between the predicted output and the true label calculated by the loss function in the trained data fusion model;

[0022] Initialize the gradients in the trained data fusion model;

[0023] Perform backpropagation and obtain the gradient of the loss with respect to each weight.

[0024] The step of inputting the multiple classified fusion data into the optimized data fusion model for secondary fusion comprises the following steps:

[0025] Arrange the multiple classified fusion data into three-dimensional tensors respectively, wherein the dimensions of the three-dimensional tensors are the number of samples, the time step and the number of features respectively, to obtain multiple classified three-dimensional fusion data;

[0026] Input multiple classified three-dimensional fusion data into the bidirectional long short-term memory unit Bi-LSTM layer to obtain multiple time series features;

[0027] The plurality of time series features are transmitted through residual connections, and a plurality of enhanced fusion features are obtained through skip connections and PReLU (ParametricRectified Linear Unit) activation functions;

[0028] The multiple enhanced fusion features are weightedly concatenated in the fully connected layer to obtain the final fusion feature.

[0029] The classification rule is to classify according to sensor type and data usage;

[0030] The step of obtaining the classified data comprises the following steps:

[0031] Classifying the multiple sensor data acquired in the unmanned system according to the sensor type to obtain first classified data;

[0032] The second classified data is obtained by classifying various sensor data in the unmanned system according to the data usage.

[0033] The preprocessing is performed based on the basic attributes of the classified data, including the following steps:

[0034] Perform data alignment and time synchronization on the first classification data and the second classification data respectively to obtain synchronized first classification data and synchronized second classification data;

[0035] Performing spatial registration and denoising on the synchronized first classification data and the synchronized second classification data respectively, to obtain the registered and denoised synchronized first classification data and the synchronized second classification data;

[0036] The registered and denoised synchronized first classification data and synchronized second classification data are respectively standardized, interpolated and filled with missing values ​​to obtain preprocessed first classification data and second classification data.

[0037] The data-level fusion of the preprocessed classified data includes respectively performing data-level fusion on the preprocessed first classified data and the second classified data, and includes the following steps:

[0038] Performing dynamic weighted fusion on the preprocessed first classification data and the second classification data respectively to obtain dynamically fused first classification data and second classification data;

[0039] Performing data dimension reduction and feature selection on the dynamically fused first classification data and the synchronized second classification data respectively to obtain a plurality of first features and a plurality of second features;

[0040] A plurality of first features and a plurality of second features are weightedly spliced ​​to obtain a plurality of first fused data and a plurality of second fused data, and a plurality of classified fused data are integrated.

[0041] A data intelligent fusion system based on deep learning, comprising:

[0042] The data analysis and processing module is used to classify the various sensor data obtained in the unmanned system according to the set classification rules to obtain classified data; pre-process the classified data based on the basic attributes of the classified data to obtain pre-processed classified data; and perform data-level fusion on the pre-processed classified data to obtain multiple classified fusion data;

[0043] A fusion model building module is used to build a data fusion model based on a deep recurrent neural network DRNN, combined with a bidirectional long short-term memory unit Bi-LSTM and a sparse attention mechanism, and optimize the data fusion model through model pruning to obtain an optimized data fusion model;

[0044] The fusion data acquisition module is used to input multiple classified fusion data into the optimized data fusion model for secondary fusion to obtain the final fusion features.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] In the present invention, sensor type classification and data usage classification help to classify various sensor data according to their sources and application scenarios, making subsequent data processing more targeted and efficient. The dynamic weighted fusion method is adopted to weighted splicing of different data sources after preprocessing. This solves the problem of data heterogeneity and reasonably distributes the influence of various types of data, further improving the fusion effect.

[0047] The present invention performs secondary fusion on the dynamically fused data through the constructed data fusion model. Specifically, the bidirectional processing mode of Bi-LSTM can capture past and future information at the same time. Then, the use of the introduced residual connection and jump connection enables the deep neural network to transmit information more effectively, thereby avoiding the gradient vanishing problem and enhancing the learning ability of the data fusion model. In order to automatically focus on the features that have an important impact on the prediction results and suppress the interference of irrelevant information, thereby improving the generalization ability and accuracy of the model, a sparse attention mechanism is also inserted in the jump connection. The pruning technology and fine-tuning method can effectively reduce the computing resource consumption of the data fusion model without damaging the prediction performance of the data fusion model. The data fusion model is optimized. The optimized data fusion model can effectively process and fuse multiple sensor data, thereby improving the accuracy and robustness of the data fusion results. The optimized data fusion model avoids the overfitting problem by means of residual connection, jump connection, activation function, etc., so that the optimized data fusion model is more stable and accurate on the unseen data, can adapt to multi-sensor data of different types and sources, and has strong adaptability to data intelligent fusion in complex environments such as unmanned systems.

[0048] The present invention significantly improves the accuracy and reliability of data fusion through multi-step preprocessing and data-level fusion, as well as the application of deep learning models, and provides more stable data support for unmanned systems. The final fusion features contain comprehensive information of multi-sensor data, fully tap the potential of multi-source information, and improve the accuracy and effectiveness of decision-making in complex environments based on accurate and comprehensive fusion features. Unmanned systems can make more informed decisions and achieve more precise control, thereby improving overall performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0050] Figure 1 It is a schematic diagram of the process structure of the present invention;

[0051] Figure 2 It is a schematic diagram of the process structure of building a data fusion model in the present invention;

[0052] Figure 3 It is a schematic diagram of the process structure of optimizing the data fusion model in the present invention;

[0053] Figure 4 It is a schematic diagram of the process structure of obtaining the optimized data fusion model in the present invention;

[0054] Figure 5 It is a schematic diagram of the process structure of obtaining the final fusion feature in the present invention;

[0055] Figure 6 It is a schematic diagram of the process structure of obtaining the classified data after preprocessing in the present invention;

[0056] Figure 7 It is a schematic diagram of the partial process structure of obtaining multiple multi-class fusion in the present invention;

[0057] Figure 8 It is a schematic diagram of the overall process structure for obtaining multiple multi-class fusions in the present invention. DETAILED DESCRIPTION

[0058] At present, unmanned systems face a series of technical challenges in data processing and decision-making. In particular, the problem of making full use of sensor data is particularly prominent. Faced with many types of sensors and their heterogeneous data structures, existing technical means often find it difficult to fully tap the potential of multi-source information, resulting in serious impacts on the accuracy and effectiveness of decision-making in complex environments. How to efficiently extract and utilize key information from these data has become a technical problem that needs to be solved in the field of unmanned systems.

[0059] In order to overcome the above-mentioned existing shortcomings, the main purpose of the present invention is to provide a data fusion method based on deep learning. Figure 1 , including the following steps:

[0060] According to the set classification rules, the various sensor data obtained in the unmanned system are classified to obtain classified data; based on the basic attributes of the classified data, pre-processing is performed to obtain the pre-processed classified data;

[0061] The preprocessed classification data is fused at the data level to obtain multiple classification fusion data;

[0062] Based on the deep recurrent neural network DRNN, combined with the bidirectional long short-term memory unit Bi-LSTM and sparse attention mechanism, a data fusion model is constructed, and the data fusion model is optimized through model pruning to obtain the optimized data fusion model;

[0063] Multiple classified fusion data are input into the optimized data fusion model for secondary fusion to obtain the final fusion features.

[0064] The present invention will be further described below in conjunction with the accompanying drawings and implementation modes.

[0065] Embodiment 1:

[0066] The present invention proposes a data fusion method based on deep learning, which uses a deep recursive neural network DRNN, a bidirectional long short-term memory unit Bi-LSTM and a sparse self-attention mechanism to further improve the effect of data fusion.

[0067] First, the various sensor data in the unmanned system are classified according to the set classification rules. The classification rules distinguish according to the sensor type and data usage, so as to obtain different categories of data. These data will be preprocessed separately, including data alignment, time synchronization, spatial registration, denoising, standardization, data interpolation and missing value filling, to ensure that the input data has good quality.

[0068] In this embodiment, there are three sensors, and the first classification data classified by sensor type are LiDAR (Light Detection and Ranging), RGB (Red Green Blue) camera and infrared sensor.

[0069] The specific preprocessing process of the first category data classified by sensor type is as follows: The timestamp and sampling frequency of each sensor may be different. In order to align these data, they need to be synchronized to a unified time axis based on the timestamp. The frame rate of the RGB camera is 30 Hz, the lidar is 10 Hz, and the infrared sensor is 5 Hz. The timestamp of the original data is:

[0070] RGB camera: [0.00, 0.033, 0.066, ..., 0.97]

[0071] LiDAR: [0.00, 0.1, 0.2, ..., 0.9]

[0072] Infrared sensor: [0.00, 0.2, 0.4, ..., 1.0]

[0073] To time synchronize them through interpolation or sampling techniques so that their timestamps are consistent, the sampling frequency of the IR sensor can be increased to 30 Hz to align with the RGB camera.

[0074] Sensor data is in different coordinate systems and needs to be converted to a common spatial reference frame. LiDAR data is a 3D point cloud, RGB images are 2D pixel values, and infrared images are 2D heat maps. Through specific spatial registration methods such as rotation and translation matrices, various types of data are converted to a unified 3D coordinate system.

[0075] Each sensor may generate noise, especially RGB cameras and infrared sensors. For RGB images, median filtering or Gaussian filtering can be used to remove noise, and the pixel values ​​of each channel can be normalized to a uniform value range. In this embodiment, each pixel value can be divided by 255 to normalize it to the interval [0, 1].

[0076] LiDAR data is missing, especially in the case of sensor failure or occlusion. Missing values ​​can be filled using linear interpolation or interpolation based on neighboring data. That is, for the missing LiDAR data at timestamp t=0.2, use the data at t=0.1 and t=0.3 for interpolation.

[0077] The pre-processed data is fused at the data level according to the multiple first classification data and the multiple second classification data. Specifically, it includes: performing dynamic weighted fusion between multiple data of the same category to obtain dynamically fused data; the pre-processed first classification data is:

[0078] LiDAR: 3D point cloud data, dimension is [number of points, 3].

[0079] RGB camera: image data with dimensions of [height, width, number of channels] (720×1280×3).

[0080] Infrared sensor: 2D heat map with dimensions [height, width] (240×320).

[0081] The pre-processed multiple first classification data are dynamically weighted and fused to obtain dynamically fused first classification data, that is, different weighting coefficients are assigned according to the reliability or task requirements of different sensors. The dynamic weighted fusion can be a simple weighted average, or the weight can be adjusted according to the real-time scene. For example, in a low-light environment, the weight of the infrared sensor may be increased, while the weight of the RGB camera may be reduced. In this embodiment, the weights of the multiple first classification data are:

[0082] Lidar weight: 0.5;

[0083] RGB camera weight: 0.3;

[0084] Infrared sensor weight: 0.2;

[0085] The fused data is often of high dimension, and feature selection and dimensionality reduction are required to obtain several features; that is, for RGB image data, a convolutional neural network (CNN) can be used for feature extraction to reduce the dimension of the image data from 720×1280×3 to a smaller feature space of 512-dimensional vectors. For radar point cloud data, PCA (Principal Component Analysis) or PointNet (PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation) can be used to reduce the dimensionality of the three-dimensional point cloud data to reduce the computational complexity. Then, features of different categories are integrated through weighted splicing, and finally multiple first fused data and second fused data are integrated to integrate multiple classified fused data.

[0086] In the process of building the data fusion model, the deep recurrent neural network DRNN is used as the basic architecture. In order to enhance the model's ability to process time series data, the neurons in the network are replaced with bidirectional long short-term memory units Bi-LSTM. Residual connections are added between each Bi-LSTM layer, and jump connections are introduced between every four Bi-LSTM layers to speed up the training process and avoid the gradient vanishing problem. A sparse self-attention mechanism is inserted at each jump connection to further improve the model's attention to important features. Finally, the PReLU activation function is used to enhance the network's nonlinear expression ability.

[0087] In order to improve the computational efficiency and generalization ability of the model, the data fusion model is also optimized through model pruning. First, the gradient of each weight is obtained during the training process, and the parameters that need to be pruned are selected through the gradient pruning strategy. The prunable parameters are identified according to the gradient and the set threshold, and pruned. Subsequently, the pruned model is fine-tuned and optimized through SGD (Stochastic Gradient Descent) to obtain the final optimized data fusion model.

[0088] The optimized data fusion model will be used for secondary fusion processing. Multiple classified fusion data will be organized into three-dimensional tensors and input into the Bi-LSTM layer for temporal feature extraction. Through residual connections and skip connections, combined with sparse self-attention mechanism, the fusion features are further enhanced. Finally, multiple enhanced fusion features will be weighted spliced ​​in the fully connected layer to obtain the final fusion features.

[0089] This example uses a sensor dataset of an unmanned system for experiments. The dataset contains multiple sensor types, such as lidar, RGB camera, infrared sensor, etc., covering multiple scenarios and tasks. The classification rules used in the experiment are set according to the type of sensor and the purpose of the data.

[0090] In the process of deep learning training and reasoning, the present invention has a lot of parameters, which requires the computer to have stronger computing power and higher memory. Therefore, the NVIDIA TitanX graphics card was selected in the deep learning environment construction. After configuring CUDNN and CUDA environment, the GPU acceleration was successfully used for data calculation, which improved the training speed. The system hardware configuration is shown in Table 1.

[0091] Table 1 System Configuration

[0092]

[0093] Under this configuration, the performance of this method is compared with that of traditional data fusion methods such as weighted average method and principal component analysis (PCA).

[0094] In order to comprehensively evaluate the effect of the fusion method, the following evaluation indicators are used:

[0095] Table 2 Comparison of indicators between traditional weighted average method and optimized data fusion model

[0096]

[0097] Experimental results show that the intelligent fusion method based on deep learning is superior to traditional methods in data fusion accuracy, computational efficiency and model generalization ability. Especially in the processing of complex time series data, Bi-LSTM combined with sparse self-attention mechanism can effectively capture key features and improve the quality of the final fusion features.

[0098] The proposed data fusion method based on deep learning combines deep recurrent neural network,

[0099] Bidirectional long short-term memory units and sparse self-attention mechanisms achieve efficient fusion of multiple sensor data. Experimental results show that this method has significant advantages in data fusion accuracy, computational efficiency and model generalization ability, and is suitable for complex data processing tasks in unmanned systems.

[0100] Embodiment 2:

[0101] The present invention proposes a data fusion method based on deep learning, which is particularly suitable for the efficient fusion of multiple sensor data in unmanned systems. The method combines deep recurrent neural network DRNN, bidirectional long short-term memory unit Bi-LSTM, sparse self-attention mechanism and model pruning technology, and improves the accuracy and efficiency of data through multi-stage data processing, fusion and optimization steps. Specifically, it includes the following contents:

[0102] In unmanned systems, sensors collect a wide variety of data, including vision, hearing, temperature, humidity, acceleration, magnetic field, etc. To facilitate processing, the sensor data is first classified according to the type of sensor and the purpose of the data.

[0103] Sensor type: such as visual sensor, temperature sensor, acceleration sensor.

[0104] Data usage: such as environmental monitoring, target detection, and path planning.

[0105] According to the sensor type, the data is preliminarily classified to obtain the following data set, multiple first classification data:

[0106] Visual sensor data (images, videos, etc.).

[0107] Environmental monitoring data (temperature, humidity, gas concentration, etc.).

[0108] Motion sensor data (accelerometer, gyroscope, etc.).

[0109] According to the purpose of the data, the data is classified to obtain the following data sets and multiple second classification data:

[0110] Sensor data for environmental monitoring (temperature, humidity, gas concentration, etc.).

[0111] Sensor data for object detection (visual images, radar data, etc.).

[0112] Sensor data (acceleration, position data, etc.) for path planning.

[0113] Through these classifications, different types of data can be subsequently processed and integrated in a more targeted manner.

[0114] For the classified data, necessary preprocessing is performed on each data set, including time synchronization, spatial registration, denoising, standardization, interpolation and other operations.

[0115] Time synchronization is performed on multiple first-classification data and multiple second-classification data to ensure that multiple first-class data and multiple second-class data are aligned at the same timestamp. Then a spatial registration operation is performed to ensure the spatial consistency of each data, followed by denoising to remove errors. Standardization is performed on all data, for example, all data are normalized to a range of 0 to 1, and missing values ​​are filled by interpolation methods.

[0116] On the preprocessed multiple first-classification and multiple second-classification data, different types of sensor data are merged through data-level fusion to obtain a more comprehensive feature representation. Dynamically weight the data of each category according to the different importance of the data. For example, in target detection, visual data may be more important than environmental monitoring data, so it is given a higher weight. Perform dimensionality reduction on the fused data to extract important features. For example, use PCA for data dimensionality reduction, or use feature selection methods based on mutual information to retain features that have a significant impact on classification or regression tasks. All features are weighted and spliced ​​to obtain a multi-dimensional fused data set, the fused first-classification and the fused second-classification data.

[0117] The present invention uses a deep recurrent neural network DRNN as the basic architecture, and embeds a bidirectional long short-term memory unit Bi-LSTM between its neurons. Bi-LSTM can simultaneously consider the previous and next time series information and improve the time series data modeling capability. A residual connection is added between each Bi-LSTM layer to prevent the gradient vanishing problem and improve the network training efficiency. At the same time, a jump connection is added between every four LSTM layers to reduce the computational complexity. A sparse self-attention mechanism is inserted in the jump connection so that the model can selectively focus on important features. The nonlinear mapping is enhanced by the PReLU activation function. The input data is a three-dimensional tensor with a shape of (number of samples, time step, number of features), for example (100, 10, 5), which means 100 samples, each sample has 10 time steps, and each time step has 5 features.

[0118] In order to improve the computational efficiency of the model and reduce overfitting, the model pruning technology is used to optimize the model. The specific pruning process includes:

[0119] First, the gradient of each parameter is calculated through the trained data fusion model, and the parameters with smaller gradients are selected for pruning. The pruning threshold is determined by the cumulative gradient. Then the pruned data fusion model is fine-tuned and optimized through stochastic gradient descent SGD to ensure that the pruned model can still achieve high accuracy. That is, during the training process, the gradient of some parameters is less than the set threshold, for example, the set threshold is 0.01. When the calculated parameters are less than 0.01, these parameters will be pruned to reduce the amount of calculation.

[0120] In the optimized data fusion model, multiple classified fusion data are input and secondary fusion is performed to obtain the final feature representation. The secondary fusion process specifically includes:

[0121] The multiple classified fusion data are organized into a three-dimensional tensor and input into the Bi-LSTM layer for time series feature extraction. The fusion features are further enhanced through residual connections, skip connections, and PReLU activation functions. Finally, all enhanced features are weighted and concatenated in the fully connected layer to obtain the final fusion features as input for the next decision or prediction. The input classification data may include images, temperature and humidity data, and acceleration data. After being processed by the deep learning model, the final fusion feature output may be a vector of length 256, representing the fusion features of all sensor data.

[0122] This embodiment provides a data fusion method based on deep learning, which realizes the efficient fusion of multiple sensor data in unmanned systems through classification, preprocessing, data fusion, deep learning model construction, pruning optimization and secondary fusion. This method can effectively improve the data representation ability and model performance, and is suitable for sensor data processing tasks in various intelligent systems.

[0123] Embodiment 3:

[0124] When drones are performing environmental monitoring and obtaining effective environmental information, the key challenge they face is how to quickly and accurately extract useful information from the massive data collected by multiple sensors. The following is based on the principle of data fusion method, and specifically includes the following contents:

[0125] First, the various sensor data obtained in the unmanned system are classified according to the set classification rules. According to the sensor type and data usage, they are divided into two categories, namely:

[0126] According to the sensor type, a plurality of first classification data are obtained, which are: infrared image data, visible light image data, radar echo data and sonar signal data, wherein:

[0127] Infrared image data: resolution 640x480, frame rate 30Hz.

[0128] Visible light image data: resolution 1920x1080, frame rate 60Hz.

[0129] Radar echo data: frequency range 2GHz to 3GHz, sampling rate 1MHz.

[0130] Sonar signal data: frequency range 10kHz to 50kHz, sampling rate 100kHz.

[0131] According to the data usage classification, multiple second classification data are obtained, namely: target detection and terrain analysis, among which:

[0132] Object detection: mainly uses infrared and visible light image data.

[0133] Terrain analysis: mainly uses radar echo data and sonar data.

[0134] Preprocessing is performed respectively based on basic attributes of the classification data, that is, preprocessing is performed respectively on a plurality of first classification data and a plurality of second classification data.

[0135] The first classification data preprocessing is to align the timestamps of the infrared and visible light image data to the millisecond level, and to match the infrared image with the visible light image at the pixel level through the image processing algorithm.

[0136] The specific preprocessing process is explained by taking the infrared image data and the visible light image data in the plurality of first classification data as examples.

[0137] Specifically, the preprocessing of the first classification data - infrared image data is as follows:

[0138] The infrared image data is adjusted for resolution and frame rate control to keep the resolution at 640x480 and the frame rate at 30Hz. Then, the infrared image data collected at different times are aligned through timestamp matching to ensure data synchronization. Edge detection is used for spatial registration and median filtering or wavelet transform denoising is applied to reduce image noise and achieve spatial registration and denoising. The infrared image data is standardized so that its pixel value range is between 0 and 1. For frames where infrared image data is missing, linear interpolation is used to fill the missing data to obtain a preprocessed first classification data.

[0139] The specific preprocessing of the first classification data - visible light image data is as follows:

[0140] The resolution and frame rate of the visible light image data are adjusted to keep the resolution at 1920x1080 and the frame rate at 60Hz. Then the data is aligned and synchronized through the timestamp, which is similar to the infrared image data processing part. Then the image is registered based on the feature points. And the bilateral filter is applied to denoise, which keeps the edge information while reducing the noise, and realizes spatial registration and denoising. Finally, the visible light image data is normalized. For frame loss, the interpolation method based on motion estimation is used to fill the visible light image data to obtain a pre-processed first classification data.

[0141] Specifically, the preprocessing of the first classification data - radar echo data is as follows: the frequency range and sampling rate of the radar echo data are controlled, and the frequency range is 2GHz to 3GHz, and the sampling rate is 1MHz. Then, the data alignment and time synchronization are achieved by synchronizing the sampling clock, and the Doppler effect is applied for spatial registration. Then, a bandpass filter is used to remove clutter and noise. Then, the radar echo data is normalized. Finally, the FFT-based interpolation method is used to process the missing data to obtain a first classification data - radar echo data after preprocessing.

[0142] The specific first classification data - sonar signal data preprocessing is as follows: the frequency range and sampling rate of the sonar signal data are controlled, and the frequency range is 10kHz to 50kHz, and the sampling rate is 100kHz. Then the synchronous clock source is used for data alignment and time synchronization. Then the beamforming technology is used for spatial registration. And the adaptive filter is used for denoising to achieve spatial registration and denoising. Finally, the sonar data is normalized. The missing data points are filled using the interpolation method based on adjacent data points. The first classification data - sonar signal data after preprocessing is obtained.

[0143] The specific preprocessing process of the second classification data is explained using target detection data and terrain analysis data in multiple second classification data.

[0144] The target detection data preprocessing is specifically as follows: combining infrared and visible light image data to ensure the temporal consistency of the target detection data. Using a feature matching algorithm to spatially register the infrared and visible light images, and applying joint denoising techniques such as non-local mean filtering to enhance target features, achieve spatial registration and denoising, and finally perform feature standardization on the image data. In this embodiment, principal component analysis PCA is used. For missing data in the target features, similarity metrics are used for interpolation. A second classification data after preprocessing is obtained.

[0145] The terrain analysis data preprocessing is specifically as follows: combining radar and sonar data to ensure the time synchronization of terrain analysis data. Using geographic information system (GIS) tools for spatial registration to ensure that the data is in the same coordinate system. And applying adaptive filters to remove noise in radar and sonar data. Spatial registration and denoising are achieved, and then the terrain data is standardized. In this embodiment, the height data is logarithmically transformed. For missing values ​​in the terrain data, the Kriging interpolation method is used to fill them. A second classification data after preprocessing is obtained.

[0146] The preprocessed first classification data and the second classification data are dynamically weighted and fused respectively to obtain dynamically fused first classification data and second classification data; data dimension reduction and feature selection are performed on the dynamically fused first classification data and the synchronized second classification data respectively to obtain a plurality of first features and a plurality of second features; the plurality of first features and a plurality of second features are weightedly spliced ​​to obtain a plurality of first fused data and a plurality of second fused data, and a plurality of classification fused data are integrated.

[0147] A deep recurrent neural network (DRNN) is selected as the basic architecture. A 3-layer DRNN is used, with 128 Bi-LSTM units in each layer. Residual connections are set between every two Bi-LSTM layers, and skip connections are set between every four layers. A sparse self-attention layer is inserted on each skip connection to focus on important features. PReLU is added as the activation function after the skip connection to build a data fusion model.

[0148] Using a training set of 1000 samples, the model was trained until the loss function converged. Weights with cumulative gradients below 0.01 were identified and pruned, reducing the model parameters by 30%. The pruned model was fine-tuned for 100 cycles using the SGD optimizer with a learning rate of 0.001.

[0149] Organize multiple classification fusion data into a three-dimensional tensor (100 samples, 10 time steps, 50 features).

[0150] The Bi-LSTM layer extracts the temporal features. The enhanced fusion features are obtained through residual connection and skip connection, as well as PReLU activation function. The enhanced fusion features are weighted concatenated in the fully connected layer to obtain the final fusion features.

[0151] In actual applications, drones monitor the environment and can identify ground targets with 95% accuracy while effectively avoiding obstacles in complex terrain. In one environmental monitoring, the drone successfully identified 20 required targets in a complex mountainous environment and avoided 5 unrecorded obstacle collisions, ensuring the safety and effectiveness of environmental monitoring.

[0152] It should be noted that, in the present invention, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0153] The above embodiments are merely examples of the present invention and do not limit the protection scope of the present invention. All designs that are the same or similar to the present invention fall within the protection scope of the present invention.

Claims

1. A data fusion method based on deep learning, characterized in that: The following steps are involved: Classify various sensor data acquired in the unmanned system according to the set classification rules to obtain classified data; Preprocess the classified data based on their basic attributes to obtain the preprocessed classified data; The preprocessed classification data is fused at the data level to obtain multiple classification fusion data; Based on the deep recurrent neural network DRNN, combined with the bidirectional long short-term memory unit Bi-LSTM and the sparse attention mechanism, a data fusion model is constructed, and the data fusion model is optimized by model pruning to obtain an optimized data fusion model; Inputting multiple classified fusion data into the optimized data fusion model for secondary fusion to obtain final fusion features; The construction of the data fusion model includes the following steps: Select the deep recurrent neural network DRNN as the basic architecture, replace its neurons with bidirectional long short-term memory units Bi-LSTM, and set the parameters of the bidirectional long short-term memory unit Bi-LSTM, including the weights and biases of the forget gate, input gate, cell state, and output gate; Add residual connections between each bidirectional long short-term memory unit Bi-LSTM layer formed; A skip connection is added between every four bidirectional long short-term memory unit Bi-LSTM layers, a sparse self-attention mechanism is inserted into each skip connection, and PReLU is added as the activation function after the skip connection to build a data fusion model.

2. The data fusion method based on deep learning according to claim 1, characterized in that: The step of optimizing the data fusion model by model pruning to obtain an optimized data fusion model comprises the following steps: Acquire a training set to train the data fusion model until convergence, and obtain a trained data fusion model; Based on the trained data fusion model, select the pruning strategy as gradient pruning, and obtain the gradient and cumulative gradient of each weight; Based on the accumulated gradient, set the pruning threshold and identify the parameters that can be pruned; Perform pruning processing according to the accumulated gradient and the set pruning threshold to obtain the pruned data fusion model; The pruned data fusion model is fine-tuned and optimized through SGD to obtain the optimized data fusion model.

3. The data fusion method based on deep learning according to claim 2, characterized in that: The step of obtaining the gradient of each weight comprises the following steps: Obtain the loss between the predicted output and the true label calculated by the loss function in the trained data fusion model; Initialize the gradients in the trained data fusion model; Perform backpropagation and obtain the gradient of the loss with respect to each weight.

4. The data fusion method based on deep learning according to claim 3, characterized in that: The step of inputting the multiple classified fusion data into the optimized data fusion model for secondary fusion comprises the following steps: Arrange the multiple classified fusion data into three-dimensional tensors respectively, wherein the dimensions of the three-dimensional tensors are the number of samples, the time step and the number of features respectively, to obtain multiple classified three-dimensional fusion data; Input multiple classified three-dimensional fusion data into the bidirectional long short-term memory unit LSTM layer to obtain multiple time series features; The plurality of time series features are transmitted through residual connections, and a plurality of enhanced fusion features are obtained through skip connections and PReLU activation functions; The multiple enhanced fusion features are weightedly concatenated in the fully connected layer to obtain the final fusion feature.

5. The data fusion method based on deep learning according to claim 1, characterized in that: The classification rule is to classify according to sensor type and data usage; The step of obtaining the classified data comprises the following steps: Classifying the multiple sensor data acquired in the unmanned system according to the sensor type to obtain first classified data; The second classified data is obtained by classifying various sensor data in the unmanned system according to the data usage.

6. The data fusion method based on deep learning according to claim 5, characterized in that: The method of preprocessing the classified data based on the basic attributes thereof and fusing the preprocessed classified data at the data level comprises the following steps: Perform data alignment and time synchronization on the first classification data and the second classification data respectively to obtain synchronized first classification data and synchronized second classification data; Performing spatial registration and denoising on the synchronized first classification data and the synchronized second classification data respectively, to obtain the registered and denoised synchronized first classification data and the synchronized second classification data; The registered and denoised synchronous first classification data and the synchronous second classification data are respectively standardized, interpolated and filled with missing values ​​to obtain preprocessed first classification data and second classification data; Performing dynamic weighted fusion on the preprocessed first classification data and the second classification data respectively to obtain dynamically fused first classification data and second classification data; Performing data dimension reduction and feature selection on the dynamically fused first classification data and the synchronized second classification data respectively to obtain a plurality of first features and a plurality of second features; A plurality of first features and a plurality of second features are weightedly spliced ​​to obtain first fused data and second fused data, and a plurality of classified fused data are integrated.

7. A data fusion system based on deep learning, characterized in that: include: The data analysis and processing module is used to classify the various sensor data obtained in the unmanned system according to the set classification rules to obtain classified data; Preprocess the classified data based on their basic attributes to obtain the preprocessed classified data; The preprocessed classification data is fused at the data level to obtain multiple classification fusion data; A fusion model construction module is used to select a deep recurrent neural network DRNN as a basic architecture, replace its neurons with bidirectional long short-term memory units Bi-LSTM, and set the parameters of the bidirectional long short-term memory units Bi-LSTM, including the weights and biases of the forget gate, input gate, cell state, and output gate; add a residual connection between each formed bidirectional long short-term memory unit Bi-LSTM layer; add a jump connection between every four bidirectional long short-term memory unit Bi-LSTM layers, insert a sparse self-attention mechanism into each jump connection, add PReLU as an activation function after the jump connection, build a data fusion model, optimize the data fusion model through model pruning, and obtain an optimized data fusion model; The fusion data acquisition module is used to input multiple classified fusion data into the optimized data fusion model for secondary fusion to obtain the final fusion features.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Classification and identification method for low, slow small targets

    CN112434643A

  • Method, device and equipment for determining blood pressure based on PPG signal and storage medium

    CN117257256A