A Fusion Method for Multimodal Data Complementation, Error Correction and Tolerance Loss Mechanisms
Through the combination of adaptive weighting and deep learning models, the problems of low efficiency, poor accuracy and poor robustness in traditional multimodal data fusion are solved, and efficient and stable data fusion in complex environments is achieved, suitable for object detection and object recognition tasks.
Patent Information
- Application Number
- CN202510472091.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Traditional multimodal data fusion technology has problems such as low efficiency, poor accuracy and poor robustness when processing different sensor data. It is especially difficult to meet the needs of real-time and high-resolution scenarios in complex environments, and traditional methods cannot effectively solve the problems of information missing, registration error and noise superposition.
Adaptive weighting mechanism and deep learning model are adopted, combined with Generative Adversarial Network (GAN) and Gaussian Process Regression (GPR) technology, and through Transformer's multimodal feature fusion framework, modal weights are dynamically adjusted, missing data is generated, registration errors are corrected, and the robustness and stability of the system are enhanced.
It improves the accuracy and robustness of multimodal data fusion, can operate stably in complex environments, adapt to sensor characteristics and environmental changes, and improves the accuracy of object detection and object recognition.
Smart Images

Figure CN120030499B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a fusion method for multi-modal data complementation, error correction and loss tolerance mechanisms. Background Art
[0002] A fusion method for multi-modal data complementation, error correction and loss tolerance mechanisms is an advanced technology for processing and combining data streams from multiple sensors (such as microwave, infrared, visible light, etc.). Multi-modal data fusion involves synthesizing different information obtained from different sensors to obtain a more comprehensive and accurate scene representation. Each type of sensor data has its unique advantages and limitations. For example, microwave data can provide stable target information in complex climates and low-light environments, while infrared data can show the temperature distribution of objects, providing in-depth information about the object state and being unaffected by lighting conditions.
[0003] Traditional multi-modal data fusion technologies have problems of low efficiency, poor accuracy and poor robustness when processing large-scale data from different sensors. This is mainly because data from different modalities have essential differences during the acquisition process, such as differences in resolution, frame rate, signal-to-noise ratio, etc., which may lead to problems such as information loss, registration error and noise superposition during the fusion process. In many practical applications, especially tasks involving real-time requirements and high-resolution scenarios, such as autonomous driving, intelligent monitoring and military reconnaissance, traditional fusion methods cannot meet the requirements of being fast, efficient and of high quality.
[0004] In addition, problems such as information loss, inter-modal registration error and sensor noise will also affect the effect of data fusion. When facing factors such as environmental changes, uneven lighting or sensor failures, the quality and integrity of the data may be affected, thereby reducing the stability and accuracy of the fusion result. Traditional technologies usually rely on rules or preset models to solve these problems, but these methods cannot fully utilize the data characteristics in complex environments and are also difficult to flexibly respond in dynamically changing scenarios.
[0005] To solve these problems, multi-modal data fusion methods based on deep learning have gradually become a research hotspot. Deep learning models can automatically learn the correlation and mutual complementarity between modalities, thus effectively solving problems such as information loss, registration error and noise. In addition, by introducing technologies such as adaptive weighting mechanisms, generative adversarial networks (GANs) and Gaussian process regression (GPR), deep learning-driven multi-modal data fusion methods can dynamically adjust the modal weights, generate missing modal data, correct registration errors, and improve the robustness and accuracy of the system in complex environments. Summary of the Invention
[0006] Aiming at the above-mentioned technical deficiencies, the purpose of the present invention is to provide a fusion method for multi-modal data complementary, error correction and loss tolerance mechanisms. This method can effectively fuse data from different modalities (such as microwave, infrared, visible light, etc.), solve the problems of information loss, registration error and noise interference, and improve the accuracy, robustness and stability of data fusion. Guided by the adaptive weighting mechanism and deep learning model, the present invention can operate stably in complex environments, adapt to different sensor characteristics and environmental changes, and is widely used in tasks with high-precision requirements such as target detection and object recognition.
[0007] To solve the above technical problems, the present invention adopts the following technical solutions:
[0008] The present invention provides a fusion method for multi-modal data complementary, error correction and loss tolerance mechanisms, including the following steps:
[0009] Step 1: Collect microwave data stream and infrared data stream in the same scene;
[0010] Step 2: Extract features from the microwave data and infrared data stream, and then adopt the multi-modal feature fusion framework of Transformer to perform complementary fusion of multi-modal data by adaptively learning the weights of features, ensuring that when the information of a certain modality is missing, other modalities can effectively supplement the missing part;
[0011] Step 3: Design a multi-modal generation network (MMGN) based on the generative adversarial network (GAN) to generate lost modality data and reduce the negative impact of information loss on the fusion result;
[0012] Step 4: Adopt a deep learning-driven error correction mechanism and Gaussian process regression (GPR) technology to automatically identify and correct the registration error caused by sensors, and improve the registration accuracy by optimizing the error distribution.
[0013] Step 5: Introduce a robust loss function (such as Huber loss function) to perform fault tolerance processing for information loss and error correction, and enhance the robustness and accuracy of the system.
[0014] Step 6: Through the adaptive weighting mechanism, dynamically adjust the weights of different modalities in the fusion to further enhance the stability of the system in complex environments;
[0015] Step 7: Generate the final fused data and use it for target detection and recognition tasks.
[0016] Furthermore, the steps of extracting features from the microwave data stream and infrared data stream in Step 2 are as follows:
[0017] Step 2.1-1. First, assume that the original data of the microwave data and the infrared data stream are respectively and , which respectively correspond to the microwave and infrared image sequences. Since data of different modalities may have different resolutions, scales, and noises, before processing, it is necessary to normalize the data: make the value ranges of each modality of data unified to reduce the scale differences in subsequent processing:
[0018] ;
[0019] wherein, and are the mean and standard deviation of the microwave data, and are the mean and standard deviation of the infrared data;
[0020] Step 2.1-2. For the data of each modality (microwave and infrared), use deep learning methods such as convolutional neural networks (CNNs) to extract high-dimensional features. Feature extraction is performed through multiple convolutional layers. Each layer extracts more abstract features from the input data until the final feature representation is reached;
[0021] Assume that the feature representation of the microwave data is , and the feature representation of the infrared data is . The process of their extraction through the CNN can be expressed as:
[0022] ;
[0023] wherein, and are the convolutional neural networks for the microwave and infrared data, respectively generating the feature representations of each modality.
[0024] Step 2.1-3. To ensure that the features of each modality have a unified distribution and avoid certain modality features dominating in subsequent processing, the extracted features are usually standardized, and this process is completed by subtracting the mean and dividing by the standard deviation:
[0025] ;
[0026] wherein, and are the mean and standard deviation of the microwave data features, and are the mean and standard deviation of the infrared data features;
[0027] After completing the above-mentioned feature extraction and preprocessing steps, we obtain the feature representations of the two modalities. and , which will be fused in the subsequent Transformer model. For subsequent self-attention calculations, we use the features of these two modalities as inputs and feed them into the Transformer framework for cross-modal fusion.
[0028] Furthermore, the steps of adopting the multi-modal feature fusion framework of Transformer for the microwave data stream and the infrared data stream in Step 2 are as follows:
[0029] Step 2.2-1, Microwave data features and infrared data features learn the relationship between them through self-attention calculation in the Transformer architecture to obtain adaptive weights. These weighted features will ultimately be fused to generate a joint feature representation. The goal of the self-attention mechanism is to calculate the correlation between each input feature and assign a weight to each feature, which reflects the importance of the feature for the current task. The specific calculation formula is as follows:
[0030] ;
[0031] Where: is the query matrix, representing the features of the current modality; is the key matrix, representing the features of another modality; is the value matrix, representing the actual feature information of the modality; is the dimension of the key, used to scale the inner product. Through the above calculation, the system assigns a dynamically calculated weight to each feature, and these weights change in each self-attention module, which can reflect the importance of different modalities in the final fusion.
[0032] Step 2.2-2, Learn the correlation between different modalities through the self-attention mechanism of Transformer. We design joint query, key, and value matrices for cross-modal adaptive learning. The calculation formula is as follows:
[0033] ;
[0034] ;
[0035] Where, , , are the query, key, and value of the first modality (such as microwave); , , are the query, key, and value of the second modality (e.g., infrared); , , are the weight matrices for training, used for the calculations of query, key, and value respectively.
[0036] Calculate the cross-modal correlation scores through the self-attention mechanism:
[0037] ;
[0038] Similarly, we can also calculate to achieve information flow and fusion between modalities.
[0039] Furthermore, the steps for weighting the adaptive learning features in the microwave data stream and the infrared data stream in Step 2 are as follows:
[0040] Step 2.3-1: Through the self-attention mechanism of the Transformer, we can dynamically calculate the weights of each modality. The model will automatically adjust the weights according to the contribution of each modality to the final fusion result. These weights can be expressed as:
[0041] ;
[0042] These weights and represent the "importance" of each modality - that is, the contribution ratio of each modality to the final fusion result;
[0043] Step 2.3-2: After calculating the weights of each modality, we can perform weighted fusion on the features of the two modalities to ensure that the dominant features of each modality make greater contributions in the final fusion. The features of the final fusion are expressed as:
[0044] ;
[0045] Among them, is the fused feature data, and are the modality weights calculated through the self-attention mechanism.
[0046] Furthermore, the steps for generating the missing modality data in the microwave data stream and the infrared data stream in Step 3 are as follows:
[0047] Step 3.1-1. In the multi-modal generation network (MMGN), we adopt the generative adversarial network to fill in the missing modal data. Suppose our system needs to process microwave data and infrared data, but due to certain reasons, the infrared data may be missing. In this case, we can use the MMGN to generate the missing infrared data;
[0048] The generator of the multi-modal generation network receives a noise vector and the known modal data (such as microwave data) as inputs, and generates a "fake" missing modal data (such as infrared data). Suppose the generator is , its input is microwave data and noise, and the output is the generated infrared data,
[0049] ;
[0050] Where: is the known modal (such as microwave) data; is the random noise, usually sampled from a Gaussian distribution; is the generated missing modal data (such as infrared data).
[0051] The discriminator is responsible for judging whether the generated infrared data is similar to the real infrared data. The goal of the discriminator is to distinguish whether the input data comes from the real infrared data or the infrared data generated by the generator, ;
[0052] ;
[0053] Step 3.1-2. The loss function of the generative adversarial network is usually based on the idea of adversarial training. The goal of the loss function is to minimize the difference between the "fake" data generated by the generator and the real data, while maximizing the ability of the discriminator to distinguish between real and fake data. The loss function of the generator:
[0054] ;
[0055] The loss function of the discriminator:
[0056] ;
[0057] Where: is the real distribution of the infrared data, is the generated infrared data;
[0058] Step 3.1-3. Through training, the generator continuously optimizes its parameters so that it can generate increasingly realistic missing modality data, while the discriminator continuously improves its ability to accurately distinguish between real data and generated data. In this process, the generator and the discriminator engage in a game through adversarial training, and ultimately the generator can generate high-quality missing modality data to supplement the missing parts in the original data.
[0059] Further, the steps of adopting a deep learning-driven error correction mechanism for the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.1-1. By designing a deep learning model, the system can automatically detect errors in the registration process. This process first registers and compares data from different modalities and calculates the registration difference between the two modalities. The detection of registration errors can be achieved by calculating the pixel differences or structural differences between different modalities. Suppose we have two modality images and . After their initial registration, the difference metric between them is calculated:
[0060] ;
[0061] where represents the difference metric between the two modalities, represents the L2 norm, which measures the Euclidean distance between pixels. If has a large value, it indicates a large registration error;
[0062] Step 4.1-2. Once the error is detected, the system will adjust the spatial position of the data according to these errors. A deep learning network is used to learn and optimize the correction method of these errors to align the data of the two modalities.
[0063] During the error correction process, we use a deep learning network to learn how to adjust the offset in the image and correct the alignment of the data. The corrected data can be expressed as:
[0064] ;
[0065] where is the corrected infrared data, and is the correction amount calculated by the deep learning model.
[0066] Further, the steps of adopting Gaussian process regression (GPR) technology for the microwave data stream and the infrared data stream in Step 4 are as follows:
[0067] Step 4.2-1. Gaussian Process Regression (GPR) is introduced to further optimize the registration error. GPR is a powerful non-parametric Bayesian method that can provide a smooth correction path for the error by calculating the potential distribution of the registration error.
[0068] The goal of Gaussian Process Regression is to estimate the probability distribution of the registration error by learning the distribution of historical data and correct the registration error according to this distribution. The formula of GPR is as follows:
[0069] ;
[0070] where: is the predicted registration error; is the mean function, usually taken as zero; is the covariance function (kernel function) that describes the correlation between input points. The posterior distribution of GPR allows us to optimize the registration error and thus obtain a smooth and accurate registration correction value.
[0071] Step 4.2-2. Given the training data and a new point , we hope to predict (the correction amount of the registration error). The prediction formula of GPR is:
[0072] ;
[0073] ;
[0074] where: is the predicted mean (i.e., the correction value); is the predicted variance, representing the uncertainty of the prediction; is the covariance between the new point and the training data point ; is the covariance matrix between the training data points. Through this process, GPR can generate a more accurate correction value based on the existing registration error data, improving the accuracy of registration.
[0075] Furthermore, the steps of introducing the robust loss function for the microwave data stream and the infrared data stream in Step 5 are as follows:
[0076] Step 5.1-1. When training a deep learning model, using the standard Mean Squared Error (MSE) loss function may be affected by outliers in the data because the mean squared error is very sensitive to large errors and may lead to model instability. To solve this problem, we use a robust loss function, such as the Huber loss function, to reduce the impact of outliers on the model training process.
[0077] The Huber loss function combines the advantages of the mean squared error (MSE) and the mean absolute error (MAE). It can use the squared error for small errors and the linear error for large errors, thus reducing the impact of outliers while retaining the optimization effect for small errors. The formula of the Huber loss function is as follows:
[0078] ;
[0079] where: is the true value; is the predicted value; is a hyperparameter that determines when to switch from the squared error to the linear error.
[0080] Furthermore, the steps of adopting the adaptive weighting mechanism for the microwave data stream and the infrared data stream in Step 6 are as follows:
[0081] Step 6.1-1: To dynamically adjust the weight of each modality, first, it is necessary to calculate the information quality of each modality, which can be achieved through various methods. For example, assuming for each modality , we can measure the quality of this modality through a quality assessment function . The quality assessment function may be based on multiple factors such as error detection, information loss, signal-to-noise ratio, etc.;
[0082] Step 6.1-2: According to the quality of each modality, we can calculate the weight of this modality in the fusion process. Generally, modalities with higher information quality will be assigned larger weights, while modalities with lower information quality will receive smaller weights. A common adaptive weighting mechanism is to calculate the weight of each modality by normalizing the quality index.
[0083] Let be the weight of modality , and the quality assessment function satisfies:
[0084] ;
[0085] where: is the quality score of modality ; is the total number of modalities. In this way, the weight will be dynamically adjusted according to the quality of each modality, and the sum is 1.
[0086] Step 6.1-3: After calculating the weight of each modality, weighted fusion can be performed. Assuming the feature representation of modality is , then the fused feature It can be expressed as a weighted average:
[0087] ;
[0088] Where: is the final feature representation after fusion; is the weight of modality ; is the feature representation of modality .
[0089] Through this weighted fusion, the system can automatically adjust its contribution to the final fusion result according to the quality of each modality.
[0090] The beneficial effects of the present invention are as follows: The present invention collects microwave data streams and infrared data streams under the same scene; preprocesses the collected microwave data streams and infrared data streams, including denoising and contrast enhancement operations, and uses a generative adversarial network to generate missing modality data; adopts an adaptive weighting mechanism to improve data fusion performance; performs weighted fusion on the processed microwave data and infrared data through a fusion algorithm to generate a final fusion data stream. It realizes the selection of appropriate preprocessing operations for different modality data streams, reduces noise while enhancing image details, and ensures the integrity and consistency of data; by integrating an adaptive weighting mechanism into the deep learning framework, dynamically adjusts the weight of each modality in the fusion, especially in the scenario where the quality differences of different sensor information are large, effectively improves the accuracy and robustness of data fusion; through a deep learning-driven error correction mechanism, combines Gaussian process regression to automatically correct the registration error, further improves the accuracy of the fusion result; performs a fusion operation based on a convolutional network on the corrected different modality data streams, effectively integrates microwave and infrared information, greatly improves the processing efficiency and quality of multi-modal data fusion, and expands the application scope of this technology in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0092] Figure 1 is the flow diagram of the present invention;
[0093] Figure 2 is the multi-modal feature fusion flow chart of the present invention using a Transformer network framework;
[0094] Figure 3 This is the flowchart of information loss compensation based on GAN adopted by the present invention;
[0095] Figure 4 This is the flowchart of registration error correction using Gaussian process regression and deep learning adopted by the present invention. Specific embodiments
[0096] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0097] Embodiment, as Figures 1 - 4 shown, a fusion method for multi-modal data complementation, error correction and loss tolerance mechanism includes the following steps:
[0098] Step 1. Collect microwave data stream and infrared data stream under the same scene;
[0099] Step 2. Extract features from the microwave data and infrared data streams, and then adopt the multi-modal feature fusion framework of Transformer to perform complementary fusion of multi-modal data by adaptively learning the weights of features, ensuring that when information in a certain modality is missing, other modalities can effectively supplement the missing part;
[0100] Step 3. Design a multi-modal generation network (MMGN) based on the generative adversarial network (GAN) to generate missing modality data and reduce the negative impact of information loss on the fusion result;
[0101] Step 4: Adopt a deep learning-driven error correction mechanism and Gaussian process regression (GPR) technology to automatically identify and correct the registration error caused by sensors, and improve the registration accuracy by optimizing the error distribution.
[0102] Step 5: Introduce a robust loss function (such as Huber loss function) to perform fault tolerance processing for information loss and error correction, and enhance the robustness and accuracy of the system.
[0103] Step 6. Through an adaptive weighting mechanism, dynamically adjust the weights of different modalities in the fusion to further enhance the stability of the system in complex environments;
[0104] Step 7. Generate the final fusion data and use it for target detection and recognition tasks.
[0105] As a preferred embodiment, the steps of performing feature extraction operations on the microwave data stream and the infrared data stream in Step 2 are as follows: Step 2.1-1. First, assume that the original data of the microwave data and the infrared data stream are respectively and , which respectively correspond to the microwave and infrared image sequences. Since data of different modalities may have different resolutions, scales, and noises, before processing, it is necessary to perform data normalization on them: making the value ranges of each modality data unified, so as to reduce the scale differences of the data during subsequent processing,
[0106] ;
[0107] wherein, and are the mean and standard deviation of the microwave data, and are the mean and standard deviation of the infrared data;
[0108] Step 2.1-2. For the data of each modality (microwave and infrared), use deep learning methods such as convolutional neural network (CNN) to extract high-dimensional features; feature extraction is performed through multiple convolutional layers. Each layer extracts more abstract features from the input data until the final feature representation is reached; assume that the feature representation of the microwave data is , and the feature representation of the infrared data is . The process of their extraction through the CNN can be expressed as:
[0109] ;
[0110] wherein, and are the convolutional neural networks for the microwave and infrared data, respectively generating the feature representations of each modality;
[0111] Step 2.1-3. In order to ensure that the features of each modality have a unified distribution and avoid some modality features dominating in subsequent processing, usually perform standardization processing on the extracted features. This process is completed by subtracting the mean and dividing by the standard deviation:
[0112] ;
[0113] wherein, and are the mean and standard deviation of the microwave data features, and are the mean and standard deviation of the infrared data features. After completing the above feature extraction and preprocessing steps, we obtain the feature representations of the two modalities. and , which will be fused in the subsequent Transformer model. For subsequent self-attention calculation, we use the features of these two modalities as inputs and feed them into the Transformer framework for cross-modal fusion.
[0114] As a preferred implementation, the steps of using the multi-modal feature fusion framework of Transformer for microwave data stream and infrared data stream in Step 2 are as follows:
[0115] Step 2.2-1, Microwave data features and infrared data features Learn the relationship between them through self-attention calculation in the Transformer architecture to obtain adaptive weights. These weighted features will ultimately be fused to generate a joint feature representation. The goal of the self-attention mechanism is to assign a weight to each input feature by calculating the correlation between each input feature. This weight reflects the importance of the feature for the current task. The specific calculation formula is as follows:
[0116] ;
[0117] Where: is the query matrix, representing the features of the current modality; is the key matrix, representing the features of another modality; is the value matrix, representing the actual feature information of the modality; is the dimension of the key, used to scale the inner product. Through the above calculation, the system assigns a dynamically calculated weight to each feature, and these weights change in each self-attention module, which can reflect the importance of different modalities in the final fusion.
[0118] Step 2.2-2, Learn the correlation between different modalities through the self-attention mechanism of Transformer. We design joint query, key, and value matrices for cross-modal adaptive learning. The calculation formula is as follows:
[0119] ;
[0120] ;
[0121] Where, , , are the query, key, and value of the first modality (such as microwave); , , are the query, key, and value of the second modality (e.g., infrared); , , are the weight matrices for training, used for calculating the query, key, and value respectively. The cross-modal correlation score is calculated through the self-attention mechanism:
[0122] ;
[0123] Similarly, we can also calculate to achieve information flow and fusion between modalities.
[0124] As a preferred implementation, the steps for weighting the adaptive learning features in the microwave data stream and the infrared data stream in Step 2 are as follows:
[0125] Step 2.3-1: Through the self-attention mechanism of the Transformer, we can dynamically calculate the weights of each modality. The model will automatically adjust the weights according to the contribution of each modality to the final fusion result. These weights can be expressed as:
[0126] ;
[0127] These weights and represent the "importance" of each modality - that is, the contribution ratio of each modality to the final fusion result.
[0128] Step 2.3-2: After calculating the weights of each modality, the features of the two modalities can be weighted and fused to ensure that the dominant features of each modality make a greater contribution to the final fusion. The features of the final fusion are expressed as:
[0129] ;
[0130] Among them, is the fused feature data, and are the modality weights calculated through the self-attention mechanism.
[0131] As a preferred implementation, the steps for generating missing modality data in the microwave data stream and the infrared data stream in Step 3 are as follows:
[0132] Step 3.1-1: In the multi-modal generation network (MMGN), we use a generative adversarial network to fill in the missing modality data. Suppose our system needs to process microwave data and infrared data, but due to certain reasons, the infrared data may be missing. In this case, we can use the MMGN to generate the missing infrared data;
[0133] The generator of the multi-modal generation network receives a noise vector and known modal data (such as microwave data) as inputs, and generates a "fake" missing modal data (such as infrared data). Assuming the generator is , its inputs are microwave data and noise, and the output is the generated infrared data,
[0134] ;
[0135] Where: is the known modal (such as microwave) data; is the random noise, usually sampled from a Gaussian distribution; is the generated missing modal data (such as infrared data);
[0136] The discriminator is responsible for judging whether the generated infrared data is similar to the real infrared data. The goal of the discriminator is to distinguish whether the input data comes from the real infrared data or the infrared data generated by the generator:
[0137] ;
[0138] ;
[0139] Step 3.1-2. The loss function of the generative adversarial network is usually based on the idea of adversarial training. The goal of the loss function is to minimize the difference between the "fake" data generated by the generator and the real data, while maximizing the ability of the discriminator to distinguish between real and fake data. The loss function of the generator:
[0140] ;
[0141] The loss function of the discriminator:
[0142] ;
[0143] Where: is the true distribution of the infrared data, is the generated infrared data;
[0144] Step 3.1-3. Through training, the generator continuously optimizes its parameters so that it can generate more and more realistic missing modal data, and the discriminator continuously optimizes its ability to ensure that it can accurately distinguish between real and generated data. In this process, the generator and the discriminator play against each other through adversarial training. Eventually, the generator can generate high-quality missing modal data to supplement the missing part of the original data;
[0145] As a preferred implementation, the steps of adopting a deep learning-driven error correction mechanism for the microwave data stream and the infrared data stream in Step 4 are as follows:
[0146] Step 4.1-1: By designing a deep learning model, the system can automatically detect errors in the registration process. This process first registers and compares data from different modalities, calculates the registration difference between the two modalities, and the detection of registration errors can be achieved by calculating the pixel difference or structural difference between different modalities. Suppose we have two-modal images and , after their preliminary registration, calculate the difference metric between the two:
[0147] ;
[0148] Among them, represents the difference metric between the two modalities, represents the L2 norm, which measures the Euclidean distance between pixels. If has a large value, it indicates that there is a large registration error;
[0149] Step 4.1-2: Once the error is detected, the system will adjust the spatial position of the data according to these errors. Use a deep learning network to learn and optimize the correction method of these errors to align the data of the two modalities.
[0150] During the error correction process, we use a deep learning network to learn how to adjust the offset in the image, correct the alignment of the data, and the corrected data can be expressed as:
[0151] ;
[0152] Among them, is the corrected infrared data, is the correction amount calculated by the deep learning model.
[0153] As a preferred implementation, the steps of adopting the Gaussian process regression (GPR) technology for the microwave data stream and the infrared data stream in Step 4 are as follows:
[0154] Step 4.2-1: Gaussian process regression (GPR) is introduced to further optimize the registration error. GPR is a powerful non-parametric Bayesian method. By calculating the potential distribution of the registration error, it can provide a smooth correction path for the error. The goal of Gaussian process regression is to estimate the probability distribution of the registration error by learning the distribution of historical data and correct the registration error according to this distribution. The formula of GPR is as follows:
[0155] ;
[0156] Wherein: is the predicted registration error; is the mean function, usually taken as zero; is the covariance function (kernel function), which describes the correlation between input points. The posterior distribution of GPR allows us to optimize the registration error, and then obtain a smooth and accurate registration correction value.
[0157] Step 4.2-2. Given the known training data and a new point , we hope to predict (the correction amount of the registration error). The prediction formula of GPR is:
[0158] ;
[0159] ;
[0160] Wherein: is the predicted mean (i.e., the correction value); is the predicted variance, indicating the uncertainty of the prediction; is the covariance between the new point and the training data point ; is the covariance matrix between the training data points. Through this process, GPR can generate more accurate correction values based on the existing registration error data, improving the accuracy of registration.
[0161] As a preferred implementation manner, the steps of introducing a robust loss function into the microwave data stream and the infrared data stream in Step 5 are as follows:
[0162] Step 5.1-1. When training a deep learning model, using the standard mean squared error (MSE) loss function may be affected by outliers in the data, because the mean squared error is very sensitive to large errors and may cause the model to be unstable. To solve this problem, we use a robust loss function, such as the Huber loss function, to reduce the impact of outliers on the model training process. The Huber loss function combines the advantages of the mean squared error (MSE) and the mean absolute error (MAE). It can use the squared error for small errors and the linear error for large errors, thereby reducing the impact of outliers while retaining the optimization effect for small errors. The formula of the Huber loss function is as follows:
[0163] ;
[0164] Wherein: is the true value; is the predicted value; is a hyperparameter that determines when to switch from squared error to linear error.
[0165] As a preferred implementation, the step of adopting an adaptive weighting mechanism for the microwave data stream and the infrared data stream in Step 6 is as follows:
[0166] Step 6.1-1: To dynamically adjust the weight of each modality, it is first necessary to calculate the information quality of each modality. This can be achieved through various methods. For example, assuming for each modality , we can use a quality assessment function to measure the quality of that modality. The quality assessment function may be based on multiple factors such as error detection, information loss, signal-to-noise ratio, etc.;
[0167] Step 6.1-2: According to the quality of each modality, we can calculate the weight of that modality in the fusion process. Generally, modalities with higher information quality will be assigned larger weights, while modalities with lower information quality will receive smaller weights. A common adaptive weighting mechanism is to calculate the weight of each modality by normalizing the quality index;
[0168] Let be the weight of modality , and the quality assessment function satisfies:
[0169] ;
[0170] where: is the quality score of modality ; is the total number of modalities. In this way, the weight will be dynamically adjusted according to the quality of each modality, and the sum is 1.
[0171] Step 6.1-3: After calculating the weight of each modality, weighted fusion can be performed. Assume that the feature representation of modality is , then the fused feature can be expressed as a weighted average:
[0172] ;
[0173] where: is the final feature representation after fusion; is the weight of modality ; is the feature representation of modality .
[0174] Through this weighted fusion, the system can automatically adjust its contribution to the final fusion result according to the quality of each modality.
[0175] Through the deep learning-driven multi-modal data fusion method of the present invention, an efficient complementary, error-correcting and loss-tolerant mechanism is proposed, which solves the deficiencies of traditional technologies in information loss, registration error and noise processing. The feature fusion framework based on the Transformer architecture realizes information complementarity between modalities, the generative adversarial network is used to make up for the lack of modal data, the deep learning-driven error-correcting mechanism combined with Gaussian process regression optimizes the registration accuracy, and at the same time, the fault tolerance and stability of the system are enhanced through the robust loss function and the dynamic weighting mechanism. This method dynamically adjusts the modal weights to ensure the dominant role of high-quality data in the fusion result, significantly improves the accuracy and robustness of data fusion, is widely applicable to fields such as target detection, environmental perception and scene analysis, provides a reliable solution for multi-modal data processing in complex tasks, and has important academic value and application prospects.
[0176] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A fusion method for multi-modal data complementarity, error correction and loss tolerance mechanism, characterized in that, Including the following steps: Step 1. Collect microwave data stream and infrared data stream under the same scene; Step 2. Extract features from the microwave data and infrared data stream, and then adopt the multi-modal feature fusion framework of Transformer to perform complementary fusion of multi-modal data by adaptively learning the weights of features. The steps for extracting features from the microwave data stream and infrared data stream are as follows: Step 2.1-1. First, assume that the original data of the microwave data and the infrared data stream are respectively and , corresponding to the microwave and infrared image sequences respectively. Since data of different modalities may have different resolutions, scales, and noises, before processing, perform data normalization on them: ; Among them, and are the mean and standard deviation of microwave data, and are the mean and standard deviation of infrared data; Step 2.1-2. For the data of each modality, use the deep learning method of convolutional neural network to extract high-dimensional features; Suppose the feature representation of microwave data is , and the feature representation of infrared data is . The process of extraction by CNN is represented as: ; Among them, and are convolutional neural networks for microwave and infrared data, respectively generating feature representations for each modality; Step 2.1-3. Standardize the extracted features, and this process is completed by subtracting the mean and dividing by the standard deviation: ; Among them, and are the mean and standard deviation of microwave data features, and are the mean and standard deviation of infrared data features; After completing the above-mentioned feature extraction and preprocessing steps, feature representations of two modalities are obtained and , which will be fused in the subsequent Transformer model. Taking the features of these two modalities as inputs, they are fed into the Transformer framework for cross-modal fusion; The steps for adopting the multi-modal feature fusion framework of Transformer for the microwave data stream and infrared data stream are as follows: Step 2.2-1, Microwave data features and infrared data features Learn the relationship between them through self-attention calculation in the Transformer architecture to obtain adaptive weights. These weighted features are finally fused to generate a joint feature representation. The goal of the self-attention mechanism is to assign a weight to each feature by calculating the correlation between each input feature. The specific calculation formula is as follows: ; Wherein: is a query matrix, representing the features of the current modality; is a key matrix, representing the features of another modality; is a value matrix, representing the actual feature information of the modality; is the dimension of the key, used to scale the inner product; Step 2.2-2. Learn the correlation between different modalities through the self-attention mechanism of Transformer, design joint query, key and value matrices, and perform cross-modal adaptive learning. The calculation formula is as follows: ; ; Among them, , , are the query, key, and value of the first modality; , , are the query, key, and value of the second modality; , , are the weight matrices for training, which are used for the calculations of the query, key, and value respectively; calculate the cross-modal correlation score through the self-attention mechanism: ; Similarly, by calculating , the information flow and fusion between modalities are achieved; Step 3. Design a multi-modal generation network based on the generative adversarial network to generate missing modality data and reduce the negative impact of information loss on the fusion result; Step 4: Adopt a deep learning-driven error correction mechanism and Gaussian process regression technology to automatically identify and correct the registration error caused by the sensor, and improve the registration accuracy by optimizing the error distribution; Step 5: Introduce a robust loss function to perform fault tolerance processing for information loss and error correction, and enhance the robustness and accuracy of the system; Step 6. Through an adaptive weighting mechanism, dynamically adjust the weights of different modalities in the fusion to further enhance the stability of the system in a complex environment; Step 7. Generate the final fusion data and use it for target detection and recognition tasks.
2. The fusion method of a multi-modal data complementation, error correction and loss tolerance mechanism according to claim 1, characterized in that, The steps for adaptively learning the weights of features in the microwave data stream and infrared data stream in Step 2 are as follows: Step 2.3-1. Dynamically calculate the weights of each modality through the self-attention mechanism of Transformer. The model will automatically adjust the weights according to the contribution of each modality to the final fusion result. These weights are expressed as: ; These weights and represent the importance of each modality, that is, the contribution ratio of each modality to the final fusion result; Step 2.3-2. After calculating the weights of each modality, perform weighted fusion on the features of the two modalities. The finally fused features are expressed as: ; Among them, is the fused feature data, and are the modality weights calculated by the self-attention mechanism.
3. The fusion method of a multimodal data complementary, error correction and loss tolerance mechanism according to claim 1, characterized in that, The steps for generating missing modality data in the microwave data stream and infrared data stream in Step 3 are as follows: Step 3.1-1. In the multimodal generation network, a generative adversarial network is used to fill in the data of the missing modality. The generator of the multimodal generation network receives a noise vector and known modality data as inputs, and generates a "fake" missing modality data. Suppose the generator is , its input is microwave data and noise, and the output is the generated infrared data: ; Wherein: is known modal data; is random noise, usually sampled from a Gaussian distribution; is the generated missing modal data; Discriminator Responsible for judging the generated infrared data Whether it is similar to the real infrared data. The goal of the discriminator is to distinguish whether the input data comes from real infrared data or infrared data generated by the generator. The formula is as follows: ; ; Step 3.1-2. The loss function of the generative adversarial network is based on the idea of adversarial training. The goal of the loss function is to minimize the difference between the "fake" data generated by the generator and the real data, and at the same time maximize the ability of the discriminator to distinguish between real and fake data. The loss function of the generator is: ; The loss function of the discriminator is: ; Wherein: is the true distribution of infrared data, is the generated infrared data; Step 3.1-3. Through training, the generator continuously optimizes its parameters, and the discriminator continuously optimizes its ability. The generator and the discriminator play against each other through adversarial training. Eventually, the generator can generate high-quality missing modality data to supplement the missing part of the original data.
4. The fusion method of a multi-modal data complementary, error-correcting and loss-tolerant mechanism according to claim 1, characterized in that, The steps of using a deep learning-driven error correction mechanism for the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.1-1. The system automatically detects errors in the registration process by designing a deep learning model. This process first registers and compares data from different modalities, calculates the registration difference between the two modalities, and the detection of registration errors is achieved by calculating the pixel difference or structural difference between different modalities. Suppose we have images of two modalities and . After their preliminary registration, the difference metric between the two is calculated: ; Among them, represents the difference metric between two modalities, represents the L2 norm, which measures the Euclidean distance between pixels. If has a large value, it indicates a large registration error; Step 4.1-2: Once an error is detected, the system will adjust the spatial position of the data according to these errors, and use a deep learning network to learn and optimize the correction method of these errors, so that the data of the two modalities are aligned; during the error correction process, a deep learning network is used to learn how to adjust the offset in the image, correct the alignment of the data, and the corrected data is expressed as: ; Among them, is the corrected infrared data, is the correction amount calculated by the deep learning model.
5. A fusion method for a multimodal data complementary, error correction and error tolerance mechanism according to claim 1, characterized in that, The steps of using the Gaussian process regression technology for the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.2-1: Gaussian process regression is introduced to further optimize the registration error. Gaussian process regression estimates the probability distribution of the registration error by learning the distribution of historical data, and corrects the registration error according to this distribution. The formula of GPR is as follows: ; Wherein: is the predicted registration error; is the mean function, taken as zero; is the covariance function, which describes the correlation between input points; the posterior distribution of GPR allows the optimization of the registration error, and then a smooth and accurate registration correction value can be obtained; Step 4.2-2. Given the known training data , a new point is given to predict , that is, the correction amount of the registration error. The prediction formula of GPR is as follows: ; ; Wherein: is the predicted mean, i.e., the correction value; is the predicted variance, representing the uncertainty of the prediction; is the new point and the training data point the covariance between; is the covariance matrix between the training data points.
6. The fusion method of a multimodal data complementary, error correction and loss tolerance mechanism according to claim 1, characterized in that, The steps of introducing a robust loss function for the microwave data stream and the infrared data stream in Step 5 are as follows: Step 5.1-1: When training a deep learning model, a robust loss function is used to reduce the impact of outliers on the model training process: The loss function combines the advantages of the mean square error and the absolute error. It can use the square error for small errors and the linear error for large errors, thereby reducing the impact of outliers and retaining the optimization effect for small errors. The formula of the loss function is as follows: ; Wherein: is the true value; is the predicted value; is a hyperparameter that determines when to switch from squared error to linear error.
7. The fusion method of a multi-modal data complementary, error-correcting and loss-tolerant mechanism according to claim 1, characterized in that The steps of using an adaptive weighting mechanism for the microwave data stream and the infrared data stream in Step 6 are as follows: Step 6.1-1. To dynamically adjust the weights of each modality, it is first necessary to calculate the information quality of each modality. For each modality , a quality assessment function is used to measure the quality of the modality. The quality assessment function is based on multiple factors such as error detection, information loss, and signal-to-noise ratio; Step 6.1-2: Calculate the weight of each modality in the fusion process according to the quality of each modality. Let be the weight of the modality , and the quality evaluation function satisfies: ; Wherein: is the quality score of the modality ; is the total number of modalities; in this way, the weight will be dynamically adjusted according to the quality of each modality, and the sum is 1; Step 6.1-3. After calculating the weights of each modality, perform weighted fusion. Assume that the feature representation of modality is . Then the fused feature is expressed as a weighted average: ; Wherein: is the final feature representation after fusion; is the modality weight; is the modality feature representation.
Citation Information
Patent Citations
Multi-modal fusion method and device based on normalized mutual information, medium and equipment
CN111461176A
Brain dysfunction auxiliary evaluation method based on multi-modal data fusion
CN115553752A
Cited By
Feature analysis complementation system for multi-modal data
CN120930070A
Feature analysis complementary system of multi-modal data
CN120930070B
Intelligent detection method for processing cooking degree of fried basa fish based on multi-modal data fusion
CN121434911A
A multi-modal data fusion intelligent detection method for processing doneness of bocachico
CN121434911B