Multi-modal data complementation, error correction and capacitance loss mechanism fusion method
Through the deep learning-driven multimodal data fusion method, combined with Transformer, GAN, deep learning error-calculating mechanism and Gaussian process regression, the problems of low efficiency, poor accuracy and poor robustness of multimodal data fusion in traditional technologies are solved, and data fusion effect with high accuracy, robustness and stability are achieved.
Patent Information
- Application Number
- CN202510472091.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Traditional multimodal data fusion technology has problems of low efficiency, poor accuracy and poor robustness when processing large-scale data from different sensors, especially in complex environments and dynamically changing scenarios, which are difficult to meet the needs of fast, efficient and high quality.
The multimodal data fusion method driven by deep learning is adopted to achieve feature fusion between modals through the Transformer architecture, and the lost modal data is generated using a generative adversarial network (GAN), combining deep learning-driven error calibration mechanism and Gaussian process regression (GPR) to optimize registration accuracy, and enhance the fault tolerance and stability of the system through robust loss functions and adaptive weighting mechanisms.
It significantly improves the accuracy, robustness and stability of data fusion, and can dynamically adjust modal weights in complex environments to ensure the dominant role of high-quality data on the fusion results. It is suitable for the fields of object detection, environmental perception and scenario analysis.
Smart Images

Figure CN120030499A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for integrating multimodal data complementation, error correction and loss tolerance mechanisms. Background Art
[0002] A fusion method of multimodal data complementation, error correction and loss tolerance mechanism is an advanced technology for processing and combining data streams from multiple sensors (such as microwave, infrared, visible light, etc.). Multimodal data fusion involves synthesizing different information obtained by different sensors to obtain a more comprehensive and accurate scene representation. Each sensor data has its own unique advantages and limitations. For example, microwave data can provide stable target information in complex climate and low light environments, while infrared data can show the temperature distribution of objects and provide in-depth information about the state of objects without being affected by lighting conditions.
[0003] Traditional multimodal data fusion technology has problems of low efficiency, poor accuracy and poor robustness when processing large-scale data from different sensors. This is mainly due to the essential differences in the acquisition process of data of different modalities, such as differences in resolution, frame rate, signal-to-noise ratio, etc., which may lead to information loss, registration errors and noise superposition problems in the fusion process. In many practical applications, especially those involving real-time requirements and high-resolution scenes, such as autonomous driving, intelligent monitoring and military reconnaissance, traditional fusion methods cannot meet the requirements of fast, efficient and high quality.
[0004] In addition, problems such as missing information, inter-modal registration errors, and sensor noise can also affect the effectiveness of data fusion. When faced with factors such as environmental changes, uneven lighting, or sensor failures, the quality and integrity of the data may be affected, thereby reducing the stability and accuracy of the fusion results. Traditional technologies usually rely on rules or preset models to solve these problems, but these methods cannot fully utilize the data characteristics in complex environments, and it is difficult to respond flexibly in dynamically changing scenarios.
[0005] In order to solve these problems, multimodal data fusion methods based on deep learning have gradually become a hot topic of research. Deep learning models can automatically learn the correlation and complementarity between modalities, thereby effectively solving problems such as information loss, registration errors and noise. In addition, by introducing technologies such as adaptive weighting mechanisms, generative adversarial networks (GANs) and Gaussian process regression (GPR), deep learning-driven multimodal data fusion methods can dynamically adjust modal weights, generate missing modal data, correct registration errors, and improve the robustness and accuracy of the system in complex environments. Summary of the invention
[0006] In view of the above-mentioned technical deficiencies, the purpose of the present invention is to provide a fusion method of multimodal data complementarity, error correction and loss tolerance mechanism, which can effectively fuse data from different modes (such as microwave, infrared, visible light, etc.), solve the problems of information missing, registration error and noise interference, and improve the accuracy, robustness and stability of data fusion. Through the guidance of adaptive weighting mechanism and deep learning model, the present invention can operate stably in complex environments, adapt to different sensor characteristics and environmental changes, and is widely used in tasks with high precision requirements such as target detection and object recognition.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions: The present invention provides a method for integrating multimodal data complementation, error correction and loss tolerance mechanisms, comprising the following steps: Step 1: Collect microwave data stream and infrared data stream in the same scene; Step 2: Extract features from microwave data and infrared data streams, and then use the Transformer multimodal feature fusion framework to adaptively learn feature weights to perform complementary fusion of multimodal data, ensuring that when information from one modality is missing, other modalities can effectively supplement the missing part. Step 3. Design a multimodal generative network (MMGN) based on a generative adversarial network (GAN) to generate the missing modal data and reduce the negative impact of information loss on the fusion results. Step 4: Use deep learning-driven error correction mechanism and Gaussian process regression (GPR) technology to automatically identify and correct the registration errors caused by the sensor, and improve the registration accuracy by optimizing the error distribution.
[0008] Step 5: Introduce a robust loss function (such as the Huber loss function) to perform fault-tolerance processing for information missing and error correction, thereby enhancing the robustness and accuracy of the system.
[0009] Step 6: Dynamically adjust the weights of different modes in fusion through an adaptive weighting mechanism to further enhance the stability of the system in complex environments; Step 7. Generate the final fused data and use it for target detection and recognition tasks.
[0010] Furthermore, the steps of performing feature extraction operations on the microwave data stream and the infrared data stream in Step 2 are: Step 2.1-1, first assume that the original data of microwave data and infrared data stream are and , which correspond to microwave and infrared image sequences respectively. Since data of different modes may have different resolutions, scales and noises, they need to be normalized before processing: the value range of each modality data is unified so as to reduce the scale difference of the data in subsequent processing: ; in, and are the mean and standard deviation of the microwave data, and are the mean and standard deviation of the infrared data; Step 2.1-2: For each modality of data (microwave and infrared), deep learning methods such as convolutional neural networks (CNN) are used to extract high-dimensional features. Feature extraction is performed through multiple convolutional layers, each of which extracts more abstract features from the input data until the final feature representation is reached. Assume that the feature representation of microwave data is , the characteristic representation of infrared data is , and their extraction process through CNN can be expressed as: ; in, and are convolutional neural networks for microwave and infrared data, which generate feature representations for each modality respectively.
[0011] Step 2.1-3, in order to ensure that the features of each mode have a uniform distribution and avoid certain modal features dominating in subsequent processing, the extracted features are usually standardized. This process is done by subtracting the mean and dividing by the standard deviation: ; in, and are the mean and standard deviation of the microwave data features, and are the mean and standard deviation of the infrared data features; After completing the above feature extraction and preprocessing steps, we obtain the feature representation of the two modalities and , they will be fused in the subsequent Transformer model. In order to perform subsequent self-attention calculations, we take the features of these two modalities as input and pass them into the Transformer framework for cross-modal fusion.
[0012] Furthermore, the steps of using the Transformer multimodal feature fusion framework for the microwave data stream and the infrared data stream in Step 2 are as follows: Step 2.2-1. Microwave data characteristics and infrared data features The relationship between them is learned through the self-attention calculation in the Transformer architecture to obtain adaptive weights. These weighted features will eventually be fused to generate a joint feature representation. The goal of the self-attention mechanism is to assign a weight to each feature by calculating the correlation between each input feature. The weight reflects the importance of the feature to the current task. The specific calculation formula is as follows: ; in: is the query matrix, which represents the characteristics of the current modality; is the key matrix, which represents the characteristics of another mode; It is the value matrix, which represents the actual characteristic information of the mode; is the dimension of the key, which is used to scale the inner product. Through the above calculation, the system assigns a dynamically calculated weight to each feature, which changes in each self-attention module and can reflect the importance of different modalities in the final fusion.
[0013] Step 2.2-2, learn the correlation between different modalities through the self-attention mechanism of Transformer. We designed a joint query, key, and value matrix for cross-modal adaptive learning. The calculation formula is as follows: ; ; in, , , are the query, keys, and values of the first modality (e.g., microwave); , , are the query, keys, and values of the second modality (e.g., infrared); , , are the trained weight matrices, which are used for the calculation of queries, keys, and values respectively.
[0014] Calculate the cross-modal relevance score through the self-attention mechanism: ; Similarly, we can also calculate , to achieve information flow and fusion between modalities.
[0015] Furthermore, the step of adaptively learning the weights of the features in the microwave data stream and the infrared data stream in Step 2 is: Step 2.3-1, through the self-attention mechanism of Transformer, we can dynamically calculate the weight of each modality, and the model will automatically adjust the weight according to the contribution of each modality to the final fusion result. These weights can be expressed as: ; These weights and Represents the “importance” of each modality, that is, the contribution ratio of each modality to the final fusion result; Step 2.3-2, after calculating the weight of each modality, the features of the two modalities can be weighted fused to ensure that the dominant features of each modality make a greater contribution in the final fusion. The final fused features are expressed as: ; in, is the fused feature data, and It is the modal weight calculated by the self-attention mechanism.
[0016] Furthermore, the step of generating the missing modal data in the microwave data stream and the infrared data stream in Step 3 is: Step 3.1-1, in the Multimodal Generation Network (MMGN), we use a generative adversarial network to fill in the missing modal data. Suppose our system needs to process microwave data and infrared data, but for some reason, the infrared data may be lost. In this case, we can use MMGN to generate the missing infrared data; The generator of the multimodal generative network receives a noise vector and known modal data (e.g. microwave data) as input and generate a "fake" missing modal data (e.g. infrared data). Assume the generator is , whose input is microwave data and noise, and whose output is the generated infrared data, ; in: It is the data of known mode (such as microwave); is random noise, usually sampled from a Gaussian distribution; is the generated missing modality data (such as infrared data).
[0017] Discriminator Responsible for judging the generated infrared data Is it similar to the real infrared data? The goal of the discriminator is to distinguish whether the input data comes from the real infrared data or the infrared data generated by the generator. ; ; Step 3.1-2, the loss function of the generative adversarial network is usually based on the idea of adversarial training. The goal of the loss function is to minimize the difference between the "fake" data generated by the generator and the real data, while maximizing the ability of the discriminator to distinguish between real and fake data. The loss function of the generator is: ; The loss function of the discriminator is: ; in: is the true distribution of infrared data, is the generated infrared data; Step 3.1-3, through training, the generator continuously optimizes its parameters so that it can generate more and more realistic missing modal data, and the discriminator continuously optimizes its ability to ensure that it can accurately distinguish between real data and generated data. In this process, the generator and the discriminator compete with each other through adversarial training, and finally the generator can generate high-quality missing modal data, thereby supplementing the missing parts of the original data.
[0018] Furthermore, the steps of using a deep learning-driven error correction mechanism for the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.1-1. By designing a deep learning model, the system can automatically detect errors in the registration process. The process first aligns and compares the data from different modalities and calculates the registration difference between the two modalities. The detection of registration errors can be achieved by calculating the pixel difference or structural difference between different modalities. Suppose we have images of two modalities and , after they are preliminarily aligned, the difference metric between them is calculated: ; in, represents the difference measure between the two modalities, represents the L2 norm, which measures the Euclidean distance between pixels. A larger value of indicates a larger registration error; Step 4.1-2: Once errors are detected, the system will adjust the spatial position of the data based on these errors. A deep learning network is used to learn and optimize the correction methods for these errors so that the data of the two modalities are aligned.
[0019] In the error correction process, we use a deep learning network to learn how to adjust the offset in the image and correct the alignment of the data. The corrected data can be expressed as: ; in, is the corrected infrared data, is the correction calculated by the deep learning model.
[0020] Furthermore, the steps of applying Gaussian process regression (GPR) technology to the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.2-1, Gaussian process regression (GPR) is introduced to further optimize the registration error. GPR is a powerful non-parametric Bayesian method that can provide a smooth correction path for the error by calculating the potential distribution of the registration error.
[0021] The goal of Gaussian process regression is to estimate the probability distribution of the registration error by learning the distribution of historical data and correct the registration error based on this distribution. The formula of GPR is as follows: ; in: is the predicted registration error; is the mean function, usually taken as zero; is the covariance function (kernel function) that describes the correlation between input points. The posterior distribution of GPR allows us to optimize the registration error and obtain a smooth and accurate registration correction value.
[0022] Step 4.2-2, in the known training data Given a new point , we hope to predict (Correction of registration error), the prediction formula of GPR is: ; ; in: is the predicted mean (i.e., revised value); is the variance of the prediction, which indicates the uncertainty of the prediction; It's new and training data points The covariance between is the covariance matrix between the training data points. Through this process, GPR can generate more accurate correction values based on the existing registration error data and improve the accuracy of registration.
[0023] Furthermore, the step of introducing a robust loss function into the microwave data stream and the infrared data stream in Step 5 is: Step 5.1-1, when training deep learning models, using the standard mean square error (MSE) loss function may be affected by outliers in the data, because the mean square error is very sensitive to large errors and may cause model instability. To solve this problem, we use a robust loss function, such as the Huber loss function, to reduce the impact of outliers on the model training process.
[0024] The Huber loss function combines the advantages of mean square error (MSE) and absolute error (MAE), and can use square error for small errors and linear error for large errors, thereby reducing the impact of outliers while retaining the optimization effect when the error is small. The formula of the Huber loss function is as follows: ; in: is the true value; is the predicted value; is a hyperparameter that determines when to switch from squared error to linear error.
[0025] Furthermore, the steps of using the adaptive weighting mechanism for the microwave data stream and the infrared data stream in Step 6 are: Step 6.1-1. In order to dynamically adjust the weight of each modality, we first need to calculate the information quality of each modality. This can be achieved in a variety of ways. For example, assuming that for each modality , we can use a quality evaluation function To measure the quality of the modality. The quality assessment function may be based on multiple factors such as error detection, information loss, signal-to-noise ratio, etc. Step 6.1-2, based on the quality of each modality, we can calculate the weight of the modality in the fusion process. Generally speaking, modalities with higher information quality are assigned larger weights, while modalities with lower information quality are assigned smaller weights. A common adaptive weighting mechanism is to calculate the weight of each modality by normalizing the quality index.
[0026] set up For modal The weight of the quality assessment function satisfy: ; in: is modal Quality rating of is the total number of modes, in this way, the weight It is dynamically adjusted based on the quality of each mode, and the total is 1.
[0027] Step 6.1-3, after calculating the weight of each modality, weighted fusion can be performed. Assume that the modality The characteristic is expressed as , then the fused features It can be expressed as a weighted average: ; in: is the final feature representation after fusion; is modal The weight of is modal The feature representation of .
[0028] Through this weighted fusion, the system is able to automatically adjust the contribution of each modality to the final fusion result based on its quality.
[0029] The beneficial effects of the present invention are as follows: the present invention collects microwave data streams and infrared data streams in the same scene; pre-processes the collected microwave data streams and infrared data streams, including denoising and contrast enhancement operations, and uses a generative adversarial network to generate the missing modal data; uses an adaptive weighting mechanism to improve data fusion performance; and uses a fusion algorithm to weightedly fuse the processed microwave data and infrared data to generate a final fused data stream. It realizes the selection of appropriate pre-processing operations for different modal data streams, reduces noise and enhances image details, and ensures the integrity and consistency of the data; by incorporating an adaptive weighting mechanism into a deep learning framework, the weight of each modality in the fusion is dynamically adjusted, especially in the face of scenes with large differences in information quality of different sensors, the accuracy and robustness of data fusion are effectively improved; through a deep learning-driven error correction mechanism, the registration error is automatically corrected in combination with Gaussian process regression, further improving the accuracy of the fusion result; a fusion operation based on a convolutional network is used for the corrected different modal data streams, effectively integrating microwave and infrared information, greatly improving the processing efficiency and quality of multimodal data fusion, and expanding the scope of application of this technology in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0031] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 The present invention adopts a multimodal feature fusion flow chart based on the Transformer network framework; Figure 3 The present invention adopts a GAN-based information loss compensation flow chart; Figure 4 The present invention adopts Gaussian process regression and deep learning to correct the registration error flow chart. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] Examples, such as Figure 1-Figure 4 As shown, a method for integrating multimodal data complementation, error correction and loss tolerance mechanism includes the following steps: Step 1: Collect microwave data stream and infrared data stream in the same scene; Step 2: Extract features from microwave data and infrared data streams, and then use the Transformer multimodal feature fusion framework to adaptively learn feature weights to perform complementary fusion of multimodal data, ensuring that when information from one modality is missing, other modalities can effectively supplement the missing part. Step 3. Design a multimodal generative network (MMGN) based on a generative adversarial network (GAN) to generate the missing modal data and reduce the negative impact of information loss on the fusion results. Step 4: Use deep learning-driven error correction mechanism and Gaussian process regression (GPR) technology to automatically identify and correct the registration errors caused by the sensor, and improve the registration accuracy by optimizing the error distribution.
[0034] Step 5: Introduce a robust loss function (such as the Huber loss function) to perform fault-tolerance processing for information missing and error correction, thereby enhancing the robustness and accuracy of the system.
[0035] Step 6: Dynamically adjust the weights of different modes in fusion through an adaptive weighting mechanism to further enhance the stability of the system in complex environments; Step 7. Generate the final fused data and use it for target detection and recognition tasks.
[0036] As a preferred implementation, the steps of performing feature extraction operation on the microwave data stream and the infrared data stream in Step 2 are as follows: Step 2.1-1, firstly, assuming that the original data of the microwave data stream and the infrared data stream are and , which correspond to microwave and infrared image sequences respectively. Since data of different modes may have different resolutions, scales and noises, they need to be normalized before processing: the value range of each modality data is unified so as to reduce the scale difference of the data in subsequent processing. ; in, and are the mean and standard deviation of the microwave data, and are the mean and standard deviation of the infrared data; Step 2.1-2: For each modality of data (microwave and infrared), use deep learning methods such as convolutional neural networks (CNN) to extract high-dimensional features. Feature extraction is performed through multiple convolutional layers, each of which extracts more abstract features from the input data until the final feature representation is reached. Assume that the feature representation of microwave data is , the characteristic representation of infrared data is , and their extraction process through CNN can be expressed as: ; in, and are convolutional neural networks for microwave and infrared data, which generate feature representations for each modality respectively; Step 2.1-3, in order to ensure that the features of each mode have a uniform distribution and avoid certain modal features dominating in subsequent processing, the extracted features are usually standardized. This process is done by subtracting the mean and dividing by the standard deviation: ; in, and are the mean and standard deviation of the microwave data features, and is the mean and standard deviation of the infrared data features. After completing the above feature extraction and preprocessing steps, we obtain the feature representation of the two modalities and , they will be fused in the subsequent Transformer model. In order to perform subsequent self-attention calculations, we take the features of these two modalities as input and pass them into the Transformer framework for cross-modal fusion.
[0037] As a preferred implementation, the steps of using the Transformer multimodal feature fusion framework for the microwave data stream and the infrared data stream in Step 2 are as follows: Step 2.2-1. Microwave data characteristics and infrared data features The relationship between them is learned through the self-attention calculation in the Transformer architecture to obtain adaptive weights. These weighted features will eventually be fused to generate a joint feature representation. The goal of the self-attention mechanism is to assign a weight to each feature by calculating the correlation between each input feature. The weight reflects the importance of the feature to the current task. The specific calculation formula is as follows: ; in: is the query matrix, which represents the characteristics of the current modality; is the key matrix, which represents the characteristics of another mode; It is the value matrix, which represents the actual characteristic information of the mode; is the dimension of the key, which is used to scale the inner product. Through the above calculation, the system assigns a dynamically calculated weight to each feature, which changes in each self-attention module and can reflect the importance of different modalities in the final fusion.
[0038] Step 2.2-2, learn the correlation between different modalities through the self-attention mechanism of Transformer. We designed a joint query, key, and value matrix for cross-modal adaptive learning. The calculation formula is as follows: ; ; in, , , are the query, keys, and values of the first modality (e.g., microwave); , , are the query, keys, and values of the second modality (e.g., infrared); , , is the trained weight matrix, which is used for the calculation of query, key and value respectively. The cross-modal relevance score is calculated through the self-attention mechanism: ; Similarly, we can also calculate , to achieve information flow and fusion between modalities.
[0039] As a preferred implementation, the step of adaptively learning the weights of features in the microwave data stream and the infrared data stream in Step 2 is: Step 2.3-1, through the self-attention mechanism of Transformer, we can dynamically calculate the weight of each modality, and the model will automatically adjust the weight according to the contribution of each modality to the final fusion result. These weights can be expressed as: ; These weights and Represents the “importance” of each modality, that is, the contribution ratio of each modality to the final fusion result.
[0040] Step 2.3-2, after calculating the weight of each modality, the features of the two modalities can be weighted fused to ensure that the dominant features of each modality make a greater contribution in the final fusion. The final fused features are expressed as: ; in, is the fused feature data, and It is the modal weight calculated by the self-attention mechanism.
[0041] As a preferred implementation, the step of generating the missing modal data in the microwave data stream and the infrared data stream in Step 3 is: Step 3.1-1, in the multimodal generation network (MMGN), we use the generative adversarial network to fill the missing modal data. Suppose our system needs to process microwave data and infrared data, but for some reason, the infrared data may be lost. In this case, we can use MMGN to generate the missing infrared data; The generator of the multimodal generative network receives a noise vector and known modal data (such as microwave data) as input, and generate a "fake" missing modal data (such as infrared data). Assume that the generator is , whose input is microwave data and noise, and whose output is the generated infrared data, ; in: It is the data of known mode (such as microwave); is random noise, usually sampled from a Gaussian distribution; is the generated missing modality data (e.g. infrared data); Discriminator Responsible for judging the generated infrared data Is it similar to the real infrared data? The goal of the discriminator is to distinguish whether the input data comes from the real infrared data or the infrared data generated by the generator: ; ; Step 3.1-2, the loss function of the generative adversarial network is usually based on the idea of adversarial training. The goal of the loss function is to minimize the difference between the "fake" data generated by the generator and the real data, while maximizing the discriminator's ability to distinguish between real and fake data. The loss function of the generator is: ; The loss function of the discriminator is: ; in: is the true distribution of infrared data, is the generated infrared data; Step 3.1-3, through training, the generator continuously optimizes its parameters so that it can generate more and more realistic missing modal data, and the discriminator continuously optimizes its ability to ensure that it can accurately distinguish between real data and generated data. In this process, the generator and the discriminator compete with each other through adversarial training, and finally the generator can generate high-quality missing modal data, thereby supplementing the missing parts in the original data; As a preferred implementation, the steps of using a deep learning driven error correction mechanism for the microwave data stream and the infrared data stream in Step 4 are: Step 4.1-1. By designing a deep learning model, the system can automatically detect errors in the registration process. The process first compares the data from different modalities and calculates the registration difference between the two modalities. The detection of registration errors can be achieved by calculating the pixel difference or structural difference between different modalities. Suppose we have images of two modalities and , after they are preliminarily aligned, the difference metric between them is calculated: ; in, represents the difference measure between the two modalities, represents the L2 norm, which measures the Euclidean distance between pixels. A larger value of indicates a larger registration error; Step 4.1-2: Once errors are detected, the system will adjust the spatial position of the data based on these errors. Use a deep learning network to learn and optimize the correction method for these errors so that the data of the two modalities are aligned.
[0042] In the error correction process, we use a deep learning network to learn how to adjust the offset in the image and correct the alignment of the data. The corrected data can be expressed as: ; in, is the corrected infrared data, is the correction calculated by the deep learning model.
[0043] As a preferred implementation, the steps of applying Gaussian process regression (GPR) technology to the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.2-1, Gaussian process regression (GPR) is introduced to further optimize the registration error. GPR is a powerful non-parametric Bayesian method that can provide a smooth correction path for the error by calculating the potential distribution of the registration error. The goal of Gaussian process regression is to estimate the probability distribution of the registration error by learning the distribution of historical data, and correct the registration error according to the distribution. The formula of GPR is as follows: ; in: is the predicted registration error; is the mean function, usually taken as zero; is the covariance function (kernel function) that describes the correlation between input points. The posterior distribution of GPR allows us to optimize the registration error and obtain a smooth and accurate registration correction value.
[0044] Step 4.2-2, in the known training data Given a new point , we hope to predict (Correction of registration error). The prediction formula of GPR is: ; ; in: is the predicted mean (i.e., revised value); is the variance of the prediction, which indicates the uncertainty of the prediction; It's new and training data points The covariance between is the covariance matrix between the training data points. Through this process, GPR can generate more accurate correction values based on the existing registration error data and improve the accuracy of registration.
[0045] As a preferred implementation, the step of introducing a robust loss function into the microwave data stream and the infrared data stream in Step 5 is: Step 5.1-1, when training deep learning models, using the standard mean square error (MSE) loss function may be affected by outliers in the data, because the mean square error is very sensitive to large errors and may cause model instability. To solve this problem, we use robust loss functions, such as the Huber loss function, to reduce the impact of outliers on the model training process. The Huber loss function combines the advantages of the mean square error (MSE) and the absolute error (MAE). It can use square errors for small errors and linear errors for large errors, thereby reducing the impact of outliers while retaining the optimization effect when the error is small. The formula of the Huber loss function is as follows: ; in: is the true value; is the predicted value; is a hyperparameter that determines when to switch from squared error to linear error.
[0046] As a preferred implementation, the step of using the adaptive weighting mechanism for the microwave data stream and the infrared data stream in Step 6 is: Step 6.1-1, in order to dynamically adjust the weight of each modality, we first need to calculate the information quality of each modality. This can be achieved in a variety of ways, for example: assuming that for each modality , we can use a quality evaluation function To measure the quality of the modality. The quality assessment function may be based on multiple factors such as error detection, information loss, signal-to-noise ratio, etc. Step 6.1-2, based on the quality of each modality, we can calculate the weight of the modality in the fusion process. Generally speaking, modalities with higher information quality will be assigned larger weights, while modalities with lower information quality will receive smaller weights. A common adaptive weighting mechanism is to calculate the weight of each modality by normalizing the quality index; set up For modal The weight of the quality assessment function satisfy: ; in: is modal Quality rating of is the total number of modes. In this way, the weight It is dynamically adjusted based on the quality of each mode, and the total is 1.
[0047] Step 6.1-3, after calculating the weight of each modality, weighted fusion can be performed. Assume that the modality The characteristic is expressed as , then the fused features It can be expressed as a weighted average: ; in: is the final feature representation after fusion; is modal The weight of is modal The feature representation of .
[0048] Through this weighted fusion, the system is able to automatically adjust the contribution of each modality to the final fusion result based on its quality.
[0049] The present invention proposes an efficient complementation, error correction and loss tolerance mechanism through a multimodal data fusion method driven by deep learning, which solves the shortcomings of traditional technologies in information loss, registration error and noise processing. The feature fusion framework based on the Transformer architecture realizes information complementarity between modalities, and the generative adversarial network is used to make up for the missing modal data. The deep learning-driven error correction mechanism combines with Gaussian process regression to optimize the registration accuracy, while enhancing the system's fault tolerance and stability through robust loss functions and dynamic weighting mechanisms. This method dynamically adjusts the modal weights to ensure the dominant role of high-quality data in the fusion results, significantly improving the accuracy and robustness of data fusion. It is widely applicable to target detection, environmental perception, scene analysis and other fields, and provides a reliable solution for multimodal data processing in complex tasks. It has important academic value and application prospects.
[0050] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for integrating multimodal data complementation, error correction and loss tolerance mechanisms, characterized in that: The following steps are involved: Step 1: Collect microwave data stream and infrared data stream in the same scene; Step 2: Extract features from microwave data and infrared data streams, and then use the Transformer multimodal feature fusion framework to adaptively learn feature weights to perform complementary fusion of multimodal data. Step 3: Design a multimodal generative network based on a generative adversarial network to generate the missing modal data and reduce the negative impact of information loss on the fusion results; Step 4: Use deep learning-driven error correction mechanism and Gaussian process regression technology to automatically identify and correct the registration errors caused by the sensor, and improve the registration accuracy by optimizing the error distribution; Step 5: Introduce a robust loss function to perform fault tolerance processing for information loss and error correction, thereby enhancing the robustness and accuracy of the system; Step 6: Dynamically adjust the weights of different modes in fusion through an adaptive weighting mechanism to further enhance the stability of the system in complex environments; Step 7. Generate the final fused data and use it for target detection and recognition tasks.
2. A method for integrating multimodal data complementation, error correction and loss tolerance mechanisms as claimed in claim 1, characterized in that: The steps of performing feature extraction operation on the microwave data stream and the infrared data stream in Step 2 are as follows: Step 2.1-1. First, assume that the original data of the microwave data stream and the infrared data stream are and , corresponding to microwave and infrared image sequences respectively. Since data of different modalities may have different resolutions, scales and noises, the data is normalized before processing: ; in, and are the mean and standard deviation of the microwave data, and are the mean and standard deviation of the infrared data; Step 2.1-2: For each modality of data, use the convolutional neural network deep learning method to extract high-dimensional features; Assume that the feature representation of microwave data is , the characteristic representation of infrared data is , the process of extraction by CNN is expressed as: ; in, and It is a convolutional neural network for microwave and infrared data, generating feature representations for each modality respectively; Step 2.1-3, standardize the extracted features by subtracting the mean and dividing by the standard deviation: ; in, and are the mean and standard deviation of the microwave data features, and are the mean and standard deviation of the infrared data features; After completing the above feature extraction and preprocessing steps, the feature representations of the two modalities are obtained. and , they will be fused in the subsequent Transformer model, and the features of these two modalities will be taken as input and passed into the Transformer framework for cross-modal fusion.
3. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of using the Transformer multimodal feature fusion framework for the microwave data stream and the infrared data stream in Step 2 are as follows: Step 2.2-1. Microwave data characteristics and infrared data features The relationship between them is learned through the self-attention calculation in the Transformer architecture to obtain adaptive weights. These weighted features are finally fused to generate a joint feature representation. The goal of the self-attention mechanism is to assign a weight to each feature by calculating the correlation between each input feature. The specific calculation formula is as follows: ; in: is the query matrix, which represents the characteristics of the current modality; is the key matrix, representing the characteristics of the other mode; is a value matrix, representing the actual characteristic information of the mode; is the dimension of the key, used to scale the inner product; Step 2.2-2, learn the correlation between different modalities through the self-attention mechanism of Transformer, design the joint query, key and value matrix, and perform cross-modal adaptive learning. The calculation formula is as follows: ; ; in, , , are the query, key, and value of the first modal; , , are the query, key, and value of the second modal; , , is the trained weight matrix, which is used for the calculation of query, key and value respectively; the cross-modal relevance score is calculated through the self-attention mechanism: ; Similarly, by calculating , to achieve information flow and fusion between modalities.
4. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of adaptively learning the weights of the features in the microwave data stream and the infrared data stream in Step 2 are: Step 2.3-1, through the self-attention mechanism of Transformer, the weight of each modality is dynamically calculated. The model will automatically adjust the weight according to the contribution of each modality to the final fusion result. These weights are expressed as: ; These weights and Represents the importance of each modality, that is, the contribution ratio of each modality to the final fusion result; Step 2.3-2, after calculating the weight of each modality, the features of the two modalities are weighted fused, and the final fused features are expressed as: ; in, is the fused feature data, and It is the modal weight calculated by the self-attention mechanism.
5. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of generating the missing modal data in the microwave data stream and the infrared data stream in Step 3 are as follows: Step 3.1-1. In the multimodal generative network, a generative adversarial network is used to fill in the missing modal data. The generator of the multimodal generative network receives a noise vector and known modal data as input, and generate a "fake" missing modal data, assuming the generator is , whose input is microwave data and noise, and whose output is the generated infrared data: ; in: is the known modal data; is random noise, usually sampled from a Gaussian distribution; is the generated missing modal data; Discriminator Responsible for judging the generated infrared data Is it similar to the real infrared data? The goal of the discriminator is to distinguish whether the input data comes from the real infrared data or the infrared data generated by the generator. The formula is as follows: ; ; Step 3.1-2, the loss function of the generative adversarial network is based on the idea of adversarial training. The goal of the loss function is to minimize the difference between the "fake" data generated by the generator and the real data, while maximizing the ability of the discriminator to distinguish between real and fake data. The loss function of the generator is: ; The loss function of the discriminator is: ; in: is the true distribution of infrared data, is the generated infrared data; Step 3.1-3. Through training, the generator continuously optimizes its parameters, and the discriminator continuously optimizes its capabilities. The generator and the discriminator compete with each other through adversarial training. Finally, the generator can generate high-quality missing modal data, thereby supplementing the missing parts in the original data.
6. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of using the deep learning driven error correction mechanism for the microwave data stream and the infrared data stream in Step 4 are as follows: Step 4.1-1. By designing a deep learning model, the system automatically detects errors in the registration process. The process first compares the data from different modalities and calculates the registration difference between the two modalities. The detection of registration errors is achieved by calculating the pixel difference or structural difference between different modalities. Suppose we have images of two modalities. and , after they are preliminarily aligned, the difference metric between them is calculated: ; in, represents the difference measure between the two modalities, represents the L2 norm, which measures the Euclidean distance between pixels. A larger value of indicates a larger registration error; Step 4.1-2: Once errors are detected, the system will adjust the spatial position of the data based on these errors, and use the deep learning network to learn and optimize the correction methods of these errors so that the data of the two modalities are aligned; in the error correction process, the deep learning network is used to learn how to adjust the offset in the image and correct the alignment of the data. The corrected data is expressed as: ; in, is the corrected infrared data, is the correction calculated by the deep learning model.
7. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of using Gaussian process regression technology on microwave data stream and infrared data stream in Step 4 are as follows: Step 4.2-1, Gaussian process regression is introduced to further optimize the registration error. Gaussian process regression estimates the probability distribution of the registration error by learning the distribution of historical data, and corrects the registration error according to the distribution. The formula of GPR is as follows: ; in: is the predicted registration error; is the mean function, which is taken as zero; is the covariance function, describing the correlation between input points; the posterior distribution of GPR allows the optimization of the registration error, thereby obtaining a smooth and accurate registration correction value; Step 4.2-2, in the known training data Given a new point , to predict , which is the correction amount of the registration error, the prediction formula of GPR is: ; ; in: is the predicted mean, i.e., the revised value; is the variance of the prediction, which indicates the uncertainty of the prediction; It's new and training data points The covariance between is the covariance matrix between the training data points.
8. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of introducing a robust loss function to the microwave data stream and the infrared data stream in Step 5 are: Step 5.1-1: When training deep learning models, use a robust loss function to reduce the impact of outliers on the model training process: The loss function combines the advantages of mean square error and absolute error. It can use square error for small errors and linear error for large errors, thereby reducing the impact of outliers while retaining the optimization effect when the error is small. The formula of the loss function is as follows: ; in: is the true value; is the predicted value; is a hyperparameter that determines when to switch from squared error to linear error.
9. The method for integrating multimodal data complementation, error correction and loss tolerance mechanism as claimed in claim 1, characterized in that: The steps of using the adaptive weighting mechanism for the microwave data stream and the infrared data stream in Step 6 are as follows: Step 6.1-1: In order to dynamically adjust the weight of each modality, we first need to calculate the information quality of each modality. , through a quality assessment function To measure the quality of the modality, the quality assessment function is based on multiple factors such as error detection, information loss, and signal-to-noise ratio; Step 6.1-2: Calculate the weight of each modality in the fusion process according to its quality. set up For modal The weight of the quality assessment function satisfy: ; in: is modal Quality rating of is the total number of modes; in this way, the weight It will be dynamically adjusted according to the quality of each mode, and the sum is 1; Step 6.1-3, after calculating the weight of each mode, perform weighted fusion, assuming that the mode The characteristic is expressed as , then the fused features Expressed as a weighted average: ; in: is the final feature representation after fusion; is modal The weight of is modal The feature representation of .
Citation Information
Patent Citations
Multi-modal human face recognition method based on deep learning
CN106909905A
Multi-modal fusion method and device based on normalized mutual information, medium and equipment
CN111461176A
Brain dysfunction auxiliary evaluation method based on multi-modal data fusion
CN115553752A