A deep metric-based airborne multimodal collaborative target detection method
By using the method of depth measurement combined with a weighted attention fusion mechanism in multimodal object detection, the problem of insufficient data alignment accuracy between different platforms is solved, and higher object detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510153045.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-12
AI Technical Summary
In the prior art, data alignment accuracy between different platforms is insufficient, and the difference between modes is difficult to eliminate, resulting in the limitation of the accuracy and robustness of multimodal object detection.
The airborne multimodal collaborative object detection method based on depth metrics is adopted, combined with the weighted attention fusion mechanism, the data of different platforms are mapped to the shared feature space, and the alignment process is optimized through deep metric learning, and the weight of each modal data is dynamically adjusted to improve the alignment accuracy.
It effectively reduces the differences between modes, improves the accuracy and robustness of data alignment, and significantly improves the accuracy of object detection.
Smart Images

Figure CN119622388B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing technology, and in particular to an airborne multi-modal collaborative target detection method based on depth measurement. Background Art
[0002] Multimodal data refers to observation information in multiple forms from different sensors or data sources, such as optical images, radar data, infrared data, etc. These data usually have different physical properties, information expression methods and perception capabilities, and can achieve comprehensive perception and accurate identification of targets in various complex scenarios.
[0003] Multimodal target detection refers to the process of detecting targets in the same scene by combining information from different types of sensors or data sources. Its main advantage is that it can make up for the limitations of single-modal data. For example, optical images have extremely high resolution and clarity during the day or in good weather conditions, but are almost ineffective in bad weather or at night; while radar data has the ability to penetrate clouds and rain, and can provide stable information output in complex environments. In addition, the coordinated use of multimodal data can significantly improve the robustness of target detection and reduce detection errors caused by noise or failure of a single sensor. For example, in marine target detection, the joint use of optical and radar data can simultaneously capture the geometric features and dynamic information of the target, thereby achieving a more comprehensive perception.
[0004] Although multimodal technology has significant advantages, it faces many challenges in practical applications. First, there are huge differences between different modal data, including acquisition time, spatial resolution, information expression, and perspective differences. These differences will lead to poor quality of directly fused data. Secondly, the huge differences between different modal data make data alignment and feature extraction more difficult, and put forward higher requirements on the performance of the system. Current researchers have proposed many solutions to the above problems, such as feature fusion technology based on deep learning, which jointly analyzes multimodal data through shared feature space; spatial alignment technology based on geometric models, which is used to reduce the geometric differences between different modalities. These methods have alleviated the difficulties of multimodal data fusion to a certain extent, but there are still obvious problems, such as: insufficient data alignment accuracy, it is difficult to completely eliminate the differences between modalities; the traditional process of fusion first and then detection is easily affected by noise or low-quality modal data, etc.
[0005] To address the above problems, the present invention aims to achieve effective alignment of data between different platforms and fusion of target detection results through an innovative depth metric combined with a weighted attention fusion mechanism, thereby improving the accuracy of target detection. Summary of the invention
[0006] The purpose of the present invention is to provide an airborne multimodal collaborative target detection method based on depth measurement, so as to solve the problems existing in the prior art that the data alignment accuracy between different platforms is insufficient and the differences between modes are difficult to eliminate.
[0007] To achieve the above object, the present invention provides an airborne multimodal collaborative target detection method based on depth measurement, comprising the following steps:
[0008] Step 1: Use the optical camera, radar sensor and radiation source sensor carried on the aircraft to synchronously collect data on the target area, obtain optical images, radar data and radiation source data, and then pre-process them;
[0009] Step 2: Perform target detection on the preprocessed optical image, radar data and radiation source data on their respective platforms; use a deep learning-based target detection algorithm to detect targets on optical images, use a constant false alarm rate (CFAR) detection algorithm combined with the target's motion characteristics to detect targets on radar data, and use signal matching and cluster analysis methods to detect targets on radiation source data;
[0010] Step 3: Based on deep metric learning combined with a weighted attention fusion mechanism, the target detection results of the optical image, radar data, and radiation source data obtained in step 2 are aligned and fused to finally obtain a comprehensive target detection result.
[0011] Preferably, the optical image, radar data and radiation source data are preprocessed as follows: first, the optical image, radar data and radiation source data are standardized and denoised; then the data acquired by each platform are time synchronized; finally, the optical image is resolution enhanced, the radar data is Doppler filtered, and the radiation source data is normalized for radiation intensity to enhance its contrast.
[0012] Preferably, the target detection algorithm based on deep learning is used to perform target detection on the optical image, specifically: construct a neural network architecture including a convolution layer, a pooling layer and a fully connected layer to perform target detection on the optical image; wherein, the feature map of the image is extracted by the convolution layer, down-sampling is performed by the pooling layer to reduce the amount of calculation, and the position and category information of the target is classified and regressed by the fully connected layer.
[0013] Preferably, the target detection on radar data is performed using a constant false alarm rate detection algorithm in combination with the target's motion characteristics as follows: first, the radar data is divided into a number of distance units and azimuth units, and the statistical characteristics of background clutter, such as the mean and variance, are calculated in each unit; then, the target's track is initiated and associated in combination with the target's velocity and acceleration motion characteristics, and the target's motion state is estimated and predicted by a Kalman filter method, thereby achieving continuous tracking and detection of radar targets, and outputting the target's position, velocity and motion trajectory information.
[0014] Preferably, the signal matching and clustering analysis methods are used to perform target detection on radiation source data as follows: first, a radiation source signal feature library is constructed, and the collected radiation source signal features are matched with the known radiation source features in the library; for several similar radiation source signals, a clustering algorithm (such as K-Means clustering) is used to divide them into different target categories, and the position, frequency and power information of the radiation source target is output.
[0015] Preferably, the specific process of step 3 is as follows:
[0016] S31. Alignment optimization based on deep metric learning. The process is as follows:
[0017] S311, set the target detection result of each platform to ,in is the jth detected target on platform i, , respectively represent optical images, radar data and radiation source data; , each target Contains location coordinates , target category and its confidence ;
[0018] S312, using a shared neural network embedding module, similar targets from different platforms are brought closer in the embedding space and dissimilar targets are moved away through training optimization; the embedding feature vector of the target is obtained, and the expression is as follows:
[0019] ;
[0020] in, represents the feature extraction network for platform i; through this network, each target is mapped to a vector of fixed dimension In it, d is the dimension of the embedding space;
[0021] S313, cosine similarity is used to measure the similarity of targets from different platforms in the embedding space, and the calculation expression is as follows:
[0022] ;
[0023] in, Represents two target feature vectors and The similarity of is the dot product of two vectors, and are the norms of the vectors respectively; by calculating the similarity between multiple target pairs, targets with high similarity are selected for alignment, and targets with low similarity are excluded;
[0024] S314. Construct loss function For training optimization, the expression is as follows:
[0025] ;
[0026] in, is the indicator function, when the target and If they belong to the same category, the value is 1, otherwise it is 0; is a set threshold that controls the maximum tolerance value of similarity;
[0027] S32. After alignment optimization based on deep metric learning, the targets of different platforms are fused through the weighted attention fusion mechanism. The process is as follows:
[0028] S321. For each platform i, the target , calculate its corresponding weight , the expression is as follows:
[0029] ;
[0030] in, is the query vector, which represents the target feature vector that needs to be aligned. is the dot product between the target and query vectors, indicating the relevance between the target and the query; this formula calculates the relevance of each target through the dot product and normalizes it through the softmax function to ensure that the sum of the weights is 1;
[0031] S322, the weights calculated based on each platform are fused by weighting, and the expression is as follows:
[0032] ;
[0033] in, represents the fusion function, which is to learn the target fusion method of different platforms through a neural network model; Represents the final fusion result, including the spatial position of the target, target category and confidence.
[0034] Therefore, the present invention adopts the above-mentioned airborne multimodal collaborative target detection method based on depth measurement, which has the following beneficial effects:
[0035] (1) By combining deep metric learning and attention fusion mechanism, the problem of data alignment between different platforms is effectively solved. Data from different platforms can be automatically mapped to a shared feature space, thereby significantly reducing the differences between modalities.
[0036] (2) The weights of data of different modalities are dynamically adjusted through the attention fusion mechanism, so that high-quality data occupies a larger proportion in the alignment process. At the same time, the interference of low-quality data can be effectively suppressed, thereby greatly improving the accuracy and robustness of data alignment.
[0037] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is an overall flow chart of an airborne multimodal collaborative target detection method based on depth measurement of the present invention. DETAILED DESCRIPTION
[0039] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] See also Figure 1 , an airborne multimodal collaborative target detection method based on deep metrics, comprising the following steps:
[0041] Step 1: Use the optical camera, radar sensor and radiation source sensor carried on the aircraft to synchronously collect data on the target area, obtain optical images, radar data and radiation source data, and then preprocess them; the preprocessing of optical images, radar data and radiation source data is specifically as follows: first, standardize and denoise the optical images, radar data and radiation source data; then perform time synchronization on the data obtained by each platform; finally, enhance the resolution of the optical image, perform Doppler filtering on the radar data, and normalize the radiation intensity of the radiation source data to enhance its contrast.
[0042] Step 2, perform target detection on the preprocessed optical image, radar data and radiation source data on their respective platforms; use a deep learning-based target detection algorithm to detect targets on optical images, use a constant false alarm rate (CFAR) detection algorithm combined with the target's motion characteristics to detect targets on radar data, and use signal matching and clustering analysis methods to detect targets on radiation source data; wherein, using a deep learning-based target detection algorithm to detect targets on optical images is as follows: construct a neural network architecture including a convolutional layer, a pooling layer and a fully connected layer to detect targets on optical images; wherein, the feature map of the image is extracted by the convolutional layer, down-sampling is performed by the pooling layer to reduce the amount of calculation, and the position and category information of the target is classified and regressed by the fully connected layer. The specific method of using constant false alarm rate detection algorithm combined with the target's motion characteristics to detect targets on radar data is as follows: first, the radar data is divided into several distance units and azimuth units, and the statistical characteristics of background clutter, such as mean and variance, are calculated in each unit; then, the target's track is started and associated in combination with the target's velocity and acceleration motion characteristics, and the target's motion state is estimated and predicted by the Kalman filter method to achieve continuous tracking and detection of radar targets, and output the target's position, velocity and motion trajectory information. The specific method of using signal matching and clustering analysis methods to detect targets on radiation source data is as follows: first, a radiation source signal feature library is constructed, and the collected radiation source signal features are matched with the known radiation source features in the library; for several similar radiation source signals, a clustering algorithm (such as K-Means clustering) is used to divide them into different target categories, and the position, frequency and power information of the radiation source target is output.
[0043] Step 3: Based on deep metric learning combined with a weighted attention fusion mechanism, the target detection results of the optical image, radar data, and radiation source data obtained in step 2 are aligned and fused to finally obtain a comprehensive target detection result. The specific process is as follows:
[0044] S31. Alignment optimization based on deep metric learning. The process is as follows:
[0045] S311, set the target detection result of each platform to ,in is the jth detected target on platform i, , respectively represent optical images, radar data and radiation source data; , each target Contains location coordinates , target category and its confidence ;
[0046] S312, using a shared neural network embedding module, similar targets from different platforms are brought closer in the embedding space and dissimilar targets are moved away through training optimization; the embedding feature vector of the target is obtained, and the expression is as follows:
[0047] ;
[0048] in, represents the feature extraction network for platform i; through this network, each target is mapped to a vector of fixed dimension In it, d is the dimension of the embedding space;
[0049] S313, cosine similarity is used to measure the similarity of targets from different platforms in the embedding space, and the calculation expression is as follows:
[0050] ;
[0051] in, Represents two target feature vectors and The similarity of is the dot product of two vectors, and are the norms of the vectors respectively; by calculating the similarity between multiple target pairs, targets with high similarity are selected for alignment, and targets with low similarity are excluded;
[0052] S314. Construct loss function For training optimization, the role of this loss function is to maximize the similarity of targets of the same category and minimize the similarity between targets of different categories, thereby effectively achieving spatial alignment of targets; the expression is as follows:
[0053] ;
[0054] in, is the indicator function, when the target and If they belong to the same category, the value is 1, otherwise it is 0; is a set threshold that controls the maximum tolerance value of similarity;
[0055] S32. After alignment optimization based on deep metric learning, the targets of different platforms are fused through the weighted attention fusion mechanism. The process is as follows:
[0056] S321. For each platform i, the target , calculate its corresponding weight , the expression is as follows:
[0057] ;
[0058] in, is the query vector, which represents the target feature vector that needs to be aligned. is the dot product between the target and query vectors, indicating the relevance between the target and the query; this formula calculates the relevance of each target through the dot product and normalizes it through the softmax function to ensure that the sum of the weights is 1;
[0059] S322, the weights calculated based on each platform are fused by weighting, and the expression is as follows:
[0060] ;
[0061] in, represents the fusion function, which is to learn the target fusion method of different platforms through a neural network model; Represents the final fusion result, including the spatial position of the target, target category and confidence.
[0062] Therefore, the present invention adopts the above-mentioned airborne multimodal collaborative target detection method based on depth measurement. For the multimodal data obtained from different platforms, target detection is first performed independently based on each data to obtain a preliminary target detection result for each modality; then, data alignment is performed on the preliminary target detection result of each modality to eliminate the influence of differences in spatial position, viewing angle, sensor characteristics, etc.; finally, the aligned target detection results of the three different modalities are fused to finally obtain a comprehensive target detection result; in this way, the accuracy of target detection is improved, and at the same time, the robustness and adaptability of the system are improved.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for airborne multimodal collaborative target detection based on depth metrics, characterized in that: The following steps are involved: Step 1: Use the optical camera, radar sensor and radiation source sensor carried on the aircraft to synchronously collect data on the target area, obtain optical images, radar data and radiation source data, and then pre-process them; Step 2: Perform target detection on the preprocessed optical image, radar data and radiation source data on their respective platforms; use a deep learning-based target detection algorithm to detect targets on optical images, use a constant false alarm rate detection algorithm combined with the target's motion characteristics to detect targets on radar data, and use signal matching and cluster analysis methods to detect targets on radiation source data; Step 3: Based on the method of deep metric learning combined with weighted attention fusion mechanism, the target detection results of the optical image, radar data and radiation source data obtained in step 2 are aligned and fused to finally obtain a comprehensive target detection result; The specific process of step 3 is as follows: S31. Alignment optimization based on deep metric learning. The process is as follows: S311, set the target detection result of each platform to ,in is the jth detected target on platform i, , respectively represent optical images, radar data and radiation source data; , each target Contains location coordinates , target category and its confidence ; S312, using a shared neural network embedding module to bring similar targets from different platforms closer together in the embedding space and to move dissimilar targets further apart; and obtaining an embedding feature vector of the target; S313, cosine similarity is used to measure the similarity of targets from different platforms in the embedding space; S314. Construct loss function Conduct training optimization; S32. After alignment optimization based on deep metric learning, the targets of different platforms are fused through the weighted attention fusion mechanism. The process is as follows: S321. For each platform i, the target , calculate its corresponding weight ; S322, the weights calculated based on each platform are fused by weighting, and the expression is as follows: ; in, represents the fusion function, which is to learn the target fusion method of different platforms through a neural network model; Represents the final fusion result, including the spatial position of the target, target category and confidence.
2. The method for airborne multimodal collaborative target detection based on depth metric according to claim 1, characterized in that: The preprocessing of optical images, radar data and radiation source data is as follows: first, the optical images, radar data and radiation source data are standardized and denoised; then the data obtained from each platform are time synchronized; finally, the optical image is resolution enhanced, the radar data is Doppler filtered, and the radiation source data is normalized for radiation intensity.
3. The method for airborne multimodal collaborative target detection based on depth metric according to claim 2, characterized in that: The target detection algorithm based on deep learning is used to detect the target in the optical image. Specifically, a neural network architecture including convolution layer, pooling layer and fully connected layer is constructed to detect the target in the optical image; the feature map of the image is extracted by the convolution layer, down-sampling is performed by the pooling layer, and the position and category information of the target are classified and regressed by the fully connected layer.
4. The method for airborne multimodal collaborative target detection based on depth metric according to claim 3, characterized in that: The constant false alarm rate detection algorithm is used to detect targets on radar data in combination with the target's motion characteristics. Specifically, the radar data is divided into several range units and azimuth units, and the statistical characteristics of background clutter are calculated in each unit. Then, the target's track is started and associated in combination with the motion characteristics of the target's velocity and acceleration. The target's motion state is estimated and predicted through the Kalman filter method, achieving continuous tracking and detection of radar targets, and outputting the target's position, velocity and motion trajectory information.
5. The method for airborne multimodal collaborative target detection based on depth metric according to claim 4, characterized in that: The signal matching and clustering analysis methods are used to detect targets on radiation source data. Specifically, first, a radiation source signal feature library is constructed, and the collected radiation source signal features are matched with the known radiation source features in the library. For several similar radiation source signals, a clustering algorithm is used to divide them into different target categories, and the location, frequency and power information of the radiation source target are output.
6. The method for airborne multimodal collaborative target detection based on depth metric according to claim 5, characterized in that: The specific process of step 3 is as follows: S31. Alignment optimization based on deep metric learning. The process is as follows: S311, set the target detection result of each platform to ,in is the jth detected target on platform i, , respectively represent optical images, radar data and radiation source data; , each target Contains location coordinates , target category and its confidence ; S312, using a shared neural network embedding module, similar targets from different platforms are brought close together in the embedding space, and dissimilar targets are moved away; the embedding feature vector of the target is obtained, and the expression is as follows: ; in, represents the feature extraction network for platform i; through this network, each target is mapped to a fixed-dimensional target feature vector In it, d is the dimension of the embedding space; S313, cosine similarity is used to measure the similarity of targets from different platforms in the embedding space, and the calculation expression is as follows: ; in, Represents two target feature vectors and The similarity of is the dot product of two vectors, and are the norms of the vectors respectively; by calculating the similarity between multiple target pairs, targets with high similarity are selected for alignment, and targets with low similarity are excluded; S314. Construct loss function For training optimization, the expression is as follows: ; in, is the indicator function, when the target and If they belong to the same category, the value is 1, otherwise it is 0; is a set threshold that controls the maximum tolerance value of similarity; S32. After alignment optimization based on deep metric learning, the targets of different platforms are fused through the weighted attention fusion mechanism. The process is as follows: S321. For each platform i, the target , calculate its corresponding weight , the expression is as follows: ; in, is the query vector, which represents the target feature vector that needs to be aligned. is the dot product between the target and query vectors, indicating the degree of association between the target and the query; S322, the weights calculated based on each platform are fused by weighting, and the expression is as follows: ; in, represents the fusion function, which is to learn the target fusion method of different platforms through a neural network model; Represents the final fusion result, including the spatial position of the target, target category and confidence.
Citation Information
Patent Citations
Pedestrian re-identification method of twin generative adversarial network based on attitude guidance pedestrian image generation
CN110427813A
Deep learning driven small target tracking method and system
CN118587253A