A multi-source target recognition method and system based on deep feature fusion and a medium

CN122815408APending Publication Date: 2026-09-25NANJING RES INST OF ELECTRONICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611007208.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]目前,对低空飞行器的识别依赖单一模态数据,存在受环境、气候等边界条件影响大、识别不稳定的问题

Benefits of technology

[0043]现有技术中,对低空飞行器的识别依赖单一模态数据,存在受环境、气候等边界条件影响大、时常无法识别目标的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122815408A_ABST
    Figure CN122815408A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of radar signal target recognition, and discloses a multi-source target recognition method and system based on deep feature fusion and a medium. In the application, a neural network is used to extract a track mode deep feature and a PD graph mode deep feature of a radar respectively, and the two are fused to be used for a final target recognition task. Compared with a single mode recognition method, multi-mode recognition can complement information between different modes, effectively overcome environmental interference, and thus significantly improve the stability and accuracy of recognition, and realize more stable and reliable recognition in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar signal target recognition technology, specifically to a multi-source target recognition method, system, and medium based on deep feature fusion. Background Technology

[0002] In recent years, with the rapid development of drone technologies such as flight control and navigation, drones have been widely used in civilian industries such as logistics, security, and environmental monitoring. Whether it is to ensure flight safety or to strengthen low-altitude surveillance, high-precision identification of drones and other low-altitude aircraft is required.

[0003] Currently, the identification of low-altitude aircraft relies on single-modal data, which suffers from significant influence from environmental and climatic boundary conditions, resulting in unstable identification. Therefore, a precise and efficient target identification method urgently needs to be researched. Summary of the Invention

[0004] To address the problems in existing technologies and improve the accuracy of target recognition, this application provides a multi-source target recognition method, system, and medium based on deep feature fusion.

[0005] Firstly, a multi-source target recognition method based on deep feature fusion is provided, adopting the following technical solution:

[0006] Obtain the original features and labels for the pairing, wherein the original features include the target PD map and target track information;

[0007] The original features are preprocessed to construct a neural network model dataset; the preprocessed target PD map is denoted as the PD map modal feature, and the preprocessed target trajectory information is denoted as the trajectory modal feature;

[0008] A single-modal feature extractor is constructed and pre-trained, comprising a PD image modal feature extractor and a track modal feature extractor; the PD image modal feature extractor extracts the PD image modal features into PD image modal depth features, and the track modal feature extractor extracts the track modal features into track modal depth features;

[0009] A deep feature fusion module is constructed and pre-trained, wherein the deep features include the PD map modal depth features and the track modal depth features;

[0010] Inference is performed based on the pre-trained deep feature fusion module.

[0011] Furthermore, when preprocessing the original features to construct the neural network model dataset, the specific steps are as follows:

[0012] The preprocessing of the target trajectory information is as follows: The target trajectory information includes information related to length and information related to angle; the information related to length is denoted as the first feature, and the information related to angle is denoted as the second feature; the first feature is normalized and mapped to a predetermined target interval; the second feature is converted to an angle value, that is, the angle unit is converted from degrees to radians; using the maximum and minimum values ​​of the first feature in the training set, all length-related features in the dataset are normalized; using the maximum and minimum values ​​of the second feature in the training set, all angle-related features in the dataset are normalized; the processed target trajectory information is the trajectory modal feature.

[0013] The preprocessing of the PD image is as follows: the PD image is converted to decibels, the obtained decibel values ​​are normalized, and mapped to a predetermined target interval; the processed PD image is the PD image modal feature.

[0014] When constructing the dataset, the dataset is divided into a training set, a cross-validation set, and a test set, with each track as a unit; all track points on a track can only be assigned to one of the training set, the cross-validation set, or the test set.

[0015] Furthermore, the construction of the single-modal feature extractor is as follows:

[0016] Constructing a PD image modality feature extractor: A convolutional neural network is used, with the network structure containing three convolutional layers. The specific parameters are as follows:

[0017] The first convolutional layer is configured to receive input data with 1 channel and perform convolution operation using a convolutional kernel of size 11×11 with padding size of 2, and output a feature map with 16 channels.

[0018] The first max pooling layer is configured to apply a 3×3 max pooling operation to the output of the first convolutional layer, with a stride of 2.

[0019] The second convolutional layer is configured to receive the output of the first max pooling layer and perform a convolution operation using a 3×3 convolutional kernel with a padding size of 1, outputting a feature map with 32 channels.

[0020] The second max pooling layer is configured to apply a 3×3 max pooling operation to the output of the second convolutional layer, with a stride of 2.

[0021] The third convolutional layer is configured to receive the output of the second max pooling layer and perform a convolution operation using a 3×3 convolutional kernel with a padding size of 1, outputting a feature map with 64 channels.

[0022] The third max pooling layer is configured to apply a 3×3 max pooling operation to the output of the third convolutional layer, with a stride of 2.

[0023] An adaptive max-pooling layer is configured to perform an adaptive max-pooling operation on the output of the third max-pooling layer, and the output space is a feature map of size 1×1.

[0024] A fully connected layer is configured to receive the output of the adaptive max-pooling layer and output a 64-dimensional feature vector.

[0025] Pre-trained PD image modality feature extractor: The classifier of the PD image modality feature extractor is a fully connected linear classifier. The loss function is selected as the cross-entropy loss function, and the training period is set to 200. The expression of the cross-entropy loss function (CE) is as follows:

[0026]

[0027] in This is the sample index, representing the sample number; The number of samples; This is a category index, representing the category number; The total number of categories; One-hot encoding for the tag; This is the model's predicted value, i.e., the model's predicted value for the [number]th [year]. The sample belongs to the first The confidence level of each category is set; the early termination condition is set to CE not decreasing for 20 consecutive training cycles; the learning rate is set to 0.0005 during pre-training; the PD graph modal features and labels are input into the PD graph modal feature extractor to obtain the PD graph modal depth features;

[0028] Constructing a track modality feature extractor: A gated recurrent unit neural network is used to build the feature vector of the entire sequence based on the input of track point features from M consecutive frames, where M is a positive integer greater than 0. The hidden layer size is consistent with the dimension of the feature vector output by the fully connected layer in the PD graph modality feature extractor.

[0029] Pre-trained track modality feature extractor: The classifier of the track modality feature extractor is a fully connected linear classifier. The training period, loss function, and early termination condition are the same as those of the pre-training of the PD graph modality feature extractor. The learning rate is set to 0.00005 during the pre-training process. The track modality features and labels are input into the track modality feature extractor to obtain the track modality depth features.

[0030] Furthermore, when constructing and pre-training the deep feature fusion module, the specific steps are as follows:

[0031] Constructing a deep feature fusion module: A deep feature fusion module is built using a graph attention network for modal features. Modal features also exist The features obtained after modal fusion are obtained. ;

[0032]

[0033] in, It is a non-linear activation function; The number of attention heads introduced; For attention head index; For nodes The neighborhood group, The meaning of node It is a node Neighbor set One of the elements; It is the first The learning weight matrix for each attention head, for each node Features Perform a linear transformation, projecting it onto a new feature space, with each weight being... ;

[0034] The pre-trained deep feature fusion module has the same training period, loss function, and early termination condition as the PD graph modality feature extractor. During pre-training, the learning rate is set to 0.0005, while the fine-tuning learning rate of the PD graph modality feature extractor and the track modality feature extractor is reduced to 0.00005. The track modality deep features and the PD graph modality deep features are input into the deep feature fusion module to classify the samples.

[0035] Secondly, a multi-source target recognition system based on deep feature fusion is provided to implement the multi-source target recognition method based on deep feature fusion as described in the first aspect, including:

[0036] Data acquisition module: acquires paired raw features and labels, the raw features including target PD map and target track information;

[0037] Preprocessing module: preprocesses the raw features acquired by the data acquisition module to construct a neural network model dataset; the preprocessed target PD map is denoted as PD map modal engineering feature, and the preprocessed target trajectory information is denoted as trajectory modal engineering feature;

[0038] Single-modal feature extraction module: The single-modal feature extractor includes a PD image modal feature extractor and a track modal feature extractor; the PD image modal feature extractor extracts the PD image modal engineering features into PD image modal depth features, and the track modal feature extractor extracts the track modal engineering features into track modal depth features;

[0039] Feature fusion module: fuses the PD map modal depth features and track modal depth features extracted by the single-modal feature extraction module, and outputs the fused features;

[0040] Target recognition module: Uses a linear classifier to classify the fused features and identify the target.

[0041] Thirdly, a storage medium is provided, the storage medium storing a computer program, which, when executed by a processor, implements the multi-source target recognition method based on deep feature fusion as described in the first aspect.

[0042] The beneficial effects of this invention are as follows:

[0043] In existing technologies, the identification of low-altitude aircraft relies on single-mode data, which is greatly affected by boundary conditions such as environment and climate, and often fails to identify targets.

[0044] In this application, neural networks are used to extract radar track mode depth features and PD map mode depth features respectively, and the two are fused for the final target recognition task. Compared with single-modality recognition methods, multimodal recognition can effectively overcome environmental interference through information complementarity between different modes, thereby significantly improving the stability and accuracy of recognition. Attached Figure Description

[0045] Figure 1 This is a technical principle diagram of the multi-source target recognition method based on deep feature fusion according to an embodiment of the present invention;

[0046] Figure 2 This is a flowchart of a multi-source target recognition method based on deep feature fusion according to an embodiment of the present invention;

[0047] Figure 3 This is the sample serial number in the embodiments of the present invention. A schematic diagram illustrating the principle of multimodal feature fusion in real time;

[0048] Figure 4 This is a schematic diagram illustrating the weight generation principle of the attention head in an embodiment of the present invention. Detailed Implementation

[0049] The present invention will now be described in further detail.

[0050] A multi-source target recognition method based on deep feature fusion, as shown in the technical principle diagram below. Figure 1 As shown. Real-time alignment data is input to obtain the paired target's original features, including target track information and target PD image. The original features are then preprocessed: the target track information is processed into track modal features, input to the track modal feature extractor to extract track modal depth features; the target PD image is processed into PD image modal features, input to the PD image modal feature extractor to extract PD image modal depth features; then the track modal depth features and PD image modal depth features are input to the depth feature fusion module for information interaction and fusion at the feature level. The fused features are then sent to classifier three for classification, outputting the fused feature classification confidence score, and selecting the category with the highest confidence score as the final target recognition result. After obtaining the target's original features, a multimodal target recognition method using depth feature fusion can be executed on a single machine. The term "multimodal" refers to the signal source used for UAV recognition encompassing modal data in various data formats, including radar PD images and track features.

[0051] When only a single modality exists, the recognition task can be completed independently, as follows: When only track modality features are available, the track modality depth features are fed into classifier one for classification, the track modality classification confidence score is output, and the category with the highest confidence score is selected as the final target recognition result; when only PD image modality features are available, the PD image modality depth features are fed into classifier two for classification, the PD image modality classification confidence score is output, and the category with the highest confidence score is selected as the final target recognition result.

[0052] This invention provides a multi-source target recognition method based on deep feature fusion, the process of which is as follows: Figure 2 As shown, the specific steps include the following:

[0053] Step 1: Obtain the original features and labels of the pairs.

[0054] In this embodiment, the original features for pairing include the target PD map and the target track information. The original features are multimodal data, meaning they come from different data sources: the target PD (Pulse-Doppler) map belongs to perceptual modality data, while the track information belongs to abstract semantic modality data. To ensure perceptual consistency and decision-making accuracy across multimodal data, real-time aligned data is required as input. The following are the specific steps for obtaining the original features and labels for pairing.

[0055] Step 1-1: Obtain the complete PD image of the target.

[0056] The raw echo data is parsed from the message received by the radar, and then pulse compression parameters are used to perform pulse compression processing on the raw echo data. Finally, slow-time FFT processing is performed on the pulse-compressed data to obtain the complete received PD map. Different targets have different PD map characteristics. UAV targets often rely on rotors, so they will show a certain Doppler broadening in the Doppler dimension of the PD map, while vehicle targets are larger, so they show a longer broadening in the range dimension.

[0057] Steps 1-2: Obtain target track information

[0058] By setting thresholds for the PD image and using methods such as CFAR (Constant False Alarm Rate) for target detection, the position and velocity information of the target can be obtained. Then, by combining the radar's operating parameters and main lobe orientation information, the target's corresponding velocity, altitude, heading angle, etc., can be calculated to form the target trajectory information corresponding to the current moment.

[0059] Steps 1-3: Obtain the target local PD image

[0060] The target track information can be obtained from steps 1-2. However, in addition to the target track information, there is also target information in the complete target PD image. In order not to lose target information other than the target track information, a partial target PD image is extracted from the complete target PD image, as follows:

[0061] Centered on the target, a K*K image patch is cropped from the complete PD image obtained in step 1-1. The acquired K*K image patch is single-channel image data. To avoid the influence of background noise, the value of K is usually between 16 and 64. The selection of K needs to ensure that the target is within the range of the complete PD image. The complete PD image reflects all received information of the beam emitted by the radar in a certain direction, rather than just the target information in that direction. Therefore, it is necessary to obtain a local PD image of the target based on the complete PD image to avoid the influence of background noise on identification.

[0062] Steps 1-4: Annotate the original features

[0063] The target trajectory information and target PD map content are manually verified, and labeled with binary classification tags for UAVs / non-UAVs. The target PD map includes a complete target PD map and a partial target PD map.

[0064] Step 2: Preprocess the original features to construct the neural network model dataset.

[0065] The original features include target track information and target PD map. There is no fixed execution order requirement for the preprocessing of target track information (steps 2-1 to 2-2) and the preprocessing of target PD map (steps 2-3). They are described in sequence below for ease of description, but the order should not be construed as a limitation of the present invention.

[0066] Step 2-1: Construct track mode features

[0067] Since the statistical features of a single waypoint may contain measurement errors, using the statistical features of multiple waypoints helps reduce measurement errors and improve recognition accuracy. Therefore, based on a preset sliding window of size M, where M is a positive integer greater than 0, the statistical features of the waypoints within the window are calculated in each dimension. The dimensions include F features such as altitude, speed, heading angle, longitude, and latitude, where F is a positive integer greater than 0; the statistical features include at least the maximum value, minimum value, average value, and standard deviation. Each dimension's features generate corresponding four statistical measures, which are finally aggregated into a feature vector of dimension 4F, serving as the model input. The 20 waypoint modal features involved in this embodiment are shown in Table 1. For example, if 10 waypoints are input, the dimension of the input matrix is ​​10*20, with 10 waypoints and 20 statistical features.

[0068] Table 1 List of Track Modal Features

[0069]

[0070] Step 2-2: Preprocess track modal features

[0071] The length-related track modal features are denoted as the first feature. The first feature is then normalized and mapped to a predetermined target interval. Specifically: For length-related track modal features such as speed and altitude, based on the distribution in the training data, the largest 1% of values ​​are denoted as H, and the smallest 1% are denoted as L. The input features are then... Cropping to between H and L yields the cropped features. ,Right now:

[0072]

[0073] The trajectory modal features involving angles are denoted as the second feature. An angle value conversion is then performed on the second feature, that is, the angle unit is converted from degrees to radians. Specifically, for trajectory modal features such as longitude, latitude, and heading angle, which are in angle units, a function is used... Convert angle values ​​from degrees to radians without cropping the values. That is:

[0074]

[0075] Using the maximum and minimum values ​​of the first feature in the training set, all length-related features in the dataset are normalized; using the maximum and minimum values ​​of the second feature in the training set, all angle-related features in the dataset are normalized; after obtaining the processed data of different feature types... Then, calculate the maximum and minimum values ​​over the entire training set, and use these two values ​​to apply to all training sets. Min-Max normalization is used to obtain the trajectory modal features of the final input model. It can be predicted that this value will be between 0 and 1 on the training set, while on the test set, this value may exceed the range of 0 to 1. The preprocessed trajectory modal features are denoted as the engineered trajectory modal features.

[0076] Steps 2-3: Preprocessing PD map modal features

[0077] The PD image is converted to decibels, and the resulting decibel values ​​are normalized and mapped to a predetermined target range. The PD image is a two-dimensional matrix whose elements are linear amplitude values ​​with a large dynamic range. The PD spectrum is a row in the PD matrix, representing a Doppler amplitude distribution over a range gate, and its values ​​are also linear. The PD spectrum is converted to decibels (dB processing), converting each linear value in the PD spectrum to a decibel value. By performing dB preprocessing on the PD image modal features, the problem of excessively large numerical distributions in the PD image is solved. Then, Min-Max normalization is used to clip the dB values ​​to between 0 and 100. The preprocessed PD image modal features are denoted as PD image modal engineering features and serve as the final input to the model's PD image modal features. The Min-Max normalization operation method is the same as in step 2-2.

[0078] Steps 2-4: Split the dataset

[0079] When constructing the dataset, it is divided into training, cross-validation, and test sets based on flight paths. Each flight path point can only be assigned to one of these sets. The dataset is divided into these sets according to a 50%-25%-25% ratio, with the training set used in step three, the cross-validation set in step four, and the test set in step five. This ensures that all flight path points on a given path belong to a specific set. Then, the PD graph modal engineering features (if any) and the flight path modal engineering features corresponding to each flight path point are used as feature inputs to the model for that point. It should be noted that due to data acquisition issues, the PD graph modal engineering features for a specific flight path point are often missing.

[0080] Step 3: Construct and pre-train a single-modal feature extractor

[0081] In this embodiment, the single-modal feature extractors that need to be constructed and pre-trained include a PD graph modal feature extractor and a track modal feature extractor. The pre-training of the PD graph modal feature extractor (steps 3-1 to 3-2) and the pre-training of the track modal feature extractor (steps 3-3 to 3-4) do not have a fixed execution order. They are described sequentially below for ease of description, but this order should not be construed as a limitation of the invention.

[0082] Step 3-1: Construct a PD graph modal feature extractor

[0083] The PD image modal feature extractor is built using a CNN (Convolutional Neural Network). The network input is the original PD image signal, i.e., the raw echo amplitude information of the target, which is an image with one channel, where the vertical axis represents the range gate and the horizontal axis represents the Doppler frequency. Its network structure includes three convolutional layers, employing an adaptive max-pooling layer to obtain the feature vector of the entire image. Based on the feature vector of the entire image and subsequent fully connected layers, the PD image modal depth features are obtained. , It is a one-dimensional vector, which is regarded as the encoded information of the PD map. It will be used for subsequent multimodal feature fusion and single-modal pre-training. The classifier used for single-modal pre-training is a linear classifier with a fully connected layer.

[0084] The PD image modality feature extractor's network structure consists of three convolutional layers, with the following specific parameters:

[0085] The first convolutional layer is configured to receive input data with 1 channel and perform convolution operation using a convolutional kernel of size 11×11 with padding size of 2, and output a feature map with 16 channels.

[0086] The first max pooling layer is configured to apply a 3×3 max pooling operation to the output of the first convolutional layer, with a stride of 2.

[0087] The second convolutional layer is configured to receive the output of the first max pooling layer and perform a convolution operation using a 3×3 convolutional kernel with a padding size of 1, outputting a feature map with 32 channels.

[0088] The second max pooling layer is configured to apply a 3×3 max pooling operation to the output of the second convolutional layer, with a stride of 2.

[0089] The third convolutional layer is configured to receive the output of the second max pooling layer and perform a convolution operation using a 3×3 convolutional kernel with a padding size of 1, outputting a feature map with 64 channels.

[0090] The third max pooling layer is configured to apply a 3×3 max pooling operation to the output of the third convolutional layer, with a stride of 2.

[0091] An adaptive max-pooling layer is configured to perform an adaptive max-pooling operation on the output of the third max-pooling layer, and the output space is a feature map of size 1×1.

[0092] The fully connected layer is configured to receive the output of the adaptive max-pooling layer and output a 64-dimensional feature vector.

[0093] The hyperparameter configuration table of the three-layer convolutional network for the PD image modality feature extractor is shown in Table 2.

[0094] Table 2. Hyperparameter configuration of the three-layer convolutional network for the PD graph modality feature extractor.

[0095]

[0096] Step 3-2: Pre-trained PD graph modal feature extractor

[0097] The training parameters, such as the training period, loss function, and early termination condition, are set. Since the recognition task in this invention is a classification problem, the cross-entropy loss function is typically chosen. The cross-entropy loss function is used to calculate the degree of deviation between the model's predicted probability distribution and the true probability distribution. The larger the loss value, the greater the prediction bias; the smaller the loss value, the more accurate the prediction. The goal of training the model is to minimize the loss value. In this embodiment, a training period of 200 is selected, and the cross-entropy loss function CE (Cross-Entropy Loss, representing the average prediction error of a batch of N samples) is chosen, as shown in the following formula.

[0098]

[0099] in This is the sample index, representing the sample number; The number of samples; This is a category index, representing the category number; The total number of categories; One-hot encoding for the tag; This is the model's predicted value, i.e., the model's predicted value for the [number]th [year]. The sample belongs to the first The confidence level for each category was set. The early termination condition was set to ensure that the confidence level (CE) did not decrease for 20 consecutive training epochs. The learning rate was set to 0.0005 during pre-training.

[0100] Step 3-3: Construct a trajectory modal feature extractor

[0101] The track modal feature extractor is built using a GRU (Gated Recurrent Unit) neural network. It continuously inputs the features described in Table 1, such as velocity and altitude, from track points in M ​​frames. Based on the input of the entire sequence, the feature vector of the entire sequence can be obtained. , It is a one-dimensional vector, which is regarded as the encoded information of the trajectory and will be used for subsequent multimodal feature fusion and single-modal pre-training. The classifier used for single-modal pre-training is also a fully connected linear classifier. In this embodiment, the number of input features of GRU is 20 (the specific information of the input features is shown in Table 1), and the hidden layer size is 64 (consistent with the output channels of the PD graph modality feature extractor designed in step 3-1).

[0102] Steps 3-4: Pre-trained track modal feature extractor

[0103] The same method as the PD graph modality extractor pre-training was used, setting training parameters such as the number of training iterations, loss function, and early termination condition. The same cross-entropy function and labels were used to construct the loss function (same as in step 3-2), except that the learning rate was set to 0.00005 during pre-training. This is because the track modality extractor model uses a GRU model structure, which is susceptible to gradient explosion. Therefore, compared to the PD graph modality extractor pre-training, the track modality extractor pre-training uses a smaller learning rate in hopes of achieving more stable training.

[0104] Step 4: Construct and pre-train a deep feature fusion module

[0105] Step 4-1: Construct a deep feature fusion module

[0106] Multimodal deep features (including track modality deep features and PD graph modality deep features) are incorporated into a graph model for information interaction and fusion at the feature level, outputting fused potential target features. The deep feature fusion module in this embodiment is built using GAT (Graph Attention Network), which handles features from any modality. It simultaneously possesses multiple modal features, and integrates the features of multiple modalities. The features obtained after multimodal fusion are obtained. .

[0107]

[0108] in, It is a non-linear activation function; The number of attention heads to be introduced is predetermined manually; in this embodiment, it is set to 2. This is the index of the attention head, representing the sequence number of the attention head; For nodes The neighborhood group, The meaning of node It is a node Neighbor set One of the elements; It is the first The learning weight matrix for each attention head, for each node Features Perform a linear transformation, projecting it onto a new feature space, with each weight being... Weight A schematic diagram of the generation principle is shown below. Figure 4 As shown, the input variable is the current mode. and other modes The parameters that can be optimized are: and , It is a learning weight matrix. It is a length of Twice the column vector, It is an activation function. It's a vector concatenation operation, from which weights can be calculated. .

[0109]

[0110] Because of the introduction of the Softmax (flexible maximum activation function) operation, when there are multiple modalities, the information between different modalities will interact through a weighted average, while when there is only one modality in the graph, If the value is always equal to 1, the depth information of that node will not be updated. (Fused features) It was used for a subsequent task, which in this case is a classification task, where a linear classifier is still used to classify the samples.

[0111] like Figure 3 As shown, when At that time, in the current mode For example, features of other modalities exist simultaneously. , , , , Assuming These are the deep features of the trajectory mode in step 3-3. , , , , The deep features obtained from the target PD images of several consecutive frames are calculated in step 3-1 and then obtained by concatenating them pairwise. (specifically) , , , , , ), and then according to different and the corresponding and The updated fusion features are calculated. This is used for subsequent classification tasks.

[0112] Step 4-2: Pre-trained deep feature fusion module

[0113] Different modal feature extractors converge at different speeds. Therefore, the training of a deep feature fusion module with multimodal features differs from the pre-training of a single-modal feature extractor. It cannot be trained synchronously with the single-modal feature extractor; instead, training should begin only after the single-modal feature extractor has completed its training. Otherwise, instability in the single-modal features may occur. During training, the learning rate of the deep feature fusion module with multimodal features is 0.0005, while the fine-tuning learning rate of the feature extractor decreases to 0.00005. Other settings are the same as the pre-training of the single-modal feature extractor. The cross-entropy function is chosen as the loss function, and 200 training iterations are selected. If the loss function value does not decrease for 20 consecutive training epochs, training is terminated prematurely.

[0114] Step 5: Perform inference based on the pre-trained deep feature fusion module.

[0115] The paired raw features and labels are input into a unimodal feature extractor to obtain multimodal deep features. These deep features are then fused and used for identification, outputting the final target recognition result. The details are as follows:

[0116] To obtain the original features and labels for pairing, refer to step one to obtain target track information, target PD map and corresponding labels;

[0117] The trajectory modal feature extractor (pre-training completed in step 3-4) extracts trajectory modal depth features; the PD graph modal feature extractor (pre-training completed in step 3-2) extracts PD graph modal depth features.

[0118] The trajectory modal depth features and PD map modal depth features extracted by the single modal feature extractor are used as inputs and fused based on the deep feature fusion module (pre-trained in step 4-2). After fusion, a linear classifier is used to classify the samples.

[0119] To verify the superiority of the method in this application and quantify its performance gap with the single-modal method, a comparison was conducted on the same test dataset, and the quantification results are shown in Table 3.

[0120] Table 3. Comparison of quantization performance between the recognition method of this application and the single-modal recognition method.

[0121]

[0122] As shown in Table 3, the identification method of this application significantly outperforms any single-modality method in all key performance indicators, including detection rate and false alarm rate (calculated as the average accuracy of all categories under a positioning accuracy threshold of 0.5). This demonstrates that multi-source information fusion can effectively overcome the perception limitations of a single modality and achieve more stable and reliable identification in complex environments through complementary advantages.

[0123] The present invention also provides a multi-source target recognition system based on deep feature fusion, comprising the following modules:

[0124] Data Acquisition Module: Acquires paired raw features and labels. The raw features include the target PD image and target track information, specifically as follows: The data acquisition module parses the raw echo data from the radar received messages, then performs pulse compression processing on the raw radar echo data using pulse compression parameters. Finally, it performs slow-time FFT processing on the pulse-compressed data to obtain the received complete PD image. The data acquisition module sets a threshold for the PD image and uses methods such as CFAR (Constant False Alarm Rate) for target detection. This allows it to obtain the target's position and velocity information. Combined with radar operating parameters and main lobe orientation information, it calculates the target's corresponding velocity, altitude, heading angle, etc., constituting the target track information for the current moment. To avoid losing target information other than the target track information, the data acquisition module extracts a partial PD image of the target from the complete target PD image. The content of the target track information and target PD image is manually verified, and labels are assigned for annotation.

[0125] Preprocessing module: preprocesses the raw features acquired by the data acquisition module to construct a neural network model dataset; the preprocessed target PD map is denoted as the PD map modal engineering feature, and the preprocessed target track information is denoted as the track modal engineering feature, as detailed below:

[0126] When constructing the dataset, it is divided into training, cross-validation, and test sets based on flight paths. Each flight path point can only be assigned to one of these sets. The dataset is divided into these sets according to a 50%-25%-25% ratio, ensuring that all flight path points belong to a specific set. Then, the PD graph modal engineering features (if any) and the flight path modal engineering features corresponding to each flight path point are used as feature inputs to the model. It should be noted that due to data acquisition issues, the PD graph modal engineering features for a specific flight path point are often missing.

[0127] When constructing the PD graph modal engineering features, track modal features are constructed to obtain the statistical features of track points in various dimensions. These dimensions include F features such as altitude, speed, heading angle, longitude, and latitude, where F is a positive integer greater than 0. The statistical features include at least the maximum, minimum, average, and standard deviation. Each dimension's features generate corresponding four statistical measures, which are ultimately aggregated into a 4F-dimensional feature vector, serving as the model input. Track modal features involving length are designated as the first feature, and are normalized and mapped to a predetermined target interval. Track modal features involving angles are designated as the second feature, and are converted from degrees to radians. Using the maximum and minimum values ​​of the first feature in the training set, all length-related features in the dataset are normalized. Similarly, using the maximum and minimum values ​​of the second feature in the training set, all angle-related features in the dataset are normalized. After obtaining the processed data for different feature types, the maximum and minimum values ​​across the entire training set are calculated, and these two values ​​are used to apply Min-Max normalization to all datasets to obtain the final trajectory modal features input to the model. It can be anticipated that this value will be between 0 and 1 on the training set, while on the test set, it may exceed this range. The preprocessed trajectory modal features are denoted as the trajectory modal engineering features.

[0128] When constructing the engineered features of the flight path modalities, the PD map is converted to decibels, and the resulting decibel values ​​are normalized and mapped to a predetermined target range. The PD map is a two-dimensional matrix whose elements are linear amplitude values ​​with a large dynamic range. The PD spectrum is a row in the PD matrix, representing a Doppler amplitude distribution over a range gate, and its values ​​are also linear. The PD spectrum is converted to decibels (dB processing), converting each linear value in the PD spectrum to a decibel value. This dB preprocessing of the PD map modal features solves the problem of excessively large numerical distributions in the PD map. Then, Min-Max normalization is used to clip the dB values ​​to between 0 and 100. The preprocessed PD map modal features are denoted as the engineered PD map modal features and serve as the final input to the model's PD map modal features.

[0129] Single-modal feature extraction module: The single-modal feature extractor includes a PD image modal feature extractor and a track modal feature extractor; the PD image modal feature extractor extracts the PD image modal engineering features into PD image modal depth features, and the track modal feature extractor extracts the track modal engineering features into track modal depth features;

[0130] The PD image modality feature extractor is built using a CNN (Convolutional Neural Network). Its network structure includes three convolutional layers, and an adaptive max-pooling layer is used to obtain the feature vector of the entire image. Based on the feature vector of the entire image and subsequent fully connected layers, the PD image modality depth features are obtained. , This will be used for subsequent multimodal feature fusion and single-modal pre-training. The classifier used for single-modal pre-training is also a fully connected linear classifier. Specific parameters are shown in Table 2. The loss function is chosen as the cross-entropy loss function, and the learning rate is set to 0.0005 during pre-training. The early termination condition is set to CE (Equal Access) without decreasing for 20 consecutive training epochs. The track modality feature extractor is built using a GRU (Gated Recurrent Unit) neural network. It continuously inputs track point features from M frames, and based on the input of the entire sequence, it can obtain the feature vector of the entire sequence. , This will be used for subsequent multimodal feature fusion and single-modal pre-training. The classifier used for single-modal pre-training is also a fully connected linear classifier. In this embodiment, the GRU has 20 input features (the specific information of the input features is shown in Table 1), and the hidden layer size is 64.

[0131] Feature fusion module: fuses the PD image modal depth features and track modal depth features extracted by the PD image modal feature extractor in the single modal feature extraction module, and outputs the fused features;

[0132] The feature fusion module is built using GAT (Graph Attention Network). For any feature of a given modality, multiple modalities exist simultaneously. By combining the features from multiple modalities, a multimodal fused feature is obtained. The learning rate of the feature fusion module is 0.0005, while the fine-tuning learning rate of the feature extractor is reduced to 0.00005. The loss function is the cross-entropy function, with 200 training iterations. The early termination condition is set to ensure that the loss does not decrease for 20 consecutive training epochs.

[0133] Target recognition module: Uses a linear classifier to classify the fused features output by the feature fusion module to identify targets.

[0134] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, implements the multi-source target recognition method based on deep feature fusion of the present application.

[0135] While the present invention has been disclosed above with reference to preferred embodiments, these embodiments are not intended to limit the invention. Any equivalent changes or modifications made without departing from the spirit and scope of the invention are also within the scope of protection of the invention. Therefore, the scope of protection of the present invention should be determined by the claims of this application.

Claims

1. A multi-source target recognition method based on deep feature fusion, characterized in that, Includes the following steps: Obtain the original features and labels for the pairing, wherein the original features include the target PD map and target track information; The original features are preprocessed to construct a neural network model dataset; The preprocessed target PD image is denoted as PD image modal feature, and the preprocessed target track information is denoted as track modal feature; A single-modal feature extractor is constructed and pre-trained, comprising a PD image modal feature extractor and a track modal feature extractor; the PD image modal feature extractor extracts the PD image modal features into PD image modal depth features, and the track modal feature extractor extracts the track modal features into track modal depth features; A deep feature fusion module is constructed and pre-trained, wherein the deep features include the PD map modal depth features and the track modal depth features; Inference is performed based on the pre-trained deep feature fusion module.

2. The multi-source target recognition method based on deep feature fusion according to claim 1, characterized in that, When preprocessing the original features to construct the neural network model dataset, the specific steps are as follows: The preprocessing of the target trajectory information is as follows: The target trajectory information includes information related to length and information related to angle; the information related to length is denoted as the first feature, and the information related to angle is denoted as the second feature; the first feature is normalized and mapped to a predetermined target interval; the second feature is converted to an angle value, that is, the angle unit is converted from degrees to radians; using the maximum and minimum values ​​of the first feature in the training set, all length-related features in the dataset are normalized; using the maximum and minimum values ​​of the second feature in the training set, all angle-related features in the dataset are normalized; the processed target trajectory information is the trajectory modal feature. The preprocessing of the PD image is as follows: the PD image is converted to decibels, the obtained decibel values ​​are normalized, and mapped to a predetermined target interval; the processed PD image is the PD image modal feature. When constructing the dataset, the dataset is divided into a training set, a cross-validation set, and a test set, with each track as a unit; all track points on a track can only be assigned to one of the training set, the cross-validation set, or the test set.

3. The multi-source target recognition method based on deep feature fusion according to claim 1, characterized in that, When constructing a single-modal feature extractor, the specific steps are as follows: Constructing a PD image modality feature extractor: A convolutional neural network is used, with the network structure containing three convolutional layers. The specific parameters are as follows: The first convolutional layer is configured to receive input data with 1 channel and perform convolution operation using a convolutional kernel of size 11×11 with padding size of 2, and output a feature map with 16 channels. The first max pooling layer is configured to apply a 3×3 max pooling operation to the output of the first convolutional layer, with a stride of 2. The second convolutional layer is configured to receive the output of the first max pooling layer and perform a convolution operation using a 3×3 convolutional kernel with a padding size of 1, outputting a feature map with 32 channels. The second max pooling layer is configured to apply a 3×3 max pooling operation to the output of the second convolutional layer, with a stride of 2. The third convolutional layer is configured to receive the output of the second max pooling layer and perform a convolution operation using a 3×3 convolutional kernel with a padding size of 1, outputting a feature map with 64 channels. The third max pooling layer is configured to apply a 3×3 max pooling operation to the output of the third convolutional layer, with a stride of 2. An adaptive max-pooling layer is configured to perform an adaptive max-pooling operation on the output of the third max-pooling layer, and the output space is a feature map of size 1×1. A fully connected layer is configured to receive the output of the adaptive max-pooling layer and output a 64-dimensional feature vector. Pre-trained PD image modality feature extractor: The classifier of the PD image modality feature extractor is a fully connected linear classifier. The loss function is selected as the cross-entropy loss function, and the training period is set to 200. The expression of the cross-entropy loss function (CE) is as follows: in This is the sample index, representing the sample number; The number of samples; This is a category index, representing the category number; The total number of categories; One-hot encoding for the tag; This is the model's predicted value, i.e., the model's predicted value for the [number]th [year]. The sample belongs to the first The confidence level of each category is set; the early termination condition is set to CE not decreasing for 20 consecutive training cycles; the learning rate is set to 0.0005 during pre-training; the PD graph modal features and labels are input into the PD graph modal feature extractor to obtain the PD graph modal depth features; Constructing a track modality feature extractor: A gated recurrent unit neural network is used to build the feature vector of the entire sequence based on the input of track point features from M consecutive frames, where M is a positive integer greater than 0. The hidden layer size is consistent with the dimension of the feature vector output by the fully connected layer in the PD graph modality feature extractor. Pre-trained track modality feature extractor: The classifier of the track modality feature extractor is a fully connected linear classifier. The training period, loss function, and early termination condition are the same as those of the pre-training of the PD graph modality feature extractor. The learning rate is set to 0.00005 during the pre-training process. The track modality features and labels are input into the track modality feature extractor to obtain the track modality depth features.

4. The multi-source target recognition method based on deep feature fusion according to claim 3, characterized in that, When constructing and pre-training the deep feature fusion module, the specific steps are as follows: Constructing a deep feature fusion module: A deep feature fusion module is built using a graph attention network for modal features. Modal features also exist The features obtained after modal fusion are obtained. ; in, It is a non-linear activation function; The number of attention heads introduced; For attention head index; For nodes The neighborhood group, The meaning of node It is a node Neighbor set One of the elements; It is the first The learning weight matrix for each attention head, for each node Features Perform a linear transformation, projecting it onto a new feature space, with each weight being... ; The pre-trained deep feature fusion module has the same training period, loss function, and early termination condition as the PD graph modality feature extractor. During pre-training, the learning rate is set to 0.0005, while the fine-tuning learning rate of the PD graph modality feature extractor and the track modality feature extractor is reduced to 0.00005. The track modality deep features and the PD graph modality deep features are input into the deep feature fusion module to classify the samples.

5. A multi-source target recognition system based on deep feature fusion, used to implement the multi-source target recognition method based on deep feature fusion as described in claim 1, characterized in that, include: Data acquisition module: acquires paired raw features and labels, the raw features including target PD map and target track information; Preprocessing module: preprocesses the raw features acquired by the data acquisition module to construct the neural network model dataset; The preprocessed target PD image is denoted as the PD image modal engineering feature, and the preprocessed target track information is denoted as the track modal engineering feature; Single-modal feature extraction module: The single-modal feature extractor includes a PD image modal feature extractor and a track modal feature extractor; the PD image modal feature extractor extracts the PD image modal engineering features into PD image modal depth features, and the track modal feature extractor extracts the track modal engineering features into track modal depth features; Feature fusion module: fuses the PD map modal depth features and track modal depth features extracted by the single-modal feature extraction module, and outputs the fused features; Target recognition module: Uses a linear classifier to classify the fused features and identify the target.

6. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the multi-source target recognition method based on deep feature fusion as described in any one of claims 1-4.