Multi-modal hierarchical hybrid fusion radar target identification method
Through the multimodal hierarchical hybrid fusion method, the hierarchical feature extraction and fusion network are used to solve the problem of unstable multimodal data recognition performance in the existing technology, and achieve higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510156229.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-12
AI Technical Summary
In the multimodal data recognition scenario, the existing radar target recognition method has unstable recognition performance and poor promotion due to semantic mismatch and modal imbalance of each modal data.
A multimodal hierarchical hybrid fusion radar target recognition method is proposed, and a feature extraction module including track feature extraction network, HRRP feature extraction network and JEM feature extraction network is constructed to perform hierarchical feature extraction and alignment. Then, a feature fusion module for feature hierarchical fusion enhancement network, key-value attention network, aggregate feature classification network, single-modal classification network and hybrid fusion network is constructed to perform feature fusion and decision-making fusion.
It effectively improves the feature alignment quality of heterogeneous data, alleviates semantic deviations between different modal data, improves the accuracy and robustness of recognition, and has better promotion and stability.
Smart Images

Figure CN120028767A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of radar target recognition, and in particular relates to a multi-modal hierarchical hybrid fusion radar target recognition method. Background Art
[0002] Radar target recognition uses the radar echo signal of the target to determine the target type. Compared with the traditional target detection task, the target recognition task requires measuring more in-depth target details. In order to achieve this goal, features can be obtained from multiple modes, including radar high-resolution range profile, modulation spectrum and track information. Among them, the radar high-resolution range profile (HRRP) is the projection vector of the target scattering point echo in the direction of the radar ray obtained by broadband radar signal, which provides the distribution of the target scattering point along the distance direction and contains rich structural information such as the target geometric size. The modulation spectrum (Jet Engine Modulation, JEM) obtains spectral information by analyzing the micro-motion characteristics of the target, and uses the interaction between the radar signal and the target to extract the micro-motion characteristics of the target, such as rotation and vibration, and provides the distribution of the target in the frequency domain, including rich information such as the target's motion mode and structural characteristics. Track information is the trajectory data obtained by the radar system continuously tracking the change of the target position, recording the target's motion path in space, and providing dynamic information such as the target's speed, acceleration and direction. The track information contains the target's motion mode and behavior characteristics, which is of great value for target recognition and prediction tasks. At present, the study of separate recognition methods based on radar HRRP or JEM is one of the important ways to achieve radar target recognition.
[0003] Traditional target recognition methods are mostly based on single-modal data, which may not be robust enough in practical applications. Single-modal data is easily disturbed and restricted in complex environments, resulting in reduced recognition performance. In contrast, multimodal fusion methods can provide more comprehensive target feature information by combining data from different modalities.
[0004] For multi-modal radar target recognition, the existing technology provides the following methods:
[0005] Prior art 1: The University of Electronic Science and Technology of China proposed a time-frequency fusion method in its patent application “Multimodal radar active deception interference identification method based on small samples” (CN202310004984.X). This method extracts feature parameters in the time domain, frequency domain, and time-frequency domain, and constructs a multimodal fusion prototype network for feature fusion and classification.
[0006] Prior art 2: China Electronics Technology Group Corporation and Nanjing Institute of Electronic Technology proposed a fusion method for identifying aircraft targets by comprehensively utilizing the wide-band and narrow-band features of the aircraft targets in their published article “Aerial target identification based on radar wide-band and narrow-band multi-feature information fusion” (DOI: 10.16592 / j.cnki.1004-7859.2015.07.005). The method analyzed the characteristics of the wide-band and narrow-band identification system using the power spectrum of HRRP and the target speed and altitude as identification features, and realized the decision-making layer fusion target identification based on the wide-band and narrow-band identification results by using the DS evidence theory.
[0007] Prior art 3: Chongqing University proposed a mathematical model for aircraft behavior recognition using a joint data approach in its article “Using a multimodal approach to aircraft behavior recognition based on trajectory data” (DOI: 10.3390 / electronics13020367). This method designs a deep network including feature abstraction, cross-modal fusion, and classification layers to obtain multi-scale features, and enhances the extracted features with the longitude and latitude in the track information, and finally classifies the enhanced features of the track information.
[0008] Prior art 4: Lanzhou University of Technology proposed a dual-modal network that fuses time domain data and frequency domain data in its patent application "A multi-modal radar HRRP target recognition method and system" (CN202311592869.5). This method first obtains dual-modal data by performing frequency domain analysis on the original HRRP data, then uses a collaborative attention block to perform feature fusion on the modal data, and finally outputs the recognition result.
[0009] However, the above prior art still has the following disadvantages:
[0010] The disadvantages of the prior art 1 are: the method is mainly based on time-frequency analysis of data to enhance information mining of radar data. The data alignment is not considered in the modal fusion process, so there may be a problem of semantic mismatch between the modal data.
[0011] The disadvantages of the existing technology 2 are: this method uses the recognition results of wide and narrow bands for decision fusion, and does not consider the semantic deviation and modal imbalance of multimodal data. When faced with modality loss in actual application scenarios, the generalization performance of the model is easily affected, and its manual selection of track information features for fusion relies heavily on the prior information of expert knowledge and has poor adaptability.
[0012] The disadvantages of existing technology 3 are: this method only uses the longitude dimension in the track information to enhance the features of other modal data, and its utilization of track information is relatively insufficient. In addition, the track information is directly integrated with other modal features without considering the distribution differences between modal data. Therefore, the recognition performance of this method is unstable and its generalizability is poor.
[0013] The disadvantages of the prior art 4 are as follows: This method performs modal expansion by obtaining the frequency-domain data of the original HRRP data. Theoretically, it is a deep feature extraction of single-modal HRRP data, and does not fully integrate multi-modal radar data such as narrowband data and track data. Therefore, the feasibility of this method is poor.
[0014] In summary, the existing radar target recognition methods have technical problems of unstable recognition performance and poor generalization ability caused by semantic mismatch and modal imbalance of each modal data in the multi-modal data recognition scenario. Summary of the Invention
[0015] In order to solve the above problems existing in the prior art, the present invention provides a multi-modal hierarchical hybrid fusion radar target recognition method. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0016] In a first aspect, the present invention proposes a multi-modal hierarchical hybrid fusion radar target recognition method, including:
[0017] Construct a multi-modal data set including multiple category targets; the multi-modal data set includes track information, HRRP echo signals, and JEM modulation spectra;
[0018] Construct a feature extraction module including a track feature extraction network, an HRRP feature extraction network, and a JEM feature extraction network; wherein, the track feature extraction network, the HRRP feature extraction network, and the JEM feature extraction network are respectively used to perform hierarchical feature extraction on the track information, HRRP echo signals, and JEM modulation spectra, and correspondingly obtain the hierarchical features of each modality;
[0019] Based on the multi-modal data set, using the track information as a link, apply multi-modal joint learning for hierarchical feature alignment to train the feature extraction module to obtain a trained feature extraction module;
[0020] Use the trained feature extraction module to process the multi-modal data set to obtain the aligned hierarchical features of each modality;
[0021] Construct a feature fusion module including a feature hierarchical fusion enhancement network, a keyless attention network, an aggregated feature classification network, a single-modal classification network, and a hybrid fusion network; wherein the feature hierarchical fusion enhancement network is used to fuse the hierarchical features of the aligned modes to obtain fused enhanced features; the keyless attention network is used to aggregate the fused enhanced features to obtain aggregated features; the aggregated feature classification network is used to classify the aggregated features to obtain a first classification result; the single-modal classification network is used to process the hierarchical features of the aligned modes separately to obtain a second classification result of each mode; the hybrid fusion network is used to perform weighted fusion of the first classification result and the second classification result to obtain a recognition result;
[0022] The feature fusion module is trained using the aligned hierarchical features of each modality to obtain a trained feature fusion module, which together with the trained feature extraction module forms a multimodal hierarchical hybrid fusion radar target recognition model, thereby realizing multimodal radar target recognition.
[0023] Beneficial effects of the present invention:
[0024] The multimodal hierarchical hybrid fusion radar target recognition method provided by the present invention constructs a feature extraction module including a track feature extraction network, a HRRP feature extraction network, and a JEM feature extraction network, and a feature fusion module including a feature hierarchical fusion enhancement network, a keyless attention network, an aggregated feature classification network, a single-modal classification network, and a hybrid fusion network; in the feature extraction module, the hierarchical features of each modal data are obtained through each network, the feature alignment quality of the heterogeneous data is effectively improved, the rich information of each modal data from details to the whole is captured, and the understanding and processing capabilities of the data are enhanced; at the same time, with the track data that is most easily obtained in the radar data as a link, the data of each modality is embedded into a unified feature space, the feature alignment of the multimodal heterogeneous data is completed, the semantic deviation between the modal data is effectively alleviated, and the method has good generalizability. In the feature fusion module, through feature fusion enhancement and keyless attention network, the model can flexibly handle different numbers of modalities, which alleviates the impact of modality loss on recognition to a certain extent; and while realizing the hierarchical feature fusion of each modality, the individual classification results of each modality are integrated with the classification results of the aggregated features, thereby realizing the hybrid fusion of feature fusion and decision fusion, so that the model can more comprehensively utilize the advantages of multimodal data and improve the accuracy and robustness of recognition. This method clearly takes into full account the impact of modality loss on target recognition in real scenarios, and is more potential than traditional methods, with better generalization performance and better stability.
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flowchart of a multi-modal hierarchical hybrid fusion radar target recognition method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] See also Figure 1 , Figure 1 1 is a flow chart of a multi-modal hierarchical hybrid fusion radar target recognition method provided by an embodiment of the present invention, and the method mainly includes the following steps:
[0029] Step 1: Construct a multimodal dataset including multiple categories of targets; the multimodal dataset includes track information, HRRP echo signals, and JEM modulation spectra.
[0030] Specifically, the radar HRRP echo signals, JEM modulation spectra and track information of N categories of targets are sorted to form a multimodal data set. The multimodal data is based on the time of the HRRP echo signal, the JEM modulation spectrum is the JEM modulation spectrum of the same batch as the current HRRP data and with the smallest time difference; the track information is the historical M multidimensional track data points of the same batch as the current HRRP data and with the smallest time difference. Each category contains at least 1000 radar echo signals, where N≥5 and M≥15.
[0031] Step 2: Construct a feature extraction module including a track feature extraction network, a HRRP feature extraction network, and a JEM feature extraction network.
[0032] Among them, the track feature extraction network, HRRP feature extraction network, and JEM feature extraction network are used to perform hierarchical feature extraction on the track information, HRRP echo signal, and JEM modulation spectrum, respectively, and obtain the hierarchical features of each mode accordingly.
[0033] Specifically, in this embodiment, the track feature extraction network, the HRRP feature extraction network, and the JEM feature extraction network are all multi-layer networks, and the three have the same network structure, and each layer network includes a first Transformer encoder network;
[0034] The structure of the first Transformer encoder network includes a first input layer, a first position encoding layer, and a first Transformer encoder layer in sequence; the first Transformer encoder layer includes 4 layers of first Transformer encoders, each of which includes a 4-head attention mechanism and a feedforward neural network;
[0035] Among them, the first input layer is used to input data of the corresponding modality;
[0036] The first position encoding layer is used to perform position encoding on the input features;
[0037] The first Transformer encoder layer is used to perform hierarchical feature extraction on the features after position encoding, and obtains corresponding hierarchical features.
[0038] Optionally, as an implementation method, the track feature extraction network, HRRP feature extraction network and JEM feature extraction network constructed in this embodiment are all three-layer networks. The three feature extraction networks are respectively introduced in detail below.
[0039] 1. Track Feature Extraction Network
[0040] The track feature extraction network constructed in this embodiment gradually extracts track features in a hierarchical manner, from fine-grained to coarse-grained. This design can effectively capture information at different levels and is suitable for processing complex time series data. The network is a three-layer network, each layer network is a Transformer encoder network, which is called the first Transformer encoder network in this embodiment. The structure of each first Transformer encoder network is the first input layer, the first position encoding layer, and the first Transformer encoder layer in sequence; wherein,
[0041] The functions of each layer of the first level Transformer encoder network are as follows:
[0042] The first input layer: the input is the track information of a certain category of targets, that is, the multidimensional track data sequence, and the sequence length is len; Patch segmentation layer: the input track information sequence is segmented according to a smaller window size P1=16 and a step size S=16 to form multiple patch blocks; the first position encoding layer: position encoding is performed on each patch block, using a trainable position encoding matrix Wpos; the first Transformer encoder layer: contains 4 layers of the first Transformer encoder, each encoder contains an 8-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp, with a dimension of 256. At this point, len / 16 256-dimensional fine-grained feature patch blocks are obtained.
[0043] The functions of each layer of the second-level first Transformer encoder network are as follows:
[0044] First input layer: input is the 256-dimensional fine-grained feature patch block output by the first-level first Transformer encoder network; Patch aggregation layer: concatenate the two adjacent blocks of the input fine-grained feature patch block to obtain len / 8 patch blocks; First position encoding layer: position encode each patch block using the trainable position encoding matrix Wpos; First Transformer encoder layer: contains 4 layers of first Transformer encoders, each encoder contains an 8-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp with a dimension of 256. So far, len / 8 medium-grained feature extraction representations are obtained.
[0045] The functions of each layer of the third-level first Transformer encoder network are as follows:
[0046] First input layer: input is the 256-dimensional fine-grained feature patch block output by the second-level first Transformer encoder network; Patch aggregation layer: concatenate two adjacent blocks of the input fine-grained feature patch block to obtain len / 4 patch blocks; First position encoding layer: position encode each patch block using a trainable position encoding matrix Wpos; First Transformer encoder layer: contains 4 layers of first Transformer encoders, each encoder contains an 8-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp with a dimension of 256. At this point, len / 4 coarse-grained feature extraction representations are obtained.
[0047] 2. HRRP feature extraction network
[0048] The HRRP feature extraction network constructed in this embodiment also gradually extracts HRRP features in a hierarchical manner, from fine-grained to coarse-grained. This design can effectively capture information at different levels and is suitable for processing complex sequence data. Like the track feature extraction network, the HRRP feature extraction network is also a three-layer network. Each layer network is a Transformer encoder, which is called the first Transformer encoder network in this embodiment. Each first Transformer encoder structure is the first input layer, the first position encoding layer, and the first Transformer encoder layer in sequence; wherein,
[0049] The functions of each layer of the first level Transformer encoder network are as follows:
[0050] First input layer: input is HRRP data of a certain category of targets, and the length of the HRRP sequence is len; Patch segmentation layer: the input HRRP data is segmented according to a smaller window size P1=16 and a step size S=16 to form multiple Patch blocks; First position encoding layer: position encoding is performed on each Patch block, using a trainable position encoding matrix Wpos; First Transformer encoder layer: Contains 4 layers of first Transformer encoders, each encoder contains an 8-head attention mechanism and a feedforward neural network. The Patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp, with a dimension of 256. At this point, len / 16 256-dimensional fine-grained feature Patch blocks are obtained.
[0051] The functions of each layer of the second-level first Transformer encoder network are as follows:
[0052] First input layer: input is the 256-dimensional fine-grained feature patch block output by the first-level first Transformer encoder network; Patch aggregation layer: concatenate two adjacent blocks of the input fine-grained feature patch block to obtain len / 8 patch blocks; Position encoding layer: position encode each patch block using a trainable position encoding matrix Wpos; First Transformer encoder layer: contains 4 layers of first Transformer encoders, each encoder contains an 8-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp with a dimension of 256. So far, len / 8 medium-grained feature extraction representations are obtained.
[0053] The functions of each layer of the third-level first Transformer encoder network are as follows:
[0054] First input layer: input is the 256-dimensional fine-grained feature patch block output by the second-level first Transformer encoder network; Patch aggregation layer: concatenate two adjacent blocks of the input fine-grained feature patch block to obtain len / 4 patch blocks; First position encoding layer: position encode each patch block using a trainable position encoding matrix Wpos; First Transformer encoder layer: contains 4 layers of first Transformer encoders, each encoder contains an 8-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp with a dimension of 256. At this point, len / 4 coarse-grained feature extraction representations are obtained.
[0055] 3. JEM feature extraction network
[0056] The JEM feature extraction network constructed in this embodiment also gradually extracts JEM features in a hierarchical manner, from fine-grained to coarse-grained. This design can effectively capture information at different levels and is suitable for processing complex sequence data. Like the track feature extraction network and the HRRP feature extraction network, the JEM feature extraction network is also a three-layer network. Each layer network is a Transformer encoder, which is called the first Transformer encoder network in this embodiment. Each first Transformer encoder structure is sequentially the first input layer, the first position encoding layer, and the first Transformer encoder layer; wherein,
[0057] The functions of each layer of the first level Transformer encoder network are as follows:
[0058] First input layer: input is JEM data of a certain category of targets, and the length of the JEM sequence is len; Patch segmentation layer: the input JEM data is segmented according to a smaller window size P1=16 and a step size S=16 to form multiple Patch blocks; First position encoding layer: position encoding is performed on each Patch block, using a trainable position encoding matrix Wpos; First Transformer encoder layer: contains 2 layers of first Transformer encoders, each encoder contains a 4-head attention mechanism and a feedforward neural network. The Patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp, with a dimension of 256. At this point, len / 16 256-dimensional fine-grained feature Patch blocks are obtained.
[0059] The functions of each layer of the second-level first Transformer encoder network are as follows:
[0060] First input layer: input is the 256-dimensional fine-grained feature patch block output by the first-level first Transformer encoder network; Patch aggregation layer: concatenate two adjacent blocks of the input fine-grained feature patch block to obtain len / 8 patch blocks; First position encoding layer: position encode each patch block, using the trainable position encoding matrix Wpos; First Transformer encoder layer: contains 2 layers of first Transformer encoders, each encoder contains a 4-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp, with a dimension of 256. So far, len / 8 medium-grained feature extraction representations are obtained.
[0061] The functions of each layer of the third-level first Transformer encoder network are as follows:
[0062] First input layer: input is the 256-dimensional fine-grained feature patch block output by the second-level first Transformer encoder network; Patch aggregation layer: concatenate two adjacent blocks of the input fine-grained feature patch block to obtain len / 4 patch blocks; First position encoding layer: position encode each patch block using a trainable position encoding matrix Wpos; First Transformer encoder layer: contains 2 layers of first Transformer encoders, each encoder contains a 4-head attention mechanism and a feedforward neural network. The patch block is mapped to the Transformer input potential space through a trainable linear parameter matrix Wp with a dimension of 256. At this point, len / 4 coarse-grained feature extraction representations are obtained.
[0063] The feature extraction network designed for different modal data in this embodiment uses a hierarchical feature extraction method to obtain multi-granular features of each modal data, effectively improves the feature alignment quality of heterogeneous data, captures rich information of each modal data from details to the whole, and enhances the ability to understand and process data.
[0064] Step 3: Based on the multimodal dataset and with track information as the link, multimodal joint learning is applied to perform hierarchical feature alignment to train the feature extraction module and obtain a trained feature extraction module.
[0065] Specifically, this embodiment uses a contrastive learning method, using track features of the same granularity as a link, so that matching features are closer in the embedding space, while unmatched pairs are farther apart. The specific steps are as follows:
[0066] 31) The multimodal data set is input into the feature extraction module, and each feature extraction network is used to extract features of the data of the corresponding modality, and the corresponding track features, HRRP features and JEM features are obtained, and positive and negative sample pairs of different modalities are constructed.
[0067] Specifically, a batch of data is obtained, including n samples, and the different modal data of each sample are input into the corresponding extraction network, that is, the data of three modalities of each sample (track data, HRRP, JEM), and the track feature extraction network, HRRP feature extraction network and JEM feature extraction network are used to obtain feature extraction representations of three granularities. The HRRP features and JEM features of each granularity are aligned with the track features of the corresponding granularity. For example, for each coarse-grained track feature, the coarse-grained HRRP features and coarse-grained JEM features that match it are constructed as positive sample pairs. in represents the track data characteristics of the i-th sample, represents the broadband data characteristics of the i-th sample, Represent the modal JEM feature of the i-th sample; construct the broadband data and modulation spectrum data features that do not match it as negative sample pairs. Negative sample pairs refer to the representations of different samples in different modes, for example, the track feature of the i-th sample and the HRRP feature data of the j-th sample (i≠j) constitute a negative sample pair.
[0068] 32) Taking the track features as the link, the cosine similarity is used to calculate the similarity between each positive and negative sample pair of different modalities, and the similarity matrix of different modalities is defined.
[0069] First, for all samples in a batch, the similarity between the track features and the HRRP features, as well as the similarity between the track features and the JEM features, is calculated using the following formula:
[0070]
[0071]
[0072] In the formula, represents the cosine similarity between the track feature of the i-th sample and the HRRP feature of the j-th sample, s ij 13 represents the cosine similarity between the track feature of the i-th sample and the JEM feature of the j-th sample, represents the track characteristics of the i-th sample, represents the HRRP feature of the jth sample, represents the JEM feature of the jth sample;
[0073] Then, define that the elements on the diagonal of the similarity matrix are positive sample pairs, and the rest are negative sample pairs, so as to obtain the similarity matrix S of the track features and the HRRP features 12 , and the similarity matrix S of the track features and the JEM features 13 .
[0074] 33) Calculate the cross-entropy loss between the similarity matrices of different modalities to obtain the total loss function
[0075] For the similarity matrix S of the track features and the HRRP features 12 , calculate the cross-entropy loss loss 12 , and the calculation formula is
[0076]
[0077] In the formula, crossentropyloss is the loss calculation of cross-entropy; labels 1 is the label, which is an integer sequence from 0 to the number of batches - 1, indicating the correct matching of each track feature and HRRP feature pair; axis 1 = 0 means aligning the track features with the HRRP features, and axis 1 = 1 means aligning the HRRP features with the track features; by maximizing the similarity of the correct track-HRRP pairs and minimizing the similarity of the incorrect pairs, the contrastive learning realizes learning the association between modal features in an unsupervised manner
[0078] For the similarity matrix S of the track features and the JEM features 13 , calculate the cross-entropy loss loss 13 , and the calculation formula is
[0079]
[0080] In the formula, crossentropyloss is the loss calculation of cross-entropy; labels 2 is the label, which is an integer sequence from 0 to the number of batches - 1, indicating the correct matching of each track feature and JEM feature pair; axis 2 = 0 means aligning the track features with the JEM features, and axis 2 = 1 means aligning the JEM features with the track features; by maximizing the similarity of the correct track-JEM pairs and minimizing the similarity of the incorrect pairs, the contrastive learning realizes learning the association between modal features in an unsupervised manner
[0081] Add the two calculated cross-entropy losses loss 12 and loss 13Perform weighted averaging to obtain the final coarse-grained feature association loss function, expressed as:
[0082]
[0083] Add the associated loss functions of each granularity to obtain the total loss function.
[0084] 34) Based on the total loss function, the back propagation algorithm is used to iteratively update the parameters of each network in the feature extraction module until the training is completed to obtain a trained feature extraction module.
[0085] Compared with traditional feature-level fusion or decision-level fusion, this embodiment uses the track data that is most easily obtained in radar data as a link, adopts the idea of contrastive learning to carry out multimodal data feature alignment, and embeds each modal data into a unified feature space. The feature alignment of multimodal heterogeneous data is completed, which effectively alleviates the semantic deviation between modal data and improves the rationality and effectiveness of fusion; and this embodiment uses the track data that is most easily obtained in radar data as a link for modal feature alignment, which has better feasibility.
[0086] Step 4: Use the trained feature extraction module to process the multimodal dataset to obtain the aligned hierarchical features of each modality.
[0087] Specifically, a multimodal data set including track information, HRRP echo signal and JEM modulation spectrum is input into the trained feature extraction module, and the corresponding track feature extraction network, HRRP feature extraction network and JEM feature extraction network are used to perform hierarchical feature extraction respectively, and the corresponding hierarchical features of each aligned modality are obtained.
[0088] Step 5: Construct a feature fusion module including a feature hierarchical fusion enhancement network, a keyless attention network, an aggregated feature classification network, a single-modal classification network, and a hybrid fusion network.
[0089] Among them, the feature hierarchical fusion enhancement network is used to fuse the hierarchical features of the aligned modes to obtain fused enhanced features; the keyless attention network is used to aggregate the fused enhanced features to obtain aggregated features; the aggregated feature classification network is used to classify the aggregated features to obtain the first classification result; the single-modal classification network is used to process the hierarchical features of the aligned modes separately to obtain the second classification results of each modality; the hybrid fusion network is used to perform weighted fusion of the first classification result and the second classification result to obtain the recognition result.
[0090] The following is an introduction to each network in the feature fusion module.
[0091] 1. Feature Hierarchical Fusion Enhanced Network
[0092] In this embodiment, after the modal feature alignment is completed, the attention mechanism needs to be applied to fuse and enhance the hierarchical features of each modality. Three granular features of three modal data are input, and nine hierarchical fusion and enhancement features are output. In this regard, this embodiment designs a feature hierarchical fusion and enhancement network.
[0093] The feature hierarchical fusion enhancement network constructed in this embodiment includes a second Transformer encoder network;
[0094] The structure of the second Transformer encoder network includes a second input layer, an embedding layer, a second position encoding layer, and a second Transformer encoder layer in sequence; the second Transformer encoder layer includes 6 layers of second Transformer encoders, each of which includes an 8-head attention mechanism and a feedforward neural network;
[0095] Among them, the second input layer is used to input the hierarchical features of each aligned modality;
[0096] The embedding layer is used to map the input features into a high-dimensional space;
[0097] The second position encoding layer is used to perform position encoding on the input features;
[0098] The second Transformer encoder layer is used to fuse and enhance the features after position encoding to obtain fused enhanced features.
[0099] Specifically, the functions and parameters of the feature hierarchical fusion enhancement network can be set as follows:
[0100] The second input layer: The input is three granular features of three modalities (a total of nine features), and the dimension of each feature vector is d. Embedding layer: Map the input modal features to a high-dimensional space, embedding dimension d model Set to 512. Second position encoding layer: Position encoding is performed on each input feature using a trainable position encoding matrix Wpos. The position encoding dimension is the same as the embedding dimension, i.e. d model . Second Transformer Encoder Layer: Contains 6 layers of second Transformer encoders, each of which contains an 8-head attention mechanism and a feedforward neural network. The number of heads of the multi-head self-attention mechanism is set to 8. The hidden layer dimension of the feedforward neural network is set to 2048. The input features are mapped to the Transformer input latent space through a trainable linear parameter matrix with dimension d model. Flatten layer: flattens the output of the Transformer encoder. Fully connected layer: contains 3 fully connected layers, the input dimension is equal to the output dimension of the Flatten layer, and the output dimension is 128. The number of neurons in the first layer is 512, the number of neurons in the second layer is 256, and the number of neurons in the third layer is 128. The activation function uses the ReLU activation function.
[0101] 2. Key-less Attention Network
[0102] After obtaining the fused enhanced features, in order to effectively handle the missing modality, a keyless attention module is designed in this embodiment to aggregate the fused enhanced features. Multiple fused enhanced features are input to the keyless attention network, and one aggregated feature is output accordingly. When there is a missing modality, this network only performs weighted fusion on the existing enhanced features to obtain an aggregated feature for classification.
[0103] As an implementation method, the keyless attention network constructed in this embodiment includes a third input layer, a weight calculation layer, a weighted summation layer and a third output layer; wherein,
[0104] The third input layer is used to input the fusion enhanced features obtained by the feature hierarchical fusion enhancement network;
[0105] The weight calculation layer is used to calculate the weight of each input feature;
[0106] The weighted summation layer is used to perform weighted summation of all calculated weights to obtain the final aggregated features;
[0107] The third output layer is used to output aggregated features.
[0108] Specifically, the functions and parameters of the keyless attention network can be set as follows:
[0109] The weight calculation layer calculates the weight of each input feature vector as follows:
[0110] w i =softmax(W·T i ′);
[0111] In the formula, w i represents the weight of the i-th input feature vector, W is a trainable weight matrix, T i It is the i-th fused and enhanced feature of the input. The subscript represents vector transposition. Here, the row vector is transposed to a column vector for operation.
[0112] The formula used by the weighted summation layer for weighted summation is:
[0113]
[0114] In the formula, Output represents the aggregated features of the final output.
[0115] 3. Aggregate feature classification network
[0116] In this embodiment, the aggregated feature classification network is mainly used for classification based on aggregated features.
[0117] Optionally, the aggregated feature classification network constructed in this embodiment includes a fourth input layer, a first hidden layer, and a fourth output layer; wherein,
[0118] The fourth input layer is used to input the aggregated features obtained by the keyless attention network;
[0119] The first hidden layer includes two fully connected layers, and each fully connected layer is followed by a ReLU activation function; the number of neurons in the first fully connected layer is set to 512, and the number of neurons in the second fully connected layer is set to 256.
[0120] The fourth output layer uses the Softmax activation function to calculate the probability of each category, obtain the first classification result, and output it. The number of neurons in the fourth output layer is equal to the number of target categories.
[0121] 4. Unimodal Classification Network
[0122] In view of the multi-granularity features of each single modality, this embodiment also constructs a corresponding single modality classification network, concatenates the multi-granularity features of each modality, and sends each to the single modality feature classification network to output the classification result.
[0123] Optionally, the single-modal classification network constructed in this embodiment includes three classification networks with the same structure, which are respectively used for classification based on track features, HRRP features and JEM features, and corresponding classification results of each modality are obtained;
[0124] The classification network includes a fifth input layer, a second hidden layer, and a fifth output layer;
[0125] The fifth input layer is used to input the feature representation after the corresponding modality aggregation;
[0126] The second hidden layer includes three fully connected layers, and each fully connected layer is followed by a ReLU activation function; the number of neurons in the first fully connected layer is set to 1024, the number of neurons in the second fully connected layer is set to 512, and the number of neurons in the third fully connected layer is set to 256.
[0127] The fifth output layer uses the Softmax activation function to calculate the probability of each category, obtain the second classification result, and output it. The number of neurons in the fifth output layer is equal to the number of target categories.
[0128] 5. Hybrid Converged Network
[0129] In this embodiment, the hybrid fusion network is based on learnable weight parameters and combines the softmax function to convert the weights into probability distributions, thereby achieving weighted fusion of the first classification result and the second classification result to obtain a recognition result.
[0130] Specifically, this embodiment proposes a weighted fusion method for integrating the output probabilities of multiple single-modal classification networks and aggregated feature classification networks. The method includes the following steps:
[0131] First, the output category probabilities of three unimodal classification networks and the output category probabilities of an aggregated feature classification network are obtained. In order to effectively fuse these probability outputs, a set of learnable weight parameters is introduced in this embodiment. These weight parameters are optimized by back propagation during the model training process to automatically adjust the importance of each network output in the final decision. In order to ensure that the sum of the weights is 1, this embodiment uses the softmax function to normalize these weights. The softmax function converts the weights into probability distributions so that each weight can reasonably reflect the relative importance of its corresponding network output during the fusion process. Through this weighted fusion strategy, the model can dynamically adjust the contribution of each network output, thereby improving the overall classification performance and robustness. This method not only enhances the model's adaptability to multimodal data, but also effectively utilizes the strengths of each network to achieve more accurate classification results.
[0132] This embodiment integrates the individual classification results of each modality with the classification results of the aggregated features while realizing the hierarchical feature fusion of each modality, thereby realizing the hybrid fusion of feature fusion and decision fusion, so that the model can more comprehensively utilize the advantages of multimodal data and improve the accuracy and robustness of recognition. At the same time, feature fusion enhancement and keyless attention network also enable the model to flexibly handle different numbers of modalities, which to some extent alleviates the impact of modality loss on recognition results.
[0133] Step 6: Use the aligned hierarchical features of each mode to train the feature fusion module to obtain a trained feature fusion module, so as to form a multi-modal hierarchical hybrid fusion radar target recognition model together with the trained feature extraction module, thereby realizing multi-modal radar target recognition.
[0134] After constructing the feature fusion module, the aligned features of each modality are input, and the prediction results are obtained through the feature fusion enhancement network, the keyless attention network and the classification network. The cross entropy loss function is used to calculate the loss value between the predicted label and the true category label. The back propagation algorithm is used to iteratively update the parameters of the feature fusion enhancement network, the keyless attention network, the aggregated feature classification network, the single modal classification network and the hybrid fusion network until the training is completed to obtain a trained feature fusion module.
[0135] The trained feature extraction module and feature fusion module are combined to form a multimodal hierarchical hybrid fusion radar target recognition model. The multimodal signal of each target to be identified is input into the model to obtain the recognition result.
[0136] The multimodal hierarchical hybrid fusion radar target recognition method provided by the present invention constructs a feature extraction module including a track feature extraction network, a HRRP feature extraction network, and a JEM feature extraction network, and a feature fusion module including a feature hierarchical fusion enhancement network, a keyless attention network, an aggregated feature classification network, a single-modal classification network, and a hybrid fusion network; in the feature extraction module, the hierarchical features of each modal data are obtained through each network, the feature alignment quality of the heterogeneous data is effectively improved, the rich information of each modal data from details to the whole is captured, and the understanding and processing capabilities of the data are enhanced; at the same time, with the track data that is most easily obtained in the radar data as a link, the data of each modality is embedded into a unified feature space, the feature alignment of the multimodal heterogeneous data is completed, the semantic deviation between the modal data is effectively alleviated, and the method has good generalizability. In the feature fusion module, through feature fusion enhancement and keyless attention network, the model can flexibly handle different numbers of modalities, which alleviates the impact of modality loss on recognition to a certain extent; and while realizing the hierarchical feature fusion of each modality, the individual classification results of each modality are integrated with the classification results of the aggregated features, thereby realizing the hybrid fusion of feature fusion and decision fusion, so that the model can more comprehensively utilize the advantages of multimodal data and improve the accuracy and robustness of recognition. This method clearly takes into full account the impact of modality loss on target recognition in real scenarios, and is more potential than traditional methods, with better generalization performance and better stability.
[0137] The following comparative experiments are conducted on measured data to verify and illustrate the beneficial effects of the present invention.
[0138] 1. Experimental data:
[0139] The data used in this experiment are 17 types of aircraft targets collected by radar in a certain area. The number of collected data for each type is shown in Table 1 below.
[0140] Table 1 Data quantity
[0141] Label Batch number Track Points HRRP quantity JEM quantity 0 3 2718 2148 1843 1 24 6035 2076 0 2 14 5638 1877 653 3 17 7246 1785 175 4 3 1004 0 200 5 12 4623 1370 827 6 16 4847 1454 337 7 7 2621 606 272 8 24 11864 2496 0 9 31 10127 5957 1401 10 1 1577 608 279 11 1 851 0 26 12 9 1748 1156 0 13 55 33820 4404 2517 14 9 3829 1144 494 15 4 1616 325 169 16 93 25973 4191 419 total 323 126137 31635 9713
[0142] According to Table 1, due to the serious imbalance of samples among various targets, some targets have modal missing problems.
[0143] 2. Experimental content and results analysis
[0144] In this experiment, traditional feature fusion, traditional decision fusion, prior art 1, prior art 2, prior art 4 mentioned in the background technology and the multimodal hierarchical hybrid fusion technology of the present invention are selected for target recognition respectively. Since the data does not support the accurate longitude and latitude information required by prior art 3, it is not included in the comparison. The recognition results are shown in Table 2 below.
[0145] Table 2 Comparative experimental results
[0146]
[0147] It can be seen from Table 2 that the average recognition rate of the multi-modal hierarchical hybrid fusion radar target recognition method proposed in the present invention reaches the optimal recognition result of 84.23%.
[0148] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0149] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A multi-modal hierarchical hybrid fusion radar target recognition method, characterized in that: include: Constructing a multimodal data set including multiple categories of targets; the multimodal data set includes track information, HRRP echo signals, and JEM modulation spectra; Constructing a feature extraction module including a track feature extraction network, a HRRP feature extraction network and a JEM feature extraction network; wherein the track feature extraction network, the HRRP feature extraction network and the JEM feature extraction network are used to perform hierarchical feature extraction on the track information, the HRRP echo signal and the JEM modulation spectrum, respectively, to obtain hierarchical features of each mode; Based on the multimodal data set, with the track information as a link, multimodal joint learning is applied to perform hierarchical feature alignment to train the feature extraction module to obtain a trained feature extraction module; Processing the multimodal dataset using the trained feature extraction module to obtain aligned hierarchical features of each modality; Construct a feature fusion module including a feature hierarchical fusion enhancement network, a keyless attention network, an aggregated feature classification network, a single-modal classification network, and a hybrid fusion network; wherein the feature hierarchical fusion enhancement network is used to fuse the hierarchical features of the aligned modes to obtain fused enhanced features; the keyless attention network is used to aggregate the fused enhanced features to obtain aggregated features; the aggregated feature classification network is used to classify the aggregated features to obtain a first classification result; the single-modal classification network is used to process the hierarchical features of the aligned modes separately to obtain a second classification result of each mode; the hybrid fusion network is used to weightedly fuse the first classification result and the second classification result to obtain a recognition result; The feature fusion module is trained using the aligned hierarchical features of each mode to obtain a trained feature fusion module, so as to form a multi-modal hierarchical hybrid fusion radar target recognition model together with the trained feature extraction module, thereby realizing multi-modal radar target recognition.
2. A multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: The track feature extraction network, the HRRP feature extraction network, and the JEM feature extraction network are all multi-layer networks, and the three have similar network structures, and each layer network includes a first Transformer encoder network; The structure of the first Transformer encoder network includes a first input layer, a first position encoding layer, and a first Transformer encoder layer in sequence; the first Transformer encoder layer includes multiple layers of first Transformer encoders, and each of the first Transformer encoders includes a multi-head attention mechanism and a feedforward neural network; Wherein, the first input layer is used to input data of the corresponding modality; The first position encoding layer is used to perform position encoding on the input features; The first Transformer encoder layer is used to perform hierarchical feature extraction on the position-encoded features to obtain corresponding hierarchical features.
3. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: Based on the multimodal data set, with the track information as a link, multimodal joint learning is applied to perform hierarchical feature alignment to train the feature extraction module to obtain a trained feature extraction module, specifically including: The multimodal data set is input into the feature extraction module, and each feature extraction network is used to extract features of the data of the corresponding modality, so as to obtain track features, HRRP features and JEM features, and construct positive and negative sample pairs of different modalities; Using the track features as a link, the cosine similarity is used to calculate the similarity between each positive and negative sample pair of different modalities, and a similarity matrix of different modalities is defined; Calculate the cross entropy loss between the similarity matrices of different modalities to obtain the total loss function; Based on the total loss function, the back propagation algorithm is used to iteratively update the parameters of each network in the feature extraction module until the training is completed to obtain a trained feature extraction module.
4. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 3 is characterized in that: Based on the track features, the cosine similarity is used to calculate the similarity between each positive and negative sample pair of different modalities, and the similarity matrix of different modalities is defined, including: The similarity between the track feature and the HRRP feature, and the similarity between the track feature and the JEM feature are calculated using the following formula: In the formula, represents the cosine similarity between the track feature of the i-th sample and the HRRP feature of the j-th sample, s ij 13 represents the cosine similarity between the track feature of the i-th sample and the JEM feature of the j-th sample, represents the track characteristics of the i-th sample, represents the HRRP feature of the jth sample, represents the JEM feature of the jth sample; Define the elements on the diagonal of the similarity matrix as positive sample pairs, and the rest as negative sample pairs, so as to obtain the similarity matrix S between the track features and the HRRP features 12 , and the similarity matrix S between the track features and the JEM features 13 .
5. A multi-modal hierarchical hybrid fusion radar target recognition method according to claim 4, characterized in that: Calculate the cross entropy loss between the similarity matrices of different modalities to obtain the total loss function, including: For the similarity matrix S between track features and HRRP features 12 , calculate the cross entropy loss loss 12 , the calculation formula is: In the formula, crossentropyloss is the loss calculation of cross entropy; labels1 is the label, which is an integer sequence from 0 to batch number - 1, indicating the correct match of each track feature and HRRP feature pair; axis1 = 0 means aligning the track feature with the HRRP feature, and axis1 = 1 means aligning the HRRP feature with the track feature; For the similarity matrix S between the track features and the JEM features 13 , calculate the cross entropy loss loss 13 , the calculation formula is: In the formula, crossentropyloss is the loss calculation of cross entropy; labels2 is the label, which is an integer sequence from 0 to batch number - 1, indicating the correct match of each track feature and JEM feature pair; axis2 = 0 means aligning the track feature with the JEM feature, and axis2 = 1 means aligning the JEM feature with the track feature; The two cross entropy losses calculated are 12 and loss 13 Perform weighted averaging to obtain the final coarse-grained feature association loss function, expressed as: Add the associated loss functions of each granularity to obtain the total loss function.
6. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: The feature hierarchical fusion enhancement network includes a second Transformer encoder network; The structure of the second Transformer encoder network includes a second input layer, an embedding layer, a second position encoding layer, and a second Transformer encoder layer in sequence; the second Transformer encoder layer includes 6 layers of second Transformer encoders, each of which includes an 8-head attention mechanism and a feedforward neural network; Wherein, the second input layer is used to input the hierarchical features of each aligned modality; The embedding layer is used to map the input features to a high-dimensional space; The second position encoding layer is used to perform position encoding on the input features; The second Transformer encoder layer is used to fuse and enhance the features after position encoding to obtain fused and enhanced features.
7. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: The keyless attention network includes a third input layer, a weight calculation layer, a weighted summation layer and a third output layer; wherein, The third input layer is used to input the fusion enhancement features obtained by the feature hierarchical fusion enhancement network; The weight calculation layer is used to calculate the weight of each input feature; The weighted summation layer is used to perform weighted summation of all calculated weights to obtain the final aggregated features; The third output layer is used to output the aggregated features.
8. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: The aggregate feature classification network includes a fourth input layer, a first hidden layer, and a fourth output layer; wherein, The fourth input layer is used to input the aggregated features obtained by the keyless attention network; The first hidden layer includes two fully connected layers, and each fully connected layer is followed by a ReLU activation function; The fourth output layer uses a Softmax activation function to calculate the probability of each category, obtain the first classification result, and output it.
9. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: The single-modal classification network includes three classification networks with the same structure, which are respectively used for classification based on track features, HRRP features and JEM features, and correspondingly obtain separate classification results for each mode; Wherein, the classification network includes a fifth input layer, a second hidden layer and a fifth output layer; wherein, The fifth input layer is used to input the feature representation after the corresponding modality aggregation; The second hidden layer includes three fully connected layers, and each fully connected layer is connected to a ReLU activation function; The fifth output layer uses a Softmax activation function to calculate the probability of each category, obtain the second classification result, and output it.
10. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that: The hybrid fusion network is based on learnable weight parameters and combines the softmax function to convert the weights into probability distributions, thereby achieving weighted fusion of the first classification result and the second classification result to obtain a recognition result.
Citation Information
Patent Citations
Multi-modal radar active deception jamming identification method based on small samples
CN116047418A
Multi-mode radar HRRP target identification method and system
CN117420552A
Transform-based radar target identification method
CN115047421A
HRRP fusion identification method and device based on CPSA-Conformer
CN118228193A
Intelligent target fusion identification method and system based on radar multi-modal data
CN118885974A
Cited By
HRRP large model multi-scene identification method based on double-domain expert knowledge
CN121388821A
Hrrp large model multi-scene recognition method based on double-domain expert knowledge
CN121388821B
Multi-modal target identification method, system and equipment based on infrared radar composite signal
CN121743842A
Multi-modal target recognition method, system and device based on infrared radar composite signals
CN121743842B