A multi-modal hierarchical hybrid fusion radar target recognition method

By constructing a multimodal hierarchical hybrid fusion radar target recognition method, and utilizing hierarchical feature extraction and feature fusion networks, the problems of unstable recognition performance and poor generalization in multimodal data scenarios are solved, and more efficient radar target recognition is achieved.

CN120028767BActive Publication Date: 2025-12-26XIDIAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510156229.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-12-26
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Existing radar target recognition methods suffer from unstable recognition performance and poor generalization in multimodal data scenarios due to semantic mismatch and modal imbalance among different modal data.

Method used

A multimodal hierarchical hybrid fusion radar target recognition method is constructed. Hierarchical feature extraction is performed through track feature extraction network, HRRP feature extraction network and JEM feature extraction network. Hierarchical feature alignment is performed by multimodal joint learning. Feature fusion and decision fusion are performed by combining feature hierarchical fusion enhancement network, keyless attention network, aggregated feature classification network and hybrid fusion network.

Benefits of technology

It effectively improves the feature alignment quality of multimodal data, alleviates semantic bias between modal data, improves the accuracy and robustness of recognition, and has better stability and generalizability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120028767B_ABST
    Figure CN120028767B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal hierarchical hybrid fusion radar target identification methods, comprising: constructing including multiple class targets multi-modal data set, including track information, HRRP echo signal and JEM modulation spectrum;Characteristic extraction module including track feature extraction network, HRRP feature extraction network and JEM feature extraction network is constructed;With track information as link, hierarchical feature alignment is carried out by applying multi-modal joint learning, to train feature extraction module;Characteristic fusion module including feature hierarchical fusion enhancement network, keyless attention network, aggregation feature classification network, single mode classification network and hybrid fusion network is constructed;Characteristic fusion module is trained, to form multi-modal hierarchical hybrid fusion radar target identification model with feature extraction module, to realize multi-modal radar target identification.The method improves the accuracy and robustness of identification, with better generalization performance and stability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of radar target recognition, and particularly relates to a multi-modal hierarchical hybrid fusion radar target recognition method. BACKGROUND

[0002] Radar target recognition is to determine the target type by using the radar echo signal of the target. Compared with the traditional target detection task, the target recognition task needs to measure more deep-level target detail features. In order to achieve this goal, features can be obtained from multiple modalities, including radar high resolution range profile (HRRP), modulation spectrum (JEM) and track information. Among them, the radar high resolution range profile (HRRP) is the projection vector of the target scattering point echo in the radar ray direction obtained by the wideband radar signal, which provides the distribution of the target scattering point along the distance direction and contains rich structural information such as target geometric size. The modulation spectrum (JEM) obtains the frequency spectrum information by analyzing the micro-motion characteristics of the target, extracts the micro-motion features of the target such as rotation and vibration by using the interaction between the radar signal and the target, and provides the distribution of the target in the frequency domain, which contains rich information such as the motion mode and structural characteristics of the target. The track information is the trajectory data obtained by continuously tracking the position change of the target by the radar system, which records the motion path of the target in space and provides dynamic information such as the speed, acceleration and direction of the target. The track information contains the motion mode and behavior characteristics of the target, which has important value for target recognition and prediction tasks. At present, the single recognition method based on radar HRRP or JEM is one of the important ways to realize radar target recognition.

[0003] Most of the traditional target recognition methods are based on single modal data, which may not be robust enough in actual application. Single modal data is easily disturbed and limited in complex environment, resulting in decreased recognition performance. In contrast, the multi-modal fusion method can provide more comprehensive target feature information by combining data from different modalities.

[0004] For multi-modal radar target recognition, the existing technology provides the following methods:

[0005] Prior art 1: In the patent document "Multi-modal radar active deception jamming recognition method based on small sample" (CN202310004984.X) applied by University of Electronic Science and Technology, a time-frequency fusion method is proposed. The method extracts feature parameters in time domain, frequency domain and time-frequency domain, constructs a multi-modal fusion prototype network for feature fusion and classification.

[0006] Prior Art 2: China Electronics Technology Group and Nanjing Electronic Technology Institute proposed a fusion method that comprehensively utilizes the wide and narrow band features of aircraft targets in their published article "Air Target Recognition Based on Radar Wide and Narrow Band Multi-feature Information Fusion" (DOI: 10.16592 / j.cnki.1004-7859.2015.07.005). The article analyzes the characteristics of a wide and narrow band recognition system that uses the power spectrum of HRRP and target speed and height as identification features. It uses DS evidence theory to achieve decision layer fusion target recognition based on wide and narrow band recognition results.

[0007] Prior Art 3: Chongqing University proposed a mathematical model for aircraft behavior recognition using joint data in their published article "Aircraft Behavior Recognition Using Trajectory Data with Multi-modal Method" (DOI: 10.3390 / electronics13020367). The method designs a deep network containing feature abstraction, cross-modal fusion, and classification layers to obtain multi-scale features. It enhances the extracted features using longitude and latitude in the trajectory information, and finally classifies the trajectory information enhanced features.

[0008] Prior Art 4: Lanzhou University of Technology proposed a dual-modal network that fuses time domain data and frequency domain data in their patent application "Multi-modal Radar HRRP Target Recognition Method and System" (CN202311592869.5). The method first performs frequency domain analysis on the original HRRP data to obtain dual-modal data, then uses collaborative attention blocks to fuse the modal data, and finally outputs the recognition result.

[0009] However, the above prior arts still have the following shortcomings:

[0010] The disadvantage of Prior Art 1 is that the method is mainly based on time-frequency analysis of data to enhance information mining of radar data. In the modal fusion process, data alignment is not considered, so there may be problems of semantic mismatch between modal data.

[0011] The disadvantage of Prior Art 2 is that the method uses the recognition results of wide and narrow bands for decision fusion, without considering the semantic deviation and modal imbalance of multi-modal data. In actual application scenarios, the generalization performance of the model is easily affected when modal is missing, and the method relies heavily on prior information of expert knowledge and has poor adaptability.

[0012] The disadvantage of Prior Art 3 is that the method only uses longitude and latitude in the trajectory information to enhance the features of other modal data, which is insufficient in utilizing trajectory information. Moreover, directly fusing trajectory information with other modal features does not consider the distribution difference between modal data, so the recognition performance of the method is unstable and has poor generalizability.

[0013] The disadvantage of the prior art 4 is that the method performs modal expansion by obtaining the frequency domain data of the original HRRP data, which is theoretically deep feature extraction of single modal HRRP data, and does not sufficiently fuse multi-modal radar data such as narrowband data and track data, so the implementability of the method is poor.

[0014] In summary, the existing radar target recognition method has the technical problems of unstable recognition performance and poor generalization in the multi-modal data recognition scene due to the semantic mismatch and modal imbalance of each modal data. SUMMARY

[0015] In order to solve the above problems existing in the prior art, the present application provides a multi-modal hierarchical hybrid fusion radar target recognition method. The technical problems to be solved by the present application are realized by the following technical solutions:

[0016] In the first aspect, the present application provides a multi-modal hierarchical hybrid fusion radar target recognition method, comprising:

[0017] A multi-modal data set including a plurality of categories of targets is constructed; the multi-modal data set includes track information, HRRP echo signals and JEM modulation spectrum;

[0018] A feature extraction module including a track feature extraction network, an HRRP feature extraction network and a JEM feature extraction network is constructed; wherein the track feature extraction network, the HRRP feature extraction network and the JEM feature extraction network are respectively used for hierarchical feature extraction of the track information, the HRRP echo signal and the JEM modulation spectrum, and corresponding hierarchical features of each modal are obtained;

[0019] Based on the multi-modal data set, the hierarchical features are aligned by applying multi-modal joint learning with the track information as the link, so as to train the feature extraction module and obtain the trained feature extraction module;

[0020] The trained feature extraction module is used to process the multi-modal data set to obtain aligned hierarchical features of each modal;

[0021] The feature fusion module includes a feature hierarchical fusion enhancement network, a key-value-free attention network, an aggregated feature classification network, a single-modal classification network and a hybrid fusion network.

[0022] The feature fusion module is trained by using the aligned hierarchical features of each modality, and a trained feature fusion module is obtained to form a multi-modal hierarchical hybrid fusion radar target recognition model together with the trained feature extraction module, so as to realize multi-modal radar target recognition.

[0023] The present application has the following advantages:

[0024] The multi-modal hierarchical hybrid fusion radar target recognition method provided by the present application constructs a feature extraction module including a track feature extraction network, an HRRP feature extraction network and a JEM feature extraction network, and a feature fusion module including a feature hierarchical fusion enhancement network, a key-value-free attention network, an aggregated feature classification network, a single-modal classification network and a hybrid fusion network. In the feature extraction module, the hierarchical features of each modality data are obtained through each network, which effectively improves the feature alignment quality of heterogeneous data, captures rich information of each modality data from details to the whole, and enhances the understanding and processing ability of data. Meanwhile, taking the track data in radar data as a link, each modality data is embedded into a unified feature space, the feature alignment of multi-modal heterogeneous data is completed, the semantic deviation between each modality data is effectively alleviated, and the method has good generalization. In the feature fusion module, the feature fusion enhancement and the key-value-free attention network enable the model to flexibly process different numbers of modalities, to a certain extent, alleviate the influence of modality loss on recognition. While realizing hierarchical feature fusion of each modality, the single classification result of each modality and the classification result of the aggregated feature are integrated, so as to realize hybrid fusion of feature fusion and decision fusion, so that the model can more comprehensively utilize the advantages of multi-modal data, and improve the recognition accuracy and robustness. The method fully considers the influence of modality loss on target recognition in real scenes, has more potential and better generalization performance than traditional methods, and has better stability.

[0025] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 is a flowchart of a multi-modal hierarchical hybrid fusion radar target recognition method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0028] Please refer to Figure 1 , Figure 1 is a flowchart of a multi-modal hierarchical hybrid fusion radar target recognition method provided by an embodiment of the present application. The method mainly includes the following steps:

[0029] Step 1, constructing a multi-modal data set including a plurality of category targets; the multi-modal data set includes track information, HRRP echo signals and JEM modulation spectrum.

[0030] Specifically, the radar HRRP echo signals, JEM modulation spectrum and track information containing N category targets are sorted to form a multi-modal data set. Among them, the multi-modal data is based on the time of the HRRP echo signal, the JEM modulation spectrum is the JEM modulation spectrum with the same batch and the smallest time difference as the current HRRP data; the track information is the historical M multi-dimensional track data points with the same batch and the smallest time difference as the current HRRP data. Each category contains at least 1000 radar echo signals, wherein N≥5 and M≥15.

[0031] Step 2, constructing a feature extraction module including a track feature extraction network, an HRRP feature extraction network and a JEM feature extraction network.

[0032] Among them, the track feature extraction network, the HRRP feature extraction network and the JEM feature extraction network are respectively used for hierarchical feature extraction on the track information, the HRRP echo signal and the JEM modulation spectrum, and the hierarchical features of each modality are obtained correspondingly.

[0033] Specifically, in the present embodiment, the track feature extraction network, the HRRP feature extraction network and the JEM feature extraction network are all multi-level networks, and the three have the same network structure. Each level network includes a first Transformer encoder network;

[0034] The structure of the first Transformer encoder network comprises, in sequence, a first input layer, a first position encoding layer, and a first Transformer encoder layer; the first Transformer encoder layer comprises four first Transformer encoders, each of which comprises four attention mechanisms and a feedforward neural network;

[0035] The first input layer is configured to input data of a corresponding modality.

[0036] The first position encoding layer is configured to perform position encoding on the input features.

[0037] The first Transformer encoder layer is configured to perform hierarchical feature extraction on the position-encoded features, thereby obtaining hierarchical features.

[0038] Optionally, as an implementation manner, the track feature extraction network, the HRRP feature extraction network, and the JEM feature extraction network constructed in the embodiment are all three-level networks, and the three feature extraction networks will be introduced in detail as follows.

[0039] I. Track feature extraction network

[0040] The track feature extraction network constructed in the embodiment extracts track features in a hierarchical manner, from fine granularity to coarse granularity. This design can effectively capture information at different levels and is suitable for processing complex time-series data. The network is a three-level network, and each level network is a Transformer encoder network, which is referred to as a first Transformer encoder network in the embodiment. The structure of each first Transformer encoder network comprises, in sequence, a first input layer, a first position encoding layer, and a first Transformer encoder layer; wherein,

[0041] The functions of the layers of the first level first Transformer encoder network are as follows:

[0042] The first input layer: the input is the track information of a certain type of target, that is, a multi-dimensional track data sequence, and the sequence length is denoted as len; the Patch segmentation layer: the input track information sequence is segmented according to a smaller window size P1=16 and a step size S=16 to form a plurality of Patch blocks; the first position encoding layer: each Patch block is positionally encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: containing 4 layers of first Transformer encoders, each encoder containing 8 attention mechanisms and a feedforward neural network. The Patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 16 256-dimensional fine-grained feature Patch blocks are obtained.

[0043] The functions of the layers of the second-level first Transformer encoder network are as follows:

[0044] The first input layer: the input is the 256-dimensional fine-grained feature Patch block output by the first-level first Transformer encoder network; the Patch aggregation layer: adjacent two of the input fine-grained feature Patch blocks are spliced to obtain len / 8 Patch blocks; the first position encoding layer: each Patch block is positionally encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: containing 4 layers of first Transformer encoders, each encoder containing 8 attention mechanisms and a feedforward neural network. The Patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 8 medium-grained feature extraction representations are obtained.

[0045] The functions of the layers of the third-level first Transformer encoder network are as follows:

[0046] The first input layer: the input is the 256-dimensional fine-grained feature Patch block output by the second-level first Transformer encoder network; the Patch aggregation layer: adjacent two of the input fine-grained feature Patch blocks are spliced to obtain len / 4 Patch blocks; the first position encoding layer: each Patch block is positionally encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: containing 4 layers of first Transformer encoders, each encoder containing 8 attention mechanisms and a feedforward neural network. The Patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 4 coarse-grained feature extraction representations are obtained.

[0047] II. HRRP feature extraction network

[0048] The HRRP feature extraction network built in this embodiment also extracts HRRP features step by step in a hierarchical manner, from fine granularity to coarse granularity. This design can effectively capture information at different levels and is suitable for processing complex sequence data. Like the track feature extraction network, the HRRP feature extraction network is also a three-level network, and each level network is a Transformer encoder, which is referred to as a first Transformer encoder network in this embodiment. Each first Transformer encoder structure is in turn a first input layer, a first position encoding layer, and a first Transformer encoder layer; wherein,

[0049] The functions of each layer of the first level first Transformer encoder network are as follows:

[0050] First input layer: the input is HRRP data of a certain type of target, denoted as HRRP sequence length len; Patch segmentation layer: the input HRRP data is segmented according to a small window size P1 = 16 and a step size S = 16 to form multiple Patch blocks; first position encoding layer: each Patch block is position encoded using a trainable position encoding matrix Wpos; first Transformer encoder layer: contains 4 first Transformer encoders, each encoder contains 8 attention mechanisms and a feedforward neural network. The Patch blocks are mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, with a dimension of 256. Thus, len / 16 256-dimensional fine-grained feature Patch blocks are obtained.

[0051] The functions of each layer of the second level first Transformer encoder network are as follows:

[0052] First input layer: the input is the 256-dimensional fine-grained feature Patch blocks output by the first level first Transformer encoder network; Patch aggregation layer: adjacent two of the input fine-grained feature Patch blocks are spliced to obtain len / 8 Patch blocks; position encoding layer: each Patch block is position encoded using a trainable position encoding matrix Wpos; first Transformer encoder layer: contains 4 first Transformer encoders, each encoder contains 8 attention mechanisms and a feedforward neural network. The Patch blocks are mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, with a dimension of 256. Thus, len / 8 medium-grained feature extraction representations are obtained.

[0053] The functions of the layers of the third-level first Transformer encoder network are as follows:

[0054] The first input layer: the input is the 256-dimensional fine-grained feature patch block output by the second-level first Transformer encoder network; the patch aggregation layer: two adjacent input fine-grained feature patch blocks are spliced to obtain len / 4 patch blocks; the first position encoding layer: each patch block is position encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: contains 4 first Transformer encoders, and each encoder contains 8 attention mechanisms and a feedforward neural network. The patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 4 coarse-grained feature extraction representations are obtained.

[0055] III. JEM feature extraction network

[0056] The JEM feature extraction network built in this embodiment also extracts JEM features in a hierarchical manner from fine-grained to coarse-grained. This design can effectively capture information at different levels and is suitable for processing complex sequence data. Like the track feature extraction network and the HRRP feature extraction network, the JEM feature extraction network is also a three-level network, and each level network is a Transformer encoder, which is referred to as a first Transformer encoder network in this embodiment. The structure of each first Transformer encoder is in turn a first input layer, a first position encoding layer, and a first Transformer encoder layer; wherein,

[0057] The functions of the layers of the first-level first Transformer encoder network are as follows:

[0058] The first input layer: the input is the JEM data of a certain type of target, and the JEM sequence length is denoted as len; the patch segmentation layer: the input JEM data is segmented according to a small window size P1=16 and a step size S=16 to form multiple patch blocks; the first position encoding layer: each patch block is position encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: contains 2 first Transformer encoders, and each encoder contains 4 attention mechanisms and a feedforward neural network. The patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 16 256-dimensional fine-grained feature patch blocks are obtained.

[0059] The functions of the layers of the second-level first Transformer encoder network are as follows:

[0060] The first input layer: the input is the 256-dimensional fine-grained feature patch block output by the first-level first Transformer encoder network; the patch aggregation layer: the adjacent two of the input fine-grained feature patch blocks are spliced to obtain len / 8 patch blocks; the first position encoding layer: each patch block is position encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: contains 2 layers of first Transformer encoders, and each encoder contains 4 heads of attention mechanisms and feedforward neural networks. The patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 8 medium-grained feature extraction representations are obtained.

[0061] The functions of the layers of the third-level first Transformer encoder network are as follows:

[0062] The first input layer: the input is the 256-dimensional fine-grained feature patch block output by the second-level first Transformer encoder network; the patch aggregation layer: the adjacent two of the input fine-grained feature patch blocks are spliced to obtain len / 4 patch blocks; the first position encoding layer: each patch block is position encoded using a trainable position encoding matrix Wpos; the first Transformer encoder layer: contains 2 layers of first Transformer encoders, and each encoder contains 4 heads of attention mechanisms and feedforward neural networks. The patch block is mapped to the Transformer input latent space through a trainable linear parameter matrix Wp, and the dimension is 256. Thus, len / 4 coarse-grained feature extraction representations are obtained.

[0063] The feature extraction network designed for different modal data in this embodiment uses a hierarchical feature extraction method to obtain multi-grained features of each modal data, effectively improves the feature alignment quality of heterogeneous data, captures rich information of each modal data from details to the whole, and enhances the understanding and processing ability of the data.

[0064] Step 3: Based on the multi-modal data set, the track information is used as a link, and multi-modal joint learning is applied for hierarchical feature alignment to train the feature extraction module.

[0065] Specifically, the method of contrast learning is used in this embodiment, and the same-grained track features are used as a link, so that the matched features are closer in the embedding space, and the unmatched pairs are farther apart. The specific steps are as follows:

[0066] 31) The multi-modal data set is input into a feature extraction module, and each feature extraction network is used to extract features from the data of the corresponding mode, to obtain track features, HRRP features and JEM features, and to construct positive and negative sample pairs of different modes.

[0067] Specifically, a batch of data containing n samples is obtained, and the different modal data of each sample is input into the corresponding extraction network, i.e. the data of three modes (track data, HRRP, JEM) of each sample is used to obtain three granularity feature extraction representations by using a track feature extraction network, an HRRP feature extraction network and a JEM feature extraction network respectively. The HRRP feature and the JEM feature of each granularity are aligned with the track feature of the corresponding granularity, for example: for each coarse-grained track feature, the coarse-grained HRRP feature and the coarse-grained JEM feature matched therewith are constructed as a positive sample pair wherein represents the track data feature of the i-th sample, represents the broadband data feature of the i-th sample, represents the modal JEM feature of the i-th sample; and the broadband data and the modulation spectrum data features that do not match the i-th sample are constructed as a negative sample pair. The negative sample pair refers to the representation of different samples in different modes, for example, the track feature of the i-th sample and the HRRP feature data of the j-th sample (i≠j) constitute a negative sample pair.

[0068] 32) Using the track feature as a link, the similarity between each positive and negative sample pair of different modes is calculated using the cosine similarity, and a similarity matrix of different modes is defined.

[0069] First, for all samples in a batch, the similarity between the track feature and the HRRP feature, and the similarity between the track feature and the JEM feature are calculated, and the calculation formula is:

[0070]

[0071]

[0072] In the formula, represents the cosine similarity between the track feature of the i-th sample and the HRRP feature of the j-th sample, s ij 13 represents the cosine similarity between the track feature of the i-th sample and the JEM feature of the j-th sample, represents the track feature of the i-th sample, represents the HRRP feature of the j-th sample, represents the JEM feature of the j-th sample;

[0073] Then, the elements on the diagonal of the similarity matrix are defined as positive sample pairs, and the rest are negative sample pairs, so as to obtain the similarity matrix S of the track feature and the HRRP feature 12 , and the similarity matrix S of the track feature and the JEM feature 13 .

[0074] 33) Calculate the cross-entropy loss between the similarity matrices of different modalities to obtain the total loss function.

[0075] For the similarity matrix S of the track feature and the HRRP feature 12 , the cross-entropy loss loss 12 is calculated, and the calculation formula is:

[0076]

[0077] In the formula, crossentropyloss is the loss calculation of cross-entropy; labels1 is a label, which is an integer sequence from 0 to batch-1, indicating the correct matching of each track feature and HRRP feature pair; axis1=0 indicates aligning the track feature with the HRRP feature, and axis1=1 indicates aligning the HRRP feature with the track feature; by maximizing the similarity of correct track-HRRP pairs and minimizing the similarity of error pairs, the correlation between modal features is learned in an unsupervised manner through contrastive learning.

[0078] For the similarity matrix S of the track feature and the JEM feature 13 , the cross-entropy loss loss 13 is calculated, and the calculation formula is:

[0079]

[0080] In the formula, crossentropyloss is the loss calculation of cross-entropy; labels2 is a label, which is an integer sequence from 0 to batch-1, indicating the correct matching of each track feature and JEM feature pair; axis2=0 indicates aligning the track feature with the JEM feature, and axis2=1 indicates aligning the JEM feature with the track feature; by maximizing the similarity of correct track-JEM pairs and minimizing the similarity of error pairs, the correlation between modal features is learned in an unsupervised manner through contrastive learning.

[0081] The two calculated cross-entropy losses loss 12 and loss 13 are weighted and averaged to obtain the final coarse-grained feature correlation loss function loss, which is expressed as:

[0082]

[0083] The correlation loss functions of each granularity are added to obtain a total loss function.

[0084] 34) Based on the total loss function, the parameters of each network in the feature extraction module are iteratively updated using a back propagation algorithm until the training is completed, and a trained feature extraction module is obtained.

[0085] Compared with traditional feature-level fusion or decision-level fusion, the embodiment takes the most easily obtained track data in radar data as a link, adopts the idea of contrastive learning to carry out multi-modal data feature alignment, and embeds each modal data into a unified feature space. The feature alignment of multi-modal heterogeneous data is completed, effectively alleviating the semantic deviation between modal data, and improving the rationality and effectiveness of fusion; and the embodiment uses the most easily obtained track data in radar data as a link for modal feature alignment, which has better implementability.

[0086] Step 4, processing the multi-modal data set by using the trained feature extraction module to obtain aligned hierarchical features of each modal.

[0087] Specifically, the multi-modal data set including track information, HRRP echo signal and JEM modulation spectrum is input into the trained feature extraction module, and hierarchical feature extraction is carried out by using the corresponding track feature extraction network, HRRP feature extraction network and JEM feature extraction network respectively, to obtain aligned hierarchical features of each modal.

[0088] Step 5, constructing a feature fusion module including a feature hierarchical fusion enhancement network, a keyless value attention network, an aggregated feature classification network, a single modal classification network and a hybrid fusion network.

[0089] The feature hierarchical fusion enhancement network is used for fusion processing of the aligned hierarchical features of each modal to obtain fusion enhancement features; the keyless value attention network is used for aggregation processing of the fusion enhancement features to obtain aggregated features; the aggregated feature classification network is used for classification of the aggregated features to obtain a first classification result; the single modal classification network is used for separate processing of the aligned hierarchical features of each modal to obtain a second classification result of each modal; and the hybrid fusion network is used for weighted fusion of the first classification result and the second classification result to obtain a recognition result.

[0090] The following describes each network in the feature fusion module.

[0091] 1. Feature hierarchical fusion enhancement network

[0092] In the embodiment, when the modal feature alignment is completed, an attention mechanism needs to be applied to fuse and enhance the hierarchical features of each modality. Three granularity features of three modal data are input, and nine hierarchical fusion and enhancement features are output. For this purpose, the embodiment designs a feature hierarchical fusion and enhancement network.

[0093] The feature hierarchical fusion and enhancement network built in the embodiment includes a second Transformer encoder network.

[0094] The structure of the second Transformer encoder network includes a second input layer, an embedding layer, a second position encoding layer, and a second Transformer encoder layer in sequence. The second Transformer encoder layer includes six second Transformer encoders, and each second Transformer encoder includes eight attention mechanisms and a feedforward neural network.

[0095] The second input layer is used to input the aligned hierarchical features of each modality.

[0096] The embedding layer is used to map the input features to a high-dimensional space.

[0097] The second position encoding layer is used to position encode the input features.

[0098] The second Transformer encoder layer is used to fuse and enhance the position-encoded features to obtain the fusion and enhancement features.

[0099] Specifically, the functions and parameters of the feature hierarchical fusion and enhancement network can be set as follows:

[0100] The second input layer: the input is three granularity features of three modalities (a total of nine features), and the dimension of each feature vector is d. The embedding layer: maps the input modal features to a high-dimensional space, and the embedding dimension d model is set to 512. The second position encoding layer: position encodes each input feature using a trainable position encoding matrix Wpos. The position encoding dimension is the same as the embedding dimension, i.e., d model . The second Transformer encoder layer: includes six second Transformer encoders, and each encoder includes eight attention mechanisms and a feedforward neural network. The number of heads of the multi-head self-attention mechanism is set to 8. The hidden layer dimension of the feedforward neural network is set to 2048. The input features are mapped to the Transformer input latent space through a trainable linear parameter matrix, and the dimension is d modelFlatten layer: the output of the Transformer encoder is flattened. Fully connected layer: contains 3 fully connected layers, the input dimension is equal to the output dimension of the Flatten layer, and the output dimension is 128. The first layer has 512 neurons, the second layer has 256 neurons, and the third layer has 128 neurons. The activation function is ReLU activation function.

[0101] 2. Keyless attention network

[0102] After obtaining the fused enhanced features, in order to effectively process the missing modalities, a keyless attention module is designed in this embodiment to aggregate the fused enhanced features. The keyless attention network inputs multiple fused enhanced features and outputs an aggregated feature. When there is a missing modality, the network only weights and fuses the existing enhanced features to obtain an aggregated feature for classification.

[0103] As an implementation manner, the keyless attention network built in this embodiment includes a third input layer, a weight calculation layer, a weighted summation layer, and a third output layer; wherein,

[0104] The third input layer is used to input the fused enhanced features obtained by the feature hierarchical fusion enhanced network.

[0105] The weight calculation layer is used to calculate the weight of each input feature.

[0106] The weighted summation layer is used to weight and sum all the calculated weights to obtain the final aggregated feature.

[0107] The third output layer is used to output the aggregated feature.

[0108] Specifically, the functions and parameters of the keyless attention network can be set as follows:

[0109] The calculation formula of the weight calculation layer for calculating the weight of each input feature vector is:

[0110] w i = softmax(W·T i ′);

[0111] In the formula, w i represents the weight of the i-th input feature vector, W is a trainable weight matrix, T i is the i-th fused enhanced feature input, and the subscript represents vector transposition, which is transposed from a row vector to a column vector for operation.

[0112] The formula for weighted summation of the weighted summation layer is:

[0113]

[0114] In the formula, Output represents the final output of the aggregated feature.

[0115] 3. Aggregated feature classification network

[0116] In the embodiment, the aggregated feature classification network is mainly used for classification based on the aggregated feature.

[0117] Optionally, the aggregated feature classification network built in the embodiment includes a fourth input layer, a first hidden layer, and a fourth output layer; wherein,

[0118] The fourth input layer is used for inputting the aggregated feature obtained by the keyless value attention network.

[0119] The first hidden layer includes two fully connected layers, and a ReLU activation function is connected after each fully connected layer; wherein, the number of neurons of the first fully connected layer is set to 512, and the number of neurons of the second fully connected layer is set to 256.

[0120] The fourth output layer adopts a Softmax activation function, is used for calculating the probability of each category, obtaining a first classification result, and outputting. The number of neurons of the fourth output layer is equal to the number of target categories.

[0121] 4. Single-modal classification network

[0122] For the multi-granularity features of each single modal, the embodiment also correspondingly builds a single-modal classification network for each single modal. The multi-granularity features of each modal are spliced and input into the single-modal feature classification network to output a classification result.

[0123] Optionally, the single-modal classification network built in the embodiment includes three classification networks with the same structure, which are respectively used for classification based on the track feature, the HRRP feature, and the JEM feature, and correspondingly obtain a classification result of each modal.

[0124] The classification network includes a fifth input layer, a second hidden layer, and a fifth output layer; wherein,

[0125] The fifth input layer is used for inputting the feature representation of the corresponding modal after aggregation.

[0126] The second hidden layer includes three fully connected layers, and a ReLU activation function is connected after each fully connected layer; wherein, the number of neurons of the first fully connected layer is set to 1024, the number of neurons of the second fully connected layer is set to 512, and the number of neurons of the third fully connected layer is set to 256.

[0127] The fifth output layer adopts a Softmax activation function, is used for calculating the probability of each category, obtaining a second classification result, and outputting. The number of neurons of the fifth output layer is equal to the number of target categories.

[0128] 5. Hybrid fusion network

[0129] In the embodiment, the hybrid fusion network converts the weights into a probability distribution based on learnable weight parameters and in combination with a softmax function, so as to realize weighted fusion of the first classification result and the second classification result and obtain a recognition result.

[0130] Specifically, the embodiment proposes a weighted fusion method for integrating the output probabilities of multiple single-modal classification networks and an aggregated feature classification network. The method includes the following steps:

[0131] First, the output class probabilities of three single-modal classification networks and the output class probability of an aggregated feature classification network are obtained. In order to effectively fuse these probability outputs, a set of learnable weight parameters is introduced in the embodiment. These weight parameters are optimized through backpropagation during model training, so as to automatically adjust the importance of each network output in the final decision. In order to ensure that the sum of the weights is 1, the softmax function is used to normalize these weights in the embodiment. The softmax function converts the weights into a probability distribution, so that each weight can reasonably reflect the relative importance of the corresponding network output in the fusion process. Through this weighted fusion strategy, the model can dynamically adjust the contribution of each network output, thereby improving the overall classification performance and robustness. This method not only enhances the adaptability of the model to multi-modal data, but also effectively utilizes the strengths of each network to achieve more accurate classification results.

[0132] The embodiment realizes hierarchical feature fusion of each modality, integrates the individual classification results of each modality and the classification results of the aggregated features, thereby realizing hybrid fusion of feature fusion and decision fusion, so that the model can more comprehensively utilize the advantages of multi-modal data, and improve the accuracy and robustness of recognition. At the same time, the feature fusion enhancement and the keyless attention network also enable the model to flexibly handle different numbers of modalities, to some extent, alleviating the impact of modality missing on the recognition result.

[0133] Step 6, training the feature fusion module using the aligned hierarchical features of each modality to obtain a trained feature fusion module, so as to form a multi-modal hierarchical hybrid fusion radar target recognition model together with the trained feature extraction module, thereby realizing multi-modal radar target recognition.

[0134] After the feature fusion module is constructed, the aligned modal features are input, and a prediction result is obtained through the feature fusion enhancement network, the key-value attention-free network and the classification network. A loss value between the predicted label and the real class label is calculated by using a cross-entropy loss function. The parameters of the feature fusion enhancement network, the key-value attention-free network, the aggregated feature classification network, the single-modal classification network and the hybrid fusion network pair are iteratively updated by using a back propagation algorithm until the training is completed, and a trained feature fusion module is obtained.

[0135] The trained feature extraction module and the feature fusion module are used to form a multi-modal hierarchical hybrid fusion radar target recognition model. Each target multi-modal signal to be recognized is input into the model, and a recognition result can be obtained.

[0136] The multi-modal hierarchical hybrid fusion radar target recognition method provided by the application constructs a feature extraction module including a track feature extraction network, an HRRP feature extraction network and a JEM feature extraction network, and a feature fusion module including a feature hierarchical fusion enhancement network, a key-value attention-free network, an aggregated feature classification network, a single-modal classification network and a hybrid fusion network. In the feature extraction module, the hierarchical features of the modal data are obtained through the networks, the feature alignment quality of the heterogeneous data is effectively improved, the rich information of the modal data from details to the whole is captured, and the understanding and processing capability of the data is enhanced. Meanwhile, the track data in the radar data which is the most easily obtained data is used as a link to embed the modal data into a unified feature space, the feature alignment of the multi-modal heterogeneous data is completed, the semantic deviation between the modal data is effectively alleviated, and the method has good generalization. In the feature fusion module, the feature fusion enhancement and the key-value attention-free network enable the model to flexibly process different numbers of modal, and to a certain extent, alleviate the influence of modal loss on recognition. While realizing the hierarchical feature fusion of the modal, the single classification result of each modal and the classification result of the aggregated feature are integrated, thereby realizing the hybrid fusion of feature fusion and decision fusion, enabling the model to more comprehensively utilize the advantages of multi-modal data, and improving the recognition accuracy and robustness. The method fully considers the influence of modal loss on target recognition in the real scene, has more potential and better generalization performance than the traditional method, and has better stability.

[0137] The beneficial effects of the application are verified and illustrated by comparative experiments on actual measurement data.

[0138] 1. Experimental data:

[0139] The data used in this experiment is 17 types of aircraft targets collected by a radar in a certain area, and the number of collected data of each type is shown in Table 1.

[0140] Table 1 Data quantity

[0141] Label Lot Number of track points Number of HRRPs Number of JEMs 0 3 2718 2148 1843 1 24 6035 2076 0 2 14 5638 1877 653 3 17 7246 1785 175 4 3 1004 0 200 5 12 4623 1370 827 6 16 4847 1454 337 7 7 2621 606 272 8 24 11864 2496 0 9 31 10127 5957 1401 10 1 1577 608 279 11 1 851 0 26 12 9 1748 1156 0 13 55 33820 4404 2517 14 9 3829 1144 494 15 4 1616 325 169 16 93 25973 4191 419 Total 323 126137 31635 9713

[0142] According to Table 1, it can be seen that due to the serious sample imbalance of various targets, the existence mode of some targets is missing.

[0143] 2. Experimental content and result analysis

[0144] In this experiment, the traditional feature fusion, the traditional decision fusion, the existing technology 1, the existing technology 2, the existing technology 4 and the multi-modal hierarchical hybrid fusion technology of the present application are selected for target recognition. Since the data does not support the accurate latitude and longitude information required by the existing technology 3, the existing technology 3 is not added for comparison. The recognition results are shown in Table 2.

[0145] Table 2 Comparison of experimental results

[0146]

[0147] From Table 2, it can be seen that the average recognition rate of the multi-modal hierarchical hybrid fusion radar target recognition method proposed in the present application reaches the optimal recognition result of 84.23%.

[0148] In the description of the present application, the terms "first", "second" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0149] The above is a further detailed description of the present application in combination with specific preferred embodiments, and the specific implementation of the present application cannot be limited to these descriptions. For ordinary skilled persons in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or substitutions can be made, which should be considered as belonging to the protection scope of the present application.

Claims

1. A multi-modal hierarchical hybrid fusion radar target recognition method, characterized in that, The method comprises the following steps: constructing a multi-modal data set comprising multiple category targets; the multi-modal data set comprises track information, HRRP echo signals, and JEM modulation spectrum; constructing a feature extraction module comprising a track feature extraction network, an HRRP feature extraction network, and a JEM feature extraction network; wherein the track feature extraction network, the HRRP feature extraction network, and the JEM feature extraction network are respectively used for hierarchical feature extraction on the track information, the HRRP echo signals, and the JEM modulation spectrum, and correspondingly obtain hierarchical features of each modality; based on the multi-modal data set, applying multi-modal joint learning for hierarchical feature alignment with the track information as the link to train the feature extraction module, and obtaining a trained feature extraction module; processing the multi-modal data set by using the trained feature extraction module to obtain aligned hierarchical features of each modality; constructing a feature fusion module comprising a feature hierarchical fusion enhancement network, a keyless value attention network, an aggregated feature classification network, a single modality classification network, and a hybrid fusion network; wherein the feature hierarchical fusion enhancement network is used for fusion processing on the aligned hierarchical features of each modality to obtain fusion enhancement features; the keyless value attention network is used for aggregation processing on the fusion enhancement features to obtain aggregated features; the aggregated feature classification network is used for classifying the aggregated features to obtain a first classification result; the single modality classification network is used for separately processing the aligned hierarchical features of each modality to correspondingly obtain a second classification result of each modality; and the hybrid fusion network is used for weighted fusion on the first classification result and the second classification result to obtain a recognition result; training the feature fusion module by using the aligned hierarchical features of each modality to obtain a trained feature fusion module, so as to form a multi-modal hierarchical hybrid fusion radar target recognition model together with the trained feature extraction module, thereby realizing multi-modal radar target recognition.

2. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, The track feature extraction network, the HRRP feature extraction network, and the JEM feature extraction network are all multi-level networks, and have similar network structures; each level network comprises a first Transformer encoder network; The structure of the first Transformer encoder network comprises a first input layer, a first position encoding layer, and a first Transformer encoder layer in sequence; the first Transformer encoder layer comprises multiple first Transformer encoders, and each first Transformer encoder comprises a multi-head attention mechanism and a feedforward neural network; The first input layer is used for inputting data of a corresponding modality; The first position encoding layer is used for position encoding on the input features; The first Transformer encoder layer is used for hierarchical feature extraction on the position-encoded features, and correspondingly obtains hierarchical features.

3. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, Based on the multi-modal data set, the track information is used as a link, multi-modal joint learning is applied for hierarchical feature alignment, the feature extraction module is trained, and a trained feature extraction module is obtained, specifically including: The multi-modal data set is input into the feature extraction module, and each feature extraction network is used to extract features of data of a corresponding mode, to obtain track features, HRRP features and JEM features, and to construct positive and negative sample pairs of different modes; The track features are used as a link, the cosine similarity is used to calculate the similarity between each positive and negative sample pair of different modes, and a similarity matrix of different modes is defined; The cross-entropy loss between the similarity matrices of different modes is calculated to obtain a total loss function; Based on the total loss function, the parameters of each network in the feature extraction module are iteratively updated using a back propagation algorithm until the training is completed, and a trained feature extraction module is obtained.

4. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 3, characterized in that, The track features are used as a link, the cosine similarity is used to calculate the similarity between each positive and negative sample pair of different modes, and a similarity matrix of different modes is defined, including: The similarity between the track features and the HRRP features, and the similarity between the track features and the JEM features are calculated, and the calculation formula is: wherein, denotes the cosine similarity between the track feature of the i-th sample and the HRRP feature of the j-th sample, s ij 13 denotes the cosine similarity between the track feature of the i-th sample and the JEM feature of the j-th sample, denotes the track feature of the i-th sample, denotes the HRRP feature of the j-th sample, denotes the JEM feature of the j-th sample; The elements on the diagonal of the similarity matrix are defined as positive sample pairs, and the rest are negative sample pairs, so as to obtain the similarity matrix S of the track feature and the HRRP feature 12 , and the similarity matrix S of the track feature and the JEM feature 13 .

5. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 4, characterized in that, The cross-entropy loss between the similarity matrices of different modes is calculated to obtain a total loss function, including: For the similarity matrix S of the track features and the HRRP features 12 , the cross-entropy loss loss is calculated 12 , and the calculation formula is: In the formula, crossentropyloss is the loss calculation of cross-entropy; labels1 is a label, which is an integer sequence from 0 to batch number-1, indicating the correct matching of each track feature and HRRP feature pair; axis1=0 indicates that the track features are aligned with the HRRP features, and axis1=1 indicates that the HRRP features are aligned with the track features; For the similarity matrix S of the track features and the JEM features 13 , the cross-entropy loss loss is calculated 13 , and the calculation formula is: In the formula, crossentropyloss is the loss calculation of cross-entropy; labels2 is a label, which is an integer sequence from 0 to batch number-1, indicating the correct matching of each track feature and JEM feature pair; axis2=0 indicates that the track features are aligned with the JEM features, and axis2=1 indicates that the JEM features are aligned with the track features; The two cross-entropy losses loss 12 and loss 13 calculated are weighted and averaged to obtain the final coarse-grained feature association loss function loss, which is expressed as: The correlation loss functions of each granularity are added to obtain a total loss function.

6. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, The feature hierarchical fusion enhancement network includes a second Transformer encoder network; The structure of the second Transformer encoder network includes a second input layer, an embedding layer, a second position encoding layer and a second Transformer encoder layer in sequence; the second Transformer encoder layer includes 6 second Transformer encoders, and each second Transformer encoder includes 8 attention mechanisms and a feedforward neural network; The second input layer is used to input the aligned hierarchical features of each mode; The embedding layer is used to map the input features to a high-dimensional space; The second position encoding layer is used to position encode the input features; The second Transformer encoder layer is configured to fuse and enhance the position-encoded features to obtain fused and enhanced features.

7. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, The keyless attention network comprises a third input layer, a weight calculation layer, a weighted summation layer, and a third output layer. The third input layer is configured to input the fused and enhanced features obtained by the feature hierarchical fusion and enhancement network. The weight calculation layer is configured to calculate a weight for each input feature. The weighted summation layer is configured to perform weighted summation on all the calculated weights to obtain final aggregated features. The third output layer is configured to output the aggregated features.

8. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, The aggregated feature classification network comprises a fourth input layer, a first hidden layer, and a fourth output layer. The fourth input layer is configured to input the aggregated features obtained by the keyless attention network. The first hidden layer comprises two fully connected layers, and each fully connected layer is followed by a ReLU activation function. The fourth output layer adopts a Softmax activation function, is configured to calculate a probability of each category to obtain a first classification result, and outputs the first classification result.

9. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, The single-modal classification network comprises three classification networks with the same structure, which are respectively configured to perform classification based on the track features, the HRRP features, and the JEM features to obtain separate classification results of each modal. The classification network comprises a fifth input layer, a second hidden layer, and a fifth output layer. The fifth input layer is configured to input the feature representation aggregated for the corresponding modal. The second hidden layer comprises three fully connected layers, and each fully connected layer is followed by a ReLU activation function. The fifth output layer adopts a Softmax activation function, is configured to calculate a probability of each category to obtain a second classification result, and outputs the second classification result.

10. The multi-modal hierarchical hybrid fusion radar target recognition method according to claim 1, characterized in that, The hybrid fusion network combines the first classification result and the second classification result by using the learned weight parameters and a softmax function to convert the weights into a probability distribution, so as to realize weighted fusion of the first classification result and the second classification result and obtain a recognition result.

Citation Information

Patent Citations

  • Multi-modal radar active deception jamming identification method based on small samples

    CN116047418A

  • Multi-mode radar HRRP target identification method and system

    CN117420552A

  • Transform-based radar target identification method

    CN115047421A

  • HRRP fusion identification method and device based on CPSA-Conformer

    CN118228193A