A Radar Target Classification Method Based on Dual-Stream Spatiotemporal Cross-Attention Graph Convolution Network

Through the dual-stream space-time cross-attention graph convolution network fusion of the Doppler spectrum and physical motion parameters, the problems of insufficient spatio-temporal feature capture and noise sensitivity in traditional radar target classification methods are solved, and the accuracy and anti-interference of radar target classification are improved.

CN120145200BActive Publication Date: 2025-07-18NAVAL AVIATION UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510622118.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-18
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Traditional radar target classification methods are difficult to capture space-time joint features, weak graph construction capabilities and high noise sensitivity, resulting in a decrease in classification accuracy, and time-frequency analysis methods have insufficient time-frequency energy diffusion and feature representation capabilities.

Method used

Using a dual-stream space-time cross attention graph convolution network (DS-TGCN) based on the graph structure model, fuses the features of Doppler spectrum data and physical motion parameters, uses the bidirectional cross attention layer and the timing attention layer for feature fusion and aggregation, and combines the graph convolution layer for target classification.

Benefits of technology

It significantly improves the accuracy and anti-interference of radar target classification, realizes deep complementarity and fine-grained alignment of cross-modal features, enhances target identification capabilities, and takes into account both computing efficiency and noise anti-rotability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145200B_ABST
    Figure CN120145200B_ABST
Patent Text Reader

Abstract

The present application discloses a radar target classification method based on a two-stream spatio-temporal cross-attention graph convolutional network, which relates to the cross technical field of radar signal processing and artificial intelligence. The method includes processing the acquired radar echo signal sequence to obtain one or more data frame sequences; for any data frame sequence, performing a first operation, specifically: using each data frame in the target data frame sequence as a graph node to construct a graph structure model and determining the adjacency matrix corresponding to the graph structure model; using the target data frame sequence and the adjacency matrix as inputs, and using the trained two-stream spatio-temporal cross-attention graph convolutional network to determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; determining the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence. The present application can improve the accuracy and anti-interference ability of radar target classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the cross - technical field of radar signal processing and artificial intelligence, and particularly to a radar target classification method based on a dual - stream spatio - temporal cross - attention graph convolutional network. Background Art

[0002] For radar target classification, traditional classification methods face the following problems: First, relying solely on Doppler spectrum data or track data, it is difficult to capture spatio - temporal joint features; second, the graph construction ability is weak, and traditional graph construction methods cannot effectively capture the dynamic spatio - temporal dependence relationships between target nodes; third, the noise sensitivity is high. In a low - signal - to - noise ratio environment, the micro - motion features of targets are easily submerged by noise, and the classification accuracy drops significantly. In addition, in the prior art, time - frequency analysis methods based on the short - time Fourier transform (STFT) and single - stream convolutional networks have defects such as time - frequency energy diffusion and insufficient feature representation ability. Therefore, there is still room for improvement in the accuracy and anti - interference ability of current radar target classification methods. Summary of the Invention

[0003] The purpose of this application is to provide a radar target classification method based on a dual - stream spatio - temporal cross - attention graph convolutional network, which can improve the accuracy and anti - interference ability of radar target classification.

[0004] To achieve the above purpose, this application provides the following solutions:

[0005] In a first aspect, this application provides a radar target classification method based on a dual - stream spatio - temporal cross - attention graph convolutional network, including:

[0006] Process the acquired radar echo signal sequence to obtain one or more data - frame sequences; each data - frame sequence includes multiple data frames, and each data frame includes Doppler spectrum data and physical motion parameters;

[0007] For any one of the data - frame sequences, perform a first operation;

[0008] The first operation is:

[0009] Construct a graph - structure model with each data frame in the target data - frame sequence as a graph node, and determine the adjacency matrix corresponding to the graph - structure model; the target data - frame sequence is any one of the data - frame sequences;

[0010] Using the target data - frame sequence and the adjacency matrix as inputs, and using a trained dual - stream spatio - temporal cross - attention graph convolutional network, determine the probability of each target category monitored by the radar corresponding to the target data - frame sequence;

[0011] Determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data - frame sequence;

[0012] The dual-stream spatio-temporal cross-attention graph convolutional network includes a dual-stream feature extraction layer, a bidirectional cross-attention layer, a temporal attention layer, a graph convolutional layer, and a classification decision layer connected in sequence; wherein, the dual-stream feature extraction layer is used to extract the features of Doppler spectrum data in each data frame and the features of physical motion parameters in each data frame; the bidirectional cross-attention layer is used to fuse the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fusion feature, fuse the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fusion feature, and splice the first fusion feature and the second fusion feature to obtain a spliced feature; the temporal attention layer is used to extract the temporal information of the spliced feature to obtain a temporal feature; the graph convolutional layer is used to aggregate the temporal feature of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix to obtain the aggregated feature of each graph node; the classification decision layer outputs the probability of each target category monitored by the radar according to the aggregated features of all graph nodes.

[0013] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0014] The present application provides a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network, which identifies the target categories monitored by the radar corresponding to the target data frame sequence based on the constructed dual-stream spatio-temporal cross-attention graph convolutional network. Specifically, the bidirectional cross-attention layer is used to fuse the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fusion feature, fuse the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fusion feature, and then splice the first fusion feature and the second fusion feature to obtain a spliced feature. By fusing the features of Doppler spectrum data and physical motion parameters as described above, deep complementarity and fine-grained alignment of cross-modal features are achieved, which can significantly improve the accuracy of radar target classification and the target discrimination ability in a complex electromagnetic environment; then, the temporal attention layer is used to extract the temporal information of the spliced feature to obtain a temporal feature, and the graph convolutional layer is used to aggregate the temporal feature of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix corresponding to the target data frame sequence to obtain the aggregated feature of each graph node, realizing hierarchical modeling of global time dependence and local spatial correlation, further improving the accuracy of radar target classification, and being able to balance computational efficiency and anti-noise robustness, effectively suppressing clutter interference. Description of the Drawings

[0015] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0016] Figure 1 It is an application environment diagram of a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network in an embodiment of the present application;

[0017] Figure 2 It is a schematic flowchart of a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network provided by an embodiment of the present application;

[0018] Figure 3 It is a schematic diagram of the DS-TGCN network structure provided by an embodiment of the present application;

[0019] Figure 4 It is a schematic diagram of generating graph nodes and edges provided by an embodiment of the present application;

[0020] Figure 5 It is a schematic diagram of the Doppler-to-physical motion attention principle provided by an embodiment of the present application;

[0021] Figure 6 It is a schematic diagram of the physical motion-to-Doppler attention principle provided by an embodiment of the present application;

[0022] Figure 7 It is a schematic flowchart of another radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network provided by an embodiment of the present application;

[0023] Figure 8 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0025] To address the problems of difficult multi-modal feature fusion and insufficient spatio-temporal dependence modeling in radar target classification, this application proposes a radar target classification method based on a Dual-Stream Temporal Graph Convolutional Network with Cross-Attention (DS-TGCN). By means of multi-modal feature fusion, a bidirectional cross-attention mechanism, and fixed-edge-weight spatio-temporal graph convolution, it effectively overcomes the deficiencies of traditional methods in terms of classification accuracy and noise resistance, achieving high-precision and high-robustness target classification, and is applicable to the classification and recognition of radar sea and air targets.

[0026] To make the above objects, features, and advantages of this application more apparent and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and specific embodiments.

[0027] The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the radar echo signal to the server 104. After receiving the radar echo signal, the server 104 processes the obtained radar echo signal sequence according to the radar echo signal to obtain one or more data frame sequences; for any data frame sequence, perform a first operation, specifically: use each data frame in the target data frame sequence as a graph node to construct a graph structure model, and determine the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; use the target data frame sequence and the adjacency matrix as inputs, and use the trained dual-stream spatio-temporal cross-attention graph convolutional network to determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence. The server 104 can feedback the target category monitored by the radar corresponding to the obtained target data frame sequence to the terminal 102. In addition, in some embodiments, the radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly process the radar echo signal, or the server 104 can obtain the radar echo signal from the data storage system and process the radar echo signal.

[0028] Among them, the terminal 102 can be but is not limited to various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0029] In an exemplary embodiment, refer to Figure 2 and Figure 3 , a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiment of this application, taking this method applied to Figure 1 the server 104 as an example for illustration, it includes the following steps 101 to 102. Among them:

[0030] Step 101, process the obtained radar echo signal sequence to obtain one or more data frame sequences; each of the data frame sequences includes multiple data frames, and each data frame includes Doppler spectrum data and physical motion parameters.

[0031] Step 102, for any one of the data frame sequences, perform a first operation; the first operation is: using each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; using the target data frame sequence and the adjacency matrix as inputs, and using the trained dual-stream spatio-temporal cross-attention graph convolutional network, determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence.

[0032] The dual-stream spatio-temporal cross-attention graph convolutional network (also known as the DS-TGCN network model, model) includes a dual-stream feature extraction layer, a bidirectional cross-attention layer, a temporal attention layer, a graph convolutional layer, and a classification decision layer connected in sequence; among them, the dual-stream feature extraction layer is used to extract the features of the Doppler spectrum data in each data frame and the features of the physical motion parameters in each data frame; the bidirectional cross-attention layer is used to fuse the features of the physical motion parameters in the features of the Doppler spectrum data to obtain a first fusion feature, fuse the features of the Doppler spectrum data in the features of the physical motion parameters to obtain a second fusion feature, and splice the first fusion feature and the second fusion feature to obtain a spliced feature; the temporal attention layer is used to extract the temporal information of the spliced feature to obtain a temporal feature; the graph convolutional layer is used to aggregate the temporal features of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix to obtain the aggregated feature of each graph node; the classification decision layer outputs the probability of each target category monitored by the radar according to the aggregated features of all graph nodes.

[0033] The data frames in each data frame sequence are arranged in the time order of (radar echo signal) acquisition, or in the time step order. Each data frame corresponds to an acquisition period, that is, the time required for the radar to complete the acquisition of one frame of data, and its size depends on the working parameters and system of the radar system. In this article, each graph node corresponds to a certain data frame in the radar observation sequence. The number of time steps T is equal to the number of nodes N, that is, the number of data frames. For example, if the radar observation sequence contains T consecutive data frames, there are T graph nodes in the constructed graph structure, that is, each graph node corresponds to a data frame.

[0034] In this article, the features corresponding to the Doppler spectrum data are called Doppler spectrum features (Doppler features), and the features corresponding to the physical motion parameters are called physical motion features.

[0035] As an optional implementation manner, step 101 specifically includes:

[0036] Step 101.1, preprocess the obtained radar echo signal sequence to obtain a preprocessed signal sequence; the preprocessing includes pulse compression and range cell truncation.

[0037] Step 101.2, perform separation processing on the preprocessed signal sequence to obtain one or more signal sequences.

[0038] Step 101.3, perform an extraction operation on any one of the signal sequences; the extraction operation is:

[0039] Step 101.3-1: Perform an extraction operation on the target signal sequence to obtain the original data frame sequence corresponding to the target signal sequence; the original data frame sequence includes a plurality of original data frames; each original data frame includes original Doppler spectrum data and original physical motion parameters , where is the batch size; is the number of time steps (also the number of data frames), equal to the number of graph nodes , and also equal to the number of data frames in the data frame sequence; is the dimension of the Doppler spectrum data, which can be adjusted according to the radar system parameters; is the number of physical motion parameters; the target signal sequence is any one of the signal sequences.

[0040] Step 101.3-2: Perform normalization processing on the original data frame sequence corresponding to the target signal sequence to obtain the data frame sequence corresponding to the target signal sequence.

[0041] Step 101.3-2 normalizes the original Doppler spectrum data and the original physical motion parameters to the interval respectively, and the mathematical expression of the normalization operation is as follows:

[0042] ;

[0043] ;

[0044] where represents the normalized Doppler spectrum data corresponding to the target data frame sequence, respectively represent the normalized Doppler spectrum data corresponding to the 1st to the N th data frames in the target data frame sequence; represents the normalized physical motion parameters corresponding to the target data frame sequence, respectively represent the normalized physical motion parameters corresponding to the 1st to the N th data frames in the target data frame sequence; represents the original Doppler spectrum data corresponding to the target data frame sequence, respectively represent the original Doppler spectrum data corresponding to the 1st to the N th data frames in the target data frame sequence; represents the original physical motion parameters corresponding to the target data frame sequence, represents the original physical motion parameters corresponding to the i th to the N th data frames in the target data frame sequence; min() represents selecting the minimum; max() represents selecting the maximum.

[0045] That is to say, after the radar echo signal is pulse-compressed and range-cell intercepted, the target area signal is separated to obtain the signal sequences of various radar targets. Then, the Doppler spectrum data and physical motion parameters (azimuth angle, elevation angle, range, speed, altitude, etc.) of various radar targets (covering targets such as yachts, birds, helicopters, etc.) with equal length are extracted, that is, two data frame sequences of various radar targets are obtained. Then, the two data frame sequences of various radar targets are normalized respectively, so as to identify the categories of various radar targets in turn.

[0046] As an alternative implementation manner, step 102 specifically includes:

[0047] Step 102.1: Taking each data frame in the target data frame sequence as a graph node, using the sliding window strategy to establish the temporal adjacency edges between different graph nodes; wherein, the temporal adjacency edges are established between different graph nodes within the sliding window; the weight of the temporal adjacency edge is 1.

[0048] Step 102.2: Establishing the self-loop edges of each graph node to obtain the graph structure model; the weight of the self-loop edge is 1.

[0049] Step 102.3: Determining the adjacency matrix of the graph structure model according to the graph nodes, the temporal adjacency edges, the self-loop edges, the weight of the temporal adjacency edges and the weight of the self-loop edges.

[0050] That is, taking the target state at each time step as a graph node (abbreviated as node), and the node feature is the concatenation of the Doppler spectrum and the physical motion parameters. The sliding window strategy is used to connect the graph nodes at adjacent time steps to generate the set of temporal adjacency edges. Add self-loop edges to each graph node, and the edge weights are uniformly fixed at 1.0. See Figure 4 , Figure 4 which takes the window size as 2 and the time step as 7 as an example to illustrate the edge generation process.

[0051] The radar observation sequence is modeled as a graph structure through the graph structure model. The construction of the graph structure includes the generation of temporal adjacency edges and the addition of self-loop edges. The graph structure model is further explained below.

[0052] The graph structure model can be expressed as , where is the set of nodes, is the set of edges, is the edge weight matrix. Each node corresponds to a data frame in the data frame sequence i , that is, the feature vector of each node is jointly composed of the Doppler spectrum feature and the physical motion feature:

[0053] ;

[0054] In the formula, is the eigenvector of the i th node (the i th data frame); is the Doppler spectrum feature of the i th data frame; is the physical motion feature of the i th data frame; The superscript symbol " " represents transpose.

[0055] When using the sliding window strategy to connect adjacent time-step graph nodes, assume the size of the sliding window is , then the generated set of temporal adjacency edges can be expressed as ; i and j correspond to graph node and graph node respectively.

[0056] Add self-loop edges to each node to retain the node's own feature information.

[0057] The established symmetric adjacency matrix (abbreviated as the adjacency matrix) is expressed as . Since the weights of both the temporal adjacency edges and the self-loop edges are 1.0, the elements in the adjacency matrix are expressed as .

[0058] As an alternative implementation, the two-stream feature extraction layer includes two feature extraction layers; each of the feature extraction layers includes a one-dimensional convolution and an activation function; among them, one feature extraction layer is used to extract the features of the Doppler spectrum data in each of the

[0059] Specifically, for the Doppler spectrum feature stream, the one-dimensional convolution (1D convolution with 64 channels, kernel size of 3, and padding of 1) in one feature extraction layer is used to extract the time-frequency local features and capture the micro-motion characteristics of the target (such as rotor vibration, flapping period, etc.). The mathematical expression of the one feature extraction layer is as follows:

[0060] ;

[0061] where represents the features of the Doppler spectrum data corresponding to the target data frame sequence, respectively represent the features of the Doppler spectrum data corresponding to the 1st to the N th data frames in the target data frame sequence; represents the one-dimensional convolution, which is a 64-channel one-dimensional convolution with a kernel size of 3 and a padding of 1; is the activation function.

[0062] For the physical motion feature stream, when extracting time-frequency local features through convolution mapping of the same dimension (i.e., another feature extraction layer with the same structure as a feature extraction layer), the macroscopic motion trajectory of the target is modeled. The mathematical expression of the other feature extraction layer is as follows:

[0063] ;

[0064] Among them, represents the feature of the physical motion parameters corresponding to the target data frame sequence, respectively represent the features of the physical motion parameters corresponding to the first to the N th data frames in the data frame sequence.

[0065] As an alternative implementation, the bidirectional cross-attention layer includes a bidirectional cross-attention mechanism and a splicing operation; the bidirectional cross-attention mechanism is used to fuse the features of the physical motion parameters in the features of the Doppler spectrum data to obtain a first fused feature, and to fuse the features of the Doppler spectrum data in the features of the physical motion parameters to obtain a second fused feature; the splicing operation is used to splice the first fused feature and the second fused feature to obtain a spliced feature.

[0066] This application overcomes the information isolation problem of traditional feature splicing through the bidirectional cross-attention mechanism.

[0067] Among them, the operation process of fusing the features of the physical motion parameters in the features of the Doppler spectrum data to obtain the first fused feature (Doppler to physical motion attention) is:

[0068] ;

[0069] ;

[0070] ;

[0071] The operation process of fusing the features of the Doppler spectrum data in the features of the physical motion parameters to obtain the second fused feature (physical to Doppler attention) is:

[0072] ;

[0073] ;

[0074] ;

[0075] Among them, 、 respectively represent the features of the Doppler spectrum data corresponding to the i th data frame in the target data frame sequence, thej The characteristics of the Doppler spectrum data corresponding to a data frame; , respectively represent the characteristics of the physical motion parameters corresponding to the i th data frame in the target data frame sequence, and the characteristics of the physical motion parameters corresponding to the j th data frame; , respectively are the learnable projection matrices corresponding to the i th data frame in the target data frame sequence; , respectively are the learnable projection matrices corresponding to the j th data frame in the target data frame sequence; is the query (Doppler query) corresponding to the Doppler spectrum data in the i th data frame in the target data frame sequence, with a dimension of ; respectively are the key (Doppler key) and value (Doppler value) corresponding to the Doppler spectrum data in the j th data frame in the target data frame sequence, with a dimension of ; is the query (physical query) corresponding to the physical motion parameters in the i th data frame in the target data frame sequence, with a dimension of ; respectively are the key (material key) and value (physical value) corresponding to the physical motion parameters in the j th data frame in the target data frame sequence, with a dimension of ; represents the activation function; is the attention scaling factor, with a value of 64; the superscript symbol " " represents transpose; represents the first fusion feature corresponding to the i th data frame in the target data frame sequence; represents the second fusion feature corresponding to the i th data frame in the target data frame sequence; and are the attention weights (similarity weights), represents and 's attention weight; represents and 's attention weight.

[0076] In the Doppler-to-physical motion attention, the Doppler feature is used as the query (Query), and the physical feature is used as the key-value (Key-Value), so as to enhance the sensitivity to motion mutations (such as sudden turns of drones), such asFigure 5 as shown

[0077] In the physical-to-Doppler attention, the physical features are used as the query, and the Doppler features are used as the key-value, so as to suppress the interference of spectral noise (such as environmental clutter), as Figure 6 shown

[0078] , are the first fusion feature and the second fusion feature corresponding to the target data frame sequence, respectively, and both are also called the bidirectional attention output features

[0079] is a set of learnable projection matrices

[0080] is another set of learnable projection matrices

[0081] All the Doppler queries constitute the global projection matrix of the Doppler queries ; all the Doppler keys constitute the global projection matrix of the Doppler keys ; all the Doppler values constitute the global projection matrix of the Doppler values ; all the physical queries constitute the global projection matrix of the physical queries ; all the physical keys constitute the global projection matrix of the physical keys ; all the physical values constitute the global projection matrix of the physical values .

[0082] Thus, for the attention from Doppler to physical motion:

[0083] ;

[0084] For the attention from physical to Doppler:

[0085] .

[0086] The , generated by the above formula are the global projection results, covering the features of all time steps

[0087] In this article, , , , are collectively referred to as the i th time step slice

[0088] and is a relationship of the whole and the part, representing the projection results of the physical motion characteristics of all time steps, is the local feature in the time dimension j , that is, the projection value of a single time step, i.e., the slice at the j -th time step. For other same parameters, the relationship between the one with subscript and the one without subscript is also a relationship of the whole and the part, such as and , and etc.

[0089] The local slice represents:

[0090] In the attention calculation, the query, key, and value of each time step are obtained through slicing:

[0091] Query slice: (Doppler query at the i -th time step).

[0092] Key slice: (Physical key at the j -th time step).

[0093] Value slice: (Physical value at the j -th time step).

[0094] These local slices are the local expressions of the global projection matrix in the time dimension, i.e.:

[0095] is 's i -th time step slice (dimension: ).

[0096] is 's j -th time step slice (dimension: ).

[0097] Similarly.

[0098] Regarding the attention weights, it is further explained below.

[0099] Taking the Doppler-to-physical motion attention as an example, each query position i (corresponding to time step i ) will calculate the similarity weights with all key positions j (time step j ) 。 As value vectors, they are weighted and summed through weights to finally generate a fused feature related to the query position i . For example, if the physical motion feature at time step j is highly correlated with the Doppler feature at time step i (such as a sharp turn), then is larger, and its contribution to the output is enhanced. Taking as an example, for each time step i , the model dynamically aggregates the physical motion features of all time steps ( j ) through to achieve feature interaction across time steps.

[0100] The dynamic interaction in the bidirectional cross-attention mechanism is further explained below.

[0101] 1) Cross-time-step correlation:

[0102] When calculating the attention weights (such as ), the model calculates the correlation strength between time step and time step through local slicing ( i and j ).

[0103] If the physical key ( j ) at time step is highly correlated with the Doppler query ( i ) at time step , then is larger, indicating a significant correlation between the two features (such as a sharp turn of a drone).

[0104] 2) Feature fusion:

[0105] In the output , the local values of all time steps ( ) are weighted and summed to dynamically fuse local physical motion features and enhance the sensitivity to key time segments (such as motion mutations).

[0106] Finally, the bidirectional attention output features are concatenated (i.e., concatenating the first fused feature and the second fused feature), and the mathematical expression is as follows:

[0107] ;

[0108] where represents the concatenated feature; the symbol represents concatenation along the feature dimension, and the dimension of the concatenated feature is 。

[0109] As an alternative implementation, the temporal attention layer is a multi-head self-attention mechanism.

[0110] After concatenating the bidirectional attention outputs, the multi-head self-attention mechanism is used to capture cross-time-step dependencies and identify key time segments (such as the hovering phase of the drone). The mathematical expression is as follows:

[0111] ;

[0112] Where is the multi-head self-attention mechanism. The multi-head attention (8 heads) splits the input features into subspaces for parallel computation ( is the number of attention heads).

[0113] The specific process and principle are as follows:

[0114] The multi-head attention splits the input features (i.e., the concatenated features) along the channel dimension into subspaces, and the dimension of each attention head is , obtaining 。

[0115] For each subspace of each attention head , independent linear projections are performed to generate queries (Q), keys (K), and values (V):

[0116] ;

[0117] Where is the learnable parameter matrix corresponding to the k th attention head; , respectively represent the queries for the 1st to the k th data frames corresponding to the target data frame sequence in the N th attention head (i.e., time steps 1 to time step N ); , respectively represent the keys for the 1st to the k th data frames corresponding to the target data frame sequence in the N th attention head; , respectively represent the values for the 1st to the k th data frames corresponding to the target data frame sequence in the N th attention head.

[0118] For each attention head, calculate the time step and the time step Association weight (similarity weight):

[0119] ;

[0120] Among them, represents the dependence strength of time step k in the th attention head on time step . If is relatively large, it indicates that the features of time step make a key contribution to the context understanding of time step (such as the stable features in the hovering stage); is the activation function.

[0121] Aggregate the values (V) of all time steps through the attention weights:

[0122] ;

[0123] ;

[0124] Output features: The k th attention head outputs to capture the dependence patterns in a specific subspace.

[0125] Finally, concatenate the outputs of all attention heads and restore the original dimension through linear projection:

[0126] ;

[0127] Among them, represents the temporal features output by the temporal attention layer; represents the concatenation operation; is the output projection matrix, which is a learnable parameter in the model and is randomly generated during model initialization. Its role is to integrate multi-head information and retain the global representation ability of features; is the output features of the 1st to h th attention heads.

[0128] The temporal attention layer realizes cross-time-step dependence modeling through the multi-head self-attention mechanism, which is mainly reflected in two aspects:

[0129] 1) Global interaction: Query, Key, and Value all come from the same input sequence, but the Query at each time step will calculate the similarity weight with the Keys of all time steps. The weight reflects the importance of different time steps. For example, the time steps in the hovering stage may obtain higher weights.

[0130] 2) Multi-head decomposition: Split the input features into multiple subspaces (heads), learn different dependency patterns in parallel, and enhance the model’s sensitivity to different time segments.

[0131] As an optional implementation, the graph convolution layer includes two layers of graph convolution networks connected in sequence; wherein the first layer of graph convolution network (first layer GCN) is used to aggregate the temporal features of each graph node and the temporal features of the adjacent nodes of each graph node according to the adjacency matrix to obtain the original aggregated features of each graph node; the second layer of graph convolution network (second layer GCN) is used to compress the feature dimension of the original aggregated features of each graph node according to the adjacency matrix to obtain the aggregated features of each graph node.

[0132] Specifically, through the first layer of GCN, under the constraints of temporal adjacency edges and self-loop edges, the features of adjacent time steps are aggregated to model the continuity of target motion. The mathematical expression is as follows:

[0133] ;

[0134] in, Represents the output features of the first layer of GCN, that is, the original aggregated features; Represents a graph convolution operation, which does not change the dimension of the input feature. The dimensions of the input feature and the output feature are both 128; Represents the activation function.

[0135] Input Features The dimension is B × T ×128 (B is the batch size, T is the number of time steps, and 128 is the feature dimension). This layer does not compress the feature dimension, but only performs feature transformation.

[0136] The spatiotemporal neighborhood information defined by the adjacency matrix A is aggregated through the first layer of graph convolution (GCNConv), and the output dimension is still 128.

[0137] The above graph convolution process is further explained through an example.

[0138] 1) Assumptions:

[0139] Input features: , .

[0140] Adjacency Matrix: , the window size is 2, self-loop edge + temporal adjacent edge (adjacent edge).

[0141] Weight matrix: .

[0142] 2) Calculation steps:

[0143] Adjacent matrix aggregation: .

[0144] Features corresponding to the aggregated nodes . .

[0145] Features corresponding to the aggregated nodes . .

[0146] Features corresponding to the aggregated nodes . .

[0147] The feature dimension is further compressed through the second-layer GCN to enhance the discriminative representation of the spatial relationship. The mathematical expression is as follows:

[0148] ;

[0149] where, represents the output feature of the second-layer GCN, that is, the aggregated feature; represents the graph convolution operation, which changes the dimension of the input feature. The dimension of the input feature is 128, and the dimension of the output feature is 64; represents the activation function.

[0150] As an alternative implementation, the classification decision layer includes a global average pooling layer, a first fully connected layer, and a second fully connected layer connected in sequence.

[0151] In the global average pooling layer, the features are aggregated along the time step dimension, and global average pooling is performed to retain the overall statistical characteristics of the target motion. The mathematical expression is as follows:

[0152] ;

[0153] where, represents the output feature of the global average pooling layer; represents the feature part corresponding to the time step t (data frame t ) in the output feature of the second-layer GCN.

[0154] For the two fully connected layers, the input feature dimension of the first fully connected layer is 64, and the output feature dimension is 32; the input feature dimension of the second fully connected layer is 32, and the output dimension is C, that is, the probabilities of C target categories are output. The mathematical expressions of the two fully connected layers are as follows:

[0155] ;

[0156] where, , The probabilities for the 1st to the C th target categories respectively; denotes a two-layer fully connected operation; denotes an activation function.

[0157] Regarding the training and validation of the DS-TGCN network model.

[0158] After constructing the DS-TGCN network model, the graph structure (model) dataset for training is divided into a training set and a validation set. Cross-entropy loss, label smoothing, and regularization are adopted, and the optimizer AdamW and early stopping strategy are used; Dropout (probability 0.3) and BatchNorm are used to prevent overfitting. The model is trained and validated. After the DS-TGCN network model completes training and validation, it is loaded for radar target classification of the measured target data. Refer to Figure 7 , Figure 7 is a schematic flowchart of the method including the model training process. Among them, the graph structure model dataset can be established based on historical radar monitoring data, and the construction method of each graph structure model in the dataset is referred to the previous text.

[0159] As an optional implementation, the label smoothed cross-entropy loss is defined as:

[0160] ;

[0161] where, is the one-hot encoding of the true label corresponding to the c th target category, is the model prediction probability corresponding to the c th target category, and 0.1 is the label smoothing coefficient.

[0162] Refer to Figure 7 , the DS-TGCN network model completes training and validation, and then is loaded for radar target classification of the measured target data, specifically as follows:

[0163] 1) Model loading and initialization: Load the parameters of the trained DS-TGCN network model from the storage medium into the inference environment.

[0164] 2) Preprocessing of the measured data.

[0165] 3) Graph data construction: Convert the preprocessed measured data into a graph structure model for model input.

[0166] 4) Model inference: Input the constructed graph data into the DS-TGCN network (model), perform forward propagation, and output the category probability .

[0167] 5) Category determination: Take the index of the maximum value in the probability vector as the final classification result:

[0168] ;

[0169] That is, determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence.

[0170] The output of the last fully connected layer of the DS-TGCN network model is the probability distribution of each target category, not the probability of a certain target category. It is to take the index of the maximum value in the probability vector as the final classification result.

[0171] Example:

[0172] 1) Radar data of 3 target categories are trained through the DS-TGCN network. Let the number of time steps .

[0173] 2) The measured radar echo signal generates graph data after preprocessing and is input into the trained DS-TGCN network model.

[0174] 3) After model inference, the probability vector is output, and the number of categories C = 3, and the corresponding labels are 0, 1, 2).

[0175] 4) Take the index of the maximum probability as the final classification result. Based on the above probability vector, the classification result is Class = 1. If Class = 1 corresponds to the "helicopter" category, then the finally recognized target category is a helicopter.

[0176] The present application also provides an application scenario, which applies the above-mentioned radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network. Specifically: The radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network provided in this embodiment can be applied in a radar target classification scenario. The radar target classification scenario includes a radar echo signal acquisition link and a radar target classification link; the radar echo signal enters the radar target classification link from the radar echo signal acquisition link. The radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network provided in this embodiment belongs to the radar target classification link. Specifically, in the radar target classification link, the obtained radar echo signal sequence is processed to obtain one or more data frame sequences; for any one of the data frame sequences, a first operation is performed, specifically: taking each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; taking the target data frame sequence and the adjacency matrix as inputs, using the trained dual-stream spatio-temporal cross-attention graph convolutional network, determining the probability of each target category monitored by the radar corresponding to the target data frame sequence; and determining the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence.

[0177] Compared with traditional radar target classification methods, the present application shows the following advantages in the radar target classification task:

[0178] The present application uses a dual-stream spatio-temporal cross-attention graph convolutional network to fuse Doppler spectra and physical motion parameters, realizing deep complementarity and fine-grained alignment of cross-modal features, significantly improving the target discrimination ability in complex electromagnetic environments; combining temporal attention and fixed-edge-weight graph convolution to hierarchically model global time dependencies and local spatial correlations, taking into account computational efficiency and anti-noise robustness, and effectively suppressing clutter interference; through an interpretable edge weight matrix and attention weight visualization, intuitively revealing the target motion law and modal interaction mechanism, providing a high-precision, high-efficiency, and high-reliability solution for radar target classification, and adapting to the real-time processing requirements of radar embedded platforms.

[0179] The present application deeply integrates multi-modal collaborative perception, spatio-temporal joint modeling, and lightweight computing architectures, overcomes the core pain points of traditional methods such as low classification accuracy, strong noise sensitivity, and insufficient interpretability in complex environments, provides an innovative solution for the intelligent perception of radar targets, and is applicable to key fields such as military security and airspace monitoring.

[0180] This application extracts target features through a dual-branch structure capable of independently processing Doppler spectra (data) and physical motion parameters, and combines a bidirectional cross-attention mechanism to achieve fine-grained alignment and interaction of cross-modal features. A fixed edge weight strategy based on a sliding window strategy to connect adjacent time-step nodes is introduced to suppress noise interference and significantly improve computational efficiency. Finally, a lightweight classification module outputs highly reliable target categories.

[0181] Experiments show that the DS-TGCN network proposed in this application significantly improves the classification performance of radar targets by fusing multi-modal data and spatio-temporal attention mechanisms, and demonstrates excellent classification performance and real-time performance on the measured radar target dataset.

[0182] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store radar target classification-related data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network.

[0183] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor implements the steps in the above method embodiments.

[0184] In an exemplary embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the steps in the above method embodiments.

[0185] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0186] In this text, specific examples are used to illustrate the principle and implementation of this application. The description of the above embodiments is only for helping to understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation and application scope. To sum up, the content of this specification should not be construed as a limitation to this application.

Claims

1. A radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network, characterized in that Including: Processing the acquired radar echo signal sequence to obtain one or more data frame sequences; each of the data frame sequences includes a plurality of data frames, and each data frame includes Doppler spectrum data and physical motion parameters; Performing a first operation on any one of the data frame sequences; The first operation is: Taking each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; Taking the target data frame sequence and the adjacency matrix as inputs, and using the trained dual-stream spatio-temporal cross-attention graph convolutional network to determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; Determining the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence; The dual-stream spatio-temporal cross-attention graph convolutional network includes a dual-stream feature extraction layer, a bidirectional cross-attention layer, a temporal attention layer, a graph convolutional layer, and a classification decision layer connected in sequence; wherein, the dual-stream feature extraction layer is used to extract the features of the Doppler spectrum data in each data frame, and extract the features of the physical motion parameters in each data frame; the bidirectional cross-attention layer is used to fuse the features of the physical motion parameters in the features of the Doppler spectrum data to obtain a first fusion feature, fuse the features of the Doppler spectrum data in the features of the physical motion parameters to obtain a second fusion feature, and splice the first fusion feature and the second fusion feature to obtain a spliced feature; The temporal attention layer is used to extract the temporal information of the spliced feature to obtain a temporal feature; the graph convolutional layer is used to aggregate the temporal feature of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix to obtain the aggregated feature of each graph node; the classification decision layer outputs the probability of each target category monitored by the radar according to the aggregated features of all graph nodes.

2. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 1, wherein, Taking each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining the adjacency matrix corresponding to the graph structure model specifically includes: Taking each data frame in the target data frame sequence as a graph node, and using a sliding window strategy to establish temporal adjacency edges between different graph nodes; wherein, temporal adjacency edges are established between different graph nodes within the sliding window; the weight of the temporal adjacency edge is 1; Establishing a self-loop edge for each graph node to obtain a graph structure model; the weight of the self-loop edge is 1; Determining the adjacency matrix of the graph structure model according to the graph nodes, the temporal adjacency edges, the self-loop edges, the weight of the temporal adjacency edges, and the weight of the self-loop edges.

3. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 1, characterized in that The dual-stream feature extraction layer includes two feature extraction layers; each feature extraction layer includes a one-dimensional convolution and an activation function; wherein, one feature extraction layer is used to extract the features of the Doppler spectrum data in each data frame; the other feature extraction layer is used to extract the features of the physical motion parameters in each data frame.

4. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 1, wherein The bidirectional cross-attention layer includes a bidirectional cross-attention mechanism and a splicing operation; the bidirectional cross-attention mechanism is used to fuse the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fused feature, and fuse the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fused feature; the splicing operation is used to splice the first fused feature and the second fused feature to obtain a spliced feature.

5. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 4, wherein The operation process of fusing the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fused feature is as follows: ; ; ; The operation process of fusing the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fused feature is as follows: ; ; ; Among them, and respectively represent the characteristics of the Doppler spectrum data corresponding to the i -th data frame in the target data frame sequence, and the characteristics of the Doppler spectrum data corresponding to the j -th data frame; and respectively represent the characteristics of the physical motion parameters corresponding to the i -th data frame in the target data frame sequence, and the characteristics of the physical motion parameters corresponding to the j -th data frame; and are respectively the learnable projection matrices corresponding to the i -th data frame in the target data frame sequence; and are respectively the learnable projection matrices corresponding to the j -th data frame in the target data frame sequence; is the query corresponding to the Doppler spectrum data in the i -th data frame of the target data frame sequence; are respectively the key and value corresponding to the Doppler spectrum data in the j -th data frame of the target data frame sequence; is the query corresponding to the physical motion parameters in the i -th data frame of the target data frame sequence; are respectively the key and value corresponding to the physical motion parameters in the j -th data frame of the target data frame sequence; represents the activation function; is the attention scaling factor; the superscript symbol " " represents transpose; represents and 's attention weight; represents and 's attention weight; represents the first fusion feature corresponding to the i -th data frame in the target data frame sequence; represents the second fusion feature corresponding to the i -th data frame in the target data frame sequence.

6. The radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network according to claim 1, characterized in that The temporal attention layer is a multi-head self-attention mechanism.

7. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 6, characterized in that, Extracting the temporal information of the spliced feature to obtain a temporal feature, specifically including: Splitting the spliced feature along the channel dimension to obtain a subspace corresponding to each attention head; For the subspace corresponding to each attention head, the operation process of each attention head is as follows: ; ; ; ; Among them, represents the subspace corresponding to the k th attention head, h where is the number of attention heads; k are the learnable parameter matrices corresponding to the th attention head respectively; respectively represent the queries for the 1st to the k th data frames corresponding to the target data frame sequence in the N th attention head; and respectively represent the keys for the 1st to the k th data frames corresponding to the target data frame sequence in the N th attention head; and respectively represent the values for the 1st to the k th data frames corresponding to the target data frame sequence in the N th attention head; is the activation function; represents the output feature of the k th attention head. For the output features of all attention heads, the operations of splicing and restoring the dimension are as follows: ; Among them, is the temporal feature output by the temporal attention layer; represents the concatenation operation; are respectively the h output features of the 1st to attention heads; is the output projection matrix.

8. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The graph convolutional layer includes two graph convolutional networks connected in sequence; the first graph convolutional network is used to aggregate the temporal features of each graph node and the temporal features of the adjacent nodes of each graph node according to the adjacency matrix to obtain the original aggregated feature of each graph node; The second graph convolutional network is used to compress the feature dimension of the original aggregated feature of each graph node according to the adjacency matrix to obtain the aggregated feature of each graph node.

9. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 1, characterized in that The classification decision layer includes a global average pooling layer, a first fully connected layer, and a second fully connected layer connected in sequence.

10. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network according to claim 1, characterized in that Processing the obtained radar echo signal sequence to obtain one or more data frame sequences, specifically including: Preprocessing the obtained radar echo signal sequence to obtain a preprocessed signal sequence; the preprocessing includes pulse compression and range cell truncation; Performing separation processing on the preprocessed signal sequence to obtain one or more signal sequences; For any one of the signal sequences, performing an extraction operation; The extraction operation is: Performing an extraction operation on the target signal sequence to obtain the original data frame sequence corresponding to the target signal sequence; the original data frame sequence includes a plurality of original data frames; each original data frame includes original Doppler spectrum data and original physical motion parameters; the target signal sequence is any one of the signal sequences; Performing normalization processing on the original data frame sequence corresponding to the target signal sequence to obtain the data frame sequence corresponding to the target signal sequence.

Citation Information

Patent Citations

  • PD radar target detection method based on graph attention network and transfer learning

    CN114814776A

  • Radar multi-angle multi-feature fusion target classification method

    CN119992228A