Radar target classification method based on double-flow space-time intersection attention map convolutional network
By adopting a dual-current spatiotemporal cross-attention graph convolution network in radar target classification, integrating the characteristics of Doppler spectrum and physical motion parameters, the problem of insufficient spatiotemporal feature capture and noise immunity in the prior art is solved, and a high precision and high robust target classification is achieved.
Patent Information
- Application Number
- CN202510622118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing radar target classification methods have shortcomings in capturing space-time joint features, graph construction capabilities and noise immunity. Especially in low signal-to-noise environments, the target micro-movement features are easily overwhelmed by noise, and the classification accuracy is significantly reduced.
The method based on the dual-stream space-time cross attention graph convolution network is adopted, and the characteristics of Doppler spectrum data and physical motion parameters are fused through the bidirectional cross attention layer, the timing attention layer extracts timing information, and the graph convolution layer aggregates graph node features to achieve deep complementarity and fine-grained alignment across modal features, and model global time dependence and local spatial correlation through the graph structure model.
It significantly improves the accuracy of radar target classification and target identification capabilities in complex electromagnetic environments, taking into account calculation efficiency and noise resistance robustness, and effectively suppresses clutter interference.
Smart Images

Figure CN120145200A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the cross - technical field of radar signal processing and artificial intelligence, and particularly to a radar target classification method based on a two - stream spatio - temporal cross - attention graph convolutional network. Background Art
[0002] For radar target classification, traditional classification methods face the following problems: First, relying solely on Doppler spectrum data or track data, it is difficult to capture spatio - temporal joint features; second, the graph construction ability is weak, and traditional graph construction methods cannot effectively capture the dynamic spatio - temporal dependence relationships between target nodes; third, the noise sensitivity is high. In a low signal - to - noise ratio environment, the micro - motion features of the target are easily submerged by noise, and the classification accuracy drops significantly. In addition, in the prior art, time - frequency analysis methods based on the short - time Fourier transform (STFT) and single - stream convolutional networks have defects such as time - frequency energy diffusion and insufficient feature representation ability. Therefore, there is still room for improvement in the accuracy and anti - interference ability of current radar target classification methods. Summary of the Invention
[0003] The purpose of this application is to provide a radar target classification method based on a two - stream spatio - temporal cross - attention graph convolutional network, which can improve the accuracy and anti - interference ability of radar target classification.
[0004] To achieve the above - mentioned purpose, this application provides the following solutions: In the first aspect, this application provides a radar target classification method based on a two - stream spatio - temporal cross - attention graph convolutional network, including: Processing the obtained radar echo signal sequence to obtain one or more data - frame sequences; each data - frame sequence includes multiple data frames, and each data frame includes Doppler spectrum data and physical motion parameters; For any data - frame sequence, perform a first operation; The first operation is: Taking each data frame in the target data - frame sequence as a graph node, constructing a graph - structure model, and determining the adjacency matrix corresponding to the graph - structure model; the target data - frame sequence is any one of the data - frame sequences; Taking the target data - frame sequence and the adjacency matrix as inputs, and using the trained two - stream spatio - temporal cross - attention graph convolutional network to determine the probability of each target category monitored by the radar corresponding to the target data - frame sequence; Determining the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data - frame sequence; The dual-stream spatio-temporal cross-attention graph convolutional network includes a dual-stream feature extraction layer, a bidirectional cross-attention layer, a temporal attention layer, a graph convolutional layer, and a classification decision layer connected in sequence; wherein, the dual-stream feature extraction layer is used to extract the features of Doppler spectrum data in each data frame and the features of physical motion parameters in each data frame; the bidirectional cross-attention layer is used to fuse the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fused feature, fuse the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fused feature, and splice the first fused feature and the second fused feature to obtain a spliced feature; the temporal attention layer is used to extract the temporal information of the spliced feature to obtain a temporal feature; the graph convolutional layer is used to aggregate the temporal feature of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix to obtain the aggregated feature of each graph node; the classification decision layer outputs the probability of each target category monitored by the radar according to the aggregated features of all graph nodes.
[0005] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application provides a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network, which identifies the target categories monitored by the radar corresponding to the target data frame sequence based on the constructed dual-stream spatio-temporal cross-attention graph convolutional network. Specifically, the bidirectional cross-attention layer fuses the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fused feature, and fuses the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fused feature, and then splices the first fused feature and the second fused feature to obtain a spliced feature. By fusing the features of Doppler spectrum data and physical motion parameters as described above, deep complementarity and fine-grained alignment of cross-modal features are achieved, which can significantly improve the accuracy of radar target classification and the target discrimination ability in complex electromagnetic environments; then, the temporal attention layer extracts the temporal information of the spliced feature to obtain a temporal feature, and the graph convolutional layer aggregates the temporal feature of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix corresponding to the target data frame sequence to obtain the aggregated feature of each graph node, realizing hierarchical modeling of global time dependence and local spatial correlation, further improving the accuracy of radar target classification, and being able to balance computational efficiency and anti-noise robustness, effectively suppressing clutter interference. Description of the Drawings
[0006] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0007] Figure 1 It is an application environment diagram of a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network in an embodiment of the present application; Figure 2 It is a schematic flowchart of a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the DS-TGCN network structure provided by an embodiment of the present application; Figure 4 It is a schematic diagram of generating graph node edges provided by an embodiment of the present application; Figure 5 It is a schematic diagram of the Doppler-to-physical motion attention principle provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the physical motion-to-Doppler attention principle provided by an embodiment of the present application; Figure 7 It is a schematic flowchart of another radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network provided by an embodiment of the present application; Figure 8 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0008] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0009] Aiming at the problems of difficult multi-modal feature fusion and insufficient spatio-temporal dependence modeling in radar target classification, this application proposes a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network (Dual-Stream Temporal Graph Convolutional Network with Cross-Attention, DS-TGCN). Through multi-modal feature fusion, bidirectional cross-attention mechanism and fixed-edge-weight spatio-temporal graph convolution, it effectively overcomes the deficiencies of traditional methods in classification accuracy and noise resistance, realizes high-precision and high-robustness target classification, and is applicable to the classification and recognition of radar sea and air targets.
[0010] To make the above objects, features and advantages of this application more obvious and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and specific embodiments.
[0011] The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the radar echo signal to the server 104. After receiving the radar echo signal, the server 104 processes the obtained radar echo signal sequence according to the radar echo signal to obtain one or more data frame sequences; for any data frame sequence, perform a first operation, specifically: using each data frame in the target data frame sequence as a graph node, construct a graph structure model, and determine the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; using the target data frame sequence and the adjacency matrix as inputs, and using the trained dual-stream spatio-temporal cross-attention graph convolutional network, determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence. The server 104 can feedback the target category monitored by the radar corresponding to the obtained target data frame sequence to the terminal 102. In addition, in some embodiments, the radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly process the radar echo signal, or the server 104 can obtain the radar echo signal from the data storage system and process the radar echo signal.
[0012] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0013] In an exemplary embodiment, refer to Figure 2 and Figure 3 , a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in as an example for illustration, it includes the following steps 101 to step 102. Among them: Step 101, process the acquired radar echo signal sequence to obtain one or more data frame sequences; each of the data frame sequences includes multiple data frames, and each data frame includes Doppler spectrum data and physical motion parameters.
[0014] Step 102, for any one of the data frame sequences, perform a first operation; the first operation is: taking each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; taking the target data frame sequence and the adjacency matrix as inputs, using the trained dual-stream spatio-temporal cross-attention graph convolutional network, determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence.
[0015] The dual-stream spatio-temporal cross-attention graph convolutional network (also known as the DS-TGCN network model, model) includes a dual-stream feature extraction layer, a bidirectional cross-attention layer, a temporal attention layer, a graph convolutional layer, and a classification decision layer connected in sequence; among them, the dual-stream feature extraction layer is used to extract the features of Doppler spectrum data in each data frame, and extract the features of physical motion parameters in each data frame; the bidirectional cross-attention layer is used to fuse the features of physical motion parameters in the features of Doppler spectrum data to obtain a first fusion feature, fuse the features of Doppler spectrum data in the features of physical motion parameters to obtain a second fusion feature, and splice the first fusion feature and the second fusion feature to obtain a spliced feature; the temporal attention layer is used to extract the temporal information of the spliced feature to obtain a temporal feature; the graph convolutional layer is used to aggregate the temporal features of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix to obtain the aggregated feature of each graph node; the classification decision layer outputs the probability of each target category monitored by the radar according to the aggregated features of all graph nodes.
[0016] The data frames in each data frame sequence are arranged in the order of acquisition time (radar echo signal), or in the order of time steps. Each data frame corresponds to an acquisition period, that is, the time required for the radar to complete the acquisition of one frame of data, and its size depends on the working parameters and system of the radar system. In this article, each graph node corresponds to a certain data frame in the radar observation sequence. The number of time steps T is equal to the number of nodes N, that is, the number of data frames. For example, if the radar observation sequence contains T consecutive data frames, then there are T graph nodes in the constructed graph structure, that is, each graph node corresponds to a data frame.
[0017] In this article, the features corresponding to Doppler spectrum data are called Doppler spectrum features (Doppler features), and the features corresponding to physical motion parameters are called physical motion features.
[0018] As an optional implementation manner, step 101 specifically includes: Step 101.1, preprocess the obtained radar echo signal sequence to obtain a preprocessed signal sequence; the preprocessing includes pulse compression and range cell truncation.
[0019] Step 101.2, perform separation processing on the preprocessed signal sequence to obtain one or more signal sequences.
[0020] Step 101.3, for any one of the signal sequences, perform an extraction operation; the extraction operation is: Step 101.3-1: Extract the target signal sequence to obtain the original data frame sequence corresponding to the target signal sequence; the original data frame sequence includes multiple original data frames; each original data frame includes original Doppler spectrum data and original physical motion parameters , where is the batch size; is the number of time steps (also the number of data frames), equal to the number of graph nodes , and also equal to the number of data frames in the data frame sequence; is the dimension of the Doppler spectrum data, which can be adjusted according to the radar system parameters; is the number of physical motion parameters; the target signal sequence is any one of the signal sequences.
[0021] Step 101.3-2: Normalize the original data frame sequence corresponding to the target signal sequence to obtain the data frame sequence corresponding to the target signal sequence.
[0022] Step 101.3-2 normalizes the original Doppler spectrum data and the original physical motion parameters to the interval respectively, and the mathematical expression of the normalization operation is as follows: ; ; where represents the normalized Doppler spectrum data corresponding to the target data frame sequence, respectively represent the normalized Doppler spectrum data corresponding to the 1st to the N th data frames in the target data frame sequence; represents the normalized physical motion parameters corresponding to the target data frame sequence, respectively represent the normalized physical motion parameters corresponding to the 1st to the N th data frames in the target data frame sequence; represents the original Doppler spectrum data corresponding to the target data frame sequence, respectively represent the original Doppler spectrum data corresponding to the 1st to the N th data frames in the target data frame sequence; represents the original physical motion parameters corresponding to the target data frame sequence, represents the original physical motion parameters corresponding to the i th to the N th data frames in the target data frame sequence; min() represents selecting the minimum; max() represents selecting the maximum.
[0023] That is to say, after the radar echo signal is pulse-compressed and the range cell is intercepted, the signal in the target area is separated to obtain the signal sequences of various radar targets. Then, the Doppler spectrum data and physical motion parameters (azimuth angle, elevation angle, range, speed, altitude, etc.) of the same length of various radar targets (covering targets such as yachts, birds, and helicopters) are extracted, that is, two data frame sequences of various radar targets are obtained. Then, the two data frame sequences of various radar targets are normalized respectively, so as to identify the categories of various radar targets in turn.
[0024] As an alternative implementation manner, step 102 specifically includes: Step 102.1: Taking each data frame in the target data frame sequence as a graph node, using the sliding window strategy to establish the temporal adjacency edges between different graph nodes; wherein, the temporal adjacency edges are established between different graph nodes within the sliding window; the weight of the temporal adjacency edge is 1.
[0025] Step 102.2: Establishing the self-loop edges of each graph node to obtain the graph structure model; the weight of the self-loop edge is 1.
[0026] Step 102.3: Determining the adjacency matrix of the graph structure model according to the graph nodes, the temporal adjacency edges, the self-loop edges, the weight of the temporal adjacency edge, and the weight of the self-loop edge.
[0027] That is, taking the target state at each time step as a graph node (abbreviation: node), the node feature is the concatenation of the Doppler spectrum and the physical motion parameters. The sliding window strategy is used to connect the graph nodes at adjacent time steps to generate the set of temporal adjacency edges. Add self-loop edges to each graph node, and the edge weights are uniformly fixed at 1.0. See Figure 4 , Figure 4 Taking the window size as 2 and the time step as 7 as an example to illustrate the generation process of the edges.
[0028] Model the radar observation sequence as a graph structure through the graph structure model. The construction of the graph structure includes the generation of temporal adjacency edges and the addition of self-loop edges. The graph structure model is further explained below.
[0029] The graph structure model can be expressed as ,wherein, is the set of nodes, is the set of edges, is the edge weight matrix. Each node corresponds to a data frame in the data frame sequence i ,that is, the feature vector of each node is jointly composed of the Doppler spectrum feature and the physical motion feature: ; In the formula, is the iThe feature vector of i data frames of a node; is the Doppler spectrum feature of the i th data frame; is the physical motion feature of the i th data frame; The superscript symbol " " represents transpose.
[0030] When connecting adjacent time-step graph nodes using the sliding window strategy, assume the sliding window size is , then the generated set of temporal adjacency edges can be expressed as ; i and j correspond to graph nodes and graph node respectively.
[0031] Add self-loop edges to each node to retain the node's own feature information.
[0032] The established symmetric adjacency matrix (abbreviated as the adjacency matrix) is expressed as . Since the weights of both the temporal adjacency edges and the self-loop edges are 1.0, the elements in the adjacency matrix are expressed as .
[0033] As an alternative implementation, the two-stream feature extraction layer includes two feature extraction layers; each of the feature extraction layers includes a one-dimensional convolution and an activation function; among them, one feature extraction layer is used to extract the features of the Doppler spectrum data in each of the
[0034] data frames; the other feature extraction layer is used to extract the features of the physical motion parameters in each of the Specifically, for the Doppler spectrum feature stream, the time-frequency local features are extracted through a one-dimensional convolution (1D convolution with 64 channels, kernel size of 3, and padding of 1) in a feature extraction layer to capture the micro-motion characteristics of the target (such as rotor vibration, flapping period, etc.). The mathematical expression of the one feature extraction layer is as follows: where represents the feature of the Doppler spectrum data corresponding to the target data frame sequence, respectively represent the features of the Doppler spectrum data corresponding to the 1st to the N th data frames in the target data frame sequence; represents the one-dimensional convolution, which is a 64-channel convolution with a kernel size of 3 and a padding of 1; is the activation function.
[0035] For the physical motion feature stream, when extracting time-frequency local features through a convolutional mapping of the same dimension (i.e., another feature extraction layer with the same structure as a feature extraction layer), the macroscopic motion trajectory of the target is modeled. The mathematical expression of the other feature extraction layer is as follows: ; Wherein, represents the feature of the physical motion parameters corresponding to the target data frame sequence, respectively represent the features of the physical motion parameters corresponding to the 1st to the N th data frames in the data frame sequence.
[0036] As an optional implementation manner, the bidirectional cross-attention layer includes a bidirectional cross-attention mechanism and a splicing operation; the bidirectional cross-attention mechanism is used to fuse the features of the physical motion parameters in the features of the Doppler spectrum data to obtain a first fusion feature, and fuse the features of the Doppler spectrum data in the features of the physical motion parameters to obtain a second fusion feature; the splicing operation is used to splice the first fusion feature and the second fusion feature to obtain a spliced feature.
[0037] This application overcomes the information isolation problem of traditional feature splicing through the bidirectional cross-attention mechanism.
[0038] Wherein, the operation process of fusing the features of the physical motion parameters in the features of the Doppler spectrum data to obtain the first fusion feature (Doppler-to-physical motion attention) is: ; ; ; The operation process of fusing the features of the Doppler spectrum data in the features of the physical motion parameters to obtain the second fusion feature (physical-to-Doppler attention) is: ; ; ; Wherein, , respectively represent the features of the Doppler spectrum data corresponding to the i th data frame in the target data frame sequence, the features of the Doppler spectrum data corresponding to the j th data frame; , respectively represent the features of the physical motion parameters corresponding to the i th data frame in the target data frame sequence, the features of the physical motion parameters corresponding to the j th data frame; , are respectively the learnable projection matrices corresponding to the i th data frame in the target data frame sequence; and are respectively the learnable projection matrices corresponding to the j th data frame in the target data frame sequence; is the query (Doppler query) corresponding to the Doppler spectrum data in the i th data frame in the target data frame sequence, with a dimension of ; are respectively the key (Doppler key) and value (Doppler value) corresponding to the Doppler spectrum data in the j th data frame in the target data frame sequence, with a dimension of ; is the query (physical query) corresponding to the physical motion parameters in the i th data frame in the target data frame sequence, with a dimension of ; are respectively the key (material key) and value (physical value) corresponding to the physical motion parameters in the j th data frame in the target data frame sequence, with a dimension of ; represents an activation function; is the attention scaling factor, with a value of 64; the superscript symbol " " represents transpose; represents the first fusion feature corresponding to the i th data frame in the target data frame sequence; represents the second fusion feature corresponding to the i th data frame in the target data frame sequence; and are attention weights (similarity weights), represents and 's attention weight; represents and 's attention weight.
[0039] In the Doppler-to-physical motion attention, the Doppler feature is used as the query (Query), and the physical feature is used as the key-value (Key-Value), so as to enhance the sensitivity to motion mutations (such as sudden turns of drones), as Figure 5 shown.
[0040] In the physical-to-Doppler attention, the physical feature is used as the query (Query), and the Doppler feature is used as the key-value (Key-Value), so as to suppress the interference of spectral noise (such as environmental clutter), as Figure 6 shown.
[0041] , are the first fusion feature and the second fusion feature corresponding to the target data frame sequence, and both are also called bidirectional attention output features.
[0042] is a set of learnable projection matrices.
[0043] is another set of learnable projection matrices.
[0044] All Doppler queries constitute the global projection matrix of Doppler queries ; all Doppler keys constitute the global projection matrix of Doppler keys ; all Doppler values constitute the global projection matrix of Doppler values ; all physical queries constitute the global projection matrix of physical queries ; all physical keys constitute the global projection matrix of physical keys ; all physical values constitute the global projection matrix of physical values .
[0045] Thus, for the attention from Doppler to physical motion: ; for the attention from physical to Doppler: .
[0046] The , generated by the above formula are global projection results, covering the features of all time steps.
[0047] In this article, , , , are collectively referred to as the i th time step slice.
[0048] and are in the relationship of the whole and the part. represents the projection result of the physical motion features of all time steps. is in the time dimension j the local feature, that is, the projection value of a single time step, that is, at the jSlices for each time step. For other parameters with the same name, the relationship between the subscripted and non-subscripted ones is also the relationship between the whole and the part, such as and 、 and etc.
[0049] Local slice representation: In the attention calculation, the query, key, and value for each time step are obtained through slicing: Query slice: (Doppler query for the i th time step).
[0050] Key slice: (Physical key for the j th time step).
[0051] Value slice: (Physical value for the j th time step).
[0052] These local slices are the local expressions of the global projection matrix in the time dimension, i.e.: is the th time step slice of i (Dimension: ).
[0053] is the th time step slice of j (Dimension: ).
[0054] Similarly.
[0055] Regarding the attention weights, further explanations are given below.
[0056] Taking the Doppler-to-physical motion attention as an example, each query position i (corresponding to time step i ) calculates the similarity weight j with all key positions j (time step ). As the value vector, it is weighted and summed through the weight to finally generate the fused feature related to the query position i . For example, if the physical motion feature at time step j is highly correlated with the Doppler feature at time step i (such as a sharp turn), then is larger, and 's contribution to the output is enhanced. Taking For example, for each time step i , the model passes through dynamically aggregates the physical motion features of all time steps j ), realizing feature interaction across time steps. ).
[0057] Regarding the dynamic interaction in the bidirectional cross-attention mechanism, it will be further explained below.
[0058] 1) Cross-time step correlation: When calculating the attention weights (such as ), the model calculates the correlation strength between time step and through local slicing ( i and time step j ).
[0059] If the physical key of time step j ( ) is highly correlated with the Doppler query of time step i ( ), then is larger, indicating a significant correlation between the two features (such as a sharp turn of a drone).
[0060] 2) Feature fusion: In the output , by weighted summing the local values of all time steps ( ), the local physical motion features are dynamically fused to enhance the sensitivity to key time segments (such as motion mutations).
[0061] Finally, the bidirectional attention output features are concatenated (i.e., concatenating the first fused feature and the second fused feature), and the mathematical expression is as follows: ; where represents the concatenated feature; the symbol represents concatenation along the feature dimension, and the dimension of the concatenated feature is .
[0062] As an optional implementation manner, the temporal attention layer is a multi-head self-attention mechanism.
[0063] After concatenating the bidirectional attention outputs, the multi-head self-attention mechanism is used to capture cross-time step dependencies and identify key time segments (such as the hovering phase of a drone), and the mathematical expression is as follows: ; where is the multi-head self-attention mechanism, and the multi-head attention (8 heads) splits the input feature into Subspace parallel computing ( is the number of attention heads).
[0064] The specific process and principle are as follows: The multi-head attention splits the input features (i.e., the concatenated features) along the channel dimension into subspaces, and the dimension of each attention head is , obtaining .
[0065] Perform independent linear projections on the subspaces of each attention head to generate queries (Q), keys (K), and values (V): ; Among them, is the learnable parameter matrix corresponding to the k th attention head; , respectively represent the queries of the 1st to the k rd data frames (i.e., time steps 1 to time step N ) in the corresponding target data frame sequence of the N th attention head; , respectively represent the keys of the 1st to the k rd data frames in the corresponding target data frame sequence of the N th attention head; , respectively represent the values of the 1st to the k rd data frames in the corresponding target data frame sequence of the N th attention head.
[0066] For each attention head, calculate the association weight (similarity weight) between time step and time step : ; Among them, represents the dependence intensity of time step k on time step in the th attention head. If is large, it indicates that the features at time step make a key contribution to the context understanding at time step (such as the stable features in the hovering phase); is the activation function.
[0067] Aggregate the values (V) of all time steps through the attention weights: ; ; Output feature: The output of the k th attention head , to capture the dependency patterns in specific subspaces.
[0068] Finally, the outputs of all attention heads are concatenated and the original dimension is restored through linear projection: ; where represents the temporal features output by the temporal attention layer; represents the concatenation operation; is the output projection matrix, a learnable parameter in the model, randomly generated during model initialization, whose role is to integrate multi-head information and retain the global representational ability of features; is the output feature of the 1st to h th attention heads.
[0069] The temporal attention layer realizes cross-time-step dependency modeling through the multi-head self-attention mechanism, which is mainly reflected in two aspects: 1) Global interaction: The query, key, and value all come from the same input sequence, but the query at each time step calculates the similarity weights with the keys at all time steps. The weights reflect the importance of different time steps. For example, the time steps in the hovering phase may obtain higher weights.
[0070] 2) Multi-head decomposition: The input features are split into multiple subspaces (heads), and different dependency patterns are learned in parallel, enhancing the model's sensitivity to different time segments.
[0071] As an optional implementation manner, the graph convolutional layer includes two graph convolutional networks connected in sequence; wherein the first graph convolutional network (the first GCN) is used to aggregate the temporal features of each graph node and the temporal features of the adjacent nodes of each graph node according to the adjacency matrix to obtain the original aggregated features of each graph node; the second graph convolutional network (the second GCN) is used to compress the feature dimension of the original aggregated features of each graph node according to the adjacency matrix to obtain the aggregated features of each graph node.
[0072] Specifically, through the first GCN, under the constraints of temporal adjacent edges and self-loops, the features of adjacent time steps are aggregated to model the continuity of the target motion. The mathematical expression is as follows: ; where represents the output feature of the first GCN, that is, the original aggregated features; Represents a graph convolution operation that does not change the dimension of the input features. The dimensions of both the input features and the output features are 128; Represents an activation function.
[0073] Input features The dimension of B × T × 128 (B is the batch size, T is the number of time steps, and 128 is the feature dimension). This layer does not compress the feature dimension but only performs feature transformation.
[0074] Aggregate the spatio-temporal neighborhood information defined by the adjacency matrix A through the first layer of graph convolution (GCNConv), and the output dimension remains 128.
[0075] The above graph convolution process is further illustrated by an example below.
[0076] 1) Assumptions: Input features: , .
[0077] Adjacency matrix: , with a window size of 2, self-loop edges + temporal adjacency edges (adjacent edges).
[0078] Weight matrix: .
[0079] 2) Calculation steps: Adjacency matrix aggregation: .
[0080] The feature corresponding to the aggregated node : .
[0081] The feature corresponding to the aggregated node : .
[0082] The feature corresponding to the aggregated node : .
[0083] Further compress the feature dimension through the second layer of GCN to enhance the discriminative representation of the spatial relationship. The mathematical expression is as follows: ; where represents the output features of the second layer of GCN, that is, the aggregated features; represents the graph convolution operation, which changes the dimension of the input features. The dimension of the input features is 128, and the dimension of the output features is 64; represents the activation function.
[0084] As an alternative implementation, the classification decision layer includes a global average pooling layer, a first fully-connected layer, and a second fully-connected layer connected in sequence.
[0085] In the global average pooling layer, features are aggregated along the time step dimension, global average pooling is performed, and the overall statistical characteristics of the target motion are retained. The mathematical expression is as follows: ; where represents the output feature of the global average pooling layer; represents the feature part corresponding to the time step t (data frame t ) in the output feature of the second layer of GCN.
[0086] There are two fully-connected layers. The input feature dimension of the first fully-connected layer is 64, and the output feature dimension is 32; the input feature dimension of the second fully-connected layer is 32, and the output dimension is C, that is, the probabilities of C target categories are output. The mathematical expressions of the two fully-connected layers are as follows: ; where , are the probabilities of the 1st to the C th target categories respectively; represents the two fully-connected operations; represents the activation function.
[0087] Regarding the training and validation of the DS-TGCN network model.
[0088] After constructing the DS-TGCN network model, the graph structure (model) dataset for training is divided into a training set and a validation set. Cross-entropy loss with label smoothing and regularization is adopted, the optimizer AdamW and early stopping strategy are used; Dropout (probability 0.3) and BatchNorm are used to prevent overfitting. The model is trained and validated. After the DS-TGCN network model completes training and validation, it is loaded for radar target classification of measured target data. Refer to Figure 7 , Figure 7 is a schematic diagram of the method process including the model training process. Among them, the graph structure model dataset can be established based on historical radar monitoring data, and the construction method of each graph structure model in the dataset is referred to the previous text.
[0089] As an alternative implementation, the label-smoothing cross-entropy loss used during training is defined as: ; where is the cOne-hot encoding of the true labels corresponding to each target category For the c model prediction probabilities corresponding to each target category, with 0.1 as the label smoothing coefficient.
[0090] See Figure 7 , after the DS-TGCN network model is trained and verified, it is then loaded for radar target classification of the measured target data, specifically as follows: 1) Model loading and initialization: Load the parameters of the trained DS-TGCN network model from the storage medium into the inference environment.
[0091] 2) Preprocessing of the measured data.
[0092] 3) Graph data construction: Convert the preprocessed measured data into a graph structure model for model input.
[0093] 4) Model inference: Input the constructed graph data into the DS-TGCN network (model), perform forward propagation, and output the category probabilities .
[0094] 5) Category determination: Take the index of the maximum value in the probability vector as the final classification result: ; That is, determine the target category corresponding to the maximum probability as the target category monitored by the radar for the target data frame sequence.
[0095] The output of the last fully connected layer of the DS-TGCN network model is the probability distribution of each target category, not the probability of a certain target category. It is to take the index of the maximum value in the probability vector as the final classification result.
[0096] Example: 1) Radar data of 3 target categories are trained through the DS-TGCN network. Let the number of time steps .
[0097] 2) The measured radar echo signal is preprocessed to generate graph data and input into the trained DS-TGCN network model.
[0098] 3) After model inference, the probability vector is output, and the number of categories C = 3, corresponding to the labels 0, 1, 2).
[0099] 4) Take the index of the maximum probability as the final classification result. Based on the above probability vector, the classification result is Class = 1. If Class = 1 corresponds to the "helicopter" category, then the finally identified target category is a helicopter.
[0100] The present application also provides an application scenario, which applies the above-mentioned radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network. Specifically: The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network provided in this embodiment can be applied in a radar target classification scenario. The radar target classification scenario includes a radar echo signal acquisition link and a radar target classification link; the radar echo signal enters the radar target classification link from the radar echo signal acquisition link. The radar target classification method based on the dual-stream spatio-temporal cross-attention graph convolutional network provided in this embodiment belongs to the radar target classification link. Specifically, in the radar target classification link, the obtained radar echo signal sequence is processed to obtain one or more data frame sequences; for any one of the data frame sequences, a first operation is performed, specifically: taking each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining the adjacency matrix corresponding to the graph structure model; the target data frame sequence is any one of the data frame sequences; using the trained dual-stream spatio-temporal cross-attention graph convolutional network with the target data frame sequence and the adjacency matrix as inputs, determining the probability of each target category monitored by the radar corresponding to the target data frame sequence; and determining the target category corresponding to the target data frame sequence as the target category with the maximum probability.
[0101] Compared with traditional radar target classification methods, the present application shows the following advantages in the radar target classification task: The present application uses a dual-stream spatio-temporal cross-attention graph convolutional network to fuse Doppler spectra and physical motion parameters, realizing deep complementarity and fine-grained alignment of cross-modal features, and significantly improving the target discrimination ability in complex electromagnetic environments; combining temporal attention and fixed-edge-weight graph convolution to hierarchically model global time dependencies and local spatial correlations, taking into account computational efficiency and anti-noise robustness, and effectively suppressing clutter interference; through an interpretable edge weight matrix and attention weight visualization, intuitively revealing the target motion law and modal interaction mechanism, providing a high-precision, high-efficiency, and high-reliability solution for radar target classification, and adapting to the real-time processing requirements of radar embedded platforms.
[0102] The present application overcomes the core pain points of traditional methods, such as low classification accuracy, strong noise sensitivity, and insufficient interpretability in complex environments, through the deep integration of multi-modal collaborative perception, spatio-temporal joint modeling, and lightweight computing architectures, providing an innovative solution for the intelligent perception of radar targets, and is applicable to key fields such as military security and airspace monitoring.
[0103] This application extracts target features through a dual-branch structure capable of independently processing Doppler spectra (data) and physical motion parameters, and combines a bidirectional cross-attention mechanism to achieve fine-grained alignment and interaction of cross-modal features; introduces a fixed edge weight strategy that connects adjacent time-step nodes based on a sliding window strategy to suppress noise interference and significantly improve computational efficiency; and finally outputs a high-confidence target category through a lightweight classification module.
[0104] Experiments show that the DS-TGCN network proposed in this application significantly improves the classification performance of radar targets by fusing multi-modal data and spatio-temporal attention mechanisms, and demonstrates excellent classification performance and real-time performance in the measured radar target dataset.
[0105] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to radar target classification. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a radar target classification method based on a dual-stream spatio-temporal cross-attention graph convolutional network.
[0106] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, implements the steps in the above method embodiments.
[0107] In an exemplary embodiment, a computer program product is provided, including a computer program, which when executed by a processor, implements the steps in the above method embodiments.
[0108] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0109] In this text, specific examples are used to illustrate the principle and implementation of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A radar target classification method based on a dual-stream spatiotemporal cross-attention graph convolutional network, characterized in that: include: Processing the acquired radar echo signal sequence to obtain one or more data frame sequences; each of the data frame sequences includes a plurality of data frames, and each data frame includes Doppler spectrum data and physical motion parameters; For any data frame sequence, performing a first operation; The first operation is: Taking each data frame in the target data frame sequence as a graph node, constructing a graph structure model, and determining an adjacency matrix corresponding to the graph structure model; the target data frame sequence is any of the data frame sequences; Taking the target data frame sequence and the adjacency matrix as input, using the trained two-stream spatiotemporal cross-attention graph convolutional network, determine the probability of each target category monitored by the radar corresponding to the target data frame sequence; Determine the target category corresponding to the maximum probability as the target category monitored by the radar corresponding to the target data frame sequence; The dual-stream spatiotemporal cross-attention graph convolution network includes a dual-stream feature extraction layer, a bidirectional cross-attention layer, a temporal attention layer, a graph convolution layer and a classification decision layer connected in sequence; wherein the dual-stream feature extraction layer is used to extract the features of Doppler spectrum data in each data frame, and extract the features of physical motion parameters in each data frame; the bidirectional cross-attention layer is used to fuse the features of physical motion parameters with the features of Doppler spectrum data to obtain a first fused feature, fuse the features of Doppler spectrum data with the features of physical motion parameters to obtain a second fused feature, and splice the first fused feature and the second fused feature to obtain a spliced feature; The temporal attention layer is used to extract the temporal information of the splicing features to obtain the temporal features; the graph convolution layer is used to aggregate the temporal features of each graph node and the temporal features of the adjacent graph nodes of each graph node according to the adjacency matrix to obtain the aggregated features of each graph node; the classification decision layer outputs the probability of each target category monitored by the radar according to the aggregated features of all graph nodes.
2. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: Take each data frame in the target data frame sequence as a graph node, build a graph structure model, and determine the adjacency matrix corresponding to the graph structure model, including: Taking each data frame in the target data frame sequence as a graph node, using the sliding window strategy, establish temporal adjacent edges between different graph nodes; wherein temporal adjacent edges are established between different graph nodes in the sliding window; and the weight of the temporal adjacent edges is 1; A self-loop edge of each graph node is established to obtain a graph structure model; the weight of the self-loop edge is 1; An adjacency matrix of the graph structure model is determined according to the graph nodes, the temporal adjacent edges, the self-loop edges, the weights of the temporal adjacent edges and the weights of the self-loop edges.
3. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The dual-stream feature extraction layer includes two feature extraction layers; each of the feature extraction layers includes a one-dimensional convolution and an activation function; wherein one feature extraction layer is used to extract the features of the Doppler spectrum data in each of the data frames; and the other feature extraction layer is used to extract the features of the physical motion parameters in each of the data frames.
4. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The bidirectional cross-attention layer includes a bidirectional cross-attention mechanism and a splicing operation; the bidirectional cross-attention mechanism is used to fuse the features of the physical motion parameters with the features of the Doppler spectrum data to obtain a first fused feature, and to fuse the features of the Doppler spectrum data with the features of the physical motion parameters to obtain a second fused feature; the splicing operation is used to splice the first fused feature and the second fused feature to obtain a spliced feature.
5. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 4 is characterized in that: The operation process of fusing the features of the physical motion parameters with the features of the Doppler spectrum data to obtain the first fusion feature is: ; ; ; The operation process of fusing the features of Doppler spectrum data with the features of physical motion parameters to obtain the second fusion feature is: ; ; ; in, , Respectively represent the first i The characteristics of the Doppler spectrum data corresponding to the data frame, j The characteristics of the Doppler spectrum data corresponding to the data frames; , Respectively represent the first i The characteristics of the physical motion parameters corresponding to the data frame, j The characteristics of the physical motion parameters corresponding to the data frames; , are the first i The learnable projection matrix corresponding to the data frame; , The first j The learnable projection matrix corresponding to the data frame; is the first i Query corresponding to Doppler spectrum data in data frames; are the first j The key and value corresponding to the Doppler spectrum data in each data frame; is the first i Query corresponding to physical motion parameters in data frames; are the first j The keys and values corresponding to the physical motion parameters in each data frame; represents the activation function; is the attention scaling factor; superscript symbol" ” indicates transposition; express and The attention weight of express and The attention weight of Indicates the first i The first fusion feature corresponding to the data frame; Indicates the first i The second fusion feature corresponding to the data frame.
6. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The temporal attention layer is a multi-head self-attention mechanism.
7. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 6 is characterized in that: The temporal information of the splicing features is extracted to obtain the temporal features, including: Split the concatenated features along the channel dimension to obtain the subspace corresponding to each attention head; For the subspace corresponding to each attention head, the calculation process of each attention head is: ; ; ; ; in, Indicates k The subspace corresponding to the attention head is h is the number of attention heads; Respectively k The learnable parameter matrix corresponding to the attention heads; , Respectively represent k The corresponding target data frames in the attention heads are from the first to the N Query of data frames; , Respectively represent k The attention heads correspond to the target data frame sequence from the 1st to the N The key of the data frame; , Respectively represent k The attention heads correspond to the target data frame sequence from the 1st to the N The value of a data frame; is the activation function; Indicates k Output features of the attention heads; The output features of all attention heads are concatenated and dimension restored. The specific operation process is as follows: ; in, It is the temporal features output by the temporal attention layer; Represents a splicing operation; The first to h Output features of the attention heads; is the output projection matrix.
8. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The graph convolution layer includes two layers of graph convolution networks connected in sequence; wherein the first layer of graph convolution network is used to aggregate the temporal features of each graph node and the temporal features of the adjacent nodes of each graph node according to the adjacency matrix to obtain the original aggregated features of each graph node; The second-layer graph convolutional network is used to compress the feature dimension of the original aggregated features of each graph node according to the adjacency matrix to obtain the aggregated features of each graph node.
9. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The classification decision layer includes a global mean pooling layer, a first fully connected layer, and a second fully connected layer which are connected in sequence.
10. The radar target classification method based on dual-stream spatiotemporal cross-attention graph convolutional network according to claim 1 is characterized in that: The acquired radar echo signal sequence is processed to obtain one or more data frame sequences, including: Preprocessing the acquired radar echo signal sequence to obtain a preprocessed signal sequence; the preprocessing includes pulse compression and range unit interception; performing separation processing on the preprocessed signal sequence to obtain one or more signal sequences; For any signal sequence, an extraction operation is performed; The extraction operation is: An extraction operation is performed on a target signal sequence to obtain an original data frame sequence corresponding to the target signal sequence; the original data frame sequence includes a plurality of original data frames; each of the original data frames includes original Doppler spectrum data and original physical motion parameters; the target signal sequence is any one of the signal sequences; The original data frame sequence corresponding to the target signal sequence is normalized to obtain a data frame sequence corresponding to the target signal sequence.
Citation Information
Patent Citations
PD radar target detection method based on graph attention network and transfer learning
CN114814776A
Radar multi-angle multi-feature fusion target classification method
CN119992228A
Human-robot collaboration method based on multi-scale graph convolutional neural network
US12159486B1
Cited By
Moving target identification method based on bidirectional cross attention
CN120953899A
Radar clutter map target detection method and device based on double-branch cross attention neural network, equipment and medium
CN122110009A
Radar clutter map target detection method and device based on double-branch cross-attention neural network, equipment and medium
CN122110009B