Low altitude radar ground clutter suppression method based on machine learning
By constructing an RD-A data cube and an improved dual-branch Transformer encoder, combined with physical prior masks and space-time adaptive processing, the decoupled feature representation of targets and background and dynamic clutter suppression in low-altitude radar systems were realized. This solved the target detection problem of low-altitude radar systems in complex environments and improved the detection effect and adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-03-31
AI Technical Summary
Existing low-altitude radar systems struggle to effectively separate and detect weak, low-speed targets in strong, non-stationary, and time-varying low-altitude background clutter environments, resulting in high false alarm and missed detection rates. Furthermore, traditional methods are unable to maintain clutter suppression effectiveness across multiple scenarios.
A machine learning-based low-altitude radar clutter suppression method is adopted. By constructing an RD-A data cube and combining terrain data and platform attitude information, a physical prior mask is generated. An improved dual-branch Transformer encoder is used for feature extraction and interactive attention to generate a background adversarial filter mask. A joint filter is constructed by combining spatiotemporal adaptive processing to achieve decoupled feature representation of target and background and dynamic clutter suppression.
It improves the separation and detection rate of small targets, reduces the false alarm rate of the system, can effectively suppress clutter in highly dynamic environments, maintains no loss of target main lobe energy, and has strong adaptability.
Smart Images

Figure CN120993364B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-altitude radar technology, and in particular to a method for suppressing ground clutter in low-altitude radar based on machine learning. Background Technology
[0002] With the rapid growth in demand for low-altitude defense, border and coastal surveillance, and target detection in complex scenarios, low-altitude radar systems are increasingly being used in ground, sea, and urban environments. For ultra-low-altitude, small target detection scenarios against ground features, sea surfaces, and building backgrounds, the system typically needs to extract weak, low-velocity, near-zero Doppler targets from the array coherent echo. Existing technologies have significant technical bottlenecks in strong, non-stationary, and time-varying low-altitude background clutter environments.
[0003] On the one hand, traditional linear filtering and covariance modeling methods rely on a large number of targetless training samples and stable background statistical characteristics, which makes it difficult to fully characterize clutter characteristics under complex spatiotemporal changes. When the sample is limited or the environment changes rapidly, existing methods are prone to the phenomenon that the target signal is masked by the background or misjudged as clutter, resulting in a significant increase in the false alarm rate and missed alarm rate of target detection.
[0004] On the other hand, in recent years, some studies have introduced deep learning to model the spatiotemporal sequence of radar data, but most of them rely on single-branch network structures, fixed window modeling or simple feature splicing, which makes it difficult to decouple the sparse dynamic features of the target from the steady-state change features of the background. In addition, traditional background suppression methods are difficult to achieve selective dynamic suppression of clutter energy and cannot maintain clutter suppression effect under low sample and multi-scene switching conditions. Summary of the Invention
[0005] One objective of this invention is to propose a low-altitude radar ground clutter suppression method based on machine learning, which improves the separation and detection rate of small targets.
[0006] A low-altitude radar ground clutter suppression method based on machine learning according to an embodiment of the present invention includes the following steps:
[0007] Receive array coherent echo I / Q signal sequences, construct RD-A data cubes, and generate representation-learned RD-A label sequences based on the RD-A data cubes;
[0008] Based on RD-A data cube estimation, the STAP covariance matrix and STAP subspace of spatiotemporal adaptive processing are estimated, and a set of physical prior masks is generated by combining terrain data, platform attitude information and background Doppler spectrum distribution.
[0009] The RD-A tag sequence and the physical prior mask set are input into the improved dual-branch Transformer encoder to initialize the parallel encoding process of the background modeling branch and the target modeling branch, and obtain the initial values of the background embedding representation and the target embedding representation.
[0010] The cross-branch interactive attention module is used to perform information interaction and feature alignment between the initial values of the background embedding representation and the initial values of the target embedding representation, resulting in interactively enhanced background embedding representation and interactively enhanced target embedding representation. Orthogonal constraint terms and mutual information constraint terms are applied to the interactively enhanced background embedding representation and the interactively enhanced target embedding representation to obtain orthogonally decoupled background embedding representation and orthogonally decoupled target embedding representation.
[0011] The orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation are input into the background adversarial filtering module, and the output is a background adversarial filtering mask for background suppression. The background adversarial filtering mask is fused with the STAP covariance matrix and the STAP subspace of the space-time adaptive processing to construct a joint filter. The joint filter is then used to filter the RD-A data cube to obtain the clutter-suppressed RD-A data cube.
[0012] Optionally, the construction of an RD-A data cube comprising distance, Doppler, and orientation dimensions, and the generation of RD-A label sequences for representation learning based on the RD-A data cube, includes:
[0013] The array coherent echo I / Q signal sequence is acquired, and time synchronization processing, frequency correction processing, and channel amplitude-phase consistency calibration are performed on the array coherent echo I / Q signal sequence to obtain the calibrated array coherent echo I / Q signal sequence.
[0014] The pulse compression process is performed on the calibrated array coherent echo I / Q signal sequence to obtain the pulse compression output signal tensor.
[0015] The pulse compression output signal tensor is processed for moving target display in the pulse dimension to generate a data tensor after moving target suppression, and then Doppler transformation is performed in the pulse dimension to output the Doppler spectrum data tensor.
[0016] The Doppler spectrum data tensor is processed by array beamforming in the spatial dimension to obtain the RD-A data cube;
[0017] The RD-A data cube is mapped to an RD-A tag sequence. Each tag in the RD-A tag sequence is composed of the complex spectrum value, phase, distance coordinate, Doppler frequency coordinate and pitch direction coordinate of the corresponding RD-A data cube.
[0018] Optionally, the step of generating a set of physical prior masks by combining terrain data, platform attitude information, and background Doppler spectrum distribution includes:
[0019] A training region containing background clutter but not moving targets is selected from the RD-A data cube. The RD-A data within the training region is flattened, and the three-dimensional complex spectral tensor is combined into a two-dimensional complex sample matrix according to the pitch direction and Doppler frequency dimension.
[0020] The STAP covariance matrix of the space-time adaptive processing is calculated based on the two-dimensional complex sample matrix. The STAP covariance matrix of the space-time adaptive processing is decomposed into eigenvalues, and the top few principal eigenvectors with the largest eigenvalues are selected to form the STAP subspace of the space-time adaptive processing.
[0021] Based on the platform's flight attitude parameters, spatial orientation masks for each pitch direction and each Doppler frequency are calculated.
[0022] The background spectral energy at each pitch direction and Doppler frequency in the RD-A data cube is statistically analyzed, and the squared values of the spectral amplitude on all range cells are averaged to obtain the background spectral intensity map.
[0023] Set a background spectral energy threshold and compare each element in the background spectral intensity map with the background spectral energy threshold:
[0024] If the average background energy level under a certain pitch direction and Doppler frequency combination is higher than the background spectrum energy threshold, then the element of the corresponding Doppler spectrum intensity mask is assigned a value of 1; otherwise, it is assigned a value of 0.
[0025] By performing a logical AND operation between the spatial orientation mask and the Doppler spectrum intensity mask at corresponding positions, the final set of physical prior masks is obtained.
[0026] Each element in the physical prior mask set is assigned a value of 1 only when both the spatial orientation mask and the Doppler spectral intensity mask are 1, indicating that the terrain occlusion and background spectral energy salience conditions are satisfied simultaneously under the combination of pitch direction and Doppler frequency.
[0027] Optionally, the improved dual-branch Transformer encoder includes:
[0028] The RD-A tag sequence obtained by mapping the RD-A data cube is matched one-to-one with the physical prior mask set at the corresponding position, and the input feature alignment matrix is obtained for each RD-A tag;
[0029] The input feature alignment matrix is input into the background modeling branch and the target modeling branch respectively, and the parallel encoding process of the improved dual-branch Transformer encoder is started.
[0030] In the background modeling branch, the input feature alignment matrix is used as the initial input, and the attention guidance mask of the physical prior mask set is applied so that the self-attention only focuses on the RD-A position of the background region marked by the physical prior mask set. In the encoding process, a multi-level Transformer structure is used to model the input features layer by layer, and the output feature sequence is the initial value of the background embedding representation.
[0031] In the target modeling branch, the input feature alignment matrix is used as the initial input. The relationship between different physical features is captured through the fully connected attention mechanism and the sparse connection strategy. The multi-scale window attention mechanism is used to progressively aggregate the input features at each coding layer. The output feature sequence is the initial value of the target embedding representation.
[0032] Optionally, the background modeling branch includes:
[0033] The input feature alignment matrix is used as the initial input for the background modeling branch. For the initial input feature alignment matrix, an attention mask matrix is constructed using the physical prior mask set.
[0034] In each layer of the Transformer encoder, the attention mask matrix is used in the self-attention module, so that each feature vector is only used for attention calculation with the feature vectors marked as background regions by the physical prior mask set, and only the background region is considered in the background modeling process.
[0035] A multi-layered Transformer structure is used for progressive modeling. In each Transformer encoder layer, the feature sequence is sequentially subjected to linear transformation, mask self-attention mechanism, feedforward network activation, normalization operation and residual connection. The output of the previous layer is the input of the next layer. The output of each layer contains the multi-dimensional feature aggregation effect under mask constraints. In the feature sequence output by the last Transformer encoder layer, the aggregation result of all feature vectors is defined as the initial value of the background embedding representation.
[0036] Background embedding means that each feature vector in the initial value corresponds one-to-one with an RD-A position in the input feature alignment matrix that is marked as a background region by the physical prior mask set. The content is the deep background feature learned by the position after being encoded by multiple Transformers. All attention calculations, feature aggregations and output vector constraints are strictly limited by the physical prior mask set, and only rely on background region information without containing any target branch features.
[0037] Optionally, the target modeling branch includes:
[0038] The target modeling branch takes the input feature alignment matrix as the starting point for modeling. In each Transformer encoding layer, all feature vectors are simultaneously subjected to a fully connected attention mechanism and a sparse connection strategy, and the output is a weighted aggregated feature representation matrix of all target region features in the current encoding layer.
[0039] In each Transformer encoding layer, a multi-scale window attention mechanism is applied based on the weighted aggregated feature representation matrix of the target region output by the previous layer. The weighted aggregated feature representation matrix is divided into several local windows according to different scale granularities. Self-attention is calculated independently in each window, and the weighted aggregated result of the target features under the multi-scale window is output.
[0040] The target feature weighted aggregation results of each window at all scales are merged across scales and residually connected with the target region features of the original input feature alignment matrix. The input is fed forward neural network module for nonlinear transformation and outputs a new target region feature weighted aggregation feature representation matrix, which is used as the input of the next coding layer. In all coding layers, the target region feature weighted aggregation feature representation matrix output by the last layer is defined as the initial value of the target embedding representation.
[0041] Optionally, the cross-branch interaction attention module and orthogonal decoupling mechanism include:
[0042] Initial values for the background embedding representation and the target embedding representation are input into the cross-branch interactive attention module, serving as the query vector sequence and key-value vector sequence, respectively. Through the interactive attention mechanism, feature dependencies are established between the two branches. The relevance weight of each initial feature point of the background embedding representation to all initial feature points of the target embedding representation is calculated, and the initial feature points of the target embedding representation are weighted and aggregated based on the relevance weight to form the interactively enhanced background embedding representation. Similarly, the relevance weight of each initial feature point of the target embedding representation to all initial feature points of the background embedding representation is calculated, and the initial feature points of the background embedding representation are weighted and aggregated based on the relevance weight to form the interactively enhanced target embedding representation.
[0043] Orthogonal constraints are applied to the interaction-enhancing background embedding representation and the interaction-enhancing target embedding representation, respectively;
[0044] Construct mutual information constraints between the interaction-enhancing background embedding representation and the interaction-enhancing target embedding representation;
[0045] The orthogonal constraint term and the mutual information constraint term are weighted and combined to form the overall optimization objective function. After jointly optimizing the objective function, the orthogonal decoupled background embedding representation and the orthogonal decoupled objective embedding representation are output.
[0046] Optionally, obtaining the clutter-suppressed RD-A data cube includes:
[0047] In the background adversarial filtering module, the orthogonal decoupled background embedding representation and orthogonal decoupled target embedding representation corresponding to each RD-A tag are concatenated by the background discriminant mapping. The confidence score of the background is output by the background discriminant mapping. The background confidence scores of all RD-A tags are arranged in order to form a background confidence sequence.
[0048] The background confidence sequence is grouped according to the pitch direction and Doppler frequency. The background confidence of all range cells under the same pitch direction and the same Doppler frequency is averaged to obtain the background adversarial filter mask defined on the pitch-Doppler plane.
[0049] Construct a weighted clutter subspace projection matrix based on the STAP covariance matrix and the STAP subspace of space-time adaptive processing;
[0050] The background adversarial filtering mask is unfolded into a vector, and a diagonal weight matrix corresponding one-to-one with the pitch-Doppler flattening domain is generated. A joint filter is constructed based on the diagonal weight matrix and the weighted clutter subspace projection matrix. The joint filter outputs a joint filter matrix consistent with the pitch-Doppler flattening domain by subtracting the product of the diagonal weight matrix and the weighted clutter subspace projection matrix from the identity matrix.
[0051] The pitch-Doppler flattening domain of each range cell in the RD-A data cube is flattened into a vector and used as the input of the joint filter. After transformation by the joint filter matrix, the output vector is obtained. The output vector is restored to the pitch-Doppler flattening domain structure and filled back into the position of the corresponding range cell. This process is repeated for all range cells to generate the clutter-suppressed RD-A data cube.
[0052] The beneficial effects of this invention are:
[0053] (1) The present invention adopts an improved dual-branch Transformer architecture, which sends the RD-A data cube into the background modeling branch and the target modeling branch respectively through the input feature alignment matrix, and uses physical prior masks and multi-scale attention mechanisms for feature extraction. The background modeling branch focuses on the spatial structure and slowly varying characteristics of background clutter, while the target modeling branch focuses on the sparse motion and micro-Doppler features of weak maneuvering targets. Through cross-branch interactive attention, orthogonal constraints and mutual information constraints, the two types of features are ensured to be maximized and minimized in the expression space, thereby achieving decoupled feature expression of target-background and improving the separation and detection rate of small targets.
[0054] (2) In this invention, a physical prior mask based on terrain, platform attitude and STAP covariance estimation is applied to the background modeling branch to realize the selection of physical constraint regions in the RD-A domain. Through the background adversarial filtering module, the background adversarial filtering mask is generated by using orthogonal decoupling features. Combined with the spatiotemporal adaptive processing of the STAP subspace, a joint filter is constructed. It not only dynamically suppresses the main background energy but also is compatible with nonlinear distribution and abrupt clutter. It has better adaptability to highly dynamic changing environments, realizes adaptive and interpretable dynamic clutter suppression, and has no loss of target main lobe energy. The clutter residual rate is significantly reduced, and the overall false alarm rate of the system is significantly reduced. Attached Figure Description
[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0056] Figure 1 This is a flowchart of a low-altitude radar ground clutter suppression method based on machine learning proposed in this invention. Detailed Implementation
[0057] Example 1:
[0058] refer to Figure 1 A machine learning-based method for suppressing low-altitude radar clutter includes the following steps:
[0059] Receive array coherent echo I / Q signal sequences, construct RD-A data cubes, and generate representation-learned RD-A label sequences based on the RD-A data cubes;
[0060] In this embodiment, an RD-A data cube containing distance, Doppler, and orientation dimensions is constructed, and an RD-A label sequence for representation learning is generated based on the RD-A data cube, including:
[0061] The array coherent echo I / Q signal sequence is acquired, and time synchronization processing, frequency correction processing, and channel amplitude-phase consistency calibration are performed on the array coherent echo I / Q signal sequence to obtain the calibrated array coherent echo I / Q signal sequence.
[0062] The array coherent echo I / Q signal sequence is numbered according to array elements, pulses, and range cells to form a three-dimensional complex signal tensor. Each complex value in the array coherent echo I / Q signal sequence corresponds to the received signal of a certain array element, a certain pulse time, and a certain range cell.
[0063] After calibration, each complex value of the array coherent echo I / Q signal sequence is equal to the corresponding complex value in the original array coherent echo I / Q signal sequence multiplied by the amplitude gain coefficient of the array element, and multiplied by a phase correction factor representing the carrier frequency deviation, to ensure that the amplitude and phase of all channels are consistent. The phase correction factor is set manually.
[0064] The pulse compression process is performed on the calibrated array coherent echo I / Q signal sequence to obtain the pulse compression output signal tensor.
[0065] Pulse compression processing achieves main lobe gain and side lobe suppression by matching and filtering the echo signal of each distance unit with the complex conjugate inversion of the transmitted pulse waveform.
[0066] The pulse compression output signal tensor is processed for moving target display in the pulse dimension to generate a data tensor after moving target suppression, and then Doppler transformation is performed in the pulse dimension to output the Doppler spectrum data tensor.
[0067] The moving target display processing is obtained by the output value at each time step being equal to the difference between the pulse compression output signal at the current time step and the previous time step; the Doppler transformation is calculated using the fast Fourier transform method, which transforms a set of time-series signals on each array element and each range unit into complex spectrum values in the Doppler frequency domain.
[0068] The Doppler spectrum data tensor is processed by array beamforming in the spatial dimension to obtain the RD-A data cube;
[0069] Array beamforming achieves energy focusing in each elevation direction by multiplying the array complex value on each range-Doppler cell with the digital beam weight vector in a one-to-one correspondence and then summing the results, resulting in an RD-A data cube. Each complex spectral value of the RD-A data cube corresponds to the joint component of a certain elevation direction, a certain Doppler frequency, and a certain range cell.
[0070] The RD-A data cube is mapped to an RD-A tag sequence. Each tag in the RD-A tag sequence is composed of the complex spectrum value, phase, distance coordinate, Doppler frequency coordinate and pitch direction coordinate of the corresponding RD-A data cube.
[0071] Phase is obtained by calculating the argument of the complex spectrum value. The phase of each point in the DA data cube is the phase part of the corresponding complex spectrum value. Each layer of the RD-A data cube corresponds to a unique distance coordinate. Each frequency point obtained by performing a fast Fourier transform on the pulse dimension of the data tensor after suppressing the moving target is the Doppler frequency coordinate. Each row of the RD-A data cube corresponds to a unique Doppler frequency coordinate. During array beamforming processing, a set of weighted outputs is generated for each specified pitch direction. Each direction corresponds to a weight vector, and the output is the pitch direction coordinate corresponding to that direction. Each block of the RD-A data cube corresponds one-to-one with a unique pitch direction coordinate.
[0072] Based on RD-A data cube estimation, the STAP covariance matrix and STAP subspace of spatiotemporal adaptive processing are estimated, and a set of physical prior masks is generated by combining terrain data, platform attitude information and background Doppler spectrum distribution.
[0073] In this embodiment, a set of physical prior masks is generated by combining terrain data, platform attitude information, and background Doppler spectrum distribution, including:
[0074] A training region containing background clutter but not moving targets is selected from the RD-A data cube. The RD-A data within the training region is flattened, and the three-dimensional complex spectral tensor is combined into a two-dimensional complex sample matrix according to the pitch direction and Doppler frequency dimension.
[0075] Each sample vector in the two-dimensional complex sample matrix represents the complex spectral value of all pitch directions and all Doppler frequency combinations under a specific range cell.
[0076] The STAP covariance matrix of the space-time adaptive processing is calculated based on the two-dimensional complex sample matrix. The STAP covariance matrix of the space-time adaptive processing is decomposed into eigenvalues, and the top few principal eigenvectors with the largest eigenvalues are selected to form the STAP subspace of the space-time adaptive processing.
[0077] The STAP covariance matrix of the space-time adaptive processing describes the correlation characteristics of background clutter in the joint space-time dimension. Each element of the STAP covariance matrix of the space-time adaptive processing is equal to the average of the complex product of all training sample vectors in the corresponding two dimensions. The STAP subspace of the space-time adaptive processing represents the principal component direction in which the energy of the background clutter is most concentrated in the space-time domain.
[0078] Based on the platform's flight attitude parameters, spatial orientation masks for each pitch direction and each Doppler frequency are calculated.
[0079] The platform's flight attitude parameters include pitch angle, heading angle, roll angle, aircraft altitude, and terrain elevation model. For each pitch direction and each Doppler frequency, a spatial orientation mask is obtained. Based on the platform's current position, flight attitude parameters, and radar installation parameters, the spatial pointing vector of the radar beam corresponding to each pitch direction and Doppler frequency is calculated. Then, a ray projection is made from the platform's current altitude along the spatial pointing vector, traversing each range cell on the range axis to obtain the corresponding three-dimensional spatial coordinates. The terrain elevation model is used to find the surface elevation at the three-dimensional coordinates, and the radar beam height at the range cell is compared with the corresponding surface elevation.
[0080] When the radar beam is at a distance cell at a height higher than the ground elevation, that distance cell is considered an unobstructed area, and the corresponding element of the spatial direction mask is assigned a value of 1 under that elevation direction and Doppler frequency combination. Conversely, when the radar beam is at a distance cell at a height equal to or lower than the ground elevation, that distance cell is considered an obstructed area, and the corresponding element of the spatial direction mask is assigned a value of 0. For the entire beam, if all distance cells are unobstructed areas, the spatial direction mask is assigned a value of 1; otherwise, if even one distance cell is obstructed, the spatial direction mask is assigned a value of 0.
[0081] In Example 1, assuming a platform is at a height of 100 meters, a pitch angle of 5 degrees, a yaw angle of 30 degrees, and a roll angle of 0 degrees, the corresponding spatial pointing vector is known. Under a beam with a Doppler frequency of f1, the spatial coordinates of different distance units such as 0 meters, 100 meters, and 200 meters are calculated sequentially from the platform's starting point. At a distance of 200 meters, the ray height is 95 meters. The terrain elevation model is consulted. If the ground elevation at that location is 80 meters, then the point is unobstructed and assigned a value of 1. If the ray height in the next distance unit is 90 meters and the ground elevation is 92 meters, then it is obstructed and assigned a value of 0. Finally, if all distance units under a certain pitch direction and Doppler frequency combination are unobstructed, the combination is assigned a value of 1; otherwise, it is assigned a value of 0.
[0082] The background spectral energy at each pitch direction and Doppler frequency in the RD-A data cube is statistically analyzed, and the squared values of the spectral amplitude on all range cells are averaged to obtain the background spectral intensity map.
[0083] For the RD-A data cube, for each combination of pitch direction and Doppler frequency, all range cells are traversed. For the complex spectral values of all range cells under the same pitch direction and Doppler frequency, the amplitude square of the complex spectral value of each range cell is calculated. The amplitude squares of all range cells under the same pitch direction and Doppler frequency are summed, and the sum is averaged using the number of range cells to obtain the average background spectral energy under the pitch direction and Doppler frequency.
[0084] Each element of the background spectral intensity map is equal to the sum of the squares of the complex spectral values of all range cells corresponding to the pitch direction and Doppler frequency, divided by the total number of range cells; each element of the background spectral intensity map reflects the average background energy level of the entire range cell at a specific pitch direction and a specific Doppler frequency.
[0085] Set a background spectral energy threshold and compare each element in the background spectral intensity map with the background spectral energy threshold:
[0086] If the average background energy level under a certain pitch direction and Doppler frequency combination is higher than the background spectrum energy threshold, then the element of the corresponding Doppler spectrum intensity mask is assigned a value of 1; otherwise, it is assigned a value of 0.
[0087] By performing a logical AND operation between the spatial orientation mask and the Doppler spectrum intensity mask at corresponding positions, the final set of physical prior masks is obtained.
[0088] Each element in the physical prior mask set is assigned a value of 1 only when both the spatial orientation mask and the Doppler spectral intensity mask are 1, indicating that the terrain occlusion and background spectral energy salience conditions are satisfied simultaneously under the combination of pitch direction and Doppler frequency.
[0089] The RD-A tag sequence and the physical prior mask set are input into the improved dual-branch Transformer encoder to initialize the parallel encoding process of the background modeling branch and the target modeling branch, and obtain the initial values of the background embedding representation and the target embedding representation.
[0090] In this embodiment, the improved dual-branch Transformer encoder includes:
[0091] The RD-A tag sequence obtained by mapping the RD-A data cube is matched one-to-one with the physical prior mask set at the corresponding position, and the input feature alignment matrix is obtained for each RD-A tag;
[0092] Each eigenvector of the input feature alignment matrix is composed of complex spectral values, phase, distance coordinates, Doppler frequency coordinates, and pitch coordinates in a uniform order. All input feature vectors are normalized and position-encoded to ensure that each feature vector is comparable in different physical meanings and meets the model input format requirements.
[0093] The input feature alignment matrix is input into the background modeling branch and the target modeling branch respectively, and the parallel encoding process of the improved dual-branch Transformer encoder is started.
[0094] In the background modeling branch, the input feature alignment matrix is used as the initial input, and the attention guidance mask of the physical prior mask set is applied so that the self-attention only focuses on the RD-A position of the background region marked by the physical prior mask set. In the encoding process, a multi-level Transformer structure is used to model the input features layer by layer, and the output feature sequence is the initial value of the background embedding representation.
[0095] The initial value of the background embedding representation represents the joint representation of all background region features learned by the background modeling branch after multiple layers of encoding. Each vector of the initial value of the background embedding representation corresponds to a background region feature point in the input features.
[0096] In this embodiment, the background modeling branch includes:
[0097] The input feature alignment matrix is used as the initial input for the background modeling branch. For the initial input feature alignment matrix, an attention mask matrix is constructed using the physical prior mask set.
[0098] Each element of the attention mask matrix indicates whether the pitch direction and Doppler frequency position corresponding to the two feature vectors are both 1 in the physical prior mask set. If they are both 1, the value is assigned to 1; otherwise, the value is assigned to 0. The attention mask matrix is used to mask the target region during the attention calculation process.
[0099] In each layer of the Transformer encoder, the attention mask matrix is used in the self-attention module, so that each feature vector is only used for attention calculation with the feature vectors marked as background regions by the physical prior mask set, and only the background region is considered in the background modeling process.
[0100] A multi-layered Transformer structure is used for progressive modeling. In each Transformer encoder layer, the feature sequence is sequentially subjected to linear transformation, mask self-attention mechanism, feedforward network activation, normalization operation and residual connection. The output of the previous layer is the input of the next layer. The output of each layer contains the multi-dimensional feature aggregation effect under mask constraints. In the feature sequence output by the last Transformer encoder layer, the aggregation result of all feature vectors is defined as the initial value of the background embedding representation.
[0101] In Example 1, a multi-layered Transformer structure is used for progressive modeling: for the initial input feature sequence, a linear transformation is performed on each feature vector to obtain the linear expression of the features at that layer.
[0102] The mask self-attention mechanism is used to calculate the attention score between each feature vector and all other feature vectors. The association between features that do not belong to the background region is masked according to the attention mask matrix. Attention weights are accumulated only between regions marked as background in the physical prior mask set. The background feature vectors are weighted and aggregated to form an attention-weighted feature representation.
[0103] The weighted feature representation is input into the feedforward neural network for nonlinear activation to enhance feature representation capabilities. The feature vectors are normalized and the output feature vector is residually concatenated with the input feature vectors. The output of each Transformer encoder layer serves as the input to the next Transformer encoder layer, progressively fusing multidimensional feature information from different background regions. Through multi-layer modeling, a richer and more robust joint feature representation of the background is finally obtained.
[0104] In each coding layer, attention computation is based on a scaled dot product attention mechanism:
[0105]
[0106] Among them, Q, K, These represent the query matrix, key matrix, and value matrix, respectively, all obtained through a linear mapping from the current layer input. For scaling factor, log(A) mask This is used to encode the attention mask matrix as an additive bias to mask invalid attention paths, and all matrix calculations are limited to the background region.
[0107] In Example 1, a three-layer Transformer encoder is provided. The first layer takes the input feature alignment matrix as input, and outputs the first layer features after linear transformation, mask self-attention, feedforward activation, normalization, and residual connection. The second layer takes the output features of the first layer as input and repeats the above operations to obtain the second layer features. The third layer is similar. Finally, the output of the third layer is the background feature sequence under the multi-layer progressive modeling structure. All attention and feature operations within each layer are constrained by the physical prior mask set and only aggregate background region features. The initial value of the final output background embedding representation is obtained through the layer-by-layer progressive method.
[0108] Background embedding means that each feature vector in the initial value corresponds one-to-one with an RD-A position in the input feature alignment matrix that is marked as a background region by the physical prior mask set. The content is the deep background feature learned by the position after being encoded by multiple Transformers. All attention calculations, feature aggregations and output vector constraints are strictly limited by the physical prior mask set, and only rely on background region information without containing any target branch features.
[0109] In the target modeling branch, the input feature alignment matrix is used as the initial input. The relationship between different physical features is captured through the fully connected attention mechanism and the sparse connection strategy. The multi-scale window attention mechanism is used to progressively aggregate the input features at each coding layer. The output feature sequence is the initial value of the target embedding representation.
[0110] The initial value of the target embedding representation represents the initial representation of all target region features learned by the target modeling branch after multiple layers of encoding. Each vector of the initial value of the target embedding representation corresponds to a weak target region feature point in the input features.
[0111] In this embodiment, the target modeling branch includes:
[0112] The target modeling branch takes the input feature alignment matrix as the starting point for modeling. In each Transformer encoding layer, all feature vectors are simultaneously subjected to a fully connected attention mechanism and a sparse connection strategy, and the output is a weighted aggregated feature representation matrix of all target region features in the current encoding layer.
[0113] In Example 1, in each Transformer encoding layer, all input feature vectors are input into the fully connected attention mechanism in a uniform physical order. The fully connected attention mechanism calculates the correlation weight between each feature vector and all other feature vectors, and performs weighted aggregation on all input feature vectors based on the correlation weight to obtain the fully connected attention weighted feature representation of each feature point in the global physical space.
[0114] A sparse connection strategy is adopted for each feature vector, retaining only the feature paths that have sparse physical connections with it. For the retained feature paths, sparse correlation weights are calculated. Based on the sparse correlation weights, the artificially set associated features are sparsely weighted and aggregated to obtain the sparse connection weighted feature representation of each feature point.
[0115] In each encoding layer, for all target region feature points, the fully connected attention-weighted feature representation and the sparse connected weighted feature representation are merged respectively. The final weighted aggregated feature of each feature point is obtained by fusing through weighting coefficients. All feature point outputs are arranged in the original order to form the target region feature weighted aggregated feature representation matrix under the current encoding layer.
[0116] The target region feature weighted aggregation feature representation matrix represents each row of a target feature point at an RD-A labeled location, and each column corresponds to the multidimensional feature expression of the target feature point after weighted aggregation by fully connected and sparse attention mechanisms in this layer.
[0117] In each Transformer encoding layer, a multi-scale window attention mechanism is applied based on the weighted aggregated feature representation matrix of the target region output by the previous layer. The weighted aggregated feature representation matrix is divided into several local windows according to different scale granularities. Self-attention is calculated independently in each window, and the weighted aggregated result of the target features under the multi-scale window is output.
[0118] The multi-scale window attention mechanism refers to the following: In each Transformer encoding layer, the weighted aggregated feature representation matrix of the target region features is divided into multiple local windows according to different scale parameters. Each window covers a set of physically continuous target region features in the weighted aggregated feature representation matrix of the target region features. Self-attention calculation is performed independently within each window at each scale. All feature vectors are only aggregated with other feature vectors within the same window, and the windows do not affect each other. Different window sizes are set for each scale. There are small-scale windows to realize fine-grained target feature representation and large-scale windows to capture wide-range spatial relationships. The outputs of windows at all scales are merged to generate a multi-scale fused weighted aggregated feature representation matrix of the target region features, which is represented as the weighted aggregation result of the target features under the multi-scale window.
[0119] In Example 1, the input feature alignment matrix has N feature vectors. The fully connected attention mechanism establishes weighted connections between each feature vector and the other N-1 features. The sparse connection strategy stipulates that only features whose distance coordinates differ from the current feature by no more than a threshold or significant features within the Doppler frequency coordinates are weighted, while the weights of other feature connections are forced to zero. The multi-scale window attention mechanism divides the data into small windows covering 3×3 RD-A points and large windows covering 7×7 RD-A points. It performs independent local attention weighting on the features within each window and then merges the outputs of different windows into the global output.
[0120] The target feature weighted aggregation results of each window at all scales are merged across scales and residually connected with the target region features of the original input feature alignment matrix. The input is fed forward neural network module for nonlinear transformation and outputs a new target region feature weighted aggregation feature representation matrix, which is used as the input of the next coding layer. In all coding layers, the target region feature weighted aggregation feature representation matrix output by the last layer is defined as the initial value of the target embedding representation.
[0121] The target embedding means that each feature vector in the initial value corresponds one-to-one with an RD-A position in the input feature alignment matrix, and each vector represents the multi-dimensional target feature information learned by the target modeling branch through multi-layer progressive modeling.
[0122] The output target embedding represents that each feature vector in the initial value integrates the combined effects of fully connected attention mechanism, sparse connection strategy and multi-scale window attention mechanism, representing the coherence and saliency of low radar cross section targets, near-zero Doppler targets and weak sparse targets in the cross-pulse sequence.
[0123] The cross-branch interactive attention module is used to perform information interaction and feature alignment between the initial values of the background embedding representation and the initial values of the target embedding representation, resulting in interactively enhanced background embedding representation and interactively enhanced target embedding representation. Orthogonal constraint terms and mutual information constraint terms are applied to the interactively enhanced background embedding representation and the interactively enhanced target embedding representation to obtain orthogonally decoupled background embedding representation and orthogonally decoupled target embedding representation.
[0124] In this embodiment, the cross-branch interaction attention module and the orthogonal decoupling mechanism include:
[0125] Initial values for the background embedding representation and the target embedding representation are input into the cross-branch interactive attention module, serving as the query vector sequence and key-value vector sequence, respectively. Through the interactive attention mechanism, feature dependencies are established between the two branches. The relevance weight of each initial feature point of the background embedding representation to all initial feature points of the target embedding representation is calculated, and the initial feature points of the target embedding representation are weighted and aggregated based on the relevance weight to form the interactively enhanced background embedding representation. Similarly, the relevance weight of each initial feature point of the target embedding representation to all initial feature points of the background embedding representation is calculated, and the initial feature points of the background embedding representation are weighted and aggregated based on the relevance weight to form the interactively enhanced target embedding representation.
[0126] Orthogonal constraints are applied to the interaction-enhancing background embedding representation and the interaction-enhancing target embedding representation, respectively;
[0127] The orthogonality constraint term is calculated as follows: after performing the inner product calculation on all feature vectors in the interactive augmented background embedding representation, it is compared with the identity matrix, and the sum of squared differences is obtained; similarly, after performing the same inner product calculation on all feature vectors in the interactive augmented target embedding representation, it is compared with the identity matrix, and the sum of squared differences is obtained. The orthogonality constraint term is used to measure the independence between feature vectors. The constraint objective is to keep the interactive augmented background embedding representation and the interactive augmented target embedding representation as orthogonal as possible within their own dimensions, that is, the features are not redundant.
[0128] Construct mutual information constraints between the interaction-enhancing background embedding representation and the interaction-enhancing target embedding representation;
[0129] The mutual information constraint term is calculated as follows: for each feature vector of the interactively enhanced background embedding representation and each feature vector of the interactively enhanced target embedding representation, the normalized inner product is calculated, and all results are summed and negatively taken. The mutual information constraint term measures the information overlap between two feature sets.
[0130]
[0131] in, B is a mutual information constraint term. enh [i] and T enh [j] represents the i-th and j-th feature vectors in the interactive enhancement background embedding representation and the interactive enhancement target embedding representation, respectively. The numerator is the inner product and the denominator is the L2 norm product. The constraint objective is to minimize the mutual information so that the two feature embeddings are independent of each other in the vector space.
[0132] The orthogonal constraint term and the mutual information constraint term are weighted and combined to form the overall optimization objective function. After jointly optimizing the objective function, the orthogonal decoupled background embedding representation and the orthogonal decoupled objective embedding representation are output.
[0133] Each feature vector in the orthogonally decoupled background embedding representation and the orthogonally decoupled target embedding representation represents a multi-dimensional feature expression that maintains maximum discriminativeness and minimum redundancy after fusing cross-branch information in the background modeling branch and the target modeling branch, respectively. The orthogonal decoupling process improves the separability between the target region and the background region.
[0134] The orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation are input into the background adversarial filtering module, and the output is a background adversarial filtering mask for background suppression. The background adversarial filtering mask is fused with the STAP covariance matrix and the STAP subspace of the space-time adaptive processing to construct a joint filter. The joint filter is then used to filter the RD-A data cube to obtain the clutter-suppressed RD-A data cube.
[0135] In this embodiment, the obtained clutter-suppressed RD-A data cube includes:
[0136] In the background adversarial filtering module, the orthogonal decoupled background embedding representation and orthogonal decoupled target embedding representation corresponding to each RD-A tag are concatenated by the background discriminant mapping. The confidence score of the background is output by the background discriminant mapping. The background confidence scores of all RD-A tags are arranged in order to form a background confidence sequence.
[0137] In Example 1, background discriminant mapping is used in the background adversarial filtering module. Calculate the background confidence score for each RD-A tag:
[0138]
[0139] in, and These are the orthogonal decoupling background embedding representation and the orthogonal decoupling target embedding representation, respectively. [·‖·] denotes vector concatenation, and pbg [m]∈[0,1] represents the probability value that the m-th RD-A tag is judged as the background, which is used to generate the base score for the background adversarial filtering mask.
[0140] The background confidence sequence is grouped according to the pitch direction and Doppler frequency. The background confidence of all range cells under the same pitch direction and the same Doppler frequency is averaged to obtain the background adversarial filter mask defined on the pitch-Doppler plane.
[0141] Each element of the background adversarial filter mask measures the background suppression weight at the corresponding pitch direction and Doppler frequency. The background suppression weight is obtained by statistically averaging the background confidence at each distance cell.
[0142] Construct a weighted clutter subspace projection matrix based on the STAP covariance matrix and the STAP subspace of space-time adaptive processing;
[0143] In Example 1, the weighted clutter subspace projection matrix is calculated as follows: The STAP covariance matrix of the space-time adaptive processing is subjected to inverse square root transformation; energy projection is performed using the basis matrix of the STAP subspace of the space-time adaptive processing; the result is transposed using a conjugate operation; and the weighted clutter subspace projection matrix is obtained through the composite calculation of these three methods.
[0144]
[0145] in, R is the weighted clutter subspace projection matrix. STAP This represents the STAP covariance matrix for space-time adaptive processing. Describe the inverse square root matrix of the STAP covariance matrix for space-time adaptive processing and satisfy the following conditions: U STAP The basis matrix represents the spacetime adaptive processing STAP subspace.
[0146] The background adversarial filtering mask is unfolded into a vector, and a diagonal weight matrix corresponding one-to-one with the pitch-Doppler flattening domain is generated. A joint filter is constructed based on the diagonal weight matrix and the weighted clutter subspace projection matrix. The joint filter outputs a joint filter matrix consistent with the pitch-Doppler flattening domain by subtracting the product of the diagonal weight matrix and the weighted clutter subspace projection matrix from the identity matrix.
[0147] In Example 1, the background adversarial filtering mask is vectorized according to the order of elevation direction and Doppler frequency to obtain a one-dimensional background mask vector. A diagonal weight matrix consistent with the elevation-Doppler flattening domain is generated based on the length of the one-dimensional background mask vector. Each diagonal element of the diagonal weight matrix corresponds to an element in the background mask vector. The weighted clutter subspace projection matrix is calculated using the STAP covariance matrix and the STAP subspace of the space-time adaptive processing. The weighted clutter subspace projection matrix represents the principal component energy projection relationship of radar clutter in the space-time dimension. The elevation-Doppler flattening domain refers to the two-dimensional spectrum corresponding to a range cell in the RD-A data cube. Each diagonal element of the diagonal weight matrix is equal to the corresponding element of the background adversarial filtering mask vector.
[0148] Based on the identity matrix, whose dimension is consistent with the pitch-Doppler flattening domain, the joint filter matrix is obtained by subtracting the product of the diagonal weight matrix and the weighted clutter subspace projection matrix from the identity matrix. The operation means that the joint filter selectively suppresses the energy of the clutter subspace at each position of the pitch-Doppler based on the weights of the background adversarial filter mask.
[0149] Each row and column of the joint filter matrix corresponds one-to-one with a position in the pitch-Doppler flattening domain. When filtering the RD-A data cube, the joint filter matrix achieves adaptive clutter suppression based on spatial location, while taking into account both physical covariance characteristics and the target discrimination capability of the depth model.
[0150] The pitch-Doppler flattening domain of each range cell in the RD-A data cube is flattened into a vector and used as the input of the joint filter. After transformation by the joint filter matrix, the output vector is obtained. The output vector is restored to the pitch-Doppler flattening domain structure and filled back into the position of the corresponding range cell. This process is repeated for all range cells to generate the clutter-suppressed RD-A data cube.
[0151] Example 2:
[0152] During a seven-day continuous low-altitude target monitoring operation, the radar system collected a large number of array coherent echo I / Q signals. In one batch processing, the system detected an abnormal increase in the complex spectral value of an RD-A data point with an elevation of 3.2°, a Doppler frequency of -0.21Hz, and a range cell of 524, over five consecutive batches. However, this area was located near the spectral ridge of the wind farm's main lobe. Traditional STAP methods suppressed the anomalous signal with background spectral energy, and the CFAR detection threshold, exceeding the target amplitude, failed to generate any target alarm. The target modeling branch of this invention... The target embedding at the output RD-A position is significantly higher than that of the background branch. The cross-branch interactive attention weights are concentrated in the target region. After orthogonal decoupling, the cosine similarity between the target features and the background features is 0.03 (the average of the traditional method is 0.41). The background adversarial filter mask gives a high suppression weight of 0.91 at this position. After filtering, the amplitude of the RD-A point is increased by 5.2dB. The target point clearly emerges below the CFAR threshold. The system automatically outputs the suspected drone trajectory. The tracking module successfully associated the target in 5 consecutive batches, verifying the model's practical ability in complex main lobe backgrounds.
[0153] Within the same monitoring period, the system detected periodic signal enhancement in an area with a range of 317 units, an elevation of 5.6°, and a Doppler frequency of 0.03Hz. This area is located at the junction of a forest belt and a coastline, where background clutter is severely affected by wind-blown vegetation and ocean waves. Traditional methods detected only two weak targets out of ten batches, misclassifying the remaining eight as background. After applying this invention, the target modeling branch effectively separated the saliency of small targets by utilizing sparse connections and a multi-scale window mechanism. After interactive orthogonal decoupling, the maximum value of target features increased by 2.7 times, the background suppression weight increased to 0.83, and the residual energy of clutter decreased by 43%. In the detection output, the system successfully identified targets in all ten batches, the trajectory delay decreased by 1.4 seconds, the average target signal-to-noise ratio increased by 7.8dB, and the actual false alarm rate decreased from 0.23 to 0.
[0154] In a test involving scarce samples and environmental changes, the monitoring platform was adjusted from its normal posture to a 30° pitch. The number of on-site training samples was only 15% of the original, and the background main lobe frequency drifted rapidly. The target detection accuracy of the traditional method dropped to 67%, and the false alarm rate increased to 4.6×10-3. After adopting this invention, the training was supplemented by self-supervised pseudo-labels, and the target modeling branch was adapted to the new environment online. After only 80 batches of online fine-tuning, the detection accuracy recovered to 92%, and the false alarm rate decreased to 1.2×10-3. The system automatically triggered the model fine-tuning process in the early stage of platform posture change without manual intervention, and automatically identified and alarmed in 0.92 seconds, which is 63% faster than the traditional method.
[0155] In another multi-target overlap scenario, at distance cell 241, pitch direction 2.3°, and Doppler frequencies of 0.17Hz and -0.18Hz, a micro helicopter and a fixed-wing UAV appeared simultaneously, respectively. The traditional STAP method could only detect the stronger fixed-wing UAV, while the micro helicopter signal was completely submerged by the spectral main lobe clutter. By applying the target modeling branch of the method of this invention, the two targets were focused on different window features. Under the interactive orthogonal mechanism, the cosine similarity of the output feature points of the two targets was as low as 0.08. The target regions were all given high suppression weights by the background adversarial filter mask. Finally, the system simultaneously output two continuous target tracks with a false negative rate of 0. The main lobe of the target signal was clearly separated from the clutter spectrum, and the quantitative detection SNR was improved by more than 6dB.
[0156] In a comparative experiment under limited computing power, the system was set to process the same batch of tasks (256 array elements, 128 distances, 64 pulses). The traditional STAP algorithm took 97ms for total processing time, while the method of this invention took 54ms for block streaming inference and weight quantization post-processing, reducing latency by 44%. Power consumption monitoring showed that the peak power of the system dropped from 25.4W to 14.7W during processing, verifying the superiority of this invention in terms of real-time performance and energy consumption control.
[0157] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for low-altitude radar ground clutter suppression based on machine learning, characterized in that, The method comprises the following steps: receiving an array coherent echo I / Q signal sequence, constructing an RD-A data cube, and generating an RD-A label sequence based on representation learning of the RD-A data cube; estimating a space-time adaptive processing (STAP) covariance matrix and a STAP subspace based on the RD-A data cube, and generating a physical prior mask set in combination with terrain data, platform attitude information, and background Doppler spectrum distribution; inputting the RD-A label sequence and the physical prior mask set into an improved double-branch Transformer encoder to initialize parallel encoding processes of a background modeling branch and a target modeling branch, and obtaining a background embedding representation initial value and a target embedding representation initial value; performing information interaction and feature alignment between the background embedding representation initial value and the target embedding representation initial value by using a cross-branch interaction attention module, obtaining an interaction-enhanced background embedding representation and an interaction-enhanced target embedding representation, and applying an orthogonal constraint term and a mutual information constraint term to the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation to obtain an orthogonally decoupled background embedding representation and an orthogonally decoupled target embedding representation; inputting the orthogonally decoupled background embedding representation and the orthogonally decoupled target embedding representation into a background adversarial filtering module to output a background adversarial filtering mask used for background suppression, fusing the background adversarial filtering mask, the STAP covariance matrix, and the STAP subspace, constructing a joint filter, and filtering the RD-A data cube by using the joint filter to obtain a clutter-suppressed RD-A data cube.
2. The method according to claim 1, wherein, The method comprises the following steps: acquiring an array coherent echo I / Q signal sequence, performing time synchronization processing, frequency correction processing, and channel amplitude and phase consistency calibration on the array coherent echo I / Q signal sequence to obtain a calibrated array coherent echo I / Q signal sequence; performing pulse compression processing on the calibrated array coherent echo I / Q signal sequence to obtain a pulse compression output signal tensor; performing moving target indication processing on the pulse compression output signal tensor in the pulse dimension to generate a moving target-suppressed data tensor, and performing Doppler transformation on the moving target-suppressed data tensor in the pulse dimension to output a Doppler spectrum data tensor; performing array beamforming processing on the Doppler spectrum data tensor in the spatial dimension to obtain an RD-A data cube; mapping the RD-A data cube into an RD-A label sequence, and each label in the RD-A label sequence is composed of a complex spectrum value, a phase, a range coordinate, a Doppler frequency coordinate, and a pitch direction coordinate in the corresponding RD-A data cube.
3. The method of claim 1, wherein the method is based on machine learning. The method comprises the following steps: selecting a training region containing background clutter but not containing a moving target from the RD-A data cube, performing a flattening operation on the RD-A data in the training region, and combining a three-dimensional complex spectrum value tensor into a two-dimensional complex sample matrix according to the pitch direction and the Doppler frequency dimension; The STAP covariance matrix is calculated based on a two-dimensional complex sample matrix, and eigenvalue decomposition is performed on the STAP covariance matrix, and a plurality of main eigenvectors with maximum eigenvalues are selected to form a STAP subspace; A spatial direction mask under each pitch direction and each Doppler frequency is calculated based on platform flight attitude parameters; The background spectrum energy under each pitch direction and Doppler frequency in the RD-A data cube is counted, and the spectrum amplitude square values of all distance units are averaged to obtain a background spectrum intensity atlas; A background spectrum energy threshold is set, and each element in the background spectrum intensity atlas is compared with the background spectrum energy threshold: If the average background energy level under a certain pitch direction and Doppler frequency combination is higher than the background spectrum energy threshold, the element of the corresponding Doppler spectrum intensity mask is assigned a value of 1, otherwise a value of 0 is assigned; The spatial direction mask and the Doppler spectrum intensity mask are logically ANDed at the corresponding positions to obtain a final physical prior mask set; Each element in the physical prior mask set is assigned a value of 1 only when the spatial direction mask and the Doppler spectrum intensity mask are both 1, indicating that the terrain shielding and background spectrum energy conditions are both met under the pitch direction and Doppler frequency combination.
4. The method of claim 1, wherein the method is based on machine learning. The RD-A labeled sequence and the physical prior mask set are input into the improved double-branch Transformer encoder to initialize the parallel encoding process of the background modeling branch and the target modeling branch, and obtain the background embedding representation initial value and the target embedding representation initial value, including: The RD-A labeled sequence obtained by mapping the RD-A data cube is one-to-one matched with the physical prior mask set at the corresponding position, and an input feature alignment matrix is obtained for each RD-A label; The input feature alignment matrix is input into the background modeling branch and the target modeling branch respectively, and the parallel encoding process of the improved double-branch Transformer encoder is started; In the background modeling branch, the input feature alignment matrix is used as the initial input, and the attention guide mask of the physical prior mask set is applied, so that the self-attention only focuses on the RD-A positions marked as background regions by the physical prior mask set. In the encoding process, a multi-level Transformer structure is used to model the input features layer by layer, and the output feature sequence is the background embedding representation initial value; In the target modeling branch, the input feature alignment matrix is used as the initial input, and the relationship between different physical features is captured through the full-connection attention mechanism and the sparse connection strategy. A multi-scale window attention mechanism is used to progressively aggregate the input features at each encoding layer, and the output feature sequence is the target embedding representation initial value.
5. The method of claim 4, wherein the method is based on machine learning. The input feature alignment matrix is taken as the initial input in the background modeling branch, the attention guided mask of the physical prior mask set is applied, the self-attention only focuses on the RD-A positions marked as background regions by the physical prior mask set, a multi-layer hierarchical Transformer structure is adopted in the encoding process, the input features are modeled layer by layer, and the output feature sequence is the background embedding representation initial value, including: The input feature alignment matrix is taken as the initial input in the background modeling branch, the attention guided mask of the physical prior mask set is applied, the self-attention only focuses on the RD-A positions marked as background regions by the physical prior mask set, a multi-layer hierarchical Transformer structure is adopted in the encoding process, the input features are modeled layer by layer, and the output feature sequence is the background embedding representation initial value, including: In each layer of the Transformer encoder, the attention mask matrix is used in the self-attention module, so that each feature vector only performs attention calculation with the feature vectors marked as background regions by the physical prior mask set, and only the background regions are focused on in the background modeling process; A multi-layer hierarchical Transformer structure is adopted for progressive modeling, in each transformer encoder, the feature sequence is sequentially subjected to linear transformation, mask self-attention mechanism, feedforward network activation, normalization operation and residual connection, the output of the previous layer is the input of the next layer, and the output of each layer includes multi-dimensional feature aggregation effect under the mask constraint. In the feature sequence output by the last transformer encoder, the aggregation results of all feature vectors are defined as the background embedding representation initial value; Each feature vector in the background embedding representation initial value corresponds to one RD-A position marked as a background region in the input feature alignment matrix by the physical prior mask set, and the content is the deep background features learned by the position after multi-layer Transformer encoding. All attention calculation, feature aggregation and output vector constraints are strictly limited by the physical prior mask set and only rely on background region information without any target branch features.
6. The method of claim 4, wherein the method is based on machine learning. In the target modeling branch, the input feature alignment matrix is taken as the initial input, the relationship between different physical features is captured through the fully connected attention mechanism and the sparse connection strategy, and the input features are progressively aggregated in each encoding layer by using the multi-scale window attention mechanism. The output feature sequence is the target embedding representation initial value, including: The input feature alignment matrix is taken as the initial input in the background modeling branch, the attention guided mask of the physical prior mask set is applied, the self-attention only focuses on the RD-A positions marked as background regions by the physical prior mask set, a multi-layer hierarchical Transformer structure is adopted in the encoding process, the input features are modeled layer by layer, and the output feature sequence is the background embedding representation initial value, including: In each Transformer encoding layer, the multi-scale window attention mechanism is applied based on the target region weighted aggregation feature representation matrix output by the previous layer, the weighted aggregation feature representation matrix is divided into several local windows according to different scale granularities, and the self-attention calculation is independently performed in each window to output the weighted aggregation result of the target features under the multi-scale window. The cross-scale merging of the weighted aggregation results of the target features of each window at all scales is performed, and the residual connection is performed with the target region features of the original input feature alignment matrix, and the front-end neural network module is input for nonlinear transformation, and the new target region feature weighted aggregation feature representation matrix is output as the input of the next encoding layer, and in all encoding layers, the target region feature weighted aggregation feature representation matrix output by the last layer is defined as the target embedding representation initial value.
7. The method of claim 1, wherein the method is based on machine learning. The cross-branch interaction attention module is used to perform information interaction and feature alignment between the background embedding representation initial value and the target embedding representation initial value, obtain the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation, and apply the orthogonal constraint term and the mutual information constraint term to the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation, to obtain the orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation, including: The cross-branch interaction attention module is used to perform information interaction and feature alignment between the background embedding representation initial value and the target embedding representation initial value, obtain the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation, and apply the orthogonal constraint term and the mutual information constraint term to the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation, to obtain the orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation, including: The cross-branch interaction attention module is used to perform information interaction and feature alignment between the background embedding representation initial value and the target embedding representation initial value, obtain the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation, and apply the orthogonal constraint term and the mutual information constraint term to the interaction-enhanced background embedding representation and the interaction-enhanced target embedding representation, to obtain the orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation, including: The orthogonal constraint term and the mutual information constraint term are combined to form a total optimization objective function, and the orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation are output after joint optimization of the objective function. The orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation are input into the background countermeasure filtering module, and a background countermeasure filtering mask for background suppression is output. The background countermeasure filtering mask, the space-time adaptive processing (STAP) covariance matrix, and the space-time adaptive processing (STAP) subspace are fused to construct a joint filter, and the joint filter is used to filter the RD-A data cube to obtain a clutter-suppressed RD-A data cube, including:
8. The method of claim 1, wherein the method is based on machine learning. In the background countermeasure filtering module, the orthogonal decoupled background embedding representation and the orthogonal decoupled target embedding representation corresponding to each RD-A label are vector-spliced through background discrimination mapping, and the confidence that each RD-A label is background is output by the background discrimination mapping. The background confidences of all RD-A labels are sequentially arranged to form a background confidence sequence. The background confidence sequence is grouped according to the pitch direction and the Doppler frequency, and the background confidences of all distance cells in the same pitch direction and the same Doppler frequency are averaged to obtain a background countermeasure filtering mask defined on the pitch-Doppler plane. The weighted clutter subspace projection matrix is constructed based on a space-time adaptive processing (STAP) covariance matrix and a space-time adaptive processing (STAP) subspace; The background countermeasure filter mask is unfolded into a vector, and a diagonal weight matrix corresponding to the pitch-Doppler flat field is generated; a joint filter is constructed based on the diagonal weight matrix and the weighted clutter subspace projection matrix; the joint filter outputs a joint filter matrix consistent with the pitch-Doppler flat field by subtracting the product of the unit matrix and the diagonal weight matrix and the weighted clutter subspace projection matrix; The pitch-Doppler flat field on each distance unit in the RD-A data cube is unfolded into a vector as the input of the joint filter; after the joint filter matrix transformation, an output vector is obtained; the output vector is restored to the pitch-Doppler flat field structure and filled back to the corresponding distance unit position; all distance units repeat the process to generate the RD-A data cube after clutter suppression.
Citation Information
Patent Citations
Airborne radar non-uniform clutter suppression method based on deep learning
CN112612006A
Millimeter wave radar target detection method based on Transform
CN118506147A