A three-channel time-frequency and hard clutter assisted low, slow and small target recognition method

CN122776201APending Publication Date: 2026-09-18NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004137.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

仅依靠主类别交叉熵训练时,大量易分类样本主导梯度,普通杂波、硬杂波和目标之间的边界约束不足

Benefits of technology

(1)通过三通道时频结构输入动态构造,使网络能够同时利用能量分布、背景残差和梯度结构三种互补信息。第一通道保留整体能量分布和主谱位置,第二通道通过扣除各频率行的时间中值削弱稳定背景并突出局部微动扰动,第三通道将能量变化转换为梯度响应以强化边缘、条纹和纹理变化。三通道由同一幅分贝时频图同源并行生成,保持相同的时间-频率坐标并逐像素对齐,形成能量定位、扰动显化和结构刻画的互补表征链条,提高了目标与杂波之间的可分性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122776201A_ABST
    Figure CN122776201A_ABST
Patent Text Reader

Abstract

The application discloses a three-channel time-frequency and hard clutter auxiliary low, slow and small target recognition method. The method comprises the following steps: power synthesis, decibel conversion and quantile normalization are performed on radar complex time-frequency samples to obtain a normalized decibel time-frequency graph; three channels of original energy, background residual and gradient structure are generated in parallel and stacked to form a three-channel input; an auxiliary label containing three categories of ordinary clutter, hard clutter and targets is established; a recognition network with a shared feature extraction module, a main classification head and an auxiliary supervision head is constructed; the network is trained by using a weighted total loss of the main classification loss and the auxiliary classification loss; the optimal model is selected according to the comprehensive evaluation score of the target recognition rate, the false alarm rate and the missed alarm rate; and the three-channel input is constructed for the to-be-tested sample and the recognition result is output by the optimal model. The application effectively improves the low, slow and small target recognition rate and suppresses false alarms and missed alarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar target recognition technology, specifically relating to a three-channel time-frequency and hard clutter-assisted method for recognizing low, slow, and small targets. Background Technology

[0002] Low-altitude, slow-moving targets with small radar cross-sections are crucial for identification in airport airspace clearance, key area protection, and low-altitude security monitoring. Radar can continuously acquire target echoes day and night, and under complex weather conditions. Micro-Doppler modulation generated by rotors, vibrating components, and local oscillations can form sidebands, stripes, periodic scintillation, and local textures on the time-frequency map, serving as important evidence for distinguishing UAVs, birds, vehicles, personnel, and environmental clutter. As the complexity of the measured scenario increases, stable backgrounds, ground reflections, weather echoes, and hard clutter containing local micro-motions can overlap with target characteristics, leading to false alarms and missed alarms in the model. Existing methods typically transform complex radar echoes into single-channel decibel time-frequency maps before inputting them into convolutional neural networks for classification. Single-channel time-frequency maps primarily retain energy intensity, and the model easily relies on the brightness of the main spectral line or the overall energy distribution. When the target energy is weak, the background is strong, or the acquisition conditions change, local micro-motions and edge structures are easily submerged. Continuous convolution and pooling may further lose fine sidebands and short-term textures, causing weak targets to be identified as clutter. Another type of improvement method enhances feature extraction capabilities by deepening the network, changing the backbone, or adding attention modules. This type of method increases representational capacity but does not directly change the separability of the input information, nor does it guarantee a reduction in false alarms. Experiments show that deeper networks may continue to amplify high-energy edges in clutter. If the attention insertion points are too shallow or too numerous, clutter textures may be mistaken for target evidence. When training solely based on the main class cross-entropy, a large number of easily classifiable samples dominate the gradient, resulting in insufficient boundary constraints between ordinary clutter, hard clutter, and the target.

[0003] Therefore, providing a method for identifying slow, small targets that can enhance the time-frequency structure representation of the target at the input level, explicitly constrain the boundary between the target and clutter features during training, and effectively suppress false alarms has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] The technical problem to be solved by this invention is that existing single-channel time-frequency map recognition methods tend to rely on the brightness of the main spectral line. When the target energy is weak, the background is strong, or the acquisition conditions change, local micro-motions and edge structures are easily submerged. When training only by relying on the main category cross-entropy, the boundary constraints between ordinary clutter, hard clutter and the target are insufficient, resulting in a high false alarm rate.

[0005] To address the aforementioned technical problems, this invention provides a three-channel time-frequency and hard clutter-assisted method for identifying low-speed, small targets, comprising the following steps: Step S100: Obtain radar complex time-frequency samples, perform power synthesis and decibel conversion to obtain a decibel time-frequency diagram; Step S200: Perform quantile normalization on the decibel time-frequency graph to obtain a normalized decibel time-frequency graph; Step S300: Construct a three-channel time-frequency input from the normalized decibel time-frequency map, including the original energy channel, background residual channel and gradient structure channel, and stack them to form a three-channel input; Step S400: Establish auxiliary labels to classify samples into three categories: ordinary clutter, hard clutter, and target. Ordinary clutter and hard clutter both belong to the clutter category in the main classification, and they are only distinguished by auxiliary labels. Hard clutter is determined by historical false alarm records, manually reviewed samples, sample source markings, or continuous high-confidence false alarms on independently selected sample sets by pre-trained models. The auxiliary labels for target samples are uniformly set as target class. Step S500: Construct a recognition network, which includes a shared feature extraction module, a main classification head, and an auxiliary supervision head, with three-channel input as the network input; Step S600: Using the three-channel input and auxiliary labels, train the recognition network with the total loss formed by the weighted sum of the main classification loss and the auxiliary classification loss; Step S700: Select the optimal model during the training process based on the comprehensive evaluation score composed of target recognition rate, false alarm rate, and missed alarm rate; Step S800: For the sample to be tested, construct its three-channel input according to steps S100 to S300, input the optimal model, and output the recognition result from the main classification head.

[0006] Furthermore, in step S100, the formulas for power combining and decibel conversion are as follows: ; In the formula, Z is the complex Doppler time-frequency matrix, Re(Z) and Im(Z) are the real and imaginary parts of the complex matrix Z, respectively, and ε is a preset positive number; the decibel time-frequency diagram D is scaled to a preset size H×W.

[0007] Furthermore, in step S200, the formula for quantile normalization is: ; In the formula, Let A be the normalization function, and A be the matrix to be normalized. Restrict the elements in A to the 1st percentile. to Within the interval, δ is a preset positive number used to prevent the denominator from being zero. Quantile truncation reduces the occupancy of the dynamic range by extreme strong scattering points, giving different batches of samples a more consistent numerical scale.

[0008] Furthermore, step S300 further includes: Step S310: Generate the first channel , N is the normalization function defined in step S200; Step S320: Generate the second channel , , This indicates that the median of each row of the decibel time-frequency plot D is calculated along the time dimension; Step S330: Generate the third channel , ,, For the first channel The gradient obtained by performing horizontal convolution with the Sobel operator. To The gradient obtained by performing Sobel convolution in the vertical direction; Step S340: Stack to obtain three-channel input .

[0009] Furthermore, step S500 further includes: Step S510: The shared feature extraction module contains four levels of convolutional units with 32, 64, 128 and 256 channels respectively. After the third and fourth level convolutional units, there is a channel-spatial joint attention module, which forms the shared feature h after flattening and full connection. Step S520: Channel attention to input F Calculate channel weighted features Where AvgPool and MaxPool are global average and max pooling, respectively, MLP is a shared multilayer perceptron, and σ is the Sigmoid function; Step S530: Spatial attention pair according to Computational spatial weighted features ,in , Avgc represents the average operation along the channel dimension, and Maxc represents the maximum operation along the channel dimension. It is a 7×7 convolution; Step S540: The main classification head uses h as input and outputs K original class predictions, and the auxiliary supervision head uses h as input and outputs predictions of three classes: ordinary clutter, hard clutter, and target.

[0010] Furthermore, the total loss function L in step S600 is: ; In the formula, B is the batch size, and CE is the cross-entropy loss. and These are the main category prediction and the true label, respectively. and λ represents the auxiliary prediction and auxiliary label, respectively, and λ is the auxiliary loss weight.

[0011] Furthermore, λ is set to 0.3.

[0012] Furthermore, the comprehensive evaluation score S in step S700 is calculated using the following formula: ; In the formula, R, F, and M are the target recognition rate, false alarm rate, and false alarm rate, respectively, α and β are weight coefficients, and the model is saved when S is increased.

[0013] Furthermore, α is set to 1.5 and β to 0.5.

[0014] Furthermore, in step S100 Pick In step S200, δ is a positive number; in step S700, the AdamW optimizer is used for training, with an initial learning rate of 0.001, weight decay of 0.0001, learning rate decay factor of 0.5, scheduling patience rounds of 2, maximum training rounds of 100, and early stopping patience rounds of 5.

[0015] Compared with the prior art, the present invention has the following advantages: (1) By dynamically constructing a three-channel time-frequency structure input, the network can simultaneously utilize three complementary information types: energy distribution, background residuals, and gradient structure. The first channel retains the overall energy distribution and main spectrum position, the second channel weakens the stable background and highlights local micro-movements by subtracting the time median of each frequency row, and the third channel converts energy changes into gradient responses to enhance edge, stripe, and texture changes. The three channels are generated in parallel from the same decibel time-frequency map, maintaining the same time-frequency coordinates and aligning pixel by pixel, forming a complementary representation chain of energy localization, perturbation manifestation, and structure characterization, which improves the separability between the target and clutter.

[0016] (2) The boundaries between targets and clutter with shared features are explicitly constrained through three types of auxiliary supervision: ordinary clutter, hard clutter, and targets. Hard clutter is determined by historical false alarm records, manually reviewed samples, sample source labels, or continuous high-confidence false alarms of pre-trained models on independently selected sample sets. The auxiliary loss participates in the joint optimization of the total loss but does not replace the main classification task, so that the network can identify clutter structures that are prone to triggering false alarms while learning the original category, thereby effectively suppressing false alarms.

[0017] (3) By setting a channel-spatial joint attention module after the third and fourth level convolutional units of the shared feature extraction module, the high-level time-frequency features are recalibrated. Compared with stacking attention modules in shallow layers or simply increasing the network depth, this can highlight local structures with class significance with lower additional overhead and improve feature discrimination ability.

[0018] (4) The model checkpoint is saved and early stop is controlled by using the target recognition rate, false alarm rate and false alarm rate to form a comprehensive evaluation score. This avoids false alarm anomalies caused by saving only the verification loss, and ensures that the selected model can effectively control false alarms and false alarms while improving the recognition rate, thereby improving the reliability of the project deployment.

[0019] The present invention will now be further described with reference to the accompanying drawings. Attached Figure Description

[0020] Figure 1 This is a flowchart of the overall process for identifying small, slow targets according to the present invention.

[0021] Figure 2 This is a schematic diagram of the three-channel time-frequency input structure.

[0022] Figure 3 A comparison of the three-channel time-frequency characterizations of the target sample and the clutter sample.

[0023] Figure 4 This is a schematic diagram of a network structure with shared features and dual output heads. Detailed Implementation

[0024] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0025] This invention provides a three-channel time-frequency and hard clutter-assisted method for identifying low-altitude, slow-moving, small targets, used in complex clutter environments. (Refer to...) Figure 1 The method includes the following steps: Step S100: Obtain radar complex time-frequency samples, perform power synthesis on the complex Doppler time-frequency matrix in the samples and convert it into a decibel time-frequency map, and scale the decibel time-frequency map to a preset size.

[0026] The complex Doppler time-frequency matrix of the radar complex time-frequency samples is denoted as Z. Its power is synthesized and converted into a decibel time-frequency graph D. ; In the formula, Re(Z) and Im(Z) are the real and imaginary parts of the complex matrix Z, respectively, and ε is a preset positive number used to prevent the value from being undefined. In one embodiment, ε takes... The decibel time-frequency graph D is scaled to a preset size H×W. In one embodiment, both H and W are 128.

[0027] Step S200: Perform quantile normalization processing on the decibel time-frequency graph, calculate the low quantile value and high quantile value of the decibel time-frequency graph, truncate the pixel values ​​outside the interval and linearly map them to the preset interval to obtain the normalized decibel time-frequency graph.

[0028] Calculate the lower quantile of the decibel time-frequency plot D. and high quantile Pixels outside the specified range are truncated and linearly mapped to [0,1]: ; In the formula, This means restricting the element values ​​in matrix A to a certain range. to Within the interval, less than The setting is greater than The setting is δ is a preset positive number used to prevent the denominator from being zero. Quantile truncation reduces the occupancy of the dynamic range by extreme strong scattering points, giving different batches of samples a more consistent numerical scale.

[0029] Step S300: Based on the normalized decibel time-frequency map, dynamically construct a three-channel time-frequency structure input. The three channels are generated in parallel from the same decibel time-frequency map D, maintaining the same time-frequency coordinates and aligning pixel by pixel.

[0030] Reference Figure 2 The three-channel time-frequency input structure is constructed as follows: Step S310: Generate the first channel First Channel The original decibel time-frequency plot after normalization: ; In the formula, N is the normalization function defined in step S200, and D is the decibel time-frequency diagram. First channel Preserve the overall energy distribution and main spectrum position.

[0031] Step S320: Generate the second channel Second Channel Background residual plot for each frequency row after subtracting the median time value of that row: ; In the formula, This indicates that the median of each row in the decibel time-frequency plot D is calculated along the time dimension. Second channel. By subtracting the time median of each frequency row, the background components that are stable over time are weakened, making the periodic fluctuations, intermittent sidebands, and local micro-movements around the main spectrum more prominent.

[0032] Step S330: Generate the third channel Third Channel In the first channel The above calculates the combined magnitude of the Sobel horizontal and vertical gradients: ; In the formula, For the first channel The gradient obtained by performing horizontal convolution with the Sobel operator. For the first channel The gradient obtained by performing Sobel convolution in the vertical direction. (Third channel) Energy changes are converted into gradient responses, which further reveal stripe boundaries, contour directions, and texture variations.

[0033] Step S340: Stack the three channels to form a three-channel input X: ; X is the three-channel time-frequency structure input, stack( ) indicates a stacking operation along the channel dimension. The three channels form a complementary characterization chain for energy localization, perturbation manifestation, and structural characterization.

[0034] Reference Figure 3 Target samples typically exhibit a relatively stable main spectral region and local structural variations, while clutter samples may display scattered textures, strong local backgrounds, or irregular edge responses. (First channel) Provides overall energy evidence, second channel Highlighting the changes in the relative background, the third channel Further characterizing stripe boundaries and texture extension, the three types of information remain complementary, improving the separability between the target and clutter.

[0035] Step S400: Establish auxiliary labels to classify samples into three categories: ordinary clutter, hard clutter, and target. Ordinary clutter and hard clutter both belong to the clutter category in the main classification, and they are only distinguished by auxiliary labels. Hard clutter is determined by historical false alarm records, manually reviewed samples, sample source markings, or continuous high-confidence false alarms of pre-trained models on independently selected sample sets. The auxiliary labels for target samples are uniformly set as target class.

[0036] auxiliary tags Defined as: ; In the formula The auxiliary label for the i-th sample is denoted as 0, corresponding to ordinary clutter, 1 to hard clutter, and 2 to the target. Hard clutter is determined by historical false alarm records, manually reviewed samples, sample source markers, or continuous high-confidence false alarms from a pre-trained model on an independently selected sample set. The auxiliary label for target samples is uniformly set as the target class. The main classification label and auxiliary labels are independent of each other. The main classification head learns the original class, and the auxiliary supervision head learns the clutter difficulty attribute, constraining the boundary between the target and clutter sharing features without changing the original class definition.

[0037] Step S500: Construct a recognition network comprising a shared feature extraction module, a main classification head, and an auxiliary supervision head. (Refer to...) Figure 4 Step S500 further includes: Step S510: Construct the shared feature extraction module. The shared feature extraction module consists of four levels of convolutional units. Each level of convolutional unit contains a convolutional layer, a batch normalization layer, an activation layer, and a downsampling layer. The number of channels in the four levels of convolutional units are 32, 64, 128, and 256, respectively. Channel-spatial joint attention modules are set after the output features of the third and fourth level convolutional units, respectively. The features output by the shared feature extraction module are flattened and then passed through a 512-dimensional fully connected layer and a random deactivation layer to form a shared feature h, where h is the shared feature vector.

[0038] Step S520: Construct the channel attention module. The channel attention module calculates the channel-weighted features of the input feature F using the following formula. : ; In the formula, F is the input feature map. denoted as channel-weighted feature map, AvgPool is the global average pooling operation, MaxPool is the global max pooling operation, MLP is a multilayer perceptron with shared weights, σ ​​is the sigmoid activation function, and ⊙ is the element-wise multiplication operation.

[0039] Step S530: Construct the spatial attention module. The spatial attention module applies channel-weighted features. Calculate the spatial weighted features using the following formula : ; In the formula, For channel-weighted feature maps, For spatially weighted feature maps, Averaging along the channel dimension To represent the maximum operation along the channel dimension, the two items within the square brackets are separated by a semicolon, indicating that the two results are concatenated along the channel dimension. This represents a convolution operation with a kernel size of 7×7, where σ is the Sigmoid activation function and ⊙ represents element-wise multiplication.

[0040] Step S540: Construct the main classification head and the auxiliary supervision head. The main classification head takes the shared feature h as input and outputs prediction results for K original categories, where K is the number of target categories. The auxiliary supervision head takes the shared feature h as input and outputs prediction results for three categories: ordinary clutter, hard clutter, and target.

[0041] Step S600: Train the recognition network using a joint optimization method of main classification loss and auxiliary classification loss.

[0042] For a training batch containing B samples, the total loss function is: ; In the formula, L is the total loss function, B is the number of samples in the training batch, and CE represents the cross-entropy loss function. Predict the probability of the main classifier for the i-th sample. Let i be the true class label of the i-th sample. Let the auxiliary supervisor predict the probability for the i-th sample. The auxiliary label λ for the i-th sample is the auxiliary loss weight coefficient. In one implementation, λ is set to 0.3.

[0043] Step S700: Select the optimal model during the training process using a multi-index joint model selection criterion.

[0044] During training, a comprehensive evaluation score S is used to save model checkpoints and control early stopping. ; In the formula, S is the comprehensive evaluation score, R is the target recognition rate, F is the false alarm rate, M is the missed alarm rate, and α and β are weighting coefficients. In one implementation, α is 1.5 and β is 0.5. The model is saved when the verification score improves; otherwise, the early stopping count is accumulated.

[0045] During training, the AdamW optimizer is used to update parameters with an initial learning rate of 0.001 and a weight decay factor of 0.0001. The learning rate is reduced when the validation loss fails to improve for several consecutive rounds. In one implementation, the learning rate decay factor is 0.5, the scheduling patience rounds are 2, the maximum number of training rounds is 100, and the early stopping patience rounds are 5.

[0046] Step S800: In the sample identification stage, a three-channel time-frequency structure input is constructed according to steps S100 to S300. The category probability is output by the shared feature extraction module and the main classification head, and the category with the highest probability is taken as the identification result.

[0047] The auxiliary supervisory head can simultaneously output hard clutter probabilities for subsequent false alarm analysis and sample library maintenance. In deployment scenarios where only category output is required, only the main classification head results can be used.

[0048] Example 1 uses an anonymized radar time-frequency dataset A as the experimental object. The dataset contains target and various clutter samples. Five fixed random seeds are used for training, and the method of this invention is compared with a benchmark single-channel model. The benchmark single-channel model uses a single-channel decibel time-frequency plot as input, and the remaining network structure is the same as that of this invention. Test results show that the benchmark single-channel model has a target recognition rate of 97.73%, a false alarm rate of 6.97%, and a missed alarm rate of 2.27%; the method of this invention has a target recognition rate of 98.72%, a false alarm rate of 2.59%, and a missed alarm rate of 1.28%. The method of this invention improves the target recognition rate by 0.99 percentage points, reduces the false alarm rate by 4.38 percentage points, and reduces the missed alarm rate by 0.99 percentage points.

[0049] Example 2: To verify the applicability across datasets, the method of this invention was applied to generalized datasets B, C, and D. On generalized dataset B, the false alarm rate decreased from 5.53% of the baseline single-channel model to 3.05%; on generalized dataset C, the false alarm rate decreased from 2.03% to 1.27%; and on generalized dataset D, the target recognition rate increased from 95.07% to 97.04%, while the false alarm rate decreased from 4.93% to 2.96%. The results show that the method of this invention can enhance the time-frequency structure representation of the target and improve target recognition or false alarm control performance under different acquisition conditions.

[0050] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural modifications made based on the concept of the present invention and the content of this specification, or any direct or indirect applications in other related technical fields, are included within the scope of patent protection of the present invention.

Claims

1. A three-channel time-frequency and hard clutter-assisted method for identifying low-speed, small targets, characterized in that, Includes the following steps: Step S100: Obtain radar complex time-frequency samples, perform power synthesis and decibel conversion to obtain a decibel time-frequency diagram; Step S2 00: Perform quantile normalization on the decibel time-frequency graph to obtain the normalized decibel time-frequency graph; Step S300: Construct a three-channel time-frequency input from the normalized decibel time-frequency map, including the original energy channel, background residual channel and gradient structure channel, and stack them to form a three-channel input; Step S400: Establish auxiliary labels to classify samples into three categories: ordinary clutter, hard clutter, and target. Ordinary clutter and hard clutter both belong to the clutter category in the main classification, and they are only distinguished by auxiliary labels. Hard clutter is determined by historical false alarm records, manually reviewed samples, sample source markings, or continuous high-confidence false alarms on independently selected sample sets by pre-trained models. The auxiliary labels for target samples are uniformly set as target class. Step S500: Construct a recognition network, which includes a shared feature extraction module, a main classification head, and an auxiliary supervision head, with three-channel input as the network input; Step S600: Using the three-channel input and auxiliary labels, train the recognition network with the total loss formed by the weighted sum of the main classification loss and the auxiliary classification loss; Step S700: Select the optimal model during the training process based on the comprehensive evaluation score composed of target recognition rate, false alarm rate, and missed alarm rate; Step S800: For the sample to be tested, construct its three-channel input according to steps S100 to S300, input the optimal model, and output the recognition result from the main classification head.

2. The method for identifying low-speed, small targets according to claim 1, characterized in that, In step S100, the formulas for power combining and decibel conversion are as follows: ; In the formula, Z is the complex Doppler time-frequency matrix, Re(Z) and Im(Z) are the real and imaginary parts of the complex matrix Z, respectively, and ε is a preset positive number; the decibel time-frequency diagram D is scaled to a preset size H×W.

3. The method for recognizing low-speed, small targets according to claim 1, characterized in that, In step S200, the formula for quantile normalization is: ; In the formula, Let A be the normalization function, and A be the matrix to be normalized. Restrict the elements in A to the 1st percentile. to Within the interval, δ is a preset positive number used to prevent the denominator from being zero. Quantile truncation reduces the occupancy of the dynamic range by extreme strong scattering points, giving different batches of samples a more consistent numerical scale.

4. The method for recognizing small, slow targets according to claim 1, characterized in that, Step S300 further includes: Step S310: Generate the first channel , N is the normalization function defined in step S200; Step S320: Generate the second channel , , This indicates that the median of each row of the decibel time-frequency plot D is calculated along the time dimension; Step S330: Generate the third channel , , For the first channel The gradient obtained by performing horizontal convolution with the Sobel operator. To The gradient obtained by performing Sobel convolution in the vertical direction; Step S340: Stack to obtain three-channel input .

5. The method for identifying low-speed, small targets according to claim 1, characterized in that, Step S500 further includes: Step S510: The shared feature extraction module contains four levels of convolutional units with 32, 64, 128 and 256 channels respectively. After the third and fourth level convolutional units, there is a channel-spatial joint attention module, which forms the shared feature h after flattening and full connection. Step S520: Channel attention to input F Calculate channel weighted features Where AvgPool and MaxPool are global average and max pooling, respectively, MLP is a shared multilayer perceptron, and σ is the Sigmoid function; Step S530: Spatial attention pair according to Computational spatial weighted features ,in , Avgc represents the average operation along the channel dimension, and Maxc represents the maximum operation along the channel dimension. It is a 7×7 convolution; Step S540: The main classification head uses h as input and outputs K original class predictions, and the auxiliary supervision head uses h as input and outputs predictions of three classes: ordinary clutter, hard clutter, and target.

6. The method for identifying low-speed, small targets according to claim 1, characterized in that, The total loss function L in step S600 is: ; In the formula, B is the batch size, and CE is the cross-entropy loss. and These are the main category prediction and the true label, respectively. and λ represents the auxiliary prediction and auxiliary label, respectively, and λ is the auxiliary loss weight.

7. The method for identifying low-speed, small targets according to claim 6, characterized in that, λ is set to 0.

3.

8. The method for identifying low-speed, small targets according to claim 1, characterized in that, In step S700, the comprehensive evaluation score S is calculated using the following formula: ; In the formula, R, F, and M are the target recognition rate, false alarm rate, and false alarm rate, respectively, α and β are weight coefficients, and the model is saved when S is increased.

9. The method for identifying small, slow targets according to claim 8, characterized in that, α is set to 1.5, and β is set to 0.

5.

10. The method for recognizing small, slow targets according to claim 1, characterized in that, In step S100 Pick In step S200, δ is a positive number; in step S700, the AdamW optimizer is used for training, with an initial learning rate of 0.001, weight decay of 0.0001, learning rate decay factor of 0.5, scheduling patience rounds of 2, maximum training rounds of 100, and early stopping patience rounds of 5.