SAR long-tail target detection method based on double-drive equalization loss and implicit feature enhancement
Patent Information
- Application Number
- CN202311691468.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-11
AI Technical Summary
[0004]本发明的目的在于针对上述现有技术的不足,提出一种基于双驱动均衡损失及隐性特征增强的SAR长尾目标检测方法,用于解决现有技术存在的SAR长尾检测中尾部类别特征数量稀少,检测精度低的问题
[0018]第一,本发明通过SAR长尾目标检测模型中的隐性特征增强单元,有效增加了尾部类别样本的特征数量,丰富了各类样本的特征集,克服了现有技术中无法有效增加尾部类别样本特征的问题,使得本发明在面对样本数稀少的尾部类别时也可得到充分训练,从而提升了尾部类别目标的检测精度。
Smart Images

Figure CN117710783B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of radar technology, and more specifically relates to a long-tail target detection method for Synthetic Aperture Radar (SAR) based on dual-drive equalization loss and latent feature enhancement in the field of radar remote sensing detection technology. This invention can be used for detecting tail-type targets in long-tail SAR datasets in complex scenarios. Background Technology
[0002] Synthetic Aperture Radar (SAR) can achieve all-weather, long-range, high-resolution imaging of targets on land and sea, regardless of lighting or weather conditions. Due to the complex imaging environment, targets in the obtained SAR images are often located in complex scenes, making target detection in SAR images difficult. Furthermore, SAR image datasets often exhibit a long-tailed distribution due to observational limitations. With the application of deep learning-based target detection and recognition methods to SAR image interpretation, significant progress has been made in SAR target detection and recognition. Currently, most mainstream SAR target detection and recognition methods are based on general target detection algorithms, such as single-stage algorithms SSD and YOLO, and two-stage algorithms Faster R-CNN and Cascade R-CNN. These algorithms, based on large-scale deep networks, have good detection and recognition performance when training samples are sufficient. However, for certain tail-type targets, the limited number of samples leads to insufficient model training, resulting in low detection accuracy. Currently, the field of computer vision has proposed methods such as class balancing and detection and recognition model improvement to address the long-tailed distribution problem. Long-tailed target detection and recognition aims to construct a reasonable class balancing method, which strengthens the learning of tail-type samples to enable correct detection and recognition of tail-type targets.
[0003] Zhengzhou University in Henan Province disclosed a method for detecting long-tailed targets in its patent application, "A Long-Tail Target Detection Method Based on Deep Learning." The method's implementation consists of four parts: First, acquiring an image dataset: obtaining an image dataset conforming to a long-tailed distribution, dividing the dataset into training and test sets; second, preprocessing the dataset: calculating the number of valid samples for each category; third, outputting the logit: training a pre-trained model on the training set to obtain the logit output by the trained network; fourth, filtering semantically similar categories: setting a threshold on the network's output logit, suppressing only tail categories semantically similar to the head category, thus increasing the network's attention to tail categories; and finally, adaptively adjusting the suppression gradient between semantically similar categories based on the network's output logit to enhance the differentiation of tail categories. While this method can improve the detection accuracy of tail categories in long-tail detection and recognition, it still has the following shortcomings: 1) Because it only addresses the long-tail problem from the perspective of loss function weighting, it does not increase the number of features for tail category samples. Therefore, it cannot effectively extract features for tail categories, leading to insufficient tail class learning and affecting subsequent detection and recognition tasks for tail category targets. 2) Since this method only reweights the loss from the single perspective of the output probability threshold, it results in low detection accuracy and poor robustness for tail category samples that are difficult to detect in complex scenarios. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the existing technology by proposing a SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement, which solves the problem of the scarcity of tail category features and low detection accuracy in the existing SAR long-tail detection technology.
[0005] The specific approach to achieving the objective of this invention is as follows: First, the preprocessed image is input into the network for training to obtain the model parameters of the network classification layer and the output features of the target samples. Using these model parameters and sample features, the intra-class covariance of different samples is estimated online, thereby obtaining the effective semantic directions for different classes. Then, the sample features are implicitly expanded based on the effective semantic directions to obtain an augmented feature set to enrich the features of the tail class samples. Finally, a reweighted loss driven by both class and sample is constructed to comprehensively balance the losses of different classes from multiple perspectives: for class reweighting, the loss is reweighted based on class gradient, class difficulty, and class accuracy; for sample reweighting, the predicted probabilities of different classes are reweighted based on false detection rate and class relevance. Simultaneously, the SmoothL1 loss is combined with the network for classification and regression tasks respectively, thereby obtaining the final detection results.
[0006] The specific implementation steps of this invention are as follows:
[0007] Step 1, Generate training set:
[0008] A sample set is formed by selecting N SAR images containing M types of aircraft targets. The number of tail-type targets in the sample set is at least equal to the number of head-type targets. Where M≥7, N≥2000; the location and category of the target in each image in the sample set are labeled, and the sample set composed of the labeled images is used as the training set;
[0009] Step 2: Construct a SAR long-tail target detection model consisting of a feature extraction network, a feature fusion network, a region proposal generation network, an ROI network, and a classification and regression network with implicit feature enhancement, and connect them sequentially. The classification and regression network with implicit feature enhancement is composed of a 3×3 convolutional layer, an implicit feature enhancement unit, and a classification and regression sub-network connected sequentially. The implicit feature enhancement unit is composed of a class covariance calculation layer and an implicit feature estimation layer connected in series. The classification and regression sub-network is composed of two 1×1 convolutional layers connected in parallel.
[0010] Step 3, construct the dual-drive equalization loss function as follows:
[0011]
[0012] Where L represents the dual-drive equalization loss function, I represents the total number of samples in each training batch, C represents the total number of classes in the training set, and ∑ represents the summation operation. w represents the j-th element of the one-hot vector converted from the true label of the i-th sample. j,CLS This represents the class-based reweighting factor for the j-th category. This represents the probability that the i-th sample is predicted to be of the j-th class. Represents the sample-based reweighting factor for the i-th sample;
[0013] Step 4, train the SAR long-tail target detection model:
[0014] The training set is input into the SAR long-tail target detection model in batches. The gradient descent method is used to update the network parameters during backpropagation until the SmoothL1 loss function and the dual-drive equalization loss function converge, thus obtaining the trained SAR long-tail target detection model.
[0015] Step 5, detect targets in the SAR long-tail distribution image:
[0016] The SAR long-tail image to be detected is input into the trained model, and the target detection result is output.
[0017] Compared with the prior art, the present invention has the following advantages:
[0018] First, this invention effectively increases the number of features of tail category samples by using the latent feature enhancement unit in the SAR long-tail target detection model, enriching the feature set of various types of samples. This overcomes the problem in the prior art that it is impossible to effectively increase the features of tail category samples, enabling this invention to be fully trained even when facing tail categories with a small number of samples, thereby improving the detection accuracy of tail category targets.
[0019] Second, this invention performs class reweighting and sample reweighting from the perspectives of category and sample, respectively, comprehensively considering various imbalanced states of categories and samples under long-tail detection. This overcomes the problems of single weighting method and poor detection results of tail category samples in the prior art, enabling this invention to have good detection results for tail category samples in complex scenarios. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the implementation of the present invention;
[0021] Figure 2 This is a diagram of the SAR target detection model based on dual-drive equalization loss and latent feature enhancement in this invention.
[0022] Figure 3 This is a comparison chart of the detection and identification results of the present invention and existing technical solutions. Detailed Implementation
[0023] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0024] Combined with appendix Figure 1 The steps for implementing the embodiments of the present invention will be described in further detail below.
[0025] Step 1: Generate the training set.
[0026] In this embodiment of the invention, a large number of SAR images containing aircraft targets were collected, totaling 2000 images, with the image sizes including 600×600, 1024×1024, and 2048×2048.
[0027] The sample set contains targets of seven aircraft types: A220, A320-321, A330, ARJ21, Boeing 737-800, Boeing 787, and others. The location and category of the targets in each SAR image are labeled. The SAR images are from the onboard radar of the Gaofen-3 satellite, and these SAR images are combined into the sample set.
[0028] The sample set was randomly divided proportionally, with 1400 images assigned to the training set and the remaining 600 images assigned to the test set.
[0029] Step 2: Construct a SAR long-tail target detection model.
[0030] A SAR long-tail target detection model is constructed by sequentially connecting a feature extraction network, a feature fusion network, a region proposal generation network, an ROI network, and a classification and regression network with implicit feature enhancement.
[0031] The classification and regression network with implicit feature enhancement consists of a 3×3 convolutional layer, an implicit feature enhancement unit, and a classification and regression subnetwork connected in sequence.
[0032] The latent feature enhancement unit is composed of a class covariance calculation layer and a latent feature estimation layer connected in series; the classification and regression sub-network is composed of two 1×1 convolutional layers connected in parallel.
[0033] The feature extraction network consists of a 3×3 convolutional layer, a first-scale feature extraction layer, a second-scale feature extraction layer, a third-scale feature extraction layer, and a fourth-scale feature extraction layer connected in series. The feature extraction layers at the first to fourth scales are composed of 3, 4, 6, and 3 identical Bottleneck blocks, respectively. Each Bottleneck block is composed of a convolutional block with residual structure and a feature fusion convolutional layer connected in series. The convolutional block with residual structure includes two branches: the first branch consists of a 3×3 convolutional layer, and the second branch consists of two 3×3 convolutional layers connected in series. The feature fusion convolutional layer consists of a 3×3 convolutional layer. The outputs of the first and second branches are added together and then input into the feature fusion convolutional layer. For the feature extraction layers at the first to fourth scales, the stride of the first convolutional layer of the last Bottleneck's two branches is 2, and the stride of all other convolutional layers is 1.
[0034] The feature fusion network upsamples the outputs of the four scale feature extraction layers by 2x using bilinear interpolation, then sums them with the feature outputs of the previous layer, and finally obtains four multi-scale fused feature outputs through a 3×3 convolution.
[0035] The region proposal generation network includes a 3×3 convolution followed by two 1×1 convolutions, which are used for the output of the category and the box offset, respectively. The region proposal generation network is used to generate candidate boxes for the fused multi-scale features, and positive and negative samples are assigned according to the intersection-union ratio of the candidate boxes and the ground truth boxes.
[0036] The scale generated at each pixel location in each multi-scale fused feature map is 8. 2 16 2 and 32 2For rectangular anchor frames with aspect ratios of 1:2, 1:1, and 2:1, anchor frames with an intersection-union ratio (IU) greater than 0.7 that are in the same category as the label and are assigned as positive samples. Anchor frames with an IU of less than 0.3 that are in the same category as the label and are assigned as negative samples. The remaining anchor frames are assigned as irrelevant samples.
[0037] The ROI network aligns the positive and negative samples obtained by the region proposal generation network with features, and unifies them to a size of 7×7.
[0038] The latent feature enhancement unit consists of a class covariance calculation layer and a latent feature estimation layer connected in series.
[0039] The specific implementation steps of the covariance calculation layer are as follows:
[0040] The online estimate of the class covariance of class j in the t-th batch is:
[0041]
[0042] The mean value of the j-th class in the t-th batch is estimated as follows:
[0043]
[0044] The statistic for the total number of samples in the j-th class of the first t batches is:
[0045]
[0046] in, Let the class covariance matrix represent the features of the j-th class samples in the t-th batch of training samples. This represents the total number of samples belonging to the j-th class among all training samples in the first t-1 batches. This represents the total number of samples belonging to the j-th class in the t-th batch of training samples. This represents the covariance of the j-th class target feature in the t-th batch of training samples. Let represent the mean of the features of samples belonging to the j-th class among all training samples in the first t-1 batches. Let T represent the mean of the features of the j-th category samples in the t-th batch of training samples, and let T represent the matrix transpose operation.
[0047] The specific implementation steps of the latent feature estimation layer are as follows:
[0048] Set in the training set Train a model G with network parameters Θ, where y i ∈{1,...,C} represents the label of the i-th sample, a i =[a i1,...,a iA ] T =G(x) i ,Θ) represents G learning features about sample i, in order to obtain semantic orientation to enhance feature a. i ,right Random sampling is performed; the class covariance matrix obtained from the class covariance calculation layer is passed through the latent feature estimation layer to obtain the latent feature augmentation set, thus expanding the features. It is in a i Based on the mean, along Translate in a random direction.
[0049] By enhancing a i The enhanced feature set is obtained by performing the Mth iteration of the feature set:
[0050]
[0051] After increasing the augmentation order to ∞, the optimal augmented feature is obtained by minimizing the upper bound of the loss. The steps are as follows:
[0052] Assume the expected value of the augmented feature loss function obtained after infinitely many feature augmentations is:
[0053]
[0054] Finding the upper bound of its loss yields:
[0055]
[0056]
[0057] Among them, L ∞ (W,b,Θ) represents the expectation of the cross-entropy loss function using augmented ∞-order features. Let I represent the upper bound of the expected loss, I represent the total number of samples in the training batch, and E represent the expected loss operation.
[0058] in It follows a Gaussian distribution, therefore:
[0059]
[0060] Where w represents the weight parameters of the 1×1 convolutional layer in the classification branch of the classification regression network, b represents the bias of the 1×1 convolutional layer, and a i Represents the features of the i-th sample. Let represent the latent augmentation feature of the i-th sample, N represent a Gaussian distribution, and λ represent the positive coefficient of latent augmentation, which is set to 1.5 in this embodiment. This represents the covariance matrix of the category to which the i-th sample belongs, calculated by the covariance calculation unit.
[0061] Step 3, construct the dual-drive equalization loss function as follows:
[0062]
[0063] Where L represents the dual-drive equalization loss function, I represents the total number of samples in each training batch, C represents the total number of classes in the training set, and ∑ represents the summation operation. w represents the j-th element of the one-hot vector converted from the true label of the i-th sample. j,CLS This represents the class-based reweighting factor for the j-th category. This represents the probability that the i-th sample is predicted to be of the j-th class. It represents the sample-based reweighting factor for the i-th sample.
[0064] The class-based reweighting factor is implemented by the following formula:
[0065] w j,CLS =w j,G ×w j,CD ×w j,AC
[0066] Among them, w j,CLS w represents the class-based weight factor for the j-th category. j,G w represents the gradient-based weight factor for the j-th category. j,CD w represents the class-difficulty-based weight factor for the j-th category. j,AC W represents the class-based accuracy weight factor for the j-th category. CLS A vector composed of class-based weighting factors for each category represents a class-based reweighted vector.
[0067] The specific calculation steps for the gradient-based weighting factor are as follows:
[0068]
[0069]
[0070] Among them, g j,pos Let g represent the sum of the positive gradients of the j-th class sample. j,reg Let I represent the sum of the negative gradients of the j-th class samples, and let I represent the total number of samples. Let z be the probability of predicting the i-th sample as belonging to the j-th class, z be the output value of the sample after inputting it into the network, and p be the probability output of z after activation. For partial derivative operations, ∑ is the summation operation;
[0071]
[0072] Where α represents a hyperparameter that limits the range of weighting factors, and is set to 0.8 in this embodiment of the invention, and T represents the current training batch.
[0073] The specific calculation steps for the weight factor based on class difficulty are as follows:
[0074] Calculate the average correct prediction probability for each class, which is the sum of all correct prediction probabilities divided by the total number of samples. The one-hot encoding of the label of the i-th sample is... The predicted probability of the i-th sample is Where C represents the total number of sample categories.
[0075]
[0076] d j Let represent the mean probability of correctly predicting the j-th class sample, where This represents the predicted probability of the j-th class for the i-th sample. This represents the j-th code of the i-th sample, and ε is a very small value, set to 0.001 in this embodiment of the invention.
[0077] By limiting the weight parameter to a certain range, we obtain the class-accuracy-based weight factor for the j-th class:
[0078]
[0079] Wherein, γ is the intensity of difficulty adjustment, which is set to 1.25 in this embodiment of the invention.
[0080] The specific calculation steps for the weighting factor based on class accuracy are as follows:
[0081]
[0082] Among them, A j N represents the classification accuracy of category j. j n represents the total number of samples belonging to the j-th class in the mini-batch training samples. j This represents the number of correctly classified samples in the j-th category of the mini-batch.
[0083] After performing range-limited mapping on the weight parameter, the weight factor based on class accuracy is obtained as follows:
[0084] w j,AC =(1-A j ) τ
[0085] Where τ is a given hyperparameter, which is automatically updated according to different datasets. Its update formula during training is:
[0086]
[0087]
[0088] Where b is the classification bias, max represents the operation of finding the maximum value, min represents the operation of finding the minimum value, and ε is a small positive number, which is set to 0.001 in this embodiment of the invention.
[0089] The sample-based reweighting factor is implemented by the following formula:
[0090]
[0091] in, Let represent the sample-based reweighting factor for the j-th class of the i-th sample. Let represent the false detection rate-based weighting factor for the j-th class of the i-th sample. Let represent the class-related weighting factor for the j-th category of the i-th sample. It is a sample-based reweighted vector composed of sample-based reweighting factors.
[0092] The specific calculation steps for the weighting factor based on the false detection rate are as follows:
[0093]
[0094]
[0095] Where, n j→k n represents the number of samples in the mini-batch training samples that were originally classified as class j but were detected as class k. j y represents the total number of samples belonging to the j-th class in the mini-batch training samples. i This represents the category number to which the i-th sample belongs. β represents the probability that the i-th sample is falsely detected as the j-th category, and β is an exponential limiting term, which is set to 1.5 in this embodiment of the invention.
[0096] The specific calculation steps for the weighting factor based on class relevance are as follows:
[0097]
[0098] in, This represents the probability that the i-th sample is detected correctly. This parameter is obtained by averaging the probabilities of correctly detecting samples belonging to the same category as the i-th sample in the training samples. This parameter represents the probability that the i-th sample is detected as the j-th category. It is obtained by averaging the probabilities of samples belonging to the same category as the i-th sample in the training samples being detected as the j-th category. μ is a range limiting parameter, which is set to 1.1 in this embodiment of the invention.
[0099] Step 4: Train the SAR long-tail target detection model.
[0100] The training set is input into the SAR long-tail target detection model in batches. The gradient descent method is used to update the network parameter weights during backpropagation until both the SmoothL1 loss function and the dual-drive equalization loss function converge, thus obtaining the trained SAR long-tail target detection model.
[0101] The SmoothL1 loss function is as follows:
[0102]
[0103] Where x represents the deviation between the true bounding box and the candidate box of the category object.
[0104] Step 5: Detect targets in the SAR long-tail distribution image.
[0105] The model takes images from the test set as input, and outputs the class score and location of the target.
[0106] The effects of this invention will be further illustrated below with simulation experiments:
[0107] 1. Simulation experimental conditions:
[0108] The software platform for the simulation experiment of this invention is: Ubuntu 18.04 operating system and PyTorch 1.8.0, and the hardware configuration is: Xeon Silver 4214 CPU and NVIDIA GeForce RTX 2080Ti GPU.
[0109] The simulation experiment of this invention uses the measured data of SAR on the Gaofen-3 satellite. The scene type is an airport, the image resolution is 1m×1m, the number of SAR images is 2000, the image size is 600×600, 1024×1024, and 2048×2048, the number of target categories is 7, the total number of targets is 6556, the number of training set images is 1400, and the number of test set images is 600.
[0110] 2. Simulation content and result analysis:
[0111] The simulation experiment of this invention involves inputting training set images into a model constructed according to this invention and a model constructed using existing technology, respectively, for training to obtain trained models. Two SAR images from the test set are then input into the models of this invention and the existing technology for testing. The resulting target category scores and locations are visualized on the original SAR images, such as... Figure 3 As shown.
[0112] In simulation experiments, the existing technologies used refer to:
[0113] The method for detecting long-tailed objects proposed by Zhengzhou University in Henan Province in its patent application document "A method for detecting long-tailed objects based on deep learning" (application number: 2022116774312, application publication number: CN116129215A) is as follows.
[0114] The following is combined with Figure 3 The simulation diagrams further illustrate the effects of the present invention.
[0115] Figure 3 (a) is a visualization of the detection results of two SAR images under the existing technical model. Figure 3 (b) is a visualization of the detection results under the model of the present invention. Figure 3 The red solid boxes represent the detection results of each model, the green solid boxes represent missed rare category targets, the blue solid boxes represent missed targets other than rare categories, and the yellow dashed boxes represent false positives. Figure 3 (a) It can be seen that the existing technology has a large number of cases of missed detections and false detections. Figure 3 (a) The icon below illustrates situations where rare category targets are easily misdetected as other categories.
[0116] contrast Figure 3 As can be seen from (a) and 3(b), the present invention can solve the problem of the large number of missed detections and false detections of rare class targets in long-tail datasets in the prior art.
[0117] To evaluate the detection performance of the method of this invention compared to existing methods, the average detection accuracy (AP) for each target category of the two methods in the simulation experiment is calculated according to the following formula:
[0118]
[0119]
[0120] Where P represents precision, R represents recall, TP represents the number of correctly classified positive samples, FP represents the number of misclassified positive samples, and FN represents the number of misclassified negative samples.
[0121] AP represents the area under the PR curve under each threshold condition, used to represent the effect of different target detection classes in a balanced way. mAP represents the mean of AP for all classes, and mAP is the standard metric for target detection models.
[0122] The standard detection metrics mAP of this invention and existing technologies were compared on all test set images. Furthermore, based on the number of samples of different categories in the dataset, the seven target classes were divided into frequent, normal, and rare classes, and their mAP values were compared. The results are shown in Table 1.
[0123] Table 1
[0124] Existing technology 87.8 81.4 89.04 88.0 This invention 91.1 95.8 90.38 90.0
[0125] In Table 1, mAP represents the overall class average precision, mAPr represents the class average precision of rare classes, mAPc represents the normal class average precision, and mAPf represents the frequent class average precision.
[0126] As can be seen from Table 1, the overall mAP of the present invention is higher than that of the prior art. In particular, the class accuracy mAPr for rare categories is improved by 14.4 points compared with the prior art, indicating that the detection performance of the present invention for long-tailed targets is significantly better than that of the prior art.
Claims
1. A SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement, characterized in that, The detection method enriches the features of tail category samples by using latent feature enhancement units in the SAR long-tail target detection model, and balances the loss calculations of different categories using a dual-drive equalization loss function. The steps of this detection method include: Step 1, Generate training set: Select containing Aircraft-like targets A sample set is composed of SAR images, and the number of all tail category targets in the sample set is at least equal to the number of head category targets. ,in, , The target location and category are labeled in each image in the sample set, and the sample set composed of the labeled images is used as the training set. Step 2: Construct a SAR long-tail target detection model consisting of a feature extraction network, a feature fusion network, a region proposal generation network, an ROI network, and a classification and regression network with implicit feature enhancement, and connect them sequentially. The classification and regression network with implicit feature enhancement is composed of a 3×3 convolutional layer, an implicit feature enhancement unit, and a classification and regression sub-network connected sequentially. The implicit feature enhancement unit is composed of a class covariance calculation layer and an implicit feature estimation layer connected in series. The classification and regression sub-network is composed of two 1×1 convolutional layers connected in parallel. Step 3, construct the dual-drive equalization loss function as follows: , Where L represents the dual-drive equalization loss function, I represents the total number of samples in each training batch, and C represents the total number of classes in the training set. This represents the summation operation. This represents the j-th element of the one-hot vector converted from the true label of the i-th sample. This represents the class-based reweighting factor for the j-th category. This represents the probability that the i-th sample is predicted to be of the j-th class. This represents the reweighting factor of the i-th sample based on the sample; Step 4, train the SAR long-tail target detection model: The training set is input into the SAR long-tail target detection model in batches. The gradient descent method is used to update the network parameter weights during backpropagation until the SmoothL1 loss function and the dual-drive equalization loss function converge, thus obtaining the trained SAR long-tail target detection model. Step 5: Input the SAR long-tail image to be detected into the trained model and output the target detection result.
2. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The feature extraction network described in step 2 consists of a 3×3 convolutional layer, a first-scale feature extraction layer, a second-scale feature extraction layer, a third-scale feature extraction layer, and a fourth-scale feature extraction layer connected in series. The feature extraction layers at the first to fourth scales are composed of 3, 4, 6, and 3 identical Bottleneck blocks, respectively. Each Bottleneck block is composed of a convolutional block with residual structure and a feature fusion convolutional layer connected in series. The convolutional block with residual structure includes two branches: the first branch consists of a 3×3 convolutional layer, and the second branch consists of two 3×3 convolutional layers connected in series. The feature fusion convolution is composed of a 3×3 convolutional layer. The outputs of the first and second branches are added together and then input into the feature fusion convolutional layer. For the feature extraction layers at the first to fourth scales, the stride of the first convolutional layer in the last Bottleneck is 2, and the stride of all other convolutional layers is 1.
3. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The feature fusion network described in step 2 upsamples the outputs of the four scale feature extraction layers by bilinear interpolation to twice the original value, then sums them with the feature outputs of the previous layer, and finally obtains four multi-scale fused feature outputs through a 3×3 convolution.
4. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The region proposal generation network described in step 2 includes a 3×3 convolution followed by two 1×1 convolutions, which are used for the output of the category and the box offset, respectively. The region proposal generation network is used to generate candidate boxes for the fused multi-scale features, and positive and negative samples are assigned according to the intersection-union ratio of the candidate boxes and the ground truth boxes.
5. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, Step 2 describes the ROI network aligning the positive and negative samples obtained by the region proposal generation network to a uniform size of 7×7.
6. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The covariance calculation layer described in step 2 is implemented using the following formula: , in, Let the class covariance matrix represent the features of the j-th class samples in the t-th batch of training samples. Let represent the total number of samples belonging to the j-th class among all training samples in the first t-1 batches. Let represent the total number of samples belonging to the j-th category in the t-th batch of training samples. This represents the covariance of the j-th class target feature in the t-th batch of training samples. Let represent the mean of the features of samples belonging to the j-th class among all training samples in the first t-1 batches. Let T represent the mean of the features of the training samples in the t-th batch belonging to the j-th category, and let T represent the matrix transpose operation.
7. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The latent feature estimation layer described in step 2 is implemented by the following formula: , Where w represents the weight parameters of the 1×1 convolutional layer in the classification branch of the classification and regression network, and b represents the bias of the 1×1 convolutional layer. This indicates the category to which the i-th sample belongs. Represents the features of the i-th sample. Let N represent the latent augmented feature of the i-th sample, and let N denote a Gaussian distribution. Indicates the latent enhancement strength. This represents the class covariance matrix of the class to which the i-th sample belongs, calculated by the covariance calculation unit.
8. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The class-based reweighting factor described in step 3 is implemented by the following formula: , in, This represents the class-based weight factor for the j-th category. Let represent the gradient-based weight factor for the j-th category. This represents the class-difficulty-based weight factor for the j-th category. This represents the class-based accuracy weight factor for the j-th category. A vector composed of class-based weighting factors for each category represents a class-based reweighted vector. , in, This represents a hyperparameter that limits the range of weighting factors. This represents the sum of the negative sample gradients of the j-th class samples in a total of T batches. Let represent the sum of positive sample gradients for the j-th class of the total T batches, and C represent the total number of sample classes in the training set; , in, Let represent the class-based difficulty weighting factor for the j-th category, and γ be the intensity factor for difficulty adjustment, ranging from [1,3]. This represents the mean probability of correctly detecting the j-th class of samples; , in, This represents the class-based accuracy weight factor for the j-th category. This represents the detection accuracy for the j-th category. This indicates the exponentially weighted strength.
9. The SAR long-tail target detection method based on dual-drive equalization loss and latent feature enhancement according to claim 1, characterized in that, The sample-based reweighting factor described in step 3 is implemented by the following formula: , in, Let represent the sample-based reweighting factor for the j-th class of the i-th sample. This represents the sample reweighting factor based on the false detection rate for the j-th class of the i-th sample. Let represent the class-related sample reweighting factor for the j-th class of the i-th sample. It is a sample-based reweighted vector composed of sample-based reweighting factors; , , in, Let represent the probability that the i-th sample is falsely detected as the j-th class. It is the exponential intensity coefficient; Indicates a range-limiting parameter. This represents the probability that the i-th sample is detected correctly. This represents the probability that the i-th sample is detected as belonging to the j-th class. The range is [1,3]. The range is [1,2].
Citation Information
Patent Citations
Long-tail target detection method based on deep learning
CN116129215A
Video SAR moving target shadow detection method
CN114511504A
Multi-modal data-based heavy balance long-tail image data classification method
CN115205592A