Remote sensing directed target detection method based on Gaussian distribution perception label distribution strategy
By employing a Gaussian distribution-based label assignment strategy and utilizing Gaussian modeling and a hybrid scoring mechanism, the problems of poor adaptability and weight imbalance in label assignment strategies for remote sensing directed target detection are solved. This enables efficient and robust detection of directed targets in remote sensing images, improving detection accuracy and robustness.
Patent Information
- Application Number
- CN202511259996.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-18
AI Technical Summary
In existing remote sensing directed target detection, the label assignment strategy relies on manually set scoring functions, which have poor adaptability, leading to an imbalance between classification and localization weights. Furthermore, traditional evaluation methods are biased and cannot comprehensively assess the matching quality between samples and real targets.
A label assignment strategy based on Gaussian distribution perception is adopted. Through Gaussian modeling and a mixed scoring mechanism, the similarity between candidate anchor boxes and real targets is calculated using Kullback-Leibler divergence and generalized Jensen-Shannon divergence. The classification and regression loss weights are dynamically adjusted to achieve adaptive assignment of positive and negative samples.
It achieves efficient and robust detection of directed targets in remote sensing images, improves detection accuracy and robustness, avoids the problems of weight imbalance and selection bias in traditional methods, and enhances the generalization ability of the model.
Smart Images

Figure CN120976583A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target detection technology, and particularly relates to a remote sensing oriented target detection method. BACKGROUND
[0002] With the continuous development of deep learning technology, remote sensing target detection technology is also advancing with the times. Remote sensing images are increasingly widely used in various fields such as city monitoring, disaster prevention and control, and resource management. In these application scenarios, remote sensing oriented target detection, as a key technology, not only enables accurate positioning of target positions, but also identifies direction information (such as heading angle or orientation) of targets in images, further improving the accuracy and practicality of detection. However, this task still faces many technical challenges, mainly including: 1) targets have arbitrary directions; 2) targets have large aspect ratios; and 3) there are a large number of dense small targets.
[0003] Label assignment is a key preparatory work in the target detection task, and its role is particularly important in the oriented target detection scenario in remote sensing images. Through a reasonable label assignment strategy, the detection model can more effectively select suitable preset anchor boxes and real targets for matching in the training stage, and accordingly divide positive and negative samples to provide reliable supervision information for the model, thereby improving its target classification and positioning ability. Therefore, how to accurately evaluate the quality of anchor boxes and how to efficiently and reasonably assign positive and negative samples become one of the core problems affecting the detection performance. The existing mainstream label assignment strategies mostly use IoU as the basic standard for evaluating the quality of anchor boxes. The earliest MaxIoU strategy selects matching targets by the maximum intersection over union. However, the targets in remote sensing images have significant characteristics such as large scale difference, dense distribution, and arbitrary direction, which leads to obvious limitations of the MaxIoU strategy in practical applications. Therefore, researchers have proposed various improved strategies, such as improving the calculation method of IoU to consider more complex geometric relationships between targets, or introducing a dynamic threshold mechanism to adapt to the distribution characteristics of multi-scale and multi-class targets in images.
[0004] For example, the adaptive training sample selection (ATSS) strategy uses the mean and variance of the IoU of cross-level anchor boxes to perform dynamic threshold calculation to fit the anchor box quality distribution; the dynamic anchor box learning (DAL) method introduces prior and posterior IoU to construct a dynamic scoring function; the shape adaptive selection and measurement (SASM) further introduces a center point distance index and formulates strategies under the anchor-based and anchor-free frameworks, effectively improving the detection effect of large aspect ratio targets. In addition, the SALA strategy sets a restrictive sampling region and performs threshold compensation for small targets and slender targets to improve the accuracy of sample selection.
[0005] Although the above methods improve the traditional label assignment mechanism to some extent, they mostly rely on manually set scoring functions and cannot flexibly adapt to the complex and diverse distribution of directional targets in remote sensing images. At the same time, these strategies usually use fixed score weights to combine classification scores and regression quality, ignoring the dynamic relationship between the two during the training process. This static scoring method may lead to an imbalance in weight distribution when classifying and locating targets, thereby affecting the final detection accuracy.
[0006] In addition, the traditional evaluation method based on IoU or center point distance has a selection bias, and some high IoU samples may correspond to low classification confidence, making it difficult to provide high-quality training signals. To overcome these problems, a more comprehensive and fine-grained label assignment strategy is needed to evaluate the matching quality between samples and real targets from multiple dimensions. SUMMARY
[0007] To solve the technical problems of existing label assignment strategies, such as poor adaptability due to reliance on manually set scoring functions, imbalance in classification and positioning weights caused by static scoring, and selection bias of evaluation methods, the present application proposes a remote sensing directional target detection method based on a Gaussian distribution-aware label assignment strategy, which realizes dynamic selection of positive and negative samples, dynamic adjustment of classification and regression loss weights, and comprehensive evaluation of sample and target matching relationships.
[0008] To achieve the above purpose, the technical solution of the present application is as follows:
[0009] A remote sensing directional target detection method based on a Gaussian distribution-aware label assignment strategy, comprising the following steps:
[0010] S1: Obtain a remote sensing image dataset and perform preprocessing;
[0011] S2: Construct a GDALA-Net network with ResNet-50 as the backbone network and a feature pyramid as the neck network, and pre-train the GDALA-Net network;
[0012] S3: Input the preprocessed remote sensing image, and train the GDALA-Net network based on the Gaussian distribution-aware label assignment strategy;
[0013] S4: Input the remote sensing image to be detected, and use the GDALA-Net network trained in step S3 to obtain directional detection boxes and confidence levels containing directional target categories.
[0014] Further, the training processing method of step S3 is as follows:
[0015] Input the training set in the preprocessed remote sensing image, and use the ResNet-50 backbone network to extract multi-scale feature maps from the input remote sensing image;
[0016] The multi-level feature maps are generated by using a feature pyramid network, and on each level feature map, a plurality of dense prior anchor boxes are generated at each feature point position based on a preset aspect ratio and a scale factor;
[0017] The detection head outputs classification scores of each prior anchor box through a classification branch, and predicts boundary box parameters of each prior anchor box through a regression branch, and generates a candidate anchor box set in combination with the classification scores and the boundary box parameters;
[0018] Based on the candidate anchor boxes, the classification scores and the labeled real oriented target boxes in the training set, label assignment is performed based on Gaussian modeling and a mixed scoring mechanism;
[0019] According to the label assignment result, the GDALA-Net network parameters are updated through back propagation based on a focal loss and a KLD loss.
[0020] Further, the method for performing label assignment based on Gaussian modeling and a mixed scoring mechanism is:
[0021] 1), respectively, the candidate anchor boxes and the real oriented target boxes are modeled by an oriented target Gaussian distribution;
[0022] 2), based on Kullback-Leibler divergence and generalized Jensen-Shannon divergence, Gaussian distribution similarity measurement calculation is performed on the Gaussian distribution corresponding to the candidate anchor boxes and the Gaussian distribution corresponding to the real oriented target boxes;
[0023] 3), based on the candidate anchor boxes and the classification scores, a mixed score is calculated, and a positioning and classification combined loss cost is calculated in combination with the mixed score and the Gaussian distribution similarity measurement;
[0024] 4), label assignment is performed based on the positioning and classification combined loss cost.
[0025] Further, the center point of the candidate anchor box is taken as the mean value of the Gaussian distribution, the width, height and rotation angle of the candidate anchor box are taken to calculate the Gaussian distribution covariance matrix, and the Gaussian distribution N p corresponding to the candidate anchor box is constructed. The center point of the real oriented target box is taken as the mean value of the Gaussian distribution, the width, height and rotation angle of the real oriented target box are taken to calculate the Gaussian distribution covariance matrix, and the Gaussian distribution N t corresponding to the real oriented target box is constructed.
[0026] Further, the implementation method of step 2) is: calculating the Kullback-Leibler divergence of different directions between the Gaussian distribution corresponding to the candidate anchor frame and the Gaussian distribution corresponding to the real directional target frame based on the closed-form solution formula of the Kullback-Leibler divergence between the multi-dimensional Gaussian distributions, calculating the generalized Jensen-Shannon divergence between the Gaussian distribution N p and the Gaussian distribution N t corresponding to the real directional target frame as the similarity measure based on the Kullback-Leibler divergence of different directions.
[0027] Further, the closed-form solution formula of the Kullback-Leibler divergence between the multi-dimensional Gaussian distributions is:
[0028]
[0029] wherein D kl (N1||N2) represents the Kullback-Leibler divergence of the direction from the Gaussian distribution N1 to the Gaussian distribution N2, Tr(·) represents the trace of the matrix, μ1 and μ2 are the mean values of the Gaussian distribution N1 and the Gaussian distribution N2 respectively, and Σ1 and Σ2 are the covariance matrices of the Gaussian distribution N1 and the Gaussian distribution N2 respectively.
[0030] The calculation method of the similarity measure is:
[0031] GJSD(N p ||N t )=(1-λ)D kl (N p ||N t )+λD kl (N t ||N p )
[0032] wherein λ is a weight coefficient for balancing the KL divergence of two directions, D kl (N p ||N t ) is the Kullback-Leibler divergence of the direction from the Gaussian distribution N p corresponding to the candidate anchor frame to the Gaussian distribution N t corresponding to the real directional target frame, and D kl (N t ||N p ) is the Kullback-Leibler divergence of the direction from the Gaussian distribution N t corresponding to the real directional target frame to the Gaussian distribution N p corresponding to the candidate anchor frame.
[0033] Further, the method for calculating the combined loss cost of positioning and classification by combining the mixed score and the similarity measure of Gaussian distribution is: calculating a mixed cost C MS according to the mixed score p ; calculating a geometric cost C t according to the similarity measure GJSD(N g ); and introducing a periodic coefficient to calculate the combined loss C(x, y) of positioning and classification according to the mixed cost and the geometric cost.
[0034] Further, the method for calculating the mixed score based on the candidate anchor frame and the classification score is:
[0035] MS i = IoU(A i , gt i )*Cls(A i )
[0036] wherein A i represents the candidate anchor frame, gt i is the real directed target frame matched with the candidate anchor frame A i , IoU(·) represents the intersection over union, and Cls(A i ) represents the classification score of the candidate anchor frame A i .
[0037] The method for calculating the mixed cost C MS is:
[0038]
[0039] The method for calculating the geometric cost C g is:
[0040]
[0041] The method for calculating the combined loss cost of positioning and classification is:
[0042]
[0043] wherein, is a periodic coefficient, iter represents the number of current iterations, iter max represents the maximum number of iterations during the model training, R g is a Gaussian modeling region, (x, y) is the center point coordinate of the candidate anchor frame, C g (x, y) is the geometric cost of the candidate anchor frame with the center position (x, y), and C MS(x,y) is a mixed cost of the candidate anchor frame with the center position (x,y), C(x,y) is a positioning and classification combined loss cost of the candidate anchor frame with the center position (x,y);
[0044] The KLD loss is:
[0045]
[0046] Wherein, τ is a hyperparameter.
[0047] Further, the method for label assignment based on the positioning and classification combined loss cost is:
[0048] For a large-scale target with a predefined size in the preprocessed remote sensing image, the first k candidate anchor frames with the lowest positioning and classification combined loss cost are selected from the Gaussian region R g as positive samples:
[0049] For a small-scale target with a predefined size in the preprocessed remote sensing image, all candidate anchor frames in the Gaussian region R g are regarded as positive samples.
[0050] For the case that the candidate anchor frame overlaps with multiple standards, it is attributed to the ignored region.
[0051] The candidate anchor frame outside the Gaussian region is attributed to the negative sample set R neg , which is used to suppress background interference.
[0052] Further, the multiple standards overlap are:
[0053] If there is a candidate anchor frame A i whose center point (x,y) is located in the Gaussian region R g , but the positioning and classification combined loss cost C(x,y) > τ1.
[0054] If there is a candidate anchor frame A i whose center point (x,y) is located outside the Gaussian region R g , but the positioning and classification combined loss cost C(x,y) > τ2; τ1 and τ2 are respectively an upper threshold and a lower threshold of the label assignment cost.
[0055] The present application has the following advantages:
[0056] The application proposes a label assignment strategy (GDALA) based on Gaussian distribution perception, by modeling the directed target as a two-dimensional Gaussian distribution, using Kullback-Leibler divergence (KLD) to construct generalized Jensen-Shannon divergence (GJSD) as a similarity measurement index, dynamically evaluating the matching quality of the candidate region and the real target, replacing the traditional scoring function which relies on manual setting, and realizing adaptive assignment of positive and negative samples.
[0057] The application constructs a unified cost function combining spatial alignment degree, classification confidence and Gaussian distribution similarity, introduces a periodic coefficient, dynamically adjusts the weight of classification and regression loss, and solves the imbalance problem caused by fixed scoring weight.
[0058] The application comprehensively considers the center position, scale, direction angle and other parameters of the target by Gaussian distribution modeling, uses GJSD as a similarity measurement, combines mixed scores to measure positioning accuracy and semantic confidence, comprehensively evaluates the matching relationship between samples and targets, and avoids the selection tendency of traditional single indicators such as IoU or center point distance. BRIEF DESCRIPTION OF DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0060] Figure 1 It is a remote sensing directed target detection method flowchart based on the label assignment strategy of the application based on Gaussian distribution perception.
[0061] Figure 2 It is a schematic diagram of the label assignment strategy based on Gaussian distribution perception.
[0062] Figure 3 It is a directed target to 2-D Gaussian diagram.
[0063] Figure 4 It is a comparison diagram of traditional regression and Gaussian distribution regression.
[0064] Figure 5 It is an ablation experiment result diagram of remote sensing directed target detection.
[0065] Figure 6 It is a class detection result diagram of remote sensing directed target detection. DETAILED DESCRIPTION
[0066] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of the present application.
[0067] A remote sensing directed target detection method based on a Gaussian distribution perception label assignment strategy, as shown in Figure 1 The steps are as follows:
[0068] S1: Obtain a remote sensing image dataset and perform preprocessing.
[0069] Specifically, in the embodiments of the present application, the dataset adopts the DIOR-R dataset. The DIOR-R is a dataset specially designed for rotating target detection. The DIOR-R dataset covers 20 categories, including airplanes, ships, sports grounds, bridges, ports and various targets. It contains 23,463 images, a total of 288,800 instances, and the image size is uniform at 800x800 pixels, with a spatial resolution in the range of 0.5m-30m.
[0070] In the present embodiment, a plurality of remote sensing images with an original resolution of 800x800 pixels are first downloaded, and are subjected to cropping processing. Each original image is cut into a fixed window size of 640x640 pixels. Specifically, an overlapping cropping method is adopted during cropping, and a stride of 320 pixels is set, that is, the original image is shifted by 320 pixels in the horizontal and vertical coordinates each time, to generate image blocks with partially overlapping adjacent regions, so as to ensure that small targets will not be truncated or missed due to the cropping boundary. Through the above operation, each original image can be cut into a plurality of 640x640 sub-image blocks, and an initial dataset is constructed. The 23,463 images are divided into a training set of 70% (16,424 images), a validation set of 15% (3,519 images), and a test set of 15% (3,520 images). The training set is used for model training, the validation set is used for parameter tuning, and the test set is used for performance evaluation and visual comparison.
[0071] After that, the remote sensing images in the training set are data enhanced, and the mosaic data enhancement method is adopted, and the specific process includes: randomly selecting four cropped images, and determining the random splicing reference point coordinates; then the four images are respectively adjusted in size and scaled, and placed in the upper left, upper right, lower left and lower right positions of the designated size large image in turn; at the same time, according to the scaling and transformation mode of each image, the corresponding target label information is mapped into the spliced new image; finally, the large image is spliced according to the preset horizontal and vertical coordinates, and the target frame exceeding the boundary is truncated or corrected. Through the above cropping and enhancement operation, the diversity and complexity of the training data can be effectively improved, so as to enhance the generalization ability of the model to different target scales and positions.
[0072] S2: Construct a GDALA-Net network with ResNet-50 as the backbone network and feature pyramid as the neck network, and pre-train the GDALA-Net network.
[0073] In the oriented target detection task, the network architecture based on Rotated RetinaNet is adopted. Since Rotated RetinaNet has introduced a rotated box regression mechanism, the present application follows this mechanism to adapt to the directional characteristics of the targets in the remote sensing image, and improves it on the basis of combining the label assignment strategy of Gaussian distribution. In this embodiment, the oriented target detection network takes ResNet-50 as the backbone feature extraction network, and introduces the feature pyramid network (FPN) as the neck structure; wherein, the FPN and the detection head are based on the Rotated RetinaNet framework to realize the fusion and expression of multi-scale features, and enhance the perception ability to targets of different sizes. The oriented target detection network can model the target angle information, thereby effectively improving the positioning accuracy and detection robustness. The ResNet-50 backbone adopts ImageNet pre-training weights, and is fine-tuned on the DIOR-R training set.
[0074] S3: Based on the Gaussian distribution perception-based label assignment strategy (GDALA), the GDALA-Net network is trained with the preprocessed remote sensing image as input.
[0075] In order to further improve the discrimination ability and geometric adaptability of label assignment, the present application designs a label assignment strategy based on Gaussian distribution perception (Gaussian-based Matching and Label Assignment, GDALA), as shown in Figure 2As shown, for replacing the traditional IoU-based label assignment method. This strategy measures the matching relationship between the candidate position and the target more accurately by modeling the candidate region and the real target box as a two-dimensional Gaussian distribution and using the Kullback-Leibler divergence (KLD) to construct the generalized Jensen-Shannon divergence (GJSD) as a similarity measure index. In remote sensing oriented object detection, the traditional label assignment strategy based on IoU or centerness has certain limitations. The IoU-based method only considers the intersection over union between the real target and the sample, ignoring the directionality of the oriented target. This rough scheme causes the screened positive samples to possibly contain too much background information, thereby interfering with model learning. While the centerness-based method introduces pixel classification scores to enhance the semantic information of the anchor box, it also ignores the importance of angle in label assignment. In contrast, the Gaussian distribution-aware label assignment strategy (GDALA) of the present application uses a multi-parameter joint evaluation method to represent the similarity between the sample and the target by Gaussian modeling, including target center point, width, height, and rotation angle. It can consider IoU, centerness method, and directionality factors comprehensively. Therefore, using Gaussian distribution to represent the oriented target is an efficient method for evaluating the similarity between the rotated boxes.
[0076] Specifically, in the embodiments of the present application, the method for training the GDALA-Net network based on the Gaussian distribution-aware label assignment strategy (GDALA) is as follows:
[0077] First, the training set in the preprocessed remote sensing image is input, and the ResNet-50 backbone network is used to extract multi-scale feature maps (C3 to C5) from the input remote sensing image.
[0078] Further, the feature pyramid network (FPN) fuses the multi-scale features by top-down feature transmission and horizontal connection to generate multi-level feature maps (P3-P7). Based on the preset aspect ratio {1:1, 1:2, 2:1} and scale factor {32, 64, 128, 256, 512}, a plurality of dense prior anchor boxes are generated at each feature point position on each level feature map, which are used to cover targets of different scales and shapes.
[0079] Further, the detection head is set with a classification branch and a regression branch based on this: the classification branch is used to output the classification scores of each prior anchor box, and the regression branch is used to predict the bounding box parameters of each prior anchor box; the combination of the two generates a candidate anchor box set (including class, classification score, and bounding box parameter), which provides a candidate basis for the subsequent Gaussian modeling and label assignment strategy.
[0080] Further, based on the candidate anchor frame, the classification score and the labeled real directed target frame in the training set, label assignment is performed based on Gaussian modeling and mixed scoring mechanism.
[0081] To solve the deficiency of the existing IoU or distance measurement strategy in direction perception and scale adaptation, the application proposes a label assignment strategy based on Gaussian distribution perception (GDALA), which realizes more robust and adaptive label assignment through Gaussian modeling and mixed scoring mechanism. Figure 2 As shown in the figure, it is a GDALA label assignment diagram. The figure shows the complete process from the target bounding box to the Gaussian label generation, including Gaussian modeling, similarity measurement calculation and label assignment strategy based on mixed cost function. Through the process, adaptive positive sample selection of dense and rotating targets can be realized, and the accuracy of training can be improved.
[0082] The label assignment strategy based on Gaussian distribution perception (GDALA) comprises:
[0083] 1. Respectively model the candidate anchor frame and the real directed target frame with directed target Gaussian distribution.
[0084] Specifically, the 2-D Gaussian distribution is used to model the candidate anchor frame and the real directed target frame respectively:
[0085]
[0086] N t =(μ t ,Σ t )
[0087] Wherein, μ p represents the mean of the Gaussian distribution corresponding to the candidate anchor frame, μ t represents the mean of the Gaussian distribution corresponding to the real directed target frame, Σ p , Σ t represents the covariance matrix.
[0088] Wherein, the mean is represented as:
[0089] μ=(x,y) T
[0090] The covariance matrix is represented as:
[0091]
[0092] Wherein, R and Λ represent the rotation matrix and the eigenvalue diagonal matrix respectively, α is the rotation angle, w represents the anchor frame width, and h represents the anchor frame height.
[0093] 2. The Kullback-Leibler divergence (KLD) and the generalized Jensen-Shannon divergence (GJSD) are used to calculate the similarity between the Gaussian distribution corresponding to the candidate anchor box and the Gaussian distribution corresponding to the real oriented target box.
[0094] Specifically, the Kullback-Leibler divergence between the Gaussian distribution corresponding to the candidate anchor box and the Gaussian distribution corresponding to the real oriented target box in different directions is calculated. The closed-form solution of the Kullback-Leibler divergence (KLD) between multi-dimensional Gaussian distributions is used to evaluate the similarity between two two-dimensional Gaussian distributions:
[0095]
[0096] where D kl (N p ||N t ), D kl (N t ||N p ) represents the Kullback-Leibler divergence between two Gaussian distributions in different directions, Tr(·) represents the trace of a matrix, i.e., the sum of the main diagonal elements of the matrix, the first term is the mean difference term, which measures the difference between the centers of the two distributions, the second term is the covariance difference term, which measures the difference between the shapes and dispersion of the two distributions, and the third term is the covariance determinant difference term, which is used to measure the difference between the overall ranges of the two distributions.
[0097] Further, the generalized Jensen-Shannon divergence between the Gaussian distribution corresponding to the candidate anchor box and the Gaussian distribution corresponding to the real oriented target box is calculated based on the Kullback-Leibler divergence as the similarity measure GJSD(N p ||N t ). Specifically, the generalized Jensen-Shannon divergence (GJSD) obtained by symmetrizing the KLD is used. Compared with the IoU, which only considers geometric overlap, the GJSD can simultaneously measure the differences in center position, scale, and direction, and is more suitable for rotating targets.
[0098] GJSD(N p ||N t ) = (1-λ)D kl (N p ||N t ) + λD kl (N t ||N p )
[0099] where λ is a weight coefficient used to balance the KL divergence in two directions, D kl (Np ||N t )direction and D kl (N t ||N p )direction. In the present application, the contributions of N p and N t are equal, so λ = 0.5 is set.
[0100] 3. Calculate a mixed score based on the candidate anchor frame and the classification score, and calculate the positioning and classification combined loss cost combining the mixed score and the Gaussian distribution similarity measure.
[0101] Specifically, a mixed score MS i is calculated based on the candidate anchor frame and the classification score: for each candidate anchor frame, a mixed score MS i is calculated to measure its positioning accuracy and semantic confidence at the same time:
[0102] MS i = IoU(A i ,gt i )*Cls(A i )
[0103] wherein A i represents the candidate anchor frame, gt i is the real directed target frame matched therewith, IoU(·) represents the intersection over union, and Cls(A i ) represents the classification score.
[0104] Further, the positioning and classification combined loss cost function is constructed combining the mixed score and the Gaussian distribution similarity measure.
[0105] The mixed cost C MS is calculated according to the mixed score MS i :
[0106]
[0107] C MS The mixed cost function can consider geometric alignment and semantic distinction under the same evaluation index by introducing the classification score of the candidate anchor frame and the Gaussian similarity measure, so as to comprehensively measure the overall quality of the candidate frame. In the candidate anchor frame screening process, the candidate anchor frame similar in geometry but unreliable in semantics is avoided to be selected as a positive sample, or the sample with high classification score but large geometric deviation is avoided to mislead the training, so as to improve the discriminability and stability of sample allocation.
[0108] Based on the construction of the aforementioned hybrid cost function, this invention further introduces a Gaussian distribution-based similarity metric, GJSD, to accurately measure the spatial alignment between candidate anchor boxes and ground truth directed target boxes. Specifically, a two-dimensional Gaussian distribution is used to model both candidate anchor boxes and ground truth directed target boxes, and the difference between them is calculated using a corresponding similarity metric function, thereby obtaining the similarity metric value GJSD(N). p ||N t Construct the geometric cost function C. g Compared to single metrics such as IoU, it captures the geometric differences between targets more comprehensively, avoiding situations where the classification score is high but the spatial position is seriously offset, especially performing better in rotated targets and targets with extreme aspect ratios.
[0109] According to the similarity metric GJSD(N) p ||N t ) Calculate the geometric cost function C g :
[0110]
[0111] The location and classification combined loss cost C(x,y) is calculated based on the hybrid cost and geometric cost. The cost score is dynamically represented and used to align the best task-oriented score in space to determine candidate locations.
[0112]
[0113] Among them, R g Let C be the region modeled by Gaussian, (x, y) be the coordinates of the center point of the candidate anchor box, and C be the region modeled by Gaussian. g (x,y) represents the geometric cost of the candidate anchor box centered at (x,y), C MS (x,y) represents the blending cost of a candidate anchor box centered at (x,y). As the periodic coefficient, in the initial stage of model training, due to inaccurate localization and classification, label assignment mainly relies on the Gaussian benchmark to ensure stability.
[0114] The combined cost C(x,y) of localization and classification is dynamically balanced by assigning weights to both categories within the cost function. Initially, geometric cost constraints are primarily used to ensure stable label assignment; later, the weight of the combined cost is gradually increased to improve accuracy. This adaptive adjustment of focus at different training stages enhances detection accuracy and model convergence speed.
[0115] This invention constructs a unified cost function that combines spatial alignment, classification confidence, and Gaussian distribution similarity. The spatial alignment corresponds to the geometric cost function C. gThis reflects the consistency between the candidate anchor frame and the target in terms of position, scale, and angle; the classification confidence corresponds to the Cls(A) output of the classification branch. i This reflects the semantic reliability of candidate anchor boxes; Gaussian distribution similarity is measured by the similarity metric GJSD(N). p ||N t The result is the core indicator of the geometric cost function.
[0116] The period coefficient is expressed as:
[0117]
[0118] Where iter represents the current iteration number, iter max This indicates the maximum number of iterations during model training.
[0119] 4. Tag allocation is performed based on the combined loss cost of positioning and classification. Figure 3 The paper presents a schematic diagram of a 2D Gaussian heatmap, in which classification scores are used for sample selection. The classification scores in this heatmap are used for sample selection, demonstrating the modeling process of rotating targets using a 2D Gaussian distribution. The diagram highlights how the Gaussian distribution is dynamically adjusted based on the target's angle, aspect ratio, and scale, thereby achieving adaptive modeling of targets in different directions. This method can more accurately characterize the spatial distribution characteristics of targets, especially demonstrating stronger modeling and detection capabilities for small targets.
[0120] Specifically, for large-scale targets, those with a pixel size of 96×96 or larger in the preprocessed image are considered large targets, measured from the Gaussian region R. g Select the k candidate anchor boxes with the lowest cost as positive samples:
[0121]
[0122] in, Let represent the set of candidate locations for large-scale targets, rank(·) represent the sorting operation, and C(i) represent the combined loss cost of localization and classification for the i-th candidate anchor box.
[0123] For small-scale targets, targets smaller than 32×32 pixels in the preprocessed image are considered as positive samples at all locations within the Gaussian region:
[0124]
[0125] in, This represents the set of candidate locations for small-scale targets.
[0126] For cases where candidate anchor boxes overlap with multiple ground truth directed target boxes, a minimum cost priority strategy is used for matching:
[0127] g * = argminC(i)
[0128] where g * the final matching target.
[0129] For the case of candidate anchor frame overlapping with multiple standards, it is attributed to the ignore area.
[0130] The candidate anchor frame outside the Gaussian area is attributed to the negative sample set R neg , used to suppress background interference.
[0131] The multiple standards overlap are:
[0132] If there is a candidate anchor frame A i The center point (x, y) is located in the Gaussian area R g , but the positioning and classification combined loss cost C(x, y) > τ1, this position may not meet the positioning and classification task, although it has a higher priority.
[0133] If there is a candidate anchor frame A i The center point (x, y) is located outside the Gaussian area R g , but the positioning and classification combined loss cost C(x, y) > τ2, this position is for low priority area, that is, too close to the joint area between the target and the background, although it obtains very low loss cost C(x, y). τ1 and τ2 are the upper threshold and lower threshold of the sample selection cost respectively.
[0134] In the above two cases, the prior priority (that is, whether the center point of the candidate anchor frame is in the Gaussian area) and the loss cost C(x, y) contradict each other, and it is not appropriate to regard the center point of the candidate anchor frame as a positive sample or a negative sample, so it is ignored and not used for network training. The ignore area processing can avoid the contradiction samples (such as the center point falling in the Gaussian area but the loss is high, or falling outside the area but the loss is low) misleading the training. It can effectively reduce the uncertainty of label assignment, avoid the model learning to noise samples, and thus improve the detection accuracy and convergence stability.
[0135] 5. Based on the label assignment result, respectively calculate the classification loss and the positioning loss and weightedly calculate the total loss, and update the model parameters of the GDALA-Net network according to the total loss through the back propagation algorithm.
[0136] To enhance detection accuracy, this invention improves the loss function by employing a KLD-based loss term to optimize the distribution consistency between candidate and target boxes, effectively improving the regression accuracy and robustness of sample matching for rotated targets. These optimizations help alleviate the shortcomings of traditional IoU metrics in detecting dense, overlapping, or rotated targets, thereby significantly improving the detection performance of oriented targets in remote sensing images.
[0137] like Figure 4 As shown in (a), traditional methods calculate the loss by independently regressing five parameters: the offset of the target center point (Δx, Δy), the changes in length and width (w', h'), and the angle offset Δθ. This approach ignores the inherent geometric correlations between parameters, such as the coupling between angle and length, width, and height, and the correlation between center point accuracy and target scale. This leads to mismatches between angle and width / height, center position drift, and even degeneration into unreasonable bounding boxes. Especially for small targets or approximately square targets, the instability of angle regression further amplifies the error, causing oscillations in the training process and fluctuations in prediction results, making it difficult to guarantee the overall geometric consistency of the bounding boxes. Figure 4 As shown in (b), this invention transforms the independent regression of real targets into a similarity assessment between Gaussian distributions, where the parameters of the Gaussian distribution include the mean and covariance matrix. The Gaussian distribution can transform the independent representation of the previous five parameters into a joint representation, eliminating potential ambiguity and inaccuracies, thus regressing the target distribution more accurately. The loss function between targets is calculated based on the similarity between the distributions. By comparison, it can be seen that the Gaussian distribution regression method used in this invention can simultaneously model the correlation between target position, scale, and angle parameters, alleviating the gradient sparsity problem in low-overlap regions, thereby improving the accuracy of bounding box prediction.
[0138] Specifically, classification loss cls Using Focal Loss:
[0139]
[0140] Where, p i Let α be the probability that the model predicts the i-th sample (i.e., the candidate anchor box) as the positive class, α be the balancing factor, and μ be a hyperparameter that adjusts the influence of easy and difficult samples. In this embodiment, μ is set to 2 and α is set to 0.25.
[0141] The positioning loss is realized based on Kullback-Leibler divergence, and is used for measuring the error degree of model prediction on target position.
[0142] The specific calculation method is:
[0143]
[0144] Wherein, f(D kl ) represents a nonlinear function used for transforming Kullback-Leibler divergence D kl (N p ||N t ), so that the loss is smoother, and tau is a hyperparameter used for loss adjustment. kl The application mainly uses two nonlinear functions, square root function sqrt(D kl ) and natural logarithm function ln(D total ) combined.
[0145]
[0146] The total loss function is:
[0147] L nwd =λ cls Loss l1 +λ loc
[0148] Wherein, lambda nwd , lambda l1 are weight coefficients, and both are 1.
[0149] S4: taking the remote sensing image to be detected as input, obtaining the directional detection frame containing the directional target category and the confidence by using the GDALA-Net network trained in step S3.
[0150] In order to verify the effectiveness of the improved strategy, the application further carries out ablation experiment, as shown in Table 1. Among them, the backbone network selects ResNet-50, a lightweight backbone structure, and the benchmark model is Rotated RetinaNet. Among them, ATSS* represents that on the basis of ATSS strategy, GJSD is used as the evaluation benchmark between target and positive sample. GDALA is the proposed Gaussian distribution perception label assignment strategy. KLD represents replacing the L1 loss function with KLD loss. The evaluation index mAP is used to verify the performance.
[0151] Table 1 ablation experiment
[0152]
[0153] Table 2 compares the results of the method of the present application with other methods. These methods include: SASM, ATSS, R-O. From the table, it is found that the method of the present application shows the best results in remote sensing target detection, in addition, the mAP of each category is also reported to measure the effect of the proposed model, wherein the data set categories are airplane (APL), airport (APO), baseball field (BF), basketball court (BC), bridge (BR), chimney (CH), dam (DAM), highway service area (ESA), highway toll station (ETS), golf course (GF), track and field (GTF), harbor (HA), overpass (OP), ship (SH), stadium (STA), storage tank (STO), tennis court (TC), train station (TS), vehicle (VE) and windmill (WM), and it is also explained that the method of the present application has good generalization ability.
[0154] Table 2 comparison of results of the method of the present application with other methods
[0155]
[0156]
[0157] In addition, the present application also carries out visualization experiment of detection result. Figure 5 is the ablation experiment result graph of remote sensing directed target detection. Figure 6 is the category detection result graph of remote sensing directed target detection. Figure 5 The comparison results with the baseline RotatedRetinaNet (abbreviated as R-O) are shown, which further illustrates the superiority of the GMLA strategy in target detection. The first column is the detection result on the baseline rotated retinanet; the second column is the detection result after adding the ATSS* module; the third column is the detection result using the GMLA strategy; and the fourth column is the joint detection result of GMLA and KLD loss. As shown in Figure 5 , in the second and third rows, the extreme detection results exist incomplete (ship) and mis-detection (ball court) cases. In Figure 6 , other category detection results on the DIOR-R data set are given. The detection effect on long and narrow targets (such as bridges) and the detection effect on small targets (such as vehicles) are shown. The visualization results show that the method of the present application has achieved obvious detection improvement on targets of different scales and directions, verifying the effectiveness of the Gaussian label assignment and regression strategy.
[0158] The above merely describes preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0159] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification and claims of the present application are intended to cover not only the inclusive but also the exclusive, for example, a process, method, system or product that comprises a list of steps or units is not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such process, method, system or product.
[0160] It should be understood that the above merely describes preferred embodiments of the present application and application of technical principles. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, reconfigurations and substitutions can be made by those skilled in the art without departing from the scope of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the specific embodiments described herein, and can include more other effective embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A remote sensing directed target detection method based on a Gaussian distribution sensing label allocation strategy, characterized in that, The steps are as follows: S1: Acquire the remote sensing image dataset and perform preprocessing; S2: Construct a GDALA-Net network with ResNet-50 as the backbone network and a feature pyramid as the neck network, and pre-train the GDALA-Net network. S3: Using the preprocessed remote sensing image as input, train the GDALA-Net network based on the label assignment strategy of Gaussian distribution perception; S4: Using the remote sensing image to be detected as input, the GDALA-Net network trained in step S3 is used to obtain directed detection boxes containing directed target categories and confidence scores.
2. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 1, characterized in that, The training processing method for step S3 is as follows: Using the training set in the preprocessed remote sensing image as input, multi-scale feature maps are extracted from the input remote sensing image using the ResNet-50 backbone network. Multi-level feature maps are generated using a feature pyramid network. At each level of the feature map, multiple dense prior anchor boxes are generated at each feature point location based on a preset aspect ratio and scale factor. The detection head outputs the classification score of each prior anchor box through the classification branch, predicts the bounding box parameters of each prior anchor box through the regression branch, and generates a candidate anchor box set by combining the classification score and the bounding box parameters. Labels are assigned based on candidate anchor boxes, classification scores, and labeled real directed target boxes in the training set, using Gaussian modeling and a mixture scoring mechanism. Based on the label assignment results, the GDALA-Net network parameters are updated via backpropagation using focus loss and KLD loss.
3. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 2, characterized in that, The method for label assignment based on Gaussian modeling and mixture scoring mechanism is as follows: 1) Model the directed target Gaussian distribution for both candidate anchor boxes and real directed target boxes; 2) Calculate the Gaussian distribution similarity measure between the Gaussian distributions corresponding to candidate anchor boxes and the Gaussian distributions corresponding to real directed target boxes based on Kullback-Leibler divergence and generalized Jensen-Shannon divergence; 3) Calculate the mixture score based on the candidate anchor boxes and classification scores, and calculate the combined localization and classification loss cost by combining the mixture score and Gaussian distribution similarity measure; 4) Label allocation is based on the combined loss cost of positioning and classification.
4. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 3, characterized in that, Using the center point of the candidate anchor frame as the mean of the Gaussian distribution, and calculating the covariance matrix of the Gaussian distribution using the width, height, and rotation angle of the candidate anchor frame, a Gaussian distribution N corresponding to the candidate anchor frame is constructed. p Using the center point of the true directed bounding box as the mean of the Gaussian distribution, and calculating the covariance matrix of the Gaussian distribution using the width, height, and rotation angle of the true directed bounding box, a Gaussian distribution N corresponding to the true directed bounding box is constructed. t .
5. The remote sensing directed target detection method based on a Gaussian distribution sensing label allocation strategy according to claim 3 or 4, characterized in that, Step 2) is implemented as follows: Based on the closed-form formula for the Kullback-Leibler divergence between multidimensional Gaussian distributions, calculate the Kullback-Leibler divergence in different directions between the Gaussian distributions corresponding to the candidate anchor boxes and the Gaussian distributions corresponding to the true directed target boxes. Then, calculate the Gaussian distribution N corresponding to the candidate anchor boxes based on the Kullback-Leibler divergence in different directions. p The Gaussian distribution N corresponding to the true directed bounding box t The generalized Jensen-Shannon divergence between them is used as the similarity measure.
6. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 5, characterized in that, The closed-form formula for the Kullback–Leibler divergence between the aforementioned multidimensional Gaussian distributions is as follows: Among them, D kl (N1||N2) represents the Kullback-Leibler divergence in the direction from Gaussian distribution N1 to Gaussian distribution N2, Tr(·) represents the trace of the matrix, μ1 and μ2 are the means of Gaussian distribution N1 and Gaussian distribution N2 respectively, and Σ1 and Σ2 are the covariance matrices of Gaussian distribution N1 and Gaussian distribution N2 respectively. The method for calculating the similarity metric is as follows: GJSD(N p ||N t )=(1-λ)D kl (N p ||N t )+λD kl (N t ||N p ) Where λ is the weighting coefficient used to balance the KL divergence in the two directions, and D kl (N p ||N t ) represents the Gaussian distribution N corresponding to the candidate anchor boxes. p The Gaussian distribution N corresponding to the true directed bounding box t Kullback–Leibler divergence in direction, D kl (N t ||N p N is the Gaussian distribution corresponding to the true directed bounding box. t The Gaussian distribution N corresponding to the candidate anchor box p Kullback–Leibler divergence in the direction.
7. The remote sensing directed target detection method based on a Gaussian distribution sensing label allocation strategy according to any one of claims 3, 4, or 6, characterized in that, The method for calculating the combined localization and classification loss cost by combining mixture score and Gaussian distribution similarity metric is as follows: calculate the mixture cost C based on the mixture score. MS According to the similarity metric GJSD(N) p ||N t )Calculate the geometric cost C g The periodic coefficient is introduced to calculate the location and classification combination loss C(x,y) based on the mixed cost and geometric cost.
8. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 7, characterized in that, The method for calculating the hybrid score based on candidate anchor boxes and classification scores is as follows: MS i =IoU(A i ,gt i )*Cls(A i ) Among them, A i Represents the candidate anchor box, gt i For candidate anchor box A i The matched true directed bounding boxes, IoU(·) represents the intersection-union ratio. Cls(A i ) represents candidate anchor box A i Category score; The calculation of the hybrid cost C MS The method is as follows: The computational geometric cost C g Method: The method for calculating the combined cost of location and classification is as follows: in, For periodic coefficients, iter represents the current iteration number. max R represents the maximum number of iterations during model training. g Let C be the region modeled by Gaussian, (x, y) be the coordinates of the center point of the candidate anchor box, and C be the region modeled by Gaussian. g (x,y) represents the geometric cost of the candidate anchor box centered at (x,y), C MS (x,y) represents the mixed cost of the candidate anchor box centered at (x,y), and C(x,y) represents the combined localization and classification loss cost of the candidate anchor box centered at (x,y). The KLD loss is: Where τ is a hyperparameter.
9. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 8, characterized in that, The method for tag allocation based on the combined loss cost of location and classification is as follows: For large-scale targets of predefined size in preprocessed remote sensing images, from the Gaussian region R g Select the top k candidate anchor boxes with the lowest combined loss cost of localization and classification as positive samples: For small-scale targets of predefined size in the preprocessed remote sensing image, the Gaussian region R g All candidate anchor boxes within the specified range are considered positive samples; If a candidate anchor frame overlaps with multiple criteria, it is classified as an ignored region. Candidate anchor boxes outside the Gaussian region are all assigned to the negative sample set R. neg It is used to suppress background interference.
10. The remote sensing directed target detection method based on Gaussian distribution sensing label allocation strategy according to claim 9, characterized in that, The multiple standards mentioned overlap as follows: If candidate anchor box A exists i The center point (x, y) lies in the Gaussian region R. g However, the cost of combining location and classification is C(x,y)>τ1; If candidate anchor box A exists i The center point (x, y) lies in the Gaussian region R. g In addition, the cost of combining location and classification is C(x,y) > τ2; τ1 and τ2 are the upper and lower thresholds of the cost for tag allocation, respectively.