Orientation target detection method based on orientation ellipse and ranking guidance alignment loss

By using the method of guiding alignment loss by directing alignment loss, the angular discontinuity and classification-positioning inconsistency in directional object detection are solved, and higher detection accuracy and stability are achieved, and are suitable for multi-directional object detection tasks such as remote sensing images.

CN120339846AActive Publication Date: 2025-07-18NANJING UNIV OF INFORMATION SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510771962.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-18
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing directional object detection model has problems such as discontinuous angles, blurred corner points and uncertain directions, and the classification and positioning tasks are inconsistent, which affects the detection accuracy and stability.

Method used

Directed ellipses are used to replace traditional rectangular bounding boxes, combined with ranking guidance alignment loss and multi-candidate matching mechanism, feature extraction and prediction are performed through Transformer model, Hungarian matching algorithm is introduced to optimize sample quality allocation, and directional ellipse regression is simplified through Cholesky decomposition.

Benefits of technology

The prediction stability and directional accuracy of directional target detection are improved, the consistency between classification and positioning is enhanced, the angular discontinuity and classification-positioning misalignment are solved, and the detection accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339846A_ABST
    Figure CN120339846A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of target detection and recognition, and discloses a directional target detection method based on directional ellipse and ranking guidance alignment loss, which comprises the following steps of: extracting multi-scale features of an image by using Transform; generating a candidate frame through query vector and feature interaction, introducing a multi-candidate matching mechanism to improve positive sample quality, calculating ellipse positioning loss by using KL divergence, constructing a loss function in combination with classification score and angle consistency, and guiding a model to focus high-quality prediction; and finally, introducing a dynamic angle adjustment factor to adapt to targets in different shapes. According to the method, the problem that the traditional bounding box angle is discontinuous and fuzzy is solved through directional ellipse representation, the consistency of classification and positioning is enhanced while the positioning precision and the direction sensing ability are improved, and the method is suitable for multi-direction target detection tasks such as remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to long-term target detection and recognition technology, and specifically relates to an oriented target detection method based on oriented ellipse and ranking-guided alignment loss. Background Art

[0002] Oriented target detection is an important research direction in the field of target detection. Compared with traditional horizontal bounding box detection methods, oriented target detection can more accurately express the position and orientation information of the target by using means such as rotated bounding boxes or point set positioning. This feature is particularly crucial in natural scenes (such as vehicles in traffic scenes) and remote sensing images (such as airplanes and ships), because the targets in these scenes often show characteristics such as multi-angle distribution, dense arrangement, or significant aspect ratio differences. Traditional horizontal detection bounding boxes are prone to introducing a large amount of background interference in this case, and it is difficult to closely fit the actual shape of the object, which limits the positioning accuracy and the ability to understand the scene.

[0003] Current oriented target detection can be divided into the following categories of methods according to the bounding box division: oriented bounding box method OBB, quadrilateral bounding box method QBB, and point set representation method PS; in the oriented target detection task, the initially proposed methods are basically implemented through rotated bounding boxes. For example, R²CNN proposed by Jiang et al. introduces rotated candidate boxes on the basis of Faster R-CNN, improving the detection accuracy of targets at any angle (such as scene text). RRPN designed by Ma et al. introduces rotated anchor boxes and rotated RoIPooling, enabling the model to directly generate inclined candidate boxes with direction information and accurately extract features. RoITransformer proposed by Ding et al. learns the offset from horizontal RoI to rotated RoI through a spatial geometric transformation module (RRoILearner) and combines rotated RoIAlign, effectively improving the positioning accuracy of rotated targets. To improve the detection accuracy and robustness, Yang et al. proposed R3Det, which alleviates the angle sensitivity problem under large aspect ratio targets through a progressive regression mechanism from horizontal boxes to rotated boxes, combined with a feature reconstruction module and an approximate SkewIoU loss. Han et al. proposed S²A-Net, aiming at the problem of misalignment between features and anchor boxes, introducing an anchor box refinement and feature alignment mechanism (FAM), and a direction detection module (ODM) with a classification / regression separation design, significantly improving the model's expression ability and positioning consistency.

[0004] However, in the above-mentioned existing methods for bounding boxes, there may be problems with discontinuous angles. For example, for bounding boxes of 0° and 359°, due to calculation losses, the final results obtained may vary greatly. The QBB method faces the problem of ambiguous corner representation. Different orders of point representation may lead to non-correspondence between predictions and GT, resulting in errors. The point set representation method PS has uncertainty in direction perception, which may lead to large regression errors.

[0005] In addition, there is a common problem in current oriented object detection models that the classification and localization tasks are inconsistent, that is, samples with high classification confidence do not necessarily have good localization quality, and vice versa, seriously affecting the accuracy and stability of the detector. Summary of the Invention

[0006] Object of the Invention: The object of the present invention is to solve the deficiencies existing in the prior art and provide an oriented object detection method based on an oriented ellipse and a ranking-guided alignment loss. The oriented ellipse is used to replace the traditional rectangular bounding box, and its advantages of continuity, rotational symmetry and parameter compactness are utilized to solve problems such as discontinuous angles, corner ambiguity and direction uncertainty. While maintaining the geometric expression ability, it provides higher prediction stability and direction accuracy. To enhance the consistency between classification and localization, a ranking-guided alignment loss is introduced as the main loss function, and a multi-candidate matching mechanism MCMS based on the Hungarian matching algorithm is used to improve the positive and negative sample quality allocation, so as to achieve end-to-end accurate detection of oriented objects.

[0007] Technical Solution: An oriented object detection method based on an oriented ellipse and a ranking-guided alignment loss of the present invention includes the following steps:

[0008] Step 1: For the input image, first perform feature extraction through a pre-trained backbone network to generate an initial feature map ;

[0009] After the obtained initial feature map is flattened and embedded with position encoding, it is input into the Transformer model. The encoder Encoder of the Transformer model models the global context relationship through the self-attention mechanism, enabling the features at each position to dynamically aggregate the information of all other positions in the image, and finally obtaining a set of image Token representations containing global context information, which are used for the localization and recognition of targets in the subsequent decoder Decoder ;

[0010] Step 2: For the image Token representation obtained in Step 1, first, in the decoder of the Transformer model, a set of learnable query vectors Query are used to interact with the output features of the encoder. That is, through the cross-attention mechanism, it focuses on potential regions of interest and gradually decodes the preliminary prediction results, including the preliminary predicted target bounding boxes and the preliminary classification scores and the localization scores ; then, the predicted target bounding boxes are inversely transformed into the oriented ellipse target bounding boxes ;

[0011] Then, for the prediction results of different layers of the decoder, a multi-candidate matching strategy MCMS is introduced, and the parameter is used to expand the number of ground-truth bounding boxes, thereby enhancing the positive sample coverage and obtaining the ground-truth boxes;

[0012] Finally, a cost matrix is constructed according to multiple cost metrics such as classification, localization, and angle, and the one-to-one matching between the predicted boxes and the ground-truth boxes is completed through the Hungarian matching method;

[0013] Step 3: For the classification scores and the localization scores obtained in Step 2, the classification loss and the localization loss are calculated and optimized respectively. The specific method is as follows:

[0014] For the obtained classification scores and localization scores, a ranking-based dynamic weight is introduced, and the results are unified through the joint alignment loss to obtain the final loss result;

[0015] At the same time, considering the possible influence of targets with different aspect ratios on the angle, a dynamic adjustment factor is used for adjustment, and the corresponding angle loss is calculated.

[0016] The oriented ellipse target bounding box of the present invention is denoted as , where is the coordinate of the center point of the oriented ellipse, and is the result of the Cholesky decomposition of the covariance matrix of the oriented ellipse using the symmetric positive definiteness of the covariance matrix, so as to better simplify the calculation and perform regression on the oriented ellipse. In addition, the present invention also proposes relevant improvement methods for the classification-localization misalignment problem, and proposes multi-candidate matching MCMS and ranking-guided classification loss to alleviate the classification-localization misalignment problem in oriented object detection.

[0017] Furthermore, the detailed process of Step 1 is as follows:

[0018] Step 1.1. Input the image Generate a set of 5-layer multi-scale feature maps through a pre-trained backbone network Backbone (such as ResNet50, Swin Transformer) , where each layer of the feature map ; To reduce the computational complexity, usually select multiple layers of feature maps and compress the channels to dimensions through convolution to obtain the initial feature map ;

[0019] The initial feature map is flattened into a sequence of dimensional space dimensions , and sinusoidal positional encoding is superimposed, specifically as follows:

[0020] ;

[0021] Step 1.2. The encoder Encoder of the Transformer model consists of 6 layers of the same structure stacked, and each layer structure includes multi-head self-attention MSA and a feed-forward network FFN. For the input of the th layer, its corresponding output is calculated as follows:

[0022] ;

[0023] ;

[0024] Among them, , is the output of the th layer encoder Encoder after adding residual and then normalizing through the multi-head self-attention mechanism, is the final result of the th layer encoder Encoder. MSA splits the dimensional input into subspaces of dimensions through

[0025] heads to calculate the attention weights

[0026] Among them, is the weight matrix for linear transformation in the self-attention mechanism;

[0027] Finally, the encoder of the Transformer model outputs , aggregating the global context information, that is, F0 outputs F6 after being processed by six layers of the encoder.

[0028] Furthermore, the detailed process of step 2 is as follows:

[0029] Step 2.1, in the decoding stage, a set of learnable query vectors interact with the multi-scale features output by the encoder through multiple layers of the Transformer Decoder. The processing flow of each layer of the decoder is as follows: first, calculate through deformable cross-attention, then update the query vectors through self-attention (SA), and finally output through a feed-forward network (FFN), specifically as follows:

[0030] ;

[0031] ;

[0032] ;

[0033] Among them, the output of each layer of the decoder is used to predict the target bounding box through a linear projection layer, forming a preliminary classification score and a localization score ; is the center point coordinate of the predicted box,

[0034] Step 2.2, for the prediction results obtained from each layer of the decoder in step 2.1, the ground-truth of the real box is copied through the multi-candidate matching mechanism MCMS to increase the number of results obtained by the subsequent Hungarian matching method, so as to achieve the expansion of positive samples and obtain the real box

[0035] ;

[0036] The ground truth of the real box after being expanded by the MCMS strategy is denoted as , , where is the set expansion multiple;

[0037] Step 2.3: After expanding the ground truth boxes in Step 2.2, construct a cost matrix for the prediction results and the expanded positive sample ground truth boxes , where denotes the th predicted bounding box and the th ground truth box matching cost, and its calculation formula is:

[0038] ;

[0039] wherein, represents the ranking-guided classification loss, represents the localization loss, represents the angle loss, respectively represent the corresponding balancing weights;

[0040] Solve the optimal bipartite matching through the Hungarian algorithm to ensure that each expanded ground truth box is assigned at most one prediction result, and the unassigned predictions are regarded as negative samples.

[0041] Furthermore, in Step 2, an oriented elliptical object bounding box is obtained based on the Gaussian distribution, specifically as follows:

[0042] For the OBB bounding box representation , first convert it to a Gaussian distribution , where represents the center coordinates of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution, and the calculation of the covariance matrix is as follows:

[0043] ;

[0044] Simplify to ;

[0045] represents the rotation matrix, represents the eigenvalue diagonal matrix;

[0046] If and only if and , is a positive definite matrix, the equiprobability contour of the Gaussian distribution is an ellipse, and directly regressing requires restricting the condition, which may lead to problems such as unstable gradients and poor training effects.

[0047] To avoid the constrained optimization problem, consider using Cholesky decomposition to parameterize the matrix, specifically as follows:

[0048] ;

[0049] ;

[0050] By performing Cholesky decomposition, it is only necessary to ensure that to satisfy positive definiteness.

[0051] Furthermore, the detailed process of step 3 is as follows:

[0052] Step 3.1: Calculate the distance between the predicted bounding box and the ground truth bounding box using the Kullback-Leibler divergence to obtain the localization loss , and the formula is as follows:

[0053] ; represents the predicted bounding box, represents the ground truth bounding box;

[0054] where and represent the coordinates of the predicted bounding box respectively, and represent the corresponding covariance matrices respectively, is the operator for calculating the trace of the matrix;

[0055] The coordinates of the predicted bounding box and the corresponding covariance matrix are both calculated based on the oriented ellipse, as shown in the following formula:

[0056] ;

[0057] ;

[0058] Step 3.2: The calculation formula for the ranking-guided classification loss is:

[0059] ;

[0060] In the above formula, is the balanced weight of the ranking-based IoU-Classification, and the calculation formula is:

[0061] ;

[0062] In the above formula, is the introduced angle consistency term, which takes into account the influence of the angle on the result, represents the ranking of the joint quality metric among all predictions, represents the dynamic weight based on the ranking . is a parameter that controls the exponential decay, is a joint quality metric for balancing classification and localization;

[0063] ;

[0064] Among them, is a joint quality metric for balancing classification and localization, is a hyperparameter for balancing the classification score and the localization score, represents the classification score, represents the localization score;

[0065] The calculation of is as follows: ;

[0066] Step 3.3. For the angular loss , consider dynamically adjusting the target with different aspect ratios through the dynamic adjustment factor as follows:

[0067] ;

[0068] Among them, and The calculation formulas of are as follows:

[0069] ;

[0070] In the above formula, and are the two eigenvalues corresponding to the covariance matrix .

[0071] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:

[0072] 1. In terms of the bounding box representation, by using the oriented ellipse, the present invention can more clearly and explicitly represent the position, size and direction of the target, effectively alleviating the problems that may occur in other types of bounding boxes under oriented object detection, such as the angle discontinuity problem and the square-like problem that may occur in the OBB model, the fuzzy representation problem that may occur in the QBB model, and the problem of difficult handling of isotropic objects that may occur in the point set model. It can be seen that the use of the oriented ellipse target box can effectively improve the ability of the model in bounding box alignment.

[0073] 2. In terms of the classification-localization misalignment problem, through the RGA Loss, the present invention correlates the classification score with the localization score, realizing the IoU-awareness of the classification score. By introducing the ranking-based dynamic weight, the bounding boxes with relatively better joint quality can be effectively screened out, and the bounding boxes with only a single good result can be filtered out.

[0074] 3. In terms of the positive and negative sample allocation problem, based on the original model, the present invention expands positive samples through the MCMS strategy, effectively alleviating the problem of positive sample shortage caused by the DETR model. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 is the model structure diagram of the present invention;

[0076] Figure 2 is the schematic diagram of the oriented ellipse representation of the present invention;

[0077] Figure 3 is the calculation flow chart of the classification loss of the present invention;

[0078] Figure 4 is the target detection result diagram obtained by using the prior art in the embodiment;

[0079] Figure 5 is the target detection result diagram obtained by using the technical solution of the present invention in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0080] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0081] The present invention replaces the traditional rectangular bounding box structure with an oriented ellipse, and utilizes its advantages of continuity, rotational symmetry and parameter compactness to solve problems such as discontinuous angles, corner point ambiguity and direction uncertainty. Compared with the OBB model being easily affected by angle jumps, the QBB model having problems with point order confusion, and the Point Set method being difficult to stably perceive directionality, while maintaining the geometric expression ability, it provides higher prediction stability and direction accuracy. In addition, to enhance the consistency between classification and positioning, the Rank-Guided Align Loss (RGA Loss) is introduced as the main loss function, and a multi-candidate matching mechanism (Multi-Candidate Matching Strategy, MCMS) based on the Hungarian algorithm is designed to improve the positive and negative sample quality allocation, so as to achieve end-to-end accurate detection of oriented targets.

[0082] As Figure 1 shown, the oriented target detection method based on an oriented ellipse and a rank-guided alignment loss of the present invention includes the following steps:

[0083] Step 1. For the input image, first perform feature extraction through a pre-trained backbone network to generate an initial feature map ;

[0084] The obtained initial feature map After being flattened and embedded with positional encoding, it is input into the Transformer model. The encoder of the Transformer model models the global context relationship through the self-attention mechanism, enabling the features at each position to dynamically aggregate the information of all other positions in the image, and finally obtaining a set of image Token representations containing global context information. ;

[0085] Step 2: For the image Token representations obtained in Step 1, first in the decoder of the Transformer model, a set of learnable query vectors Query are used to interact with the output features of the encoder, that is, through the cross-attention mechanism, focus on potential regions of interest, and gradually decode the preliminary prediction results, including the preliminary predicted target bounding boxes. And the preliminary classification scores and localization scores ; Then, the predicted target bounding boxes are inversely transformed into oriented elliptical target bounding boxes ;

[0086] Then, for the prediction results of different layers of the decoder, a multi-candidate matching strategy MCMS is introduced, and the parameter is used to expand the number of ground-truth bounding boxes, thereby enhancing the positive sample coverage and obtaining the ground-truth boxes;

[0087] Finally, a cost matrix (matching matrix) is constructed according to multiple cost metrics such as classification, localization, and angle, and the one-to-one matching between the predicted boxes and the ground-truth boxes is completed through the Hungarian matching method;

[0088] Step 3: For the classification scores and localization scores obtained in Step 2, the classification loss and localization loss are calculated respectively for optimization. The specific method is as follows:

[0089] For the obtained classification scores and localization scores, a ranking-based dynamic weight is introduced, and the results are unified through the joint alignment loss to obtain the final loss result; at the same time, considering the influence of targets with different aspect ratios on the angle, a dynamic adjustment factor is used for adjustment, and the corresponding angle loss is calculated.

[0090] In this embodiment, Transformer is used as the backbone network to enhance the understanding of rotated targets in complex scenes through global context modeling. A directional ellipse in the form of a Gaussian distribution is used to replace the traditional anchor box or rotated box, and the covariance matrix is parameterized and regressed using Cholesky decomposition, effectively improving the model convergence stability and direction expression ability. At the same time, a Rank-Guided Align Loss that combines KL divergence, ranking guidance mechanism, and angle consistency modeling is designed, and a multi-candidate matching strategy (MCMS) is introduced to solve problems such as classification-localization inconsistency, discontinuous angles, and insufficient positive samples in traditional methods, achieving a comprehensive breakthrough in the stability, accuracy, and geometric consistency of remote sensing target detection.

[0091] The detailed process of step 1 in this embodiment is as follows:

[0092] Step 1.1. Input image Generate a set of 5-layer multi-scale feature maps through the pre-trained backbone network Backbone , where each layer of feature map ; Select multiple layers of feature maps and compress the channels to dimensions through convolution to obtain the initial feature map ;

[0093] The initial feature map is flattened into a sequence of dimensions in the spatial dimension , and sinusoidal positional encoding is superimposed, specifically as follows:

[0094] ;

[0095] Step 1.2. The encoder Encoder of the Transformer model consists of 6 identical structures stacked, and each layer structure includes multi-head self-attention MSA and a feed-forward network FFN. For the input of the layer, its corresponding output is calculated as follows:

[0096] ;

[0097] ;

[0098] Among them, , is the output of the m-th layer encoder Encoder after adding residual and then normalizing through the multi-head self-attention mechanism, is the final result of the m-th layer encoder Encoder, and MSA is passed through The head splits the -dimensional input into -dimensional subspaces and calculates the attention weights :

[0099] ;

[0100] Among them, is the weight matrix for linear transformation in the self-attention mechanism;

[0101] Finally, the encoder Encoder of the Transformer model outputs , aggregating global context information.

[0102] The detailed process of step 2 of this embodiment is as follows:

[0103] Step 2.1. In the decoding stage, a set of learnable query vectors interact with the multi-scale features output by the encoder through multiple layers of Transformer Decoder. The processing flow of each layer of the decoder is as follows: First, through deformable cross-attention calculation, then update the query vectors through self-attention, and finally output through the feed-forward network FFN, specifically as follows:

[0104] ;

[0105] ;

[0106] ;

[0107] Among them, the output of each layer of the decoder is used to predict the target bounding box through a linear projection layer, forming a preliminary classification score and a localization score ; is the center point coordinate of the predicted box, is the parameter obtained by Cholesky decomposition of the covariance matrix of the predicted box;

[0108] Step 2.2. For the prediction results obtained from each layer of the Decoder in step 2.1, the true box Ground-truth is copied through the multi-candidate matching mechanism MCMS to achieve the augmentation of positive samples, and then the true box is obtained, specifically as follows:

[0109] Denote the ground truth boxes augmented by the Multi-Candidate Matching Mechanism (MCMS) as , ;

[0110] Step 2.3: After augmenting the ground truth boxes in Step 2.2, construct a cost matrix for the prediction results and the augmented positive ground truth boxes , where represents the -th predicted box and the -th ground truth box , and its calculation formula is:

[0111] ;

[0112] where represents the ranking-guided classification loss, represents the localization loss, represents the angular loss, respectively represent the corresponding balance weights;

[0113] Solve the optimal bipartite matching through the Hungarian matching method to ensure that each augmented ground truth box is assigned at most one prediction result, and the unassigned predictions are regarded as negative samples.

[0114] As Figure 2 shown, in this embodiment, an oriented ellipse target bounding box is obtained based on Gaussian distribution transformation , specifically as follows:

[0115] For the OBB bounding box representation , first convert it into a Gaussian distribution , where represents the center coordinates of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution, and the calculation of the covariance matrix is as follows:

[0116] ;

[0117] Simplify to ;

[0118] represents the rotation matrix, represents the diagonal matrix of eigenvalues;

[0119] If and only if and , is a positive definite matrix, and the equiprobability contour of the Gaussian distribution is an ellipse;

[0120] To avoid the constrained optimization problem, consider using Cholesky decomposition to parameterize the matrix as follows:

[0121] ;

[0122] ;

[0123] Through Cholesky decomposition, only need to ensure that , then it satisfies is positive definite.

[0124] The above process uses the oriented ellipse representation modeled by Gaussian distribution, expresses the target bounding box in the form of a covariance matrix, and parameterizes the covariance matrix through Cholesky decomposition, which not only effectively guarantees the positive definiteness and the stability of model regression, but also naturally embeds the direction information into the overall shape expression of the ellipse, avoiding the discontinuous problem caused by explicit angle regression. In addition, introducing the Kullback-Leibler (KL) divergence as a distribution difference metric can accurately measure the comprehensive difference in shape and direction between the predicted ellipse and the true ellipse, thus realizing a more stable, continuous and geometrically consistent target regression optimization process.

[0125] As Figure 3 shown, the detailed process of step 3 in this embodiment is as follows:

[0126] Step 3.1: Calculate the distance between the predicted box and the true box using the Kullback-Leibler divergence to obtain the localization loss , and the formula is as follows:

[0127] ;

[0128] Among them, and respectively represent the coordinates of the predicted box, and respectively represent the corresponding covariance matrices, is the operator for calculating the trace of the matrix;

[0129] The coordinates of the predicted box and the corresponding covariance matrix are both calculated based on the oriented ellipse, and the specific formulas are as follows:

[0130] ;

[0131] ;

[0132] Step 3.2: The calculation formula of the ranking-guided classification loss is:

[0133] ;

[0134] In the above formula, is the balanced weight of ranking-based IoU-Classification, and the calculation formula is:

[0135] ;

[0136] In the above formula, is the introduced angle consistency term, represents the ranking of the joint quality metric among all predictions, represents based on the ranking of the dynamic weight, is a parameter that controls the exponential decay, is the joint quality metric that balances classification and localization;

[0137] ;

[0138] Among them, is the joint quality metric that balances classification and localization, is the hyperparameter that balances the classification score and the localization score, represents the classification score, represents the localization score;

[0139] The calculation of is:

[0140] Step 3.3. For the angle loss , consider dynamically adjusting different aspect ratio targets through the dynamic adjustment factor as follows:

[0141] ;

[0142] Among them, and The calculation formulas are as follows:

[0143] ;

[0144] In the above formula, and are the two eigenvalues corresponding to the covariance matrix .

[0145] To verify the technical effect of the present invention, in this embodiment, the technical solution of the present invention is compared with the prior art based on the DOTA dataset, and the comparison results are as shown in Figure 4 , Figure 5 and Table 1.

[0146] Table 1

[0147] Figure 4 is the target detection result obtained by using the existing technology, Figure 5 is the target detection obtained by using the technical solution of the present invention.

[0148] It can be seen from the above embodiments that the present invention introduces a multi-candidate matching strategy MCMS in the positive and negative sample allocation. By expanding multiple candidate matching pairs on the basis of the original ground truth boxes and combining with the Hungarian matching method, a one-to-one global optimal allocation of the prediction results and the ground truth boxes is realized. It avoids the rigid dependence of the fixed threshold on the sample selection and can achieve a more reasonable matching between targets of different scales and different angles. In addition, after the positive sample allocation is completed, a ranking guidance mechanism is introduced to calculate the joint quality index according to the classification score, localization accuracy and angle consistency, and different optimization weights are assigned to the positive samples, so as to highlight the training value of high-quality positive samples and suppress the negative impact of mis-matched samples. The present invention not only improves the accuracy and stability of sample allocation, but also significantly enhances the model's attention ability to high-confidence targets, providing stronger guarantee for the final detection performance.

Claims

1. A directional object detection method based on directional ellipse and ranking-guided alignment loss, characterized in that Including the following steps: Step 1. For the input image, first perform feature extraction through a pre-trained backbone network to generate an initial feature map ; The obtained initial feature map After being flattened and embedded with positional encoding, it is input into the Transformer model. The encoder of the Transformer model models the global context relationship through the self-attention mechanism, enabling the features at each position to dynamically aggregate the information of all other positions in the image, and finally obtaining a set of image token representations containing global context information ; Step 2. For the image Token representation obtained in Step 1, first in the decoder of the Transformer model, a set of learnable query vectors Query are used to interact with the output features of the encoder, that is, through the cross-attention mechanism, focus on potential regions of interest, and gradually decode the preliminary prediction results, which include the preliminary predicted target bounding boxes and the preliminary classification scores and the localization scores ; then the predicted target bounding boxes are inversely transformed into the oriented ellipse target bounding boxes ; Then, for the prediction results of different layer decoders Decoder, a multi-candidate matching strategy MCMS is introduced, and the parameter is used to expand the number of ground truth bounding boxes, thereby enhancing the coverage of positive samples to obtain ground truth boxes. Finally, a cost matrix is constructed based on the classification, localization, and angular cost metrics , and one-to-one matching between the predicted bounding boxes and the ground truth bounding boxes is completed through the Hungarian matching method; Step 3: For the classification scores obtained in Step 2 and the localization scores , calculate the classification loss and the localization loss respectively for optimization. The specific method is as follows: For the obtained classification scores and localization scores, introduce dynamic weights based on ranking into them, unify the results through the joint alignment loss to obtain the final loss result; at the same time, considering the influence of targets with different aspect ratios on the angle, use a dynamic adjustment factor to make adjustments and calculate the corresponding angle loss .

2. The oriented object detection method based on the oriented ellipse and the ranking-guided alignment loss according to claim 1, wherein The detailed process of step 1 is as follows: Step 1.1: Input the image Generate a set of 5 - layer multi - scale feature maps through the pre - trained backbone network Backbone , where each layer of the feature map ; Select multiple layers of feature maps and pass them through convolution to compress the channels to dimensions to obtain the initial feature map ; Initial feature map is flattened into a sequence of dimensional space dimensions , and sinusoidal positional encoding is superimposed , specifically as follows: ; Step 1.

2. The encoder of the Transformer model consists of 6 identical structures stacked, and each layer structure includes a multi-head self-attention (MSA) and a feed-forward network (FFN). For the input of the th layer, its corresponding output is calculated as follows: ; ; Among them, , is the output after adding residual and normalization to the output of the m-th layer encoder Encoder through the multi-head self-attention mechanism, is the final result of the m-th layer encoder Encoder. MSA splits the -dimensional input into subspaces through heads and calculates the attention weights : : ; Among them, is the weight matrix for linear transformation in the self-attention mechanism; Finally, the encoder of the Transformer model outputs , aggregating the global context information.

3. The oriented object detection method based on an oriented ellipse and a ranking-guided alignment loss according to claim 1, wherein The detailed process of step 2 is as follows: Step 2.1, in the decoding stage, a set of learnable query vectors interact with the multi-scale features output by the encoder through a multi-layer Transformer Decoder. The processing flow of each layer of the decoder is as follows: first, calculate through deformable cross-attention, then update the query vectors through self-attention, and finally output through the feed-forward network FFN, specifically as follows: ; ; ; Among them, the output of each layer of decoder predicts the target bounding box through a linear projection layer to form a preliminary classification score and a localization score ; is the center point coordinate of the predicted box, and is the parameter obtained by Cholesky decomposition of the covariance matrix of the predicted box; Step 2.

2. For the prediction results obtained by each layer of the Decoder in Step 2.1, the true box Ground-truth is copied through the multi-candidate matching mechanism MCMS , to achieve the augmentation of positive samples, and then the true box is obtained. The specific steps are as follows: ; The true bounding boxes augmented using the multi-candidate matching mechanism MCMS are denoted as , ; Step 2.3: After expanding the ground truth boxes in Step 2.2, construct a cost matrix for the prediction results and the expanded positive sample ground truth boxes , where represents the th predicted box and the th ground truth box matching cost, and its calculation formula is as follows: ; Among them, represents the classification loss guided by ranking, represents the localization loss, represents the angle loss, respectively represent the corresponding balance weights; Solve the optimal bipartite graph matching through the Hungarian matching method to ensure that each augmented true box is assigned at most one prediction result, and the unmatched predictions are regarded as negative samples.

4. The oriented object detection method based on oriented ellipse and ranking-guided alignment loss according to claim 1 or 3, characterized in that Step 2 is based on the Gaussian distribution transformation to obtain an oriented elliptical object bounding box, specifically as follows: For the OBB bounding box representation , first convert it to a Gaussian distribution , where represents the center coordinates of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution, and the covariance matrix is calculated as follows: ; Convert to ; represents a rotation matrix, represents a diagonal matrix of eigenvalues; if and only if and then is a positive definite matrix, and the equiprobability contour of the Gaussian distribution is an ellipse; Finally, the Cholesky decomposition is used to parameterize the matrix, specifically as follows: ; ; Through Cholesky decomposition, it is only necessary to ensure that , then it satisfies is positive definite.

5. The directional object detection method based on directional ellipse and ranking-guided alignment loss according to claim 1, wherein The detailed process of step 3 is as follows: Step 3.

1. Calculate the distance between the predicted bounding box and the ground truth bounding box using the Kullback-Leibler divergence to obtain the localization loss , and the formula is as follows: ; Among them, and respectively represent the coordinates of the prediction box, and respectively represent the corresponding covariance matrix, is the operator for calculating the trace of the matrix; represents the prediction box, represents the ground truth box; The coordinates of the prediction box and the corresponding covariance matrix are both calculated based on the oriented ellipse, as shown in the following formula: ; ; Step 3.2, Ranking-guided Classification Loss The calculation formula is as follows: ; In the above formula, is the balanced weight of ranking-based IoU-Classification, and its calculation formula is: ; In the above formula, is the introduced angle consistency term, represents the ranking of the joint quality metric among all predictions, represents the dynamic weight based on the ranking and is the parameter that controls the exponential decay, is the joint quality metric that balances classification and localization, as follows: ; Among them, is a hyperparameter for balancing the classification score and the localization score, represents the classification score, represents the localization score; The calculation of ; Step 3.

3. For the angle loss , consider dynamically adjusting the targets with different aspect ratios through a dynamic adjustment factor as follows: ; Among them, and The calculation formulas are as follows: ; In the above formula, and are the two eigenvalues corresponding to the covariance matrix .

Citation Information

Patent Citations

  • Methods and system for multi-target tracking

    CN111527463A

  • Multi-target tracking method and device under aerial view angle

    CN115984586A

  • Cow single target tracking method

    CN117671565A

  • Remote sensing image target detection method based on alignment convolution and ellipse loss function

    CN118506196A

  • Methods and system for multi-target tracking

    US20200126239A1