Rotating target detection method based on dynamic reference axis

By introducing dynamic reference axis, region division mechanism and angle conversion module in rotation object detection, the boundary discontinuity problem in directional object detection is solved, the accuracy and stability of angle prediction are improved, and the robustness and generalization ability of the model are enhanced.

CN120219729AActive Publication Date: 2025-06-27XIAMEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510704489.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-06-27
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

When the prior art deals with the problem of boundary discontinuity caused by angle periodicity in directional object detection, it is difficult for the model to stably predict the angle of the target near the boundary during training, resulting in low prediction accuracy and stability.

Method used

Using a rotation object detection method based on the dynamic reference axis, the angle domain is divided into multiple regions by introducing a region division mechanism and an angle conversion module, and the target belongs to the region is classified through the region predictor, so as to perform angle regression in the stable region and dynamically adjust the reference axis to alleviate the problem of boundary discontinuity.

Benefits of technology

It significantly improves the accuracy and stability of angle prediction, enhances the robustness and generalization ability of the model in complex scenarios, and effectively alleviates the problem of boundary discontinuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219729A_ABST
    Figure CN120219729A_ABST
Patent Text Reader

Abstract

The invention relates to a rotating target detection method based on a dynamic reference axis, which decomposes an original angle prediction problem based on a single reference axis into an angle prediction problem based on a plurality of reference axes, and transfers a boundary discontinuity problem to a stable region for processing through region division and an angle conversion mechanism. Through region division, the model can carry out angle regression in a stable region, angle prediction errors near a boundary are avoided, the angle prediction precision is significantly improved, and a region division mechanism enables the model to better adapt to targets of different angles, and the generalization ability of the model in a complex scene is enhanced. And flexible switching among different reference axes can be realized, so that the problem of boundary discontinuity can be further alleviated, and the flexibility of angle prediction can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically relates to a method for rotating target detection based on a dynamic reference axis. Background Art

[0002] Compared with horizontal target detection, the major challenge faced by oriented target detection is the problem of boundary discontinuity caused by angular periodicity. Currently, the solutions to this problem can be roughly divided into two categories: The first category of methods focuses on position representation and models the position , size and angle . For example, the CSL method introduces a cyclic smooth label loss to convert the angle regression task into a classification task. DCL accelerates CSL by using dense codes, while GF-CSL stabilizes CSL by using dynamic weights. In addition, the sliding vertex method redefines the rotating box prediction task as a four-sided prediction relative to the horizontal box offset.

[0003] The second category of methods uses an IoU-like loss to address the limitation of the smooth L1 loss in accurately describing the relationship between the ground truth box and the predicted box. For example, GWD and KLD convert the rotating box into a smooth Gaussian distribution and measure the distribution distance to maintain consistency with IoU. KFIoU also uses Gaussian modeling but directly approximates IoU, eliminating the hyperparameters required for measuring the distribution distance. Oblique IoU derives the IoU calculation formula between rotating rectangles. The ACM module improves the IoU-like loss by addressing the inconsistency between the target angle and the annotation box angle.

[0004] Existing angle regression methods still have deficiencies in dealing with the problem of boundary discontinuity. For methods based on position representation, when dealing with complex scenes and multi-class targets, the stability and accuracy of the model still need to be improved. For example, the CSL method may exhibit unstable predictions near the angle boundary. Although the IoU-like loss function alleviates the problem of boundary discontinuity to a certain extent, in practical applications, due to the regression model needing to predict highly divergent angle values for similar features, the problem of boundary discontinuity still exists. In addition, existing methods based on position representation also face challenges in terms of model stability and accuracy when dealing with complex scenes and multi-class targets.

[0005] In the regression head of an Oriented Object Detection (OOD) model, the input is the image features extracted from the proposals, and the output is in The represented bounding box is used to represent the position. However, we observe that for targets close to the angular boundary, the image features received by the regression head are very similar, but two significantly different angular values need to be predicted. In contrast, in non-boundary cases, similar image features correspond to similar angular values. This difference causes the regression process to mutate near the boundary. Especially near the boundary, it is difficult for the regression head to produce continuous and accurate angular predictions. Therefore, the angular prediction of targets near the boundary becomes highly unstable during the training process, which significantly exacerbates the Boundary Discontinuity Problem (BDP), resulting in low prediction accuracy and stability. Summary of the Invention

[0006] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a rotation target detection method based on a dynamic reference axis to improve the prediction accuracy and stability.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is: A rotation target detection method based on a dynamic reference axis, which is implemented by using a rotation target detection model. The rotation target detection model introduces a region division mechanism to divide the angular domain into multiple regions, and classifies the region to which the target belongs through a region predictor, so as to perform angular regression within a stable region; the rotation target detection model introduces an angular conversion module to map the reference angle based on the x-axis to the reference angle based on the y-axis, enabling the rotation target detection model to flexibly switch between different reference axes.

[0008] Specifically, the method includes a training stage and an inference stage. The training stage is divided into two stages, specifically as follows: The first stage: The backbone network and the region proposal network of the rotation target detection model remain unchanged, and the classification head also remains unchanged. The regression head is optimized: In addition to predicting outside, the model also predicts , where x and y respectively represent the horizontal and vertical coordinates of the center point of the bounding box predicted by the model, w and h respectively represent the width and height of the bounding box predicted by the model, and θ x represents the rotation angle of the bounding box predicted by the model based on the original x-axis, and θ y represents the rotation angle of the bounding box predicted by the model based on the y-axis; and the regression head introduces an angular conversion module to adjust the angular value; Specifically, the training set is input into the rotation target detection model. Feature vectors are extracted by the backbone network and the region proposal network. The feature vectors are input into the classification head to obtain class information, and the feature vectors are input into the regression head to predict the position information , and the predicted Compare with the corresponding ground truth to calculate the loss of the trained model; The second stage: Add the region predictor to the regression head and only train its parameters to keep the angle regression result obtained in the first stage unaffected by the second stage training; The region predictor adopts the same structure as the classification head, except that the output is a one-dimensional binary classification, and the region predictor accurately divides the region; Specifically, the feature vector is input into the region predictor to obtain the predicted region score , and the region encoder encodes the proposed true rotation angle into the corresponding region encoding to obtain the true belonging region ; and calculate the binary cross-entropy loss according to the true belonging region and the predicted region score ; In the inference stage, the image of the target to be detected is input into the rotation target detection model. After the feature vector is extracted by the backbone network and the region proposal network, the feature vector is passed through the region predictor to obtain the classification score, and the angle regression value is adjusted according to the score. The rotation target detection model finally outputs the category and position information of the target to be detected as the detection result.

[0009] The region division mechanism is as follows: The angle of the object bounding box can be restricted within , select axis as the reference axis, and define the angle of the object as one side of the y-axis, that is, the range of the rotation angle; Divide the rotation angle range into four regions in the clockwise order, namely Region 1, Region 2, Region 3 and Region 4, and the angle range of each region is ; Region 2 and Region 3 are stable regions close to the reference axis, and Region 1 and Region 4 are regions where boundary discontinuity problems may occur; Construct the y-axis as the second reference axis. At this time, Region 1 and Region 4 are stable regions close to the second reference axis. In order to unify Region 1 and Region 4, Region 4 is centered symmetric about the origin to obtain Region 5; At this time, Region 1 and Region 5 are continuous regions.

[0010] The regression head models the position regression as a classification task with two regression components: 1) the region where the target appears, and 2) the position and angle regression values of the target based on the x-axis, and the angle regression value based on the y-axis, as shown in the following equation:

[0011] where, is the binary cross-entropy loss, is the input image feature; is based on The ground truth angle value with the axis as the reference axis; θ is the predicted angle of the original model for the bounding box with the x-axis as the reference axis, is the output of the region predictor for predicting the region where the target is located; is the output of the region encoder for encoding the ground truth angle into region information; , is the output of the regression head, containing the angle prediction values based on the x-axis and y-axis; is the distance metric used to measure the change between the predicted value and the ground truth value, is an angle change function for mapping the position parameter to the corresponding region position on the y-axis; is the angle conversion module for converting the reference angle based on axis to the reference angle based on axis.

[0012] The calculation formula for mapping the reference angle based on the x-axis to the reference angle based on the y-axis is as follows:

[0013] where, is the angle with axis as the reference axis, which is changed to obtain the corresponding angle with axis as the reference axis, and maps the original angle of region four to region five.

[0014] The specific steps of the inference stage are as follows: Region classification: The feature vector passes through the region predictor to obtain the classification score of the region to which the target belongs; Angle selection: According to the region classification result, select the angle prediction value corresponding to the reference axis; if the target is classified as belonging to axis region, then retain the prediction based on axis ; otherwise, select the prediction based on axis , and output after being processed by the angle conversion module.

[0015] After adopting the above scheme, the present invention introduces a dynamic reference axis mechanism, which dynamically adjusts the reference axis according to the approximate direction of the target, and decomposes the angle prediction problem based on a single reference axis into an angle prediction problem based on multiple reference axes. The dynamic reference axis mechanism maps the regions prone to boundary discontinuity problems to the stable regions of another reference axis, effectively alleviating the boundary discontinuity problem, and at least one angle prediction is stable in any angle region, significantly improving the robustness of the model in complex scenarios.

[0016] Specifically, on the one hand, the present invention introduces a region division mechanism, divides the angular domain into multiple regions, and classifies the region to which the target belongs through a region predictor, so as to perform angular regression within a stable region. Through region division, the model can perform angular regression within a stable region, avoiding angular prediction errors near the boundary, significantly improving the accuracy of angular prediction, and the region division mechanism enables the model to better adapt to targets at different angles, enhancing the generalization ability of the model in complex scenarios. On the other hand, the present invention introduces an angle conversion module to map the reference angle based on the x-axis to the reference angle based on the y-axis, enabling the model to flexibly switch between different reference axes and further alleviating the problem of boundary discontinuity. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of the rotation target detection method of the present invention; Figure 2 are different representations that may occur in angular regression under different circumstances; Figure 3 is a schematic structural diagram of the training and inference of the present invention; Figure 4 is a schematic diagram of the detection result of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] As Figures 1-3 shown, where Figure 2 (1) indicates that for a target close to the angular boundary, the angles are adjacent but exceed the boundary limit, Figure 2 (2) indicates that crossing the boundary will cause a sudden change in the angle, forcing the model to regress very different angles from similar features, Figure 2 (3) indicates that the boundary limit is avoided by transforming the reference axis. The present invention discloses a rotation target detection method based on a dynamic reference axis, which decomposes the angular prediction problem based on a single reference axis (such as the x-axis) into an angular prediction problem based on multiple reference axes. On the one hand, a region division mechanism is introduced to divide the angular domain into multiple regions, and the region to which the target belongs is classified through a region predictor, so as to perform angular regression within a stable region; on the other hand, an angle conversion module is introduced to map the reference angle based on the x-axis to the reference angle based on the y-axis, enabling the rotation target detection model to flexibly switch between different reference axes.

[0019] The present invention is implemented using a rotation target detection model, which specifically includes a training stage and an inference stage. The training stage is divided into two stages, specifically as follows: The first stage: The backbone network and the region proposal network of the rotation target detection model remain unchanged, and the classification head also remains unchanged. The regression head is optimized: The regression head predicts In addition, the model also predicts Since the ground truth angle values are normalized with the axis as the reference axis during the preprocessing to unify the angle annotations, an Angle Transformation Module (ATM) is introduced in the regression head to adjust the corresponding angle values. Here, x and y represent the horizontal and vertical coordinates of the center point of the bounding box predicted by the model, w and h represent the width and height of the bounding box predicted by the model, and θ x represents the rotation angle of the bounding box predicted by the model based on the original prediction with respect to the x-axis, and θ y represents the rotation angle of the bounding box predicted by the model based on the y-axis.

[0020] The training process of the first stage is specifically as follows: The training set is input into the rotated object detection model. Feature vectors are extracted by the backbone network and the Region Proposal Network. The feature vectors are input into the classification head to obtain class information, and the feature vectors are input into the regression head to predict the position information , and the predicted is compared with the corresponding ground truth value to calculate the loss of the training model. The loss function can be of any type, such as Smooth L1, KFIOU, or GWD. According to the requirements of the selected loss, after passing through the angle transformation module, the mapped target value can be adjusted accordingly. Since has a clear mapping relationship, the two can use the same features for regression. In addition, from the perspective of computational cost, the regression head only needs to add an additional output dimension, so the modification is very small.

[0021] The second stage: Add a binary region predictor (RP) to the regression head and only train its parameters to keep the angle regression result obtained in the first stage unaffected by the second stage training. To minimize the computational cost, the region predictor adopts the same structure as the classification head, except that the output is a one-dimensional binary classification. Since the features extracted after the first stage training can already enable the classification head to correctly classify the proposals, it can be reasonably assumed that the same features can accurately divide the regions through a region predictor with a similar structure.

[0022] Since the angle of the object bounding box can be restricted within , any straight line can be arbitrarily selected as the reference axis ( axis), take this reference axis as axis one, and define the angle of the object as one side of the y-axis, that is, the range of the rotation angle. We divide this angle range into four equal parts, and at this time each angle range is . In the clockwise order, these four regions are divided into Region One, Region Two, Region Three, and Region Four. At this time, Region Two and Region Three are stable regions close to the reference axis, and Region One and Region Four are regions where boundary discontinuity problems may occur. Then construct the second reference axis ( Axis), taking this reference axis as Axis Two, Axis Two is perpendicular to Axis One and is the boundary line between Region Two and Region Three. When taking Axis Two as the reference axis, Region One and Region Four are stable regions close to the reference axis, and Region Two and Region Three are regions where boundary discontinuity problems may occur. In order to unify Region One and Region Four, Region Four can be made centrosymmetric about the origin to obtain Region Five. At this time, Region One and Region Five are continuous regions, which is convenient for subsequent calculations. Therefore, under the combined action of the two reference axes, all regions can become stable regions. For the specific illustration, see Figure 2 (3).

[0023] Taking the common setting as an example. This setting takes axis as the reference axis, and the rotation angle range is . The angle boundary is along axis, which is called Case One. Similarly, when the reference axis is axis, the angle boundary is along axis, which is called Case Two. As shown in Figure 2 (1), using and as the dividing lines, the angle domain of is divided into regions and , where region is equivalent to region in terms of angle. In Case One, regions and will not have boundary discontinuity problems and belong to stable regions. Similarly, in Case Two, regions and are stable regions.

[0024] The training in the second stage is as follows: The feature vector is input into the region predictor to obtain the predicted region score , and the region encoder (Region Encoder, RE) encodes the proposed true rotation angle into the corresponding region encoding to obtain the true region to which it belongs (a 0 / 1 vector); and binary cross-entropy loss calculation is performed according to the true region to which it belongs and the predicted region score .

[0025] In the present invention, the regression head models the position regression as a classification task including two regression components: 1) the region where the target appears, and 2) the position and angle regression values of the target based on the x-axis, and the angle regression value based on the y-axis, as shown in the following equations:

[0026] Among them, is the input image feature; is based on axis as the reference axis for the ground truth angle value; θ is the predicted angle of the original model for the bounding box when taking the x-axis as the reference axis, is the output of the region predictor, used to predict the region where the target is located; is the output of the region encoder, used to encode the ground truth angle into region information; , is the output of the regression head, containing the angle prediction values based on the x-axis and y-axis; is the distance metric used to measure the variation between the predicted value and the ground truth value, is an angle variation function, used to map the position parameter to the corresponding region position on the y-axis. is the Angle Transfer Module (ATM), used to map the reference angle based on axis to the reference angle based on axis, and the specific formula is as follows:

[0027] where, is the angle based on axis, and through variation, the corresponding angle based on axis is obtained, and the original angle of region four is mapped to region five.

[0028] The boundary discontinuity problem is not directly solved through the above two-stage training; the predicted still performs poorly near the boundary. However, within their respective regions, that is, Figure 2 in (3) the axis region and axis region, their performance remains stable, which is consistent with the previous research observations. This indicates that a fully trained ARF-DRA has two key capabilities: 1) ensuring that at least one angle prediction is stable within any angle region; 2) distinguishing the angle region to which the target belongs.

[0029] In the inference stage, the image of the target to be detected is input into the rotated object detection model. After the feature vector is extracted by the backbone network and the region proposal network, the feature vector is passed through the region predictor to obtain the classification score, and the angle regression value is adjusted according to this score. The specific steps are as follows: Region classification: The feature vector passes through the region predictor (RP) to obtain the classification score of the region to which the target belongs.

[0030] Angle selection: According to the region classification result, select the angle prediction value corresponding to the reference axis. If the target is classified as belonging to For the axis region, the prediction based on the axis is retained ; otherwise, the prediction based on the axis is selected , and after being processed by the angle conversion module and output, the rotation target detection model finally outputs the category and position information of the target to be detected as the detection result.

[0031] In the present invention, the regions prone to boundary discontinuity problems are mapped to the stable regions of another reference axis by dynamically adjusting the reference axis. An important consideration is whether boundary problems will occur at the division of the axis and axis regions. From the perspective of the region classifier, this boundary represents a challenging situation because the classification result may become ambiguous due to highly similar features. However, the division line of the region boundary does not coincide with the actual position of the angle discontinuity. This means that in this region, whether taking the axis or axis as the reference, the regression value of the angle remains stable. Therefore, even if a region classification error occurs near the boundary, the impact on the model performance is small because the difference between and is negligible and is closely aligned with the ground truth. This mechanism effectively alleviates the boundary discontinuity problem and improves the stability and accuracy of angle prediction.

[0032] In summary, the key points of the present invention are: 1. The dynamic reference axis mechanism; In the prior art, there is a discontinuity problem in angle regression near the boundary, resulting in difficulty for the model to stably predict the angle of the target near the boundary during the training process, thereby affecting the overall detection accuracy. The present invention introduces a dynamic reference axis mechanism, which dynamically adjusts the reference axis according to the approximate direction of the target, decomposing the angle prediction problem based on a single reference axis into an angle prediction problem based on multiple reference axes. The dynamic reference axis mechanism maps the regions prone to boundary discontinuity problems to the stable regions of another reference axis, effectively alleviating the boundary discontinuity problem, and at least one angle prediction is stable in any angle region, significantly improving the robustness of the model in complex scenarios.

[0033] 2. Region division and classification; In the prior art, when the angle regression head processes targets near the boundary, due to similar image features but large differences in the angle values that need to be predicted, the regression process is unstable, which in turn affects the generalization ability and detection accuracy of the model. The present invention introduces a region division mechanism, divides the angle domain into multiple regions, and classifies the region to which the target belongs through a region predictor (RP), so as to perform angle regression within a stable region. Through region division, the model can perform angle regression within a stable region, avoiding angle prediction errors near the boundary, significantly improving the accuracy of angle prediction, and the region division mechanism enables the model to better adapt to targets at different angles, enhancing the generalization ability of the model in complex scenarios.

[0034] III. Angle Transformation Module (ATM) In the prior art, angle prediction based on a single reference axis has a discontinuity problem near the boundary, resulting in difficulty for the model to stably predict the angle of the target during training and inference. The present invention introduces an angle transformation module (ATM) to map the reference angle based on the x-axis to the reference angle based on the y-axis, enabling the model to flexibly switch between different reference axes and further alleviating the boundary discontinuity problem.

[0035] The ATM module enables the model to flexibly switch between different reference axes, further alleviating the boundary discontinuity problem, enhancing the flexibility of angle prediction, and the introduction of the ATM module only requires adding an additional output dimension to the regression head, with low computational overhead and easy to implement.

[0036] The present invention has the following beneficial effects: (1) Significantly improve the accuracy of angle prediction: Through the dynamic reference axis mechanism and the region division mechanism, at least one angle prediction in any angle region of the present invention is stable, significantly improving the accuracy of angle prediction, especially in the detection of targets near the boundary.

[0037] (2) Enhance the robustness of the model: The dynamic reference axis mechanism and the region division mechanism enable the model to better adapt to targets at different angles in complex scenarios, significantly enhancing the robustness of the model and improving the reliability of the model in practical applications.

[0038] (3) Improve the generalization ability of the model: The region division mechanism enables the model to better adapt to targets at different angles, enhancing the generalization ability of the model in complex scenarios, so that the model can maintain a high detection accuracy in different environments and scenarios.

[0039] (4) Simplify the calculation process: Introducing the angle transformation module (ATM) only requires adding an additional output dimension to the regression head, with low computational overhead and easy to implement, making the present invention have high efficiency and scalability in practical applications.

[0040] (5) Strong compatibility: The framework of the present invention is theoretically compatible with most existing object detection models, has good scalability, and can be easily integrated into existing object detection systems, providing new technical means for research and applications in related fields.

[0041] To better illustrate the effects achieved by the present invention, further explanations will be given through simulation experiments below.

[0042] I. Simulation conditions The present invention is developed on the Ubuntu platform, and the developed deep learning framework is based on Pytorch. The main language used in the present invention is Python.

[0043] II. Simulation content Take the commonly used rotated object dataset DOTA in the field of remote sensing images, train the network according to the above steps and use the test set for testing. Figure 4 The detection results of the present invention and other methods on the DOTA dataset are shown. It can be found that the present invention consistently improves the performance of the baseline algorithm, and compared with other methods, the present invention achieves the best results. Since specific technical details are included in the names of various algorithms, it is a common practice in the industry to use English model names, which will not be elaborated here. The detection performance of this method on the DOTA dataset reaches 66.53%, which is higher than that of other methods, proving that the present invention has better effects in rotated object detection.

[0044] As described above, it is only an embodiment of the present invention, and it does not impose any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes and decorations made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A rotating target detection method based on a dynamic reference axis, characterized in that: It is implemented by using a rotated object detection model. The rotated object detection model introduces a region division mechanism, divides the angular domain into multiple regions, and classifies the region to which the object belongs through a region predictor, so as to perform angular regression within a stable region; The rotated object detection model introduces an angle conversion module to map the reference angle based on the x-axis to the reference angle based on the y-axis, enabling the rotated object detection model to flexibly switch between different reference axes; Specifically, the method includes a training stage and an inference stage. The training stage is divided into two stages, which are as follows: The first stage: Keep the backbone network and region proposal network of the rotated object detection model unchanged, and also keep the classification head unchanged. Optimize the regression head: In addition to predicting the model also predicts , where x and y respectively represent the horizontal and vertical coordinates of the center point of the bounding box predicted by the model, and w and h respectively represent the width and height of the bounding box predicted by the model. θ x represents the rotation angle of the bounding box predicted by the model based on the original prediction of the x-axis, and θ y represents the rotation angle of the bounding box predicted by the model based on the y-axis; and the regression head introduces an angle conversion module to adjust the angle value; Specifically, the training set is input into the rotation object detection model, and the feature vectors are extracted by the backbone network and the region proposal network. The feature vectors are input into the classification head to obtain the category information, and the feature vectors are input into the regression head to predict the position information , and the predicted is compared with the corresponding ground truth to calculate the loss of the training model; The second stage: Add the region predictor to the regression head and only train its parameters to ensure that the angular regression results obtained in the first stage are not affected by the training in the second stage; The region predictor adopts the same structure as the classification head, except that the output is a one-dimensional binary classification, and the region predictor accurately divides the regions; Specifically, the feature vector input region predictor obtains the predicted region score , and the region encoder encodes the proposed true rotation angle into the corresponding region encoding to obtain the true belonging region ; and performs binary cross-entropy loss calculation according to the true belonging region and the predicted region score ; In the inference stage, the image of the target to be detected is input into the rotation target detection model. After the feature vectors are extracted by the backbone network and the region proposal network, the feature vectors are passed through the region predictor to obtain the classification scores, and the angular regression values are adjusted according to these scores. The rotation target detection model finally outputs the category and location information of the target to be detected. As the detection result.

2. The rotation target detection method based on a dynamic reference axis according to claim 1, wherein: The area division mechanism is as follows: The angle of the object bounding box can be restricted within and the axis is selected as the reference axis. The angle of the object is defined as one side of the y-axis, that is, the range of the rotation angle. The range of the rotation angle is divided into four regions in the clockwise order, namely Region 1, Region 2, Region 3, and Region 4. The angle range of each region is ; Region 2 and Region 3 are stable regions close to the reference axis, and Region 1 and Region 4 are regions where boundary discontinuity problems may occur. The y-axis is constructed as the second reference axis. At this time, Region 1 and Region 4 are stable regions close to the second reference axis. To unify Region 1 and Region 4, Region 4 is centered symmetric about the origin to obtain Region 5. At this time, Region 1 and Region 5 are continuous regions.

3. A method for detecting a rotating target based on a dynamic reference axis according to claim 1, characterized in that: The regression head models the position regression as a classification task with two regression components: 1) the region where the object appears, and 2) the position and angular regression values of the object based on the x-axis, and the angular regression value based on the y-axis, as shown in the following equation: Among them, is the binary cross-entropy loss, is the input image feature; is the ground truth angle value with the axis as the reference axis; θ is the predicted angle of the original model for the bounding box when the x-axis is the reference axis, is the output of the region predictor for predicting the region where the target is located; is the output of the region encoder for encoding the ground truth angle into region information; , which is the output of the regression head and contains the angle prediction values based on the x-axis and y-axis; is a distance metric used to measure the change between the predicted value and the ground truth, is an angle change function used to map the position parameter to the corresponding regional position on the y-axis; is an angle conversion module used to map the reference angle based on axis to the reference angle based on axis.

4. A method for detecting a rotating target based on a dynamic reference axis according to claim 3, characterized in that: The calculation formula for mapping the reference angle based on the x-axis to the reference angle based on the y-axis is as follows: Among them, is the angle with axis as the reference axis. By changing, the corresponding angle with axis as the reference axis is obtained, and the original angle in Region Four is mapped to Region Five.

5. A method for detecting a rotating target based on a dynamic reference axis according to claim 1, characterized in that: The specific steps of the inference stage are as follows: Region classification: The feature vector passes through the region predictor to obtain the classification score of the region to which the object belongs; Angle selection: According to the region classification result, select the angle prediction value corresponding to the reference axis; If the target is classified as belonging to the axis region, retain the prediction based on the axis ; otherwise, select the prediction based on the axis , and output it after being processed by the angle conversion module.

Citation Information

Patent Citations

  • Remote sensing image target detection method based on directional rotation equivariant network

    CN117830868A

  • Shooting control method for differential robot, differential robot and shooting system

    CN118042281A

  • Remote sensing image directed target detection method based on Faster R-CNN

    CN119963981A

  • Method for planning optimal path having minimum target positioning error of autonomous surface vehicle

    WO2023221658A1