Method for designing SAR target orientation frame detection model based on continuous coding

By designing a SAR target orientation box detection model based on continuous encoding, the continuous coding method of coordinate decomposition and joint optimization loss function are used to solve the discontinuity problem of the OBB model at the boundary, and the accuracy and stability of ship direction detection are improved.

CN120107783APending Publication Date: 2025-06-06AIR FORCE UNIV PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510160620.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The OBB-based DL model has boundary discontinuity problems when the ship target rotates, resulting in angle prediction errors and training that are difficult to converge.

Method used

A SAR target orientation box detection model based on continuous encoding is designed, and the boundary discontinuity problem is solved through the continuous coding method of coordinate decomposition and the joint optimization loss function.

Benefits of technology

It effectively solves the problem of boundary discontinuity in ship direction detection, avoids angle prediction errors, and improves the performance of the detector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107783A_ABST
    Figure CN120107783A_ABST
Patent Text Reader

Abstract

The invention discloses a design method of an SAR target orientation frame detection model based on continuous coding, and belongs to the technical field of computer vision, and the design method specifically comprises the following steps: S1, determining the boundary discontinuity problem of an SAR target orientation frame, and analyzing the boundary continuity condition; s2, designing a continuous coding method and a loss function based on coordinate decomposition based on an analysis result; s3, constructing an SAR target orientation frame detection model, and training and testing the model by using a public SAR ship detection data set; compared with a previous detection model, the method can effectively solve the problem of boundary discontinuity in ship direction detection, and the angle value at the boundary does not suddenly change, so that the angle prediction error at the boundary is avoided, and the performance of the detector is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular to a design method of a SAR target directional frame detection model based on continuous coding. Background Art

[0002] As a precursor to SAR (Synthetic Aperture Radar) image recognition, traditional SAR ship detection models based on DL (Deep Learning) usually use HBB (Horizontal Bounding Box) to represent the position and size of the ship. The loss function of the model is generally expressed as the IoU (Intersection over Union) loss between the actual HBB and the predicted HBB. Although HBBs can effectively surround the ship area, they cannot capture the direction and aspect ratio of the ship. In addition, HBBs inevitably introduce a large amount of background clutter, which is not conducive to the subsequent SAR image recognition processing.

[0003] To address this challenge, researchers have recently shifted their focus to directional SAR ship detection algorithms that employ OBBs (Oriented Bounding Boxes). Unlike HBBs, OBBs are rectangular boxes with rotation angles that allow the OBB to closely conform to the outline of the ship. By using OBBs, background clutter can be minimized and a more accurate representation of the ship's aspect ratio and orientation can be generated.

[0004] However, DL models based on OBB often have discontinuity problems. That is, when the ship target rotates near the boundary angle, the angle prediction will change suddenly due to the periodicity of the angle, which is usually called the boundary discontinuity problem. Boundary discontinuity leads to two problems: (1) When fitting discontinuous functions, the DL model may have numerical instability, resulting in angle prediction errors; (2) The inconsistency between the L1-based loss function and the evaluation index makes it difficult for the deep learning model to converge and train effectively. Summary of the invention

[0005] The purpose of the present invention is to solve the defects in the prior art and propose a design method for a SAR target directional frame detection model based on continuous coding.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The design method of the SAR target directional frame detection model based on continuous coding, the specific steps of the design method are as follows:

[0008] S1: Determine the boundary discontinuity problem of the SAR target orientation box and analyze the boundary continuity conditions;

[0009] S2: Design a continuous encoding method and loss function based on coordinate decomposition based on the analysis results;

[0010] S3: Build a SAR target directional box detection model and use the public SAR ship detection dataset to train and test the model.

[0011] As a further solution of the present invention, the specific steps of determining the SAR target orientation box boundary discontinuity problem in S1 are as follows:

[0012] S1.1: Collect the coordinates (cx, cy) of the center of the SAR target directional box OBB, the shortest side w and the longest side h in the OBB bounding box, and the rotation angle θ of the OBB relative to the horizontal axis, and use the five-parameter method to represent the OBB based on the collected parameter information.

[0013] S1.2: Use the long edge definition method and OpenCV definition method to determine the value range of the rotation angle θ of OBB respectively, then set a blue predicted OBB and a red label OBB at the boundary position, and calculate the SkewIoU solution result of OBB at this time, and based on the periodicity of the angle, the L1 loss value of OBB at this time;

[0014] S1.3: Comparing the SkewIoU solution result with the L1 loss value, it can be seen that when the target angle exceeds the defined range, the value of θ will change suddenly, resulting in a sudden increase in the model loss value. Due to the angle mutation, there is an inconsistency between the L1 loss function and the SkewIoU, which makes it difficult for the training process to converge. In addition, the OpenCV definition method has both angle prediction breakpoints and edge prediction breakpoints at the boundary angles.

[0015] S1.4: Use the vertex offset method and the center point offset method to represent the OBB, and analyze the results of the two methods. It can be seen that the vertex offset method and the center point offset method will also lead to inconsistency between the loss and the measurement. The L1 loss function is replaced by the joint optimization loss function, and the deviation between the expected prediction angle and the actual prediction angle is calculated through the updated loss function. It can be seen that the two sets of angles have large errors in the boundary angles.

[0016] As a further solution of the present invention, the specific steps of analyzing the boundary continuity condition in S1 are as follows:

[0017] S2.1: Define E(x) as the parameter space P from the SAR ship target x to the OBB n The encoding function of , where n represents the parameter dimension, defines T(x; θ) as the transformation that rotates x counterclockwise by an angle θ, and S is the state space set of x;

[0018] S2.2: To ensure the continuity of coding, it is necessary to ensure that slight rotation has the least impact on OBB parameters, that is:

[0019]

[0020] Then define D() as the decoding function, p as a set of OBB parameters. In addition to satisfying the encoding continuity, it is also necessary to satisfy the decoding completeness to ensure that every state of x can be accurately represented, that is:

[0021]

[0022] S2.3: Define L separately E () and L D () is the loss function calculated using encoding parameters and decoding parameters. During the rotation of the ship target, the loss function should not have a sudden change and should have rotation continuity, that is:

[0023]

[0024] The conditions corresponding to formulas (1), (2) and (3) are called coding continuity, decoding integrity and loss function rotation continuity respectively. When only coding parameters or only decoding parameters are used, only the corresponding conditions need to be met. Then, the detection model in the field of SAR ship detection is analyzed. It can be seen that each model cannot simultaneously meet the three conditions of coding continuity, decoding integrity and loss function rotation continuity.

[0025] As a further solution of the present invention, the specific steps of designing the continuous encoding method and loss function based on coordinate decomposition described in S2 are as follows:

[0026] S3.1: Change the angle range of OBB in the long edge definition method to [0,π), and denote the rotation angle of OBB as θ obb , the rotation angle of the ship target is represented as θ t , and θ obb With θ t The value ranges of are [0,π) and [0,2π) respectively. Then θ obb With θ t The conversion relationship between can be expressed as:

[0027]

[0028] From formula (4), we can know that θ obb The period is π, and at θ t = discontinuous at 0 and π, corresponding to π / 2 and 3π / 2 in polar and rectangular coordinate systems;

[0029] S3.2: The coordinates of a point on the unit circle in a rectangular coordinate system can be expressed as Its angle It can be expressed as:

[0030]

[0031] In formula (5), mod represents the modulo operation, arctan2 represents a quadrant function, and then the angle information of the point in the rectangular coordinate system is calculated according to the coordinates. Its value range is [0,2π). The specific calculation formula is as follows:

[0032]

[0033] When ω=2, formula (5) becomes:

[0034]

[0035] when When the value range of is [0,2π], formula (7) can be expressed as:

[0036]

[0037] S3.2: From formula (4) to formula (8) in step S3.1, it can be seen that when ω = 2, and When the value range of is [0,2π], formula (8) is equal to formula (4), so the coordinates of the point on the unit circle can be used to encode θ obb , its encoding function can be expressed as:

[0038] E θ (θ t )=(α 1 ,α 2 )=(cos2θ t ,sin2θ t ) (9)

[0039] Among them, α 1 and α 2 Respectively represent the encoding parameters, and the decoding function can be expressed as:

[0040]

[0041] In formula (9) and formula (10), E θ (·) and D θ (·) respectively represent the functions for encoding and decoding the angle parameter of x. Since the other four parameters (cx, cy, w, h) are not encoded or decoded, the encoding functions E(·) and D(·) for the angle parameters only are written as E θ (·) and D θ (·);

[0042] S3.3: Plot θ in polar coordinates obb , α 1 and α 2 Rotate with the target angle θ t, in order to intuitively illustrate the change of θ obb , α 1 and α 2 The change of the value in the polar coordinate system, where the azimuth axis represents the rotation angle θ of the target t , the radial axis represents θ obb , α 1 and α 2 The corresponding value, based on the drawn polar coordinate information, can be seen as the azimuth angle θ t The rotation of α 1 and α 2 Keep continuous, there is no breakpoint, and for any OBB angle θ obb , there is always a pair (α 1 ,α 2 ) corresponds to it, therefore, the encoding function satisfies the rotation continuity and guarantees the decoding integrity;

[0043] S3.4: Design a joint optimization loss function L using the form of a joint optimization loss function obb , and its specific calculation formula is expressed as:

[0044] L obb =L SkewIoU (B p (cx p ,cy p ,w p ,h p ,θ p ),B GT (cx t ,cy t ,w t ,h t ,θ t )) (11)

[0045] Among them, B p Represents the predicted OBB, with parameters (cx p ,cy p ,w p ,h p ,θ p ), B GT Represents the label OBB, the parameter is (cx t ,cy t ,w t ,h t ,θ t ), (cx, cy) represents the center position parameter, (w, h) represents the width and height parameters, θ represents the rotation angle parameter, L SkewIoU Represents the loss function approximating SkewIoU, θ p Obtained by decoding the encoding parameters, expressed as:

[0046]

[0047] in, and Represents the predicted coding parameters, and then an angle coding loss is added to better constrain the predicted coding parameters and achieve faster convergence. The angle coding loss is expressed as:

[0048]

[0049] in, and Indicates the rotation angle θ of the label OBB t Through the encoding function E θ (·) The target encoding parameters obtained, L 1 () represents the L1 loss function, and the total loss of the model is L total It can be expressed as:

[0050] L total =w cls L cls +w obb L obb +w ang L ang (14)

[0051] Among them, L cls represents the classification loss, w cls 、w obb and w ang L cls , L obb and L ang , and L total The value of is not fixed and needs to be adjusted according to the detection network.

[0052] As a further solution of the present invention, the specific steps of constructing the SAR target directional frame detection model in S3 are as follows:

[0053] S4.1: When constructing the SAR target directional frame detection model, the backbone network selects ResNet, ReResNet or ARC detection models, and the feature fusion network selects feature pyramid network, LR-FPN or BFPN detection models. Then, an angle coding prediction branch is added to the detection head of the backbone network and the feature fusion network. The output of the detection head is a two-dimensional angle coding parameter β 1 and β 2 ;

[0054] S4.2: According to formula (9) in step S3.2, determine that the value range of the angle encoding parameter is [-1, 1], and then transform the output parameter to obtain the final angle encoding parameter and To make the training more stable, the specific transformation formula is as follows:

[0055]

[0056] After the predicted coding angle parameters are obtained, the decoding function D θ (·) Encoding parameters of the predicted angle and Decode to get θ p and compared with the predicted HBB parameters (cx p ,cy p ,w,h) parameters that make up OBB;

[0057] S4.3: Use the approximate SkewIoU loss to calculate the loss between the predicted OBB and the label OBB, through the encoding function E θ (·) for target angle θ t Encode and get the target angle encoding parameters and And the L1 loss function is used to calculate the difference between the predicted parameters and the target parameters to constrain the angle encoding;

[0058] S4.4: The backbone network extracts features from the SAR image, and the extracted feature maps at different stages are recorded as C 1 , C 2 , C 3 , C 4 and C 5 ,Afterwards, as the backbone network performs multi-stage convolution and pooling operations on SAR images, the number of channels of its feature map gradually increases, and the size gradually decreases;

[0059] S4.5: Feature fusion network for feature map C i The convolution operation is performed, and then the newly obtained feature map is upsampled twice and added to the previous feature map to fuse features of different scales. Finally, the fused feature map is convolved to obtain the final feature map P 2 , P 3 , P 4 and P 5 .

[0060] As a further solution of the present invention, the detection head part described in S4.1 has three detection branches, which are responsible for category parameter prediction, HBB parameter prediction and angle encoding parameter prediction, respectively, where C represents the number of categories, K represents the number of anchor boxes corresponding to each anchor point, and if it is an anchor-free network, K=1;

[0061] The specific expression of the feature fusion network convolution operation described in S4.5 is as follows:

[0062]

[0063] Where Conv2d(·) represents a two-dimensional convolution operation, and Upasme(·) represents a two-fold upsampling operation.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] The design method of the SAR target directional frame detection model based on continuous coding collects various parameter information of the SAR target directional frame OBB, and determines the boundary discontinuity problem of the SAR target directional frame through the long edge definition method, OpenCV definition method, vertex offset method and center point offset method, and analyzes the boundary continuity conditions. After that, when constructing the SAR target directional frame detection model, an angle coding prediction branch is added to the detection head of the backbone network and the feature fusion network respectively. The output of the detection head is a two-dimensional angle coding parameter. The value range of the angle coding parameter is determined, and then the output parameter is transformed to obtain the final angle coding parameter. After obtaining the predicted coding angle parameter, the predicted angle coding parameter is decoded by the decoding function and combined with the predicted HBB parameter to form the OBB parameter. The loss between the predicted OBB and the label OBB is calculated using the approximate SkewIoU loss. The target angle is encoded by the encoding function to obtain the target angle coding parameter, and the L1 loss function is used to calculate the difference between the predicted parameter and the target parameter. The backbone network extracts features from the SAR image and extracts feature maps of different stages. Then, as the backbone network performs multi-stage convolution pooling operations on the SAR image, the number of channels of the feature map gradually increases, and the size gradually decreases. The feature fusion network performs convolution operations on the feature map, and then upsamples the newly obtained feature map by two times and adds it to the previous feature map to fuse features of different scales. Finally, the fused feature map is subjected to convolution operations to obtain the final feature map. Compared with previous detection models, the present invention can effectively solve the problem of boundary discontinuity in ship direction detection, and the angle value at the boundary will not mutate, thereby avoiding angle prediction errors at the boundary and improving the performance of the detector. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0067] Figure 1 A flowchart of a design method for a SAR target directional frame detection model based on continuous coding proposed by the present invention;

[0068] Figure 2A schematic diagram of the boundary discontinuity problem in the long side definition method and the OpenCV definition method of the design method of the SAR target directional frame detection model based on continuous coding proposed by the present invention;

[0069] Figure 3 A diagram of boundary discontinuity problems in a vertex offset definition method and a center point offset definition method of a design method for a SAR target directional frame detection model based on continuous coding proposed by the present invention;

[0070] Figure 4 A schematic diagram of an ideal prediction angle and an actual prediction angle in a polar coordinate system in the design method of a SAR target directional frame detection model based on continuous coding proposed in the present invention;

[0071] Figure 5 It is a schematic diagram of the ideal prediction angle and the actual prediction angle in a rectangular coordinate system in the design method of the SAR target directional frame detection model based on continuous coding proposed by the present invention;

[0072] Figure 6 A coordinate representation diagram of a point on a unit circle in a plane rectangular coordinate system in the design method of a SAR target directional frame detection model based on continuous coding proposed by the present invention;

[0073] Figure 7 The polar coordinate system θ in the design method of the SAR target directional frame detection model based on continuous coding proposed by the present invention is obb , α 1 and α 2 Rotate with the target angle θ t Schematic diagram of the changes;

[0074] Figure 8 A schematic diagram of a network design of a design method for a SAR target directional frame detection model based on continuous coding proposed by the present invention;

[0075] Fig. 9 This is a detailed structural diagram of the network of the design method of the SAR target directional frame detection model based on continuous coding proposed by the present invention. DETAILED DESCRIPTION

[0076] Reference Figure 1-9 , a design method for a SAR target directional frame detection model based on continuous coding, the specific steps of the design method are as follows:

[0077] Determine the boundary discontinuity problem of the SAR target orientation box and analyze the boundary continuity conditions.

[0078] Specifically, refer to Figure 1-Figure 3It can be seen that the coordinates of the center of the frame (cx, cy) of the SAR target directional frame OBB, the shortest side w and the longest side h in the OBB boundary box, and the rotation angle θ of the OBB relative to the horizontal axis are collected, and the five-parameter method is used to represent the OBB according to the collected sets of parameter information. The value range of the rotation angle θ of the OBB is determined by the long edge definition method and the OpenCV definition method respectively. Then, a blue predicted OBB and a red label OBB are set at the boundary position, and the SkewIoU solution result of the OBB is calculated at this time. Based on the periodicity of the angle, the L1 loss value of the OBB at this time is compared with the SkewIoU solution result and the L1 loss value. It can be seen that when the target angle When it exceeds the defined range, the value of θ will change suddenly, resulting in a sharp increase in the model loss value. Due to the angle mutation, there is an inconsistency between the L1 loss function and SkewIoU, which makes it difficult for the training process to converge. In addition, the OpenCV definition method has both angle prediction breakpoints and edge prediction breakpoints at the boundary corners. The vertex offset method and the center point offset method are used to represent OBB, and the results of the two methods are analyzed. It can be seen that the vertex offset method and the center point offset method will also lead to inconsistency between loss and measurement, and the L1 loss function is replaced by the joint optimization loss function. The deviation between the expected prediction angle and the actual prediction angle is calculated through the updated loss function. It can be seen that the two sets of angles have large errors at the boundary angles.

[0079] Specifically, refer to Figure 1-Figure 3 It can be seen that E(x) is defined as the parameter space P from the SAR ship target x to the OBB n The encoding function is , where n represents the parameter dimension, T(x; θ) is defined as the transformation that rotates x counterclockwise by an angle of θ, S is the state space set of x, and the goal is to ensure the continuity of the encoding. It is necessary to satisfy that the slight rotation has the least effect on the OBB parameters, that is:

[0080]

[0081] Then define D() as the decoding function, p as a set of OBB parameters. In addition to satisfying the encoding continuity, it is also necessary to satisfy the decoding completeness to ensure that every state of x can be accurately represented, that is:

[0082]

[0083] Define L E () and L D () is the loss function calculated using encoding parameters and decoding parameters. During the rotation of the ship target, the loss function should not have a sudden change and should have rotation continuity, that is:

[0084]

[0085] The conditions corresponding to formulas (1), (2) and (3) are called coding continuity, decoding integrity and loss function rotation continuity respectively. When only coding parameters or only decoding parameters are used, only the corresponding conditions need to be met. Then, the detection model in the field of SAR ship detection is analyzed. It can be seen that each model cannot simultaneously meet the three conditions of coding continuity, decoding integrity and loss function rotation continuity.

[0086] Based on the analysis results, a continuous encoding method and loss function based on coordinate decomposition are designed.

[0087] Specifically, refer to Figure 1 , Figure 4-7 It can be seen that the angle range of OBB in the long side definition method is changed to [0,π), and the rotation angle of OBB is expressed as θ obb , the rotation angle of the ship target is represented as θ t , and θ obb With θ t The value ranges of are [0,π) and [0,2π) respectively. Then θ obb With θ t The conversion relationship between can be expressed as:

[0088]

[0089] From formula (4), we can know that θ obb The period is π, and at θ t =0 and π are discontinuous, corresponding to π / 2 and 3π / 2 in the polar coordinate system and the rectangular coordinate system. The coordinates of a point on the unit circle in the plane rectangular coordinate system can be expressed as Its angle It can be expressed as:

[0090]

[0091] In formula (5), mod represents the modulo operation, arctan2 represents a quadrant function, and then the angle information of the point in the rectangular coordinate system is calculated according to the coordinates. Its value range is [0,2π). The specific calculation formula is as follows:

[0092]

[0093] When ω=2, formula (5) becomes:

[0094]

[0095] when When the value range of is [0,2π], formula (7) can be expressed as:

[0096]

[0097] S3.2: From formula (4) to formula (8) in step S3.1, it can be seen that when ω = 2, and When the value range of is [0,2π], formula (8) is equal to formula (4), so the coordinates of the point on the unit circle can be used to encode θ obb , its encoding function can be expressed as:

[0098] E θ (θ t )=(α 1 ,α 2 )=(cos 2θ t ,sin 2θ t ) (9)

[0099] Among them, α 1 and α 2 Respectively represent the encoding parameters, and the decoding function can be expressed as:

[0100]

[0101] In formula (9) and formula (10), E θ (·) and D θ (·) respectively represent the functions for encoding and decoding the angle parameter of x. Since the other four parameters (cx, cy, w, h) are not encoded or decoded, the encoding functions E(·) and D(·) for the angle parameters only are written as E θ (·) and D θ (·);

[0102] Plot θ in polar coordinates obb , α 1 and α 2 Rotate with the target angle θ t , in order to intuitively illustrate the change of θ obb , α 1 and α 2 The change of the value in the polar coordinate system, where the azimuth axis represents the rotation angle θ of the target t , the radial axis represents θ obb , α 1 and α 2 The corresponding value, based on the drawn polar coordinate information, can be seen as the azimuth angle θ t The rotation of α 1 and α 2 Keep continuous, there is no breakpoint, and for any OBB angle θ obb , there is always a pair (α 1 ,α 2) corresponds to it, therefore, the encoding function satisfies the rotation continuity and ensures the decoding integrity. Then, the joint optimization loss function L is designed in the form of a loss function based on joint optimization. obb , and its specific calculation formula is expressed as:

[0103] L obb =L SkewIoU (B p (cx p ,cy p ,w p ,h p ,θ p ),B GT (cx t ,cy t ,w t ,h t ,θ t )) (11)

[0104] Among them, B p Represents the predicted OBB, with parameters (cx p ,cy p ,w p ,h p ,θ p ), B GT Represents the label OBB, the parameter is (cx t ,cy t ,w t ,h t ,θ t ), (cx, cy) represents the center position parameter, (w, h) represents the width and height parameters, θ represents the rotation angle parameter, L SkewIoU Represents the loss function approximating SkewIoU, θ p Obtained by decoding the encoding parameters, expressed as:

[0105]

[0106] in, and Represents the predicted coding parameters, and then an angle coding loss is added to better constrain the predicted coding parameters and achieve faster convergence. The angle coding loss is expressed as:

[0107]

[0108] in, and Indicates the rotation angle θ of the label OBB t Through the encoding function E θ (·) The target encoding parameters obtained, L 1() represents the L1 loss function, and the total loss of the model is L total It can be expressed as:

[0109] L total =w cls L cls +w obb L obb +w ang L ang (14)

[0110] Among them, L cls represents the classification loss, w cls 、w obb and w ang L cls , L obb and L ang , and L total The value of is not fixed and needs to be adjusted according to the detection network.

[0111] A SAR target directional box detection model was built, and the model was trained and tested using a public SAR ship detection dataset.

[0112] Specifically, refer to Figure 1 , Figure 8-9 It can be seen that when constructing the SAR target directional frame detection model, the backbone network selects ResNet, ReResNet or ARC detection models, and the feature fusion network selects feature pyramid network, LR-FPN or BFPN detection models. Then, an angle coding prediction branch is added to the detection head of the backbone network and the feature fusion network. The output of the detection head is a two-dimensional angle coding parameter β 1 and β 2 According to formula (9) in step S3.2, the value range of the angle encoding parameter is determined to be [-1, 1], and then the output parameter is transformed to obtain the final angle encoding parameter and To make the training more stable, the specific transformation formula is as follows:

[0113]

[0114] After the predicted coding angle parameters are obtained, the decoding function D θ (·) Encoding parameters of the predicted angle and Decode to get θ p and compared with the predicted HBB parameters (cx p ,cy p ,w,h) constitute the parameters of OBB, and use the loss of approximate SkewIoU to calculate the loss between the predicted OBB and the label OBB, through the encoding function Eθ (·) for target angle θ t Encode and get the target angle encoding parameters and The L1 loss function is used to calculate the difference between the predicted parameters and the target parameters. The backbone network extracts features from the SAR image with constrained angle encoding, and the extracted feature maps at different stages are recorded as C 1 , C 2 , C 3 , C 4 and C 5 Then, as the backbone network performs multi-stage convolution and pooling operations on SAR images, the number of channels of its feature map gradually increases, and its size gradually decreases. The feature fusion network performs multi-stage convolution and pooling operations on the feature map C. i The convolution operation is performed, and then the newly obtained feature map is upsampled twice and added to the previous feature map to fuse features of different scales. Finally, the fused feature map is convolved to obtain the final feature map P 2 , P 3 , P 4 and P 5 .

[0115] It should be further explained that the detection head part of the backbone network and the feature fusion network has three detection branches, which are responsible for category parameter prediction, HBB parameter prediction and angle encoding parameter prediction, respectively. C represents the number of categories, and K represents the number of anchor boxes corresponding to each anchor point. If it is an anchor-free network, K = 1.

[0116] The specific expression of the feature fusion network convolution operation is as follows:

[0117]

[0118] Where Conv2d(·) represents a two-dimensional convolution operation, and Upasme(·) represents a two-fold upsampling operation.

[0119] In addition, it should be further explained that the effectiveness of the SAR target directional frame detection model proposed in the present invention in solving the boundary discontinuity problem is verified by conducting experiments on two public SAR ship detection datasets. The RSDD-SAR dataset is a SAR ship target detection dataset, which consists of 84 scenes collected by the Gaofen-3 satellite, 41 scenes from the TerraSAR-X satellite, and 2 large uncropped images, with a total of 127 scene data. These scenes include different imaging modes, polarizations, and resolutions. Specifically, it includes 7,000 SAR images and 10,263 ship targets. Among them, 5,000 images are used for training and 2,000 images are used for testing. The image resolution of the RSDD-SAR dataset ranges from 2 meters to 20 meters, and the image size is 512×512 pixels. The RSSDD dataset consists of 1,160 SAR images with different resolutions, polarization modes, and sea conditions. It covers a range of scenes, including offshore scenes with considerable land clutter and scenes with less clutter. For the division of the training set, the default principle is adopted, including a training set of 928 images and a test set of 232 images. The image sizes of this dataset range from 271×241 to 526×646. In order to standardize the size, they are uniformly adjusted to 512×512 pixels;

[0120] The MMRotate platform was used to complete the comparative experiment, with epoch set to 36, batch size set to 4, optimizer set to Adam, and initial learning rate set to 0.0001. All experiments were completed on the Pytorch platform, with the graphics card being NVIDIA GeForce RTX A6000.

[0121] The three evaluation indicators in the MS COCO dataset are used to evaluate the detection accuracy of the model, namely: average precision (AP), AP 50 and AP 75 The AP is defined as follows:

[0122]

[0123] Among them, P(R) represents the precision-recall curve, and AP is defined as the area under the PRC curve. The higher the AP value, the better the detection performance of the method. In the MS COCO dataset, AP is calculated according to the average precision of IoU values ​​​​of 0.5 to 0.95, with a step size of 0.05. This paper uses the same calculation method. P represents precision, that is, the ratio of correctly detected ship samples to all predicted positive ship samples; R represents recall, that is, the ratio of correctly detected ship samples to all marked positive ship samples, expressed as:

[0124]

[0125] Among them, FN is false negative, which means samples that are mistakenly identified as negative; FP is false positive, which means samples that are mistakenly identified as positive; TP is true positive, which means samples that are correctly identified as positive.

[0126] AP 50 and AP 75 They represent the recognition accuracy when the SkewIoU threshold is 0.5 and 0.75 respectively. 50 In comparison, AP 75 It can better reflect the accuracy of the network's high-precision detection.

[0127] The proposed CDM is compared with the commonly used LE90 and CSL encoding methods. At the same time, the angle encoding length of CSL is set to 180, that is, the angle of the regression box is divided into 180 categories, each category represents one degree.

[0128] Since R-RetinaNet and R-FCOS are the baseline networks for many SAR ship detection methods, ablation experiments were performed on these two networks. RetinaNet is a baseline network based on anchor detection, while FCOS is a baseline network for anchor-free detection. The difference between R-RetinaNet and RetinaNet is that R-RetinaNet has an additional angle branch in the detection head to predict the direction of the OBB, and the same is true for R-FCOS. Both networks use ResNet-50 as the backbone network, and the feature fusion network uses the official FPN of MMRotate. These settings are followed in the experiments that follow this section. In the RSDD-SAR and RSSDD datasets, the detection accuracy comparison of different encoding methods is shown in Tables 1 and 2 below:

[0129] Table 1 Comparison of detection accuracy of different angle encoding methods in RSDD-SAR dataset (%):

[0130]

[0131]

[0132] Table 2 Comparison of detection accuracy of different angle encoding methods on RSSDD dataset (%):

[0133]

[0134] As shown in Tables 1 and 2, the network performance using the CSL method is reduced compared to the long edge definition method. This is mainly because when the angle is encoded as a category, a large number of training samples are required to accurately classify the angle. In addition, the excessively long encoding length increases the difficulty and complexity of training. Finally, discretizing the angle for classification will bring additional discretization errors.

[0135] The CDM coding method of the present invention only increases the coding length by one bit, but the detection accuracy of the two detection networks is greatly improved. This is because the long side definition method will encounter the problem of discontinuous boundaries, resulting in angle prediction errors and optimization difficulties at the angle boundaries. CDM coding is a continuous coding method that solves the problem of unstable angle prediction and difficulty in optimization at the angle boundaries.

[0136] According to the actual experimental results, the detection network using the long edge definition method has obvious angle prediction errors when the ship target is located around ±90°, which is consistent with the analysis in Section 2.3. Due to the discontinuity of the boundary, the network's prediction at the angle boundary is unstable, resulting in incorrect angle prediction. After using CDM, the angle prediction error problem of the detection network is significantly alleviated, indicating that the CDM continuous coding method can effectively solve the problem of discontinuous boundaries in ship direction detection. Since the CDM method is continuous at the angle boundary, the angle value at the boundary will not change suddenly, thereby avoiding angle prediction errors at the boundary.

[0137] Analyzed different ang The detection accuracy of the detection network under the value of is shown in Table 3 on the RSDD-SAR dataset. obb Select KLD loss, L cls Select Focal loss, w cls and w obb Set to 1. The results in Table 2.4 show that when w ang = 0.1 and w ang When w = 0.15, R-RetinaNet and R-FCOS have the highest detection accuracy. Therefore, in subsequent experiments, the angle loss weights of these two networks are 0.1 and 0.15 respectively. It is worth noting that when w ang When ∈R = 0, the performance of both networks drops significantly, indicating that the angle loss is essential in the loss function.

[0138] Table 3 Detection accuracy of the network under different angle loss weights (%):

[0139]

[0140] Tables 4 and 5 show the ablation experiments on the impact of different optimization methods on network detection performance, where independent optimization is represented as IO and joint optimization is represented as JO. The independent optimization loss of R-RetinaNet is the predicted OBB parameter (cx p ,cy p ,w p ,h p ,θ p ) and label OBB(cx p ,cy p,w p ,h p ,θ p ) parameters. The independent optimization loss of R-FCOS is the sum of HBB loss and angle loss, where HBB loss is IoU loss and angle loss is L1 loss. In the joint optimization of the two detection networks, GWD loss and KLD loss are selected as L1 loss. obb .

[0141] The purpose of both GWD Loss and KLD Loss is to replace the SkewIoU loss by approximating the changing trend of the SkewIoU value. As shown in Tables 4 and 5, the joint optimization method using GWD loss and KLD loss significantly improves the detection accuracy of the two detection networks on both datasets compared with independent optimization. However, the joint optimization loss still cannot fundamentally solve the discontinuity problem at the corner boundary. When GWD loss and KLD loss are combined with the CDM method, the detection accuracy is further improved. This shows that the CDM encoding method can also improve the performance of the detector in joint optimization.

[0142] Table 4 Ablation experiments of different optimization methods on RSDD-SAR dataset:

[0143]

[0144] Table 5 Ablation experiments of different optimization methods on RSSDD dataset:

[0145]

[0146]

[0147] Table 6 gives the comparison results with existing target detection methods on two datasets. The compared methods are mainly divided into three categories:

[0148] (1) Two-stage detector: The two-stage detector divides the detection process into two stages: first, generating candidate boxes, and then classifying and detecting objects based on the features of the candidate boxes. Classic two-stage detectors mainly include FasterR-CNN, ReDet, and OrientedR-CNN.

[0149] (2) Single-stage detector: This type of detection network directly outputs the detection results and has the advantage of fast detection speed. Since SAR ship detection has high requirements for speed, a single-stage detection network is usually used to detect ship targets. The classic methods of this type mainly include S2ANet, O-Reppoints and R3Det.

[0150] (3) Encoding-based detectors: This type of method improves detection performance by modifying the encoding method of OBB. This category mainly includes GlidingVertex, PSC, EPE, LTPO and FPDDet. Due to the lack of publicly available codes for EPE, LTPO and FPDDet.

[0151] Different detection methods use different backbone networks, feature fusion networks, detection heads, and training strategies, which will affect the final detection performance. It is difficult to establish a completely fair comparison between different methods. Despite these challenges, the 'FCOS+KLD+CDM' method proposed in this chapter achieves the best detection results on the RSDD-SAR dataset, with an AP75 of 91.7% and an AP50 of 55.0%. On the RSSDD dataset, our model achieves the highest AP75 accuracy of 63.8%.

[0152] Table 6 Detection results of different networks on RSDD-SAR and RSSDD datasets:

[0153]

[0154] According to the actual experimental results, in many networks, there are errors in angle prediction due to discontinuous boundaries. The root cause is the existence of discontinuous OBB coding. The method proposed in this chapter can effectively solve the problem of angle prediction errors at the boundaries.

Claims

1. A design method for a SAR target directional frame detection model based on continuous coding, characterized in that: The specific steps of this design method are as follows: S1: Determine the boundary discontinuity problem of the SAR target orientation box and analyze the boundary continuity conditions; S2: Design a continuous encoding method and loss function based on coordinate decomposition based on the analysis results; S3: Build a SAR target directional box detection model and use the public SAR ship detection dataset to train and test the model.

2. The design method of the SAR target directional frame detection model based on continuous coding according to claim 1 is characterized in that: The specific steps for determining the discontinuity of the SAR target directional box boundary described in S1 are as follows: S1.1: Collect the coordinates (cx, cy) of the center of the SAR target directional box OBB, the shortest side w and the longest side h in the OBB bounding box, and the rotation angle θ of the OBB relative to the horizontal axis, and use the five-parameter method to represent the OBB based on the collected parameter information. S1.2: Use the long edge definition method and OpenCV definition method to determine the value range of the rotation angle θ of OBB respectively, then set a blue predicted OBB and a red label OBB at the boundary position, and calculate the SkewIoU solution result of OBB at this time, and based on the periodicity of the angle, the L1 loss value of OBB at this time; S1.3: Comparing the SkewIoU solution result with the L1 loss value, it can be seen that when the target angle exceeds the defined range, the value of θ will change suddenly, resulting in a sudden increase in the model loss value. Due to the angle mutation, there is an inconsistency between the L1 loss function and the SkewIoU, which makes it difficult for the training process to converge. In addition, the OpenCV definition method has both angle prediction breakpoints and edge prediction breakpoints at the boundary angles. S1.4: Use the vertex offset method and the center point offset method to represent the OBB, and analyze the results of the two methods. It can be seen that the vertex offset method and the center point offset method will also lead to inconsistency between the loss and the measurement. The L1 loss function is replaced by the joint optimization loss function, and the deviation between the expected prediction angle and the actual prediction angle is calculated through the updated loss function. It can be seen that the two sets of angles have large errors in the boundary angles.

3. The design method of the SAR target directional frame detection model based on continuous coding according to claim 2 is characterized in that: The specific steps for analyzing the boundary continuity conditions described in S1 are as follows: S2.1: Define E(x) as the parameter space P from the SAR ship target x to the OBB n The encoding function of , where n represents the parameter dimension, defines T(x; θ) as the transformation that rotates x counterclockwise by an angle θ, and S is the state space set of x; S2.2: To ensure the continuity of coding, it is necessary to ensure that slight rotation has the least impact on OBB parameters, that is: Then define D() as the decoding function, p as a set of OBB parameters. In addition to satisfying the encoding continuity, it is also necessary to satisfy the decoding completeness to ensure that every state of x can be accurately represented, that is: S2.3: Define L separately E () and L D () is the loss function calculated using encoding parameters and decoding parameters. During the rotation of the ship target, the loss function should not have a sudden change and should have rotation continuity, that is: The conditions corresponding to formulas (1), (2) and (3) are called coding continuity, decoding integrity and loss function rotation continuity respectively. When only coding parameters or only decoding parameters are used, only the corresponding conditions need to be met. Then, the detection model in the field of SAR ship detection is analyzed. It can be seen that each model cannot simultaneously meet the three conditions of coding continuity, decoding integrity and loss function rotation continuity.

4. The design method of the SAR target directional frame detection model based on continuous coding according to claim 3 is characterized in that: The specific steps of designing the continuous encoding method and loss function based on coordinate decomposition described in S2 are as follows: S3.1: Change the angle range of OBB in the long edge definition method to [0,π), and denote the rotation angle of OBB as θ obb , the rotation angle of the ship target is represented as θ t , and θ obb With θ t The value ranges of are [0,π) and [0,2π) respectively. Then θ obb With θ t The conversion relationship between can be expressed as: From formula (4), we can know that θ obb The period is π, and at θ t = discontinuous at 0 and π, corresponding to π / 2 and 3π / 2 in polar and rectangular coordinate systems; S3.2: The coordinates of a point on the unit circle in a rectangular coordinate system can be expressed as Its angle It can be expressed as: In formula (5), mod represents the modulo operation, arctan2 represents a quadrant function, and then the angle information of the point in the rectangular coordinate system is calculated according to the coordinates. Its value range is [0,2π). The specific calculation formula is as follows: When ω=2, formula (5) becomes: when When the value range of is [0,2π], formula (7) can be expressed as: S3.2: From formula (4) to formula (8) in step S3.1, it can be seen that when ω = 2, and When the value range of is [0,2π], formula (8) is equal to formula (4), so the coordinates of the point on the unit circle can be used to encode θ obb , its encoding function can be expressed as: E θ (i t )=(α1,α2)=(cos 2θ t ,sin 2θ t ) (9) Among them, α1 and α2 represent encoding parameters respectively, and their decoding function can be expressed as: In formula (9) and formula (10), E θ (·) and D θ (·) respectively represent the functions for encoding and decoding the angle parameter of x. Since the other four parameters (cx, cy, w, h) are not encoded or decoded, the encoding functions E(·) and D(·) for the angle parameters only are written as E θ (·) and D θ (·); S3.3: Plot θ in polar coordinates obb , α1 and α2 rotate with the target angle θ t , in order to intuitively illustrate the change of θ obb , α1 and α2 in the polar coordinate system, where the azimuth axis represents the rotation angle θ of the target t , the radial axis represents θ obb , α1 and α2, based on the polar coordinate information drawn, it can be seen that as the azimuth angle θ t α1 and α2 remain continuous without any breakpoints, and for any OBB angle θ obb , there is always a pair (α1, α2) corresponding to it, so the encoding function satisfies the rotation continuity and guarantees the decoding integrity; S3.4: Design a joint optimization loss function L using the form of a joint optimization loss function obb , and its specific calculation formula is expressed as: L obb =L SkewIoU (B p (cx p ,cy p ,w p ,h p ,θ p ),B GT (cx t ,cy t ,w t ,h t ,θ t )) (11) Among them, B p Represents the predicted OBB, with parameters (cx p ,cy p ,w p ,h p ,θ p ), B GT Represents the label OBB, the parameter is (cx t ,cy t ,w t ,h t ,θ t ), (cx, cy) represents the center position parameter, (w, h) represents the width and height parameters, θ represents the rotation angle parameter, L SkewIoU Represents the loss function approximating SkewIoU, θ p Obtained by decoding the encoding parameters, expressed as: in, and Represents the predicted coding parameters, and then an angle coding loss is added to better constrain the predicted coding parameters and achieve faster convergence. The angle coding loss is expressed as: in, and Indicates the rotation angle θ of the label OBB t Through the encoding function E θ (·) The target encoding parameters obtained, L1() represents the L1 loss function, and the total loss of the model is L total It can be expressed as: L total =w cls L cls +w obb L obb +w ang L ang (14) Among them, L cls represents the classification loss, w cls 、w obb and w ang L cls , L obb and L ang , and L total The value of is not fixed and needs to be adjusted according to the detection network.

5. The design method of the SAR target directional frame detection model based on continuous coding according to claim 4 is characterized in that: The specific steps of constructing the SAR target directional frame detection model described in S3 are as follows: S4.1: When constructing a SAR target directional frame detection model, the backbone network selects ResNet, ReResNet or ARC detection models, and the feature fusion network selects feature pyramid network, LR-FPN or BFPN detection models. Then, an angle coding prediction branch is added to the detection head of the backbone network and the feature fusion network. The output of the detection head is a two-dimensional angle coding parameter β1 and β2. S4.2: According to formula (9) in step S3.2, determine that the value range of the angle encoding parameter is [-1, 1], and then transform the output parameter to obtain the final angle encoding parameter and To make the training more stable, the specific transformation formula is as follows: After the predicted coding angle parameters are obtained, the decoding function D θ (·) Encoding parameters of the predicted angle and Decode to get θ p and compared with the predicted HBB parameters (cx p ,cy p ,w,h) parameters that make up OBB; S4.3: Use the approximate SkewIoU loss to calculate the loss between the predicted OBB and the label OBB, through the encoding function E θ (·) for the target angle θ t Encode and get the target angle encoding parameters and And the L1 loss function is used to calculate the difference between the predicted parameters and the target parameters to constrain the angle encoding; S4.4: The backbone network extracts features from the SAR image, and the extracted feature maps at different stages are recorded as C1, C2, C3, C4, and C5. Then, as the backbone network performs multi-stage convolution and pooling operations on the SAR image, the number of channels of the feature map gradually increases, and the size gradually decreases. S4.5: Feature fusion network for feature map C i A convolution operation is performed, and then the newly obtained feature map is upsampled twice and added to the previous feature map to fuse features of different scales. Finally, the fused feature map is convolved to obtain the final feature maps P2, P3, P4 and P5.

6. The design method of the SAR target directional frame detection model based on continuous coding according to claim 1 is characterized in that: The detection head described in S4.1 has three detection branches, which are responsible for category parameter prediction, HBB parameter prediction and angle encoding parameter prediction, respectively. C represents the number of categories, and K represents the number of anchor boxes corresponding to each anchor point. If it is an anchor-free network, K = 1. The specific expression of the feature fusion network convolution operation described in S4.5 is as follows: Where Conv2d(·) represents a two-dimensional convolution operation, and Upasme(·) represents a two-fold upsampling operation.