A SAR Image Oblique Frame Detection Method Based on Rotation Matrix Estimation

By constructing a neural network model and using the loss function of distance weight mask, angle weight mask and rotation matrix estimation, the problem of discontinuity of angle boundaries in SAR image oblique frame detection is solved, and the accuracy of object detection is improved.

CN119672362BActive Publication Date: 2025-07-29NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411810946.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-07-29
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

The existing SAR image oblique frame object detection methods have problems such as discontinuous angle boundaries, large impact on the angle, and unstable boundaries, resulting in low accuracy of ship target detection.

Method used

Using the SAR image oblique frame detection method based on rotation matrix estimation, a neural network model is constructed, including the backbone, neck part and head part, and a distance weight mask, an angle weight mask and a loss function calculation module based on rotation matrix estimation are used to improve the cascade detection head and improve the stability of the target bounding box.

Benefits of technology

It effectively improves the problem of angle boundary discontinuity, reduces the impact of angle on target detection, and improves the accuracy of ship target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672362B_ABST
    Figure CN119672362B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting oblique frames in SAR images based on rotation matrix estimation, which constructs an SAR image data set; constructs an SAR image oblique frame detection model, and the SAR image oblique frame detection model adopts a neural network model, and the neural network model includes: a backbone part, a neck part, and a head part, wherein the backbone part is used to extract features of the input image, the neck part is used to fuse multi-scale features, and the head part realizes the detection of targets; uses the constructed SAR image data set to train the constructed neural network model; uses the trained neural network to realize the oblique frame detection of SAR image targets. The present invention improves the cascaded detection head to make the edges of the target bounding box more stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to SAR image detection technology, in particular to a SAR image oblique frame detection method based on rotation matrix estimation. Background Art

[0002] Decades of development have spawned numerous methods for target detection in synthetic aperture radar (SAR) images. Compared to optical images, SAR images suffer from poor visual readability, low resolution, speckle noise, and geometric distortion, making target detection difficult and challenging. Currently, mainstream methods include those based on structural features, grayscale features, texture features, and deep learning, with deep learning-based target detection algorithms being the most widely used. Conventional deep learning algorithms typically use horizontal boxes to detect ship targets. However, horizontal boxes contain a significant amount of background area and result in numerous redundant pixels, making classification difficult. Furthermore, densely packed ships can cause non-maximum suppression algorithms to suppress valid targets, leading to missed detections. Oblique box target detection algorithms can effectively address these drawbacks, but they also introduce new challenges.

[0003] The existing SAR image oblique frame target detection method has a series of problems such as angular boundary discontinuity, target detection is greatly affected by angles, and boundary instability. Currently, there is still a lack of a method to improve the above problems and enhance the accuracy of ship target detection. Summary of the Invention

[0004] The purpose of the present invention is to provide a SAR image slant frame detection method based on rotation matrix estimation.

[0005] The technical solution for achieving the purpose of the present invention is: a SAR image slant frame detection method based on rotation matrix estimation, comprising:

[0006] Construct a SAR image dataset, where the images in the dataset are SAR images of the target to be detected, and the targets in the SAR images are annotated using tilted boxes;

[0007] Constructing a SAR image oblique frame detection model, wherein the SAR image oblique frame detection model adopts a neural network model, and the neural network model includes: a backbone part, a neck part, and a head part, wherein the backbone part is used to extract features of the input image, the neck part is used to fuse multi-scale features, and the head part is used to detect the target;

[0008] Use the constructed SAR image dataset to train the constructed neural network model;

[0009] The trained neural network is used to realize the oblique frame detection of SAR image targets.

[0010] Preferably, the head part includes a convolution module, a distance weight mask generation module, an angle weight mask generation module, and a loss function calculation module based on rotation matrix estimation; the processing process of the neural network model for the image is as follows:

[0011] Step S2-1, the SAR image with a resolution of is successively subjected to feature extraction and multi-scale feature fusion through the backbone part and the neck part to obtain feature maps with channels and a resolution of , where is a constant, and the feature maps are sent to the convolution module;

[0012] The convolution module decodes the feature maps to obtain the center point coordinate information of the target bounding box, and the center point coordinate information contains the positions of the center points of the bounding boxes of N possible targets;

[0013] Step S2-2, the distance weight mask generation module uses the center point coordinate information output by the convolution module to generate a distance weight mask;

[0014] Step S2-3, the distance weighted convolution module uses the feature maps output by the convolution module in Step S2-1 and the distance weight mask output by the distance weight mask generation module to perform distance weighting on the feature maps output by the convolution module, that is, multiplying each channel of the feature maps output by the convolution module by the pixel value at the corresponding position of the distance weight mask to obtain the distance weighted feature maps, and decoding the distance weighted feature maps to obtain the angle information of the target bounding box;

[0015] Step S2-4, the angle weight mask generation module uses the angle information output by the distance weighted convolution module to generate an angle weight mask;

[0016] Step S2-5, the angle weighted convolution module uses the distance weighted feature maps output by the distance weighted convolution module in Step S2-3 and the angle weight mask output by the angle weight mask generation module to perform angle weighting on the distance weighted feature maps, that is, multiplying each channel of the distance weighted feature maps by the pixel value at the corresponding position of the angle weight mask to obtain new feature maps, and decoding the distance weighted and angle weighted feature maps to obtain the length and width information of the target bounding box;

[0017] Step S2-6, the loss function calculation module based on rotation matrix estimation uses the center point coordinate information output by the convolution module, the angle information output by the distance weighted convolution module, and the length and width information output by the angle weighted convolution module to calculate the loss value.

[0018] Preferably, the method for the distance weight mask generation module to generate a distance weight mask by using the center point coordinate information output by the convolution module is as follows:

[0019] Assume that the center point position of the bounding box of the th possible target is ; then the pixel coordinates of the feature map where the th possible target is located are , where , , ; create a blank single-channel image with the same resolution as the feature map for the th possible target, and the pixel value of each pixel of satisfies the formula

[0020] ,

[0021] where , is the horizontal and vertical coordinates of the th pixel of the single-channel image , and is a constant;

[0022] Add the corresponding pixels of the single-channel images corresponding to the targets, and pass the result through the sigmoid function to obtain a unique single-channel image, which is the distance weight mask.

[0023] Preferably, the method for the angle weight mask generation module to generate an angle weight mask by using the angle information output by the distance weighted convolution module is as follows:

[0024] Assume that the bounding box angle of the th possible target is , where ;

[0025] Create a blank single-channel image with the same resolution as the feature map after distance weighting by the distance weighted convolution module for the th possible target, and the pixel value of each pixel of satisfies the formula

[0026] ,

[0027] where ,

[0028] In the formula, ( , is a single-channel image The horizontal and vertical coordinates of the th pixel point and are the horizontal and vertical coordinates of the pixel point in the feature map channel where the th target is located, and

[0029] is a constant; The single-channel images of the targets are added for the corresponding pixels and the result is passed through the sigmoid function to obtain a unique single-channel image, which is the angular weight mask.

[0030] Preferably, the loss function calculation module based on the rotation matrix estimation uses the center point coordinate information output by the convolution module, the angle information output by the distance weighted convolution module, and the length and width information output by the angular weighted convolution module. The specific method for calculating the loss value is as follows:

[0031] The target oblique box information includes ( , , , , , ), where is a flag bit used to calculate the value, where , , are respectively the horizontal and vertical coordinates of the target box, , are respectively the width and length of the target box, is the angle information of the target box, and the loss function is

[0032] ,

[0033] where , is the center distance, is the aspect ratio loss, is the weight factor, is the square of the diagonal length, are respectively the minimum and maximum horizontal and vertical coordinates among the four vertices of the target box, is the intersection over union between the predicted box and the ground truth box; ; is the class loss; , , , , are respectively the horizontal and vertical coordinates, width, length, and angle of the center point of the predicted box.

[0034] Compared with the prior art, the remarkable advantages of the present invention are as follows: by improving the cascaded detection head, the edge of the target bounding box becomes more stable; by adopting an improved loss function, the problem of discontinuous angle boundaries is effectively improved, and the influence of angles on target detection is significantly reduced. Brief Description of the Drawings

[0035] Figure 1 It is a training and detection flow chart of an inclined box detection model for SAR images based on rotation matrix estimation.

[0036] Figure 2 It is a schematic diagram of the method for generating a distance weight mask by the distance weight mask generation module.

[0037] Figure 3 It is a schematic diagram of the method for generating an angle weight mask by the angle weight mask generation module.

[0038] Figure 4 It is the SAR image used in the embodiment.

[0039] Figure 5 It is a schematic diagram of the generated single-channel distance weight mask.

[0040] Figure 6 It is a schematic diagram of the generated single-channel angle weight mask.

[0041] Figure 7 It is the labeled SAR image to be detected and the detection result image. Detailed Embodiment

[0042] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the accompanying drawings. The present invention can be implemented in different forms and is not limited to the embodiments described in the text. On the contrary, the embodiments are provided to make the disclosure of the present invention more thorough and comprehensive. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0043] In this embodiment, it is assumed that the RSDD-SAR image dataset is selected, and a certain image such as Figure 4 (left), with a resolution of 512*512, and the feature map generated by the SAR image is as Figure 4 (right), with a resolution of 32*32, having 64 channels, containing one target. From the annotation information of the target, it can be known that the position of the target on the feature map is (16, 16), and the angle of the target is 45°, and the prior calculations are 8 and 12 respectively.

[0044] Step S1, select training samples to construct an SAR image dataset. The images in the dataset are the images of the targets to be detected, and the targets in the SAR images are labeled with inclined boxes.

[0045] Step S2, construct an SAR image inclined box detection model. The SAR image inclined box detection model adopts a neural network model. The neural network model is divided into three parts: backbone, neck, and head. Among them, the backbone part realizes the feature extraction of the input image, the neck part realizes the fusion of multi-scale features, and the head part realizes the detection of targets; the head part includes a convolution module, a distance weight mask generation module, an angle weight mask generation module, and a loss function calculation module based on rotation matrix estimation;

[0046] The processing process of the neural network model for images is as follows:

[0047] Step S2-1, input an SAR image with a resolution of . After the feature extraction and multi-scale feature fusion of the backbone and neck parts of the neural network, obtain feature maps with channels and a resolution of . Here,

[0048] is a constant, and the feature maps are sent to the head part. The feature maps are sent to the convolution module. The convolution module realizes the decoding of the feature maps and obtains the center point coordinate information of the target bounding box. The center point coordinate information contains the positions of the center points of the bounding boxes of multiple possible targets. The center point coordinate information is respectively sent to the distance weight mask generation module, the angle weight mask generation module, and the loss function calculation module based on rotation matrix estimation, and the feature maps decoded by the convolution module are sent to the distance weighted convolution module;

[0049] Step S2-2, the distance weight mask generation module uses the center point coordinate information output by the convolution module to generate a distance weight mask and outputs it to the distance weighted convolution module; among them, the method of generating the distance weight mask is as Figure 2 shown. The center point coordinate information contains the positions of the center points of the bounding boxes of multiple possible targets. Assume that the position of the center point of the bounding box of the th possible target is . The pixel coordinates of the pixels in the feature map where this target is located can be obtained as . Here, , , . Create a blank single-channel image with the same resolution as the feature map for the th possible target. The pixel value of each pixel in satisfies the formula

[0050] ,

[0051] Among them, , is the horizontal and vertical coordinates of the th pixel point,

[0052] The single-channel images corresponding to the targets obtained by decoding the feature map in step S2-1 are pixel-added for the corresponding pixels and the result is passed through the sigmoid function to obtain a unique single-channel image, which is the distance weight mask of the feature map.

[0053] Step S2-3, the distance-weighted convolution module uses the feature map output by the convolution module in step S2-1 and the distance weight mask output by the distance weight mask generation module to perform distance weighting on the feature map output by the convolution module, that is, each channel of the feature map output by the convolution module is multiplied by the pixel value at the corresponding position of the distance weight mask to obtain a distance-weighted feature map, and the distance-weighted feature map is decoded to obtain the angle information of the target bounding box. The angle information contains the angle values of multiple possible targets, which are respectively sent to the angle weight mask generation module and the loss function calculation module based on the rotation matrix estimation, and the distance-weighted feature map is sent to the angle-weighted convolution module;

[0054] Step S2-4, the angle weight mask generation module uses the angle information output by the distance-weighted convolution module to generate an angle weight mask and outputs it to the angle-weighted convolution module; among them, the method for generating the angle weight mask is as Figure 3 , as follows, the angle information of the input target box contains angle values of multiple possible targets. Assuming that the bounding box angle of the th possible target is , where .

[0055] Create a blank single-channel image with the same resolution as the distance-weighted feature map output by the distance-weighted convolution module in step S2-3 for the th possible target , The pixel values of each pixel point of satisfy the formula

[0056] ,

[0057] Among them, , , is The horizontal and vertical coordinates of the first pixel point, , are the horizontal and vertical coordinates of the pixel points in the feature map channel where the first target is located, and

[0058] is a constant. The single-channel images of the targets obtained by decoding the distance-weighted feature map in step S2-3 are added pixel by pixel, and the result is passed through the sigmoid function to obtain a unique single-channel image, which is the angular weight mask of the feature map.

[0059] In step S2-5, the angular weighted convolution module uses the distance-weighted feature map output by the distance weighted convolution module in step S2-3 and the angular weight mask output by the angular weight mask generation module to perform angular weighting on the distance-weighted feature map, that is, each channel of the distance-weighted feature map is multiplied by the pixel value at the corresponding position of the angular weight mask to obtain a new feature map, and the distance-weighted and angular-weighted feature map is decoded to obtain the length and width information of the target bounding box, which is output to the loss function calculation module based on the rotation matrix estimation.

[0060] In step S2-6, the loss function calculation module based on the rotation matrix estimation uses the center point coordinate information output by the convolution module, the angular information output by the distance weighted convolution module, and the length and width information output by the angular weighted convolution module to calculate and output the loss value. The loss function calculation method is as follows: the rotation matrix is used to replace the target angular information, and the original target oblique box information contains five items ( , , , , ), , where and are the horizontal and vertical coordinates of the target box respectively, and are the width and length of the target box respectively, is the angular information of the target box, which is now replaced by six new elements ( , , , , , ), where is a flag bit used to calculate the

[0061] value, and the original loss function is

[0062] Now, add the sine-cosine conversion symbol loss, and the new loss function is

[0063] ,

[0064] where , is the center distance, is the aspect ratio loss, is the weight factor, is the square of the diagonal length, , , , are respectively the minimum and maximum horizontal and vertical coordinates among the four vertices of the target bounding box, is the intersection over union (IoU) between the predicted bounding box and the ground truth bounding box; ; is the class loss; , , , , are respectively the horizontal coordinate, vertical coordinate, width, length, and angle of the center point of the predicted bounding box.

[0065] Step S3: Use the constructed SAR image dataset to train the constructed neural network model;

[0066] Divide the SAR images in the dataset into multiple batches and send them into the constructed neural network to obtain the loss value. Use the loss value to complete backpropagation and update the parameters. Repeat the above steps until the model meets the accuracy requirements.

[0067] Step S4: Use the trained neural network to implement the detection of oblique bounding boxes for SAR image targets. The specific method is as follows:

[0068] Step S4-1: Input the SAR image to be detected into the trained neural network. After feature extraction by the backbone and neck parts of the neural network and multi-scale feature fusion, obtain feature maps with channels and a resolution of where

[0069] Step S4-2: The distance weight mask generation module generates a distance weight mask by using the center point coordinate information output by the convolution module, and outputs it to the distance weighted convolution module;

[0070] Step S4-3: The distance weighted convolution module uses the feature map output by the convolution module in Step S4-1 and the distance weight mask output by the distance weight mask generation module to perform distance weighting on the feature map, that is, multiplies each channel of the feature map by the pixel value at the corresponding position of the distance weight mask one by one to obtain a distance weighted feature map, and decodes the distance weighted feature map to obtain the angle information of the target bounding box. The angle information contains the angle values of multiple possible targets, and is sent to the angle weight mask generation module respectively. The distance weighted feature map is sent to the angle weighted convolution module;

[0071] Step S4-4: The angle weight mask generation module generates an angle weight mask by using the angle information output by the distance weighted convolution module, and outputs it to the angle weighted convolution module;

[0072] Step S4-5: The angle weighted convolution module uses the distance weighted feature map output by the distance weighted convolution module in Step S4-3 and the angle weight mask output by the angle weight mask generation module to perform angle weighting on the distance weighted feature map, that is, multiplies each channel of the feature map by the pixel value at the corresponding position of the angle weight mask one by one to obtain an angle weighted feature map, and decodes the distance weighted and angle weighted feature map to obtain the length and width information of the target bounding box.

[0073] Step S4-6: According to the center point coordinate information of the target bounding box output by the convolution module in S4-1, the angle information of the target bounding box output by the distance weighted convolution module in S4-3, and the length and width information of the target bounding box output by the angle weighted convolution module in S4-5, determine the position and shape of the target to achieve target detection. The detected target is as Figure 7 (right). It can be Figure 7 observed that the target detection is correct and the model performance is good.

Claims

1. A method for detecting oblique frames in SAR images based on rotation matrix estimation, characterized in that Including: Construct an SAR image dataset, where the images in the dataset are SAR images of the target to be detected, and the targets in the SAR images are labeled with inclined boxes; Construct an SAR image inclined box detection model. The SAR image inclined box detection model uses a neural network model, and the neural network model includes: a backbone part, a neck part, and a head part. Among them, the backbone part is used to extract features of the input image, the neck part is used to fuse multi-scale features, and the head part is used to detect targets; the head part includes a convolution module, a distance weight mask generation module, an angle weight mask generation module, and a loss function calculation module based on rotation matrix estimation; the processing process of the neural network model for the image is as follows: Step S2-1, SAR images with a resolution of are successively subjected to feature extraction and multi-scale feature fusion through the backbone part and the neck part to obtain feature maps with a resolution of channels, where is a constant, and the feature maps are sent to the convolution module; The convolution module decodes the feature map to obtain the center point coordinate information of the target bounding box, and the center point coordinate information contains the positions of the center points of the bounding boxes of N possible targets; Step S2-2: The distance weight mask generation module uses the center point coordinate information output by the convolution module to generate a distance weight mask; Step S2-3: The distance weighted convolution module uses the feature map output by the convolution module in step S2-1 and the distance weight mask output by the distance weight mask generation module to perform distance weighting on the feature map output by the convolution module, that is, multiply each channel of the feature map output by the convolution module by the pixel value at the corresponding position of the distance weight mask to obtain a distance-weighted feature map, and decode the distance-weighted feature map to obtain the angle information of the target bounding box; Step S2-4: The angle weight mask generation module uses the angle information output by the distance weighted convolution module to generate an angle weight mask; Step S2-5: The angle weighted convolution module uses the distance-weighted feature map output by the distance weighted convolution module in step S2-3 and the angle weight mask output by the angle weight mask generation module to perform angle weighting on the distance-weighted feature map, that is, multiply each channel of the distance-weighted feature map by the pixel value at the corresponding position of the angle weight mask to obtain a new feature map, and decode the distance-weighted and angle-weighted feature map to obtain the length and width information of the target bounding box; Step S2-6: The loss function calculation module based on rotation matrix estimation uses the center point coordinate information output by the convolution module, the angle information output by the distance weighted convolution module, and the length and width information output by the angle weighted convolution module to calculate the loss value; Use the constructed SAR image dataset to train the constructed neural network model; Use the trained neural network to implement the inclined box detection of the SAR image target.

2. The SAR image skew frame detection method based on rotation matrix estimation according to claim 1, wherein, The specific method for the distance weight mask generation module in step S2-2 to generate a distance weight mask using the center point coordinate information output by the convolution module is as follows: Assume that the center point position of the bounding box of the -th possible target is , then the pixel coordinates of the feature map where the -th possible target is located are , where , , ; create a blank single-channel image with the same resolution as the feature map for the -th possible target , and the pixel values of each pixel of satisfy the formula , Among them, , is the horizontal and vertical coordinates of the th pixel of the single-channel image, is a constant; Add the corresponding pixels of the single-channel images corresponding to the targets, and pass the result through the sigmoid function to obtain a unique single-channel image, which is the distance weight mask.

3. The SAR image skew frame detection method based on rotation matrix estimation according to claim 1, characterized in that The specific method for the angle weight mask generation module to generate an angle weight mask using the angle information output by the distance weighted convolution module is as follows: Suppose the bounding box angle of the th possible target is , where ; Create a blank single-channel image with the same resolution as the distance-weighted feature map of the distance-weighted convolution module for the th possible target , The pixel values of each pixel Satisfy the formula , Among them, , In the formula, ( , ) is a single-channel image The horizontal and vertical coordinates of the th pixel point, , are the horizontal and vertical coordinates of the pixel points of the feature map channel where the th target is located; is a constant; Add the corresponding pixels of the single-channel images of the targets, pass the result through the sigmoid function to obtain a unique single-channel image, which is the angular weight mask.

4. The SAR image skew frame detection method based on rotation matrix estimation according to claim 1, characterized in that The loss function calculation module based on rotation matrix estimation uses the center point coordinate information output by the convolution module, the angle information output by the distance weighted convolution module, and the length and width information output by the angle weighted convolution module. The specific method for calculating the loss value is as follows: The target oblique box information includes ( , , , , , ), where is a flag bit used to calculate the value. Among them, , are respectively the abscissa and ordinate of the target box, Among them are the abscissa and ordinate of the target box respectively, , are respectively the width and length of the target box, is the angle information of the target box, and the loss function is , Among them , is the center distance, is the aspect ratio loss, is the weight factor, is the square of the diagonal length, are respectively the minimum and maximum horizontal and vertical coordinates among the four vertices of the target box, is the intersection over union between the predicted box and the ground truth box; ; is the class loss; , , , , are respectively the horizontal coordinate, vertical coordinate, width, length and angle of the center point of the predicted box.

Citation Information

Patent Citations

  • Lightweight multi-scale remote sensing image rotating target detection method and system

    CN116030360A