SAR image arbitrary direction ship target detection method based on visual filtering mechanism

By introducing a filtering module based on the visual filtering mechanism and polar coordinate encoding method in SAR image detection, the problems of noise interference and unclear ship target boundaries in SAR images are solved, and efficient ship target detection is achieved.

CN120236120APending Publication Date: 2025-07-01DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263066.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The noise interference in the SAR images is severe, the ship's target boundaries are unclear, and the traditional methods have low detection accuracy when changing complex backgrounds and large-scale scales.

Method used

A filter module based on the visual filtering mechanism and a five-parameter encoding method based on the polar coordinate system are adopted. The filtering module extracts the characteristics of the ship target and suppresses noise interference through bottom-up and top-down filtering processes; the five-parameter encoding method effectively represents the spatial position, size and rotation angle of the ship target in any direction through polar coordinate encoding.

Benefits of technology

It significantly improves the accuracy and efficiency of ship target detection in SAR images, solves the problems of noise interference, unclear boundaries and large-scale changes, and improves the performance of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236120A_ABST
    Figure CN120236120A_ABST
Patent Text Reader

Abstract

The invention relates to an SAR image arbitrary direction ship target detection method based on a visual filtering mechanism. The method comprises the following steps: obtaining an SAR image data set; dividing the SAR image data set into a training set and a test set; constructing a ship target detection network model for detecting ship targets in any direction in the image; training the ship target detection network model based on the training set data to obtain a trained ship target detection network model; and inputting the images in the test set into the trained ship target detection network model to predict the position of the ship target in any direction. According to the method, the ship target detection efficiency in a near-shore complex scene can be effectively improved, simple and efficient label coding is realized, the overall performance of SAR image ship target detection can be enhanced, and a new technical thought and powerful support are provided for ship target detection in any direction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of SAR image processing, and in particular to a method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism. Background Art

[0002] Synthetic Aperture Radar (SAR) is an active microwave imaging system that can image ground targets with high resolution under various weather conditions, day or night. SAR images have a wide range of applications in military monitoring, marine resource management, disaster warning, and environmental monitoring. Among them, ship target detection is an important part of the application of SAR images and has important military and civilian significance. Real-time monitoring of ships is an important task in the military field. By detecting ship targets in SAR images, it is possible to monitor and track the activities of enemy ships, provide intelligence support for military decision-making, and enhance maritime security.

[0003] However, due to the unique imaging mechanism of SAR images, problems such as noise interference, unclear boundaries of ship targets, and difficulty in distinguishing near-shore objects from ship targets occur. Therefore, to ensure the accurate detection of ship targets in any direction in SAR images, a filtering mechanism needs to be introduced to effectively suppress the interference of complex backgrounds and thereby improve the detection accuracy. In addition, in the case where ship targets are densely arranged, the traditional horizontal bounding box detection method will cause the overlap of target boundaries and the background. Therefore, the use of a rotated bounding box detection method can more accurately determine the position of ship targets.

[0004] In recent years, with the continuous progress of deep learning technology, the application of ship target detection methods based on deep learning in the offshore environment has gradually matured. However, existing methods often face the problem of low detection accuracy when dealing with ship targets with complex backgrounds and large-scale variations. To address this problem, a rotated SAR image detection model based on a visual filtering mechanism is proposed, which aims to filter out non-target areas in SAR images, enhance ship target areas, and use a five-parameter coding method based on polar coordinates to achieve efficient encoding and decoding of ship target bounding boxes. Specifically, this method uses five parameters in the polar coordinate system to represent ship targets at multiple angles. Compared with the horizontal box annotation method, it can effectively reduce the overlap of bounding boxes and background redundancy during the detection process, and solve the problem of sudden increase in loss and difficult convergence during the training process. This method provides an innovative technical solution for ship target detection in complex backgrounds in offshore SAR images. Summary of the Invention

[0005] The present invention is a method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism. On the one hand, it can significantly improve the detection performance of the model in the case of nearshore or complex backgrounds. On the other hand, it can solve the problem that the loss suddenly increases and the convergence is difficult during the training process based on the Cartesian coordinate system encoding method. The present invention adopts the following technical solutions: A method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism, including the following steps:

[0006] Obtain the SAR image dataset;

[0007] Divide the SAR image dataset into a training set and a test set;

[0008] Construct a ship target detection network model for detecting ship targets in any direction in the image;

[0009] Train the ship target detection network model based on the training set data to obtain a trained ship target detection network model;

[0010] Input the images in the test set into the trained ship target detection network model to predict the positions of ship targets in any direction.

[0011] Furthermore: The ship target detection network model includes a filtering network for extracting the features of ship targets in the image and outputting a feature map with enhanced target features and suppressed interference information;

[0012] A detection network for polar coordinate encoding of the predicted ship target features based on the five-parameter encoding method in the polar coordinate system based on the feature map output by the filtering network to achieve the prediction of the positions of ship targets.

[0013] Furthermore: The filtering network includes:

[0014] A bottom-up filtering module for imitating the process of the human brain to enhance significant targets and automatically filter noise according to the bottom-up filtering process of the visual filtering mechanism, and roughly filter the image to generate a low-level filtered feature map, process based on the gray value of the image, use average pooling to smooth the image, and then use threshold segmentation to extract high-gray value regions;

[0015] A top-down filtering module for extracting semantic information, imitating the process of the human brain to purposefully search for targets according to the top-down filtering process of the visual filtering mechanism, implementing this visual filtering behavior through a trainable neural network, and finally generating a high-level filtered feature map, imitating the mechanism of the extrastriate cortex using a trainable neural network model, and adopting a skip fusion method to lightweight the model training;

[0016] The brain-like fusion module uses the attention mechanism to fuse the original image, low-level filtered feature maps, and high-level filtered feature maps, and outputs a feature map with enhanced target features and suppressed interference information.

[0017] Further: The process of using the attention mechanism to fuse the original image, low-level filtered feature maps, and high-level filtered feature maps, and output a feature map with enhanced target features and suppressed interference information is as follows:

[0018] Use the attention mechanism to integrate the features filtered by the bottom-up filtering module into the original image;

[0019] Use the attention mechanism to integrate the features output by the bottom-up filtering module into the features output by the top-down filtering module, and further integrate the fused features with the original image through the attention mechanism to enhance the feature information;

[0020] Use the attention mechanism again to integrate the features filtered by the top-down into the original image;

[0021] Output the original image to retain the original information;

[0022] By using trainable hyperparameters, fuse the features output in the first four steps, and finally generate a feature map with the same size as the input image.

[0023] Further: It also includes using the elliptical label inscribed in the target bounding box as the label of the neural network in the top-down filtering module for training, learning to retain the ship target area and suppress the non-ship target area.

[0024] Further: The detection network consists of three networks with different functions, including a feature extraction network, a feature fusion network, and an output network;

[0025] The feature extraction network is used to extract ship target features from the output of the brain-like filtering module;

[0026] The feature fusion network is used to fuse the features of different scales output by the feature extraction network into a feature map containing various ship feature information through bilinear interpolation and channel superposition operations, for predicting the ship target bounding box;

[0027] The output network is used to accurately predict the rotated bounding box of the ship target using the features extracted by the feature extraction network and the features fused by the feature fusion network.

[0028] Further: The five-parameter coding method based on the polar coordinate system realizes the prediction of the position of the ship target by performing polar coordinate coding on the output predicted ship target features. The process is as follows:

[0029] Using the four corner coordinates (x i , y i ) of the original target bounding box to calculate the center point of the target bounding box in the polar coordinate system;

[0030] Calculate the polar radius according to the distances from the center point to the four corner points, and calculate the polar angles of each corner point in the polar coordinate system, and select the two smaller polar angles within the range of [0, π];

[0031] Normalize the polar radius of the bounding box, that is, divide it by the mean value of the short side of the bounding box and map it to the logarithmic space;

[0032] Limit the change range of the polar angle within [-1, 1] to ensure the standardization during the network training process.

[0033] Finally, output five parameters These parameters can represent the target bounding box in the polar coordinates.

[0034] The five-parameter encoding method based on the polar coordinate system in this application is designed according to the rotated bounding box labels of ship targets. In the polar coordinate system, the spatial position, size, and rotation angle of a ship target in any direction can be jointly determined by the following five parameters: the coordinates of the center point, the distance from the center point to the corner points of the rotated bounding box (polar radius), and two polar angles within the range of [0, π]. By encoding the rotated bounding box of the ship target in the polar coordinate system, the problems of sudden increase in loss and difficult convergence that occur during the training process based on the Cartesian coordinate system can be effectively solved. Therefore, introducing the five-parameter encoding method in the polar coordinate system can effectively represent the spatial position, size, and rotation angle of ship targets in any direction, thereby improving the performance of ship target detection.

[0035] This application comprehensively uses a filtering module based on a visual filtering mechanism and a five-parameter encoding method based on the polar coordinate system. On the one hand, this method draws on the process of the human eye observing targets and introduces a filtering module based on a visual filtering mechanism to accurately extract target features; on the other hand, this method uses a five-parameter encoding method based on the polar coordinate system to encode the spatial position, size, and rotation angle of ship targets in any direction, which can effectively solve the problem of boundary overlap of horizontal bounding boxes and can also solve the problems of sudden increase in loss and difficult convergence in the traditional Cartesian coordinate system during the training process. This method can not only effectively improve the detection efficiency of ship targets in complex near-shore scenes and achieve simple and efficient label encoding, but also enhance the overall performance of ship target detection in SAR images, providing new technical ideas and strong support for ship target detection in any direction.

[0036] Compared with the prior art, the present invention has the following beneficial technical effects:

[0037] 1. The present invention aims to solve the problems of difficult noise elimination and difficult foreground-background distinction in SAR images, and proposes a method for filtering SAR images before the detection network. Most of the existing filtering methods only perform simple denoising operations and fail to effectively handle the semantic relationship between targets and non-targets. In view of the above problems, the present invention proposes a filtering module based on the human eye vision filtering mechanism, which can effectively suppress the interference of complex backgrounds and non-target areas while enhancing the feature information of the target area.

[0038] 2. In order to effectively fuse the original image with the outputs of the two filtering modules while keeping the input-output scale unchanged, the present invention proposes a brain-inspired fusion module. This module fuses features through five steps and uses a hyperparameter in the trainable [0,1] interval to dynamically adjust the weights between the above three feature maps.

[0039] 3. The present invention aims to solve the overlapping problem caused by horizontal bounding boxes and the defect of containing a large amount of background information, so a five-parameter encoding method based on the polar coordinate system is introduced. Existing bounding box encoding methods have problems such as bounding box overlap and difficult training convergence. For this reason, the present invention proposes a method for encoding the spatial position, size, and rotation angle of a ship target in any direction in the polar coordinate system through five parameters to improve the accuracy and stability of encoding.

[0040] 4. In order to solve the problem of difficult regression of bounding boxes with large scale changes, the present invention processes the polar radius of the original ship target in the training stage: divides the polar radius by the mean value of the short side of the bounding box and maps it through a logarithmic function. This method effectively improves the detection ability of the detection model for ship targets of different scales. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, additional drawings can be obtained based on these drawings.

[0042] Figure 1 This is the overall framework of the method. The upper part is a brain-inspired filtering module, and the lower part is a detection module. The two parts form a complete method for detecting ship targets in any direction in SAR images based on the vision filtering mechanism;

[0043] Figure 2 This is the overall framework of the brain-inspired filtering module;

[0044] Figure 3The ship target bounding box area visualized after detector decoding. (a) is the nearshore label of the SAR image, (b) is the detection result of the detector of the present invention, (c) is the farshore label of the SAR image, and (d) is the detection result of the detector of the present invention. Detailed implementation manners

[0045] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0046] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The description of at least one exemplary embodiment below is actually only illustrative and in no way restrictive of the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners of the present invention. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0048] A method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism, comprising the following steps:

[0049] S1: Obtain an SAR image data set;

[0050] S2: Divide the SAR image data set into a training set and a test set;

[0051] S3: Construct a ship target detection network model for detecting ship targets in any direction in the image;

[0052] S4: Train the ship target detection network model based on the training set data to obtain a trained ship target detection network model;

[0053] S5: Input the images in the test set into the trained ship target detection network model to predict the positions of ship targets in any direction.

[0054] Steps S1 / S2 / S3 / S4 / S5 are executed in sequence;

[0055] The ship target detection network model includes a filtering network, which is used to extract the features of ship targets in the image and output a feature map with enhanced target features and suppressed interference information.

[0056] A detection network, which is used to perform polar coordinate encoding on the predicted ship target features based on the feature map output by the filtering network and the five-parameter encoding method in the polar coordinate system, so as to realize the prediction of the position of the ship target.

[0057] On the one hand, the ship target detection network model introduces a brain-inspired filtering network by imitating the process of human eyes observing targets, which can solve the problem of low detection efficiency of ship targets in complex scenarios such as near the shore; on the other hand, it introduces a five-parameter encoding method based on the polar coordinate system to solve the problems of horizontal bounding box target overlap and background redundancy, providing technical support for the ship target detection task in any direction of SAR images; this filtering network can effectively enhance the detection efficiency of the detection model in complex scenarios; on the other hand, to solve the problems of horizontal bounding box boundary overlap and background redundancy, a five-parameter encoding method based on the polar coordinate system is introduced to solve the problem that the loss suddenly increases and the convergence is difficult during the training process of the Cartesian coordinate system encoding method.

[0058] The filtering module is designed based on the process of the human brain observing targets. The image captured by the human eye first passes through the retina, and after the transmission of nerve signals, it finally reaches the visual cortex. The visual cortex, as the central processor of visual information, is mainly composed of the primary visual cortex (also known as the striate cortex, the first visual area, i.e., V1) and multiple extrastriate cortices (such as the second, third, fourth, and fifth visual areas, i.e., V2, V3, V4, V5).

[0059] The filtering network includes:

[0060] A bottom-up filtering module, which is used to imitate the process of the human brain enhancing significant targets and automatically filtering noise according to the bottom-up filtering process of the visual filtering mechanism, and perform rough filtering on the image to generate a low-level filtering feature map. It processes based on the gray value of the image, uses average pooling to smooth the image, and then uses threshold segmentation to extract high-gray-value regions.

[0061] A top-down filtering module, which is used to extract semantic information, imitate the process of the human brain purposefully searching for targets according to the top-down filtering process of the visual filtering mechanism, and realize this visual filtering behavior through a trainable neural network. Finally, a high-level filtering feature map is generated. It uses a trainable neural network model to imitate the mechanism of the extrastriate cortex and adopts a skip fusion method to lightweight the model training.

[0062] Among them, the top-down filtering module has a supervised learning task, so an ellipse is used as a label. The bottom-up filtering module has an unsupervised learning task, so there is no label;

[0063] It also includes using the ellipse label inscribed in the target bounding box as the label of the neural network in the top-down filtering module for training, learning to retain the ship target area and suppressing the non-ship target area.

[0064] The bottom-up filtering module and the top-down filtering module can effectively suppress the influence of complex backgrounds and noise.

[0065] The brain-inspired fusion module uses the attention mechanism to fuse the original image, low-level filtered feature map, and high-level filtered feature map, and outputs a feature map with enhanced target features and suppressed interference information.

[0066] In visual processing, the high-level nervous system coordinates the primary visual cortex with other high-level cortices to jointly complete complex visual tasks. The process of using the attention mechanism to fuse the original image, low-level filtered feature map, and high-level filtered feature map, and output a feature map with enhanced target features and suppressed interference information is as follows:

[0067] S31: Use the attention mechanism to integrate the features filtered by the bottom-up filtering module into the original image;

[0068] S32: Use the attention mechanism to integrate the features output by the bottom-up filtering module into the features output by the top-down filtering module, and further integrate the fused features with the original image through the attention mechanism to enhance the feature information;

[0069] S33: Use the attention mechanism again to integrate the features filtered by the top-down into the original image;

[0070] S34: Output the original image to retain the original information;

[0071] S35: By using trainable hyperparameters, fuse the features output in the previous four steps, and finally generate a feature map with the same size as the input image.

[0072] Furthermore, the detection network consists of three networks with different functions, including a feature extraction network, a feature fusion network, and an output network;

[0073] The feature extraction network is used to extract ship target features from the output of the brain-inspired filtering module; the feature extraction network uses ResNet50;

[0074] The feature fusion network is used to fuse features of different scales output by the feature extraction network into a feature map containing various ship feature information through operations such as bilinear interpolation and channel superposition, for predicting the bounding box of the ship target.

[0075] The output network is used to accurately predict the rotated bounding box of the ship target by using the features extracted by the feature extraction network and the features after fusion by the feature fusion network.

[0076] Calculate the loss between the bounding box of the ship target predicted by the ship target detection network model and the true bounding box encoded based on the polar coordinate system.

[0077] The five-parameter encoding method based on the polar coordinate system is designed according to the rotated bounding box label of the ship target. In the polar coordinate system, the spatial position, size, and rotation angle of the ship target in any direction can be jointly determined by the following five parameters: the coordinates of the center point, the distance from the center point to the corner point of the rotated bounding box (polar radius), and two polar angles within the range of [0, π].

[0078] By encoding the rotated bounding box of the ship target in the polar coordinate system, the problems of sudden increase in loss and difficult convergence during the training process based on the Cartesian coordinate system can be effectively solved. Therefore, introducing the five-parameter encoding method in the polar coordinate system can effectively represent the spatial position, size, and rotation angle of the ship target in any direction, thereby improving the performance of ship target detection.

[0079] The five-parameter encoding method based on the polar coordinate system realizes the prediction of the ship target position by polar coordinate encoding of the output predicted ship target features. The specific process is as follows:

[0080] Use the four corner point coordinates (x i , y i ) of the original target bounding box to calculate the center point of the target bounding box in the polar coordinate system;

[0081] Calculate the polar radius according to the distance from the center point to the four corner points, and calculate the polar angle of each corner point in the polar coordinate system, and select the two smaller polar angles within the range of [0, π];

[0082] Normalize the polar radius of the bounding box, that is, divide it by the mean value of the short side of the bounding box and map it to the logarithmic space;

[0083] Limit the change range of the polar angle within [-1, 1] to ensure the normality during the network training process.

[0084] Finally, output five parameters These parameters can represent the target bounding box in the polar coordinates.

[0085] During the training phase, when making labels, the original data (the four corner coordinates xi, yi of the bounding box) is first processed, transformed into a representation method in the polar coordinate system, and the radial distance and polar angle are normalized, and then used as training labels (polar coordinate encoding). Each image passes through the network and outputs three feature maps, namely the center point coordinates, the center point offset, and the distance from the rotation bounding box represented by polar coordinates to the corner points and two polar angles (represented by five parameters <x, y, r, angle1, angle2>, which can represent a ship target bounding box). The label data and the output data are used to calculate the loss to train the model.

[0086] Example 1: The design of the rotating object detection network is described below in combination with Figure 1 the network structure and specific examples in

[0087] As Figure 1 shown, the size of the input image is adjusted to 608×608 and fed into a brain-inspired filtering network to enhance the ship target information in the SAR image and suppress complex background information.

[0088] As Figure 2 shown, in order to hierarchically filter the SAR image features, first the image is fed into a bottom-up filtering module. To remove speckle noise, the blur function in the OpenCV library is used to smooth the image, and then a threshold segmentation method is used to filter out the targets with high gray values while retaining details, thus achieving the preliminary filtering of the image. During this process, a trainable variable is used as the threshold to dynamically adjust the gray value range in the image.

[0089] As Figure 2 shown, the preprocessed image is input into the top-down filtering module to learn the ship target area. During the downsampling process, a 4×4 convolution operation is performed on the input feature map to adjust the scale of the feature map, and two 3×3 convolutions are sequentially used for channel adjustment to output two feature maps with different channel numbers but the same scale, as shown in formula (1):

[0090]

[0091] Among them, M down represents the input feature map of downsampling, R * , R represents the feature output during downsampling.

[0092] As Figure 2 shown, during the upsampling process, an attention mechanism is used to dynamically allocate the ship target area, as shown in formula (2):

[0093] A = xf avgpool (x)f sigmoid (Conv3×3 (x)) (2)

[0094] Among them, x is the feature input, and A is the feature after attention enhancement.

[0095] As Figure 2 shown, during the upsampling process, in order to reduce the computational amount, a skip fusion method is used to fuse features of different scales, and an attention mechanism is used to allocate the ship target area. In order to fuse the feature outputs during the downsampling process, first, a 3×3 convolution operation is performed on the input features of the upsampling to adjust the number of channels, and the feature map output by the downsampling is concatenated in channels, and then a 3×3 convolution and a bilinear interpolation operation are used to realize the upsampling of the feature map scale, as specifically shown in formula (3):

[0096]

[0097] Among them, AM up is the input of the upsampling, R down is the output of the downsampling, and f inte is the bilinear interpolation function.

[0098] As Figure 2 shown, the feature map filtered from top to bottom, the feature map filtered from bottom to top, and the preprocessed original image are fused through the attention mechanism, and the weights of different information are dynamically regulated by imitating the processing process of the advanced nervous system, and finally the fused feature map fusion1 is output, as specifically shown in formula (4):

[0099]

[0100] Among them, t * is a trainable variable taken from [0 to 1], fusion * is various fusion methods, and f attention is the attention mechanism, and its formula is as follows:

[0101] f attention (x1,x2) = x2conv 1×1 (cat(avg(x1),max(x1))) (5)

[0102] As Figure 1 shown, sending fusion1 into the feature extraction network outputs three downsampled feature maps, and the sizes of the output feature maps are 512×76×76, 1024×38×38, and 2048×19×19 respectively. In order to fuse features of different scales, the method used is as Figure 1The feature fusion network shown above fuses operations such as feature map upsampling, channel concatenation, and channel adjustment to output a feature map with a size of 256×76×76. Among them, bilinear interpolation is used for upsampling the features, and a 3×3 convolution is used for channel adjustment. Next, a channel concatenation operation is performed with the next-layer feature map, and finally, a 3×3 convolution is used for channel adjustment. The specific algorithm is shown in formula (5):

[0103]

[0104] Where: f r* is the output of the feature extraction network, f r2 is the output of 512×76×76, f r3 is the output of 1024×38×38, f r4 is the output of 2048×19×19, f fusion2 is the output of the feature fusion network.

[0105] To predict the rotated bounding box, the output network first uses two 3×3 convolutions and a 1×1 convolution to output the predicted Gaussian heat map, polar coordinate encoding, and center point offset. Among them, the feature map size of the Gaussian heat map is 1×76×76, the feature size of the polar coordinate encoding is 3×76×76, and the feature size of the center point offset is 2×76×76.

[0106] The five-parameter encoding method based on the polar coordinate system is designed according to the rotated ship targets in SAR images. As Figure 1 shown, where (x, y) represents the center coordinates of the bounding box, P r represents the polar radii of the four corner points with the center of the bounding box as the pole, P α1 , P α2 represent two polar angles in [0, π]. In order to encode the eight parameters in the rectangular coordinate system into five parameters in the polar coordinate system.

[0107] For the five-parameter encoding method based on the polar coordinate system described above, the formula process (7)-(10) for polar coordinate encoding of the output predicted ship target features is used to transform the four corner point coordinates (x i , y i ) of the bounding box into Specifically as follows:

[0108] First, calculate the center (x, y) of the bounding box according to the four corner point coordinates, and establish a polar coordinate system with the center of the bounding box as the pole. Represent the four corner point coordinates in the polar coordinate system, as shown in formula (7):

[0109]

[0110] Where \(w\) and \(h\) are the width and height of the bounding box respectively. \(\beta\) is the angle between the long side of the bounding box and the horizontal direction.

[0111] Calculate the polar radius from the corner coordinates of the bounding box after center transfer, as shown in formula (8):

[0112]

[0113] Calculate the polar angle of each corner point in the polar coordinate system. To increase the stability of bounding box training, select two polar angles \(P\) α1 , \(P\) α2 in \([0, \pi]\). To solve the problem caused by the large scale variation of ship targets, first divide the polar radius by the mean of the short side, and then map it to the logarithmic space to limit the variation range of the polar radius, thereby effectively reducing the impact of the large scale variation of the target bounding box. At the same time, the range of the polar angle is mapped to \([-1, 1]\). Specifically, as shown in formula (9) and formula (10):

[0114]

[0115] Where \(\lambda\) r is the mean of the short side of the bounding box.

[0116]

[0117] During the model training process, to dynamically allocate positive and negative samples of the target, a two-dimensional Gaussian heatmap based on limited covariance is used to represent the positive samples of ship targets. As Figure 1 shown, take the \(w\) and \(h\) of the restricted bounding box as the covariance of the two-dimensional Gaussian function. This method can effectively solve the problem of blurred boundaries between ship targets when ship targets are densely arranged.

[0118] The two-dimensional Gaussian heatmap based on limited covariance can effectively distinguish the boundaries of ship targets. To obtain the rotated bounding box, perform a rotation transformation on the covariance matrix, as shown in formula (11):

[0119]

[0120] Where: \(\sigma\) 11 , \(\sigma\) 22 is the limited covariance, \(M\) h is the rotated limited covariance matrix, and \(R\) is

[0121]

[0122] Where: \(\alpha\) is the angle between the long side of the bounding box and the positive half-axis of the coordinate axis.

[0123] To obtain the covariance value of the Gaussian heatmap, it is determined by the major axis and minor axis when the original bounding box decays to a fixed value towards the center. Specifically, as shown in formulas (13), (14), and (15):

[0124]

[0125]

[0126] where P b is the point on the boundary of the bounding box, and λ s = 0.001 represents the fixed value of decay.

[0127]

[0128] Substituting the limited covariance matrix into the two-dimensional Gaussian function gives

[0129]

[0130] where, (x c , y c ) represents the center point of the ship target.

[0131] In the top-down filtering module, the cross-entropy loss function is used to calculate the loss between the predicted ship target region and the true ellipse label, and the Gaussian heatmap is used to dynamically allocate positive and negative samples, as shown in formula (17) specifically.

[0132]

[0133] where, N represents the number of positive samples, represents the predicted filtered feature map. Samples with values greater than the confidence value λ th = 0.50 in the Gaussian heatmap are set as positive samples of the ship target, and the positive sample weight W fij is generated. The value of the ship target region is 1, the background is 0, and η is a positive number approaching 0. Here, 0.0001 is taken.

[0134] For the loss function corresponding to the predicted Gaussian heatmap in the output network, formulas (18) and (19) are used to calculate the loss of the heatmap.

[0135]

[0136] where, N represents the number of positive samples, represents the predicted Gaussian heatmap, GM ∈ [0, 1] H '×W' ×1 represents the generated heatmap, W hij represents that the ship target foreground is 1 and the background is 0, α * = 2, β *= 0.05.

[0137] For the polar coordinate loss in the output network, Smooth L1 is used to calculate the loss of the mapped radial distance and polar angle.

[0138]

[0139] Among them, N represents the number of positive samples, represents the predicted radial distance and polar angle, M P represents the true radial distance and polar angle, and beta = 0.11.

[0140] For the center point offset loss of the output network, formula (21) is used to calculate the loss of the center point offset.

[0141]

[0142] Among them, N represents the number of positive samples, represents the predicted center point offset, M o represents the true center point offset, and beta = 0.11.

[0143] The loss function of the entire network can be expressed as:

[0144] L all = L F + L H + L O + L P . (22)

[0145] The present invention relates to a method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism. In order to verify the processing ability of the detection method proposed by the present invention in complex scenarios such as near the shore and the detection performance in comprehensive scenarios, near-shore data tests and comprehensive data tests are respectively carried out on the RSSDD dataset, and compared with currently relatively advanced ship target detectors in any direction in SAR images (single-stage detectors: Polar Encodings and Sasm Reppoints, two-stage detectors: ReDet and FPDDet). The results are as follows in the table:

[0146]

[0147] The experimental results show that the filtering method based on the visual filtering mechanism and the five-parameter encoding method based on the polar coordinate system adopted by the present invention significantly improve the performance of the detection model in complex near-shore environments and comprehensive scenarios, and effectively solve problems such as background interference in complex environments. Such as Figure 3As shown, the effectiveness of the present invention in detecting ship targets in any direction of SAR images is demonstrated by visualizing the detection results, where (a) is the nearshore label of the SAR image, (b) is the detection result of the detector of the present invention, (c) is the offshore label of the SAR image, Figure 3 (d) is the detection result of the detector of the present invention.

[0148] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting ship targets in arbitrary directions in SAR images based on a visual filtering mechanism, characterized by: The following steps are involved: Obtain SAR image dataset; Divide the SAR image dataset into training set and test set; Construct a ship target detection network model for detecting ship targets in any direction in the image; The ship target detection network model is trained based on the training set data to obtain a trained ship target detection network model; The images in the test set are input into the trained ship target detection network model to predict the position of ship targets in any direction.

2. According to claim 1, a method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism is characterized by: The ship target detection network model includes a filtering network for extracting features of the ship target in the image and outputting a feature map in which the target features are enhanced and interference information is suppressed; The detection network is used to predict the position of the ship target by performing polar coordinate encoding on the output predicted ship target features based on the feature map output by the filtering network and a five-parameter encoding method based on a polar coordinate system.

3. The method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism according to claim 2 is characterized by: The filtering network comprises: The bottom-up filtering module is used to imitate the human brain to enhance salient targets and automatically filter noise according to the bottom-up filtering process of the visual filtering mechanism, and roughly filter the image to generate a low-level filtering feature map. It processes the image based on its grayscale value, uses average pooling to smooth the image, and then uses threshold segmentation to extract high grayscale value areas. The top-down filtering module is used to extract semantic information. It imitates the process of the human brain purposefully searching for targets according to the top-down filtering process of the visual filtering mechanism. This visual filtering behavior is realized through a trainable neural network, and finally a high-level filtering feature map is generated. A trainable neural network model is used to imitate the mechanism of the extrastriate cortex, and a jump fusion method is used to lighten the model training. The brain-like fusion module uses the attention mechanism to fuse the original image, low-level filtered feature map and high-level filtered feature map, and outputs a feature map with enhanced target features and suppressed interference information.

4. The method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism according to claim 3 is characterized by: The process of using the attention mechanism to fuse the original image, the low-level filter feature map and the high-level filter feature map to output a feature map with enhanced target features and suppressed interference information is as follows: Use the attention mechanism to integrate the features filtered by the bottom-up filtering module into the original image; By using the attention mechanism, the features output by the bottom-up filtering module are integrated into the features output by the top-down filtering module, and the fused features are further fused with the original image through the attention mechanism to enhance the feature information; The attention mechanism is used again to integrate the top-down filtered features into the original image; Output the original image to preserve the original information; By using trainable hyperparameters, the features output from the first four steps are fused to finally generate a feature map with the same size as the input image.

5. The method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism according to claim 3 is characterized by: It also includes using the ellipse label inscribed in the target bounding box as the label of the neural network in the top-down filtering module for training, learning to retain the ship target area and suppress the non-ship target area.

6. The method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism according to claim 2, characterized in that: The detection network consists of three networks with different functions. Including feature extraction network, feature fusion network and output network; The feature extraction network is used to extract ship target features from the output of the brain-based filtering module; The feature fusion network is used to fuse the features of different scales output by the feature extraction network into a feature map containing various ship feature information through bilinear interpolation and channel superposition operations, and is used to output the prediction of the ship target bounding box; The output network is used to accurately predict the rotation bounding box of the ship target by using the features extracted by the feature extraction network and the features fused by the feature fusion network.

7. The method for detecting ship targets in any direction in SAR images based on a visual filtering mechanism according to claim 2 is characterized by: The five-parameter encoding method based on the polar coordinate system realizes the prediction of the ship target position by performing polar coordinate encoding on the output predicted ship target features. The specific process is as follows: Using the coordinates of the four corner points of the original target bounding box (x i ,y i ) Calculate the center point of the target bounding box in the polar coordinate system; Calculate the polar diameter based on the distance from the center point to the four corner points, and calculate the polar angle of each corner point in the polar coordinate system, and select the two smaller polar angles in the range [0,π]; The polar diameter of the bounding box is normalized by dividing it by the mean of the short side of the bounding box and mapped to logarithmic space; Limit the range of polar angle variation to [-1,1] to ensure the standardization of network training process; Finally, five parameters are output These parameters represent the object bounding box in polar coordinates.