WSSRNet-based SAR image ship small target detection method

By adopting a multi-scale wavelet convolutional residual module and shallow jump residual structure in SAR image ship detection, combined with a bounding box regression strategy with shape adaptive bounding box regression problems, the missed detection and positioning deviation problems caused by scattering effects and noise interference in SAR images are solved, and the detection accuracy and recall rate are significantly improved.

CN119942092AActive Publication Date: 2025-05-06WUXI UNIV

Patent Information

Application Number
CN202510421867.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In SAR images, small ship targets are missed detection and positioning deviations due to scattering effects and noise interference. Especially in dense scenarios, it is difficult for the prior art to effectively extract the context information and spatial characteristics of small targets.

Method used

Using the SAR image ship small object detection method based on WSSRNet, the feature extraction and positioning accuracy of small objects is enhanced by constructing a multi-scale wavelet convolutional residual module (MWTRM) and shallow jump residual structure (SSRM).

Benefits of technology

It significantly improves the recall and detection accuracy of small targets in SAR images, reduces missed detection and positioning deviations, and ensures real-time detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942092A_ABST
    Figure CN119942092A_ABST
Patent Text Reader

Abstract

The invention discloses a WSSRNet-based SAR image ship small target detection method. The method comprises the steps of obtaining an SAR ship image data set and performing preprocessing; building an SAR image vessel detection network based on the WSSRNet; inputting the training and verification set into the network for training, calculating a loss function and carrying out back propagation to obtain a trained optimal parameter network; and inputting the test set into the trained optimal parameter network, and outputting a vessel detection graph of the SAR vessel image. According to the method, the context feature extraction capability of the small target is enhanced by constructing a multi-scale wavelet convolution residual module, the spatial positioning precision is improved by designing a shallow jump residual structure, and the problems of missing detection and positioning deviation caused by a scattering effect and noise interference of the ship small target in an SAR image are effectively solved by combining a shape-adaptive bounding box regression strategy, so that the accuracy of positioning of the ship small target in the SAR image is improved. And the recall rate and the detection precision of the small target in the dense scene are remarkably improved while the real-time detection efficiency is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a SAR image ship small target detection method based on WSSRNet, and belongs to the technical field of small target detection in image processing. Background Art

[0002] Synthetic Aperture Radar (SAR) is an active radar sensor system that uses the principle of synthetic aperture to generate high-resolution SAR images by emitting electromagnetic waves on a high-altitude platform and collecting return signals at different locations. Compared with optical sensors, this system is not affected by factors such as weather, clouds, and light, and can monitor resources in real time around the clock. Using SAR images for ship target detection is of great significance in the fields of marine environment monitoring, maritime traffic management, and marine resource development.

[0003] With the continuous development of deep learning technology, convolutional neural networks have shown great potential in the field of SAR target detection. Target detection algorithms based on convolutional neural networks can be divided into two-stage detection algorithms and single-stage detection algorithms. Two-stage algorithms are mainly represented by R-CNN, Faster R-CNN, and Mask R-CNN, which generate candidate box areas in the image and then use classification and regression techniques to classify and locate the candidate box areas. It has the advantages of high detection accuracy and good results, but because the two-stage algorithm separates the generation of candidate box areas from target classification, it has a large amount of calculation and slow detection speed.

[0004] However, in the SAR ship target detection task, the specific shape, size and texture features of the ship in the SAR image are easily affected by background noise and clutter, which to some extent affects the feature extraction effect of the above algorithms. To this end, Zhu et al. used a repeated bilateral feature pyramid network (DB-FPN) to enhance the fusion effect of spatial and semantic details. Peng et al. used an adaptive multi-scale feature fusion structure and a channel attention mechanism to improve the network's sensitivity to features. Yan et al. used a selective coordinate attention mechanism to enhance the extraction and representation capabilities of ship target features. Although the above studies have made some progress, the detection performance for small targets is still limited. To this end, Su et al. used the branch structure of the inception module in the shallow feature enhancement network structure, and used a dilated convolution to expand the visual receptive field of the feature map, thereby enhancing the network's adaptability to small-sized ship targets.

[0005] Although the above methods have improved the recognition accuracy of small targets in SAR images to a certain extent, due to background interference, SAR scattering effect and resolution limitation, the characteristics of ship targets in near-shore scenes are more easily confused with background interference, especially when the ship targets are smaller and densely distributed, the problems of missed detection and false detection are more serious. Therefore, how to better focus on the contextual information and spatial feature information of small targets, improve the feature richness while enriching the gradient flow, and further enhance the feature expression ability of the target has become a key issue in SAR image ship detection. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a SAR image ship small target detection method based on WSSRNet, enhance the small target context feature extraction capability by constructing a multi-scale wavelet convolution residual module, design a shallow jump residual structure to improve the spatial positioning accuracy, and combine with a shape-adaptive bounding box regression strategy to effectively solve the problems of missed detection and positioning deviation of small ship targets in SAR images caused by scattering effects and noise interference, and significantly improve the recall rate and detection accuracy of small targets in dense scenes while ensuring real-time detection efficiency.

[0007] The present invention adopts the following technical solutions to solve the above technical problems: The SAR image ship small target detection method based on WSSRNet includes the following steps: Step 1, obtain a SAR ship image dataset, preprocess the dataset, and divide it into a training set, a validation set, and a test set according to a set ratio; Step 2: Build a SAR image ship detection network based on WSSRNet, including the backbone network, neck network and detection layer; A multi-scale wavelet convolution residual module is constructed in the backbone network to extract feature maps C2, C3, C4 and C5 of four different scales; The neck network includes upsampling and downsampling stages. In the upsampling stage, the four feature maps C2, C3, C4 and C5 of different scales extracted by the backbone network are upsampled, concatenated and partially fused across stages to obtain the feature map P2. In the downsampling stage, a shallow skip residual structure is used to perform skip connections and feature fusion on the same-level features extracted by the backbone network and the features fused by the neck network to obtain feature maps P3 and P4. The detection layer is used to detect ship targets of small, medium and large scales based on feature maps P2, P3 and P4; Step 3, input the training set and the validation set into the SAR image ship detection network based on WSSRNet built in step 2 for training and validation, calculate the loss function and perform back propagation, update the network parameters, and obtain the trained optimal parameter network; Step 4: Input the test set into the trained optimal parameter network and output the ship detection map of the SAR ship image.

[0008] As a preferred solution of the present invention, in the SAR image ship detection network based on WSSRNet, The backbone network includes the first to fifth convolution modules, the first to fourth multi-scale wavelet convolution residual modules and the fast spatial pyramid pooling module; the SAR ship image input by the backbone network is sequentially subjected to feature extraction by the first convolution module, the second convolution module and the first multi-scale wavelet convolution residual module to obtain a feature map C2; C2 is sequentially subjected to feature extraction by the third convolution module and the second multi-scale wavelet convolution residual module to obtain a feature map C3; C3 is sequentially subjected to feature extraction by the fourth convolution module and the third multi-scale wavelet convolution residual module to obtain a feature map C4; C4 is sequentially subjected to feature extraction by the fifth convolution module, the fourth multi-scale wavelet convolution residual module and the fast spatial pyramid pooling module to obtain a feature map C5; The neck network includes the sixth to seventh convolution modules, the first to fifth cross-stage partial fusion modules, and the first to third upsampling modules; the first upsampling module is used to upsample C5, and the upsampling result is channel-joined with C4, and the joint result is subjected to feature fusion by the first cross-stage partial fusion module to obtain the feature map F0; the second upsampling module is used to upsample F0, and the upsampling result is channel-joined with C3, and the joint result is subjected to feature fusion by the second cross-stage partial fusion module to obtain the feature map F1; the third upsampling module is used to perform feature fusion on F1 Upsampling, the upsampling result is channel-joined with C2, and the joint result is subjected to feature fusion by the third cross-stage partial fusion module to obtain feature map P2; P2 is subjected to feature extraction by the sixth convolution module, and then channel-joined with F1 and the output of the third convolution module, and the joint result is subjected to feature fusion by the fourth cross-stage partial fusion module to obtain feature map P3; P3 is subjected to feature extraction by the seventh convolution module, and then channel-joined with F0 and the output of the fourth convolution module, and the joint result is subjected to feature fusion by the fifth cross-stage partial fusion module to obtain feature map P4; The detection layer includes the first to third detection heads. The feature map P2 output by the third cross-stage partial fusion module is used as the input of the first detection head for detecting small-scale ship targets; the feature map P3 output by the fourth cross-stage partial fusion module is used as the input of the second detection head for detecting medium-scale ship targets; the feature map P4 output by the fifth cross-stage partial fusion module is used as the input of the third detection head for detecting large-scale ship targets.

[0009] As a preferred embodiment of the present invention, the specific process of step 1 is as follows: The HRSID ship image dataset was obtained, including 5604 SAR images of 800×800 pixels in different sea conditions, including ocean, offshore, coastal ports, and different sea conditions. There were 16951 ship targets in total, of which small, medium, and large ship targets accounted for 54.5%, 43.5%, and 2% of the total, respectively. The original images and the corresponding label images in the HRSID ship image dataset are cropped into images of 640×640 pixels in sequence. The cropped HRSID ship image dataset is randomly divided into training set, validation set and test set in the ratio of 7:2:1.

[0010] As a preferred solution of the present invention, the first and fourth multi-scale wavelet convolution residual modules have the same structure, and both include the eighth to tenth convolution modules, the first wavelet transform convolution module and the first to second bottleneck modules; The input of the first or fourth multi-scale wavelet convolution residual module is evenly divided into four feature subsets after passing through the eighth convolution module, and the second feature subset is sequentially passed through the first wavelet transform convolution module and the ninth convolution module to obtain the output of the ninth convolution module; the output of the ninth convolution module is added element by element to the third feature subset and used as the input of the first bottleneck module; the output of the first bottleneck module is added element by element to the fourth feature subset and used as the input of the second bottleneck module; the output of the second bottleneck module is channel-joined with the output of the first bottleneck module, the output of the ninth convolution module and the first feature subset and used as the input of the tenth convolution module; the output of the tenth convolution module is channel-joined with the input of the first or fourth multi-scale wavelet convolution residual module to obtain the output of the first or fourth multi-scale wavelet convolution residual module.

[0011] As a preferred solution of the present invention, the second and third multi-scale wavelet convolution residual modules have the same structure, and both include the eighth to tenth convolution modules, the first wavelet transform convolution module and the first to sixth bottleneck modules; The input of the second or third multi-scale wavelet convolution residual module is evenly divided into four feature subsets after passing through the eighth convolution module, and the second feature subset is sequentially passed through the first wavelet transform convolution module and the ninth convolution module to obtain the output of the ninth convolution module; the output of the ninth convolution module is added element by element to the third feature subset and used as the input of the first bottleneck module; the output of the first bottleneck module is sequentially passed through the third bottleneck module and the fourth bottleneck module to obtain the output of the fourth bottleneck module; the output of the fourth bottleneck module is added element by element to the fourth feature subset and used as the input of the second bottleneck module; the output of the second bottleneck module is sequentially passed through the fifth bottleneck module and the sixth bottleneck module to obtain the output of the sixth bottleneck module; the output of the sixth bottleneck module is channel-joined with the output of the fourth bottleneck module, the output of the ninth convolution module and the first feature subset to serve as the input of the tenth convolution module; the output of the tenth convolution module is channel-joined with the input of the second or third multi-scale wavelet convolution residual module to obtain the output of the second or third multi-scale wavelet convolution residual module.

[0012] As a preferred solution of the present invention, the first to sixth bottleneck modules have the same structure, and all include a second wavelet transform convolution module and an eleventh convolution module. The input of each bottleneck module is sequentially passed through the second wavelet transform convolution module and the eleventh convolution module to obtain the output of the eleventh convolution module; the output of the eleventh convolution module is channel-joined with the input of each bottleneck module to obtain the output of each bottleneck module; The convolution kernel sizes of the eighth to eleventh convolution modules are all 1×1, and the convolution kernel sizes of the first and second wavelet transform convolution modules are both 3×3.

[0013] As a preferred solution of the present invention, the fast spatial pyramid pooling module includes twelfth to thirteenth convolution modules and first to third maximum pooling layers. The input of the fast spatial pyramid pooling module is sequentially passed through the twelfth convolution module, the first maximum pooling layer, the second maximum pooling layer and the third maximum pooling layer to obtain the output of the third maximum pooling layer. The output of the third maximum pooling layer is channel-joined with the output of the first maximum pooling layer and the output of the second maximum pooling layer, and then enters the thirteenth convolution module. The output of the thirteenth convolution module is the output of the fast spatial pyramid pooling module.

[0014] As a preferred solution of the present invention, in step 3, the loss function adopts the Inner-ShapeIoU loss function, combines ShapeIoU with Inner-IoU, embeds the dynamic auxiliary frame mechanism into the weighted loss of the geometric features of the frame, and generates the auxiliary frame as follows: , , , , , , , in, and Respectively represent the value of the center point of the real frame on the x and y axes, and Represent the width and height of the real box respectively, , , , Respectively represent the left, right, top, and bottom boundary coordinates of the dynamic auxiliary frame of the real frame, and Respectively represent the value of the center point of the prediction box on the x and y axes, and Represents the width and height of the prediction box, , , , Respectively represent the left, right, top, and bottom boundary coordinates of the dynamic auxiliary box of the prediction box, is the scale factor, which is used to control the scale of the dynamic auxiliary frame to adjust the sensitivity of the regression gradient; Represents the intersection-over-union ratio of the predicted box and the true box; The shape weight is calculated based on the aspect ratio of the real box, the center point offset and size difference are constrained, and finally the regression loss of the border is obtained. The implementation formula is as follows: , , , , , , , in, and are shape weights, is the scaling factor, represents the shape distance loss, represents the shape similarity loss, Represents the diagonal length of the minimum enclosing box covering the predicted box and the true box, represents the Inner-ShapeIoU loss, Based on shape weights The weighted term calculated about the difference between the predicted box and the true box width, Represents shape weights A weighted term for calculating the height difference between the predicted box and the true box.

[0015] A computer device comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the steps of the WSSRNet-based SAR image ship small target detection method are implemented.

[0016] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the WSSRNet-based SAR image ship small target detection method are implemented.

[0017] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects: 1. The present invention constructs a multi-scale wavelet convolution residual module (MWTRM) in the backbone feature extraction network, obtains a richer receptive field by introducing wavelet transform convolution (WTConv), combines the residual structure to fuse multi-scale features, pays more attention to the contextual information and spatial feature information of small targets, and reduces the computational overhead.

[0018] 2. The present invention proposes a shallow skip residual structure (SSRM), which enhances the utilization of position information and shallow detail features by designing a shallow feature pyramid structure, thereby improving the positioning accuracy and recall rate of small targets. At the same time, it uses cross-stage skip connections to integrate multi-level features, enriching the gradient flow while improving the feature richness, and further enhancing the feature expression ability of the target.

[0019] 3. The present invention adopts Inner-ShapeIoU as the loss function of the network. By introducing the ideas of auxiliary boxes and Shape-IoU, the scale of the auxiliary box is adaptively adjusted and combined with the border shape information to alleviate the scale difference problem in border regression, thereby improving the accuracy and convergence speed of border regression. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flowchart of the WSSRNet-based SAR image ship small target detection method proposed by the present invention; Figure 2 This is the network structure diagram of SAR image ship small target detection based on WSSRNet proposed by the present invention; Figure 3 It is a network structure diagram of the MWTRM module proposed in the present invention; Figure 4 This is a comparison test chart between the WSSRNet proposed in this invention and other Yolo series algorithms. DETAILED DESCRIPTION

[0021] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.

[0022] Accurate SAR image ship detection is of great significance for marine surveillance and management, such as maritime traffic supervision, illegal fishing monitoring, marine resource exploration and maritime search and rescue. How to effectively identify small-sized ship targets in dense nearshore scenes, suppress background clutter interference and capture multi-scale scattering features has become a key issue in SAR image ship target detection. Aiming at the problems of background interference, feature confusion and missed detection in SAR image nearshore dense ship small target detection, a SAR image ship small target detection method based on WSSRNet is proposed. Figure 1 As shown, the specific steps are as follows: Step 1: Obtain the HRSID ship detection dataset, which contains SAR images generated in different scenes such as open sea and near ports and multiple polarizations. After preprocessing the dataset, divide it into training set, validation set and test set according to the set ratio.

[0023] The data set is obtained as follows: The HRSID ship detection data set is obtained, which contains 5604 SAR images of size 800×800 generated by different scenes such as open sea and near ports and multiple polarizations. There are 16951 ship targets in total, of which the number of small, medium and large ship targets accounts for 54.5%, 43.5% and 2% of the total, respectively.

[0024] The data set preprocessing is as follows: the original image and the corresponding label image in the HRSID data set are cropped into images of size 640×640 pixels in sequence to avoid overlapping areas.

[0025] The specific division of the data set is as follows: the 640×640 data set is randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0026] Step 2: Build a SAR image ship detection network based on Wavelet Shallow Residual Network (WSSRNet), such as Figure 2 As shown, it includes the backbone network (Backbone), the neck network (Neck) and the detection layer (Head).

[0027] Step 2.1: Build the backbone network (Backbone), which includes 5 CBS modules, 4 multi-scale wavelet transform convolutional residual modules (MWTRM) and 1 fast spatial pyramid pooling (Spatial Pyramid Pooling-Fast, SPPF) module. The CBS module consists of a convolutional layer (Conv), batch normalization (BN) and SiLu activation function, which is used for the preliminary extraction of image features and the size transformation of the image. The convolution kernel size of the convolution layer is 3 and the step size is 2.

[0028] like Figure 3 As shown in the figure, the MWTRM module uses a multi-branch residual structure to perform jump connections on different feature extraction stages. It is composed of three CBS modules with a convolution kernel size of 1, one WTCBS module with a convolution kernel size of 3, and several bottleneck modules stacked with one WTCBS module and one CBS module. The input feature first passes through a 1×1 CBS to double its channel number, and then is evenly divided into four feature subsets. The first feature subset is used to retain low-level features; the second feature subset passes through two convolution modules consisting of a convolution layer, batch normalization, and an activation function. The first convolution module uses a wavelet transform convolution. convolutional, WTConv) as the convolution layer, by using Haar wavelet transform to perform multi-level decomposition of the input image to retain more spatial resolution, thereby increasing the receptive field of convolution and avoiding over-parameterization of the network; the second one uses point-by-point convolution; the processing process of the third and fourth feature subsets is similar to that of the second feature subset, the difference is that before the third and fourth feature subsets perform the operation of the two convolution modules, the feature map is first added to the feature map processed by the previous branch, and then input into the bottleneck block composed of two different convolution modules and skip layer connections; finally, the feature maps of the four branches are spliced, the number of channels is compressed and the number of parameters is reduced by point-by-point convolution, and the gradient flow is further enriched by skip connections on the periphery of the module. The Bottleneck module is composed of two CBS residuals with a convolution kernel size of 3.

[0029] The SPPF module is composed of 2 CBS with a convolution kernel size of 1 and 3 maximum pooling layers with a convolution kernel size of 5 through residuals. The 2 CBS with a convolution kernel size of 1 are used to change the number of channels. After the above operations, the original image in the backbone network obtains the features of interest through each MWTRM module and further extracts effective features, which improves the network's feature extraction ability for small targets.

[0030] Step 2.2: Build the Neck network (Neck), using the output feature map of the SPPF module at the end of the backbone network in step 2.1 as the input feature map of the Neck part. Neck consists of 3 upsampling operations, 5 Concat operations, 5 cross-stage partial fusion (Cross Stage Partial Bottleneck with 2 convolutions, C2f) modules and 2 CBS modules. The output feature maps of the first three MWTRMs and the third and fourth CSB modules in the backbone network are then input into the Neck part layer by layer for high-dimensional and low-dimensional feature concatenation to reduce feature loss during feature transmission, namely the Concat operation. The upsampling operation is called Up Sample, and the C2f module is a cross-stage partial network fusion layer composed of two convolution kernels of size 1 and two Bottleneck modules through residuals.

[0031] Step 2.3: Build the detection layer (Head). The Neck part built in step 2.2 has three output feature maps of different sizes, which correspond to the inputs of three detection heads of three different sizes. The detection head is Detect: It consists of 4 CBS modules with a convolution kernel of 3 and 2 convolutions with a convolution kernel of 1. After the original features are processed by the backbone network and the Neck part, the output feature map will be input to the detection head for prediction. The three detection heads are P2 / 4, P3 / 8, and P4 / 16, which are used to cover targets of different scales. The P2 / 4 small target detection layer is used to reduce deep semantic interference and reduce computational overhead, so that the network can further focus on the shallow feature extraction of small targets. When the detection head makes a prediction, it calculates the difference between the real target and the predicted target, and calculates the loss through the Inner-ShapeIoU loss function for back propagation, so as to correct the network parameters for the next time.

[0032] Step 2.4: Build a shallow skip residual module (SSRM). In the SAR image ship detection task, the pixel ratio of small targets is low and the feature information is scarce. The high-dimensional abstraction of the deep feature map easily leads to the loss of spatial details, especially in complex scenes such as near-shore occlusion and dense arrangement. The deep detection head with a large receptive field will introduce background interference, aggravating the missed detection and false detection of small targets. The shallower high-resolution feature map contains more spatial information and includes more small-sized position details. Therefore, the present invention first introduces a P2 small target detection layer with a feature map size of 160×160 in the Neck built in step 2.2 and the Head built in step 2.3, and directly uses high-resolution features to capture the spatial position and edge texture and other detailed feature information of the target; at the same time, the redundant P5 detection layer is removed to reduce deep semantic interference and reduce computational overhead, so that the network can further focus on the shallow feature extraction of small targets. Secondly, a skip residual connection is established between the shallower feature layers (C3 / P3, C4 / P4) of the same level between the Backbone built in step 2.1 and the Neck built in step 2.2 to compensate for the weakening of the fusion effect between low-level features and high-level semantic information caused by the change in network depth. Specifically, first connect the output of the third CBS of Backbone to the fourth Concat operation in Neck, and then connect the output of the fourth CBS of Backbone to the fifth Concat operation in Neck, so as to realize the skip connection of features across stages.

[0033] Step 2.5: After deploying steps 2.1, 2.2, 2.3, and 2.4, the processed SAR ship image JPG picture with a size of 640×640×3 is used as input. It first passes through two CBS modules in the backbone network with a convolution kernel size of 3 and a step size of 2, and then passes through a MWTRM module. At this time, the image size becomes 160×160×32. The output at this time is recorded as the first stage C2, which is the input of the third CBS module with a convolution kernel size of 3 and a step size of 2. The output of the third CBS module is then used as the input of the second MWTRM module. At this time, the output image size becomes 80×80×6 4. The output at this time is recorded as the second stage C3. After passing through the fourth CBS module with a convolution kernel of 3 and a step size of 2, the output image is used as the input of the third MWTRM module. At this time, the size of the output image becomes 40×40×128. The output at this time is recorded as the third stage C4, which is the input of the fifth CBS module with a convolution kernel of 3 and a step size of 2. The output image then passes through the fourth MWTRM module and an SPPF module in sequence. At this time, the size of the output image becomes 20×20×256. The output at this time is recorded as the fourth stage C5. The output feature map of this stage is transmitted to the Neck part. After the first Up Sample, its output is concatenated (Concat) with the output C4 of the third stage. The concatenated feature map is used as the input to the first C2f module. The output feature map is used as the fifth stage. The output of this stage is subjected to the second Up After Sample, the output is concat-operated with the output C3 of the second stage, and after splicing, it is input to the second C2f module. The output at this time is used as the sixth stage, and the third UpSample is performed. The output is concat-operated with the output C2 of the first stage as the input of the third C2f module. At this time, the output feature map size is 160×160×65, which is input to the first Detect to complete the prediction and accuracy calculation of the object, and the loss is calculated by the Inner-ShapeIoU loss function for back propagation, and the network parameters are updated. The output of the third C2f module is used as the input of the first CBS module of the Neck part, and its output result is concat-operated with the output of the third CBS module in Backbone and the output of the sixth stage, combining the shallow spatial information and the deep semantic information to reduce the loss of feature information. The splicing result is used as the input of the fourth C2f module. At this time, the output feature map size is 80×80×65, and it is input to the second Detect to complete the prediction and the second CBS module of the Neck part. After completing the prediction of the object, the loss is calculated by the Inner-ShapeIoU loss function for back propagation, and the network parameters are updated.The output of the second CBS module in the Neck part is concat-operated with the output of the fourth CBS module in the Backbone and the output of the fifth stage. The concatenated feature map is used as the input of the last C2f module. The output feature map size is 40×40×65. The feature map is input to the third Detect to complete the prediction and the Inner-ShapeIoU loss function is used to calculate the loss for back propagation, update the network parameters, and obtain the experimental indicators through multiple rounds of training.

[0034] There are several Bottleneck modules in the MWTRM module. In the MWTRM modules of the C2, C3, C4, and C5 layers, the number of Bottleneck modules is 1, 3, 3, and 1, respectively.

[0035] Step 3: Input the SAR ship images of the training set preprocessed in step 1 into the SAR image ship detection network in step 2 for training, calculate the Inner-ShapeIoU loss function and perform back propagation, update the network parameters, and obtain the optimal parameter model.

[0036] The Inner-ShapeIoU loss function combines ShapeIoU with Inner-IoU and embeds the dynamic auxiliary box mechanism into the weighted loss of the geometric features of the border. The dynamic auxiliary box mechanism is embedded in the weighted loss of the geometric features of the border. First, the auxiliary box is generated according to the following formula: , , , , , , , in, and Represents the coordinates of the center point of the real frame, and represents the width and height of the real box, , , , Respectively represent the left, right, top, and bottom boundary coordinates of the dynamic auxiliary frame of the real frame, and Represents the coordinates of the center point of the prediction box, and Represents the width and height of the prediction box, , , , Respectively represent the left, right, top, and bottom boundary coordinates of the dynamic auxiliary box of the real box, and the scale factor It is used to control the scale of the dynamic auxiliary frame to adjust the sensitivity of the regression gradient. .

[0037] Then, the weight coefficient is calculated based on the aspect ratio of the original ground truth (GT) box, the center point offset and size difference are constrained, and finally the regression loss of the box is obtained. The implementation formula is as follows: , , , , , , , in, is the scaling factor, the present invention sets , and Shape weight is calculated based on the aspect ratio of the GT box. The introduction of shape weight enhances the differential penalty for the offset in the length and width directions, which helps to improve the positioning accuracy of non-square targets. and shape similarity loss They are used to impose weighted constraints on center point offset and size difference respectively.

[0038] The experimental indicators are accuracy (Precision), recall rate (Recall), average precision at IOU 0.5 (average precision at IOU 0.5, mAP 0.5), which represents the average accuracy when the IOU detection threshold is 0.5, and average precision at IOU 0.95 (average precision at IOU 0.95, mAP 0.5:0.95), which represents the average accuracy when the IOU detection threshold ranges from 0.5 to 0.95. The formula is as follows: , , , in, Represents a true positive sample (True Positive), Indicates a false positive sample (FalsePositive), Indicates a positive sample that has not been detected (False Negative).

[0039] IOU stands for Intersection over Union, which is defined as a geometric overlap metric between the predicted bounding box and the annotated bounding box. It quantifies the detection accuracy by calculating the area ratio of the spatial intersection area and the union area of ​​the two. It is one of the commonly used performance indicators in evaluating target detection tasks. In addition, in target detection tasks, IOU is also used as the optimization target of the loss function during model training, and drives the adjustment of model parameters through the back-propagation mechanism to improve the bounding box regression accuracy. It is also used to filter redundant detection boxes in the non-maximum suppression algorithm to ensure that the final output box has the optimal spatial coverage characteristics.

[0040] Step 4: Input the preprocessed test set in step 1 into the optimal parameter model trained in step 3, and output the ship detection map of the SAR ship image.

[0041] In order to prove the effectiveness of the SAR image ship small target detection method based on WSSRNet provided by the present invention, the HRSID ship detection dataset is used to train, verify and test the model. As can be seen from Table 1, the evaluation indicators of the present invention are higher than those of the existing detection network, and the detection effect is closest to the true value. The detection results are compared as follows: Figure 4 shown.

[0042] Table 1

[0043] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the aforementioned WSSRNet-based SAR image ship small target detection method are implemented.

[0044] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the aforementioned WSSRNet-based SAR image ship small target detection method are implemented.

[0045] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0046] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0047] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0049] The above embodiments are only for illustrating the technical idea of ​​the present invention, and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. A SAR image ship small target detection method based on WSSRNet is characterized by: The steps include: Step 1, obtain a SAR ship image dataset, preprocess the dataset, and divide it into a training set, a validation set, and a test set according to a set ratio; Step 2: Build a SAR image ship detection network based on WSSRNet, including the backbone network, neck network and detection layer; A multi-scale wavelet convolution residual module is constructed in the backbone network to extract feature maps C2, C3, C4 and C5 of four different scales; The neck network includes upsampling and downsampling stages. In the upsampling stage, the four feature maps C2, C3, C4 and C5 of different scales extracted by the backbone network are upsampled, concatenated and partially fused across stages to obtain the feature map P2. In the downsampling stage, a shallow skip residual structure is used to perform skip connections and feature fusion on the same-level features extracted by the backbone network and the features fused by the neck network to obtain feature maps P3 and P4. The detection layer is used to detect ship targets of small, medium and large scales based on feature maps P2, P3 and P4; Step 3, input the training set and the validation set into the SAR image ship detection network based on WSSRNet built in step 2 for training and validation, calculate the loss function and perform back propagation, update the network parameters, and obtain the trained optimal parameter network; Step 4: Input the test set into the trained optimal parameter network and output the ship detection map of the SAR ship image.

2. The SAR image ship small target detection method based on WSSRNet according to claim 1 is characterized in that: In the SAR image ship detection network based on WSSRNet, The backbone network includes the first to fifth convolution modules, the first to fourth multi-scale wavelet convolution residual modules and the fast spatial pyramid pooling module; the SAR ship image input by the backbone network is sequentially subjected to feature extraction by the first convolution module, the second convolution module and the first multi-scale wavelet convolution residual module to obtain a feature map C2; C2 is sequentially subjected to feature extraction by the third convolution module and the second multi-scale wavelet convolution residual module to obtain a feature map C3; C3 is sequentially subjected to feature extraction by the fourth convolution module and the third multi-scale wavelet convolution residual module to obtain a feature map C4; C4 is sequentially subjected to feature extraction by the fifth convolution module, the fourth multi-scale wavelet convolution residual module and the fast spatial pyramid pooling module to obtain a feature map C5; The neck network includes the sixth to seventh convolution modules, the first to fifth cross-stage partial fusion modules, and the first to third upsampling modules; the first upsampling module is used to upsample C5, the upsampling result is channel-joined with C4, and the joint result is subjected to feature fusion by the first cross-stage partial fusion module to obtain the feature map F0; the second upsampling module is used to upsample F0, the upsampling result is channel-joined with C3, and the joint result is subjected to feature fusion by the second cross-stage partial fusion module to obtain the feature map F1; F1 is upsampled by the third upsampling module, and the upsampling result is channel-joined with C2. The joint result is subjected to feature fusion by the third cross-stage partial fusion module to obtain feature map P2; P2 is subjected to feature extraction by the sixth convolution module, and then channel-joined with F1 and the output of the third convolution module. The joint result is subjected to feature fusion by the fourth cross-stage partial fusion module to obtain feature map P3; P3 is subjected to feature extraction by the seventh convolution module, and then channel-joined with F0 and the output of the fourth convolution module. The joint result is subjected to feature fusion by the fifth cross-stage partial fusion module to obtain feature map P4; The detection layer includes the first to third detection heads. The feature map P2 output by the third cross-stage partial fusion module is used as the input of the first detection head to detect small-scale ship targets. The feature map P3 output by the fourth cross-stage partial fusion module is used as the input of the second detection head to detect medium-scale ship targets; The feature map P4 output by the fifth cross-stage partial fusion module is used as the input of the third detection head to detect large-scale ship targets.

3. The SAR image ship small target detection method based on WSSRNet according to claim 2 is characterized in that: The specific process of step 1 is as follows: The HRSID ship image dataset was obtained, including 5604 SAR images of 800×800 pixels in different sea conditions, including ocean, offshore, coastal ports, and different sea conditions. There were 16951 ship targets in total, of which small, medium, and large ship targets accounted for 54.5%, 43.5%, and 2% of the total, respectively. The original images and the corresponding label images in the HRSID ship image dataset are cropped into images of 640×640 pixels in sequence. The cropped HRSID ship image dataset is randomly divided into training set, validation set and test set in the ratio of 7:2:

1.

4. The SAR image ship small target detection method based on WSSRNet according to claim 2 is characterized in that: The first and fourth multi-scale wavelet convolution residual modules have the same structure, and both include the eighth to tenth convolution modules, the first wavelet transform convolution module, and the first to second bottleneck modules; The input of the first or fourth multi-scale wavelet convolution residual module is evenly divided into four feature subsets after passing through the eighth convolution module, and the second feature subset is sequentially passed through the first wavelet transform convolution module and the ninth convolution module to obtain the output of the ninth convolution module; the output of the ninth convolution module is added element by element to the third feature subset and used as the input of the first bottleneck module; the output of the first bottleneck module is added element by element to the fourth feature subset and used as the input of the second bottleneck module; the output of the second bottleneck module is channel-joined with the output of the first bottleneck module, the output of the ninth convolution module and the first feature subset and used as the input of the tenth convolution module; the output of the tenth convolution module is channel-joined with the input of the first or fourth multi-scale wavelet convolution residual module to obtain the output of the first or fourth multi-scale wavelet convolution residual module.

5. The SAR image ship small target detection method based on WSSRNet according to claim 4 is characterized in that: The second and third multi-scale wavelet convolution residual modules have the same structure, both including the eighth to tenth convolution modules, the first wavelet transform convolution module and the first to sixth bottleneck modules; The input of the second or third multi-scale wavelet convolution residual module is evenly divided into four feature subsets after passing through the eighth convolution module, and the second feature subset is sequentially passed through the first wavelet transform convolution module and the ninth convolution module to obtain the output of the ninth convolution module; the output of the ninth convolution module is added element by element to the third feature subset and used as the input of the first bottleneck module; the output of the first bottleneck module is sequentially passed through the third bottleneck module and the fourth bottleneck module to obtain the output of the fourth bottleneck module; the output of the fourth bottleneck module is added element by element to the fourth feature subset and used as the input of the second bottleneck module; the output of the second bottleneck module is sequentially passed through the fifth bottleneck module and the sixth bottleneck module to obtain the output of the sixth bottleneck module; the output of the sixth bottleneck module is channel-joined with the output of the fourth bottleneck module, the output of the ninth convolution module and the first feature subset to serve as the input of the tenth convolution module; the output of the tenth convolution module is channel-joined with the input of the second or third multi-scale wavelet convolution residual module to obtain the output of the second or third multi-scale wavelet convolution residual module.

6. The SAR image ship small target detection method based on WSSRNet according to claim 5 is characterized in that: The first to sixth bottleneck modules have the same structure, and all include a second wavelet transform convolution module and an eleventh convolution module. The input of each bottleneck module passes through the second wavelet transform convolution module and the eleventh convolution module in sequence to obtain the output of the eleventh convolution module. The output of the eleventh convolution module is channel-joined with the input of each bottleneck module to obtain the output of each bottleneck module. The convolution kernel sizes of the eighth to eleventh convolution modules are all 1×1, and the convolution kernel sizes of the first and second wavelet transform convolution modules are both 3×3.

7. The SAR image ship small target detection method based on WSSRNet according to claim 2 is characterized in that: The fast spatial pyramid pooling module includes twelfth to thirteenth convolution modules and first to third maximum pooling layers. The input of the fast spatial pyramid pooling module passes through the twelfth convolution module, the first maximum pooling layer, the second maximum pooling layer and the third maximum pooling layer in sequence to obtain the output of the third maximum pooling layer. The output of the third maximum pooling layer is channel-joined with the output of the first maximum pooling layer and the output of the second maximum pooling layer, and then enters the thirteenth convolution module. The output of the thirteenth convolution module is the output of the fast spatial pyramid pooling module.

8. The SAR image ship small target detection method based on WSSRNet according to claim 1 is characterized in that: In step 3, the loss function adopts the Inner-ShapeIoU loss function, combines ShapeIoU with Inner-IoU, embeds the dynamic auxiliary frame mechanism into the weighted loss of the geometric features of the frame, and generates the auxiliary frame formula as follows: , , , , , , , in, and Respectively represent the value of the center point of the real frame on the x and y axes, and Represent the width and height of the real box respectively, , , , Respectively represent the left, right, top, and bottom boundary coordinates of the dynamic auxiliary frame of the real frame, and Respectively represent the value of the center point of the prediction box on the x and y axes, and Represents the width and height of the prediction box, , , , Respectively represent the left, right, top, and bottom boundary coordinates of the dynamic auxiliary box of the prediction box, is the scale factor, which is used to control the scale of the dynamic auxiliary frame to adjust the sensitivity of the regression gradient; Represents the intersection-over-union ratio of the predicted box and the true box; The shape weight is calculated based on the aspect ratio of the real box, the center point offset and size difference are constrained, and finally the regression loss of the border is obtained. The implementation formula is as follows: , , , , , , , in, and are shape weights, is the scaling factor, represents the shape distance loss, represents the shape similarity loss, Represents the diagonal length of the minimum enclosing box covering the predicted box and the true box, represents the Inner-ShapeIoU loss, Represents shape weight based The weighted term calculated about the difference between the predicted box and the true box width, Represents shape weights A weighted term for calculating the height difference between the predicted box and the true box.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for detecting small ship targets in SAR images based on WSSRNet are implemented as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting small ship targets in SAR images based on WSSRNet are implemented as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Steel surface defect detection method and system and computer equipment

    CN116664558A

  • SAR image urban area scene classification method based on multi-scale wavelet convolution

    CN117636052A

  • KIA Net network model and image thereof, and high-precision wafer defect detection and segmentation method

    CN118762013A

  • Lightweight target detection method and system based on improved YOLOv8

    CN119323668A

  • Surface detection method and system for precise fastener

    CN119338827A

Cited By

  • Fan blade fault detection method and device, storage medium and computer equipment

    CN121010549A