Radar sea surface target detection method based on fusion of trunk shunt fidelity and adjacent layer features

By designing a lightweight network model that integrates backbone branch fidelity and neighbor layer features, the problem of false alarm and high computational complexity of SAR image ship object detection algorithm in complex environments is solved, and efficient and accurate target detection is achieved, suitable for edge-end deployment.

CN120125801APending Publication Date: 2025-06-10NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118001.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing SAR image ship object detection algorithm has false alarm problems in complex environments, and the calculation complexity is high, making it difficult to meet the real-time detection requirements.

Method used

A lightweight radar ship image target detection network model that integrates the backbone branch fidelity and adjacent layer features is designed. Through feature extraction backbone module, adjacent feature fusion module, channel shuffling and channel attention module, the amount of model parameters and calculation amount is reduced.

Benefits of technology

It effectively reduces the computational complexity and memory overhead, improves detection accuracy, is suitable for edge-end deployment, and meets real-time detection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125801A_ABST
    Figure CN120125801A_ABST
Patent Text Reader

Abstract

The invention discloses a main branch fidelity and adjacent layer feature fusion radar sea surface target detection method, which comprises the steps of obtaining a radar image to be identified, and inputting the radar image into a trained main branch fidelity and adjacent layer feature fusion lightweight radar ship image target detection network model, outputting an identification result of the ship target; wherein the target detection network model comprises a feature extraction trunk module, a proximity feature fusion module, a channel shuffling and channel attention module and a detection head; the feature extraction backbone module comprises a CBS component and four layers of feature extraction backbone units which are connected in sequence; the adjacent feature fusion module comprises a plurality of shunt connection units and three fusion feature outputs; and the channel shuffling and channel attention module comprises a series channel shuffling unit, a depth separable convolutional layer and a channel attention mechanism, and is used for processing the multi-path fusion features respectively, sending the processed multi-path fusion features to a detection head, and outputting a ship target detection result by using the detection head.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radar signal and information processing, and can be applied to the detection task of ship targets in radar images. Specifically, it relates to a method for detecting radar sea targets with backbone branch fidelity and adjacent layer feature fusion. Background Technique

[0002] Synthetic Aperture Radar (SAR) has become an important part of radar research due to its characteristics such as all-weather, all-time, and high resolution. Target detection is an important basis for interpreting and processing radar images. At the present stage, the target detection of SAR images is mainly divided into traditional methods and deep learning-based target detection methods. In the detection of SAR ship targets, the biggest feature of traditional algorithms is that they need to manually design and extract target features, such as: gray scale, contrast, shape, texture, etc. Traditional SAR ship target detection algorithms generally use the artificially designed feature differences between targets for target detection. This algorithm can often achieve satisfactory results in the case of simple background clutter and small scene interference, but in some complex environments, there will be more false alarms, thus greatly reducing the accuracy of the algorithm. Deep learning-based target detection methods do not require pre-artificial design of target features. The neural network learns and self-updates the data through the internal feature extraction structure, so as to achieve relatively accurate target detection results. With the development of deep learning technology and the update of the computing power of computer devices, it has become possible to deploy neural networks in the field of SAR image target detection.

[0003] Among traditional SAR image target detection algorithms, the most representative one is the Constant False Alarm (CFAR) detection algorithm. In the CFAR algorithm, first, the background statistical distribution in the SAR image is fitted, and then the threshold is calculated to complete target detection. However, in practical applications, the background statistical distribution is often affected by sea waves, sea clutter, etc., making it difficult to fit a suitable statistical distribution to describe the background. Traditional SAR image target detection algorithms represented by the CFAR algorithm are pixel-level detection algorithms. When performing image processing, especially for high-resolution images, the detection algorithm will result in a large amount of computation, and the detection speed will be greatly reduced, making it difficult to meet the real-time requirements.

[0004] The operation process of the neural network can be briefly summarized as follows. First, the input image undergoes feature extraction through the convolutional layer, and the output of the convolutional layer is activated using an activation function to add non-linear factors so that it can fit complex functions. Then, through layer structures such as pooling, normalization, and fully connected layers, the classification or regression results are output. The entire process is uniformly backpropagated according to the loss function generated during the calculation process to adaptively optimize the learnable parameters during the operation process, ultimately achieving the convergence of the network and ensuring a high accuracy rate. Thanks to its powerful fitting ability and the ability to extract abstract features from large-scale data, the object detection algorithm based on deep learning has achieved success. Additionally, since there is no need to artificially design feature differences in the neural network, the generalization ability of the network is also enhanced. Networks with these advantages can achieve higher accuracy in the SAR image ship target detection task.

[0005] In the current object detection field, the mainstream high-accuracy neural networks are mainly designed for optical images, and these networks often perform poorly when dealing with SAR images. There are significant differences in characteristics between SAR images and optical images. Due to the particularity of the imaging mechanism, SAR images usually contain complex background clutter and coherent noise, which greatly increase the difficulty of object detection.

[0006] In optical image detection, deep networks can extract more detailed feature information. Combining various complex feature extraction and feature fusion methods can significantly improve the detection accuracy. However, these methods require a relatively large network model and computational overhead, thereby increasing the cost of deployment on airborne, spaceborne, and mobile platforms and reducing the inference speed of the model. Therefore, while not affecting the computational performance, minimizing the model parameter size as much as possible and researching lightweight neural networks have become key issues in the field of SAR image object detection. Summary of the Invention

[0007] The purpose of the present invention is to provide a radar sea target detection method with backbone branch fidelity and adjacent layer feature fusion. By designing a more efficient network structure and algorithm, without sacrificing performance, the number of model parameters and computational amount are significantly reduced, thereby achieving convenient deployment on devices with limited computing resources.

[0008] To achieve the above task, the present invention adopts the following technical solutions:

[0009] A radar sea target detection method with backbone branch fidelity and adjacent layer feature fusion, including:

[0010] Obtain the radar image to be recognized, input the radar image into the trained lightweight radar ship image object detection network model with backbone branch fidelity and adjacent layer feature fusion, and output the recognition result of the ship target through the object detection network model; where:

[0011] The described object detection network model includes a feature extraction backbone module, a neighboring feature fusion module, a channel shuffle and channel attention module, and a detection head;

[0012] The feature extraction backbone module includes a CBS component and four layers of feature extraction backbone units connected in sequence;

[0013] In each layer of the feature extraction backbone unit, after dividing the input feature map equally, one part uses a deformable convolutional layer and a depthwise separable convolutional layer for feature enhancement processing and dimension adjustment, and the other part uses a convolutional layer for feature extraction. The processing results of the two parts are concatenated in the channel dimension and then activated using a squeeze-and-excitation attention mechanism. After the processed result is fused with the input feature map through a long skip connection, the output feature of this layer of the feature extraction backbone unit is obtained;

[0014] The neighboring feature fusion module includes multiple split connection units and three fused feature output units; the output feature of each layer of the feature extraction backbone unit is only fused with the output feature of its adjacent layer of the feature extraction backbone unit through the split connection unit in the fused feature output unit for feature fusion and channel dimension adjustment to obtain multiple fused features;

[0015] The channel shuffle and channel attention module is used in the multiple fused feature outputs of the neighboring feature fusion module; each fused feature is processed respectively through a serial channel shuffle unit, a depthwise separable convolutional layer, and a channel attention mechanism, and then sent to the detection head respectively, and the detection head outputs the ship target detection result.

[0016] Furthermore, the split connection unit first divides the input feature into two paths. One path undergoes size and dimension transformation through a depthwise separable convolutional layer, and the other path uses a convolutional layer for feature extraction. Finally, the processing results of the two paths are concatenated in the channel dimension to obtain the output feature.

[0017] Furthermore, the CBS component is a common convolutional layer and a batch normalization layer connected in series, and is set with a SiLU activation function.

[0018] Furthermore, in the deformable convolutional layer, depthwise separable convolutional layer, and convolutional layer of the feature extraction backbone unit, batch normalization processing and SiLU activation function are both used.

[0019] Furthermore, the output feature of each layer of the feature extraction backbone unit is only fused with the output feature of its adjacent layer of the feature extraction backbone unit through the split connection unit in the fused feature output unit for feature fusion and channel dimension adjustment to obtain multiple fused features. Specifically:

[0020] The output features of the second-layer feature extraction backbone unit, the output features of the first-layer feature extraction backbone unit processed by the shunt connection unit, and the output features of the third-layer feature extraction backbone unit processed by the shunt connection unit and upsampled are jointly subjected to feature fusion in the fusion feature output unit to obtain the first path of fused features;

[0021] The output features of the third-layer feature extraction backbone unit, the output features of the second-layer feature extraction backbone unit processed by the shunt connection unit, and the output features of the fourth-layer feature extraction backbone unit processed by the shunt connection unit and upsampled are jointly subjected to feature fusion in the fusion feature output unit to obtain the second path of fused features;

[0022] The output features of the fourth-layer feature extraction backbone unit and the output features of the third-layer feature extraction backbone unit processed by the shunt connection unit are jointly subjected to feature fusion in the fusion feature output unit to obtain the third path of fused features.

[0023] Further, the training process of the object detection network model is as follows:

[0024] Collect SAR images containing ship targets and corresponding annotation files, and organize them into a dataset; the dataset is converted into the YOLO format and divided into a training dataset and a validation dataset;

[0025] Input the training dataset into the object detection network model for training to obtain the trained object detection network model; input the validation dataset into the trained object detection network model for validation to obtain the validation result, use performance evaluation indicators to evaluate and analyze the validation result, determine whether the model converges and whether the performance reaches the optimum, and adjust the network hyperparameters according to the validation result; if the accuracy reaches the optimum and the model has converged, stop training to obtain the trained object detection network model.

[0026] A terminal device includes a processor, a memory, and a computer program stored in the memory; when the processor executes the computer program, it implements the radar sea target detection method with backbone shunt fidelity and adjacent layer feature fusion.

[0027] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, it implements the radar sea target detection method with backbone shunt fidelity and adjacent layer feature fusion.

[0028] Compared with the prior art, the present invention has the following technical features:

[0029] In response to the need for flexible deployment at the edge, this invention constructs a feature extraction backbone network and a feature layer connection module with the idea of shunt fidelity. The introduction of shunt processing and depthwise separable convolution effectively reduces the computational complexity of the system. In response to the need for high-precision object detection, this invention solves the problems of ship target contour, deformation, and multi-scale irregularity by introducing deformable convolution into the backbone structure, and performs feature extraction more precisely and effectively. In the feature fusion structure, to avoid the problem of loss of inter-layer connection between different feature layers caused by non-linear factors, the method of adjacent feature layer fusion is used to better ensure the detection accuracy of the network. At the same time, channel shuffle and attention mechanism are appropriately used to make up for the loss of interactive information between the channels of the feature map caused by depth convolution. This invention constructs an end-to-end ship target detection method for SAR images based on the single-stage anchor-free network YOLOv8. In actual measurement, the method of this invention effectively reduces the amount of computation and the memory overhead of each training, and better ensures the detection accuracy of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is the implementation flowchart of this invention;

[0031] Figure 2 It is the overall model structure of the network proposed by this invention;

[0032] Figure 3 It is the shunt fidelity feature extraction backbone structure and shunt connection module described by this invention;

[0033] Figure 4 It is the schematic diagram of the adjacent layer feature fusion structure described by this invention;

[0034] Figure 5 It is the performance comparison of the implementation effect of this invention and the detection effect of the existing network. DETAILED DESCRIPTION OF THE INVENTION

[0035] In order to solve the problem of convenient deployment of the ship target detection network for SAR images on airborne, spaceborne, and mobile platforms, and at the same time deal with the problems of multi-scale and irregular targets in SAR image ship targets, the purpose of the present invention is to reduce the network operation parameters, streamline the network structure, and at the same time use hybrid dual-modal features to perform more accurate feature extraction on the target, so as to achieve the balance and maintenance of the network with better detection accuracy while reducing the network parameters. The present invention is based on the single-stage, anchor-free network YOLOv8. For the task problems, the backbone network is redesigned to reduce the computational complexity, and a deformable convolution and adjacent feature layer feature fusion structure is used to achieve the enhancement of feature fidelity. It can better ensure the detection accuracy of the overall network while reducing the parameters. Experiments prove that the radar sea target detection method based on backbone branch fidelity and adjacent layer feature fusion of the present invention can effectively reduce the computational complexity and better ensure the accuracy of target detection.

[0036] See the appendix Figure 1 , a radar sea target detection method based on backbone branch fidelity and adjacent layer feature fusion provided by the present invention, includes the following steps:

[0037] Obtain the radar image to be recognized, input the radar image into the lightweight radar ship image target detection network model with backbone branch fidelity and adjacent layer feature fusion that has been trained, and output the recognition result of the ship target through the target detection network model; where:

[0038] The target detection network model includes a feature extraction backbone module, an adjacent feature fusion module, a channel shuffle and channel attention module, and a detection head;

[0039] The feature extraction backbone module includes a CBS component and four layers of feature extraction backbone units connected in sequence;

[0040] In each layer of the feature extraction backbone unit, after the input feature map is equally divided, one part uses a deformable convolutional layer and a depthwise separable convolutional layer for feature enhancement processing and dimension adjustment, and the other part uses a convolutional layer for feature extraction. The processing results of the two parts are feature concatenated in the channel dimension and then activated using a squeeze-and-excitation attention mechanism (SE). After the processed result is fused with the input feature map through a long skip connection, the output feature of this layer of the feature extraction backbone unit is obtained;

[0041] The adjacent feature fusion module includes multiple split connection units and three fusion feature output units; the output feature of each layer of the feature extraction backbone unit is only fused with the output feature of its adjacent layer of the feature extraction backbone unit in the fusion feature output unit through the split connection unit for feature fusion and channel dimension adjustment to obtain multiple paths of fused features;

[0042] The channel shuffle and channel attention module are used in the multi-path fusion feature output of the adjacent feature fusion module; each path of the fusion feature is processed by a series-connected channel shuffle unit, a depthwise separable convolutional layer, and a channel attention mechanism respectively, and then sent to the detection head, and the detection results of ship targets are output by the detection head.

[0043] In this solution, the detection head is an existing network module in YOLOv8.

[0044] 1. Feature extraction backbone module

[0045] In the feature extraction backbone module, there are CBS components and four layers of feature extraction backbone units in sequence; the input radar image is processed by the CBS components, and the extracted feature maps are input into the four layers of feature extraction backbone units in sequence; the CBS components are a common convolution (Conv, C) with a kernel size of 3*3, batch normalization (BN, B), and SiLU (S) activation function in series; for the convenience of explanation, the output features of the first to fourth layer feature extraction backbone units are denoted as H 1 to H 4 .

[0046] The structure of the feature extraction backbone unit is as Figure 3 shown; for the input feature map its specific processing process is:

[0047] First, split the input feature map in the channel dimension:

[0048] dim(X input ) = dim(X D ) + dim(X re )

[0049] X = split(X D , X re )

[0050] C = C D + C re

[0051] Among them, split(·) represents data division in the channel dimension, C represents the number of channels of the input feature map, H and W represent the height and width of the feature map respectively, C D , C re respectively represent the number of channels of the two parts of the feature map obtained after splitting of.

[0052] The feature extraction backbone unit has two branches:

[0053] In the first branch, for the first part of the split Use deformable convolution for feature extraction DConv(·) to obtain more features related to the true contour and deformation of the target, improve the detection accuracy of the network, and concatenate depthwise separable convolution DWConv(·) for dimension adjustment to obtain the output features of the first branch

[0054]

[0055] Among them, the SiLU(·) function is used for activation during both the DConv(·) and DWConv(·) processes to prevent problems such as gradient explosion and gradient vanishing. The output data is processed using batch normalization and then the activation function is applied

[0056] In the second branch, for the second part of the split Use a common convolutional layer Conv(·) for feature extraction to obtain output features Among them, Conv(·) also uses the SiLU(·) function for activation, and the data is processed using batch normalization before activation

[0057] After obtaining the processing results of the two branches Concatenate them in the channel dimension. At this time, it is necessary to ensure that the concatenation order is consistent with the split order, and sequentially use the spatial attention mechanism SE to obtain the fidelity features of the hybrid bimodal

[0058] X dual =f SE (Concat(X D_temp ,X re_temp ,dim=c))

[0059] dim(X dual )=dim(X D_temp )+dim(X re_temp )=dim(X)

[0060] Among them, f SE (·) represents the assignment process of spatial attention, and Concat(·,dim=c) represents feature concatenation in the channel dimension

[0061] After obtaining the features Through the long skip branch, directly fuse the input feature map with After that, obtain the output features of this module through downsampling operations

[0062]

[0063] dim(X out) = dim(X dual )

[0064] Among them, f 2*sample (·) represents the 2x downsampling process, represents element-wise direct addition, respectively represent the values of each pixel in X out , X, X dual in X.

[0065] See Appendix Figure 2 , the output feature H 1 of the first-layer feature extraction backbone unit is X out ; after taking H 1 as the input feature map of the second-layer feature extraction backbone unit and processing it, its output feature is H 2 ; after processing by the third-layer feature extraction backbone unit, the output feature of H 2 is H 3 , and after processing by the fourth-layer feature extraction backbone unit, the output feature of H 3 is H 4 .

[0066] 2. Adjacent Feature Fusion Module

[0067] The adjacent feature fusion module mainly includes a split connection unit; the split connection unit first splits the input feature H into two paths H DW and H C , one path H DW undergoes size and dimension transformation through a depthwise separable convolutional layer, and the other path H C uses a convolutional layer for feature extraction, and finally the processing results of the two paths are concatenated in the channel dimension to obtain the output feature H out , thereby achieving size and dimension adjustment while appropriately controlling the calculation parameters.

[0068] H = split(H DW , H C )

[0069] H out = Concat(DWConv(H DW ), Conv(H C ), dim = c)

[0070] In this solution, the output feature of each layer of the feature extraction backbone unit is only fused with the output feature of the adjacent layer of the feature extraction backbone unit through the split connection unit for feature fusion and channel dimension adjustment to obtain multiple paths of fused features, specifically:

[0071] The output feature H 2, and H 1 The output feature H of the first-layer feature extraction backbone unit after being processed by the shunt connection unit out1 , H 3 The output feature H of the third-layer feature extraction backbone unit after being processed by the shunt connection unit and upsampling out3 Perform feature fusion together in the fusion feature output unit to obtain the first-way fusion feature;

[0072] The output feature H of the third-layer feature extraction backbone unit 3 , and H 2 The output feature H of the second-layer feature extraction backbone unit after being processed by the shunt connection unit out2 , H 4 The output feature H of the fourth-layer feature extraction backbone unit after being processed by the shunt connection unit and upsampling out4 Perform feature fusion together in the fusion feature output unit to obtain the second-way fusion feature;

[0073] The output feature H of the fourth-layer feature extraction backbone unit 4 , and H 3 The output feature H of the third-layer feature extraction backbone unit after being processed by the shunt connection unit out3 Perform feature fusion together in the fusion feature output unit to obtain the third-way fusion feature.

[0074] This solution adopts the above structural design, rather than the traditional pyramid hierarchical structure, and cascades the channel shuffle and channel attention mechanism for the three-way fusion features to achieve refined feature extraction and enhanced expression for network detection.

[0075] 4. Channel Shuffle and Channel Attention Module

[0076] The channel shuffle and channel attention module has three paths, corresponding to one-way fusion features respectively; each path of the channel shuffle and channel attention module sequentially includes a series channel shuffle unit shuffle_neck, a depthwise separable convolutional layer DWConv, and a channel attention mechanism CA.

[0077] Taking the output feature H of the third-layer feature extraction backbone unit 3 as an example, its adjacent layer features are H of the third layer 2 , H of the fourth layer 4 , then its corresponding output feature H 3_out is:

[0078] H 3_temp =Concat(H 2_down ,H 3 ,H 4_up ,dim=c)

[0079] H 2_down = φ dc (H 2 )

[0080] H 4_up = f up (φ dc (H 4 ))

[0081] H 3_out = f CA (σ shuffle (H 3_temp ))

[0082] where φ dc (·) is a shunt connection unit for flexibly connecting each feature layer to ensure the unity of size and channel dimension, f up (·) represents using 2x upsampling, f CA (·) represents channel attention, and σ shuffle (·) represents the channel shuffle process.

[0083] The output feature H 3_out is fed into the detection head to achieve object detection.

[0084] 5. Model Training

[0085] Collect SAR images containing ship targets and corresponding annotation files, and organize them into a dataset; convert the dataset into the YOLO format and divide it into a training dataset, a validation dataset, and a test dataset according to the ratio of 7:2:1;

[0086] Input the training dataset into the object detection network model for training to obtain the trained object detection network model; input the validation dataset into the trained object detection network model for validation to obtain the validation results, evaluate and analyze the validation results using performance evaluation metrics to determine whether the model converges and whether the performance reaches the optimum; if the accuracy reaches the optimum and the model has converged, stop training to obtain the trained object detection network model; if the accuracy of the validation results does not reach the optimum and the model training has not converged, adjust the model hyperparameters during training and continue training; input the test dataset to evaluate the performance of the trained object detection network model.

[0087] Example:

[0088] In this embodiment, a SAR image ship target dataset for object detection is collected. For convenience, the publicly available SSDD dataset can be directly used. The environment used in this invention is NVIDIA GeForce RTX 4060 8GB. The dataset is converted into the YOLO format, and the training dataset, validation dataset, and test dataset are divided according to 7:2:1. A target detection network model as shown in Figure 2 is built.

[0089] The input radar image size of the feature extraction backbone module is 640*640*3. The input radar image first undergoes a convolution operation with a convolution kernel size of 3, a stride of 2, and an output channel number of 64 in the CBS component, and then is batch-normalized and activated by SiLU(·) to achieve preliminary feature extraction. At this time, the output feature map size is 320*320*64.

[0090] In the feature extraction backbone unit, split(·) divides the input features into two parts in a 1:1 ratio and enters the two branches for feature extraction respectively. The deformable convolution DConv(·), depthwise separable convolution DWConv(·), and ordinary convolution Conv(·) used in the two branches all have a convolution kernel size of 3*3. The input and output channel numbers are both 1 / 2 of the original radar image.

[0091] If all ordinary convolutions Conv(·) are used in the backbone structure, its computational complexity FLOPs Conv can be expressed as:

[0092] FLOPs Conv = H out · W out · C out · C in · K · K

[0093] If all deformable convolutions DConv(·) are used in the backbone structure, the offset FLOPs offset and the interpolation computational complexity FLOPs interp can be roughly expressed as:

[0094] FLOPs offset = H out · W out · C out · C in · K · K

[0095] FLOPs interp = H out · W out · C out · C in · 4

[0096] The FLOPs (floating point operations) generated by deformable convolution can be obtained DConv It can be expressed as:

[0097] FLOPs DConv = FLOPs Conv + FLOPs offset + FLOPs interp

[0098] FLOPs DConv = (2·K·K + 4)H out ·W out ·C out ·C in

[0099] In the above formula, H out , W out are the height and width of the output feature map, C in , C out represent the number of input channels and output channels of the feature map, and K represents the size of the convolution kernel.

[0100] In the method of the present invention, the FLOPs generated by the split-fidelity feature extraction backbone through split convolution operations are:

[0101]

[0102] In the method of the present invention, the input and output channels of deformable convolution and ordinary convolution are the same, both being 1 / 2 of the original input data, and the convolution kernel size is 3*3. The FLOPs can be obtained as:

[0103] FLOPs offset = 9·H out ·W out ·C out ·C in

[0104] FLOPs DConv = 22·H out ·W out ·C out ·C in

[0105]

[0106] Obviously, the computational parameters of the method of the present invention are significantly lower than those of the feature extraction method that fully uses a single convolution.

[0107] In the feature extraction backbone module, there are a total of 4 feature extraction backbone units from top to bottom. In the present invention, the 2nd, 3rd, and 4th layers are used as the effective feature inputs of the feature fusion structure for feature fusion; the output features of the first to fourth layer feature extraction backbone units are denoted as H1 to H 4 are 160*160*128, 80*80*256, 40*40*512, and 20*20*1024 respectively.

[0108] In the adjacent feature fusion module, the first-layer feature extraction backbone unit outputs the feature H 1 Although it is not used as a valid feature, in order to make full use of the feature information and maintain the small target information, H is still used in the fusion process of H 2 . 1 .

[0109] The first to the third fused features are: H Mtemp , M = 2, 3, 4;

[0110] H 2temp = Concat(H 1_down , H 2 , H 3_up , dim = c)

[0111] H 3temp = Concat(H 2_down , H 3 , H 4_up , dim = c)

[0112] H 4temp = Concat(H 3_down , H 4 , dim = c)

[0113] In the formula, H m_down , m = 1, 2, 3, 4 are the downsampled output and the 2x upsampled output of the corresponding layer output feature H m .

[0114] After obtaining the fused features of each path, channel shuffle is used to compensate for the loss of cross-channel interaction information, controlling the output channel number to 256, and finally the final feature output is generated through the channel attention mechanism:

[0115]

[0116] H final_M = f CA (σ shuffle (H M_temp ))

[0117] Among them, φ dc (·) is the split connection unit, which is used to flexibly adjust the channel dimension of the input feature and the downsampling process, f CA (·) represents the channel attention mechanism, σ shuffle(·) represents channel shuffle, and the shuffle coefficient is 2 in the present invention. N , N = 1, 2, 3......

[0118] The number of input channels of the detection head part is 256. The output features of the channel shuffle and the channel attention module are input for the detection task.

[0119] The partitioned SSDD dataset is used for network training, and the network effect is evaluated. The evaluation indicators are mainly: Precision, Recall, F1 score, and mean average precision mAP. 50 . If the model training result meets the requirements, stop training to obtain the corresponding network model. If not, adjust the hyperparameters and repeat the training process. In the experimental results shown in the present invention, the batch size is 16, the input feature grouping ratio is 1:1, the channel shuffle coefficient is 64, and the initial learning rate is 0.001.

[0120] The experimental results of the present invention on the SSDD test dataset are recorded in the following table. Compared with YOLOv8, the prediction accuracy of the method of the present invention is close, the comprehensive F1 score decreases by 1% compared with the original network, but the computational cost of the network decreases by more than 35%:

[0121] Table 1 Comparison of network effects on the SSDD test dataset

[0122]

[0123] The above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for detecting sea surface targets by using a main branch fidelity and adjacent layer feature fusion radar, characterized in that: include: The radar image to be identified is obtained, and the radar image is input into the trained lightweight radar ship image target detection network model with trunk branch fidelity and adjacent layer feature fusion, and the recognition result of the ship target is output through the target detection network model; wherein: The target detection network model includes a feature extraction backbone module, an adjacent feature fusion module, a channel shuffling and channel attention module, and a detection head; The feature extraction backbone module includes CBS components and four-layer feature extraction backbone units connected in sequence; In each layer of feature extraction backbone unit, the input feature map is divided into equal parts, one part is subjected to feature enhancement processing and dimension adjustment using a deformable convolution layer and a depthwise separable convolution layer, and the other part is subjected to feature extraction using a convolution layer. The processing results of the two parts are concatenated in the channel dimension and then activated using a compression-excitation attention mechanism. The processed results are fused with the input feature map via a long skip connection to obtain the output features of the feature extraction backbone unit of this layer. The adjacent feature fusion module includes multiple branch connection units and three fusion feature output units; the output features of each layer of feature extraction trunk unit are only fused and channel dimension adjusted in the fusion feature output unit through the branch connection unit and the output features of the feature extraction trunk unit of the adjacent layer, so as to obtain multi-channel fusion features; The channel shuffling and channel attention modules are used in the multi-channel fusion feature outputs of the adjacent feature fusion module; each fusion feature is processed by the series channel shuffling unit, the depthwise separable convolutional layer and the channel attention mechanism, and then sent to the detection head, which outputs the ship target detection results.

2. The main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to claim 1 is characterized in that: The branch connection unit first divides the input features into two paths, one path is transformed in size and dimension through a depth-separable convolutional layer, and the other path uses a convolutional layer to extract features. Finally, the processing results of the two paths are spliced ​​in the channel dimension to obtain the output features.

3. The main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to claim 1 is characterized in that: The CBS component is a convolutional layer and a batch normalization layer connected in series, and is set with a SiLU activation function.

4. The main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to claim 1 is characterized in that: Batch normalization and SiLU activation function are used in the deformable convolution layer, depth-wise separable convolution layer, and convolution layer of the feature extraction backbone unit.

5. The main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to claim 1 is characterized in that: The output features of each layer of feature extraction trunk units are only connected by a branch connection unit and the output features of the feature extraction trunk units of the adjacent layers are subjected to feature fusion and channel dimension adjustment in the fusion feature output unit to obtain multi-channel fusion features, specifically: The output features of the second-layer feature extraction trunk unit, the output features of the first-layer feature extraction trunk unit processed by the branch connection unit, and the output features of the third-layer feature extraction trunk unit processed by the branch connection unit and up-sampled are fused together in the fusion feature output unit to obtain the first fusion feature; The output features of the third-layer feature extraction backbone unit, the output features of the second-layer feature extraction backbone unit processed by the branch connection unit, and the output features of the fourth-layer feature extraction backbone unit processed by the branch connection unit and up-sampled are fused together in the fusion feature output unit to obtain the second fusion feature; The output features of the fourth-layer feature extraction backbone unit are fused with the output features of the third-layer feature extraction backbone unit after being processed by the branch connection unit in the fusion feature output unit to obtain the third fusion feature.

6. The main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to claim 1 is characterized in that: The training process of the target detection network model is: Collect SAR images containing ship targets and corresponding annotation files and organize them into a data set; convert the data set into YOLO format and divide it into training data set and verification data set; Input the training data set into the target detection network model for training to obtain the trained target detection network model; Input the validation data set into the trained target detection network model for validation, obtain validation results, use performance evaluation indicators to evaluate and analyze the validation results, determine whether the model has converged and whether the performance has reached the optimal level, and adjust the network hyperparameters according to the validation results; If the accuracy reaches the optimal level and the model has converged, the training is stopped to obtain a trained target detection network model.

7. A terminal device, comprising a processor, a memory and a computer program stored in the memory; characterized in that: When the processor executes the computer program, it implements the main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to any one of claims 1-6.

8. A computer-readable storage medium, wherein a computer program is stored in the medium; characterized in that: When the computer program is executed by a processor, the main branch fidelity and adjacent layer feature fusion radar sea surface target detection method according to any one of claims 1 to 6 is implemented.