Ship target detection method suitable for medium-resolution optical remote sensing image
By building the YOLO-Medium Resolutionship network, combining innovative activation functions and convolution kernel modules, the problem of ship object detection in medium-resolution remote sensing images is solved, and high-precision multi-classification detection effect is achieved.
Patent Information
- Application Number
- CN202510404379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, ship object detection has problems such as large in-class differences, high inter-class similarity, difficulty in feature extraction and low detection accuracy in small-scale objects in medium-resolution remote sensing images, and there are few targeted solutions for multi-classification tasks.
The YOLO-MediumResolutionship network is built, and iterative optimization training is carried out through the combination of backbone feature extraction network, feature fusion network and detection output decoupling head, combined with the Mish activation function, SPD-Conv module and PKI multi-scale convolution kernel module, and iterative optimization training is carried out to improve feature extraction and detection accuracy.
While keeping the model lightweight, the detection accuracy of ship targets in medium-resolution remote sensing images is significantly improved, achieving efficient multi-classification detection.
Smart Images

Figure CN120339833A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applied to the technical field of remote sensing image target detection, and particularly relates to a ship target detection method applicable to medium-resolution optical remote sensing images. Background Art
[0002] With the rapid development of remote sensing technology and computer technology, rich data sources and hardware support have been provided for target detection and segmentation in the field of marine security. As the main carrier of maritime transportation and an important military target, large-scale dynamic monitoring of ships is of great significance for military and civilian use.
[0003] The development of deep learning technology and earth observation technology has brought about a major revolution in remote sensing image target detection and recognition. However, compared with other target detection tasks in remote sensing images such as vehicles, buildings, and water areas, factors such as the object shape ratio and complex background of ship targets make this task still challenging. Although deep learning detection technology has achieved many successes in the field of ship target detection in remote sensing images, there are problems such as large intra-class differences, high inter-class similarities, difficult feature extraction, and small target problems in scene classification tasks, and the detection accuracy urgently needs to be improved. Moreover, most of the existing mainstream methods are developed for single-class ship detection in high-resolution remote sensing images, and there are few methods for multi-class ship target detection tasks in medium-resolution remote sensing images. Therefore, it is particularly crucial to develop a high-efficiency method applicable to multi-class ship target detection in medium-resolution remote sensing images. Summary of the Invention
[0004] Aiming at the problems of large intra-class differences, high inter-class similarities, difficult feature extraction, and small target problems in the multi-class ship target detection task of medium-resolution remote sensing images, the present invention proposes a ship target detection method applicable to medium-resolution optical remote sensing satellites.
[0005] The technical solution adopted by the present invention is as follows: The present invention includes the following steps:
[0006] Step S1. Make a ship target detection dataset for medium-resolution remote sensing images;
[0007] Step S2. Construct a YOLO-MediumResolutionship network, abbreviated as YOLO-MRship network;
[0008] Step S3. Use the ship target detection dataset for medium-resolution remote sensing images made in Step S1 to train the YOLO-MRship network constructed in Step S2, and perform forward and backward propagation based on the loss function for iterative optimization training to obtain the optimal YOLO-MRship network weights;
[0009] Step S4. Input the remotely sensed image to be detected and the optimal model YOLO-MRship network weights obtained by iterative training in Step S3 into the YOLO-MRship network constructed in Step S2 for inference, detect the ship targets in the remotely sensed image, and save the detection results as vector files.
[0010] Further, Step S1 includes the following steps:
[0011] S11. Complete preprocessing operations including band selection, radiometric correction, geometric correction, orthorectification, and color enhancement on the medium-resolution remotely sensed images used to construct the dataset, and output true-color images;
[0012] S12. Establish ship target interpretation criteria, including ship target classification criteria, ship target annotation box forms, and quality inspection processes;
[0013] S13. According to the interpretation criteria, manually visually interpret the preprocessed true-color images to draw ship target samples and obtain sample vectors;
[0014] S14. Set parameters such as cropping size, cropping overlap, target box truncation threshold, dataset training-validation-test set ratio, and output dataset organization paradigm, and synchronously crop the remotely sensed images and ship target vectors to output the original ship target detection dataset;
[0015] S15. Enhance the original ship target detection dataset using data augmentation methods such as noise filling, contrast adjustment, and angle rotation to expand the number of dataset samples.
[0016] Further, the YOLO-MRship network model in Step S2 consists of three parts: a backbone feature extraction network, a feature fusion network, and a detection output decoupling head;
[0017] The backbone feature extraction network is sequentially connected by a first CBM module, a second CBM module, a first C2F module, a third CBM module, a first SimAM module, a second C2F module, a fourth CBM module, a third C2F module, a fifth CBM module, a fourth C2F module, a first SPPF module, and a first PKI multi-scale convolution kernel module;
[0018] The feature fusion network includes feature enhancement and extraction sub-modules at three scales: high, medium, and low. The feature fusion network includes a first Upsample module, a first Concat module, a fifth C2F module, a second Upsample module, a second Concat module, a sixth C2F module, a sixth CBM module, a first SPD-Conv module, a second Concat module, a seventh C2F module, a seventh CBM module, a third Concat module, a second SPD-Conv module, and an eighth C2F module connected in sequence;
[0019] The detection output decoupling head includes detection decoupling heads at three scales. Each decoupling head is composed of an eighth CBM module, a first Conv2d module, Bbox.Loss, a ninth CBM module, a second Conv2d module, and Cls.Loss. The sixth CBM module is connected to the first decoupling head and is respectively connected to two branches in parallel. The first branch is connected in sequence to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the first decoupling head. The second branch is connected in sequence to the ninth CBM module, the second Conv2d module, and the Conv2d module of the first decoupling head; The seventh CBM module is connected to the second decoupling head and is respectively connected to two branches in parallel. The first branch is connected in sequence to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the second decoupling head. The second branch is connected in sequence to the ninth CBM module, the second Conv2d module, and the Conv2d module of the second decoupling head; The eighth C2F module is connected to the third decoupling head and is respectively connected to two branches in parallel. The first branch is connected in sequence to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the third decoupling head. The second branch is connected in sequence to the ninth CBM module, the second Conv2d module, and the Conv2d module of the third decoupling head.
[0020] Further, the first PKI multi-scale convolution kernel module is respectively connected to the first Upsample module and the third Concat module in parallel; The second C2F module is also connected to the first Concat module in parallel; The first C2F module is also connected to the second Upsample module in parallel.
[0021] Further, the first CBM module to the ninth CBM module are all composed of a first Conv2d unit, BatchNorm2d, and a Mish activation function connected in sequence.
[0022] Further, the first SPPF module is composed of five components: initial_Conv, MaxPool1, MaxPool2, MaxPool3, and Final_Conv.
[0023] Furthermore, the loss function consists of a classification loss and a regression loss, corresponding to Cls.Loss and Bbox.Loss respectively.
[0024] Furthermore, the first PKI multi-scale convolution kernel module consists of a PKI module and a CAA module.
[0025] The present invention uses remotely sensed images collected by satellites as the data source, and innovatively introduces the Mish activation function, the first SPD-Conv module, the second SPD-Conv module, and the first PKI multi-scale convolution kernel module to construct the YOLO-MRship model, which greatly improves the detection accuracy of medium-resolution optical remotely sensed ships while keeping the model lightweight. Brief Description of the Drawings
[0026] Figure 1 Flowchart of the present invention;
[0027] Figure 2 Schematic diagram of the YOLO-MRship network structure constructed by the present invention;
[0028] Figure 3 Schematic diagram of the detection effect of the present invention applied to OHS hyperspectral satellite images. Detailed Embodiments
[0029] As Figures 1 to 3 shown, in this embodiment, the present invention includes the following steps:
[0030] Step S1: Make a dataset for ship target detection in medium-resolution remotely sensed images;
[0031] The said step S1 includes the following steps: S11. Using the remote sensing images obtained by the OHS hyperspectral satellite of Zhuhai-1 and Sentinel-2 of the European Space Agency, select bands 12, 6, 3 (OHS) and 4, 3, 2 (Sentinel-2) respectively, and perform preprocessing operations such as radiometric correction, geometric correction, orthorectification and color enhancement, and output as a true color image. S12. The targets are divided into three major categories: other ships (Other-Ship), aircraft carriers (Aircraft-Carrier) and background (Backguond). The ship targets need to meet the requirements that the width ≥ 30 meters (3 pixels) and the length ≥ 80 meters (8 pixels). The annotation box adopts the drawing form of OHB and the interpretation standard of the corresponding quality inspection process. S13. According to the established interpretation standard, draw ship target samples from the above preprocessed true color image by manual visual interpretation and save them as vector files. S14. Set the cropping size to 512 pixels × 512 pixels, the cropping overlap degree to 20%, and retain the targets with the remaining area of the target box truncation ≥ 60%. Crop the remote sensing image and the vector file synchronously according to the above parameters, and output as the original ship target detection dataset in the organization paradigm of the PASCAL VOC dataset and the ratio of the training validation set to the test set of 7:2:1. S15. In order to enable the model to have better detection performance in complex scenarios, randomly enhance the dataset images through three methods: "noise filling", "contrast adjustment" and "angle rotation". Assume that the remote sensing image of the ship target detection dataset constructed above is X_(l) ∈ R^(H×W×B), where H, W, and B are the number of rows, columns and channels of the image respectively. "Noise filling" mainly uses Gaussian noise (GaussianNoise), and its formula is Y_out = X_(l) + X_Means + X_sigma × G(d), where Y_out is the resulting image after adding Gaussian noise (GaussianNoise), X_(l) is the input original image, X_Means is the average value of image pixels, X_sigma represents the standard deviation, and G(d) is the Gaussian distribution random value of random numbers; "Contrast adjustment" is to expand or shrink the difference between bright and dark pixel points of the image on the basis of keeping the average brightness of the image unchanged, and its adjustment formula is: Y_out = X_Means + (X_I - X_Means) × (1 - Q_percent), where Y_out is the resulting image after contrast adjustment, X_(l) is the input original image, X_Means is the average value of image pixels, and Q_percent is the adjustment range [-1, 1]. "Angle rotation" is represented by a rotation matrix. For the original image, the rotation matrix for rotating the angle θ around the origin is as follows:
[0032]
[0033] Among them, cos(θ) and sin(θ) are the cosine and sine values of the rotation angle θ, respectively.
[0034] The input image X l For each pixel coordinate (x, y) multiplied by the rotation matrix R(θ), the rotated pixel coordinates (x', y') can be obtained:
[0035]
[0036] Step S2: Construct the YOLO-MediumResolutionship network, abbreviated as the YOLO-MRship network. The YOLO-MRship network model consists of three parts: the backbone feature extraction network Backbone, the feature fusion network Neck, and the detection output decoupling head head. The backbone feature extraction network Backbone is composed of a first CBM module, a second CBM module, a first C2F module, a third CBM module, a first SimAM module, a second C2F module, a fourth CBM module, a third C2F module, a fifth CBM module, a fourth C2F module, a first SPPF module, and a first PKI multi-scale convolution kernel module connected in sequence;
[0037] The feature fusion network Neck contains feature enhancement extraction sub-modules at high, medium, and low scales. The feature fusion network Neck includes a first Upsample module, a first Concat module, a fifth C2F module, a second Upsample module, a second Concat module, a sixth C2F module, a sixth CBM module, a first SPD-Conv module, a second Concat module, a seventh C2F module, a seventh CBM module, a third Concat module, a second SPD-Conv module, and an eighth C2F module connected in sequence;
[0038] The detection output decoupling head "head" includes three detection decoupling heads corresponding to high, medium, and low scales respectively. Each decoupling head is composed of an eighth CBM module, a first Conv2d module, Bbox.Loss, a ninth CBM module, a second Conv2d module, and Cls.Loss. The sixth CBM module is connected to the first decoupling head and is respectively connected to two branches in parallel. The first branch is sequentially connected to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the first decoupling head. The second branch is sequentially connected to the ninth CBM module, the second Conv2d module, and Cls.Loss of the first decoupling head. The seventh CBM module is connected to the second decoupling head and is respectively connected to two branches in parallel. The first branch is sequentially connected to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the second decoupling head. The second branch is sequentially connected to the ninth CBM module, the second Conv2d module, and Cls.Loss of the second decoupling head. The eighth C2F module is connected to the third decoupling head and is respectively connected to two branches in parallel. The first branch is sequentially connected to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the third decoupling head. The second branch is sequentially connected to the ninth CBM module, the second Conv2d module, and Cls.Loss of the third decoupling head.
[0039] The first PKI multi-scale convolution kernel module is respectively connected to the first Upsample module and the third Concat module in parallel; the second C2F module is also connected to the first Concat module in parallel; the first C2F module is also connected to the second Upsample module; the first CBM module to the ninth CBM module are all composed of a first Conv2d unit, BatchNorm2d, and Mish activation function connected in sequence.
[0040] The full names of the first C2F module to the eighth C2F module are all CSPBottleneckwith2Convolutions, which are used to achieve feature cross-stage aggregation. The first C2F module to the eighth C2F module all include a Split processing unit, a first Concat unit, and several second Conv2d units; among them, the first Concat unit is one of the basic processing units, and is essentially the same as the first Concat module and the second Concat module in the model structure diagram; the first Conv2d unit and the second Conv2d unit are the most basic units in the deep learning model.
[0041] The internal calculations of the first C2F module to the eighth C2F module are as follows. Assume that the feature map input to any of the first C2F module to the eighth C2F module is X_in. First, the feature is processed by the Split processing unit, then passed through several Bottleeck modules composed of the second Conv2d units, and finally the features of each scale are processed by the first Concat unit. The entire processing process is expressed by formulas 1-3 as follows:
[0042] X split1 ,X split2 =Split(X in )#(1-3)
[0043] Among them, Split is the feature map splitting operation, and X_split1 and X_split2 are the split feature maps.
[0044] F bl =Bottleeck(Conv(X split1 ))#(1-4)
[0045] Among them, Bottleeck is for deep feature extraction, and F bl is the processed feature map.
[0046] F1 out =Conv 1×1 (Concat(X split1 ,X split2 ,F bl )#(1-5)
[0047] Among them, F1 out is the output feature map corresponding to the input of any of the first C2F module to the eighth C2F module, and the first Concat unit is for feature fusion.
[0048] The full name of the first SPPF module is SpatialPyramidPoolingFast. The first SPPF module consists of five components: initial_Conv, MaxPool1, MaxPool2, MaxPool3, and Final_Conv. Let the input feature map be X_in. First, perform scale adjustment through a 1×1 two-dimensional convolution feature, then use three MaxPool layers to pool the feature map, and finally perform feature fusion on the features of each scale. The entire processing process is expressed by formulas 1-6 to 1-10 as follows:
[0049] X1=Conv 1×1 (X in )#(1-6)
[0050] X pool1= MaxPool2d(X1) #(1 - 7)
[0051] X pool2 = MaxPool2d(X pool1 ) #(1 - 8)
[0052] X pool3 = MaxPool2d(X pool2 ) #(1 - 9)
[0053] F2 out = Conv 1×1 (Concat(X1, X pool1 , X pool2 , X pool3 ) #(1 - 10)
[0054] Among them, F2 out is the final output of the first SPPF module, Conv 1×1 is the convolution operation, MaxPool2d is the max pooling layer, and Concat is the feature concatenation layer.
[0055] The SimAM module is derived from neuroscience theory. By designing an energy function to calculate the attention weights, 3D attention weights can be derived for the feature map without additional parameters. The energy function of each neuron is expressed as follows:
[0056]
[0057] In the formula, t and x i are the input features X ∈ R C×H×W , i is the index of the spatial dimension, M = H × W is the number of neurons in each channel, and the transformation weights and biases w t and b t are expressed as follows:
[0058]
[0059] Among them, the average value here is the mean value of all neurons except t in the channel. According to formula (1 - 14), the minimum energy is:
[0060]
[0061] Assuming that all pixels of the feature map have the same distribution, based on formula (1 - 16), the weight of each neuron is
[0062] Therefore, the SimAM attention mechanism can be expressed as:
[0063]
[0064] In the formula, E is grouped in both the channel and spatial dimensions as The sigmoid function is added to limit the value of E.
[0065] The first PKI multi-scale convolution kernel module is composed of a PKI module and a CAA module. The first PKI multi-scale convolution kernel module is added to the end of the backbone network, that is, connected in series behind the first SPPF module, and the height-scale feature map is output by the PKI module. The PKI module separately obtains cross-scale context texture features through a group of parallel depth convolution layers and local feature information through a small kernel convolution. It is expressed by the mathematical formula as:
[0066]
[0067] Among them, is the local feature extracted by k s ×k s convolution, is the context texture feature extracted by the k (m) ×k (m) th depth convolution (DWConv).
[0068] In the present invention, k s = 3, k (m) = (m + 1) × 2 + 1. For n = 0, X l-1,n = X l-1 . Then, the local and context features are fused through a 1×1 Conv, and the mutual relationship between the channel feature maps is shown as follows in (1-20):
[0069]
[0070] Among them, represents the output feature. The 1×1 Conv is used to fuse the feature maps with different receptive field sizes.
[0071] In the ContextAnchorAttention (CAA) module, the feature map first passes through the Averagepooling layer, and then the local region feature is obtained through a 1×1 Conv, which is expressed by the following formula (1-21):
[0072]
[0073] Among them, P avg represents the Averagepooling layer operation, and then two Depth-wiseConv are used to approximate the standard large kernel Depth-wiseConv:
[0074]
[0075] Next, the feature map is output through a 1×1 Conv and the Sigmoid activation function:
[0076]
[0077] The feature maps output by the first SPPF module pass through the PKI module and the CAA module respectively to output P l-1,n , A l-1,n feature maps. First, multiply P l-1,n , A l-1,n element-wise, and then add the result of the multiplication to P l-1,n to obtain Finally, adjust the output through a 1×1 Conv This process is represented by the following formula:
[0078]
[0079] where ⊙ represents element-wise multiplication, represents element-wise summation. PKI_map l,n is the result output of the PKIBlock multi-scale convolution kernel module.
[0080] The first SPD-Conv module consists of a spatial-to-depth (SPD) layer and a non-strided convolution (Conv). In the SPD structure, the feature map of any size S×S×C1 is downsampled to generate four equal-sized feature maps of size S / 2×S / 2×C1. Subsequently, the individual feature maps are concatenated along the C1 dimension to obtain a feature map of size S / 2×S / 2×4C1. Then, the non-strided convolution layer is used to adjust the number of channels of the feature map, retaining all discriminative feature information, to obtain a feature map of size S / 2×S / 2×C2.
[0081] The loss function of this model consists of a classification loss and a regression loss, corresponding to Cls.Loss and Bbox.Loss respectively. The classification loss uses the cross-entropy loss function, which is represented by the following formula:
[0082]
[0083] where C represents the total number of classes, y i is the true label, is the probability that the model predicts to belong to class i.
[0084] The regression loss adopts DFLLoss+CIOULoss, and DFLLoss is represented by the following formula:
[0085] DFL(S i , S i+1 ) = -((yi+1 -y)log(S i )+(y - y i )log S i+1 )#(2 - 2)
[0086] Among them, S i is the Sigmod output of the network, and y i , y i+1 are the probabilities at the two positions on the left and right of the true label value.
[0087] After the three feature maps output by the feature fusion network (Neck) pass through the above detection decoupling head module, they are respectively output as Batch_Size×66×80×80, Batch_Size×66×40×40, and Batch_Size×66×20×20.
[0088] Step S3: Set the training hyperparameters Batch_Size to 16, Epoch to 300, and init_Lr to 0.007. Put the medium-resolution remote sensing image ship target detection dataset made in Step S1 into the YOLO-MRship model constructed in Step S2 for training. Verify and iteratively optimize the training results of each round through the loss function to obtain the optimal training weights.
[0089] Step S4: Input the remote sensing image to be detected and the optimal model weights obtained in Step S3 into the YOLO-MRship network constructed in Step S2 for inference. Detect the ship targets in the image and convert the detection result coordinates into longitude and latitude, and save them as a Shapefile vector file.
[0090] In order to explore the advantages of the YOLO-MRship network model in multi-classification of ship targets in medium-resolution remote sensing images, based on the test set part of the dataset in Step S1, this paper selects the mainstream object detection models / algorithms YOLOv5, YOLOV8, SSD, and Faster_RCNN in the current computer vision field for comparative experiments. The experimental results are shown in Table 1 below.
[0091] Table 1. Comparison of the performance of each model on the test set Tab.4 Comparison of the performance of each model on the test set
[0092]
[0093] The test results show that YOLO-MRship has obtained higher detection accuracy compared with models with the same number of parameters.
[0094] In summary, the present invention uses the Zhuhai-1 hyperspectral satellite and the Sentinel-2 multispectral satellite as data sources, innovatively introduces the Mish activation function, the SPD-Conv module, and the PKI multi-scale convolution kernel module, constructs the YOLO-MRship model, and while maintaining the lightweight of the model, greatly improves the detection accuracy of medium-resolution optical remote sensing ships. The experimental results show that the method described in the present invention has an average accuracy of mAP@0.5 of 94.48%, the model parameters are 8.27M, and it has the characteristics of high detection accuracy and lightweight model.
Claims
1. A ship target detection method applicable to medium-resolution optical remote sensing images, characterized in that, The detection method includes the following steps: Step S1. Make a ship target detection dataset for medium-resolution remote sensing images; Step S2. Construct a YOLO-Medium Resolution ship network, abbreviated as YOLO-MRship network; Step S3. Use the ship target detection dataset for medium-resolution remote sensing images made in Step S1 to train the YOLO-MRship network constructed in Step S2, and perform forward and backward propagation based on the loss function for iterative optimization training to obtain the optimal YOLO-MRship network weights; Step S4. Input the remote sensing image to be detected and the optimal model YOLO-MRship network weights obtained by iterative training in Step S3 into the YOLO-MRship network constructed in Step S2 for inference, detect the ship targets in the remote sensing image, and save the detection results as vector files.
2. The ship target detection method applicable to medium-resolution optical remote sensing images according to claim 1, wherein The said Step S1 includes the following steps: S11. Complete preprocessing operations such as band selection, radiometric correction, geometric correction, orthorectification, and color enhancement on the medium-resolution remote sensing images used to construct the dataset, and output true-color images; S12. Establish ship target interpretation criteria, including ship target classification criteria, ship target annotation box forms, and quality inspection processes; S13. According to the interpretation criteria, manually visually interpret the preprocessed true-color images to draw ship target samples and obtain sample vectors; S14. Set parameters such as cropping size, cropping overlap, target box truncation threshold, dataset training validation test set ratio, and output dataset organization paradigm, and synchronously crop the remote sensing images and ship target vectors to output the original ship target detection dataset; S15. Use data augmentation methods such as noise filling, contrast adjustment, and angle rotation to augment the original ship target detection dataset and expand the number of dataset samples.
3. The ship target detection method applicable to medium-resolution optical remote sensing images according to claim 1, wherein: The YOLO-MRship network model in the said Step S2 is composed of three parts: a backbone feature extraction network (Backbone), a feature fusion network (Neck), and a detection output decoupling head (head).
4. A ship target detection method applicable to medium-resolution optical remote sensing images according to claim 3, characterized in that: The backbone feature extraction network (Backbone) is composed of a first CBM module, a second CBM module, a first C2F module, a third CBM module, a first SimAM module, a second C2F module, a fourth CBM module, a third C2F module, a fifth CBM module, a fourth C2F module, a first SPPF module, and a first PKI multi-scale convolution kernel module connected in sequence; The feature fusion network (Neck) includes a first Upsample module, a first Concat module, a fifth C2F module, a second Upsample module, a second Concat module, a sixth C2F module, a sixth CBM module, a first SPD-Conv module, a second Concat module, a seventh C2F module, a seventh CBM module, a third Concat module, a second SPD-Conv module, and an eighth C2F module connected in sequence; The detection output decoupling head includes detection decoupling heads of three scales. Each decoupling head is composed of an eighth CBM module, a first Conv2d module, Bbox.Loss, a ninth CBM module, a second Conv2d module, and Cls.Loss. The sixth CBM module is connected to the first decoupling head and is respectively connected to two branches in parallel. The first branch is sequentially connected to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the first decoupling head, and the second branch is sequentially connected to the ninth CBM module, the second Conv2d module, and Cls.Loss of the first decoupling head; the seventh CBM module is connected to the second decoupling head and is respectively connected to two branches in parallel. The first branch is sequentially connected to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the second decoupling head, and the second branch is sequentially connected to the ninth CBM module, the second Conv2d module, and Cls.Loss of the second decoupling head; the eighth C2F module is connected to the third decoupling head and is respectively connected to two branches in parallel. The first branch is sequentially connected to the eighth CBM module, the first Conv2d module, and Bbox.Loss of the third decoupling head, and the second branch is sequentially connected to the ninth CBM module, the second Conv2d module, and Cls.Loss of the third decoupling head.
5. The ship target detection method applicable to medium-resolution optical remote sensing images according to claim 4, characterized in that: The first PKI multi-scale convolution kernel module is respectively connected to the first Upsample module and the third Concat module in parallel; the second C2F module is also connected to the first Concat module in parallel; the first C2F module is also connected to the second Upsample module.
6. The ship target detection method for medium-resolution optical remote sensing images according to claim 4, characterized in that: The first to ninth CBM modules are all composed of a first Conv2d unit, BatchNorm2d, and a Mish activation function connected in sequence.
7. A ship target detection method applicable to medium-resolution optical remote sensing images according to claim 4, characterized in that: The first SPPF module is composed of five components: initial_Conv, MaxPool1, MaxPool2, MaxPool3, and Final_Conv.
8. The ship target detection method for medium-resolution optical remote sensing images according to claim 4, characterized in that: The loss function is composed of a classification loss and a regression loss, corresponding to Bbox.Loss and Cls.Loss respectively.
9. A method for ship target detection applicable to medium-resolution optical remote sensing images according to claim 4, characterized in that: The first PKI multi-scale convolution kernel module is composed of a PKI module and a CAA module.