Method for Detecting Ships in Remote Sensing Images Based on YOLO-V Network
By improving the YOLOv5 network, adding a collaborative attention layer and a hybrid space pyramid pooling layer, using the EIOU loss function and a large-scale shallow feature map detection layer, the problems of low detection accuracy and large calculation amount of small and medium-sized ships in remote sensing images are solved, and efficient real-time detection of ships is achieved.
Patent Information
- Application Number
- CN202211339370.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-10-29
AI Technical Summary
In the detection of remote sensing images, the prior art has problems such as poor accuracy, low positioning accuracy, large missed detection, large calculation amount and long detection time, making it difficult to detect ships in real time.
Improved on the basis of the YOLOv5 network, add a collaborative attention layer and a hybrid space pyramid pooling layer, use EIOU as the bounding box loss function, and add a detection layer of large-scale shallow feature maps to the Head subnet, and introduce leap connections to the Neck subnet to build a YOLO-V network.
It improves the accuracy and positioning accuracy of small ships, reduces missed inspections, reduces calculation amount, and realizes the ability to detect ships in real time.
Smart Images

Figure CN115546650B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for detecting ships in remote sensing images based on an improved YOLO-V (You Only Look Once-V) network on the basis of a single-stage YOLOv5 network in the technical fields of deep learning and digital image processing. The present invention can be used to automatically detect ships from optical remote sensing images. Background Art
[0002] Ships are the main transportation carriers on the ocean. Accurately and quickly locating the position of ships is crucial for China's national security, marine protection, fishing vessel management, etc. Ship target detection technology based on images relies on image features to determine whether there are ships and locate the positions of the ships. With the development of remote sensing images and deep learning technology, remote sensing images can be batch input into a deep learning model, and the model can automatically detect ships in the images using the trained weights, improving the accuracy of ship detection.
[0003] Authors such as Lu Yunjiao and Luo Sha proposed a method for detecting ships in remote sensing images based on texture high-order fractal features in their published paper "Method for Detecting Sea Surface Ship Targets Based on Texture High-Order Fractal Features" (Ship Science and Technology, Volume 4, 2020, Pages 37-39). The specific steps of this method are as follows: First, process the remote sensing image containing ships; then extract the texture high-order fractal features of the ship target and introduce a convolutional neural network to analyze the change characteristics of the ship target; finally, establish a ship detection model to detect ships in the remote sensing image. The disadvantage of this method is that ships can be detected in simple remote sensing images. However, when there are many complex background interferences such as cloud cover, sea clutter, and land in the remote sensing image, there are many missed detections of ships and misdetections of identifying the background as ships, resulting in insufficient ship detection accuracy of this method.
[0004] In their paper "Ship Target Detection in Optical Images Based on Improved Mask R-CNN" (Beijing Ligong Daxue Xuebao / Transaction of Beijing Institute of Technology, pp. 734-744, 2021), Xiao Ma and Xin Jin proposed a method for detecting ships in remote sensing imagery based on Mask R-CNN. The method's steps are as follows: First, a discriminant module is added to the Mask R-CNN backbone network. This module separates the semantic labels of the ship's hull and superstructure by comparing the number of mask labels in the image and the size of the mask area. This is then input into the two subtask branches for completing the category prediction and semantic segmentation tasks, respectively. Finally, the category prediction branch obtains the detected ship's category, outline, and location. However, a drawback of this method is that Mask R-CNN is a two-stage network, requiring the generation of a large number of target pre-selected boxes within the model. This results in a high computational load, slow detection speed, and significant resource waste. This makes real-time ship detection difficult, making it difficult to apply to real-world industrial applications.
[0005] In their paper, "An Improved YOLOv4-Based SAR Ship Detection Algorithm" (Electronic Measurement Technology, Vol. 45, No. 11, 2022, pp. 120-125), authors Chen Yang, Zhang Ming, and others proposed a YOLOv4 algorithm based on multi-scale fusion and attention enhancement to detect ships in Synthetic Aperture Radar (SAR) imagery. The method's steps are as follows: First, a Convolutional Block Attention Module (CBAM) attention mechanism is added to the original YOLOv4 PANet (Path Aggregation Network). Then, an enhanced K-means clustering algorithm is used to cluster the ground truth boxes of ships in the dataset, and the resulting anchor boxes are linearly transformed. Finally, the model is trained and used to detect ships in remote sensing imagery. The shortcomings of this method are: there are a large number of small ships active in the sea areas near the coastline. These small ships are small in size and occupy very few pixels in the remote sensing image, which makes the method easily lose the ship features during continuous downsampling. Therefore, the accuracy of this method in detecting small ships is low, there are many missed detections, and the ship positioning accuracy is poor, which cannot meet industrial needs. Summary of the invention
[0006] The object of the present invention is to address the deficiencies of the above-mentioned existing technologies, and propose a method for detecting ships in remote sensing images based on the YOLO-V network, which is used to solve the problems of poor accuracy in detecting small ships, low positioning accuracy, a large number of small ships being missed in detection, as well as a large amount of network calculations, long detection time, and difficulty in real-time detecting ships from remote sensing images.
[0007] The idea to achieve the object of the present invention is as follows: The YOLO-V network designed in the present invention is improved on the basis of the existing single-stage YOLOv5 network, and it includes four sub-networks, namely the input Input sub-network, the backbone Backbone sub-network, the neck Neck sub-network, and the head Head sub-network. In the Backbone sub-network of the present invention, a coordinate attention layer CA (Coordinate Attention) is added. The CA decomposes the channel attention into two one-dimensional feature encoding processes, aggregates the feature information of the input feature map along the horizontal and vertical directions respectively. In this way, long-distance dependencies can be captured along one spatial direction, and precise position information can be located in the other direction. The two generated feature maps are respectively encoded into a pair of direction-aware and position-aware attention feature maps, which are applied to the input feature map to enhance the representation ability of the object of interest, and enhance the representation ability of the network for ship features. The hybrid spatial pyramid pooling HSPP layer (Hybrid Spatial Pyramid Pooling) is used to replace the SPP layer (Spatial Pyramid Pooling) in the YOLOv5 network. The hybrid spatial pyramid pooling layer performs a 1×1 convolution operation on the input feature map F0, converts its size from 20×20×512 to 20×20×256 to generate F1, and reduces the computational amount by reducing the number of channels of the feature map. Then, a global average pooling operation with a size of (kernel_size = 5, stride = 1, padding = 2) is applied to F1 to generate F2; global maximum pooling operations with sizes of (kernel_size = 9, stride = 1, padding = 4) and (kernel_size = 13, stride = 1, padding = 6) are respectively applied to F2 to generate F3 and F4.The sizes of F1, F2, F3, and F4 are all 20×20×256. These four feature maps are concatenated to generate a feature map F5 with a size of 20×20×1024. Finally, to ensure that the output feature map has the same size as the input feature map, a 1×1 convolution operation is performed on F5 to reduce the number of channels, obtaining an output feature map F6 with the same size as the input, which better fuses the local and global information of the feature map. In the Head sub-network of the YOLOv5 network, the present invention adds a detection layer with large-scale shallow feature maps. Therefore, there are a total of four detection layers with different feature map sizes in the Head sub-network of the YOLO-V network designed in the present invention. The detection layer with a large-size feature map is used to detect small ships, and the detection layer with a small-size feature map is used to detect large ships. A leapfrog connection is introduced in the Neck sub-network, shortening the information fusion path between the high-level and low-level feature maps and improving the network's information fusion ability. The YOLO-V network designed in the present invention uses EIOU as the bounding box loss function. The EIOU loss function consists of three parts: the bounding box overlap loss, the bounding box center point distance loss, and the bounding box width and height loss. The width and height loss directly minimizes the difference between the width and height of the predicted bounding box and the ground truth bounding box, making the positioning accuracy of the YOLO-V network better and alleviating the problems of poor detection accuracy for small ships, low positioning accuracy, and many missed detections of small ships. The YOLO-V network designed in the present invention belongs to a single-stage object detection network, which only needs to extract feature information once for ship detection, alleviating the problems of large computational load, long detection time, and difficulty in real-time ship detection.
[0008] The specific steps implemented by the present invention are as follows:
[0009] Step 1, construct the Input sub-network in the YOLO-V network:
[0010] The Input sub-network is composed of a scaling layer and an enhancement layer connected in series. The scaling layer and the enhancement layer are implemented using the image adaptive scaling algorithm and the data augmentation algorithm respectively; the data augmentation algorithm is implemented using the Mosaic function and the Mixup function.
[0011] Step 2, construct the Backbone sub-network of the YOLO-V network:
[0012] Build an 11-layer Backbone sub-network, and its structure is as follows: Focus layer, the first convolutional layer, the first C3 layer, the second convolutional layer, the second C3 layer, the third convolutional layer, the third C3 layer, the fourth convolutional layer, Hybrid Spatial Pyramid Pooling layer, Coordinate Attention layer, the fourth C3 layer; the input image size of the Focus layer is 640×640×3, the output image size is set to 320×320×32, the convolutional kernel size is set to 3×3, and the stride is set to 1; the convolutional kernel sizes of the first to fourth convolutional layers are all set to 3×3, the strides are all set to 2, the input channel numbers are set to 32, 64, 128, 256 respectively, and the output channel numbers are set to 64, 128, 256, 512 respectively; the output channel numbers of the first to fourth C3 layers are the same as the input channel numbers, which are set to 64, 128, 256, 512 respectively. Set the parameter n of the first C3 layer to 1 and the parameter shortcut to True, set the parameter n of the second and third C3 layers to 3 and the parameter shortcut to True, and set the parameter n of the fourth C3 layer to 1 and the parameter shortcut to False; set the input channel number and output channel number of the HSPP layer and the Coordinate Attention layer to 512;
[0013] Step 3, construct the Neck sub-network in the YOLO-V network:
[0014] Step 3.1, build an 11-layer Neck-A sub-network, and its structure is as follows: the first convolutional layer, the first upsampling layer, the first concatenation layer, the first C3 layer, the second convolutional layer, the second upsampling layer, the second concatenation layer, the second C3 layer, the third convolutional layer, the third upsampling layer, the third concatenation layer; the convolutional kernel sizes of the first to third convolutional layers are all set to 1×1, the strides are all set to 1, the input channel numbers are set to 512, 256, 128 respectively, and the output channel numbers are set to 256, 128, 128 respectively; the upsampling multiples of the first to third upsampling layers are all set to 2, and the upsampling method is all the nearest filling; the first to third concatenation layers concatenate the input images along the channel direction; the input channel numbers of the first C3 layer to the second C3 layer are 512, 256, and the output channel numbers are set to 256, 128, and the parameter n is set to 1 and the parameter shortcut is set to False;
[0015] Step 3.2: Build a 10-layer Neck-B sub-network, whose structure is as follows: the first C3 layer, the first convolutional layer, the first splicing layer, the second C3 layer, the second convolutional layer, the second splicing layer, the third C3 layer, the third convolutional layer, the third splicing layer, and the fourth C3 layer; set the parameter n of the first to fourth C3 layers to 1, the parameter shortcut to False, the input channel numbers to 192, 512, 1024, 1024 respectively, and the output channel numbers to 128, 128, 256, 512 respectively; set the convolutional kernel sizes of the first to third convolutional layers to 3, the strides to 2, the input channel numbers to 128, 128, 256 respectively, and the output channel numbers to 128, 256, 512 respectively; the first to third splicing layers concatenate the input images along the channel direction;
[0016] Step 3.3: Cascade the third splicing layer in the Neck-A sub-network with the first C3 layer in the Neck-B sub-network, cascade the second C3 layer in the Neck-A sub-network with the first splicing layer in the Neck-B sub-network, cascade the first C3 layer in the Neck-A sub-network with the second splicing layer in the Neck-B sub-network, and cascade the first convolutional layer in the Neck-A sub-network with the third splicing layer in the Neck-B sub-network to obtain the Neck sub-network;
[0017] Step 4: Build the Head sub-network in the YOLO-V network:
[0018] The Head sub-network is composed of the first detection layer, the second detection layer, the third detection layer, and the fourth detection layer in parallel to implement the detection layer with large-scale shallow feature maps in the Head sub-network; each detection layer sets the number of detected classes nc = 1; use the K-means++ algorithm to cluster and set 3 anchor box sizes for each detection layer;
[0019] Step 5: Build the YOLO-V network:
[0020] Step 5.1: Cascade the first C3 layer, the second C3 layer, and the third C3 layer in the Backbone sub-network with the first concatenation layer, the second concatenation layer, and the third concatenation layer in the Neck-A sub-network respectively, and cascade the fourth C3 layer with the first convolutional layer in the Neck-A sub-network; at the same time, cascade the second C3 layer and the third C3 layer in the Backbone sub-network with the first concatenation layer and the second concatenation layer in the Neck-B sub-network respectively to introduce skip connections in the Neck sub-network; cascade the third concatenation layer in the Neck-A sub-network with the first C3 in the Neck-B sub-network, and cascade the first convolutional layer, the first C3 layer, and the second C3 layer in the Neck-A sub-network with the third concatenation layer, the second concatenation layer, and the third concatenation layer in the Neck-B sub-network respectively; cascade the first C3 layer, the second C3 layer, the third C3 layer, and the fourth C3 layer in the Neck-B sub-network with the first detection layer, the second detection layer, the third detection layer, and the fourth detection layer in the Head sub-network respectively to complete the connection of the three sub-networks: the Backbone sub-network, the Neck sub-network, and the Head sub-network.
[0021] Step 5.2: Cascade the enhancement layer of the Input sub-network with the Focus layer of the Backbone sub-network to form the YOLO-V network.
[0022] Step 6: Generate the training set and the validation set:
[0023] Step 6.1: Download at least 25 large-size remote sensing images to form a sample set. The remote sensing images in the sample set contain ships of various scales such as fishing boats about 30 meters long and warships about 300 meters long, with various sea conditions such as near the coast, near the port, and in the open sea, and also have complex backgrounds such as cloud occlusion and sea clutter interference.
[0024] Step 6.2: Crop each large-size remote sensing image into an image of 512×512 pixels.
[0025] Step 6.3: Select the images containing ships from all the cropped images to form the positive sample set; form the negative sample set from the sub-images obtained by segmenting the land, sea clutter, and cloud complex backgrounds in the positive samples.
[0026] Step 6.4: Perform operations on each image in the positive and negative sample sets in turn, including changing the brightness, blurring the image, rotating the image, changing the contrast, adding noise, and random combinations of the above operations after adding noise, such as rotating the image after adding noise.
[0027] Step 6.5: Combine all the images after the transformation in Step 6.4 to form a remote sensing image ship dataset containing multi-scale ships, various sea conditions, and complex backgrounds.
[0028] Step 6.6, divide the dataset into a training set and a validation set in a ratio of 9:1;
[0029] Step 7, train the YOLO-V network:
[0030] Step 7.1, set the training parameters. Set the batch-size to 16, randomly and without repetition select 16 images from the training set each time and input them into the network. Use the yolov5s.pt as the weight parameter file; both the classification loss function and the confidence loss function use the binary cross-entropy loss function; use EIOU as the bounding box loss function;
[0031] Step 7.2, input the training set and the validation set into the YOLO-V network in turn. Convert the size of the images input to the Input sub-network to 640×640×3 through the image adaptive scaling algorithm, and enhance the images through the enhancement algorithm. The enhancement algorithm uses the Mosaic function and the Mixup function. The Mixup function adds two input images and their corresponding labels in proportion to generate a new image and the corresponding label; set the Mixup parameter to 0.8, that is, with a probability of 0.8, determine whether to perform Mixup enhancement on the input images;
[0032] Step 7.3, use the Adam optimizer and the stochastic gradient descent method to iteratively update the weights of the YOLO-V network parameters, and then verify the training effect on the validation set until the classification loss function, the confidence loss function, and the bounding box loss function all converge, and obtain the trained YOLO-V network;
[0033] Step 8, detect ships in the remote sensing images:
[0034] Step 8.1, batch input all the remote sensing images to be detected into the trained YOLO-V network, and perform scaling operations on the input images through the image adaptive scaling algorithm;
[0035] Step 8.2, perform ship detection on the images to be detected, and obtain the center point position, width and height of the ship bounding box, as well as the confidence of the ship;
[0036] Step 8.3, output the ship detection results and save them as a label file in txt format.
[0037] The present invention has the following advantages compared with the prior art:
[0038] First, the present invention adds a collaborative attention layer and a hybrid spatial pyramid pooling layer to the Backbone sub-network of the existing YOLOv5 network, uses EIOU as the bounding box loss function, adds a detection layer with large-scale shallow feature maps in the Head sub-network, and introduces a leapfrog connection in the Neck sub-network, overcoming the deficiencies of the prior art such as more undetected small ships and alleviating the poor accuracy and low positioning accuracy of detecting small ships, making the positioning bounding box of detecting ships from optical remote sensing images more accurate in the present invention and improving the accuracy of detecting ships.
[0039] Second, the YOLO-V network constructed by the present invention belongs to a single-stage object detection network, and only needs to extract feature information once to achieve ship detection. The YOLO-V network reduces the network calculation amount, solves the deficiencies of the prior art such as long detection time and difficulty in real-time ship detection, makes the time consumed by the present invention to detect ships from optical remote sensing images shorter through the YOLO-V network, improves the efficiency of detecting ships from optical remote sensing images, and can be used for real-time detection of ships in optical remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the present invention;
[0041] Figure 2 is a schematic structural diagram of the YOLO-V network constructed by the present invention;
[0042] Figure 3 is a schematic structural diagram of the HSPP layer in the Backbone sub-network of the YOLO-V network constructed by the present invention;
[0043] Figure 4 is a schematic structural diagram of the convolutional layer in the Backbone sub-network and Neck sub-network constructed by the present invention;
[0044] Figure 5 is a schematic structural diagram of the C3 layer in the Backbone sub-network and Neck sub-network constructed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] The technical solutions and effects of the present invention will be further described in detail below in conjunction with the drawings and embodiments.
[0046] Refer to Figure 1 for a further description of the implementation steps of the embodiments of the present invention.
[0047] The YOLO-V network constructed by the present invention includes four sub-networks, namely: Input sub-network, Backbone sub-network, Neck sub-network and Head sub-network, as Figure 2As shown by the four dashed boxes in []. The Input sub-network performs tasks such as data augmentation and adaptive image scaling on the input images to achieve pre-training preprocessing operations. The Backbone sub-network is responsible for feature extraction operations on the images passed in from the Input sub-network to obtain rich feature information. The Neck sub-network is responsible for feature fusion of the information extracted by the Backbone sub-network. The Head sub-network is responsible for ship prediction on the images passed in from the Neck sub-network to obtain the output of the network.
[0048] Step 1, construct the Input sub-network in the YOLO-V network.
[0049] Build an Input sub-network for preprocessing the input images. This Input sub-network is composed of a scaling layer and an enhancement layer connected in series. The scaling layer and the enhancement layer are implemented using an image adaptive scaling algorithm and a data augmentation algorithm respectively. The data augmentation algorithm is implemented using the Mosaic function and the Mixup function.
[0050] Step 2, construct the Backbone sub-network of the YOLO-V network.
[0051] Build an 11-layer Backbone sub-network. Its structure is as follows: Focus layer, first convolutional layer, first C3 layer, second convolutional layer, second C3 layer, third convolutional layer, third C3 layer, fourth convolutional layer, Hybrid Spatial Pyramid Pooling (HSPP) layer, Coordinate Attention layer, fourth C3 layer. The input image size of the Focus layer is 640×640×3, the output image size is set to 320×320×32, the convolutional kernel size is set to 3×3, and the stride is set to 1. The convolutional kernel sizes of the first to fourth convolutional layers are all set to 3×3, the strides are all set to 2, the input channel numbers are set to 32, 64, 128, 256 respectively, and the output channel numbers are set to 64, 128, 256, 512 respectively. The output channel numbers of the first to fourth C3 layers are the same as the input channel numbers, and are set to 64, 128, 256, 512 respectively. The parameter n of the first C3 layer is set to 1, the parameter shortcut is set to True. The parameter n of the second and third C3 layers is set to 3, the parameter shortcut is set to True. The parameter n of the fourth C3 layer is set to 1, the parameter shortcut is set to False. The input channel numbers and output channel numbers of the HSPP layer and the Coordinate Attention layer are both set to 512.
[0052] Refer to Appendix Figure 3 , and make a further description of the HSPP layer of the Backbone sub-network in the YOLO-V network constructed by the present invention.
[0053] Figure 3It is a schematic diagram of the HSPP layer structure. The HSPP layer includes a first convolutional layer, an average pooling layer, a concatenation layer, a second convolutional layer, a first max pooling layer, and a second max pooling layer. Among them, the first convolutional layer, the average pooling layer, the concatenation layer, and the second convolutional layer are connected in series in sequence. The first max pooling layer is bridged between the output of the average pooling layer and the input of the concatenation layer. The second max pooling layer is bridged between the output of the average pooling layer and the input of the concatenation layer. The HSPP layer first performs a 1×1 convolutional operation on the input feature map F0, converting its size from 20×20×512 to 20×20×256 to generate F1, reducing the computational amount by reducing the number of channels of the feature map. Then, an average pooling operation with a convolutional kernel size of 5×5, a stride of 1, and a padding of 2 is applied to F1 to generate F2. Max pooling with a convolutional kernel size of 9×9, a stride of 1, and a padding of 4 and max pooling with a convolutional kernel size of 13×13, a stride of 1, and a padding of 6 are respectively applied to F2 to generate F3 and F4. The sizes of F1, F2, F3, and F4 are all 20×20×256. These four feature maps are concatenated to generate a feature map F5 with a size of 20×20×1024. Finally, to ensure that the output feature map has the same size as the input feature map, a 1×1 convolutional operation is performed on F5 to reduce the number of channels, obtaining an output feature map F6 with the same size as the input.
[0054] Step 3: Construct the Neck sub-network in the YOLO-V network.
[0055] The Neck sub-network is composed of a Neck-A sub-network and a Neck-B sub-network.
[0056] Step 3.1: Build an 11-layer Neck-A sub-network. Its structure is as follows in sequence: a first convolutional layer, a first upsampling layer, a first concatenation layer, a first C3 layer, a second convolutional layer, a second upsampling layer, a second concatenation layer, a second C3 layer, a third convolutional layer, a third upsampling layer, a third concatenation layer. Set the convolutional kernel sizes of the first to third convolutional layers to 1×1, the strides to 1, the input channel numbers to 512, 256, 128 respectively, and the output channel numbers to 256, 128, 128 respectively. Set the upsampling multiples of the first to third upsampling layers to 2, and the upsampling method to nearest filling. The first to third concatenation layers concatenate the input images along the channel direction. Set the input channel numbers of the first C3 layer to the second C3 layer to 512, 256, the output channel numbers to 256, 128, the parameter n to 1, and the parameter shortcut to False.
[0057] Step 3.2, construct a 10-layer Neck-B sub-network, and its structure is as follows: the first C3 layer, the first convolutional layer, the first splicing layer, the second C3 layer, the second convolutional layer, the second splicing layer, the third C3 layer, the third convolutional layer, the third splicing layer, and the fourth C3 layer. Set the parameter n of the first to fourth C3 layers to 1, the parameter shortcut to False, the input channel numbers to 192, 512, 1024, 1024 respectively, and the output channel numbers to 128, 128, 256, 512 respectively. Set the convolutional kernel size of the first to third convolutional layers to 3, the stride to 2, the input channel numbers to 128, 128, 256 respectively, and the output channel numbers to 128, 256, 512 respectively. The first to third splicing layers concatenate the input images along the channel direction.
[0058] Step 3.3, cascade the third splicing layer in the Neck-A sub-network with the first C3 layer in the Neck-B sub-network, cascade the second C3 layer in the Neck-A sub-network with the first splicing layer in the Neck-B sub-network, cascade the first C3 layer in the Neck-A sub-network with the second splicing layer in the Neck-B sub-network, and cascade the first convolutional layer in the Neck-A sub-network with the third splicing layer in the Neck-B sub-network to obtain the Neck sub-network.
[0059] Refer to Figure 4 , and make a further description of the convolutional layer structure in the Backbone sub-network and Neck sub-network constructed by the present invention.
[0060] The structures of the convolutional layers in the Backbone sub-network and Neck sub-network of the present invention are the same. As Figure 4 shown, the convolutional layer sequentially includes a convolutional module, a batch normalization module, and an activation function module. The parameter settings of the convolutional module are set according to the layer where the convolutional layer is located, and the activation function module uses the SiLU function.
[0061] Refer to Figure 5 , and make a further description of the C3 layer in the Backbone sub-network and Neck sub-network constructed by the present invention.
[0062] The structures of the C3 layers in the Backbone sub-network and Neck sub-network of the present invention are the same. As Figure 5(a) consists of two branches. The first branch is composed of a first convolutional layer, a Bottleneck group formed by concatenating n* Bottleneck layers, a splicing layer, and a third convolutional layer in series. The second branch is composed of a second convolutional layer, which is bridged between the input and the splicing layer. Here, n* = n × the depth factor of the network, and the setting of n is based on the layer where the C3 layer is located. In the YOLO-V network designed in the present invention, the depth factor is set to 0.33. As Figure 5 (b) shows that in the Bottleneck layer, the first branch is composed of a first convolutional layer and a second convolutional layer in series. The second branch is bridged between the input and the second convolutional layer of the first branch, and the input of the Bottleneck layer is directly added to the input of the first branch through the second branch. The parameter shortcut controls the existence of the second branch. When the parameter shortcut is set to True, the second branch exists, and the input and the output of the second branch are added, and the result of the addition operation is used as the output of the Bottleneck layer. When the parameter shortcut is set to False, the second branch does not exist, and the result of the first branch is used as the output of the Bottleneck.
[0063] Step 4, construct the Head sub-network in the YOLO-V network.
[0064] The Head sub-network is composed of a first detection layer, a second detection layer, a third detection layer, and a fourth detection layer in parallel, realizing a detection layer with large-scale shallow feature maps in the Head sub-network. Each detection layer is set with the number of detection categories nc = 1, and each detection layer is set with 3 anchor box sizes. Using the K-means++ algorithm to cluster, the 3 anchor box sizes set for each detection layer are: {19, 28, 34, 22, 34, 48}, {53, 34, 24, 66, 71, 24}, {57, 59, 40, 87, 87, 41}, {62, 87, 86, 64, 102, 100}.
[0065] Step 5, construct the YOLO-V network.
[0066] Step 5.1: Cascade the first C3 layer, the second C3 layer, and the third C3 layer in the Backbone sub-network with the first splicing layer, the second splicing layer, and the third splicing layer in the Neck-A sub-network respectively, and cascade the fourth C3 layer with the first convolutional layer in the Neck-A sub-network. At the same time, cascade the second C3 layer and the third C3 layer in the Backbone sub-network with the first splicing layer and the second splicing layer in the Neck-B sub-network respectively. Cascade the third splicing layer in the Neck-A sub-network with the first C3 in the Neck-B sub-network, and cascade the first convolutional layer, the first C3 layer, and the second C3 layer in the Neck-A sub-network with the third splicing layer, the second splicing layer, and the third splicing layer in the Neck-B sub-network respectively. Cascade the first C3 layer, the second C3 layer, the third C3 layer, and the fourth C3 layer in the Neck-B sub-network with the first detection layer, the second detection layer, the third detection layer, and the fourth detection layer in the Head sub-network respectively to complete the connection of the three sub-networks: the Backbone sub-network, the Neck sub-network, and the Head sub-network.
[0067] Step 5.2: Cascade the enhancement layer of the Input sub-network with the Focus layer of the Backbone sub-network to form the YOLO-V network.
[0068] Step 6: Generate the training set, validation set, and test set.
[0069] Step 6.1: In the embodiments of the present invention, download 25 large-sized remote sensing images to form a sample set. The remote sensing images in the sample set include fishing boats about 30 meters long, warships about 300 meters long, and ships of various scales, with various sea conditions such as near the coast, near the port, and the original coast, and also have a complex background with cloud occlusion and sea clutter interference.
[0070] Step 6.2: Crop each large-sized remote sensing image into an image of 512×512 pixels.
[0071] Step 6.3: Select the images containing ships from all the cropped images to form a positive sample set; form a negative sample set from the sub-images obtained by segmenting the land, sea clutter, and complex cloud background in the positive samples.
[0072] Step 6.4: Perform operations of changing brightness, blurring the image, rotating the image, changing contrast, adding noise, and random combinations of the above operations after adding noise to each image in the positive and negative sample sets. For example, perform a transformation of rotating the image after adding noise to the image.
[0073] Step 6.5: Combine all the images after the transformation in Step 6.4 to form a remote sensing image ship dataset containing ships of multiple scales, various sea conditions, and complex backgrounds.
[0074] Step 6.6, divide the dataset into a training and validation set and a test set at a ratio of 7:3.
[0075] Step 6.7, divide the training and validation set into a training set and a validation set at a ratio of 9:1.
[0076] Step 7, train the YOLO-V network:
[0077] Step 7.1, set the training parameters. Set the batch-size to 16, randomly and without repetition select 16 images from the training set each time and input them into the network. Use the yolov5s.pt as the weight parameter file. Use the binary cross-entropy loss function for both the classification loss function and the confidence loss function; use the EIOU (Efficient IOU) loss function for the bounding box loss function.
[0078] The binary cross-entropy loss function is as follows:
[0079] Loss = -(y·log e (y*) + (1 - y)·log e (1 - y*))
[0080] Where y* represents the probability that the network predicts the target as a ship, and y represents the label of the target. If the target belongs to a ship, y takes the value of 1; if the target does not belong to a ship, y takes the value of 0.
[0081] The bounding box loss function is as follows:
[0082]
[0083] Where IOU represents the intersection over union of the predicted box and the ground truth box, ρ 2 (b, b gt ) represents the distance between b and b gt , b and b gt respectively represent the center points of the predicted box and the ground truth box, c represents the diagonal length of the minimum bounding rectangle of the predicted box and the ground truth box, w, h and w gt , h gt respectively represent the widths and heights of the predicted box and the ground truth box, C w and C h respectively represent the widths and heights of the minimum bounding rectangles of the predicted box and the ground truth box.
[0084] Step 7.2, input the training set and the validation set into the YOLO-V network in sequence. Use the image adaptive scaling algorithm to convert the size of the image input to the Input sub-network to 640×640×3 (width and height are 640 respectively, and the number of channels is 3). The specific steps are as follows: First, calculate the scaling ratio according to the size of the image input to the Input sub-network and the required image size of the network 640×640×3; then, calculate the size of the scaled image according to the size of the image input to the network and the scaling ratio; finally, pad 0 around the scaled image to transform the size of the image to 640×640×3. Use the Mosaic function in the augmentation algorithm to perform random scaling and random cropping on the four input images in sequence to obtain four images, and then randomly arrange and splice the obtained images into one image. Use the Mixup function to add the two input images and their corresponding labels in proportion to generate a new image and its corresponding label. The Mosaic parameter is set to 1.0, that is, the Mosaic augmentation is performed on all images input to the network, and the Mixup parameter is set to 0.8, that is, the probability of 0.8 is used to determine whether to perform Mixup augmentation on the input images.
[0085] The steps of the image adaptive scaling algorithm are as follows:
[0086] In the first step, calculate the scaling ratio according to the size of the image input to the Input sub-network and the required image size of the network 640×640×3;
[0087] In the second step, calculate the size of the scaled image according to the size of the image input to the network and the scaling ratio;
[0088] In the third step, pad 0 around the scaled image to transform the size of the image to 640×640×3.
[0089] The Mosaic function is to perform random scaling and random cropping on the four input images in sequence to obtain four images, and then randomly arrange and splice the obtained images into one image. The Mosaic parameter is set to 1.0, that is, the Mosaic augmentation is performed on all images input to the network.
[0090] The Mosaic function is to perform random scaling and random cropping on the four input images in sequence to obtain four images, and then randomly arrange and splice the obtained images into one image.
[0091] The calculation formula of the Mixup function is: x = λx i +(1 - λ)x j and y = λy i +(1 - λ)y j ; where, x represents the i-th image x in the training set using the Mixup functioni and the j-th image x j The generated new image, where y represents the new image generated by using the Mixup function on the i-th image x in the training set i and the j-th image x j The corresponding label of the generated new image, and λ ∈ [0,1] is a modulation factor following a Beta distribution with parameter α.
[0092] Step 7.3: Use the Adam optimizer and the stochastic gradient descent method to iteratively update the weights of the YOLO-V network parameters, and then verify the training effect on the validation set until the loss function EIOU converges, obtaining the trained YOLO-V neural network.
[0093] The EIOU loss function is as follows:
[0094]
[0095] Among them, IOU represents the intersection over union of the predicted bounding box and the ground truth bounding box, and ρ 2 (b, b gt ) represents the distance between b and b gt b and b gt respectively represent the centers of the predicted bounding box and the ground truth bounding box, c represents the diagonal length of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, w, h and w gt , h gt respectively represent the widths and heights of the predicted bounding box and the ground truth bounding box, C w and C h respectively represent the widths and heights of the minimum bounding rectangles of the predicted bounding box and the ground truth bounding box.
[0096] Step 8: Detect ships in the test set.
[0097] Step 8.1: Input the test set into the trained YOLO-V network and load the trained YOLO-V network parameter file.
[0098] Step 8.2: Perform ship detection on the images in the test set to obtain the center point position, width and height of the ship bounding box, and the confidence of the ship.
[0099] Step 8.3: Output the ship detection results and save them as a label file in txt format.
[0100] In the embodiments of the present invention, the K-means++ clustering algorithm is used to obtain more reasonable anchor box sizes by clustering the data set. The anchor box sizes of the first to fourth layers are set to {19, 28, 34, 22, 34, 48}, {53, 34, 24, 66, 71, 24}, {57, 59, 40, 87, 87, 41}, and {62, 87, 86, 64, 102, 100} respectively. Each detection layer performs a 1×1 convolution operation on the input image to expand the number of channels of the image. The expanded number of channels is 18. Then, the input image is divided into grids according to the set anchor box sizes, and then ships are detected.
Claims
1. A method for detecting ships in remote sensing images based on the YOLO-V network, characterized in that, The steps of this method are as follows: Step 1, construct the Input sub-network in the YOLO-V network: The Input sub-network is composed of a scaling layer and an enhancement layer connected in series. The scaling layer and the enhancement layer are implemented using the image adaptive scaling algorithm and the data augmentation algorithm respectively. The data augmentation algorithm is implemented using the Mosaic function and the Mixup function; Step 2, construct the Backbone sub-network of the YOLO-V network: Build an 11-layer Backbone sub-network, and its structure is in turn: Focus layer, first convolutional layer, first C3 layer, second convolutional layer, second C3 layer, third convolutional layer, third C3 layer, fourth convolutional layer, mixed spatial pyramid pooling layer, co-attention layer, fourth C3 layer; Step 3, construct the Neck sub-network in the YOLO-V network: Step 3.1, build an 11-layer Neck-A sub-network, and its structure is in turn: first convolutional layer, first upsampling layer, first splicing layer, first C3 layer, second convolutional layer, second upsampling layer, second splicing layer, second C3 layer, third convolutional layer, third upsampling layer, third splicing layer; Step 3.2, build a 10-layer Neck-B sub-network, and its structure is in turn: first C3 layer, first convolutional layer, first splicing layer, second C3 layer, second convolutional layer, second splicing layer, third C3 layer, third convolutional layer, third splicing layer, fourth C3 layer; Step 3.3, cascade the third splicing layer in the Neck-A sub-network with the first C3 layer in the Neck-B sub-network, cascade the second C3 layer in the Neck-A sub-network with the first splicing layer in the Neck-B sub-network, cascade the first C3 layer in the Neck-A sub-network with the second splicing layer in the Neck-B sub-network, and cascade the first convolutional layer in the Neck-A sub-network with the third splicing layer in the Neck-B sub-network to obtain the Neck sub-network; Step 4, construct the Head sub-network in the YOLO-V network: The Head sub-network is composed of a first detection layer, a second detection layer, a third detection layer and a fourth detection layer connected in parallel, and a detection layer with large-scale shallow feature maps in the Head sub-network is realized; Step 5, construct the YOLO-V network: Step 5.1, cascade the first C3 layer, the second C3 layer, and the third C3 layer in the Backbone sub-network with the first concatenation layer, the second concatenation layer, and the third concatenation layer in the Neck-A sub-network respectively, and cascade the fourth C3 layer with the first convolutional layer in the Neck-A sub-network; at the same time, cascade the second C3 layer and the third C3 layer in the Backbone sub-network with the first concatenation layer and the second concatenation layer in the Neck-B sub-network respectively to introduce skip connections in the Neck sub-network; cascade the third concatenation layer in the Neck-A sub-network with the first C3 in the Neck-B sub-network, and cascade the first convolutional layer, the first C3 layer, and the second C3 layer in the Neck-A sub-network with the third concatenation layer, the second concatenation layer, and the third concatenation layer in the Neck-B sub-network respectively; cascade the first C3 layer, the second C3 layer, the third C3 layer, and the fourth C3 layer in the Neck-B sub-network with the first detection layer, the second detection layer, the third detection layer, and the fourth detection layer in the Head sub-network respectively to complete the connection of the three sub-networks of the Backbone sub-network, the Neck sub-network, and the Head sub-network; Step 5.2, cascade the enhancement layer of the Input sub-network with the Focus layer of the Backbone sub-network to form the YOLO-V network; Step 6, generate the training set and the validation set: Step 7, train the YOLO-V network: Step 8, detect ships in remote sensing images.
2. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 1, wherein, In step 4, the detection layers are clustered using the K-means++ algorithm, and three anchor box sizes are set for each detection layer respectively as: {19, 28, 34, 22, 34, 48}, {53, 34, 24, 66, 71, 24}, {57, 59, 40, 87, 87, 41}, {62, 87, 86, 64, 102, 100}.
3. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 1, characterized in that, In step 7, the weight parameter file for training the YOLO-V network uses yolov5s.pt; both the classification loss function and the confidence loss function use the binary cross-entropy loss function; use EIOU as the bounding box loss function; input the training set and the validation set into the YOLO-V network in sequence, and through the image adaptive scaling algorithm, convert the size of the image input to the Input sub-network to 640×640×3, and enhance the image through the enhancement algorithm; The enhancement algorithm uses the Mosaic function and the Mixup function. The Mixup function adds two input images and their corresponding labels in proportion to generate a new image and the corresponding label.
4. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 3, characterized in that, The binary cross-entropy loss function is as follows: Loss=-(y·log e (y*)+(1-y)·log e (1-y*)) Among them, y* represents the probability that the network predicts the target as a ship, y represents the label of the target. If the target belongs to a ship, y takes the value of 1. If the target does not belong to a ship, y takes the value of 0.
5. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 3, wherein, The bounding box loss function is as follows: Among them, IOU represents the intersection over union of the predicted bounding box and the ground truth bounding box, and ρ 2 (b, b gt ) represents the distance between b and b gt , b and b gt represent the centers of the predicted bounding box and the ground truth bounding box respectively, c represents the diagonal length of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, w, h and w gt , h gt represent the widths and heights of the predicted bounding box and the ground truth bounding box respectively, C w and C h represent the widths and heights of the minimum bounding rectangles of the predicted bounding box and the ground truth bounding box respectively.
6. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 3, characterized in that, The steps of the image adaptive scaling algorithm are as follows: The first step, calculate the scaling ratio according to the size of the image input to the Input sub-network and the image size 640×640×3 required by the network; Step 2: Calculate the size of the scaled image based on the size of the image input to the network and the scaling ratio; Step 3: Pad 0 around the scaled image to transform the size of the image to 640×640×3.
7. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 3, wherein The Mosaic function randomly scales and randomly crops the four input images in sequence to obtain four images, and then randomly arranges and splices the obtained images into one image.
8. The method for detecting ships in remote sensing images based on the YOLO-V network according to claim 3, wherein, The Mixup function is as follows: x = λx i +(1 - λ)x j y = λy i +(1 - λ)y j Among them, x represents the new image generated by using the Mixup function on the i-th image x i and the j-th image x j in the training set, and y represents the new image generated by using the Mixup function on the i-th image x i and the j-th image x j in the training set. The label corresponding to the new image, and λ ∈ [0, 1] is a modulation factor that follows a Beta distribution with parameter α.
Citation Information
Patent Citations
Method for detecting ship target in remote-sensing image
CN111091095A
Remote sensing image marine ship identification system and method based on improved YOLOv4 algorithm
CN113920436A