Ship detection method

Through the DE-YOLOv8 network model, combined with deformable convolution and centralized feature pyramid, the problems of complex background and target deformation in ship detection are solved, and high-precision and efficient detection effects are achieved.

CN120088741APending Publication Date: 2025-06-03CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411309833.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing ship detection algorithms face complex background interference, target size diversity and perspective changes in complex marine environments, and it is difficult to take into account the multi-scale feature extraction and the deformation adaptability of targets.

Method used

A ship detection method is proposed, using the DE-YOLOv8 network model, including backbone network, feature enhancement network and detection head. The backbone network introduces deformable convolution through the D2_C2f module, the feature enhancement network adopts the idea of ​​path aggregation network and feature pyramid network, and the detection head adopts the decoupling head design.

Benefits of technology

It improves detection accuracy and operation speed, enhances adaptability to complex backgrounds, diverse target forms and perspective changes, reduces the probability of false detection and missed detection, and is suitable for industrial application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088741A_ABST
    Figure CN120088741A_ABST
Patent Text Reader

Abstract

The invention provides a ship detection method. The ship detection method comprises the following steps: S1, constructing a data set for ship detection; s2, a DE-YOLOv8 network model is constructed; s3, a DE-YOLOv8 network is trained; and S4, detecting the target to be detected. The DE-YOLOv8 network model is composed of a backbone network, a feature enhancement network and a detection head. Wherein the variable convolution v2 is embedded in the backbone network, a D2C2f module is provided, and the idea of CSP (compact and separate) is maintained at the same time. The design of the D2C2f module fully considers the balance between the calculation efficiency and the performance, through the feature extraction process of the compression stage and the expansion stage, the feature enhancement network part adopts the thought of a path aggregation network-feature pyramid network, and structural optimization is carried out. Meanwhile, a centralized feature pyramid is introduced, and important corner region semantic information easy to ignore in an input image is effectively captured. A detection head adopts a Decoupled-Head thought, a regression branch and a classification branch are separated, and the problems of low precision, poor real-time performance and the like when a ship target detection task is completed by an existing target detection method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ship detection, and particularly relates to a ship detection method. Background Art

[0002] In the field of ship detection, traditional object detection algorithms usually rely on fixed feature extraction methods and convolutional neural network structures. Although these methods perform well in general object detection tasks, they still face some challenges when performing ship detection in complex marine environments. Specifically, ship detection in a marine background faces the following problems: 1. Complex background interference: The marine environment has strong interference factors, such as waves, clouds, shadows, etc. These complex backgrounds are prone to confusing with ship targets, resulting in false detection or missed detection; 2. Diversity of target sizes: The sizes and shapes of ships vary greatly, from small speedboats to large cargo ships. The detection algorithm needs to have good multi-scale feature extraction capabilities to accurately identify targets of different sizes. 3. Perspective changes and deformations: Due to the shooting angle and the movement of the ship itself, the shape, size, and position of the ship will change greatly in the image, which requires the detection algorithm to have a certain degree of robustness and adaptability.

[0003] To solve the above problems, researchers have proposed various improvement methods, including improving the feature pyramid structure to enhance multi-scale feature extraction capabilities and introducing adaptive convolution to improve the detection effect of deformed targets. However, most of the existing methods only focus on improvements in a single aspect and are difficult to simultaneously take into account multi-scale feature extraction and the deformation adaptability of targets. Summary of the Invention

[0004] The purpose of the present invention is to solve the technical problems in the above background, and propose a ship detection method, including the following steps:

[0005] S1. Construct a ship detection data set;

[0006] S2. Construct a DE-YOLOv8 network model;

[0007] S3. Train the DE-YOLOv8 network;

[0008] S4. Detect the target to be detected.

[0009] In a preferred solution, the DE-YOLOv8 network model includes a backbone network, a feature enhancement network, and a detection head;

[0010] The backbone network includes a D2_C2f module;

[0011] The D2_C2f module adaptively adjusts the shape of the convolution kernel and adjusts the position of the convolution kernel by introducing an offset;

[0012] The feature enhancement network adopts the ideas of the path aggregation network and the feature pyramid network to construct a centralized feature pyramid;

[0013] The detection head part adopts the design idea of the decoupled head to separate the regression branch and the classification branch.

[0014] In the preferred solution, the network structure of deformable convolution V2 mainly includes an Offset prediction module, a Modulation Scale module, and a deformable convolution operation;

[0015] The Offset prediction module is used to predict the offset of each convolution kernel position;

[0016] The Modulation Scale module introduces a modulation scalar based on the offset to adjust the weight of each sampling position;

[0017] The deformable convolution operation extracts features at the new sampling positions according to the predicted offset and modulation scalar to complete the convolution operation.

[0018] In the preferred solution, the calculation process of the deformable convolution includes the calculation of the offset and the convolution operation.

[0019] In the preferred solution, the calculation of the offset is specifically as follows: Given an input feature map x, the offset Δ is calculated through a convolutional layer p :

[0020] Δ p = W offset * x;

[0021] Among them, Δ p is the offset of each convolution kernel position, and * represents the convolution operation.

[0022] In the preferred solution, the calculation of the modulation scalar is specifically as follows:

[0023] The modulation scalar m is calculated by another convolutional layer W modulation to control the weight of each sampling point:

[0024] m = σ(W modulation * x);

[0025] Among them, σ(·) is the Sigmoid function, which is used to limit the modulation scalar between [0,1].

[0026] In the preferred solution, the final deformable convolution operation is:

[0027]

[0028] where p 0 is the center position of the current convolution kernel, pk is the position of the k-th sampling point of the conventional convolution operation, Δp k is the corresponding offset, w k is the weight of the convolution kernel, m k is the modulation scalar.

[0029] In a preferred solution, step S4 is specifically: input the picture containing the target to be detected into the trained DE-YOLOv8 network, and output the detection results of the category of the target to be detected in the picture and the position of each circumscribed rectangle where the target is located.

[0030] The beneficial effects of the present invention are:

[0031] (1) The algorithm model of the present invention optimizes the YOLOv8 structure, introduces deformable convolution and a centralized feature pyramid, not only improves the detection accuracy, but also maintains a high running speed. This high efficiency is very suitable for industrial application scenarios that require fast and real-time response, such as autonomous driving, intelligent monitoring, etc.

[0032] (2) Target detection in industrial environments often faces complex backgrounds, different lighting conditions, and diverse target shapes. The present invention improves the adaptability to these complex situations by enhancing feature extraction and the ability to handle target deformation, reduces the probability of false detection and missed detection, and ensures stable operation in a changing industrial environment.

[0033] (3) In industrial applications, the ease of use and deployment cost of the algorithm are very crucial. The algorithm of the present invention is based on the improved YOLOv8 architecture, maintains a lightweight design, and is compatible with existing deep learning frameworks. It can be easily integrated into existing industrial systems, reduces the deployment difficulty and cost, and helps with large-scale application promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the overall framework diagram of the DE-YOLOv8 network proposed by the present invention.

[0035] Figure 2 is the framework diagram of the Dv2_C2f module.

[0036] Figure 3 is the framework diagram of the centralized feature pyramid.

[0037] Figure 4 is the schematic diagram of the deformable convolution. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present invention will be further described below with reference to the accompanying drawings.

[0039] Such as Figures 1 to 4As shown, a ship detection method based on a centralized feature pyramid and deformable convolution specifically includes the following steps:

[0040] Step 1: Construct a ship detection dataset.

[0041] Step 2: Construct a DE-YOLOv8 network model;

[0042] The DE-YOLOv8 network model described in the present invention is composed of a backbone network, a feature enhancement network, and a detection head. Among them, the backbone network embeds deformable convolution v2 and proposes a D2_C2f module. This improvement makes the network more lightweight, better adapts to the deformation and pose changes of the target, and at the same time maintains the CSP (compact and separated) idea. The design of the D2_C2f module fully considers the balance between computational efficiency and performance. Through the feature extraction process in the compression stage and the expansion stage, it not only ensures the effectiveness of feature extraction but also reduces the computational complexity. The feature enhancement network part adopts the idea of PAN-FPN (Path Aggregation Network - Feature Pyramid Network) and conducts structural optimization. At the same time, a centralized feature pyramid is introduced to effectively capture the semantic information of important corner regions that are easily overlooked in the input image. The detection head adopts the Decoupled-Head idea to separate the regression branch and the classification branch. This design makes the training and inference of the network more efficient.

[0043] The backbone network is specifically manifested as:

[0044] Responsible for extracting features from the input image. The backbone network of DE-YOLOv8 uses C2f and D2_C2f modules as basic building units, which enhance the model performance by optimizing the gradient flow. Starting from the input image, the backbone network gradually extracts high-level features of the image through a series of convolutional layers and C2f modules. These feature maps are then sent to the neck network for further fusion and processing.

[0045] The neck network is specifically manifested as:

[0046] Responsible for fusing feature maps of different scales to enhance the multi-scale detection ability of the model. The neck network of DE-YOLOv8 adopts a structure similar to PAN-FPN of YOLOv5, namely Path Aggregation Network (PANet). PANet fuses feature maps of different scales through a bottom-up path and a top-down path, realizing cross-scale information transmission. Specifically, the output feature maps of the backbone network are subjected to multi-scale feature extraction through the SPP (Spatial Pyramid Pooling) structure and then feature fusion is performed through PANet.

[0047] The detection head is specifically manifested as:

[0048] Responsible for converting the fused feature map into the final detection result. DE-YOLOv8 adopts a Decoupled Head structure similar to YOLOX, separating the regression branch and the classification branch. This design helps to improve the convergence speed and detection effect of the model. In the detection head part, the feature map undergoes a series of convolutional operations and upsampling operations to generate detection layers of multiple scales. Each detection layer contains a regression branch and a classification branch, which are used to predict the bounding box and category of the target respectively.

[0049] Specifically, the centralized feature pyramid is as follows:

[0050] The centralized feature pyramid selects the Explicit Visual Center Block (EVCBlock) in CFPNet as the feature fusion stage for multi-scale feature fusion. The EVCBlock is a key module that makes up CFPNet. It consists of a lightweight multi-layer perceptron (MLP) and a learnable visual center (LVC) mechanism. The lightweight MLP architecture can capture the deepest global information, and the parallel learnable visual center mechanism LVC aggregates the local key region information of the input image and concatenates the MLP and LVC along the channel dimension as the output of the EVCBlock, enabling the model to effectively distinguish the feature representations of different scale targets while reducing the computational complexity. Adding CFPNet to the feature fusion network, where the EVCBlock can effectively capture the semantic information of important corner regions that are easily overlooked in the input image. It can adapt to the representation of abstract shallow features in the deep model features and achieve comprehensive and differentiated feature representations.

[0051] Specifically, the deformable convolution is as follows:

[0052] The present invention uses deformable convolution v2. Deformable Convolution (DCN) is a convolution operation proposed to enhance the ability of traditional convolutional networks to handle complex visual tasks such as irregular deformations and geometric transformations. Deformable Convolution v2 is an improved version of the original deformable convolution, with further optimizations in structure and function.

[0053] In traditional convolution operations, the sampling points of the convolution kernel are fixed, usually sampling the corresponding region of the input feature map with a regular grid (such as a 3×33\times 33×3 matrix). While deformable convolution makes the sampling position of the convolution kernel flexibly change within a certain range by introducing offsets, thus adapting to the deformation and geometric transformation of the target.

[0054] In the V1 version, the convolutional kernel adjusts the sampling positions through the learned offsets, enabling the convolutional kernel to better capture irregular deformation features. Deformable Convolution V2 further introduces the learned weights and combines multi-scale features, thus enhancing the flexibility and expressive power of the convolutional operation.

[0055] The network structure of Deformable Convolution V2 mainly includes the following modules:

[0056] (1) Offset prediction module: Predicts the offsets for each convolutional kernel position. Similar to the V1 version, but the offset prediction is more efficient in V2.

[0057] (2) Modulation Scale module: Introduces a modulation scalar on the basis of the offsets, which is used to adjust the weights of each sampling position, so as to more flexibly express the feature importance of different positions.

[0058] (3) Deformable convolution operation: According to the predicted offsets and modulation scalars, the convolution operation extracts features at the new sampling positions to complete the convolution operation.

[0059] The calculation process of the deformable convolution can be divided into two steps: the calculation of the offsets and the convolution operation.

[0060] (1) Calculation of the offsets

[0061] Given an input feature map x, the offsets Δ are calculated through a convolutional layer p :

[0062] Δ p = W offset * x

[0063] where Δ p is the offset for each convolutional kernel position, and * represents the convolution operation.

[0064] (2) Calculation of the modulation scalar

[0065] The modulation scalar m is calculated by another convolutional layer W modulation and is used to control the weights of each sampling point:

[0066] m = σ(W modulation * x)

[0067] where σ(·) is the Sigmoid function, which is used to limit the modulation scalar between [0, 1].

[0068] (3) Calculation of the deformable convolution

[0069] The final deformable convolution operation can be expressed as:

[0070]

[0071] where p 0 is the center position of the current convolution kernel, p k is the position of the k-th sampling point in the conventional convolution operation, Δp k is the corresponding offset, w k is the weight of the convolution kernel, m k is the modulation scalar.

[0072] Step 3: Train the DE-YOLOv8 network

[0073] Step 4: Detect the target to be detected

[0074] Input the image containing the target to be detected into the trained network, and output the detection results of the category of the target to be detected in the image and the position of each bounding rectangle where the target is located.

[0075] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A ship detection method, characterized in that: The following steps are involved: S1. Build a dataset for ship detection; S2. Build the DE-YOLOv8 network model; S3, train the DE-YOLOv8 network; S4. Detect the target to be detected.

2. A ship detection method according to claim 1, characterized in that: The DE-YOLOv8 network model includes a backbone network, a feature enhancement network, and a detection head; The backbone network includes D2_C2f modules; The D2_C2f module adaptively adjusts the shape of the convolution kernel and adjusts the position of the convolution kernel by introducing an offset; The feature enhancement network adopts the ideas of path aggregation network and feature pyramid network to construct a centralized feature pyramid; The detection head adopts the design idea of ​​decoupling head to separate the regression branch and the classification branch.

3. A ship detection method according to claim 1, characterized in that: The network structure of Deformable Convolution V2 mainly includes Offset prediction module, Modulation Scale module and Deformable Convolution operation; The Offset prediction module is used to predict the offset of each convolution kernel position; The Modulation Scale module introduces a modulation scalar based on the offset to adjust the weight of each sampling position; The deformable convolution operation performs feature extraction at the new sampling position based on the predicted offset and modulation scalar to complete the convolution operation.

4. A ship detection method according to claim 1, characterized in that: The calculation process of deformable convolution includes the calculation of offset and convolution operation.

5. A ship detection method according to claim 4, characterized in that: the offset The calculation is as follows: given an input feature map x, the offset Δ is calculated through a convolutional layer p : Δ p =W offset *x; Among them, Δ p is the offset of each convolution kernel position, and * indicates the convolution operation.

6. A ship detection method according to claim 4, characterized in that: The calculation of the modulation scalar is as follows: The modulation scalar m is obtained by another convolutional layer W modulation Calculated to control the weight of each sampling point: m=σ(W modulation *x); Here, σ(·) is the Sigmoid function, which is used to constrain the modulation scalar to be between [0, 1].

7. A ship detection method according to claim 6, characterized in that: The final deformable convolution operation is: Where p0 is the center position of the current convolution kernel, p k is the kth sampling point position of the conventional convolution operation, Δp k is the corresponding offset, w k is the weight of the convolution kernel, m k is the modulation scalar.

8. A ship detection method according to claim 1, characterized in that: Step S4 is specifically as follows: inputting a picture containing a target to be detected into the trained DE-YOLOv8 network, and outputting the detection results of the category of the target to be detected in the picture and the position of each circumscribed rectangular frame where the target is located.