Lightweight marine ship target detection method for edge device deployment

By building a lightweight YOLO-SMB network model, the problems of large computational complexity and storage space requirements for offshore ship target detection on edge devices are solved, and the detection accuracy and multi-scale target detection capabilities are improved while maintaining high efficiency and real-time performance.

CN120635418APending Publication Date: 2025-09-12DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510770401.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing deep learning target detection algorithms require large amounts of computation and storage space when deployed on edge devices, making them difficult to adapt. At the same time, the task of detecting target ships at sea faces the complexity and real-time requirements of multi-scale target detection. Existing methods find it difficult to achieve lightweight and efficient deployment while maintaining detection accuracy.

Method used

A lightweight YOLO-SMB network model is constructed, including the improved backbone network StarNets, the neck network YOLO11n, and the improved head network MBConv. Through feature extraction, fusion, and prediction, combined with the parameter-free attention mechanism SimaM and MBConv modules, the model structure is optimized to adapt to edge devices.

Benefits of technology

It significantly reduces the number of model parameters and computational complexity, while improving detection accuracy and real-time performance, achieving efficient marine ship target detection on edge devices, maintaining a detection accuracy of 91.3% and a real-time inference performance of 246FPS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635418A_ABST
    Figure CN120635418A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight marine ship target detection method for edge device deployment, and the method comprises the steps: obtaining a ship target data set, dividing the ship target data set according to a preset proportion, and obtaining a training set, a verification set and a test set; a YOLO-SMB network model for lightweight marine ship target detection is constructed; performing model training and model verification on the YOLO-SMB network model according to the training set and the verification set to obtain an optimal YOLO-SMB network model; and deploying the optimal YOLO-SMB network model to edge equipment, and detecting a marine ship target according to the test set. The method solves the problems that when complex marine ship target detection is carried out, in order to improve the detection precision of an existing model, the structure of the model is complex, the parameter quantity and the calculated quantity are increased, and actual deployment to edge equipment is difficult.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning computer vision technology, and in particular to a lightweight marine ship target detection method for edge device deployment. Background Art

[0002] Edge deployment refers to a technical architecture that moves computing, storage, and network resources to edge devices close to data sources or users, achieving low latency, high reliability, and privacy protection through localized processing. Currently, the main challenges faced in deploying deep learning models to edge devices include:

[0003] Computing resource limitations, power consumption limitations, real-time requirements, and model size. In the maritime sector, places like ports require real-time detection of ship types, numbers, and berthing status within the port to optimize berth allocation and cargo loading and unloading processes. Camera data needs to be processed locally to avoid reliance on cloud transmission and improve the real-time nature of detection. Another example is maritime search and rescue boats. Unmanned boats have limited computing resources, limited computing power and storage space, and when performing tasks, they need to limit the power consumption of the equipment as much as possible while maintaining reliable detection accuracy for real-time detection. Therefore, it is of great significance to design a method that is suitable for edge device deployment, has reliable accuracy, and can perform real-time detection, which can address the imbalance between performance and lightweight of existing methods.

[0004] Object detection is a key research area in computer vision, aiming to identify objects in images or videos and label their locations and categories. Traditional object detection has limited performance in complex backgrounds, large datasets, and real-time applications. Manual feature extraction is time-consuming and lacks versatility. Deep learning methods, through automatic feature extraction and end-to-end detection, learn more complex representations, significantly improving their adaptability to diverse and complex scenarios. They offer higher accuracy and robustness than traditional methods. Existing deep learning-based object detection algorithms fall into two main categories: the first is two-stage object detection algorithms, such as R-CNN, Fast R-CNN, and Faster R-CNN. These methods offer improved accuracy at the expense of speed. The second category is single-stage object detection algorithms, such as YOLO, SSD, and DETR. These methods offer a speed advantage and are more suitable for real-time object detection tasks.

[0005] However, these methods are computationally intensive and require significant storage space, making them difficult to deploy on edge devices. Furthermore, the diversity of ship types presents a significant challenge in maritime object detection, and the presence of multi-scale objects further complicates detection. Furthermore, the detection system must maintain high inference speed and accuracy at varying scales to avoid missed or false detections. Therefore, designing an efficient and lightweight method for maritime vessel object detection is crucial. Summary of the Invention

[0006] The present invention provides a lightweight marine ship target detection method for edge device deployment to overcome the above technical problems.

[0007] In order to achieve the above object, the technical solution of the present invention is:

[0008] A lightweight marine ship target detection method for edge device deployment includes the following steps:

[0009] S1: Acquire marine ship target images for target detection;

[0010] Perform a screening operation on the marine ship target images to remove duplicate / low-quality images, and annotate the filtered marine ship target images with ship types and bounding boxes to obtain a ship target dataset;

[0011] Divide the ship target data set into training set, validation set and test set according to the preset ratio;

[0012] S2: Build a YOLO-SMB network model for lightweight maritime ship target detection;

[0013] The YOLO-SMB network model includes an improved backbone network for extracting different-scale feature maps of various types of maritime ship targets; a neck network for fusing the different-scale feature maps of maritime ship targets to obtain fused features; and an improved head network for predicting maritime ship targets based on the fused features.

[0014] S3: Perform model training and model verification on the YOLO-SMB network model based on the training set and validation set to obtain the optimal YOLO-SMB network model;

[0015] The model training includes: performing feature extraction training on the marine ship target images in the training set based on the improved backbone network to obtain multi-scale marine ship target feature maps; performing feature fusion training on the marine ship target feature maps based on the improved neck network to obtain fused feature maps; and performing prediction training on marine ship targets based on the fused feature maps based on the improved head network.

[0016] Model validation includes: conducting validation tests on the trained YOLO-SMB network model through the validation set, and updating the gradient based on the back-propagation method to obtain the optimal parameter weights of the network model, so as to reconstruct the YOLO-SMB network model and obtain the optimal YOLO-SMB network model;

[0017] S4: Deploy the optimal YOLO-SMB network model to the edge device and detect maritime ship targets based on the test set.

[0018] Furthermore, the improved backbone network in S2 includes an input layer, a convolutional module, a StarNets network unit, an SPPF module, and a SimaM module connected in sequence;

[0019] The StarNets network unit includes a first-stage module, a second-stage module, a third-stage module, and a fourth-stage module connected in sequence. Each stage module is equipped with a convolutional layer with different convolution kernels and a StarBlock module. At the same time, any branch in the StarBlock module is added with a ReLU6 activation function.

[0020] The input layer is used to transmit sample images in the training set to the convolution module;

[0021] The convolution module is used to extract the geometric features of the marine ship from the output of the input layer to obtain a first feature map;

[0022] The first-stage module is used to perform a convolution operation on the first feature map through a convolution layer, and to obtain a first depth-separable convolution feature map by performing a star operation on the first feature map after the convolution operation of the StarBlock module; the second-stage module is used to obtain a second depth-separable convolution feature map based on the first depth-separable convolution feature map; the third-stage module is used to obtain a third depth-separable convolution feature map based on the second depth-separable convolution feature map; the fourth-stage module is used to obtain a fourth depth-separable convolution feature map based on the third depth-separable convolution feature map;

[0023] The SPPF module is used to perform a pooling operation on the fourth depth-separable convolutional feature map to obtain a pooling feature map;

[0024] The SimaM module is used to construct an attention mechanism energy function about feature extraction neurons to determine the weight value of the neurons, and obtain an attention feature map by performing pixel feature extraction on the pooled feature map;

[0025] The third depth-wise separable convolution feature map, the fourth depth-wise separable convolution feature map and the attention feature map are feature maps of different scales extracted for various types of maritime ship targets.

[0026] Furthermore, the method for performing star operation of the StarBlock module specifically includes:

[0027] The feature map input to the StarBlock module is subjected to a depthwise separable convolution operation through the DWconv module. The results of the separable convolution operation are subjected to the ConvBN convolution operation to obtain the first convolution feature map and the second convolution feature map.

[0028] and performing full connection operations on the first convolution feature map and the second convolution feature map respectively to obtain a first fully connected feature map and a second fully connected feature map;

[0029] Performing an activation operation on the first fully connected feature map through the ReLU6 activation function, and performing an element-by-element multiplication operation on the first fully connected feature map after the activation operation and the second fully connected feature map to obtain a fused feature map;

[0030] A full connection operation is performed on the fused feature map, and a depthwise separable convolution operation is performed on the fused feature map after the full connection operation to obtain a depthwise separable convolution feature map corresponding to the first stage module, the second stage module, the third stage module, or the fourth stage module.

[0031] Furthermore, a method for constructing an attention mechanism energy function for feature extraction neurons to determine the weight value of neurons specifically includes:

[0032] S100: Define the initial energy function of the neuron;

[0033] And the expression of the initial energy function is

[0034]

[0035] Where: e t The energy function of the neuron is e t (w t ,b t ,y,x i ) is a simplified form of t represents the weight of the current neuron; b t represents the bias of the current neuron; y represents the activation value of the current neuron; x i Represents other neurons in the same dimension channel except the current target neuron; y t represents the activation value of the target neuron; represents the linearly transformed representation of the target neuron and M represents the number of neurons in a dimensional channel; y o Represents the activation values ​​of other neurons; represents the linearly transformed representation of other neurons and t represents the current neuron;

[0036] S101: Using binary labels, that is, let y in the initial energy function o =-1,y t =1, and add a regularization term to rewrite the initial energy function;

[0037] And the expression of the initial energy function after rewriting is

[0038]

[0039] Where: λ represents the regularization coefficient used to prevent overfitting;

[0040] S102: b in the rewritten initial energy function t Taking the derivative, we can get:

[0041]

[0042] S103: Order Can get:

[0043]

[0044] According to formula (4), we can get:

[0045]

[0046] S104: w in the rewritten initial energy function t Taking the derivative, we can get:

[0047]

[0048] S105: Order Combined with formula (5), we can get:

[0049]

[0050] Where: σ t represents the variance of the current target neuron; μ t Represents the mean value of the current target neuron;

[0051] S106: Assuming that all neurons in the same dimensional channel obey the same distribution, obtain the global mean and global variance of the neurons;

[0052] And the expressions of global mean and global variance are

[0053]

[0054] Where: Represents the average value of all neurons except the current target neuron in the same dimension channel, that is, the global mean; Represents the variance of other neurons in the same dimension channel except the current target neuron, that is, the global variance;

[0055] S107: Update formula (5) and formula (7) according to formula (8) to obtain:

[0056]

[0057] And by substituting formula (9) into formula (2), we can get the optimal energy function, that is, the attention mechanism energy function, and the optimal energy function The expression is

[0058]

[0059] S108: According to the optimal energy function Get the weight of each neuron, its expression is

[0060]

[0061] Where: Represents the weight value of the current neuron; E represents the optimal energy function The energy value of ; X represents the pixel feature value in the input pooling feature map.

[0062] Furthermore, the neck network described in S2 for performing feature fusion on the different scale feature maps of the marine ship target to obtain fusion features is the YOLO11n neck network;

[0063] It includes a first concat splicing layer, a second concat splicing layer, a third concat splicing layer, a fourth concat splicing layer, a first upsample sampling layer, a second upsample sampling layer, a first conv convolutional layer, a second conv convolutional layer, a first c3k2 module, a second c3k2 module, a third c3k2 module and a fourth c3k2 module;

[0064] The first upsample sampling layer is used to perform an upsampling operation on the attention feature map output by the SimaM module to obtain a first sampling feature map; the first concat splicing layer is used to splice the first sampling feature map with the fourth depth-separable convolution feature map to obtain a first splicing feature map;

[0065] The first c3k2 module is used to perform convolution and pooling operations on the first spliced ​​feature map to obtain a first c3k2 feature map; the second upsample sampling layer is used to perform an upsampling operation on the first c3k2 feature map to obtain a second sampling feature map;

[0066] The second concat splicing layer is used to splice the third depth-separable convolutional feature map and the second sampling feature map to obtain a second spliced ​​feature map; the second c3k2 module is used to perform convolution and pooling operations on the second spliced ​​feature map to obtain a second c3k2 feature map;

[0067] The first conv convolution layer is used to perform a convolution operation on the second c3k2 feature map to obtain a first conv feature map; the third concat splicing layer is used to splice the first conv feature map and the first c3k2 feature map to obtain a third splicing feature map;

[0068] The third c3k2 module is used to perform convolution and pooling operations on the third spliced ​​feature map to obtain a third c3k2 feature map; the second conv convolution layer is used to perform convolution operations on the third c3k2 feature map to obtain a second conv feature map;

[0069] The fourth concat splicing layer is used to splice the second conv feature map and the attention feature map to obtain a fourth spliced ​​feature map; the fourth c3k2 module is used to perform convolution and pooling operations on the fourth spliced ​​feature map to obtain a fourth c3k2 feature map.

[0070] Furthermore, the method for obtaining the improved head network in S2 is:

[0071] Only the two convolutional layers used for bounding box regression detection in the YOLO11n detection head are replaced with MBConv modules, and the rest of the network structure is left unchanged.

[0072] Furthermore, the S3 specifically includes the following steps:

[0073] S31: Perform model training on the YOLO-SMB network model according to the training set to obtain the trained YOLO-SMB network model;

[0074] S32: Validate the trained YOLO-SMB network model using the validation set, and confirm whether the output of the trained YOLO-SMB network model converges based on the total loss function;

[0075] And the function value of the total loss function is the sum of the loss value of the classification loss function CLS_LOSS and the loss value of the bounding box regression loss function BOX_Loss:

[0076] If convergence is confirmed, the trained YOLO-SMB network model is used as the optimal YOLO-SMB network model;

[0077] Otherwise, based on the back propagation method, the model parameter weights of the trained YOLO-SMB network model are adaptively gradient updated, and step S31 is repeated until the optimal parameter weights of the network model are obtained to reconstruct the YOLO-SMB network model, thereby obtaining the optimal YOLO-SMB network model.

[0078] Beneficial effects: The present invention provides a lightweight maritime ship target detection method for edge device deployment, by constructing a YOLO-SMB network model for lightweight maritime ship target detection; and performing model training and model verification on the YOLO-SMB network model according to the training set and the verification set to obtain the optimal YOLO-SMB network model; it solves the problem that when detecting complex maritime ship targets, the existing model has a complex model structure, an increased number of parameters and calculations in order to improve the detection accuracy of the model, and is difficult to be actually deployed on edge equipment. The present invention can map a high-dimensional feature space in a low-dimensional manner through an improved backbone network, and can significantly improve the lightweight level of the network model. At the same time, it can dynamically quantify the importance of channels through the constructed energy function to strengthen the characteristic responses of key components at zero parameter cost, and can improve target positioning accuracy while maintaining computational efficiency; the improved head network can more easily capture the detailed features of multi-scale targets to reduce information loss, thereby improving the detection capability of multi-scale targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0080] Figure 1 This is a flow chart of the lightweight marine ship target detection method for edge device deployment according to the present invention;

[0081] Figure 2 Schematic diagram of the structure of the YOLO-SMB network model for lightweight marine ship target detection constructed in this embodiment;

[0082] Figure 3 Schematic diagram of the structure of the StarNets network unit in this embodiment;

[0083] Figure 4 : is a structural diagram of the lightweight detection head of the MBConv module in this embodiment;

[0084] Figure 5 Schematic diagram of the structure of the parameter-free attention mechanism SimaM module in this embodiment;

[0085] Figure 6 Schematic diagram of the structure of the StarBlock module in this embodiment;

[0086] Figure 7 Schematic diagram of the structure of the original YOLO11n detection head in this embodiment;

[0087] Figure 8 are sample images of various marine ship targets in this embodiment;

[0088] Figure 9 This is a technical block diagram of the lightweight maritime ship target detection method deployed to edge devices in this embodiment. DETAILED DESCRIPTION

[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0090] This embodiment provides a lightweight marine ship target detection method for edge device deployment, such as Figure 1 and Figure 9 As shown, the following steps are included:

[0091] S1: Acquire marine ship target images for target detection;

[0092] Perform a screening operation on the marine ship target images to remove duplicate / low-quality images, and annotate the filtered marine ship target images with ship types and bounding boxes to obtain a ship target dataset;

[0093] Divide the ship target data set into training set, validation set and test set according to the preset ratio;

[0094] Specifically, this embodiment selects 10 types of ships from the MODD and MVDD13 public data sets as the data set, and performs operations such as deduplication and removal of low-quality images on the data to construct the final maritime ship target detection data set. The filtered images are labeled with the LabelImg tool for target detection, and the bounding boxes and category labels of the maritime targets are marked, and finally a maritime ship target detection data set is obtained; the data set includes 10 types of maritime targets currently common in the maritime field, including: ferries, cargo ships, sailboats, speedboats, container ships, cruise ships, fishing boats, warships, etc. Figure 8 Some examples are shown, and then they are divided into training, validation, and test sets according to a 7:2:1 ratio. The images are uniformly resized to 640*640 pixels. The methods for deduplicating the data and removing low-quality images are well-known technical means and will not be elaborated on here.

[0095] S2: Build a YOLO-SMB network model for lightweight maritime ship target detection;

[0096] like Figure 2 As shown, the YOLO-SMB network model includes an improved backbone network for extracting different-scale feature maps of various types of maritime ship targets; a neck network for performing feature fusion on different-scale feature maps of maritime ship targets to obtain fusion features; and an improved head network for predicting maritime ship targets based on the fusion features;

[0097] In a specific embodiment, the improved backbone network includes an input layer, a convolution module, a StarNets network unit, an SPPF module, and a SimaM module connected in sequence;

[0098] like Figure 3 As shown, the StarNets network unit includes a first-stage module, a second-stage module, a third-stage module, and a fourth-stage module connected in sequence, wherein each stage module is provided with a convolutional layer with different convolution kernels and a StarBlock module, and any branch in the StarBlock module is added with a ReLU6 activation function;

[0099] The input layer is used to transmit sample images in the training set to the convolution module;

[0100] The convolution module is used to extract the geometric features of the marine ship from the output of the input layer to obtain a first feature map;

[0101] The first-stage module is used to perform a convolution operation on the first feature map through a convolution layer, and to obtain a first depth-separable convolution feature map by performing a star operation on the first feature map after the convolution operation of the StarBlock module; the second-stage module is used to obtain a second depth-separable convolution feature map based on the first depth-separable convolution feature map; the third-stage module is used to obtain a third depth-separable convolution feature map based on the second depth-separable convolution feature map; the fourth-stage module is used to obtain a fourth depth-separable convolution feature map based on the third depth-separable convolution feature map;

[0102] Among them Figure 6 As shown, the method for performing star operation of the StarBlock module specifically includes:

[0103] The feature map input to the StarBlock module is subjected to a depthwise separable convolution operation through the DWconv module; the results after the separable convolution operation are subjected to a ConvBN convolution operation to obtain a first convolution feature map and a second convolution feature map; and the first convolution feature map and the second convolution feature map are subjected to a full connection operation to obtain a first fully connected feature map and a second fully connected feature map; the first fully connected feature map is activated by the ReLU6 activation function, and the first fully connected feature map after the activation operation is multiplied element-wise with the second fully connected feature map to obtain a fused feature map; the fused feature map is subjected to a full connection operation, and a depthwise separable convolution operation is performed on the fused feature map after the full connection operation to obtain a depthwise separable convolution feature map corresponding to the first stage module or the second stage module or the third stage module or the fourth stage module; the SPPF module is used to perform a pooling operation on the fourth depthwise separable convolution feature map to obtain a pooling feature map;

[0104] In this embodiment, the backbone network Backbone of YOLO11n is reconstructed by using the StarNets network. The network structure of StarNets is a four-stage hierarchical architecture, which uses convolution layers for downsampling to reduce the resolution, and each stage consists of a convolution operation and a star module StarBlock. Inside the star module, the input undergoes a depth-separable convolution, and then the star operation, i.e., element-wise multiplication, is used to fuse the features of two linear transformations through element-by-element multiplication. Finally, after a depth-separable convolution, the star operation calculation within the star module does not introduce additional parameters, but only uses the product or weighted product between the input features; in the two branches of the star operation, only one is added with an activation function, and the conventional activation function gelu is replaced with the ReLU6 activation function. This design significantly reduces the number of model parameters and computational overhead, improves the deployment efficiency and inference speed of the model, and enhances the scalability and adaptability of the model, enabling it to flexibly cope with diverse hardware platforms and application scenarios, ensuring that efficient and reliable performance can be maintained in different deployment environments;

[0105] The SimaM module is used to construct an attention mechanism energy function about the feature extraction neuron to determine the weight value of the neuron, and obtain an attention feature map by performing pixel feature extraction on the pooled feature map; the third depth-separable convolution feature map, the fourth depth-separable convolution feature map and the attention feature map are the extracted feature maps of different scales of various types of maritime ship targets;

[0106] In a specific embodiment, Figure 5As shown in FIG, the method for determining the weight value of the neuron by replacing the C2PSA layer in the backbone network Backbone of the original YOLO11n with the parameter-free SimaM attention mechanism and constructing an attention mechanism energy function for the feature extraction neuron specifically includes:

[0107] S100: Define the initial energy function of the neuron;

[0108] And the expression of the initial energy function is

[0109]

[0110] Where: e t The energy function of the neuron is e t (w t ,b t ,y,x i ) is a simplified form of t represents the weight of the current neuron; b t represents the bias of the current neuron; y represents the activation value of the current neuron; x i Represents other neurons in the same dimension channel except the current target neuron; y t represents the activation value of the target neuron; represents the linearly transformed representation of the target neuron and M represents the number of neurons in a dimensional channel; y o Represents the activation values ​​of other neurons; represents the linearly transformed representation of other neurons and t represents the current neuron;

[0111] S101: In order to simplify the calculation in this embodiment, a binary label is used, that is, y in the initial energy function is set to o =-1,y t =1, and add a regularization term to rewrite the initial energy function;

[0112] And the expression of the initial energy function after rewriting is

[0113]

[0114] Where: λ represents the regularization coefficient used to prevent overfitting;

[0115] S102: b in the rewritten initial energy function t Taking the derivative, we can get:

[0116]

[0117] S103: Order Can get:

[0118]

[0119] According to formula (4), we can get:

[0120]

[0121] S104: w in the rewritten initial energy function t Taking the derivative, we can get:

[0122]

[0123] S105: Order Combined with formula (5), we can get:

[0124]

[0125] Where: σ t represents the variance of the current target neuron; μ t Represents the mean value of the current target neuron;

[0126] S106: Assuming that all neurons in the same dimensional channel obey the same distribution, obtain the global mean and global variance of the neurons;

[0127] And the expressions of global mean and global variance are

[0128]

[0129] Where: Represents the average value of all neurons except the current target neuron in the same dimension channel, that is, the global mean; Represents the variance of other neurons in the same dimension channel except the current target neuron, that is, the global variance;

[0130] S107: Update formula (5) and formula (7) according to formula (8) to obtain:

[0131]

[0132] And by substituting formula (9) into formula (2), we can get the optimal energy function, that is, the attention mechanism energy function, and the optimal energy function The expression is

[0133]

[0134] S108: According to the optimal energy function Get the weight of each neuron, its expression is

[0135]

[0136] Where: Represents the weight value of the current neuron; E represents the optimal energy function The energy value of ; X represents the pixel feature value in the input pooling feature map;

[0137] The SimaM module of this embodiment can adaptively learn and utilize the similarity information between targets, accurately calculate the similarity measure between features to determine the weights of different pixel features, so that the SimaM module can focus more specifically on important features related to the target detection task and adjust the importance of these features. Based on the principle that attention regulation in the mammalian brain is usually manifested as a gain effect of neuronal response, the SimaM module can use a scaling operator instead of addition to refine the features of marine targets; the energy value E can be used to combine all Cross-channel and spatial attention are combined, and a sigmoid activation function is added to prevent the E value from being too large to ensure the relative importance of each neuron. This calculation process does not require additional parameters and only requires a single element-level multiplication operation with less computational effort, making it more efficient than traditional attention mechanisms.

[0138] In a specific embodiment, the neck network for performing feature fusion on feature maps of different scales of marine ship targets to obtain fusion features is the YOLO11n neck network; it includes a first concat splicing layer, a second concat splicing layer, a third concat splicing layer, a fourth concat splicing layer, a first upsample sampling layer, a second upsample sampling layer, a first conv convolutional layer, a second conv convolutional layer, a first c3k2 module, a second c3k2 module, a third c3k2 module and a fourth c3k2 module; the first upsample sampling layer is used to upsample the attention feature map output by the SimaM module to obtain a first sampling feature map; the first concat splicing layer is used to splice the first sampling feature map and the fourth depth-separable convolutional feature map to obtain a first splicing feature map; the first c3k2 module is used to perform convolution and pooling operations on the first splicing feature map to obtain a first c3k2 feature map; the second upsample sampling layer is used to upsample the first c3k2 feature map to obtain a second sampling feature map; the second concat splicing layer is used to splice the first sampling feature map and the fourth depth-separable convolutional feature map to obtain a first splicing feature map; The three depth-separable convolution feature maps and the second sampling feature map are used to obtain the second spliced ​​feature map; the second c3k2 module is used to perform convolution and pooling operations on the second spliced ​​feature map to obtain the second c3k2 feature map; the first conv convolution layer is used to perform convolution operations on the second c3k2 feature map to obtain the first conv feature map; the third concat splicing layer is used to splice the first conv feature map and the first c3k2 feature map to obtain the third spliced ​​feature map; the third c3k2 module is used to perform convolution and pooling operations on the third spliced ​​feature map to obtain the third c3k2 feature map ; The second conv convolution layer is used to perform a convolution operation on the third c3k2 feature map to obtain a second conv feature map; the fourth concat splicing layer is used to splice the second conv feature map and the attention feature map to obtain a fourth splicing feature map; the fourth c3k2 module is used to perform convolution and pooling operations on the fourth splicing feature map to obtain a fourth c3k2 feature map. In this embodiment, the second c3k2 feature map, the third c3k2 feature map, and the fourth c3k2 feature map are transmitted to the improved detection head network of the improved YOLO11n detection head to achieve detection of marine targets;

[0139] In a specific embodiment, the method for obtaining the improved head network is:

[0140] like Figure 7 As shown, only the two convolutional layers used for bounding box regression detection in the YOLO11n detection head are replaced with MBConv modules, and the rest of the network structure is not processed;

[0141] This embodiment improves the detection head by using the MBConv module. The basic building block of the MBConv module is a bottleneck depth-separable convolution with residual, such as Figure 4 As shown in the figure, the internal operation of the module is that the input first passes through a pointwise convolution layer with a 1x1 kernel and a stride of 1 to increase the number of channels. This layer is responsible for constructing new features by computing linear combinations of the input channels, breaking the feature channel limit and providing more information. It then passes through a depthwise separable convolution layer, which performs lightweight filtering by applying a single convolution filter to each input channel, extracting spatial features with a very low parameter count. The channel attention mechanism (SE) module is integrated in the middle, which is an attention-based feature map operation. It includes compression: global average pooling compresses the spatial dimensions to generate channel description vectors; excitation: learning channel weights through fully connected layers to strengthen key feature channels. This operation is performed after the depthwise convolution and ends with a 1x1 pointwise convolution to restore the original channel dimension. It then passes through another pointwise convolution layer to reduce the output channel dimension to similar to the input channel dimension for feature fusion. Finally, an inverted residual is introduced, an operation similar to residual chaining, which uses a short-circuit method to add the input and the output after convolution, achieving feature reuse, reducing information loss, and alleviating the vanishing gradient problem. In this embodiment, two MBConv modules are used to replace the two convolutional layers originally used for bounding box regression in the YOLO11n detection head, reducing the parameters brought by the convolutional layers.

[0142] S3: Perform model training and model verification on the YOLO-SMB network model based on the training set and validation set to obtain the optimal YOLO-SMB network model;

[0143] The model training includes: performing feature extraction training on the marine ship target images in the training set based on the improved backbone network to obtain multi-scale marine ship target feature maps; performing feature fusion training on the marine ship target feature maps based on the improved neck network to obtain fused feature maps; and performing prediction training on marine ship targets based on the fused feature maps based on the improved head network.

[0144] Model validation includes: conducting validation tests on the trained YOLO-SMB network model through the validation set, and updating the gradient based on the back-propagation method to obtain the optimal parameter weights of the network model, so as to reconstruct the YOLO-SMB network model and obtain the optimal YOLO-SMB network model;

[0145] In a specific embodiment, the S3 specifically includes the following steps:

[0146] S31: Perform model training on the YOLO-SMB network model according to the training set to obtain the trained YOLO-SMB network model;

[0147] S32: Validate the trained YOLO-SMB network model using the validation set, and confirm whether the output of the trained YOLO-SMB network model converges based on the total loss function;

[0148] And the function value of the total loss function is the sum of the loss value of the classification loss function CLS_LOSS and the loss value of the bounding box regression loss function BOX_Loss:

[0149] In this embodiment, the YOLO-SMB network model is trained on a preset server to obtain the optimal model weights, which are saved as best.pt. During the model training process, the input is the image in the training set that has undergone the preprocessing and data enhancement. The final model output is the bounding box position of the complex marine target detected in the sample image of the training set, the ship target category label, and the possible confidence score. The optimal weights are obtained through the optimization algorithm in the training process, namely the back propagation method. In the loss function and optimizer selection, the category classification loss adopts CLS_Loss loss, and the bounding box regression loss adopts BOX_Loss loss. s, the optimizer selects the Adam optimizer based on the gradient descent algorithm. In each training iteration, data is input into the model for forward propagation and the total loss function value is calculated. Then, the gradient is calculated based on the total loss function value through backpropagation. Finally, the model parameters are updated according to the gradient. This process continues for multiple training rounds until the model converges. During the training process, the model is verified using the validation set to check the performance of the model on unseen data. The optimal weight is obtained based on the performance of the model on the validation set. When the performance of the model on the validation set stops improving or starts to decline, the training can be stopped and the corresponding model parameters are selected as the optimal weight.

[0150] If convergence is confirmed, the trained YOLO-SMB network model is used as the optimal YOLO-SMB network model; otherwise, based on the back propagation method, the model parameter weights of the trained YOLO-SMB network model are adaptively updated with gradients, and step S31 is repeated until the optimal parameter weights of the network model are obtained to reconstruct the YOLO-SMB network model, thereby obtaining the optimal YOLO-SMB network model;

[0151] S4: Deploy the optimal YOLO-SMB network model to the edge device, and detect the marine ship target according to the test set; the edge device is, for example, a high-performance edge computing device Jetson Xavier.

[0152] Compared with the prior art, the method described in this embodiment has the following beneficial effects:

[0153] The method described in this embodiment is improved based on YOLO11n. First, StarNets is used to reconstruct the backbone network, and multi-branch feature interaction is realized through a star topology structure. After the input image is reduced in dimensionality by depthwise separable convolution, radially connected 1×1 convolution kernels are used to aggregate multi-scale receptive field features, namely geometric features, and combined with jump connections to retain shallow details, it can map a high-dimensional feature space in a low-dimensional space, significantly improving the lightweight level compared to the CBS module of the standard YOLO11; second, the SimaM parameter-free attention mechanism module is used to replace the C2PSA layer in the original YOLO11 backbone network, and by constructing an energy function to dynamically quantize the importance of channels, the pixel feature response of key components is strengthened at zero parameter cost, which can improve target positioning accuracy while maintaining computational efficiency; third, the MBConv module is integrated into the YOLO11 detection head, and depthwise separable convolution is used to reduce parameters and computational complexity. The SE attention mechanism in the MBConv module is used to enhance the feature extraction capability of maritime targets, making it easier to capture the detailed features of multi-scale targets and reduce information loss, thereby improving the detection capability of multi-scale targets. In addition, after training and testing, the model finally compressed the parameter amount to 1.3M and the computational amount to 3.6G while ensuring 91.3% of the map index. As shown in Table 1, Table 1 is a data comparison of the model performance and lightweight index of the YOLO-SMB model and the mainstream target detection model. For the basic model YOLO11n, while significantly reducing the model parameter amount and computational amount by 46.9%, the final model achieved 91.3% detection accuracy and 246FPS real-time inference performance, completely surpassing the YOLOv5, YOLOv8, and YOLO11n models in terms of performance and lightweight. Compared with several other models, the performance is It lags behind GhostNetV3 by only one percentage point, but its model parameters, computational complexity, and model weight files are significantly better than those of the GhostNetV3 model. Compared with EfficintNetv1, MobileNetv, MobileNetv2, and GhostNetv1, its performance is only 0.1-0.3 percentage points lower. In summary, the YOLO-SMB network model is completely ahead of other models in terms of lightweight indicators. Without significantly reducing model performance, it achieves a balance between model lightweightness and performance, providing the best lightweight solution for edge deployment and an excellent choice for edge device deployment.

[0154] Table 1. Data comparison of model performance and lightweight indicators

[0155]

[0156]

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lightweight marine ship target detection method for edge device deployment, characterized in that: The specific steps include: S1: Acquire marine ship target images for target detection; Perform a screening operation on the marine ship target images to remove duplicate / low-quality images, and annotate the filtered marine ship target images with ship types and bounding boxes to obtain a ship target dataset; Divide the ship target data set into training set, validation set and test set according to the preset ratio; S2: Build a YOLO-SMB network model for lightweight maritime ship target detection; The YOLO-SMB network model includes an improved backbone network for extracting different-scale feature maps of various types of maritime ship targets; a neck network for fusing the different-scale feature maps of maritime ship targets to obtain fused features; and an improved head network for predicting maritime ship targets based on the fused features. S3: Perform model training and model verification on the YOLO-SMB network model based on the training set and validation set to obtain the optimal YOLO-SMB network model; The model training includes: performing feature extraction training on the marine ship target images in the training set based on the improved backbone network to obtain multi-scale marine ship target feature maps; performing feature fusion training on the marine ship target feature maps based on the improved neck network to obtain fused feature maps; and performing prediction training on marine ship targets based on the fused feature maps based on the improved head network. Model validation includes: conducting validation tests on the trained YOLO-SMB network model through the validation set, and updating the gradient based on the back-propagation method to obtain the optimal parameter weights of the network model, so as to reconstruct the YOLO-SMB network model and obtain the optimal YOLO-SMB network model; S4: Deploy the optimal YOLO-SMB network model to the edge device and detect maritime ship targets based on the test set.

2. A lightweight marine ship target detection method for edge device deployment according to claim 1, characterized in that: The improved backbone network in S2 includes an input layer, a convolution module, a StarNets network unit, an SPPF module, and a SimaM module connected in sequence; The StarNets network unit includes a first-stage module, a second-stage module, a third-stage module, and a fourth-stage module connected in sequence. Each stage module is equipped with a convolutional layer with different convolution kernels and a StarBlock module. At the same time, any branch in the StarBlock module is added with a ReLU6 activation function. The input layer is used to transmit sample images in the training set to the convolution module; The convolution module is used to extract the geometric features of the marine ship from the output of the input layer to obtain a first feature map; The first-stage module is used to perform a convolution operation on the first feature map through a convolution layer, and to obtain a first depth-separable convolution feature map by performing a star operation on the first feature map after the convolution operation of the StarBlock module; the second-stage module is used to obtain a second depth-separable convolution feature map based on the first depth-separable convolution feature map; the third-stage module is used to obtain a third depth-separable convolution feature map based on the second depth-separable convolution feature map; the fourth-stage module is used to obtain a fourth depth-separable convolution feature map based on the third depth-separable convolution feature map; The SPPF module is used to perform a pooling operation on the fourth depth-separable convolutional feature map to obtain a pooling feature map; The SimaM module is used to construct an attention mechanism energy function about feature extraction neurons to determine the weight value of the neurons, and obtain an attention feature map by performing pixel feature extraction on the pooled feature map; The third depth-wise separable convolution feature map, the fourth depth-wise separable convolution feature map and the attention feature map are feature maps of different scales extracted for various types of maritime ship targets.

3. A lightweight marine ship target detection method for edge device deployment according to claim 2, characterized in that: The method for performing star operation of the StarBlock module specifically includes: The feature map input to the StarBlock module is subjected to a depthwise separable convolution operation through the DWconv module. The results of the separable convolution operation are subjected to the ConvBN convolution operation to obtain the first convolution feature map and the second convolution feature map. and performing full connection operations on the first convolution feature map and the second convolution feature map respectively to obtain a first fully connected feature map and a second fully connected feature map; Performing an activation operation on the first fully connected feature map through the ReLU6 activation function, and performing an element-by-element multiplication operation on the first fully connected feature map after the activation operation and the second fully connected feature map to obtain a fused feature map; A full connection operation is performed on the fused feature map, and a depthwise separable convolution operation is performed on the fused feature map after the full connection operation to obtain a depthwise separable convolution feature map corresponding to the first stage module, the second stage module, the third stage module, or the fourth stage module.

4. A lightweight marine ship target detection method for edge device deployment according to claim 3, characterized in that: The method of constructing an attention mechanism energy function about feature extraction neurons to determine the weight value of neurons includes: S100: Define the initial energy function of the neuron; And the expression of the initial energy function is Where: e t The energy function of the neuron is e t (w t ,b t ,y,x i ) is a simplified form of t represents the weight of the current neuron; b t represents the bias of the current neuron; y represents the activation value of the current neuron; x i Represents other neurons in the same dimension channel except the current target neuron; y t represents the activation value of the target neuron; represents the linearly transformed representation of the target neuron and M represents the number of neurons in a dimensional channel; y o Represents the activation values ​​of other neurons; represents the linearly transformed representation of other neurons and t represents the current neuron; S101: Using binary labels, that is, let y in the initial energy function o =-1,y t =1, and add a regularization term to rewrite the initial energy function; And the expression of the initial energy function after rewriting is Where: λ represents the regularization coefficient used to prevent overfitting; S102: Derivative bt in the rewritten initial energy function, we can get: S103: Order Can get: And according to formula (4), we can get: S104: w in the rewritten initial energy function t Taking the derivative, we can get: S105: Order Combined with formula (5), we can get: Where: σ t represents the variance of the current target neuron; μ t Represents the mean value of the current target neuron; S106: Assuming that all neurons in the same dimensional channel obey the same distribution, obtain the global mean and global variance of the neurons; And the expressions of global mean and global variance are Where: Represents the average value of all neurons except the current target neuron in the same dimension channel, that is, the global mean; Represents the variance of other neurons in the same dimension channel except the current target neuron, that is, the global variance; S107: Update formula (5) and formula (7) according to formula (8) to obtain: And by substituting formula (9) into formula (2), we can get the optimal energy function, that is, the attention mechanism energy function, and the optimal energy function The expression is S108: According to the optimal energy function Get the weight of each neuron, its expression is Where: Represents the weight value of the current neuron; E represents the optimal energy function The energy value of ; X represents the pixel feature value in the input pooling feature map.

5. A lightweight marine ship target detection method for edge device deployment according to claim 4, characterized in that: The neck network described in S2 for performing feature fusion on the different scale feature maps of the marine ship target to obtain fusion features is the YOLO11n neck network; It includes a first concat splicing layer, a second concat splicing layer, a third concat splicing layer, a fourth concat splicing layer, a first upsample sampling layer, a second upsample sampling layer, a first conv convolutional layer, a second conv convolutional layer, a first c3k2 module, a second c3k2 module, a third c3k2 module and a fourth c3k2 module; The first upsample sampling layer is used to perform an upsampling operation on the attention feature map output by the SimaM module to obtain a first sampling feature map; the first concat splicing layer is used to splice the first sampling feature map with the fourth depth-separable convolution feature map to obtain a first splicing feature map; The first c3k2 module is used to perform convolution and pooling operations on the first spliced ​​feature map to obtain a first c3k2 feature map; The second upsample sampling layer is used to perform an upsampling operation on the first c3k2 feature map to obtain a second sampling feature map; The second concat splicing layer is used to splice the third depth-separable convolutional feature map and the second sampling feature map to obtain a second spliced ​​feature map; the second c3k2 module is used to perform convolution and pooling operations on the second spliced ​​feature map to obtain a second c3k2 feature map; The first conv convolution layer is used to perform a convolution operation on the second c3k2 feature map to obtain the first conv feature map; The third concat splicing layer is used to splice the first conv feature map and the first c3k2 feature map to obtain a third splicing feature map; The third c3k2 module is used to perform convolution and pooling operations on the third spliced ​​feature map to obtain a third c3k2 feature map; The second conv convolution layer is used to perform a convolution operation on the third c3k2 feature map to obtain the second conv feature map; The fourth concat splicing layer is used to splice the second conv feature map and the attention feature map to obtain a fourth spliced ​​feature map; the fourth c3k2 module is used to perform convolution and pooling operations on the fourth spliced ​​feature map to obtain a fourth c3k2 feature map.

6. A lightweight marine ship target detection method for edge device deployment according to claim 5, characterized in that: The method for obtaining the improved head network in S2 is: Only the two convolutional layers used for bounding box regression detection in the YOLO11n detection head are replaced with MBConv modules, and the rest of the network structure is left unchanged.

7. A lightweight marine ship target detection method for edge device deployment according to claim 6, characterized in that: The S3 specifically includes the following steps: S31: Perform model training on the YOLO-SMB network model according to the training set to obtain the trained YOLO-SMB network model; S32: Validate the trained YOLO-SMB network model using the validation set, and confirm whether the output of the trained YOLO-SMB network model converges based on the total loss function; And the function value of the total loss function is the sum of the loss value of the classification loss function CLS_LOSS and the loss value of the bounding box regression loss function BOX_Loss: If convergence is confirmed, the trained YOLO-SMB network model is used as the optimal YOLO-SMB network model; Otherwise, based on the back propagation method, the model parameter weights of the trained YOLO-SMB network model are adaptively gradient updated, and step S31 is repeated until the optimal parameter weights of the network model are obtained to reconstruct the YOLO-SMB network model, thereby obtaining the optimal YOLO-SMB network model.

Citation Information

Cited By

  • Electric power fitting detection method and system for data center park power transmission line inspection

    CN121330451A

  • Power fitting detection method and system for data center park transmission line inspection

    CN121330451B

  • Lightweight context aware network-based ship heaving motion prediction method and system

    CN121705670A