Improved YOLOv8 underwater image target detection method, device, medium and program

By introducing a fine-grained feature extraction module in YOLOv8, the problem of insufficient detection accuracy and robustness of YOLOv8 in underwater environments is solved, the ability to identify small targets and complex backgrounds is improved, and efficient underwater target detection is achieved.

CN120279400APending Publication Date: 2025-07-08NANJING VOCATIONAL UNIV OF IND TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510291747.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing YOLOv8 has problems in underwater environments with insufficient detection capabilities for small objects, poor robustness of lighting changes, lack of long-distance dependency modeling capabilities and insufficient fusion of multi-scale features, resulting in insufficient detection accuracy and robustness.

Method used

On the basis of YOLOv8, the fine-grained feature extraction module is introduced, including the visual state space model, the local feature extraction module and the multi-scale cross-attention fusion module. The multi-scale cross-attention fusion module dynamically weighted the local and global features of different scales, and combined with depth separation convolution and SiLU activation functions to improve feature extraction capabilities.

Benefits of technology

It significantly improves the accuracy and robustness of underwater image object detection, especially the detection ability of tiny targets in low-visibility scenarios, and enhances the adaptability and noise resistance to multi-scale targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279400A_ABST
    Figure CN120279400A_ABST
Patent Text Reader

Abstract

The invention provides an improved YOLOv8 underwater image target detection method, device, medium and program, a fine-grained feature extraction module is arranged to replace a C2f module on the basis of YOLOv8, the fine-grained feature extraction module comprises a visual state space model and a local feature extraction module which are arranged in parallel, feature distribution is adjusted through linear transformation in the visual state space model, and the local feature extraction module is used for extracting local features in the visual state space model. A multi-scale space context is extracted in combination with depth separable convolution and a SiLU activation function; in the local feature extraction module, detail information of a target is focused through a convolution kernel channel attention mechanism, and noise interference is suppressed; the two outputs are subjected to element-by-element addition fusion and then enter a multi-scale cross attention fusion module, local and global features of different scales are fused in the multi-scale cross attention fusion module through dynamic weighting, and fine-grained features are output. The improved YOLOv8 is used for underwater target detection, and the identification capability of a fuzzy target and a complex background in underwater image target detection can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and deep learning, and particularly relates to an improved YOLOv8 underwater image target detection method, device, medium, and program. Background Art

[0002] Water bodies cover approximately 71% of the Earth's surface, playing a crucial role not only in regulating the global climate and absorbing carbon dioxide but also in containing rich natural resources. The health of the marine ecosystem has an important impact on the global environment. Therefore, efficient and accurate underwater target detection technologies have wide application values in fields such as marine biological monitoring, resource exploration, and underwater robot navigation. However, due to the complexity of the underwater environment, traditional computer vision technologies still face many challenges in underwater target detection tasks.

[0003] Firstly, the underwater lighting conditions are complex. The propagation of light in water is affected by factors such as scattering, absorption, and refraction, making underwater images often exhibit problems such as low contrast, blurriness, and color deviation, resulting in traditional target detection algorithms being difficult to extract effective features. Secondly, underwater targets often have characteristics such as small scale, complex morphology, and dense distribution. Many small targets (such as corals, marine organisms, shipwreck remains, etc.) are easily overlooked or misidentified in traditional target detection algorithms, leading to a decrease in detection accuracy. In addition, the noise interference in the underwater environment is severe, including water flow disturbance, plankton influence, etc., further increasing the detection difficulty.

[0004] Currently, deep learning-based target detection methods, such as the YOLO (You Only Look Once) series, Faster R-CNN, RetinaNet, etc., have been widely applied to underwater target detection tasks. Among them, YOLOv8, as the latest version of the YOLO series, inherits the characteristics of lightweight and efficient detection. However, in underwater scenarios, YOLOv8 still has the following problems: 1. Insufficient ability to detect small targets: Due to the small size of underwater targets, the feature extraction ability of YOLOv8 is limited, resulting in a low recall rate for small targets; 2. Poor robustness to light changes: In underwater environments with low light or uneven light, the traditional YOLOv8 model is difficult to adapt, prone to false detections and missed detections; 3. Lack of long-distance dependence modeling ability: YOLOv8 mainly relies on CNN (Convolutional Neural Network) for feature extraction, making it difficult to capture the global information between distant pixel points, affecting the detection accuracy; 4. Unable to fully utilize multi-scale feature fusion: The C2f module adopted by YOLOv8 still has room for improvement in adapting to multi-scale targets in the underwater environment. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides an improved YOLOv8 underwater image target detection method, device, medium, and program to solve the accuracy and robustness of the original YOLOv8 for underwater environment target detection.

[0006] The present invention achieves the above technical objectives through the following technical means.

[0007] An improved YOLOv8 underwater image target detection method adopts the following underwater image target detection network: based on YOLOv8, a fine-grained feature extraction module is set to replace the C2f module;

[0008] In the fine-grained feature extraction module, fine-grained features Y are extracted from the input feature map X:

[0009]

[0010] where Cross(·) represents the multi-scale cross-attention fusion module, LocalModule(·) represents the local feature extraction module, VisualSSM(·) represents the visual state space model, represents element-wise addition, and X′ is the feature map after element-wise addition and fusion of the outputs of the local feature extraction module and the visual state space model;

[0011] The visual state space model adjusts the feature distribution through linear transformation and extracts multi-scale spatial context by combining depthwise separable convolution and SiLU activation function;

[0012] The local feature extraction module focuses on the detailed information of the target through the convolutional kernel channel attention mechanism and suppresses noise interference;

[0013] The multi-scale cross-attention fusion module outputs fine-grained features by dynamically weighting and fusing local and global features of different scales.

[0014] Furthermore, in the visual state space model:

[0015] X1 = LN(2D_SSM(SiLU(DWConv(Linear(X)))00

[0016] X2 = SiLU(Linear(X))

[0017] VisualSSM(X) = Linear(X1⊙X2)

[0018] In the formula, LN(·) represents layer normalization, 2D_SSM(·) represents a two-dimensional state space model, SiLU(·) represents the SiLU activation function, DWConv(·) represents depthwise separable convolution, Linear(·) represents linear transformation, and "⊙" represents the Hadamard product.

[0019] Further, in the local feature extraction module:

[0020] F local = Conv 3×3 (ReLU(Conv 3×3 (X)))

[0021] α = σ(Conv 1×1 (ReLU(Conv 1×1 (AvgPool(X)))))

[0022] F fused = F local ·α

[0023] β = σ(Conv 7×7 (F fused ))

[0024] LocalModule(X) = F fused ·β

[0025] In the formula, Conv(·) represents convolution, the subscript represents the convolution kernel size, ReLU(·) represents the ReLU activation function, σ(·) represents the sigmoid activation function, and AvgPool(·) represents global average pooling.

[0026] Further, in the multi-scale cross-attention fusion module:

[0027] There are three parallel branches. In one branch: Asymmetric convolutions of two different scales are performed on the feature X′ to extract features F1 and F2, and then 1×1 convolution is used to fuse the channels of features F1 and F2 to generate the fused feature F3;

[0028] In the other two branches, 1×1 convolution is respectively performed on the feature X′ to generate features K and V:

[0029] K = Conv 1×1 (X′)

[0030] V = Conv 1×1 (X′)

[0031] The fused feature F3 is subjected to 1×1 convolution to generate the feature Q:

[0032] Q = Conv 1×1 (F3)

[0033] Feature Q is multiplied element - by - element with feature K, and the attention weight A is obtained through the Softmax activation function:

[0034] A = Softmax(Q·K)

[0035] The attention weight A is multiplied element - by - element with feature V to obtain the attention - enhanced feature F4:

[0036] F4 = A·V

[0037] The attention - enhanced feature F4 is added and fused with the fusion feature F3 to obtain feature Z, and feature Z is concatenated with the fusion feature F3 and output through a 1×1 convolution to obtain the fine - grained feature Y:

[0038] Z = F3 + F4

[0039] Y = Cross(X′) = Conv 1×1 ([F3,Z])

[0040] In the formula, Conv(·) represents convolution, its subscript represents the convolution kernel size, Softmax(·) represents the Softmax activation function, and [F3,Z] represents the concatenation of features F3 and Z.

[0041] Furthermore, in the multi - scale cross - attention fusion module:

[0042] F1 = Conv 3×1 (ReLU(Conv 1×3 (X′)))

[0043] F2 = Conv 5×1 (ReLU(Conv 1×5 (X′)))

[0044] F3 = Conv 1×1 (F1 + F2)

[0045] In the formula, ReLU(·) represents the ReLU activation function.

[0046] Furthermore, underwater target images are collected to make a dataset, and the underwater image target detection network is trained.

[0047] Furthermore, when training the underwater image target detection network, the number of training epochs Epoch = 300, the batch size BatchSize = 16, and the input image size InputSize = 640×640.

[0048] A computer device includes a memory and a processor;

[0049] The memory is used to store a computer program;

[0050] The processor is used to execute the computer program and implement the improved YOLOv8 underwater image target detection method when executing the computer program.

[0051] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the improved YOLOv8 underwater image target detection method.

[0052] A computer program product includes a computer program, and when the computer program is executed by a processor, the improved YOLOv8 underwater image target detection method is implemented.

[0053] The beneficial effects of the present invention are as follows:

[0054] (1) The present invention provides an improved YOLOv8 underwater image target detection method, device, medium, and program. In the present invention, a fine-grained feature extraction module introduced on the basis of YOLOv8 significantly improves the recognition ability of fuzzy targets and complex backgrounds in underwater image target detection through a parallel branch processing and dynamic feature fusion mechanism.

[0055] (2) In the improved YOLOv8 network of the present invention, the fine-grained feature extraction module takes lightweight parallel design and dynamic interaction as the core, avoids redundant calculations, and significantly improves the detection accuracy of small targets (such as sea urchins and sea clams) in underwater low-visibility scenes while maintaining high real-time performance.

[0056] (3) In the present invention, the visual state space model adjusts the feature distribution through linear transformation (Linear), combines depthwise separable convolution (DWConv) and SiLU activation function to extract multi-scale spatial context, and uses lightweight spatial optimization (SS2D) and layer normalization (LN) to enhance feature robustness. Description of the Drawings

[0057] Figure 1 It is the structural diagram of the improved YOLOv8 network of the present invention;

[0058] Figure 2 It is the structural diagram of the fine-grained feature extraction module network in the present invention;

[0059] Figure 3 It is the structural diagram of the local feature extraction module network in the present invention;

[0060] Figure 4 It is the structural diagram of the multi-scale cross-attention fusion module network in the present invention. Detailed Embodiments

[0061] Embodiments of the present invention will be described in detail below. Examples of the illustrated embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0062] I. Technical Solution

[0063] Traditional YOLOv8 relies on CNN for feature extraction. However, due to the local receptive field of CNN, it is difficult to capture the long-range dependencies in underwater images. As Figure 1 shown, based on the network structure of traditional YOLOv8, the present invention designs a new fine-grained feature extraction module to replace the original C2f module of YOLOv8. In the fine-grained feature extraction module, a single-variable sequence is mapped to an output sequence through an implicit latent intermediate state. This kind of hidden state models long-range dependencies and can overcome the problem of the local receptive field limitation of traditional convolution; thus, it not only bridges the relationship between the input and the output, but also encapsulates the temporal dynamics and is more suitable for capturing the global features of underwater images.

[0064] As Figure 2 shown, the fine-grained feature extraction module includes ① visual state space model, ② local feature extraction module, and ③ multi-scale cross-attention fusion module. Among them, ① visual state space model and ② local feature extraction module belong to two parallel branches. After the output results of the two branches are fused by element-wise addition, they are then processed by ③ multi-scale cross-attention fusion module to obtain the final required fine-grained feature Y. For the input feature map (where H, W, and C are the three dimensions of the input image, namely height, width, and number of channels), the processing process of the above fine-grained feature extraction module is expressed as:

[0065]

[0066] In the formula, Cross(·) represents the multi-scale cross-attention fusion module, LocalModule(·) represents the local feature extraction module, VisualSSM(·) represents the visual state space model, represents element-wise addition, and X′ is the feature map obtained by fusing the outputs of the local feature extraction module and the visual state space model through element-wise addition.

[0067] 1. Visual State Space Model

[0068] As Figure 2 shown, there are two parallel branches in the visual state space model, where:

[0069] 1) For the right branch shown in the figure, first apply a linear transformation (Linear) to the input feature X to adjust the distribution of the feature space; then extract local spatial information through depthwise separable convolution (DWConv), and introduce non-linear representation ability using the SiLU activation function (Swish activation function); then use a two-dimensional state space model (2D-SSM) for time series modeling to capture global dynamic features, and finally obtain the intermediate feature representation X1 through layer normalization (LN):

[0070] X1 = LN(2D_SSM(SiLU(DWConv(Linear(X))))

[0071] In the formula, LN(·) represents layer normalization, 2D_SSM(·) represents the two-dimensional state space model, SiLU(·) represents the SiLU activation function, DWConv(·) represents depthwise separable convolution, and Linear(·) represents linear transformation.

[0072] 2) For the left branch shown in the figure, perform a linear transformation (Linear) and the SiLU activation function on X in sequence to obtain the intermediate feature X2:

[0073] X2 = SiLU(Linear(X))

[0074] 3) For the output features of the above two branches, use the Hadamard Product for fusion and further adjust the feature distribution through a linear mapping to obtain the output of the visual state space model:

[0075] VisualSSM(X) = Linear(X1 ⊙ X2)

[0076] In the formula, "⊙" represents the Hadamard Product.

[0077] 2. Local Feature Extraction Module

[0078] As Figure 3 shown, the local feature extraction module focuses on details such as the edges and textures of underwater targets through convolution and channel attention mechanisms, suppresses noise interference, and realizes the localization and optimization of small targets. Among them, local feature extraction is first achieved through two parallel branches, including:

[0079] 1) For the right branch shown in the figure, use two 3×3 convolutions to extract local details F local :

[0080] F local = Conv 3×3 (ReLU(Conv 3×3 (X)))

[0081] In the formula, Conv(·) represents convolution, and its subscript represents the convolution kernel size, that is, Conv 3×3 (·) represents a 3×3 convolution, and ReLU(·) represents the ReLU activation function.

[0082] 2) For the left branch shown in the figure, global average pooling and 1×1 convolution are used to generate the channel attention weight α to highlight important features:

[0083] α = σ(Conv 1×1 (ReLU(Conv 1×1 (AvgPool(X)))))

[0084] In the formula, σ(·) represents the sigmoid activation function, Conv 1×1 (·) represents a 1×1 convolution, and AvgPool(·) represents global average pooling.

[0085] 3) Multiply the outputs of the above two branches to obtain the local feature F fused , then generate a spatial attention map through a 7×7 convolution to further enhance the significant region; finally, multiply it with the local feature F fused to obtain the output of the local feature extraction module:

[0086] F fused = F local ·α

[0087] β = σ(Conv 7×7 (F fused ))

[0088] LocalModule(X) = F fused ·β

[0089] In the formula, Conv 7×7 (·) represents a 7×7 convolution.

[0090] 3. Multi-scale Cross-attention Fusion Module

[0091] The outputs of the visual state space model and the local feature extraction module are fused by element-wise addition, making the local details and the global context complementary, and effectively solving the problem of feature weakening caused by scattering and uneven illumination in underwater images. The fused feature X′ is input into the multi-scale cross-attention module.

[0092] As Figure 4 shown, in the multi-scale cross-attention module, local and global features of different scales are fused through dynamic weighting to adaptively enhance the response of key regions, and finally discriminative fine-grained features are output. Among them, X′ is input into three parallel branches respectively:

[0093] 1) For the first branch on the upper side of the diagram, perform multi-scale asymmetric convolutions (1×3, 3×1, 1×5, 5×1) on feature X′ to extract features F1 and F2 of different scales, and then perform channel fusion on the multi-scale features through 1×1 convolution to generate the fused feature F3:

[0094] F1 = Conv 3×1 (ReLU(Conv 1×3 (X′)))

[0095] F2 = Conv 5×1 (ReLU(Conv 1×5 (X′)))

[0096] F3 = Conv 1×1 (F1 + F2)

[0097] In the formula, Conv 3×1 (·), Conv 1×3 (·), Conv 5×1 (·), Conv 1×5 (·) represent 3×1 convolution, 1×3 convolution, 5×1 convolution, and 1×5 convolution respectively.

[0098] 2) For the other two branches, perform 1×1 convolution on feature X′ respectively to generate features K and V:

[0099] K = Conv 1×1 (X′)

[0100] V = Conv 1×1 (X′)

[0101] 3) The fused feature F3 generates feature Q through 1×1 convolution; then first perform element-wise multiplication on feature Q and feature K, and obtain the attention weight A through the Softmax activation function; the attention weight A is then multiplied element-wise with feature V to obtain the attention-enhanced feature F4:

[0102] Q = Conv 1×1 (F3)

[0103] A = Softmax(Q·K)

[0104] F4 = A·V

[0105] In the formula, Softmax(·) represents the Softmax activation function.

[0106] 4) After adding and fusing the attention-enhanced feature F4 and the fused feature F3 to obtain feature Z, then feature Z is cascaded with the fused feature F3 and output through 1×1 convolution to complete multi-scale cross-attention fusion and obtain the required fine-grained feature Y:

[0107] Z = F3 + F4

[0108] Y = Cross(X′) = Conv 1×1 ([F3, Z])

[0109] In the formula, [F3, Z] represents the concatenation of features F3 and Z. By dynamically allocating the weights of different-scale features as described above, the adaptability of the network to multi-scale targets is enhanced.

[0110] II. Test and Verification

[0111] The improved YOLOv8 described above is used as an underwater image target detection network for underwater target image detection. Before using the target detection network, a dataset is made for training:

[0112] 1) Collect underwater target images from the Internet, public datasets, and underwater video frames collected independently. Use Labelimg software to annotate the target areas and generate.txt files in YOLO format. Each label file contains target category and bounding box coordinate information to complete the creation of the dataset.

[0113] 2) Clean the dataset, filter out low-quality, blurred, or invalid samples to ensure the availability of the data. At the same time, divide the dataset according to a ratio of 8:2, where 80% is used as the training set (Train) and 20% is used as the test set (Test).

[0114] Before training, set hyperparameters such as the number of training epochs Epoch, batch size BatchSize, and input image size InputSize of the underwater image target detection network. Through training iterations, the finally improved YOLOv8 underwater image target detection network is obtained. Then, use the trained detection network to predict the test set to obtain the final target detection results.

[0115] In this embodiment, an effectiveness experiment is carried out on the URPC2021 underwater target detection dataset. This dataset contains a total of 7,600 labeled underwater images and includes 4 types of targets: sea urchins, sea cucumbers, starfish, and scallops. Two comparative tests are carried out in the experiment: Method 1 is the target detection algorithm based on the original YOLOv8; Method 2 is the method proposed in the present invention. The number of training epochs Epoch = 300, the batch size BatchSize = 16, and the input image size InputSize = 640×640. The corresponding test results are shown in Table 1 below:

[0116] Table 1: Underwater Target Detection

[0117]

[0118]

[0119] The following 5 evaluation indicators are selected in the table: mAP50 (the average precision of the model when the IoU threshold is 50%), mAP50-95 (the average precision of the model when the IoU threshold ranges from 50% to 90%), Aps (the average precision of small targets (with an area less than 32×32 pixels)), Apm (the average precision of medium targets (with an area between 32×32 and 96×96 pixels)), and Apl (the average precision of large targets (with an area greater than 96×96 pixels)). The experimental results shown in Table 1 above indicate that the present invention improves by 1.1% in small target detection (APs) and by 1.8% in overall mAP50; compared with the original YOLOv8, the robustness of underwater target detection is significantly enhanced.

[0120] III. Device, Storage Medium, Program Product

[0121] 1. Based on the same inventive concept as the above-improved YOLOv8 underwater image target detection method, the present application also provides an electronic device, which includes a processor and a memory. The memory stores computer-readable code. When the computer-readable code is executed by the processor, the improved YOLOv8 underwater image target detection method of the present invention is implemented.

[0122] Among them, the memory includes a non-volatile storage medium and an internal memory; the non-volatile storage medium can store an operating system and computer-readable code. The computer-readable code includes program instructions, and when the program instructions are executed, the processor can be made to execute the improved YOLOv8 underwater image target detection method. The processor is used to provide computing and control capabilities to support the operation of the entire electronic device. The memory provides an environment for the operation of the computer-readable code in the non-volatile storage medium, and when the computer-readable code is executed by the processor, the processor can be made to execute the improved YOLOv8 underwater image target detection method.

[0123] It should be understood that the processor can be a central processing unit, other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.

[0124] 2. The present application also provides a readable storage medium, which can be the internal storage unit of the electronic device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card, a secure digital card, etc. equipped on the electronic device.

[0125] 3. The present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the improved YOLOv8 underwater image target detection method of the present invention.

[0126] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0127] The present invention is not limited to the above embodiments, and any obvious improvements, substitutions or deformations that those skilled in the art can make without departing from the essence of the present invention fall within the protection scope of the present invention.

Claims

1. An improved YOLOv8 underwater image target detection method, characterized in that: The following underwater image target detection network is adopted: Based on YOLOv8, a fine-grained feature extraction module is set to replace the C2f module; In the fine-grained feature extraction module, fine-grained features Y are extracted from the input feature map X: Wherein, Cross(·) represents the multi-scale cross-attention fusion module, LocalModule(·) represents the local feature extraction module, and VisualSSM(·) represents the visual state space model. represents element-wise addition, and X′ is the feature map obtained by fusing the outputs of the local feature extraction module and the visual state space model through element-wise addition. In the visual state space model, the feature distribution is adjusted through linear transformation, and multi-scale spatial context is extracted by combining depthwise separable convolution and SiLU activation function; In the local feature extraction module, the detail information of the target is focused through the convolutional kernel channel attention mechanism, and noise interference is suppressed; In the multi-scale cross-attention fusion module, local and global features of different scales are fused through dynamic weighting to output fine-grained features.

2. The improved YOLOv8 underwater image target detection method according to claim 1, characterized in that: In the visual state space model: X1 = LN(2D_SSM(SiLU(DWConv(Linear(X)))00 X2 = SiLU(Linear(X)) VisualSSM(X) = Linear(X1⊙X2) In the formula, LN(·) represents layer normalization, 2D_SSM(·) represents two-dimensional state space model, SiLU(·) represents SiLU activation function, DWConv(·) represents depthwise separable convolution, Linear(·) represents linear transformation, and "⊙” represents Hadamard product.

3. The improved YOLOv8 underwater image object detection method according to claim 1, characterized in that: In the local feature extraction module: F iocal = Conv 3×3 (ReLU(Conv 3×3 (X))) α = σ(Conv 1×1 (ReLU(Conv 1×1 (AvgPool(X))))) F fused = F local · α β = σ(Conv 7×7 (F fused )) LocalModule(X) = F fused ·β In the formula, Conv(·) represents convolution, its subscript represents the convolution kernel size, ReLU(·) represents ReLU activation function, σ(·) represents sigmoid activation function, and AvgPool(·) represents global average pooling.

4. The improved YOLOv8 underwater image target detection method according to claim 1, characterized in that: In the multi-scale cross-attention fusion module: There are three parallel branches. In one branch: Asymmetric convolutions of two different scales are performed on the feature X′ to extract features F1 and F2, and then channel fusion of features F1 and F2 is performed through 1×1 convolution to generate a fused feature F3; In the other two branches, 1×1 convolutions are respectively performed on the feature X′ to generate features K and V: K = Conv 1×1 (X′) V = Conv 1×1 (X′) The fused feature F3 generates a feature Q through 1×1 convolution: Q = Conv 1×1 (F3) The feature Q is multiplied element-wise with the feature K, and the attention weight A is obtained through the Softmax activation function: A = Softmax(Q·K) The attention weight A is multiplied element-wise with the feature V to obtain the attention-enhanced feature F4: F4 = A·V The attention-enhanced feature F4 and the fused feature F3 are added and fused to obtain the feature Z. The feature Z and the fused feature F3 are concatenated and output through 1×1 convolution to obtain the fine-grained feature Y: Z = F3 + F4 Y = Cross(X′) = Conv 1×1 ([F3,Z]) In the formula, Conv(·) represents convolution, its subscript represents the convolution kernel size, Softmax(·) represents Softmax activation function, and [F3,Z] represents concatenation of the features F3 and Z.

5. The improved YOLOv8 underwater image target detection method according to claim 4, characterized in that: In the multi-scale cross-attention fusion module: F1 = Conv 3×1 (ReLU(Conv 1×3 (X'))) F2 = Conv 5×1 (ReLU(Conv 1×5 (X'))) F3 = Conv 1×1 (F1 + F2) In the formula, ReLU(·) represents ReLU activation function.

6. The improved YOLOv8 underwater image object detection method according to claim 1, characterized in that: Underwater target images are collected to make a dataset, and the underwater image target detection network is trained.

7. The improved YOLOv8 underwater image object detection method according to claim 1, characterized in that: When training the underwater image target detection network, the number of training epochs Epoch = 300, the batch size BatchSize = 16, and the input image size InputSize = 640×640.

8. A computer device, characterized in that: It includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and, when executing the computer program, implement the improved YOLOv8 underwater image target detection method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: There is a computer program stored, and when the computer program is executed by the processor, the processor is made to execute the improved YOLOv8 underwater image target detection method according to any one of claims 1 to 7.

10. A computer program product, characterized in that: It includes a computer program, and when the computer program is executed by the processor, it implements the improved YOLOv8 underwater image target detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and system for detecting motion state of executing mechanism in cabin of underwater unmanned underwater vehicle

    CN121482106A

  • Garbage detection system and method based on computer vision

    CN121921313A

  • A computer vision-based garbage detection system and method

    CN121921313B