Target detection method for tiny organisms in underwater image

By improving the YOLOv8 network, combining the lightweight downsampling module and the small-target feature pyramid structure, the accurate positioning and efficient detection problems of underwater small-target detection are solved, the detection accuracy and robustness are improved, and it is suitable for complex underwater environments.

CN120259861APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510312966.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Small and medium-sized object detection in underwater images is difficult to accurately locate and identify, and efficient detection is difficult to achieve equipment with limited computing resources. The existing methods have poor detection performance in complex underwater environments.

Method used

A small biological dataset of underwater scenes is constructed, an improved YOLOv8 network is adopted, a lightweight downsampling module and a small-target feature pyramid structure are introduced, and feature recombination and fusion are performed through SPDConv and CSP-OmniKernel, combining bidirectional cross-scale connections to optimize feature expression and detection accuracy.

Benefits of technology

It improves the accuracy and robustness of underwater small target detection, adapts to complex underwater environments, reduces calculation overhead, and achieves efficient target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259861A_ABST
    Figure CN120259861A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method suitable for small creatures in an underwater complex scene, and aims to solve the problems that small target features in underwater image detection are fuzzy, small targets are difficult to accurately position and identify by the existing detection algorithm, and efficient detection is difficult to realize on equipment with limited computing resources. According to the method, on the basis of a YOLOv8 network and a backbone network, a P2 feature layer is adopted to enhance the expression ability of a small target, an ADwn lightweight down-sampling module is introduced, efficient feature down-sampling is achieved, a neck network constructs an SOEP small target feature enhancement pyramid module, feature recombination is conducted on P2 through SPDConv, efficient feature integration is conducted through CSP-OmniKernel, and therefore the expression ability of the small target features is improved; and feature fusion is carried out through bidirectional cross-scale connection, so that multi-scale information interaction is enhanced, and the detection capability of small targets is improved. The target detection network is more suitable for small biological target detection in an underwater complex environment, and provides powerful technical support for underwater ecological monitoring, marine organism research and other applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of object detection, and mainly relates to an object detection method for tiny organisms in underwater images. Background Art

[0002] In the underwater environment, object detection technology is widely used in underwater exploration, marine ecological monitoring, underwater robot navigation and other fields. However, due to problems such as uneven illumination, color deviation, low contrast and blur in underwater images, the detection performance of traditional object detection methods in the underwater environment is often poor. Especially for small object detection, due to water body scattering, low signal-to-noise ratio and target size limitations, existing methods are difficult to ensure the accuracy and robustness of detection. Therefore, studying a small object detection method for underwater images, which can effectively improve the recognition rate of underwater small objects, has important theoretical significance and application value for underwater exploration tasks.

[0003] With the rapid development of deep learning technology, object detection algorithms have made remarkable progress. These methods automatically learn image features by training a Convolutional Neural Network (CNN) and achieve high-precision object detection. Object detection algorithms based on deep learning can be roughly divided into single-stage and two-stage methods. Single-stage object detection algorithms use convolutional networks to extract high-level features, and directly complete object localization after fusing the feature maps. The typical representatives are the YOLO series. In contrast, two-stage object detection algorithms first generate candidate regions through a Region Proposal Network, and then classify and perform position regression on these regions to achieve more accurate object detection, mainly including methods such as R-CNN and Faster R-CNN. In recent years, the successful application of Transformer in the field of natural language processing has gradually introduced it into object detection tasks. Transformer relies on the global attention mechanism to model the relationships between regions of the image, can effectively capture the context information of the object, and establish global associations between the object and the background and between objects, thereby improving the detection ability of the model in complex scenarios.

[0004] The above algorithms improve the performance of object detection tasks from different perspectives. However, when directly applying these methods to underwater small object detection, many challenges are often faced, including image quality degradation. Underwater images suffer from color distortion, low contrast, and blurriness due to light scattering and absorption, making it difficult to distinguish small objects. The object size is small and the features are not obvious. Underwater small objects often have a small physical size, with a very low pixel proportion in the image, making them easily overlooked or confused with the background. There are many environmental interference factors. Plankton, suspended particles, bubbles, etc. in the underwater environment will interfere with object detection, increasing the false detection rate and missed detection rate. The computing resources are limited. When running the detection algorithm on an underwater robot or underwater monitoring device, limited by computing resources, the real-time performance and lightweight design of the algorithm become key issues.

[0005] To address the above problems, the present invention proposes an object detection method for small organisms in underwater images. This method can effectively improve the detection accuracy of the model in the underwater image small object detection task, and has the characteristics of strong robustness, high detection accuracy, strong adaptability to complex underwater environments, and low computational overhead, to meet the actual needs of underwater detection and monitoring tasks. Summary of the Invention

[0006] The purpose of the present invention is to propose an object detection algorithm for small organisms in underwater images to solve the problems in underwater image detection, such as the blurred features of small objects, the difficulty of accurately locating and identifying small objects by existing detection algorithms, and the difficulty of achieving efficient detection on devices with limited computing resources, in view of the numerous and difficult-to-detect underwater small objects and the lack of effective balance between accuracy and real-time performance.

[0007] To achieve the above purpose, the present invention provides the following technical solutions:

[0008] An object detection method for small organisms in underwater images, which specifically includes the following steps:

[0009] Step 1) Construct a dataset of small organisms in the underwater scene, preprocess the images in the dataset, annotate the position information and category of the objects, and generate a high-quality training sample set;

[0010] Step 2) Input the preprocessed images into an improved object detection model. This model adopts a layer-by-layer downsampling strategy to extract multi-scale features, optimizes the design of the feature pyramid structure for small object detection, introduces a P2 feature layer to enhance the expression ability of small object features, and simultaneously optimizes the computing efficiency using a lightweight downsampling module;

[0011] Step 3) Use the above-extracted feature maps as the input of the feature fusion module, construct a small target feature enhancement pyramid module, perform feature recombination on the P2 layer through SPDConv, and use CSP-OmniKernel for efficient feature integration to enhance the expression ability of small target features, and perform feature fusion through bidirectional cross-scale connections to enhance multi-scale information interaction and improve the detection ability of small targets;

[0012] Step 4) Input the fused features into the detection head, extract target features through the convolutional layer and perform position regression to achieve efficient and accurate small target detection. Finally, the detection head completes the precise positioning and class determination of the target.

[0013] Furthermore, step 1) specifically includes the following steps:

[0014] Step 11) Obtain underwater small creature image data covering different water environments, lighting conditions, and target pose changes. Use professional annotation tools to annotate the dataset for targets, generating corresponding target position and class information to ensure the accuracy and diversity of the dataset;

[0015] Step 12) Preprocess the original image data, including image normalization, size adjustment, color correction, etc., to optimize the image quality and construct a high-quality training sample set.

[0016] Furthermore, step 2) specifically includes the following steps:

[0017] Step 21) Input the preprocessed underwater image into the object detection backbone network improved based on YOLOv8, and use the lightweight downsampling (ADown) module to perform downsampling step by step to gradually extract high-level features. The ADown module adopts a multi-path processing strategy to reduce the complexity of the model, optimize the feature extraction effect, and reduce the loss of detailed information. The P3 layer initially extracts low-level features, retaining more detailed information, which helps the feature learning of small targets. The P4 layer extracts intermediate features, capturing the local structure and contour information of the target. The P5 layer extracts high-level features, focusing on the overall semantic information of the target to improve the target discrimination ability;

[0018] Furthermore, step 3) specifically includes the following steps:

[0019] Step 31) To enhance the expression ability of small targets, perform feature recombination on the P2 layer and use SPDConv for feature transformation in the spatial dimension. SPDConv extracts the key features of small targets through convolutional operations and performs scale adjustment so that these features can be effectively fused with the features of the P3 layer;

[0020] Step 32) Fuse the P2 feature layer after feature recombination with the P3 feature layer, and use the Cross-Stage OmniKernel (CSP-OmniKernel) module for efficient feature integration to obtain the enhanced P3'. The Cross-Stage OmniKernel module is improved based on the OmniKernel module, adopting a channel splitting strategy. Part of the channels are used for deep feature extraction, while the other part is directly passed as residual information to reduce computational overhead and retain key information. During the feature extraction process, global modeling is carried out by combining frequency-domain and spatial attention. At the same time, long-range context information is captured through convolution operations of different scales, and local convolution is used to maintain details. Finally, the enhanced feature P3' is generated through weighted fusion and feature splicing, which not only preserves the detailed information of small targets but also optimizes the hierarchical relationship of feature representation;

[0021] Step 33) Adopt Bidirectional Cross-Scale Connection to achieve layer-by-layer feature fusion from P3', P4, and P5, and construct a Small Object EnhancePyramid (SOEP) module. The top-down feature aggregation uses nearest-neighbor upsampling to gradually transfer high-level features (P5, P4) to low-level feature layers, enabling deep semantic information to feed back to the low levels and improving the detection ability of small targets. The bottom-up feature enhancement uses CSP residual connections to transfer low-level features (P3') upward to P4 and P5, enhancing the attention of high-level features to small targets.

[0022] Furthermore, step 4) specifically includes the following steps:

[0023] Step 41) Input the fused feature map into the detection head, optimize the computational efficiency using depthwise separable convolution, and combine Batch Normalization and SiLU activation to achieve efficient and accurate small target detection, predicting the bounding boxes and classes of the targets. Each bounding box corresponds to a class prediction and a confidence score. Finally, redundant boxes are removed through non-maximum suppression, and the final detection results are output. Description of the Drawings

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings, where:

[0025] Figure 1 is the flow chart of the present invention;

[0026] Figure 2 is the overall model framework diagram of the present invention;

[0027] Figure 3 is the structural diagram of the ADown module proposed by the present invention;

[0028] Figure 4 The structural diagram of the SPDConv module proposed by the present invention;

[0029] Figure 5 The structural diagram of the CSP-OmniKernel module proposed by the present invention;

[0030] Figure 6 The structural diagram of the SOEP module proposed by the present invention. Detailed implementation manners

[0031] In order to make the technical solutions, advantages and objectives of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments and drawings.

[0032] A target detection method for small underwater creatures in underwater images provided by the present invention, the method flow chart is as Figure 1 shown, the overall model framework diagram proposed by this method is as Figure 2 shown, and the method includes the following steps:

[0033] Step 1) Construct a target detection dataset for small underwater creatures in underwater images;

[0034] Step 2) Construct a lightweight downsampling module (ADown);

[0035] Step 3) Construct a small target backbone feature extraction network;

[0036] Step 4) Construct a small target feature enhancement pyramid module (SOEP);

[0037] Step 5) Construct a target detection network for small underwater creatures in underwater images;

[0038] Step 6) Train the target detection network;

[0039] Step 7) Test the target detection network.

[0040] Furthermore, the specific content of step 1) includes the following steps:

[0041] Step 11) Obtain underwater creature image targets using the publicly available dataset RUOD. This dataset contains 14,000 high-resolution underwater images, which include a variety of complex underwater scenes, including coral reefs, seagrass beds, rocks and sandy areas, etc., and provides a variety of lighting conditions and underwater environments, covering 10 common aquatic creature categories, namely: sea turtles, starfish, fish, sea urchins, scallops, sea cucumbers, jellyfish, corals, cuttlefish and divers. Delete the medium and large creature targets in the dataset and retain the small creature dataset;

[0042] Step 12) To increase the diversity of data, the open-source website acquires images of underwater small organisms, including different species, backgrounds, and lighting conditions. For each image, the target categories and their true positions are labeled through image annotation software, and an annotation file is generated. Together with the filtered RUOD, they are used as a dataset to ensure the richness and representativeness of the data.

[0043] Step 13) The images are resized, normalized, and the diversity of the data is increased by methods such as rotation, scaling, and cropping to prevent overfitting of the model. The obtained dataset is divided into a training set, a test set, and a validation set in a ratio of 6:2:2, which are used for training, validating, and testing the target detection network respectively.

[0044] Furthermore, the specific content of Step 2) includes the following steps:

[0045] Step 21) Combining the attached Figure 3 Detailed description, for the characteristics of high-resolution underwater images, small targets, and blurred edges, traditional downsampling (such as stride convolution, max pooling) will cause the following problems: small organisms are difficult to be captured by the subsequent network due to pixel annihilation after downsampling; the number of parameters of conventional convolution is large and it is difficult to be deployed to the edge computing unit of underwater devices; the interference of suspended particles in turbid water is easily retained by mistake during downsampling. The ADown module realizes efficient downsampling through channel attention-guided feature compression and spatial dimension reorganization. The ADown module is composed of a channel attention branch and a spatial reorganization branch in parallel. First, the input x is subjected to 2×2 average pooling for preliminary feature downsampling. Then the pooled x is split into two parts along the channel dimension (dim = 1), namely x1 and x2. x1 undergoes feature extraction through a 3×3 convolution to change its number of channels to c2 / 2. x2 first undergoes 3×3 max pooling for further feature extraction, and then the number of channels is adjusted through a 1×1 convolution to also make its number of channels c2 / 2. Then the features of the two branches are concatenated in the channel dimension, and the final output has c2 channels. ADown realizes downsampling through the combination of average pooling and max pooling, and combines 3×3 and 1×1 convolutions to adjust the number of channels. Compared with directly using ordinary stride convolution, it can retain feature information more fully, while improving the expression ability of the model, providing multi-scale features with high signal-to-noise ratio for the subsequent SOEP pyramid, and is the core module for balancing the accuracy and efficiency of the algorithm.

[0046] Furthermore, the specific content of Step 3) includes the following steps:

[0047] Step 31) Use the lightweight downsampling module constructed in Step 2 to construct a backbone feature extraction network for small target features. Perform preliminary feature extraction using Conv and enhance feature representation through the C2f structure. At the same time, introduce ADown as the downsampling module to replace the traditional Conv+Pool to optimize the feature extraction efficiency. Initially extract the features of P2, P3, P4, and P5 layers. Finally, use Spatial Pyramid Pooling (SPPF) to enhance multi-scale feature representation and ensure that key information can be effectively captured at different scales. The backbone feature extraction network is as Figure 2 shown in Backbone in

[0048] Furthermore, the specific content of Step 4) includes the following steps:

[0049] Step 41) As Figure 4 shown, use SPDConv to perform feature recombination on the P2 layer of the backbone network. The core idea of SPDConv is to map spatial information to the channel dimension. By aggregating the information of multiple spatial positions into the channels, the feature layer can adapt to the fusion between different scales while retaining details. It unfolds each 2×2 small area into 4 different channels, enabling the expression of the information of originally spatially adjacent pixels on the channels, expanding the feature map of the P2 layer to a higher channel dimension, adjusting the size to be compatible with the feature map of the P3 layer, so as to ensure that small target information can be effectively transmitted to the next layer;

[0050] Step 42) Fuse the P2 feature layer after feature recombination with the P3 feature layer, and use CSP-OmniKernel for efficient feature integration to obtain the enhanced P3'. CSP-OmniKernel is a multi-branch convolution module designed to strengthen the multi-scale representation ability of features through convolution operations with multiple different receptive fields. As Figure 5 shown, the input feature map x is adjusted in channels through cv1 (1×1 convolution layer), keeping the number of input and output channels unchanged, but preparing for subsequent channel splitting. The output of cv1 is split into two parts along the channel dimension (dim = 1): the first part accounts for e proportion (default 25%) of the total number of channels and is used for subsequent complex calculations, and the second part accounts for 1-e proportion (default 75%) of the total number of channels and is directly used as a residual connection to retain the original information. The first part is input into the OmniKernel module, as Figure 5As shown, the OmniKernel module achieves multi-scale feature fusion through the collaborative work of the global branch, large branch, and local branch. The global branch uses frequency-domain attention (FCA) and spatial-domain attention (SCA) to globally model the input features in the frequency-spatial domain. After applying channel attention weights in the frequency domain through Fourier transform and then inverse-transforming back to the spatial domain, important regions are screened through spatial attention. That is, the input feature map is transformed to the frequency domain through the fast Fourier transform (FFT):

[0051]

[0052] where represents the two-dimensional Fourier transform, and the output is the frequency-domain feature in complex form.

[0053] The space is compressed through global average pooling to generate channel-level statistics, and then attention weights are generated through 1×1 convolution and activation function:

[0054] a FCA = σ(Conv 1×1 (GAP(y in ))) ∈ R C×1×1

[0055]

[0056] where σ is the Sigmoid function, which constrains the weights to the interval [0, 1].[[]END]]

[0057] The channel attention weight a FCA is applied to the frequency-domain feature, and then the spatial feature is restored through the inverse Fourier transform (IFFT):

[0058]

[0059] where ⊙ represents element-wise multiplication, is the inverse Fourier transform.

[0060] Channel attention calculation is performed again on the frequency-domain enhanced feature y FCA :

[0061] a SCA = σ(Conv 1×1 (GAP(y FCA ))) ∈ R C×1×1

[0062] y SCA = a SCA ⊙ y FCA ∈ R C×H×W

[0063] y out = FGM(y SCA ) ∈ RC×1×1

[0064] Among them, FGM is spatial gated convolution.

[0065] The large branch uses three types of depthwise separable convolutions of 31×31, 1×31, and 31×1 in parallel to extract long-range spatial features, capturing omnidirectional, horizontal, and vertical large-scale context information respectively; the local branch retains local detailed features through 1×1 depth convolution and residual connections. Finally, the global attention calibration features, large receptive field structure features, and local fine features are added and fused, and after ReLU activation, they are output through 1×1 convolution, forming a composite feature expression system that takes into account global perception, mid-range structure modeling, and local detail retention. The processed OmniKernel output (occupying 25% of the channels) is concatenated with the unprocessed Identity branch (the remaining 75% of the channels) along the channel dimension to form a feature map P3' with the complete number of channels. The finally generated P3' feature map contains rich multi-scale information and is more suitable for further object detection;

[0066] Step 43) As Figure 6 shown, P3', P4, and P5 are used for feature fusion layer by layer through bidirectional cross-scale connections to construct the SOEP module. The bidirectional cross-scale connections enhance the interaction between multi-layer features through top-down and bottom-up alternating information transmission, enabling each layer of features to be optimized with the assistance of information from other layers, ultimately improving the detection ability of small objects.

[0067] Furthermore, the specific content of step 5) includes the following steps:

[0068] Step 51) Using the module formed in the above steps, fuse ADown lightweight downsampling, small object backbone feature extraction, and SOEP feature pyramid to optimize small object feature expression and cross-scale fusion, and improve detection accuracy. Based on the YOLOv8 optimized backbone network, enhance the P2 feature layer, perform efficient downsampling in combination with ADown, and perform deep feature extraction with C2f. SOEP reorganizes the P2 features through SPDConv, efficiently fuses information with CSP-OmniKernel, and enhances the detection ability using bidirectional cross-scale connections (P3', P4, P5). The detection head uses depthwise separable convolution to optimize the calculation efficiency, combines Batch Normalization and SiLU activation to improve classification and regression performance, realizes efficient and accurate small object detection, and constructs an object detection network for underwater small organisms.

[0069] Furthermore, the specific content of step 6) includes the following steps:

[0070] Step 61) Input the dataset into the object detection network model, configure the corresponding training environment, set the corresponding training parameters, and carry out the training task of the model;

[0071] Step 62) The operating system used in this experimental environment is Windows 11, the deep learning framework is pytorch 2.2.2, Python 3.10.14, and CUDA 12.1. The operating environment is: Intel 12th generation Core i5-12600KF central processing unit, 32GB of memory, and the GPU is NVIDIA GeForce RTX 4060Ti-8G;

[0072] Step 63) In the hyperparameter configuration of model training, the optimizer uses SGD for training, and the total number of training cycles is set to 300. The initial learning rate is set to 0.01, the momentum parameter is 0.973, and each batch contains 16 samples.

[0073] Furthermore, step 7) specifically includes the following steps:

[0074] Step 71) Load the trained weight file into the object detection network, randomly select pictures from the test dataset for detection, and return the annotated pictures with the position and category information of the objects to be detected.

[0075] The above embodiments are only used to illustrate the technical solution of the present invention, rather than limiting it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that without departing from the core idea and scope of the technical solution of the present invention, it can be modified or equivalently replaced, and these changes should be included within the scope of the claims of the present invention.

Claims

1. A target detection method for tiny organisms in underwater images, characterized in that, It includes the following steps: Step 1: Construct a dataset of small underwater organisms and preprocess the dataset. Obtain underwater biological images using the publicly available dataset RUOD, which contains 14,000 high-resolution images covering complex underwater scenes such as coral reefs, seagrass beds, rocks, and sandy areas, with various lighting conditions and underwater environments, and covering 10 common aquatic organisms such as sea turtles, starfish, fish, sea urchins, etc. Subsequently, screen the dataset, delete the medium and large biological targets, and only retain the small organism data. In addition, to enhance the diversity of the data, supplement underwater small organism images of different varieties, backgrounds, and lighting conditions from open-source websites, and use annotation software to annotate the target categories and locations to generate corresponding annotation files, and finally integrate them with the screened RUOD dataset. To optimize the model training effect, adjust the size of the data, normalize it, and expand the dataset by rotating, scaling, cropping, etc. to prevent overfitting. Finally, divide the data into a training set, a test set, and a validation set in a ratio of 6:2:2 to ensure the balance and representativeness of the training, validation, and testing of the target detection model; Step 2: Construct a lightweight downsampling module. The ADown module adopts a multi-path processing strategy, uses convolutional layers to extract useful information from the feature map, reduces the spatial dimension of the feature map by adjusting the stride of the convolutional layer, optimizes the number of parameters in the convolutional layer to reduce the complexity of the model, and optimizes the feature extraction effect and reduces the loss of detailed information; Step 3: Construct a small target backbone feature extraction network. Based on the backbone network improved from YOLOv8, introduce the P2 feature layer to enhance the small target feature expression ability, and gradually extract features in a hierarchical and progressive manner and continuously improve the feature expression ability. First, perform preliminary feature extraction through consecutive Conv layers, and at the same time use a convolutional layer (3×3) with a stride of 2 to achieve downsampling and gradually reduce the spatial resolution. Then, at different scales, introduce the C2f module for deep feature extraction respectively, and increase the number of channels after each downsampling to enhance the feature expression ability. To improve the information flow, use the lightweight downsampling module Adown for downsampling. Finally, the backbone ends with the SPPF module, which enhances the receptive field through pooling operations at different scales and further integrates multi-scale information to provide rich feature representations for the subsequent detection heads; Step 4, construct the Small Object Enhanced Feature Pyramid (SOEP) module: reorganize the features of the P2 layer, and use SPDConv to transform the features in the spatial dimension. SPDConv extracts the key features of the small target through convolution operations and adjusts the scale to make it easier to effectively fuse with the P3 layer. Furthermore, CSP-OmniKernel is used for efficient feature integration to obtain the enhanced P3'. This module combines channel segmentation strategy, frequency domain and spatial attention, and convolution operations of different scales to optimize the feature expression level while reducing computational overhead. Finally, the enhanced feature P3' is generated through weighted fusion and feature concatenation, which not only retains the detailed information of the small target, but also improves the overall feature quality. On this basis, bidirectional cross-scale connections are introduced to construct the SOEP pyramid to achieve layer-by-layer feature fusion from P3', P4, and P5. Top-down feature aggregation uses nearest neighbor upsampling to pass high-level semantic information to low-level layers to improve small target detection capabilities. Bottom-up feature enhancement uses CSP residual connections to pass low-level features, enhancing the attention of high-level features to small targets, thereby building a more accurate small target detection feature pyramid. Step 5, construct a target detection network for small organisms in underwater images: use the downsampling module constructed in the above steps, integrate the backbone feature extraction network, the SOEP pyramid module that integrates small target features, and the detection head of the YOLO network to form a target detection method network for small organisms in underwater images; Step 6, training the target detection network model: the data set is used to input the target detection network model, the Windows 11 operating system and the PyTorch2.2.2 deep learning framework are configured, and the training environment is built based on Python 3.10.14 and CUDA 12.

1. SGD is used as the optimization function in the training phase, the training cycle is set to 300 epochs, the initial learning rate is 0.01, the momentum is 0.973, and the number of single training samples is 16 to ensure efficient model convergence and improve detection performance. The target detection network for small organisms in underwater images is fully trained to obtain the trained network model weights; Step 7, test the target detection network model: load the trained weight file into the target detection network, randomly select samples from the test set for inference, and output the detection result image with the target location and category information marked. Evaluate the target detection performance of the model in the underwater small organisms dataset, including detection accuracy, inference speed, and overall average accuracy.

2. The object detection method according to claim 1, wherein The ADown module in step 2 adopts a multi-path optimization strategy for efficient downsampling, which specifically includes: The ADown module is composed of a channel attention branch and a spatial recombination branch in parallel. First, the input x is subjected to 2×2 average pooling to perform preliminary feature downsampling. Then, the pooled x is split into two parts along the channel dimension (dim = 1), namely x1 and x2. x1 undergoes feature extraction through a 3×3 convolution, making its number of channels become c2 / 2. x2 first undergoes 3×3 max pooling for further feature extraction, and then the number of channels is adjusted through a 1×1 convolution, making its number of channels also become c2 / 2. Then, the features of the two branches are concatenated in the channel dimension, and the final output has c2 channels. ADown achieves downsampling through the combination of average pooling and max pooling, and combines 3×3 and 1×1 convolutions to adjust the number of channels.

3. The object detection method according to claim 1, wherein The construction of the SOEP pyramid module in step 4 includes the SPDConv module for reorganizing the features of the P2 layer and the CSP-OmniKernel module for enhancing the P3 layer that fuses small target features. The SPDConv reorganizes the features of the P2 layer in a space-to-depth conversion manner. The space-to-depth mapping divides the input feature map into 4 sub-regions by 2×2 and concatenates them in the channel dimension, enabling small target features to be enhanced in a deeper channel dimension. Convolution processing is performed on the converted feature map to avoid a decrease in spatial resolution and retain detailed information, enabling the features of the P2 layer to be effectively fused with the P3 layer. The CSP-OmniKernel module performs efficient multi-scale feature extraction and fusion through a channel partitioning strategy to enhance the feature expression ability and reduce the computational cost. First, after the input feature map is adjusted through a 1×1 convolution, it is partitioned in a ratio of 25%:75% in terms of channels. Among them, 25% enters the OmniKernel for deep feature extraction, while 75% is directly passed as residual information to reduce the computational amount. The OmniKernel adopts global, large-branch, and local-branch for collaborative feature extraction: The global branch combines frequency domain attention (FCA) and spatial attention (SCA), enhances channel information through Fourier transform, and uses spatial attention to screen key regions; The large branch uses 31×31, 1×31, and 31×1 depthwise separable convolutions to extract omnidirectional, horizontal, and vertical long-range context information; the local branch retains local details through 1×1 depth convolution and residual connection. Finally, the features of the three branches are added and fused, processed through ReLU activation and 1×1 convolution, and concatenated with the unchanged 75% residual information in the channel dimension to generate the enhanced P3' feature map, achieving efficient multi-scale information integration. The P3', P4, and P5 are subjected to feature fusion layer by layer using bidirectional cross-scale connections to construct the SOEP pyramid module. The bidirectional cross-scale connection enhances the interaction between multi-layer features through top-down and bottom-up alternating information transfer, enabling each layer of features to be optimized with the assistance of information from other layers, ultimately improving the detection ability of small targets.

4. The object detection method according to claim 1, wherein Step 5 builds a target detection network for small organisms in underwater images. Based on the ADown lightweight downsampling module, small target backbone feature extraction network and SOEP small target enhanced feature pyramid module constructed above, the YOLOv8 detection head is further integrated to form a complete target detection network. This detection network optimizes the feature extraction, cross-scale fusion and detection accuracy of small targets based on the characteristics of small organisms in underwater images, ensuring excellent detection performance in complex underwater scenes. The backbone network is optimized based on YOLOv8, and the P2 feature layer is used to enhance the expression ability of small targets. The ADown lightweight downsampling module is introduced to perform efficient feature downsampling with a multi-path optimization strategy to reduce computational overhead while retaining key details. In addition, the C2f module is used for deep feature extraction to ensure that features at different scales have high recognition. The SOEP small target enhanced feature pyramid module is constructed, P2 is reorganized through SPDConv, and CSP-OmniKernel is used for efficient feature integration to improve the expression ability of small target features. P3', P4, and P5 use bidirectional cross-scale connections to perform feature fusion, enhance multi-scale information interaction, and improve the detection ability of small targets. The detection head of YOLOv8 is used to perform target classification and bounding box regression on features of different scales to further optimize the detection effect of underwater small organisms. The detection head is based on deep separable convolution to reduce the amount of calculation. At the same time, BatchNormalization and SiLU activation functions are used to ensure that the network can be stably trained on data with different distributions. On the basis of retaining the end-to-end detection advantages of YOLOv8, it is combined with an enhanced feature pyramid to make it more adaptable to complex underwater environments. Through the above optimization strategies, the target detection network can efficiently and accurately detect small organisms in complex underwater environments, providing better technical support for applications such as underwater ecological monitoring and marine biological research.

Citation Information

Cited By

  • Density detection and intrusion early warning method for shallow sea biological model

    CN120635687A

  • Deep-sea biological image processing method based on neural network

    CN121236572A

  • Remote sensing image target detection method based on improved RT-DETR and computer equipment

    CN121259530A

  • Lightweight target detection Transform model based on space-frequency domain joint modeling, method and application

    CN121280870A

  • Lightweight target detection transformer model, method and application based on space-frequency domain joint modeling

    CN121280870B