End-to-end lightweight underwater target detection method and system

By adopting an end-to-end lightweight underwater target detection method, combined with a lightweight image enhancement module and a low-parameter convolutional structure, the problems of cumbersome models and slow processing in underwater target detection systems are solved, achieving efficient and stable underwater target detection, which is suitable for complex environments and resource-constrained platforms.

CN120913050AActive Publication Date: 2025-11-07齐鲁空天信息研究院

Patent Information

Application Number
CN202511439489.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing underwater target detection systems suffer from excessive model parameters, high computational resource consumption, and slow inference speed, resulting in response delays and increased recognition error rates. They cannot meet the technical requirements of underwater operations for high real-time performance, lightweight deployment, and stable recognition.

Method used

An end-to-end lightweight underwater target detection method is adopted, which combines a lightweight image enhancement module, a low-parameter convolutional structure, and a multi-scale target perception mechanism with edge computing devices to achieve localized inference of the model, thereby reducing resource consumption and improving detection speed and accuracy.

Benefits of technology

A lightweight target detection system suitable for complex underwater environments was constructed, which improved the system's processing speed and detection accuracy, enhanced the model's operating efficiency and deployment flexibility on resource-constrained platforms, and achieved fast, stable, and accurate underwater target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913050A_ABST
    Figure CN120913050A_ABST
Patent Text Reader

Abstract

The invention provides an end-to-end lightweight underwater target detection method and system, and belongs to the field of target detection, and the method comprises the steps: carrying out the frame sampling of an underwater video stream at a fixed interval, and obtaining an original image as a to-be-detected image; sending the original image into a differentiable brightness enhancement module and a lightweight texture enhancement module in parallel to generate a color correction image and a texture enhancement image in sequence, and inputting the original image, the color correction image and the texture enhancement image into a space correlation gating attention fusion module to obtain a fusion enhancement image; inputting the fusion enhancement graph into a backbone network formed by connecting a plurality of feature extraction layers in series and then connecting the feature extraction layers in series with a spatial pyramid pooling layer, and outputting a multi-scale feature graph; sending the multi-scale feature map into a path aggregation network to obtain a fused feature map with consistent semantics; and inputting the fusion feature map and the output of the spatial pyramid pooling layer into a multi-branch detection head, completing category, bounding box and confidence prediction, and outputting a final detection result through non-maximum suppression.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of target detection, and particularly relates to an end-to-end lightweight underwater target detection method and system. BACKGROUND

[0002] With the progress of the net zero development concept, green development and automatic exploration of seabed resources have gradually become a hot spot of research and industry attention in recent years. As a key link to support the realization of technologies in this field, underwater image and video target detection technology has attracted widespread attention. Underwater target detection not only has important application value in marine energy exploration, pipeline inspection, ecological monitoring, etc., but also provides a technical basis for building an intelligent and unmanned seabed operation system. However, due to the special imaging environment of underwater, the detection task faces many challenges. The turbidity of seawater, the interference of suspended particles, and the strong attenuation and scattering of light in water result in serious visual distortion problems in the collected original images, such as uneven color difference, low contrast, blur and loss of details, etc. These problems significantly reduce the effectiveness of traditional image processing and detection algorithms, limiting their application performance in actual complex environments.

[0003] To address the above problems, researchers have tried to propose improvement schemes from image enhancement, feature extraction and model structure design, etc. Although some methods have achieved good performance on standard test sets, how to balance detection speed and end-to-end lightweight deployment of the system while pursuing high accuracy is still an important technical problem in current underwater target detection research.

[0004] Based on the current research results of scholars in the field of underwater target detection, although significant progress has been made in detection accuracy, model structure optimization and image preprocessing, which can improve the processing ability and adaptability of the model to complex underwater images to some extent, there are still many challenges in practical applications. Especially when facing massive underwater video streams or high-frequency image data, the existing methods often have too many model parameters, high resource occupation, slow inference speed, and limited device processing capacity, etc., which leads to response delay, rising error rate, even system freezing, frame loss or crash, etc. in the running process of the overall detection system, which cannot meet the technical requirements of high real-time, lightweight deployment and stable recognition of underwater operations, and also affects the application expandability and practicality of the system. SUMMARY

[0005] To solve the above technical problems, the present application provides an end-to-end lightweight underwater target detection method and system, and the specific technical solutions are as follows:

[0006] An end-to-end lightweight underwater target detection method, comprising the following steps:

[0007] Step 1, frame sampling is performed on the underwater video stream at a fixed interval to obtain an original image as a to-be-detected image;

[0008] Step 2, the original image is sent into a differentiable brightness enhancement module and a lightweight texture enhancement module in parallel to generate a color correction image and a texture enhancement image in sequence, and the original image, the color correction image and the texture enhancement image are input into a spatial correlation gate attention fusion module to obtain a fusion enhancement image;

[0009] Step 3, the fusion enhancement image is input into a backbone network composed of a plurality of feature extraction layers connected in series and then connected in series with a spatial pyramid pooling layer to output a multi-scale feature map;

[0010] Step 4, the multi-scale feature map is input into a path aggregation network to obtain a semantic consistent fusion feature map;

[0011] Step 5, the fusion feature map and the output of the spatial pyramid pooling layer are input into a multi-branch detection head to complete class, bounding box and confidence prediction, and then output the final detection result through non-maximum suppression.

[0012] An end-to-end lightweight underwater target detection system, comprising the following modules:

[0013] The sampling module performs frame sampling on the underwater video stream at a fixed interval to obtain an original image as a to-be-detected image;

[0014] The fusion enhancement image generation module sends the original image into a differentiable brightness enhancement module and a lightweight texture enhancement module in parallel to generate a color correction image and a texture enhancement image in sequence, and inputs the original image, the color correction image and the texture enhancement image into a spatial correlation gate attention fusion module to obtain a fusion enhancement image;

[0015] The multi-scale feature map generation module inputs the fusion enhancement image into a backbone network composed of a plurality of feature extraction layers connected in series and then connected in series with a spatial pyramid pooling layer to output a multi-scale feature map;

[0016] The fusion feature map generation module inputs the multi-scale feature map into a path aggregation network to obtain a semantic consistent fusion feature map;

[0017] The result generation module inputs the fusion feature map and the output of the spatial pyramid pooling layer into a multi-branch detection head to complete class, bounding box and confidence prediction, and then outputs the final detection result through non-maximum suppression.

[0018] An electronic device, comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method.

[0019] A computer-readable storage medium having stored thereon executable instructions that, as a result of being executed by a processor, cause the processor to implement the described method.

[0020] The present application has the following beneficial effects:

[0021] The present application proposes an end-to-end lightweight underwater target detection optimization method to solve the problems of heavy model, slow processing, complex deployment and other problems existing in traditional underwater target detection systems. The method realizes efficient processing of the whole process from input to detection output of underwater images through joint optimization in multiple dimensions such as model architecture, data preprocessing and inference process, and guarantees the stability and real-time performance of the system in complex scenes.

[0022] In the image input stage, the present application designs a lightweight image enhancement module, fully considers the characteristics of underwater images, combines with the differentiable enhancement mechanism, realizes real-time correction and contrast enhancement of color distortion and fuzzy noise; secondly, introduces a low-parameter and high-efficiency convolution structure in the backbone network, which greatly reduces the model complexity while ensuring the feature extraction capability; at the same time, the multi-scale target perception mechanism is integrated in the detection head part to enhance the recognition accuracy of small targets and fuzzy targets; finally, through end-to-end model compression and quantization technology in the whole system deployment link, the resource consumption of the model running is reduced, and the model local inference is realized combined with the edge computing device to reduce the data transmission delay. Through the above optimization strategies, the present application constructs a lightweight target detection system suitable for complex underwater environment, which not only effectively improves the overall processing speed and detection accuracy of the system, but also significantly enhances the running efficiency and deployment flexibility of the model on resource-constrained platforms, finally realizes the rapid, stable and accurate detection of underwater targets, and provides strong technical support for underwater intelligent operation system.

[0023] The lightweight image enhancement module proposed by the present application includes differentiable brightness enhancement, lightweight texture enhancement and spatial correlation gated attention fusion network, which can be directly integrated into the front end of the detection network to realize end-to-end training, image enhancement and target detection integration. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 The flow chart of the end-to-end target detection method of the present application.

[0025] Figure 2 The original image, the optimized image and the detection result of the present application. DETAILED DESCRIPTION

[0026] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other. In order to achieve the above-mentioned purpose, the technical solutions adopted by the present application are as follows.

[0027] As shown in Figure 1 The present application provides an end-to-end lightweight underwater target detection method, comprising the following steps:

[0028] Step 1, first, frame sampling is performed at an interval of every 15 frames from a video stream continuously collected by an optical camera, and a key frame image is extracted as a to-be-detected image.

[0029] Further, the to-be-detected image sampled is input into a lightweight image enhancement module to perform color deviation correction, noise suppression and contrast enhancement, etc., and an enhanced image with higher visual quality is output.

[0030] Specifically, the lightweight image enhancement module is composed of a differentiable brightness enhancement module, a lightweight texture enhancement module and a spatially related gated attention fusion module. First, the to-be-detected image is input into the differentiable brightness enhancement module and the lightweight texture enhancement module in two parallel branch processing paths, on the one hand, the differentiable brightness enhancement module performs brightness separation and color correction on the image according to the Retinex theory, compensates for the color deviation and low illumination in underwater imaging, and outputs a color-corrected image (denoted as image 1) with natural color and enhanced contrast; on the other hand, the image is input into the lightweight texture enhancement module, which extracts texture and edge features through a low-complexity guided filtering strategy and a deep separable convolution, effectively alleviates image blur and suppresses background noise, and outputs a texture-enhanced image (denoted as image 2) with clear texture. Then, the original image, image 1 and image 2 are input into the spatially related gated attention fusion module for fusion processing, which constructs a spatially sensitive attention weight to fuse the features of the three input images, dynamically adjusts the contribution weight of the features of each path according to the importance of the image region, and finally outputs the enhanced image as the input of the subsequent target detection network.

[0031] Specifically, the differentiable brightness enhancement module aims to simulate the traditional Retinex image enhancement principle by decomposing the image into a reflection component reflecting the inherent properties of the object and an illumination component reflecting the influence of light to improve the brightness and color quality of underwater images. The module adopts a fully differentiable neural network structure, which can be integrated into the entire system for end-to-end joint training. First, the input image is normalized and converted to the logarithmic space; then, a shallow convolutional network composed of a small number of lightweight convolutional layers is used to extract features of the image and estimate its illumination component; then, by subtracting the estimated illumination component from the original image in the logarithmic space, the reflection component is extracted, thereby eliminating the effects of uneven underwater lighting; subsequently, the reflection map is converted back to the original image domain, and a learnable nonlinear enhancement parameter is introduced to further improve the brightness dynamic range and visual contrast of the image; finally, the color-corrected image after brightness enhancement (denoted as image 1) is output.

[0032] The shallow convolutional network includes a depth separable convolution layer with a Relu function, a 1x1 point convolution layer, and a Sigmoid activation function for controlling the brightness range.

[0033] Further, the lightweight texture enhancement module uses the original image as a guide image to enhance the texture details of the input image based on the idea of guided filtering. First, the original image is input into the guide image feature convolutional network to extract multi-scale structural information and obtain an intermediate feature map; then, low-frequency structural information is extracted from the intermediate feature map as an RGB guide image; subsequently, based on the local consistency relationship between the RGB guide image and the original image, a texture extraction network separates the smooth structure component and the texture detail component. The extracted texture detail component is then subjected to nonlinear mapping and detail enhancement to improve the local contrast and detail perception of the image, obtaining a high-frequency texture map, i.e., a texture-enhanced image; finally, the RGB guide image and the enhanced high-frequency texture map are fused in the spatial dimension by a structure-texture fusion module to output an image with clearer texture and stronger detail performance.

[0034] The guide image feature convolutional network includes a 3x3 depth separable convolution layer with a Relu activation function, a GhostConv layer for compressing the number of channels, a large kernel convolution 5x5 depth separable convolution network with a Relu activation function for enhancing the receptive field, a 1x1 point convolution layer for controlling the texture details, and a Sigmoid activation function, finally outputting a three-channel RGB guide image.

[0035] The texture extraction network includes a GhostConv layer for expanding the number of channels, a 3x3 depth separable convolution network with a Relu activation function, a 3x3 depth separable convolution layer with a residual connection, a 1x1 point convolution layer for compressing the number of channels, and finally outputs a high-frequency texture map.

[0036] The detail enhancement refers to the scaling and adjustment of the pixel values of the high-frequency texture image by multiplying a weight, adding an offset, and then passing through a ReLU function, so as to enhance the local contrast.

[0037] The structure-texture fusion module first combines the original image, the RGB guide image extracted by the guide image feature convolution network, and the high-frequency texture image after texture extraction and detail enhancement processing according to the channel dimension, then generates a fusion weight map by a light convolution network, and finally, the structure image and the texture image are weighted and added according to the fusion weight. The light convolution network includes a layer of 1x1 convolution, a layer of 3x3 convolution, and a layer of Sigmoid activation function.

[0038] Further, the above-mentioned spatial correlation gate attention fusion module controls the weight of the fusion of features at different spatial positions by combining the spatial attention mechanism and the gate network. First, the original image, the color correction image, and the texture enhancement image are subjected to channel transformation and normalization using 1x1 convolution; then, using the original image as a reference, the spatial correlation between the three feature maps is calculated in the spatial dimension using the self-attention idea to obtain the spatial correlation weight matrix of the color feature and the texture feature; then, the gate network is used to fuse and output a fusion image as the input of target detection.

[0039] The gate network first splices the three feature maps in the channel layer, then inputs the gate network to generate the weights of the three branches, finally participates in the weighting after channel broadcasting, and performs convolution integration, normalization processing on the fused features to output the enhanced image.

[0040] The gate network includes a layer of 1x1 convolution network with ReLU activation function for compressing channels, a layer of 1x1 convolution network for mapping to three gate channels, and a layer of Sigmoid activation function for normalization operation.

[0041] Step 2, input the enhanced image into the backbone network to extract multi-scale features and obtain feature maps of different scales.

[0042] Specifically, the backbone network comprises a plurality of feature extraction layers and a spatial pyramid pooling layer, the feature extraction layers are connected in series, and the spatial pyramid pooling layer is connected in series at the rear end of the series-connected feature extraction layers. The feature extraction layer comprises a plurality of convolution modules connected in series, each convolution module comprising a 1x1 convolution layer, a depth separable convolution layer, a nonlinear activation function layer, and a batch normalization layer. The spatial pyramid pooling layer comprises a first lightweight depth separable convolution network, a plurality of series-connected pooling convolution kernels, a splicing layer, and a second lightweight depth separable convolution network; the first lightweight depth separable convolution network is used for feature extraction of the input feature map, and the extracted features are divided into two paths, one path is sequentially pooled by a plurality of parallel-connected pooling convolution kernels and transmitted to the splicing layer, the series-connected pooling convolution kernels comprise a 1x1 hollow convolution, a 3x3 hollow convolution, and a 5x5 hollow convolution, and the other path is directly transmitted to the splicing layer; the splicing layer is used for splicing all the received features, and then transmitting them to the second lightweight depth separable convolution network for further feature extraction to obtain a feature map. The first and second lightweight depth separable convolution networks are both sub-networks sequentially stacked by a plurality of lightweight convolution modules, and the lightweight convolution module comprises a series-connected 1x1 convolution layer, a depth separable convolution layer, a nonlinear activation function layer (RelU6), and a batch normalization layer.

[0043] Step 3: input the feature map output in step 2 into the path aggregation network for feature fusion. The path aggregation network adopts a feature pyramid structure and comprises a plurality of 1x1 convolution layers and a top-down multi-scale fusion connection module; wherein the 1x1 convolution layer is used for channel compression of the feature maps output by the backbone network at different scales to unify the channel dimensions, facilitating subsequent fusion processing; the top-down connection module is used to gradually sample the high semantic and low resolution feature map and fuse it with the adjacent low layer feature map in a manner of guiding low layer features by high layer features, forming a top-down multi-scale feature conduction path. In the fusion process, the output fusion feature map of each level is obtained by element-wise addition operation of the corresponding low layer feature map and the up-sampled high layer feature map, and edge refinement and semantic enhancement processing is performed through subsequent 3x3 convolution, so as to obtain multi-scale and semantic consistent fusion feature map, and improve the recognition ability of the detection model for different scale targets.

[0044] Step 4, input the fusion feature map obtained in step 3 and the feature map output by the spatial pyramid pooling layer in step 2 into a detection head network for detection to obtain a result map containing target categories and positions, and then process the result map through a non-maximum suppression method to obtain a target detection result map. The detection head network includes a plurality of detection branches arranged in parallel, each branch corresponding to a fusion feature map of a specific scale input to achieve accurate detection of targets of different scales. The branch includes a plurality of 1x1 convolution layers and 3x3 convolution layers connected in series, which are respectively used for category prediction, bounding box regression and confidence prediction.

[0045] Step 5, perform structural optimization processing on the trained target detection network, specifically including two stages of model pruning and model quantization, to reduce model complexity and improve inference speed, thereby meeting the deployment requirements on resource-constrained devices.

[0046] The model pruning operation prunes a plurality of convolution layers in the backbone network, the path aggregation network and the detection head network based on the channel importance evaluation results under the premise of maintaining the continuity of the backbone network structure, to remove redundant channels with low contribution. Through pruning processing, the number of model parameters and the calculation overhead can be significantly reduced, while the original detection accuracy is maximally preserved.

[0047] The model quantization operation is performed after pruning, and a perception quantization training strategy is adopted to convert floating-point weight parameters and intermediate activation values in the model into low-bit representations. The quantization process includes linear mapping of the convolution layer weights to INT8 representation, and setting a fixed range for the activation values and inserting a quantization pseudo-operation to realize integer arithmetic replacement in the network inference stage, thereby further reducing the model size and improving the execution efficiency.

[0048] After the above pruning and quantization processing, the target detection network can realize lightweight deployment while maintaining high detection accuracy, meeting the requirements of edge devices for real-time performance, low power consumption and low latency.

[0049] Step 6, deploy the pruned and quantized model on an embedded or edge computing platform to realize real-time inference of the model in a low-power and efficient environment. The deployment process includes model export, inference engine adaptation, hardware loading and system integration.

[0050] In the model export step, the trained target detection model is converted into a standard intermediate representation format suitable for deployment, preferably using the ONNX format to ensure the universality and portability of the model between different platforms.

[0051] In the reasoning engine adaptation step, an adapted deep learning reasoning framework is selected according to the computing resources and instruction set architecture of the target hardware platform, including but not limited to TensorRT, TFLite, OpenVINO or NCNN, and the model structure is processed for operator fusion, memory optimization and quantization scheduling, etc., to improve the execution efficiency and resource utilization of the reasoning process.

[0052] In the hardware loading step, the reasoning engine and the target model are deployed together into an edge device, which can include Jetson Nano, Jetson Xavier, Raspberry Pi, RK3588, Cambrian MLU, Ascend Atlas NPU, etc., to meet the performance and power consumption balance requirements in different scenarios.

[0053] In the system integration step, the target detection module is integrated and packaged with the image input module, the image preprocessing module, the post-processing module and the control instruction interface module to form a complete underwater target detection system, and a unified interface calling standard is provided to facilitate deployment and linkage with underwater robots, remote buoys, AUVs and other platforms, and to improve the overall stability and real-time response capability of the system.

[0054] More specifically, the present application provides an end-to-end lightweight underwater target detection method, comprising the following steps:

[0055] Step 1, input the original underwater image into the lightweight image enhancement module for color correction and texture structure enhancement processing to obtain the detected image.

[0056] Specifically, assuming that the original underwater image includes RGB three channels, and the size is , the original image is first input into the differentiable brightness enhancement module to obtain the color corrected image , and at the same time, the lightweight texture enhancement module is input in parallel to obtain the texture enhanced image , and finally the original image , the color corrected image , the texture enhanced image are input into the spatial correlation gate attention fusion module, and the fusion output is obtained to obtain the detected image .

[0057] Specifically, the differentiable brightness enhancement module first normalizes the image to [0, 1] and converts it to logarithmic space according to the following formula, and the result is:

[0058] ;

[0059] wherein, , >0 is used to avoid log zero. Then, input to the illumination component estimation network, and output an estimated illumination map :

[0060] ;

[0061] ;

[0062] wherein, denotes a depthwise separable convolution, denotes a point convolution, denotes a Sigmoid activation function.

[0063] Then, the reflectance component is estimated according to the following formula:

[0064] ;

[0065] Finally, the reflectance map is non-linearly enhanced and remapped to obtain a color-corrected image according to the following formula :

[0066] ;

[0067] wherein, , is a learnable parameter, is a weight parameter, which controls the scaling and translation of the input R, is a bias parameter, which controls the translation amount, is initially set to 1, is initially set to 0 and will be automatically adjusted through backpropagation during the training process. exp denotes an exponential function with a natural constant as the base, and ReLU denotes a ReLU6 activation function.

[0068] Further, the lightweight texture enhancement module first inputs the original image into a guide map feature extraction network to extract multi-scale structures and output a structure guide map :

[0069] ;

[0070] ;

[0071] Then, the original image I and the guide map G are input into a texture extraction network to extract a texture detail map :

[0072] ;

[0073] ;

[0074] Then, the texture detail map T is non-linearly adjusted to obtain :

[0075] ;

[0076] wherein, , are learnable parameters. are weight parameters, controlling the scaling and translation of input R, are bias parameters, controlling the translation amount, initially set to 1, initially set to 0, which will be automatically adjusted through back propagation during the training process.

[0077] Finally, the original image I, the structure guide image G and the enhanced texture image are spliced into , and input into the fusion network to obtain the texture enhanced image :

[0078] ;

[0079] ;

[0080] ;

[0081] Specifically, the spatial correlation gate attention fusion module first aligns and normalizes the feature channels of the original image I, the color enhanced image and the texture enhanced image according to the following formula:

[0082] ;

[0083] ;

[0084] ;

[0085] Then, based on the normalized original image , the similarity between the color enhanced image , the texture enhanced image and the original image is calculated respectively. The greater the similarity, the higher the credibility of the position.

[0086] ;

[0087] ;

[0088] wherein, denotes the spatial coordinates of the image; ε is a stability parameter, which is a very small positive number for numerical stability; finally, the weight W of the fusion of the three image features is obtained according to the following formula, wherein , and finally outputting the fused image according to the weight :

[0089] ;

[0090] ;

[0091] Step 2, inputting the fused image obtained in Step 1 into a multi-scale feature extraction process in a lightweight backbone network, wherein the lightweight backbone network adopts MobileNetV3_Small as the backbone network skeleton, which has the characteristics of small parameter quantity and high computational efficiency, and is suitable for image processing tasks in resource-constrained environments. Specifically, the MobileNetV3_Small network first performs an initial convolution operation on the input image to compress the spatial size and extract low-level texture features, generating a feature map , wherein is the initial channel number. Subsequently, the fused image features are sequentially passed through multiple inverted residual structure modules of different scales and receptive fields for semantic feature extraction. Each inverted residual structure module includes an expansion convolution, a depth separable convolution, and a projection convolution. Some modules have SE attention mechanisms embedded inside to enhance channel expression capabilities, and output multiple scale feature maps , , , . The above feature maps contain spatial structures and semantic information at different resolutions. After that, a spatial pyramid pooling layer is used to enhance the multi-scale receptive field. First, the feature map is input into a first lightweight depth separable convolution network to obtain basic features. Then, it is divided into two paths. One path is directly transmitted to the concatenation layer. The other path is sequentially passed through a 1x1 atrous convolution, a 3x3 atrous convolution, and a 5x5 atrous convolution in parallel to extract features at different receptive fields, and then sent to the concatenation layer to fuse all features. Finally, a second lightweight depth separable convolution network is used for further extraction and compression to obtain enhanced high-level feature maps . Subsequently, the input path aggregation network is used for feature fusion and enhancement.

[0092] Step 3, inputting the multiple different scale feature maps output by the lightweight backbone network MobileNetV3_Small in Step 2 into the path aggregation network in sequence to perform feature fusion operations, so as to further improve the model's detection capability for different scale targets.

[0093] ​The path aggregation network is constructed by using a feature pyramid network structure, which mainly includes a plurality of 1x1 convolution layers and a top-down multi-scale fusion connection module. The plurality of 1x1 convolution layers are used for channel compression processing on the input different scale feature maps } to unify the channel number to a preset channel dimension, so as to be fused step by step subsequently.

[0094] Next, the top-down multi-scale fusion connection module conducts from the low resolution, high semantic high layer feature to the high resolution, strong structure low layer feature, layer by layer restores the high layer feature map to the size of the adjacent low layer feature map by the up-sampling mode, and performs the element-by-element addition fusion operation to form the fusion feature map.

[0095] According to the above mode, the final three-layer fusion feature map set is obtained Each layer of the fusion feature retains the high layer semantic information and the bottom layer structure details, and effectively improves the expression and perception ability of the model to the multi-scale target.

[0096] Step 4, the multi-scale fusion feature map obtained by processing the path aggregation network in step 3 and the feature map output by the spatial pyramid pooling layer in step 2 are input into the detection head network, and a target detection operation is performed to output a prediction result map containing target category information and spatial position information in the image. After the prediction result map is processed by the non-maximum suppression method, the final target detection result map is obtained.

[0097] Specifically, the detection head network includes a plurality of parallel detection branches, each branch corresponding to an input feature map of a specific scale, for realizing accurate detection of multi-scale targets. Each detection branch includes a plurality of convolution structures connected in cascade, and the convolution structures include a group of 1x1 convolution layers and a group of 3x3 convolution layers in sequence, which are respectively used for target class prediction, boundary box regression and confidence score estimation. Among them:

[0098] The 1x1 convolution layer is used for channel compression and feature mapping conversion of the input feature map;

[0099] The 3x3 convolution layer is used to enhance the spatial context perception ability and further extract the structure information required for positioning and classification;

[0100] The class prediction sub-branch is used to output the probability value of the class to which each detection box belongs;

[0101] The boundary box regression sub-branch is used to predict the center point coordinates (x, y), width (w) and height (h) of the target box;

[0102] The confidence prediction sub-branch outputs the target score by combining the classification probability and the target existence probability.

[0103] The outputs of the respective detection branches are unified and summarized in the inference stage, and redundant boxes with high overlap and low confidence are removed through a non-maximum suppression algorithm, and finally the target detection result map with the highest confidence is retained.

[0104] Step 5, the target detection network model trained in step 4 is subjected to structure optimization processing. The structure optimization processing includes two operations of model pruning and model quantization.

[0105] The model pruning process refers to pruning a plurality of convolutional layers in the image enhancement network, the backbone network, the path aggregation network and the detection head network according to channel importance, so as to remove redundant calculation.

[0106] The model quantization process refers to inserting a simulated quantization operation to the floating-point number weight and activation value of the model in the training stage by using a perception quantization training strategy. The weight quantization adopts a linear symmetric quantization manner, and the activation value quantization adopts an asymmetric quantization manner.

[0107] Step 6, the model pruned and quantized in step 5 is deployed in the RK3588 edge computing platform. Specifically, the model is first exported in the ONNX format, and the model conversion and quantization adaptation are performed through the RKNN Toolkit to generate an RKNN format model that can run on the RK3588 NPU; then the model is loaded into the inference engine of the embedded device, and is integrated with the video input module, the preprocessing module, the post-processing module and the control instruction module to construct an underwater target detection system that can run on the RK3588 platform.

[0108] Another aspect of the present application provides an end-to-end lightweight underwater target detection system, comprising the following modules:

[0109] The sampling module samples the underwater video stream at a fixed interval to obtain an original image as a to-be-detected image;

[0110] The fusion enhanced image generation module sequentially generates a color correction image and a texture enhanced image by parallelly feeding the original image into a differentiable brightness enhancement module and a lightweight texture enhancement module, and inputs the original image, the color correction image and the texture enhanced image into a spatial correlation gate attention fusion module to obtain a fusion enhanced image.

[0111] The multi-scale feature map generation module inputs the fusion enhanced image into a backbone network composed of a plurality of feature extraction layers connected in series and then connected in series with a spatial pyramid pooling layer to output a multi-scale feature map.

[0112] The fusion feature map generation module inputs the multi-scale feature map into a path aggregation network to obtain a semantic consistent fusion feature map.

[0113] The results generation module inputs the fused feature map and the output of the spatial pyramid pooling layer into the multi-branch detection head to complete the prediction of category, bounding box and confidence, and then outputs the final detection result after non-maximum suppression.

[0114] Another aspect of the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method.

[0115] Another aspect of the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described thereon.

[0116] The training process of the end-to-end lightweight underwater target detection model is described below: Step 1: Create a dataset.

[0117] Step 1.1: Data Collection. Collect typical underwater datasets of RUIE as sample data, covering the main scenarios of low brightness, occlusion, and high blur.

[0118] Step 1.2: Data Partitioning. To demonstrate the effectiveness of this invention, the data in this embodiment is randomly divided into a training set, a validation set, and a test set in a 6:2:2 ratio. The training set and validation set are used to train the model parameters, and the test set is used to test the model performance, ensuring that the data from the three sets do not overlap.

[0119] Step 1.3: Data Preprocessing. To improve detection efficiency, the model input data is resampled to 640×640. Simultaneously, random cropping, flipping, mixing, and mosaic techniques are used to expand the training dataset, reduce class imbalance, and improve the robustness of the lightweight object detection model.

[0120] Step 2: Train and validate the model. Input the training dataset from Step 1 into the constructed lightweight object detection model for training. The training parameters are set as follows: input image size is 640×640, batch size is 16, optimizer is SGD, initial learning rate is 0.01, and number of iterations is 300, resulting in the trained lightweight object detection model.

[0121] Step 3: Model Export. Export the pt format weight file of the trained lightweight object detection model as an ONNX format lightweight object detection model. After compilation by the toolchain, convert it into an rknn format lightweight object detection model suitable for NPU devices.

[0122] Step 4: deploying the trained lightweight target detection model to an underwater platform. The rk3588 chip is arranged in the underwater platform, that is, the trained lightweight target detection model is deployed to a board card carrying the rk3588 chip, inference is performed through access to multiple video streams, and target coordinates and class confidence are returned after performing a non-maximum suppression algorithm.

[0123] Table 1: Performance comparison of the lightweight target detection model of the application and YOLOv8-s , According to Table 1, the end-to-end lightweight underwater target detection method is simple in principle and easy to implement. Compared with YOLOv8s, the lightweight target detection model of the application has the following advantages: The parameter amount is reduced from 11.1M to 8.9M, the calculation amount is reduced from 28.5GFLOPS to 24.6GFLOPS, and the inference speed is increased from 131FPS to 139FPS.

[0124] As shown in Figure 2 , the original image, the color correction image, the texture enhancement image, the fusion enhancement image and the detection result of the application can be seen, and the detection result has high accuracy.

Claims

1. An end-to-end lightweight underwater target detection method, characterized in that, The method comprises the following steps: Step 1: frame sampling of underwater video stream at fixed intervals to obtain original images as detection images; Step 2: the original images are sent into a differentiable brightness enhancement module and a lightweight texture enhancement module in parallel to generate color correction images and texture enhancement images in sequence, and the original images, color correction images and texture enhancement images are input into a spatial correlation gate attention fusion module to obtain a fusion enhancement image; Step 3: the fusion enhancement image is input into a backbone network composed of a plurality of feature extraction layers connected in series and then connected with a spatial pyramid pooling layer to output a multi-scale feature map; Step 4: the multi-scale feature map is input into a path aggregation network to obtain a semantic consistent fusion feature map; Step 5: the fusion feature map and the output of the spatial pyramid pooling layer are input into a multi-branch detection head to complete class, bounding box and confidence prediction, and then the non-maximum suppression is performed to output the final detection result.

2. The end-to-end lightweight underwater target detection method according to claim 1, wherein, The differentiable brightness enhancement module in step 2 is executed in the following order: the original image is normalized and logarithmized to obtain a logarithmic image; a 3x3 depth separable convolution, a 1x1 point convolution and a Sigmoid function are used to estimate an illumination component; The logarithmic image is subtracted from the illumination component to obtain a reflection component; The reflection component is enhanced by a learnable parameter nonlinear function and is exponentially mapped to output a color correction image.

3. The end-to-end lightweight underwater target detection method according to claim 2, wherein, The lightweight texture enhancement module in step 2 is executed in the following order: A 3x3 depth separable convolution, a GhostConv, a 5x5 depth separable convolution, a 1x1 convolution and a Sigmoid function are used to generate a structure guide image guided by the original image; The original image and the structure guide image are spliced and then sequentially input into a GhostConv, a 3x3 depth separable convolution, a residual 3x3 depth separable convolution and a 1x1 convolution to extract a high-frequency texture image; The high-frequency texture image is adjusted by a learnable parameter nonlinear function to obtain an adjusted high-frequency texture image; The original image, the structure guide image and the adjusted high-frequency texture image are spliced in the channel dimension and then sequentially input into a 1x1 convolution, a 3x3 convolution and a Sigmoid function to generate a fusion weight, and the fusion weight is used to output a texture enhancement image.

4. The end-to-end lightweight underwater target detection method according to claim 3, wherein, The spatial correlation gate attention fusion module in step 2 is executed in the following sub-order: The original image, the color correction map, and the texture enhancement map are respectively subjected to 1x1 convolution and normalization to obtain an original enhancement image , a color enhancement image , and a texture enhancement image ; the original enhanced image as a reference, the original enhanced image and the color enhanced image , the original enhanced image and the texture enhanced image spatial cosine similarity images Sim1, Sim2 The original enhanced image , the color enhanced image , the texture enhanced image , the spatial cosine similarity image Sim1, Sim2 are spliced, then sequentially pass through 1*1 convolution, Sigmoid function to generate three-channel weights, and output the fusion enhanced image according to the three-channel weights.

5. The end-to-end lightweight underwater target detection method according to claim 1, wherein, The spatial pyramid pooling layer in step 3 is executed in the following order: First, a first lightweight depth separable convolution is used to extract features and divide them into two paths; One path is directly sent to a splicing layer, and the other path is sent to the splicing layer after being processed by a 1x1, 3x3 and 5x5 atrous convolution in parallel; The splicing result is further fused by a second lightweight depth separable convolution to output a multi-scale feature map.

6. The end-to-end lightweight underwater target detection method according to claim 1, wherein, The feature extraction layer comprises a plurality of convolution modules connected in sequence, and each convolution module comprises a 1x1 convolution layer, a depth separable convolution layer, a nonlinear activation function layer and a batch normalization layer.

7. The end-to-end lightweight underwater target detection method according to claim 1, wherein, The detection head network comprises a plurality of detection branches arranged in parallel, and each branch corresponds to a specific scale of fusion feature input to realize detection of targets of different scales.

8. An end-to-end lightweight underwater target detection system, characterized in that, The method comprises the following modules: A sampling module for frame sampling of underwater video stream at fixed intervals to obtain original images as detection images; The fusion enhanced image generation module sequentially generates a color correction image and a texture enhanced image by feeding the original image into a differentiable brightness enhancement module and a lightweight texture enhancement module in parallel, and inputs the original image, the color correction image, and the texture enhanced image into a spatial correlation gate attention fusion module to obtain a fusion enhanced image; The multi-scale feature map generation module inputs the fusion enhanced image into a backbone network composed of a plurality of feature extraction layers connected in series and then connected with a spatial pyramid pooling layer to output a multi-scale feature map; The fusion feature map generation module inputs the multi-scale feature map into a path aggregation network to obtain a semantic consistent fusion feature map; The result generation module inputs the fusion feature map and the output of the spatial pyramid pooling layer into a multi-branch detection head to complete class, bounding box and confidence prediction, and then outputs the final detection result through non-maximum suppression.

9. An electronic device, comprising: Comprise: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A processor having stored thereon executable instructions that, when executed, cause the processor to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cross-scale context low-illumination image enhancement method based on attention mechanism

    CN113284064A

  • Lightweight underwater target detection method and system based on image enhancement

    CN114821286A

  • Lightweight target detection method and system based on unmanned aerial vehicle platform

    CN118587622A

  • Lightweight network target detection method based on structure optimization and feature fusion

    CN118710883A

  • Underwater image enhancement method based on Mama multi-feature enhancement fusion

    CN119809951A

Cited By

  • Lightweight end-to-end sea surface small target detection method and system based on original digital baseband echo signal

    CN121934041A