Multi-scale edge information enhancement method for underwater target detection

By building a cross-stage multi-scale edge information enhancement module, the problem of edge information being ignored in underwater target detection is solved, the detection accuracy and robustness are improved, the model's sensitivity to edge information is enhanced, and noise interference is reduced.

CN120339815APending Publication Date: 2025-07-18DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510306143.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing underwater target detection methods ignore the role of edge information, resulting in low target positioning accuracy and serious noise interference in complex underwater environments.

Method used

A cross-stage multi-scale edge information enhancement module is built, and edge information enhancement units and gradient flow branches are divided into channel operations, multi-scale edge information enhancement units, and gradient flow branches, edge information is extracted and enhanced, and fused with the original feature map to generate multi-scale enhancement features.

Benefits of technology

It improves the accuracy and robustness of underwater target detection, reduces noise interference, enhances the model's sensitivity to edge information, and improves the accuracy of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339815A_ABST
    Figure CN120339815A_ABST
Patent Text Reader

Abstract

The invention provides a multi-scale edge information enhancement method for underwater target detection, and belongs to the technical field of image processing and deep learning. Dividing the number of channels of the initial feature information into a main part and other parts by using channel division operation; the input main branch of the main part is subjected to feature extraction and enhancement through a plurality of multi-scale edge information enhancement units; the rest part is input into parallel gradient flow branches for transmission; splicing the output feature information of all branches to obtain spliced features; inputting the splicing features into a convolution structure to obtain edge information enhancement features; and meanwhile, dividing channel operation and cross-layer connection of the plurality of multi-scale edge information enhancement units. According to the method, the multi-scale edge information enhancement features are obtained, the edge information extracted from the shallow layer features is transmitted to the whole backbone network and is fused with the features of different scales, so that the edge information in the features extracted from each scale is enhanced, and the target detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and deep learning, and particularly to a multi-scale edge information enhancement method for underwater target detection. Background Art

[0002] Underwater target detection requires processing complex image data from vast ocean areas at different time points, spatial positions, and with limited resources. The water characteristics in these areas, including transparency, color, temperature, biological components, and the content of suspended particles, all show significant differences, which particularly tests the network's ability to distinguish targets from the background.

[0003] In the prior art, the accurate determination of the position of underwater targets depends on target detection models with strong generalization performance. Existing object detection architectures can be roughly divided into three categories: two-stage detection networks, one-stage detection networks, and Transformer-based end-to-end detection networks. Among them, two-stage detection networks have better accuracy and robustness, such as the RCNN series. Such networks first generate a series of region candidate boxes through specific strategies, and then use deep convolutional neural networks to extract features from the image regions within these candidate boxes. On this basis, the extracted features are classified and the bounding boxes are finely adjusted to achieve accurate positioning of the targets. The core advantage of one-stage object detection networks, such as SSD and YOLO series, lies in their real-time inference ability, which is suitable for application scenarios that require quick responses. The Transformer-based architecture has an advantage in dealing with long-range dependencies in its self-attention mechanism, thus enabling a more accurate understanding of the image content. However, the object detection in the prior art ignores the role of edge information. Edge information can more clearly show the contour and shape of the target, which can help the network more accurately determine the position and boundary of the target and improve the positioning accuracy. In addition, there is a large amount of noise in the underwater environment, such as suspended particles and light refraction. Edge information helps to suppress these noises and extract more crucial and clear target information.

[0004] Therefore, a multi-scale edge information enhancement method for underwater target detection is needed. Summary of the Invention

[0005] In view of this, the present invention provides a multi-scale edge information enhancement method for underwater target detection, which improves the detection performance and feature expression ability on edge information by constructing a cross-stage multi-scale edge information enhancement module.

[0006] To this end, the present invention provides the following technical solutions:

[0007] A multi-scale edge information enhancement method for underwater target detection, comprising:

[0008] The number of channels of the initial feature information is divided into a main part and the remaining part by using a channel division operation;

[0009] The main part is input into the main branch for feature extraction and enhancement through multiple multi-scale edge information enhancement units;

[0010] The remaining part is input into the parallel gradient flow branch for transmission;

[0011] The output feature information of all branches is concatenated to obtain a concatenated feature;

[0012] The concatenated feature is input into a convolutional structure to obtain an edge information enhanced feature;

[0013] At the same time, the channel division operation and multiple multi-scale edge information enhancement units are cross-connected.

[0014] Further, the multi-scale edge information enhancement unit performs feature extraction and enhancement, including:

[0015] Extract local information of different sizes through an adaptive average pooling layer to obtain feature maps at different scales;

[0016] Extract the edge information of the feature map through an edge enhancer to obtain an edge enhanced feature;

[0017] Align and concatenate the edge enhanced features at different scales to obtain a fused feature;

[0018] Input the fused feature into a convolutional layer to generate a unified feature representation.

[0019] Further, the step of extracting the edge information of the feature map through an edge enhancer to obtain an edge enhanced feature includes:

[0020] Perform a smoothing operation on the input feature map through two-dimensional average pooling;

[0021] Obtain enhanced edge information based on the input feature map and the smoothed feature map;

[0022] Perform rectangular self-calibration on the enhanced edge information to obtain calibrated enhanced edge information;

[0023] Add the calibrated enhanced edge information to the input feature map to obtain an edge enhanced feature.

[0024] Further, the step of performing rectangular self-calibration on the enhanced edge information includes:

[0025] Perform horizontal pooling and vertical pooling on the enhanced edge information to generate a horizontal axis vector and a vertical axis vector;

[0026] Construct an edge information model based on the horizontal axis vector and the vertical axis vector through the broadcasting method;

[0027] Calibrate the region of interest of the edge information model through a shape self - calibration function.

[0028] Further, the calibration of the region of interest of the edge information model through the shape self - calibration function includes:

[0029] Use horizontal bar convolution to calibrate the shape in the horizontal direction;

[0030] Use vertical bar convolution to calibrate the shape in the vertical direction;

[0031] And perform non - linear processing, batch normalization processing, Sigmoid activation, and convolution processing using the ReLU activation function to obtain calibrated enhanced edge information.

[0032] Further, splice the output feature information of all branches through bilinear interpolation to obtain a spliced feature.

[0033] Advantages and positive effects of the present invention:

[0034] The present invention obtains enhanced edge information through an edge enhancer including rectangle self - calibration; generates multi - scale enhanced edge features by cooperating the edge enhancer with an adaptive average pooling layer; performs cross - stage connection on the multi - scale enhanced edge features, transfers the edge information extracted from the shallow - layer features to the entire backbone network, and fuses with features of different scales, thereby enhancing the edge information in the features extracted at each scale. The method of the present invention can be integrated into a multi - scale edge information enhancement module and applied to different object detection models to improve the detection accuracy of the object detection model. Description of the Drawings

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is the architecture diagram of the multi - scale edge information enhancement unit in the embodiment of the present invention;

[0037] Figure 2 It is the structural diagram of the cross - stage multi - scale edge information enhancement module in the embodiment of the present invention;

[0038] Figure 3 It is the schematic diagram of the cross - stage multi - scale edge information enhancement module applied to the DETR architecture in the embodiment of the present invention;

[0039] Figure 4Schematic diagram of the cross-stage multi-scale edge information enhancement module applied to the YOLO architecture in the embodiments of the present invention;

[0040] Figure 5 Comparison chart of object detection results in the comparative experiment results of the present invention;

[0041] Figure 6 Thermal contrast chart in the comparative experiment results of the present invention. Detailed implementation manners

[0042] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0044] The present invention provides a multi-scale edge information enhancement method for underwater target detection. The main idea is to extract features from different scales, enhance the edge information, and integrate these features and output through a convolutional layer. Specifically, first, multi-scale pooling is performed to extract multi-level features. Then, an edge enhancer extracts edge information to enhance the network's sensitivity to edges. Finally, the features of different scales are aligned through interpolation, spliced and fused into a unified feature representation to improve the model's perception of multi-scale features. Among them, the edge enhancer extracts edge information: extracts and enhances the edge information in the image, obtains edge features through smoothing operations and subtraction, and then performs global context modeling in the horizontal and vertical directions. Then, the edge features are adjusted through a shape self-calibration function to make them closer to the foreground target. Finally, the processed edge information is added to the original feature map to generate an enhanced output.

[0045] Combined with Figure 2 The method of the present invention is further described;

[0046] S1 divides the number of channels of the initial feature information into two halves using a divided channel;

[0047] S2 Several multi-scale edge information enhancement units in the main branch perform feature extraction on the input information;

[0048] The remaining part is passed through a parallel gradient flow branch, and n represents that the multi-scale edge information enhancement module is repeated n times.

[0049] S3 Concatenates the output feature information of multiple branches through a Concat operation to obtain a fused feature;

[0050] S4 Inputs the fused feature into a convolutional structure to output an edge information enhanced feature.

[0051] At the same time, the divided channel operation and multiple cross-layer connections ensure that the C2f-MSEIE module obtains rich gradient flow information.

[0052] Combined with the figure, the feature extraction of the input information by the multi-scale edge information enhancement unit in S2 is further described as follows:

[0053] 1. Perform multi-scale pooling through adaptive average pooling to extract local information of different sizes for capturing multi-level features of the image:

[0054] features i (x) = Conv(Conv(AdaptiveAvgPool(x, size = bin)))

[0055] Among them, AdaptiveAvgPool represents using adaptive average pooling on the input feature map and pooling it to a feature map with a size of bin; Conv represents a convolutional operation for feature extraction.

[0056] 2. Extract edge information through an edge enhancer to enhance the network's sensitivity to edges:

[0057]

[0058] Among them, the edge enhancer represents performing edge enhancer processing on the feature map, and the number of input channels is

[0059] Obtain edge information through the edge enhancer;

[0060] 1) Use two-dimensional average pooling to smooth the input feature map and extract its low-frequency information; then subtract the original input feature map from the smoothed feature map to obtain enhanced edge information:

[0061] edge = x - AvgPool(x)

[0062] Among them, x is the input feature map, edge is the low-frequency information, AvgPool is the average pooling operation, the window size is 3×3, the stride is 1, and the padding is 1.

[0063] 2) Perform rectangular self-calibration on the enhanced edge information;

[0064] Rectangular self-calibration captures the axial global context in two directions through horizontal pooling and vertical pooling, generating a horizontal axis vector and a vertical axis vector:

[0065] x h = AdaptiveAvgPool2d(edge, height = H, width = 1)

[0066] x w = AdaptiveAvgPool2d(edge, height = 1, width = W)

[0067] Among them, AdaptiveAvgPool2d is adaptive average pooling, the output height or width after pooling is 1, and other dimensions are adaptively adjusted. x h Performs adaptive average pooling on the input feature in the height dimension, and the obtained result represents the information of the image in the height direction. x w Performs adaptive average pooling on the input feature in the width dimension, and the obtained result represents the information of the image in the width direction.

[0068] 3) Through broadcast addition, use the horizontal axis vector and the vertical axis vector to model the enhanced edge information.

[0069] x gather = x h + x w

[0070] Among them, x gather represents x w and x h are added together to aggregate the information of the image in the horizontal and vertical directions to capture more comprehensive global context information.

[0071] 4) Through the shape self-calibration function, calibrate the region of interest to make it closer to the foreground object. Use two large kernel strip convolutions to decouple the calibration of the attention map in the horizontal and vertical directions, including: use the horizontal strip convolution to calibrate the shape in the horizontal direction, adjust each row of elements to make the horizontal shape closer to the foreground object. Normalize the features with BN and use the ReLU activation function for non-linear processing. Use the vertical strip convolution to calibrate the shape in the vertical direction:

[0072] ge = Sigmod(Conv2d(ReLU(BatchNorm(x gather ))))

[0073] Among them, ReLU represents the activation function, which uses non-linear activation. BatchNorm represents batch normalization, which helps to stabilize the training process. Sigmod represents the activation function, which compresses the output value into the range of [0, 1]. Conv2d represents the convolution operation, which is used to extract features.

[0074] Through this method, the convolutions in two directions can be decoupled, adapted to any shape, and ensure more accurate calibration of the region of interest.

[0075] 5) Add the processed edge information to the original input feature map to form the enhanced output:

[0076] out = ge + x

[0077] 3. Align the features extracted at different scales to the same scale through interpolation operations, splice them together, and finally fuse them into a unified feature representation through a convolutional layer, which can improve the model's perception of multi-scale features:

[0078] out i = ees i (Interpolate(features i (x), size=(H, W), mode

[0079] = bilinear, align comers = True))

[0080] out = final conv (concat(out0, out1, …, out n ))

[0081] Among them, interpolate represents bilinear interpolation for each feature map obtained through feature extraction to make its size the same as the input; then it is processed through the edge enhancer ees i out0 is the feature map generated through local convolution, and the remaining out i are the feature maps obtained through multi-scale edge enhancement; then, a convolutional layer is applied to output the final enhanced feature map.

[0082] Application Example 1:

[0083] In the backbone architecture of RT-DETR, the network performs downsampling operations at scales of 2×, 4×, 8×, 16×, and 32× to generate five groups of feature maps with different resolutions. Taking an initial image of 640×640 pixels as an example, these feature maps are scaled down to 320×320, 160×160, 80×80, 40×40, and 20×20 pixels respectively. RT-DETR can detect objects with sizes of 8×8, 16×16, and 32×32 pixels in the original image by integrating the P3, P4, and P5 layers.

[0084] The backbone network of RT-DETR adopts a series of convolutional and deconvolutional layers, and at the same time uses residual connections and bottleneck structures to reduce the size of the network and improve performance. Among them, the C2f module, as the basic component of the backbone network, enables the model to better capture complex features in the image, thus achieving better results in object detection tasks. Specifically, the C2f module first transforms the input feature map through a convolutional layer (usually a 1x1 convolution) to extract basic features and increase the feature expression ability of the model. The generated intermediate feature map is split into two parts. One part is directly passed to the final Concat block, and the other part is passed to multiple Bottleneck blocks for further processing. These Bottleneck blocks capture more complex patterns and details through a series of convolutional, normalization, and activation operations, thereby enhancing the features. The processed feature map is concatenated with the directly passed feature map in the Concat block to achieve feature fusion. This enables the model to comprehensively utilize multi-scale and multi-level information, which helps to improve the accuracy and robustness of the model. The concatenated feature map is then processed by a convolutional block to generate the final output feature map, providing a rich feature representation for subsequent detection and classification tasks.

[0085] In this embodiment, the cross-stage multi-scale edge information enhancement module (C2f-MSEIE module) is used to replace the C2f module in the backbone network of the DETR model to construct the MSEIE-DETR model, and the structure is as shown in Figure 3;

[0086] The C2f-MSEIE module in this embodiment:

[0087] First, multi-scale pooling is performed through adaptive average pooling to extract local information of different sizes for capturing multi-level features of the image.

[0088] Then, the edge enhancer is used to extract edge information, making the network more sensitive to edges.

[0089] Next, the features extracted at different scales are aligned to the same scale through interpolation operations and then concatenated together;

[0090] Finally, it is fused into a unified feature representation through the convolutional layer, which can improve the model's perception of multi-scale features.

[0091] Among them, the edge enhancer extracts and enhances the edge information in the input feature map, improving the model's sensitivity and expressive ability to image edges. Specifically, the edge enhancer first uses two-dimensional average pooling to smooth the input feature map and extract its low-frequency information. Then, the original input feature map is subtracted from the smoothed feature map to obtain the enhanced edge information. Rectangular self-calibration is performed on the enhanced edge information. Rectangular self-calibration captures the axial global context in two directions through horizontal pooling and vertical pooling, generating two different axis vectors. Through broadcast addition, these two vectors can effectively model the enhanced edge information. A shape self-calibration function is designed to calibrate the region of interest to make it closer to the foreground object.

[0092] Calibrate the region of interest through the shape self-calibration function: First, use horizontal bar convolution to calibrate the shape in the horizontal direction, adjusting each row of elements to make the horizontal shape closer to the foreground object.

[0093] Then, normalize the features with BN and perform non-linear processing using the ReLU activation function. Subsequently, use vertical bar convolution to calibrate the shape in the vertical direction. In this way, the convolutions in two directions can be decoupled to adapt to any shape, ensuring more accurate calibration of the region of interest.

[0094] Finally, add the processed edge information to the original input feature map to form the enhanced output.

[0095] Application Example 2:

[0096] The present invention applies the cross-stage multi-scale edge information enhancement module to the YOLO object detection network to construct the MSEIE-YOLO model, and the model architecture is as Figure 4 shown.

[0097] Combining Application Example 1 and Application Example 2 with real experimental data to verify the beneficial effects of the present invention:

[0098] Using standard evaluation metrics: average precision (AP), AP@50, AP@75, AP-s, AP-m, and AP-l. In pycocotools, these metrics can provide in-depth understanding of the detector's performance at different IoU thresholds and object sizes.

[0099] Use the "Real-world Underwater Object Detection" dataset. The real-world underwater object detection dataset is characterized by marine biodiversity and complex underwater environments. It contains 14,000 high-resolution images, of which 9,800 images are dedicated to training and 4,200 images are used for testing. These images contain a total of 74,903 annotated objects in 10 common aquatic categories, making it a suitable dataset for evaluating the performance of detectors in real-world scenarios.

[0100] In the comparative experiment, the parameter volumes of the MSEIE-DETR model and the MSEIE-YOLO model match that of R18. To ensure the complete convergence of the models, all models are set to train for 300 epochs, and all other parameters are set to their default values.

[0101] The experimental results are shown in Table 1. As can be seen from Table 1:

[0102] The AP of MSEIE-DETR combined with the present invention is increased by 2%, while the computational load is 70% of ResNet-18. The AP of MSEIE-YOLO is increased by 0.7%, while the computational load is 90% of ResNet-18. To visually compare the performance of each model, as Figure 5 shown, six challenging scenarios are selected. Column A is the original underwater image, column B is the detection of the underwater image by the traditional RT-DETR method, and column C is the detection of the underwater image by the MSEIE-DETR model in the application implementation of the present invention. It can be intuitively seen that the improved method of the present invention helps to reduce the cases of false detection and missed detection compared with the traditional method, and improves the accuracy of object detection.

[0103] As Figure 6 shown, heatmaps are used to show the attention of the traditional RT-DETR method and the MSEIE-DETR model to the feature importance in the data. Column A is the original underwater image, column B is the heatmap of the traditional RT-DETR method for the underwater image, and column C is the heatmap of the MSEIE-DETR model for the underwater image. The color depth (red indicates high importance, blue indicates low importance) can intuitively reflect the importance degree of each feature. The method of the present invention can better reflect the contour and shape features of the objects in the image compared with the traditional method, and can be combined with other features to form a richer feature vector, further improving the accuracy of object detection.

[0104] Table 1

[0105] Method <![CDATA[AP 0.50:0.95 > <![CDATA[AP 0.50 > <![CDATA[AP 0.75 > <![CDATA[AP S > <![CDATA[AP M > <![CDATA[AP L > Parameters(M) GFLOPS RT-DETR-R18 50.4 83.5 54.7 19.6 45.8 57.1 19.885 57.0 MSEIE-DETR 52.5 85.5 57.2 22.0 49.1 58.3 14.476 48.3 Yolov10n 49.5 82.2 53.3 23.1 46.9 53.5 2.267 6.5 MSEIE-yolov10n 50.1 83.0 54.0 24.6 47.7 54.4 2.112 6.0

[0106] By combining the multi-scale edge information enhancement module with YOLO and DETR, a new solution is provided for resource-constrained and high-real-time scenarios, thus significantly improving the performance. The progressiveness and effectiveness of the present invention compared with traditional methods are demonstrated.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-scale edge information enhancement method for underwater target detection, characterized in that, Including: Dividing the number of channels of the initial feature information into a main part and the remaining part by using a channel division operation; Inputting the main part into the main branch for feature extraction and enhancement through multiple multi-scale edge information enhancement units; Inputting the remaining part into the parallel gradient flow branch for transmission; Concatenating the output feature information of all branches to obtain a concatenated feature; Inputting the concatenated feature into a convolutional structure to obtain an edge information enhanced feature; Meanwhile, the channel division operation and multiple multi-scale edge information enhancement units are cross-connected.

2. The multi-scale edge information enhancement method for underwater target detection according to claim 1, wherein The multi-scale edge information enhancement unit performs feature extraction and enhancement, including: Extracting local information of different sizes through an adaptive average pooling layer to obtain feature maps at different scales; Extracting the edge information of the feature map through an edge enhancer to obtain an edge enhanced feature; Aligning and concatenating the edge enhanced features at different scales to obtain a fused feature; Inputting the fused feature into a convolutional layer to generate a unified feature representation.

3. A multi-scale edge information enhancement method for underwater target detection according to claim 1, characterized in that The extracting the edge information of the feature map through the edge enhancer to obtain an edge enhanced feature includes: Smoothing the input feature map through two-dimensional average pooling; Obtaining enhanced edge information based on the input feature map and the smoothed feature map; Performing rectangular self-calibration on the enhanced edge information to obtain calibrated enhanced edge information; Adding the calibrated enhanced edge information to the input feature map to obtain an edge enhanced feature.

4. A multi-scale edge information enhancement method for underwater target detection according to claim 1, characterized in that The performing rectangular self-calibration on the enhanced edge information includes: Performing horizontal pooling and vertical pooling on the enhanced edge information to generate a horizontal axis vector and a vertical axis vector; Constructing an edge information model based on the horizontal axis vector and the vertical axis vector through the broadcasting method; Calibrating the region of interest of the edge information model through a shape self-calibration function.

5. The multi-scale edge information enhancement method for underwater target detection according to claim 1, wherein, The calibrating the region of interest of the edge information model through the shape self-calibration function includes: Calibrating the shape in the horizontal direction using a horizontal bar convolution; Calibrating the shape in the vertical direction using a vertical bar convolution; And performing non-linear processing, batch normalization processing, Sigmoid activation, and convolution processing through a ReLU activation function to obtain calibrated enhanced edge information.

6. The multi-scale edge information enhancement method for underwater target detection according to claim 1, wherein Obtaining a concatenated feature by concatenating the output feature information of all branches through bilinear interpolation.

Citation Information

Cited By

  • Target segmentation method and device, electronic equipment and storage medium

    CN121458985A