Target detection method and system for underwater sonar image, medium and equipment
By introducing a cross-scale channel attention mechanism and multi-scale fusion network layer in the RT-DETR model, the noise interference and low resolution problems of underwater sonar images are solved, and the accuracy and robustness of target detection are improved, which is suitable for underwater autonomous detection and target recognition.
Patent Information
- Application Number
- CN202510387620.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
Underwater sonar images have strong noise interference, low image resolution and high target feature similarity, resulting in insufficient target detection accuracy and robustness of underwater robots.
The cross-scale channel attention mechanism and convolution-based multi-scale fusion network layer are added to the RT-DETR model to enhance feature extraction and fusion capabilities, and improve object detection accuracy through the cross-scale channel attention mechanism and multi-scale fusion network layer.
It significantly improves the accuracy and robustness of underwater sonar image object detection, adapts to complex underwater environments, and meets the needs of efficient and accurate detection of underwater robots.
Smart Images

Figure CN120259866A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image target detection, and in particular to a method, system, medium and device for target detection of underwater sonar images. Background Art
[0002] With the rapid development of underwater detection technology, underwater robots have become important tools in the fields of ocean exploration, environmental monitoring, underwater target recognition, etc. Among them, the sonar system, as one of the core sensors, can provide stable imaging capabilities in complex environments such as low visibility and turbid water bodies. The sonar generates an underwater environment image by emitting acoustic wave pulses and receiving echo signals, enabling the underwater robot to perceive the surrounding environment and achieve autonomous navigation, target recognition, and task execution. However, due to device noise, environmental reverberation, and the inherently low resolution characteristics of underwater images, sonar images often have problems such as strong noise interference, blurred target boundaries, and large variations in target scales, seriously affecting the target detection accuracy and robustness of underwater robots.
[0003] In the case of high noise, complex backgrounds, and multi-scale targets, although traditional image processing methods and detection algorithms based on convolutional neural networks (CNNs) perform well in general target detection tasks, they still face many challenges in underwater robot sonar imaging tasks. First, traditional methods are difficult to effectively separate foreground targets and background noise, resulting in easy loss of small targets and large detection errors. Second, the limitations of CNNs in sonar image processing are mainly reflected in the difficulty of accurately positioning targets with blurred boundaries. Especially in a dynamic underwater environment, target features are greatly affected by noise and environmental factors, and the detection accuracy and robustness of existing technologies are difficult to meet the high-precision detection requirements of underwater robots.
[0004] To solve these problems, researchers have proposed improvement schemes such as attention mechanisms and multi-scale feature fusion to enhance the network's ability to capture key features. Among them, the channel attention mechanism can improve the model's attention to effective target regions, and multi-scale feature fusion can extract more discriminative features from information at different levels. However, in a complex underwater environment, existing methods still have problems such as insufficient target discrimination, feature loss, and low computational efficiency when processing sonar images. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method, system, medium and device for target detection of underwater sonar images, which can solve the common problems of strong noise interference, low image resolution, and high similarity of target features.
[0006] To achieve the above object, the present invention is implemented by the following technical solutions:
[0007] On the one hand, the present invention provides a method for target detection of underwater sonar images, including:
[0008] Obtain the original underwater sonar image;
[0009] Input the original underwater sonar image into a pre-constructed target detection model, and output the target detection result of the underwater sonar image;
[0010] Among them, the construction of the target detection model includes:
[0011] Obtain the RT-DETR model;
[0012] Add a cross-scale channel attention mechanism before the basic block in the third stage of the backbone network in the RT-DETR model, add a cross-scale channel attention mechanism before the basic block in the fourth stage, and add a convolutional-based multi-scale fusion network layer before the output layer of the hybrid encoder in the RT-DETR model to obtain the constructed target detection model.
[0013] Optionally, the processing steps of the target detection model include:
[0014] In the backbone network, perform multi-layer convolution on the original underwater sonar image to obtain an input feature map, and perform feature extraction on the input feature map through a cross-scale channel attention mechanism to obtain an output feature map;
[0015] In the hybrid encoder, perform in-scale feature interaction on the output feature map to obtain interaction features, and perform cross-layer feature fusion on the interaction features through a convolutional-based multi-scale fusion network layer to obtain a fusion feature map;
[0016] In the decoder, perform decoding and reconstruction on the fusion feature map to obtain the target detection result of the underwater sonar image.
[0017] Optionally, the cross-scale channel attention mechanism includes an SE module, a dilated convolution module, and an EMA module connected in sequence, and the processing steps of the cross-scale channel attention mechanism include:
[0018] In the SE module, extract the channel features of the input feature map to obtain sonar channel features;
[0019] In the dilated convolution module, extract the dilated features of the sonar channel features to obtain sonar dilated features;
[0020] In the EMA module, extract the global features of the sonar dilated features to obtain an output feature map.
[0021] Optionally, performing feature extraction on the input feature map through a cross-scale channel attention mechanism to obtain an output feature map includes:
[0022] ;
[0023] ;
[0024] ;
[0025] wherein, represents the sonar channel feature; represents the sonar multi-scale feature; represents the output feature map; represents the original underwater sonar image; represents the channel feature extraction operation; , , respectively represent convolution operations with dilation rates of 1, 3, and 5; represents the feature concatenation operation; represents the global feature extraction operation.
[0026] Optionally, the convolutional multi-scale fusion network layer includes multiple sequentially connected Fusion modules, and the processing steps of each Fusion module include:
[0027] Perform convolution processing on the interaction feature to obtain a convolution feature;
[0028] Perform feature recombination on the convolution feature to obtain a recombined feature;
[0029] Fuse the convolution feature and the recombined feature to obtain a fused feature map.
[0030] Optionally, cross-layer feature fusion is performed on the interaction feature through the convolutional multi-scale fusion network layer to obtain a fused feature map, including:
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] ;
[0037] ;
[0038] wherein, , , represents the features first fused by the Fusion module; , , represents the features fused by the Fusion module again; , , represents the output features of the basic blocks at different stages of the backbone network; represents the features extracted by the self-attention mechanism in the encoder; represents upsampling; represents downsampling; represents the feature concatenation operation; represents the fusion operation; represents the ordinary convolution operation; represents the reparameterized convolution operation.
[0039] In a second aspect, the present invention provides an underwater sonar image target detection system, including:
[0040] An image acquisition module that acquires the original underwater sonar image;
[0041] A target detection module that inputs the original underwater sonar image into a pre-constructed target detection model and outputs the target detection result of the underwater sonar image;
[0042] A model construction module, wherein the construction of the target detection model includes:
[0043] Obtain the RT-DETR model;
[0044] Add a cross-scale channel attention mechanism before the basic block in the third stage of the backbone network in the RT-DETR model, add a cross-scale channel attention mechanism before the basic block in the fourth stage, and add a convolutional multi-scale feature fusion network layer before the output layer of the hybrid encoder in the RT-DETR model to obtain the constructed target detection model.
[0045] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0046] In a fourth aspect, the present invention provides a computer device, including a processor and a storage medium;
[0047] The storage medium is used to store instructions;
[0048] The processor is used to operate according to the instructions to execute the method described in the first aspect.
[0049] Beneficial effects achieved by the present invention compared with the prior art:
[0050] The present invention adds a cross-scale channel attention module to the backbone network, which combines the advantages of SE and EMA to dynamically adjust the weights of feature channels and capture multi-scale spatial dependencies, enabling the model to more accurately focus on target features, thereby improving the recognition and localization accuracy of targets in environments with complex noise and backgrounds; a multi-scale feature fusion module based on CNN is added to the hybrid encoder, introducing more low-level features, fusing low-level geometric information and high-level semantic information, thereby improving the ability to capture details in low-resolution sonar images. Through cross-layer connections and feature interactions, the low-level geometric information and high-level semantic features are more effectively combined, thereby significantly improving the detection accuracy of multi-scale targets, meeting the requirements of underwater robots for efficient and accurate detection, and providing reliable technical support for underwater autonomous detection, target recognition, and environmental perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flowchart of the target detection method for underwater sonar images of the present invention in one embodiment;
[0052] Figure 2 It is a schematic structural diagram of the SE module of the present invention in one embodiment;
[0053] Figure 3 It is a schematic structural diagram of the EMA module of the present invention in one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0054] The technical solution of the present invention will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0055] The term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0056] Embodiment 1
[0057] As Figure 1As shown in the figure, this embodiment introduces a method for target detection of underwater sonar images, aiming to improve the detection accuracy and robustness of underwater robots for targets in complex environments, especially applicable to application scenarios of high-noise and weak-feature target detection, capable of adapting to a variety of underwater robot platforms, and providing an efficient and accurate intelligent detection solution for underwater autonomous detection, target recognition, and environmental perception.
[0058] The method specifically includes the following steps:
[0059] Step 1: Obtain the original underwater sonar image.
[0060] Step 2: Input the original underwater sonar image into a pre-constructed target detection model, and output the target detection result of the underwater sonar image. Specifically:
[0061] In terms of model design, as Figure 1 shown in the figure, this embodiment improves on the framework of the Real-Time Detection Transformer (RT-DETR) model. After the basic blocks in the third stage and the fourth stage of the backbone network, a Cross-scale Channel Attention (CCA) mechanism is added to enhance the feature extraction process in the third and fourth stages of the backbone network. In terms of feature fusion, this embodiment adds a CNN-based Multi-scale Feature Fusion (CMFF) network layer before the output layer of the Hybrid Encode to further optimize the feature extraction and fusion process, obtaining the constructed target detection model.
[0062] The cross-scale attention mechanism sequentially includes a Squeeze-and-Excitation (SE) module, a dilated convolution module, and an Efficient Multi-scale Attention (EMA) module; the structures of the SE module and the EMA module are respectively as Figure 2 and Figure 3 shown in the figure.
[0063] In the backbone network, perform multi-layer convolution on the original underwater sonar image to obtain an input feature map, and perform feature extraction on the input feature map through CCA to obtain an output feature map;
[0064] As Figure 2As shown, the SE module extracts the channel features of the input feature map, adaptively enhances the key channel features, improves the attention to important information, and obtains the sonar channel features; combines the dilated convolution to extract the dilated features of the sonar channel features, effectively improves the receptive field of the CCA, so as to improve the detection ability of the underwater robot for targets of different scales, and obtains the sonar dilated features; as Figure 3 As shown, the EMA module extracts the global features of the sonar dilated features, through the cross-scale attention mechanism, uses feature interaction to capture the global context information, so as to effectively cope with the noise interference and background complexity in the sonar image, and obtains the output feature map; then the calculation formula of the CCA is as follows:
[0065] ;
[0066] ;
[0067] ;
[0068] Among them, represents the sonar channel features; represents the sonar multi-scale features; represents the output feature map; represents the original underwater sonar image, , represents the real number space, represents the number of channels, height, and width; represents the channel feature extraction operation; , , respectively represent the convolution operations with dilation rates of 1, 3, and 5; represents the feature concatenation operation; represents the global feature extraction operation.
[0069] After the CCA is integrated into the backbone network, even in the case of low signal-to-noise ratio and complex background, it can still effectively strengthen the target feature expression and improve the detection accuracy.
[0070] In the hybrid encoder, the intra-scale feature interaction is performed on the output feature map to obtain the interaction features, and the cross-layer feature fusion is performed on the interaction features through the CMFF to obtain the fused feature map;
[0071] The CMFF includes multiple sequentially connected Fusion modules. In each Fusion module, the interaction features are convolutionally processed through a 1×1 convolution to obtain the convolution features; the convolution features are feature recombined through 3 reparameterized convolutions to obtain the recombined features; the convolution features and the recombined features are feature fused through a residual connection to obtain the fused feature map. In this embodiment, the number of Fusion modules is 6.
[0072] In CMFF, each Fusion module uniformly scales features of different scales to a fixed size through upsampling and downsampling. Finally, through the Concat operation, the features of each layer are concatenated by channels into a new feature as the input of the next Fusion module, that is, it receives the features processed in the second to fifth stages (P2 - P5) of the backbone network. By enhancing the information flow of features at different levels, it optimizes the network's detection ability for low - resolution and highly similar targets, and performs particularly well in the case of small targets and complex backgrounds; the combination of cross - layer connection and feature interaction effectively integrates low - level geometric details and high - level semantic information, thus significantly improving the detection accuracy.
[0073] Let the features extracted by the backbone network at different stages be represented as The features after being processed by the Fusion module for the first time are represented by The features after being processed by the Fusion module again are represented by where i represents the i - th level of features. Then the feature fusion process in CMFF can be described by the following formula:
[0074] ;
[0075] ;
[0076] ;
[0077] ;
[0078] ;
[0079] ;
[0080] ;
[0081] where, , , represent the features fused by the Fusion module for the first time; , , represent the features fused by the Fusion module again; , , represent the output features of the basic blocks at different stages of the backbone network; represents the features extracted by the self - attention mechanism in the encoder; represents upsampling; represents downsampling; represents the feature concatenation operation; represents the fusion operation; represents a normal convolution operation; represents a reparameterized convolution operation.
[0082] CMFF fuses low-level geometric details (edges, textures) and high-level semantic information (object categories, structures) through cross-layer connections, optimizes the feature representations of objects at different scales, and the rich skip connections enhance the information flow of multi-scale features, improving the detection ability for small objects and complex backgrounds.
[0083] In the decoder, the fused feature map is decoded and reconstructed to obtain the object detection results of the underwater sonar image.
[0084] To evaluate the model performance, mAP (mAP0.5) at IoU = 0.5 is used as the evaluation metric, and the mAP values of different categories at this threshold are calculated to comprehensively measure the detection accuracy; the results show that this embodiment has achieved significant performance improvement on the SCTD dataset, especially in complex underwater environments, demonstrating stronger object detection ability, being able to adapt to a variety of underwater robot platforms, and providing an efficient and accurate intelligent detection solution for underwater autonomous detection, object recognition, and environmental perception.
[0085] Embodiment 2
[0086] Based on Embodiment 1, this embodiment introduces an object detection system for underwater sonar images, including:
[0087] An image acquisition module that acquires the original underwater sonar image;
[0088] An object detection module that inputs the original underwater sonar image into a pre-constructed object detection model and outputs the object detection results of the underwater sonar image;
[0089] A model construction module, where the construction of the object detection model includes:
[0090] Obtain the RT-DETR model;
[0091] Add a cross-scale channel attention mechanism after the basic block in the third stage of the backbone network in the RT-DETR model, add a cross-scale channel attention mechanism after the basic block in the fourth stage, and add a convolutional multi-scale fusion network layer before the output layer of the hybrid encoder in the RT-DETR model to obtain the constructed object detection model.
[0092] For the specific function implementation of the above modules, refer to the relevant content in the method of Embodiment 1, which will not be elaborated here.
[0093] Embodiment 3
[0094] This embodiment introduces a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.
[0095] Embodiment 4
[0096] This embodiment introduces a computer device, including a processor and a storage medium;
[0097] The storage medium is used to store instructions;
[0098] The processor is used to operate according to the instructions to execute the method described in Embodiment 1.
[0099] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing in the processFigure 1 One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes.
[0103] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for target detection of underwater sonar images, characterized in that, Including: Obtain the original underwater sonar image; Input the original underwater sonar image into a pre-constructed object detection model to output the object detection result of the underwater sonar image; Among them, the construction of the object detection model includes: Obtain the RT-DETR model; Add a cross-scale channel attention mechanism after the basic block in the third stage of the backbone network in the RT-DETR model, add a cross-scale channel attention mechanism after the basic block in the fourth stage, and add a convolutional-based multi-scale fusion network layer before the output layer of the hybrid encoder in the RT-DETR model to obtain the constructed object detection model.
2. The target detection method for underwater sonar images according to claim 1, wherein, The processing steps of the object detection model include: In the backbone network, perform multi-layer convolution on the original underwater sonar image to obtain an input feature map, and perform feature extraction on the input feature map through the cross-scale channel attention mechanism to obtain an output feature map; In the hybrid encoder, perform in-scale feature interaction on the output feature map to obtain interaction features, and perform cross-layer feature fusion on the interaction features through the convolutional-based multi-scale fusion network layer to obtain a fusion feature map; In the decoder, perform decoding and reconstruction on the fusion feature map to obtain the object detection result of the underwater sonar image.
3. The object detection method for underwater sonar images according to claim 2, wherein The cross-scale channel attention mechanism includes an SE module, a dilated convolution module, and an EMA module connected in sequence. The processing steps of the cross-scale channel attention mechanism include: In the SE module, extract the channel features of the input feature map to obtain sonar channel features; In the dilated convolution module, extract the dilated features of the sonar channel features to obtain sonar dilated features; In the EMA module, extract the global features of the sonar dilated features to obtain an output feature map.
4. The target detection method for underwater sonar images according to claim 3, characterized in that, Performing feature extraction on the input feature map through the cross-scale channel attention mechanism to obtain an output feature map includes: ; ; ; Among them, represents the sonar channel feature; represents the sonar multi-scale feature; represents the output feature map; represents the original underwater sonar image; represents the channel feature extraction operation; 、 、 respectively represent the convolution operations with dilation rates of 1, 3, and 5; represents the feature concatenation operation; represents the global feature extraction operation.
5. The target detection method for underwater sonar images according to claim 2, characterized in that, The convolutional-based multi-scale fusion network layer includes multiple sequentially connected Fusion modules. The processing steps of each Fusion module include: Perform convolutional processing on the interaction features to obtain convolutional features; Perform feature recombination on the convolutional features to obtain recombined features; Perform feature fusion on the convolutional features and the recombined features to obtain a fusion feature map.
6. The method for target detection of an underwater sonar image according to claim 5, wherein Performing cross-layer feature fusion on the interaction features through the convolutional-based multi-scale fusion network layer to obtain a fusion feature map includes: ; ; ; ; ; ; ; Among them, , , represent the features first fused by the Fusion module; , , represent the features fused by the Fusion module again; , , represent the output features of the basic blocks at different stages of the backbone network; represents the features extracted by the self-attention mechanism in the encoder; represents upsampling; represents downsampling; represents the feature concatenation operation; represents the fusion operation; represents the ordinary convolution operation; represents the reparameterized convolution operation.
7. An object detection system for underwater sonar images, characterized in that, Including: An image acquisition module that obtains the original underwater sonar image; An object detection module that inputs the original underwater sonar image into a pre-constructed object detection model and outputs the object detection result of the underwater sonar image; A model construction module, where the construction of the object detection model includes: Obtain the RT-DETR model; Add a cross-scale channel attention mechanism after the basic block in the third stage of the backbone network in the RT-DETR model, add a cross-scale channel attention mechanism after the basic block in the fourth stage, and add a convolutional-based multi-scale fusion network layer before the output layer of the hybrid encoder in the RT-DETR model to obtain the constructed object detection model.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer device, characterized in that, It includes a processor and a storage medium; The storage medium is used for storing instructions; The processor is used for operating according to the instructions to execute the method according to any one of claims 1 to 6.
Citation Information
Cited By
Underwater target detection method based on RT-SonarNet
CN121982507A