Underwater target detection method and device

By enhancing and fusing underwater sonar images and combining channel attention and multi-scale context attention mechanisms, the accuracy problem of underwater target detection models under water scattering and noise interference is solved, and high-precision underwater target detection is achieved.

CN120747723APending Publication Date: 2025-10-03JIMEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510818063.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing CNN-based underwater target detection models cannot effectively handle water scattering and noise interference in underwater images, resulting in poor target detection accuracy, especially in low-contrast and edge-blurred conditions, and are difficult to adapt to multi-scale data.

Method used

The enhancement processing module is used to perform CLAHE enhancement, non-local mean denoising and frequency domain bandpass filtering on the sonar image. The channel attention mechanism and multi-scale context attention mechanism are combined to perform feature fusion and adaptive noise suppression, and the target detection accuracy is improved through hierarchical void convolution.

Benefits of technology

It improves the accuracy of underwater target detection and the sensitivity to faint targets, enhances the model's target recognition ability in low-contrast and severe noise interference environments, and improves feature extraction accuracy and detection speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747723A_ABST
    Figure CN120747723A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to an underwater target detection method and device, and the method comprises the steps: determining an initial parameter and a sonar image of an underwater target detection model; performing enhancement processing on the sonar image by adopting an enhancement processing module of the target detection model to obtain a processed image; performing feature fusion on the processed image based on a channel attention mechanism to obtain fused features; carrying out adaptive noise suppression and hierarchical cavity convolution on the fused features to obtain a target area; and performing underwater target detection on the sonar image based on the target area. According to the method, the sensitivity of a target detection model to weak target and detail information processing can be improved through feature fusion, and target information under different scales can be better captured by adopting a multi-scale feature extraction and fusion strategy, so that the target detection precision can be improved while the detection speed is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to an underwater target detection method and device in the field of machine learning technology. Background Art

[0002] In related technologies, Synthetic Aperture Sonar (SAS) images are usually identified and detected using CNN-based methods such as the target detection algorithm (You Only Look Once, YOLO), Single Shot MultiBox Detector (SSD), and Faster Region Convolutional Neural Network (Faster R-CNN). These CNN-based deep learning models have great advantages in the field of image recognition.

[0003] Models and algorithms used in related technologies often fail to effectively address issues such as water scattering and noise interference in underwater image data. Underwater sonar images often exhibit low contrast between targets and backgrounds, with blurred edges. Multiple reflections of sound waves can lead to speckle noise and artifacts in the images. Furthermore, underwater targets vary greatly in size within images, making models generally inadequately adapted to multi-scale data, resulting in poor underwater target detection accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and device for underwater target detection, and the technical solutions adopted are as follows:

[0005] In a first aspect, an embodiment of the present invention provides an underwater target detection method, the method comprising:

[0006] Determine the initial parameters of the underwater target detection model and sonar image;

[0007] Using the enhancement processing module of the target detection model, the sonar image is enhanced to obtain a processed image;

[0008] Performing feature fusion on the processed image based on a channel attention mechanism to obtain fused features;

[0009] Performing adaptive noise suppression and hierarchical dilated convolution on the fused features to obtain the target area;

[0010] Based on the target area, underwater target detection is performed on the sonar image.

[0011] In a second aspect, an embodiment of the present invention provides an underwater target detection device, the device comprising:

[0012] A determination module, used to determine the initial parameters of the underwater target detection model and the sonar image;

[0013] an enhancement module, configured to perform enhancement processing on the sonar image using the enhancement processing module of the target detection model to obtain a processed image;

[0014] A fusion module, configured to perform feature fusion on the processed image based on a channel attention mechanism to obtain fused features;

[0015] A convolution module is used to perform adaptive noise suppression and hierarchical dilated convolution on the fused features to obtain a target area;

[0016] A detection module is used to perform underwater target detection on the sonar image based on the target area.

[0017] In a third aspect, a computer program product is provided, which includes: a computer program code, which, when executed on a computer, causes the computer to execute the method in the first aspect.

[0018] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code is run on a computer, the computer executes the method in the first aspect.

[0019] The present invention has the following beneficial effects: After determining the initial parameters of the underwater target detection model and the sonar image, the sonar image is enhanced using the target detection model's enhancement processing module to obtain a processed image. This enhances the processed image's anti-interference capabilities, thereby further improving the accuracy of subsequent feature extraction. Subsequently, the processed image is subjected to feature fusion based on a channel attention mechanism to obtain fused features, and the fused features are subjected to adaptive noise suppression and hierarchical dilated convolution to obtain the target region; underwater target detection is performed on the sonar image using the target region. In this way, feature fusion can enhance the target detection model's sensitivity to processing faint targets and detailed information. Because underwater targets vary greatly in size within an image, a multi-scale feature extraction and fusion strategy is employed to better capture target information at different scales, thereby improving target detection accuracy while ensuring detection speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 This is a schematic diagram of an implementation flow of an underwater target detection method provided by an embodiment of the present invention;

[0022] Figure 2 This is another implementation flowchart of an underwater target detection method provided by an embodiment of the present invention;

[0023] Figure 3 This is another implementation flowchart of an underwater target detection method provided by an embodiment of the present invention;

[0024] Figure 4 1 is a schematic diagram of an implementation framework of an underwater target detection method provided by an embodiment of the present invention;

[0025] Figure 5 1 is a schematic diagram of the structure of an underwater target detection device provided by an embodiment of the present invention;

[0026] Figure 6 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of an underwater target detection method proposed in accordance with the present invention. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0028] In the description of the embodiments of the present invention, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present invention, "multiple" refers to two or more than two.

[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0030] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0031] Synthetic aperture sonar technology is an underwater imaging technique that uses the motion of a sonar platform to create a virtual, ultra-large aperture. Using a small sonar array, the system continuously transmits and receives sound waves while in motion. By precisely recording the platform's position and echo phase, the signals received along the way are coherently superimposed and processed, creating an effect equivalent to an extremely long physical aperture. This technology achieves clear, high-resolution imaging at long distances without blurring with distance, making it particularly suitable for the precise detection of targets such as mines, pipelines, and sunken ships. Due to its strong acoustic penetration and excellent resistance to environmental interference, sonar technology has become an indispensable tool for underwater detection. By actively transmitting sound waves and analyzing the echo signals, sonar systems can effectively construct information about the contours, textures, and spatial distribution of underwater targets, providing critical data for target detection and localization in complex scenarios. With the increasing demand for rapid underwater sonar image data, the speed and accuracy requirements for detection methods are also increasing. Furthermore, sonar acoustic reflection imaging poses greater challenges for high-precision target detection due to multiple reflections, significant speckle noise caused by environmental noise, and the shadow areas behind the targets created by the sonar side-scan characteristic. Deep learning has developed rapidly in recent years. A series of deep learning models such as Convolutional Neural Network (CNN) and Transformer architecture have demonstrated superior performance in processing image and video data, and are often used to solve tasks such as computer vision and natural language processing.

[0032] The present invention provides an underwater target detection method. First, to address the low background contrast and blurred edges in underwater sonar images, an enhancement processing module and data augmentation strategy are designed. This strategy combines three enhancement processing modules, namely contrast-limited adaptive histogram equalization (CLAHE), non-local mean denoising, and frequency-domain bandpass filtering, into a three-stage cascade to enhance image contrast and address edge blurring. A C2fD module is then proposed to fuse differential features with basic features, ensuring the model has stronger detail capture capabilities. This module uses spatial differential operations to extract target edge information to address edge blur and texture loss in sonar images. Furthermore, an efficient channel attention mechanism and a lightweight feature fusion strategy are used to prevent imbalance between basic and edge features, effectively enhancing the model's target recognition capabilities in low-contrast, noisy underwater environments. Finally, an underwater multi-scale contextual attention mechanism is designed to further enhance the model's sensitivity to faint targets and better adapt to multi-scale underwater targets.

[0033] The following is a detailed description of a method for underwater target detection provided by the present invention with reference to the accompanying drawings. Figure 1 , which shows a schematic diagram of an implementation flow of an underwater target detection method provided by an embodiment of the present invention, the method comprising:

[0034] 101, determine the initial parameters of the underwater target detection model and the sonar image.

[0035] Here, the basic parameters of the dataset are determined, including the size, input batch size, and important parameters in the corresponding design process. Among them, the important parameters in the design process include the limit threshold of CLAHE contrast enhancement, the intensity parameter of non-local means denoising, and the radius of the frequency domain bandpass filter.

[0036] 102. Use the enhancement processing module of the target detection model to perform enhancement processing on the sonar image to obtain a processed image.

[0037] Here, the enhancement processing module performs CLAHE, denoising, filtering and other processing on the sonar image to obtain a processed image.

[0038] In some possible implementations, first, the sonar image is enhanced using the enhancement processing module to generate a visual image; second, the visual image is subjected to non-local mean denoising to obtain a denoised image; finally, the denoised image is subjected to multi-dimensional denoising to obtain the processed image.

[0039] Here, after introducing the enhancement processing module (for example, the CLAHE enhancement module), the parameters of the enhancement processing module are debugged and a visualization image is generated to observe the image changes before and after the introduction of the enhancement processing module, so as to optimize the visualization image. That is, the visualization drawings are compared and the clearest visualization image after preprocessing is selected.

[0040] Image enhancement with CLAHE also makes noise more noticeable, so non-local means denoising is used. This module, combined with the CLAHE enhancement module, forms a collaborative "enhancement-cleanup" process. Frequency-domain bandpass filtering is also used to address noise from different dimensions, creating a dual noise reduction mechanism of "spatial-domain cleanup and frequency-domain focusing."

[0041] Because the related technology has serious problems of noise superposition and target distortion when directly stitching low-quality sonar images, a three-level preprocessing (i.e., CLAHE, denoising, and filtering) is embedded before stitching the sonar images to ensure the clarity and target integrity of the stitched sub-images. This not only ensures the diversity of the original data, but also further enhances the target features and anti-interference ability in the image through preprocessing, greatly improving the feature extraction accuracy of the subsequent network.

[0042] 103. Perform feature fusion on the processed image based on a channel attention mechanism to obtain fused features.

[0043] Here, the fused features are obtained by performing spatial differential feature extraction on the processed image and introducing a channel attention mechanism for feature fusion.

[0044] In some possible implementations, the above step 103 can be performed by Figure 2 The steps shown achieve:

[0045] 201 , performing feature extraction on the processed image based on spatial difference to obtain basic features and differential features.

[0046] Here, in order to address the problems of blurred target edges, low contrast, and background noise interference in underwater sonar images, the convolutional neural network in the related art relies on adaptive convolution kernels to implicitly extract edge features, which is easily contaminated by noise and leads to feature confusion. In an embodiment of the present invention, an explicit spatial gradient enhancement module (Optimized Spatial Difference) is introduced. Through three-stage operations of channel compression, bidirectional gradient extraction, and feature expansion, basic features and differential features are obtained, thereby improving the model's ability to model the edge features of sonar images. This part suppresses the channel-specific noise caused by scattering in the sonar image through cross-channel information fusion, while reducing the amount of calculation to 1 / C of the original input (where C represents the number of channels). In this way, the intensity of the background noise response can be reduced. In addition to compressing the input feature channels, a bidirectional gradient feature extraction method is also used to explicitly enhance the horizontal and vertical edge responses. A Sobel convolution kernel with fixed weights is used to calculate the horizontal and vertical gradients.

[0047] 202. Using a channel attention mechanism, the basic features and the differential features are fused to obtain the fused features.

[0048] Here, first, by adopting the channel attention mechanism, the weights of the basic features and differential features are adjusted to obtain adjusted basic features and adjusted differential features; for example, a dynamic feature fusion module is introduced into the underwater target detection model. The dynamic feature fusion module realizes the adaptive fusion of basic features and differential features through the dynamic allocation mechanism of attention weights, which can reduce the amount of calculation while improving the accuracy of small target detection.

[0049] Secondly, the adaptive fusion module of the underwater target detection model is used to perform feature fusion on the adjusted basic features and the adjusted differential features to obtain the fused features.

[0050] In some possible implementations, channel compression is performed on the basic features and differential features to obtain compressed basic features and compressed differential features; bidirectional gradient extraction and feature expansion are then performed on the compressed basic features and compressed differential features to obtain expanded basic features and expanded differential features; finally, a channel attention mechanism is used to adjust the weights of the bidirectional gradient extraction and feature expansion to obtain the adjusted basic features and the adjusted differential features.

[0051] Here, because spatial differentiation may enhance noise edges, but the dynamic feature fusion module can only address the problem of invalid responses, a channel attention mechanism is introduced to suppress noisy channel responses through global channel weights. While the Efficient Channel Attention mechanism can achieve lightweight channel attention, its use of one-dimensional convolution for local channel interactions makes it difficult to effectively model global channel relationships and is sensitive to high-frequency noise in sonar images. This embodiment of the present invention proposes Enhanced Efficient Channel Attention (ECA). This first models global channel interactions, using fully connected layers (FCs) to construct channel interaction paths and perform channel dimensionality reduction on global average pooled features. Secondly, a SiLU-Sigmoid hybrid activation strategy is employed to avoid the vanishing gradient problem that occurs when ECA uses only the Sigmoid function in the output layer. Subsequently, an adjustable compression ratio mechanism is introduced, using a reduction parameter (i.e., r value) to explicitly control the channel compression ratio. For example, by adjusting the r value (which can be achieved through the reduction parameter), a flexible trade-off between model lightweightness and feature expressiveness can be achieved. When the r value increases (such as r = 8), the computational complexity is reduced but high-frequency detail information may be lost; when the r value decreases (such as r = 4 in the C2fD module), more channel interaction information is retained to improve the accuracy of small target detection.

[0052] 104, adaptive noise suppression and hierarchical dilated convolution are performed on the fused features to obtain the target area.

[0053] To prevent contextual data loss, we introduce the Underwater Attention (UWA) mechanism. Adaptive noise suppression and hierarchical dilated convolutions are used to capture multi-scale context, and then channel-spatial attention is combined to collaboratively enhance the target region response, thereby identifying the target region. Dynamically gated residual fusion balances the contributions of original and enhanced features, improving the model's sensitivity to faint targets.

[0054] In some possible implementations, the above step 104 can be performed by Figure 3 The steps shown achieve:

[0055] 301 : Suppress high-frequency scattered noise in the fused features to obtain suppressed features.

[0056] Here, an adaptive noise suppression module is introduced. This module suppresses the residual high-frequency scattered noise in the fused features through a grouped convolution-dual-path denoising unit to obtain suppressed features, providing a preliminary purified feature map for subsequent modules and reducing the interference of noise on multi-scale context modeling.

[0057] 302. Determine a convolution feature of the suppressed feature.

[0058] Here, a multi-granularity receptive field parallel branch is used to perform hierarchical dilated convolution on the suppressed features to obtain convolution features containing multi-granularity contextual information. In some possible implementations, a multi-granularity receptive field parallel branch is first constructed, and then the hierarchical dilated convolution is performed on the suppressed features to obtain convolution features containing multi-granularity contextual information. The dilation rate [d1, d2, d3] of the hierarchical dilated convolution can be set to [1, 2, 3]. For example, in order to capture contextual information of different scales, a multi-granularity receptive field parallel branch is constructed to implement hierarchical dilated convolution, generate features containing multi-granularity contextual information, and provide rich spatial and semantic information for the subsequent two-dimensional attention.

[0059] 303 , performing attention coordination on the convolutional features to obtain coordinated features.

[0060] Here, the attention coordination mechanism of the channel dimension and the spatial dimension is used to perform cross-dimensional feature enhancement on the convolution feature to obtain the enhanced feature; and the cross-dimensional attention mechanism is used to suppress the noise points and noise channels in the enhanced feature to obtain the coordinated feature.

[0061] In some possible implementations, in order to further suppress background noise and enhance the response of the target area, the attention coordination of the two dimensions of channel and space is utilized to achieve cross-dimensional feature enhancement. The cross-dimensional attention is used to suppress noise channels and discrete noise points to improve the saliency of the target area.

[0062] 304 : Using a learnable gating coefficient to balance the coordinated features to obtain the target region.

[0063] Here, the coordinated features are balanced using a learnable gating coefficient to obtain a target feature, and the target area corresponding to the target feature is determined in the sonar image. In some implementations, excessive noise removal can lead to loss of detail. Therefore, to balance noise suppression with target feature preservation, a dynamic gated residual fusion module is introduced. This dynamic gated residual fusion module balances the information contribution of the original and enhanced features using a learnable gating coefficient.

[0064] 105. Perform underwater target detection on the sonar image based on the target area.

[0065] Here, after determining the target area in the sonar image, the target in the sonar image can be accurately detected by performing target detection on the target area.

[0066] In an embodiment of the present invention, the enhanced processing module of the target detection model is employed to enhance the sonar image to obtain a processed image. This enhances the processed image's anti-interference capabilities, thereby further improving the accuracy of subsequent feature extraction. Subsequently, the processed image is subjected to feature fusion based on a channel attention mechanism to obtain fused features. Adaptive noise suppression and hierarchical dilated convolution are then performed on the fused features to obtain the target region; underwater target detection is then performed on the sonar image using the target region. In this way, feature fusion can enhance the target detection model's sensitivity to processing faint targets and detailed information. Because underwater targets vary greatly in size within an image, a multi-scale feature extraction and fusion strategy is employed to better capture target information at different scales, thereby improving target detection accuracy while maintaining detection speed.

[0067] An underwater target detection method provided by an embodiment of the present invention can be Figure 4 The framework shown implements:

[0068] First, the input image 41 is subjected to sonar preprocessing 42. After that, the features are convolved through two convolutional layers, and then the features are fused through the C2fd module 43 and input into the next convolutional layer for feature convolution. After that, the features are fused through the C2f module and input into the next convolutional layer for feature convolution. The convolution results are processed by the underwater multi-scale contextual attention mechanism (UWA) 44 introduced and pooled. The spliced ​​features are upsampled and the upsampled results are spliced ​​with the output results of the second C2f module. The spliced ​​results are then fused through the C2f module, and upsampled, spliced, and processed by the C2f module again. Finally, the processed results (i.e., the target area) are used for underwater target detection. At the same time, a convolution layer can be used to perform feature convolution on the processing result, and the result of the feature convolution is spliced ​​with the output of the C2f module. After splicing, it is input into the next C2f module for feature fusion, and the result after feature fusion is used for underwater target detection; then, a convolution layer is used again for feature convolution, splicing and C2f module processing, and the result after feature fusion is used for underwater target detection.

[0069] exist Figure 4In the sonar preprocessing 42, adaptive CLAHE enhancement, non-local mean denoising, and frequency domain bandpass filtering are included. The C2fd module 43 includes feature segmentation, an optimized spatial difference module, context gating (outputting gated features), an enhanced ECA module, feature fusion, and output. The data processing process is as follows: feature segmentation is performed on the features output by the convolutional layer, and the segmentation results are input to the optimized spatial difference module and context gating (outputting gated features); the input results of the optimized spatial difference module are input to the enhanced ECA module, and the output of the enhanced ECA module is fused with the gated features to obtain the output.

[0070] UWA 44 includes an input, an adaptive noise suppression module, hierarchical dilated convolution, an attention coordination mechanism, dynamic feature fusion, and an output. The data processing process is as follows: the adaptive noise suppression module performs noise suppression on the input, feeds the output into the hierarchical dilated convolution, and then processes it using the attention coordination mechanism. Finally, the dynamic feature fusion of the attention coordination mechanism and the output of the adaptive noise suppression module is performed to obtain the output.

[0071] In this paper, experiments were conducted on a Linux system, using PyTorch 2.0.0 for training and CUDA 11.8 for acceleration. The processor used was an Intel(R) Xeon(R) Platinum 8362 CPU @ 2.80 GHz, 45 GB of memory, and an RTX 3090 graphics card. The input image size was maintained at 640 × 640 pixels and the batch size was 10 throughout the training process.

[0072] In order to evaluate the effectiveness of the underwater target detection model provided by the embodiment of the present invention, the mean average precision (mAP) of the underwater target detection model (i.e., HAUOD) is calculated. As shown in Table 1, the mean average precision can reach 95.1%, which is 8.3 percentage points higher than the baseline model YOLOv8n (mAP is 86.8%), significantly better than other models. The MAP50-95 index (70.0%) is 4.1 percentage points higher than YOLOv8n (65.9%), indicating that it is more robust under strict IoU thresholds. Both the recall rate (89.7%) and the precision rate (96.3%) are leading, verifying the dual suppression effect of multi-strategy preprocessing and attention mechanism on missed detection and false detection. In order to further verify the effectiveness of HAOOD, the performance of Faster R-CNN and SSD is compared. The experimental results are shown in Table 1.

[0073] Table 1 Model comparison between multiple models and HAUOD

[0074]

[0075] The embodiment of the present invention provides an underwater target detection device, please refer to Figure 5 , which shows a schematic structural diagram of an underwater target detection device provided by one embodiment of the present invention. The device 500 includes:

[0076] Determination module 501, used to determine the initial parameters of the underwater target detection model and the sonar image;

[0077] An enhancement module 502 is configured to perform enhancement processing on the sonar image using the enhancement processing module of the target detection model to obtain a processed image;

[0078] A fusion module 503 is configured to perform feature fusion on the processed image based on a channel attention mechanism to obtain fused features;

[0079] A convolution module 504 is configured to perform adaptive noise suppression and hierarchical dilated convolution on the fused features to obtain a target region;

[0080] The detection module 505 is configured to perform underwater target detection on the sonar image based on the target area.

[0081] In some possible implementations, the fusion module 503 is further configured to perform feature extraction on the processed image based on spatial difference to obtain basic features and differential features;

[0082] The channel attention mechanism is adopted to fuse the basic features and the differential features to obtain the fused features.

[0083] In some possible implementations, the fusion module 503 is further used to adopt a channel attention mechanism to adjust the weights of the basic features and the differential features to obtain adjusted basic features and adjusted differential features; and adopt the adaptive fusion module of the underwater target detection model to perform feature fusion on the adjusted basic features and the adjusted differential features to obtain the fused features.

[0084] In some possible implementations, the fusion module 503 is further used to perform channel compression on the basic features and differential features to obtain compressed basic features and compressed differential features; perform bidirectional gradient extraction and feature expansion on the compressed basic features and compressed differential features to obtain expanded basic features and expanded differential features; and use a channel attention mechanism to adjust the weights of the bidirectional gradient extraction and feature expansion to obtain the adjusted basic features and the adjusted differential features.

[0085] In some possible implementations, the convolution module 504 is further used to suppress high-frequency scattered noise in the fused features to obtain suppressed features; determine the convolution features of the suppressed features; perform attention coordination on the convolution features to obtain coordinated features; and use a learnable gating coefficient to balance the coordinated features to obtain the target area.

[0086] In some possible implementations, the convolution module 504 is further used to construct multi-granularity receptive field parallel branches; and the hierarchical dilated convolution is performed on the suppressed features to obtain convolution features containing multi-granularity context information.

[0087] In some possible implementations, the convolution module 504 is further used to adopt an attention coordination mechanism of the channel dimension and the spatial dimension to perform cross-dimensional feature enhancement on the convolution feature to obtain an enhanced feature; and adopt a cross-dimensional attention mechanism to suppress noise points and noise channels in the enhanced feature to obtain the coordinated feature.

[0088] In some possible implementations, the convolution module 504 is further configured to perform balancing processing on the coordinated features using a learnable gating coefficient to obtain a target feature; and determine a target area corresponding to the target feature in the sonar image.

[0089] In some possible implementations, the enhancement module 502 is further used to enhance the sonar image using the enhancement processing module to generate a visual image; perform non-local mean denoising on the visual image to obtain a denoised image; and perform multi-dimensional denoising on the denoised image to obtain the processed image.

[0090] Optionally, the transmission medium can be a wired link (for example, but not limited to, coaxial cable, optical fiber and digital subscriber line (DSL)) or a wireless link (for example, but not limited to, wireless Fidelity (WIFI), Bluetooth and mobile device network). It should be noted that the device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the method embodiments provided in the above embodiments belong to the same concept. The specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0091] Figure 6 FIG. 1 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. For example, Figure 6As shown, the computer device 600 includes: a memory 601, a processor 602, and a computer program 603 stored in the memory 601 and running on the processor 602, wherein when the processor 602 executes the computer program 603, the computer device can execute any one of the underwater target detection methods introduced above.

[0092] In addition, an embodiment of the present invention also protects a system, which may include a memory and a processor, wherein the memory stores an executable program code, and the processor is used to call and execute the executable program code to perform an underwater target detection method provided by an embodiment of the present invention. This embodiment can divide the system into functional modules according to the above-mentioned method example. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic, which is only a logical function division. There may be other division methods in actual implementation. It should be noted that all relevant contents of each step involved in the above-mentioned method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here.

[0093] It should be understood that the device provided in this embodiment is used to perform the above-mentioned underwater target detection method, and therefore can achieve the same effect as the above-mentioned implementation method. In the case of an integrated unit, the device may include a processing module and a storage module. Specifically, when the device is applied to a device, the processing module can be used to control and manage the actions of the device. The storage module can be used to support the device to execute mutual program codes, etc. Specifically, the processing module can be a processor or a controller, which can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the contents disclosed in the present invention. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module can be a memory.

[0094] In addition, the device provided by the embodiments of the present invention may specifically be a chip, component, or module. The chip may include a connected processor and memory; the memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute the underwater target detection method provided by the above embodiment. This embodiment also provides a computer-readable storage medium, which stores computer program code. When the computer program code is executed on a computer, it causes the computer to execute the above-mentioned method steps to implement the underwater target detection method provided by the above embodiment.

[0095] This embodiment also provides a computer program product. When the computer program product is executed on a computer, it causes the computer to execute the above-mentioned steps to implement the underwater target detection method provided in the above-mentioned embodiment. The device, computer-readable storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects achieved by the device, computer-readable storage medium, computer program product, or chip provided in this embodiment can refer to the beneficial effects of the corresponding method provided above and will not be repeated here. Through the description of the above embodiments, those skilled in the art will understand that for the sake of convenience and brevity, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In the embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, other division methods can be used. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not implemented. On the other hand, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, which may be electrical, mechanical or other forms.

[0096] It should be noted that the above-mentioned order of the embodiments of the present invention is for description only and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous. The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. The above content is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered within the scope of protection of the present invention.

Claims

1. A method for underwater target detection, characterized in that: The underwater target detection method comprises: Determine the initial parameters of the underwater target detection model and sonar image; Using the enhancement processing module of the target detection model, the sonar image is enhanced to obtain a processed image; Performing feature fusion on the processed image based on a channel attention mechanism to obtain fused features; Performing adaptive noise suppression and hierarchical dilated convolution on the fused features to obtain the target area; Based on the target area, underwater target detection is performed on the sonar image.

2. The underwater target detection method according to claim 1, characterized in that: The performing feature fusion on the processed image based on the channel attention mechanism to obtain fused features includes: Extracting features from the processed image based on spatial differences to obtain basic features and differential features; The channel attention mechanism is adopted to fuse the basic features and the differential features to obtain the fused features.

3. The underwater target detection method according to claim 2, characterized in that: The channel attention mechanism is used to fuse the basic features and differential features to obtain the fused features, including: Adopting a channel attention mechanism to adjust the weights of the basic features and the differential features to obtain adjusted basic features and adjusted differential features; The adaptive fusion module of the underwater target detection model is used to perform feature fusion on the adjusted basic features and the adjusted differential features to obtain the fused features.

4. The underwater target detection method according to claim 3, characterized in that: The channel attention mechanism is used to adjust the weights of the basic features and the differential features to obtain adjusted basic features and adjusted differential features, including: Performing channel compression on the basic features and the differential features to obtain compressed basic features and compressed differential features; Performing bidirectional gradient extraction and feature expansion on the compressed basic features and the compressed differential features to obtain expanded basic features and expanded differential features; A channel attention mechanism is adopted to adjust the weights of the bidirectional gradient extraction and feature expansion to obtain the adjusted basic features and the adjusted differential features.

5. The underwater target detection method according to claim 1, characterized in that: The adaptive noise suppression and hierarchical dilated convolution are performed on the fused features to obtain the target area, including: Suppressing high-frequency scattered noise in the fused features to obtain suppressed features; determining convolution features of the suppressed features; Performing attention coordination on the convolutional features to obtain coordinated features; The coordinated features are balanced using a learnable gating coefficient to obtain the target region.

6. The underwater target detection method according to claim 5, characterized in that: Determining the convolution feature of the suppressed feature includes: Construct multi-granularity receptive field parallel branches; The suppressed features are subjected to hierarchical dilated convolution to obtain convolutional features containing multi-granularity context information.

7. The underwater target detection method according to claim 5, characterized in that: The performing attention coordination on the convolutional features to obtain coordinated features includes: Adopting the attention coordination mechanism of channel dimension and spatial dimension to perform cross-dimensional feature enhancement on the convolutional features to obtain enhanced features; A cross-dimensional attention mechanism is used to suppress noise points and noise channels in the enhanced features to obtain the coordinated features.

8. The underwater target detection method according to claim 5, characterized in that: The balancing process of the coordinated features using the learnable gating coefficient to obtain the target region includes: Using a learnable gating coefficient to balance the coordinated features to obtain a target feature; In the sonar image, a target area corresponding to the target feature is determined.

9. The underwater target detection method according to claim 1, characterized in that: The enhancement processing module of the target detection model is used to perform enhancement processing on the sonar image to obtain a processed image, including: Using the enhancement processing module to enhance the sonar image to generate a visual image; Performing non-local mean denoising on the visualized image to obtain a denoised image; Multi-dimensional noise reduction is performed on the denoised image to obtain the processed image.

10. An underwater target detection device, characterized in that: The device comprises: A determination module, used to determine the initial parameters of the underwater target detection model and the sonar image; an enhancement module, configured to perform enhancement processing on the sonar image using the enhancement processing module of the target detection model to obtain a processed image; A fusion module, configured to perform feature fusion on the processed image based on a channel attention mechanism to obtain fused features; A convolution module is used to perform adaptive noise suppression and hierarchical dilated convolution on the fused features to obtain a target area; A detection module is used to perform underwater target detection on the sonar image based on the target area.