Forward-looking sonar image obstacle detection method, device and equipment and storage medium
Through the improved YOLOv8n-seg model with multi-domain adaptive threshold denoising and polar coordinate feature fusion network, the problem of balancing forward-looking sonar detection accuracy and efficiency is solved, and efficient and high-precision obstacle detection is achieved, which is suitable for underwater robots and marine exploration equipment.
Patent Information
- Application Number
- CN202511197113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing forward-looking sonar obstacle detection solutions are unable to strike a balance between detection accuracy and efficiency. Limited by the physical characteristics of sonar imaging and the bottleneck of embedded computing power, they suffer from problems such as strong speckle noise, obvious sector distortion, and limited embedded computing power.
An improved YOLOv8n-seg model with multi-domain adaptive threshold denoising and polar coordinate feature fusion network is adopted to reduce noise through multi-domain adaptive threshold denoising and correct sector distortion through polar coordinate feature fusion network, thus achieving efficient and high-precision obstacle detection.
It achieves efficient and high-precision obstacle detection in forward-looking sonar images, reduces speckle noise, enhances the ability to perceive the boundaries of arc-shaped and fan-shaped obstacles, and maintains the stability of real-time detection and low operation and maintenance costs.
Smart Images

Figure CN120726464A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, equipment and storage medium for detecting obstacles in forward-looking sonar images. Background Art
[0002] Forward-looking sonar uses acoustic imaging technology to capture real-time images of the underwater environment and is widely used in underwater robots, autonomous submersibles, and ocean exploration equipment. With the widespread adoption of underwater robots in complex tasks such as marine exploration and pipeline inspection, real-time obstacle avoidance based on forward-looking sonar has become a core requirement for ensuring navigation safety. However, due to the complexity of the underwater environment, sonar images often suffer from high noise, low contrast, and blurred details, resulting in significant challenges in accurate obstacle detection.
[0003] Among related technologies, existing forward-looking sonar (FLS) obstacle detection solutions are limited by the physical characteristics of sonar imaging and the bottleneck of embedded computing power. They suffer from problems such as strong speckle noise, obvious sector distortion, and limited embedded computing power, making it difficult to simultaneously balance detection accuracy and efficiency. Summary of the Invention
[0004] The embodiments of this specification provide a forward-looking sonar image obstacle detection method to solve the problem of difficulty in simultaneously balancing detection accuracy and efficiency in existing forward-looking sonar obstacle detection solutions.
[0005] To solve the above technical problems, the embodiments of this specification are implemented as follows: In a first aspect, an embodiment of this specification provides a forward-looking sonar image obstacle detection method, comprising: Acquire the forward-looking sonar image to be detected; Performing noise reduction processing on the forward-looking sonar image to be detected by multi-domain adaptive threshold noise reduction; Inputting the denoised forward-looking sonar image to be detected into an improved YOLOv8n-seg model to obtain a target forward-looking sonar image containing an obstacle mask output by the improved YOLOv8n-seg model; Among them, the improved YOLOv8n-seg model replaces the neck network in the original YOLOv8n-seg model with a polar coordinate feature fusion network, and the polar coordinate feature fusion network is used to correct fan-shaped distortion and output multi-scale polar coordinate features.
[0006] In a second aspect, an embodiment of this specification provides a forward-looking sonar image obstacle detection device, comprising: An acquisition module is used to acquire the forward-looking sonar image to be detected; A processing module, configured to perform noise reduction processing on the forward-looking sonar image to be detected by multi-domain adaptive threshold noise reduction; a determination module, configured to input the forward-looking sonar image to be detected after noise reduction processing into an improved YOLOv8n-seg model, and obtain a target forward-looking sonar image containing an obstacle mask output by the improved YOLOv8n-seg model; Among them, the improved YOLOv8n-seg model replaces the neck network in the original YOLOv8n-seg model with a polar coordinate feature fusion network, and the polar coordinate feature fusion network is used to correct fan-shaped distortion and output multi-scale polar coordinate features.
[0007] In a third aspect, an embodiment of this specification provides a forward-looking sonar image obstacle detection device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the forward-looking sonar image obstacle detection method in solution one.
[0008] In a fourth aspect, an embodiment of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the forward-looking sonar image obstacle detection method in solution one.
[0009] One embodiment of the present specification achieves the following beneficial effects: the forward-looking sonar image processed by multi-domain adaptive threshold denoising is input, the neck network is replaced with an improved YOLOv8n-seg model with a polar coordinate feature fusion network, and efficient and high-precision obstacle detection in the forward-looking sonar image is achieved through the coordinated optimization of multi-domain adaptive threshold denoising and the polar coordinate feature fusion network. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0011] Figure 1 A flow chart of a forward-looking sonar image obstacle detection method provided in an embodiment of this specification; Figure 2 A schematic diagram of the noise reduction process provided in the embodiments of this specification; Figure 3 A schematic diagram of the framework of the polar coordinate feature fusion network provided in the embodiments of this specification; Figure 4A logical diagram of mask scheduling provided in an embodiment of this specification; Figure 5 A schematic diagram of the structure of a forward-looking sonar image obstacle detection device provided in an embodiment of this specification; Figure 6 This is a schematic diagram of the structure of a forward-looking sonar image obstacle detection device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0012] To make the purpose, technical solutions, and advantages of one or more embodiments of this specification more clear, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of one or more embodiments of this specification.
[0013] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0014] A forward-looking sonar image obstacle detection method provided in the embodiments of the specification is described in detail with reference to the accompanying drawings.
[0015] Figure 1 A flow chart of a forward-looking sonar image obstacle detection method provided in an embodiment of this specification.
[0016] like Figure 1 As shown, the process may include the following steps: Step 110: Acquire the forward-looking sonar image to be detected.
[0017] In the embodiments of this specification, a forward-looking sonar image to be detected is obtained by a sonar acquisition board, such as an 8-bit grayscale forward-looking sonar image, which can reflect the echo intensity distribution of objects in an underwater environment. The forward-looking sonar image to be detected contains noise, which affects the image quality.
[0018] Step 120: De-noising the forward-looking sonar image to be detected by multi-domain adaptive threshold denoising.
[0019] In the embodiment of this specification, after the image is processed by multi-domain adaptive threshold noise reduction, the speckle noise is significantly reduced, and the weak echo edges are completely preserved, providing high-quality input for subsequent feature extraction and segmentation.
[0020] Step 130: Inputting the denoised forward-looking sonar image to be detected into an improved YOLOv8n-seg model to obtain a target forward-looking sonar image containing an obstacle mask output by the improved YOLOv8n-seg model; Among them, the improved YOLOv8n-seg model replaces the neck network in the original YOLOv8n-seg model with a polar coordinate feature fusion network, and the polar coordinate feature fusion network is used to correct fan-shaped distortion and output multi-scale polar coordinate features.
[0021] In the embodiments of this specification, the neck network (Neck) in the original YOLOv8n-seg model is replaced with a polar coordinate feature fusion network. The polar coordinate feature fusion network may include a polar coordinate remapping module and a multi-scale polar coordinate feature fusion module, which are used to correct fan-shaped distortion and output multi-scale polar coordinate features, thereby enhancing the network's perception of the boundaries of arc-shaped and fan-shaped obstacles.
[0022] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some steps can be omitted or deleted.
[0023] In the embodiments of this specification, the forward-looking sonar image after multi-domain adaptive threshold denoising is input, and the neck network is replaced with an improved YOLOv8n-seg model with a polar coordinate feature fusion network. Through the coordinated optimization of multi-domain adaptive threshold denoising and the polar coordinate feature fusion network, efficient and high-precision detection of obstacles in the forward-looking sonar image is achieved.
[0024] based on Figure 1 The method in this specification also provides some specific implementation plans of the method, which are described below.
[0025] Optionally, the denoising process of the forward-looking sonar image to be detected by multi-domain adaptive threshold denoising in the embodiments of this specification may specifically include: The local noise variance is calculated by pixel sliding window, and the directional gradient is calculated by convolution to generate a pixel-by-pixel threshold map; The pixel-by-pixel threshold map is subjected to two-layer discrete wavelet decomposition, and non-local block matching filtering is performed according to a three-dimensional block matching algorithm to obtain the forward-looking sonar image to be detected after noise reduction.
[0026] In the embodiments of this specification, the mean and variance of the local area are calculated through a pixel sliding window, and the variance is used to measure the noise intensity. The directional gradient is calculated by combining convolution, and the noise variance and gradient information are fused to generate a pixel-by-pixel threshold map.
[0027] A two-layer DWT decomposition is performed on the pixel-by-pixel threshold map, decomposing the image into a low-frequency approximation component and a high-frequency detail component. The high-frequency component contains noise and edge information and requires filtering. The high-frequency components after wavelet decomposition are divided into image blocks. Similar blocks are searched in three-dimensional space and aggregated into block stacks. Collaborative filtering is used to suppress noise, resulting in a denoised forward-looking sonar image to be detected.
[0028] Figure 2 This is a flowchart of the noise reduction process provided in the embodiments of this specification.
[0029] like Figure 2 As shown in the figure, the noise reduction process of the acquired forward-looking sonar image is performed as follows: Step 201: Acquire a forward-looking sonar image; Step 202: Local variance. Use a 32×32 pixel sliding window to slide on the image and calculate the variance σ of the pixel values in each sliding window to determine the local noise intensity. The larger the variance, the stronger the noise in the area.
[0030] Step 203: Directional gradient, using a 3×3 Sobel convolution kernel to calculate the directional gradient of the image , to identify edge information in the image. Regions with high gradient values usually correspond to the edges of objects and need to be specially protected to avoid being weakened by the noise reduction process.
[0031] The calculation of local variance and gradient is parallel and independent, and the variance and directional gradient are calculated at the same time.
[0032] Step 204: Adaptive thresholding: Generate a pixel-by-pixel adaptive threshold map based on the local variance and directional gradient. This threshold map can dynamically adjust the noise reduction strength according to the noise and edge characteristics of different image regions.
[0033] Step 205: Wavelet soft threshold denoising, perform two-layer discrete wavelet decomposition on the image, decompose the image into low-frequency sub-bands and high-frequency sub-bands, the low-frequency sub-bands retain the main structure of the image, the high-frequency sub-bands contain noise and edge information, and only the high-frequency sub-band coefficients are subjected to adaptive threshold soft threshold processing. , then set it to zero (to remove noise); if , then it shrinks to (Preserves edges and reduces ringing.) Soft thresholding provides a smoother transition than hard thresholding and is less likely to produce artifacts during reconstruction. Soft thresholding can effectively preserve image edge information while removing noise, avoiding artifacts that may be introduced by hard thresholding.
[0034] Step 206: Adaptive BM3D denoising, using The global noise level is estimated using the mean of σ, which drives the Block Matching 3D (BM3D) algorithm to perform non-local block matching filtering. Similar 8×8 patches are searched for in the image and stacked into a 3D array. A 3D transformation is performed on the 3D array, and hard thresholding is applied to remove noise coefficients. The filtered 3D array is inversely transformed and further optimized using a Wiener filter. The optimized patches are weighted averaged back to their original image positions to produce the final denoised result. BM3D combines 3D transformation of similar image patches with hard thresholding to further smooth residual noise while preserving image details and edges.
[0035] Step 207: Output the noise-reduced forward-looking sonar image.
[0036] Specifically, the local noise variance can represent the degree of local texture change. The larger the local noise variance, the stronger the noise in the area. The calculation formula of the local noise variance is:
[0037] in, Pixel The local noise variance at , is the variance operation, Pixels Central Image pixel values within the local window.
[0038] The formula for calculating the directional gradient is:
[0039] in, Pixel The gradient amplitude at , which measures the edge strength; is the image gradient vector, which indicates the direction and magnitude of the intensity change of the image at that point; is the image gradient in the horizontal direction, is the image gradient in the vertical direction, and The calculation formula is:
[0040] in, is the original image matrix, It is a two-dimensional convolution operation.
[0041] The calculation formula of the adaptive threshold is:
[0042] in, Pixel The adaptive threshold at To control the weight coefficient of noise suppression, The weight coefficient reserved for the control edge, Pixel The standard deviation of the local noise at Pixel The directional gradient amplitude at . In practice, .
[0043] The wavelet soft threshold denoising formula is:
[0044] in, is the coefficient after soft threshold processing; is the high frequency coefficient of the original wavelet; is the adaptive threshold of the position corresponding to the current coefficient; for The sign function of , positive is +1, negative is -1, and zero is 0; For the maximum value operation (if , then set to zero).
[0045] The wavelet reconstruction formula is:
[0046] in, is the image after wavelet soft threshold denoising (reconstruction result), IDWT is the two-dimensional inverse wavelet transform, is the low-frequency subband of wavelet transform, For the The wavelet high frequency subband of the layer, is the number of wavelet decomposition levels.
[0047] Global noise estimation can be used for subsequent BM3D filter strength adjustment. The global noise estimation formula is:
[0048] in, is the global noise estimation of the image; is the average value of all local noise in the whole image; is the scale factor, which can be set to 1.0 to adjust the estimated intensity; 255 is the maximum pixel value of the grayscale image, scaling the pixel range to [0,1].
[0049] Adaptive BM3D denoising, the BM3D (Block-Matching and 3D Filtering) algorithm includes the following steps: Block Matching: Search the entire image for image blocks similar to the reference block (such as) and stack them into a 3D data volume.
[0050] Collaborative HT (Hard Thresholding): Performs a 3D transform (such as DCT) on the 3D data volume, and then performs hard thresholding to remove noise.
[0051] Wiener re-estimation: The second stage uses Wiener filtering to make a finer estimate of the image block and improve the edge preservation ability.
[0052] Aggregation: Multiple estimated image blocks are superimposed into a complete image in a weighted average manner to reduce blocking effects.
[0053] Output The final output is a forward-looking sonar image to be detected with high signal-to-noise and clear edges. The calculation formula for the forward-looking sonar image to be detected is:
[0054] in, The final output is a forward-looking sonar image with high signal-to-noise ratio and clear edges to be detected; Used to limit the pixel value range to [0,255] to prevent BM3D output values from exceeding the image range; BM3D Express Using the BM3D algorithm, based on global noise estimation Perform denoising.
[0055] Wavelet soft-threshold denoising and adaptive BM3D denoising form a two-stage collaborative mechanism: coarse-first, fine-second. The former uses frequency-domain decomposition and local adaptive thresholding to suppress high-frequency noise, providing a smoother input image for the subsequent adaptive BM3D denoising refinement. Furthermore, the local standard deviation mean calculated in the wavelet stage serves as the global noise estimation input for the BM3D stage, enabling parameter sharing and data coupling, enhancing the robustness and structure-preserving capabilities of the overall denoising framework.
[0056] Optionally, in the embodiment of this specification, the method of inputting the denoised forward-looking sonar image to be detected into the improved YOLOv8n-seg model to obtain the target forward-looking sonar image containing the obstacle mask output by the improved YOLOv8n-seg model may specifically include: Inputting the denoised forward-looking sonar image to be detected into the backbone network of the improved YOLOv8n-seg model, performing local feature extraction on the forward-looking sonar image to be detected through a convolution operation, and generating a feature map; The polar coordinate feature fusion network remaps the feature map to the polar coordinate network, and performs multi-scale feature fusion in the polar coordinate domain to generate a multi-scale feature map; The head network receives the multi-scale feature maps and generates a target forward-looking sonar image containing an obstacle mask through linear combination.
[0057] In practice, the main components of the original YOLOv8n-seg model include the backbone network (Backbone), neck (Neck) and head (Head). A polar coordinate domain feature fusion network is constructed to replace the neck structure of the original YOLOv8n-seg model.
[0058] Figure 3 A schematic diagram of the framework of the polar coordinate feature fusion network provided in the embodiments of this specification.
[0059] like Figure 3 As shown, a forward-looking sonar image (high signal-to-noise ratio, weak edge preservation) that has undergone multi-domain adaptive threshold denoising is fed into the backbone network of the improved YOLOv8n-seg model. The backbone network extracts local features from the input image through a series of convolutional layers (e.g., the CSPDarknet architecture) to generate feature maps (P3, P4, and P5) at different scales. P3, at 1 / 8 resolution, contains high-frequency details (e.g., edges of small objects). P4, at 1 / 16 resolution, balances semantics with details. P5, at 1 / 32 resolution, contains global semantic information (e.g., obstacle categories). The feature maps extracted by the backbone network provide the basis for subsequent polar coordinate remapping and multi-scale fusion, ensuring that obstacle features at different scales are effectively captured.
[0060] The polar coordinate feature fusion network (Polar-PAN) replaces the neck network (Neck) of the original YOLOv8n-seg model. The polar coordinate feature fusion network includes a polar coordinate remapping module and a multi-scale polar coordinate feature fusion module.
[0061] Polar coordinate remapping module: Performs polar coordinate remapping (ToPolar) on the P5, P4, and P3 features output by Backbone. ToPolar can correct the fan-shaped geometric distortion of sonar images and convert feature maps in the rectangular coordinate system into the polar coordinate system.
[0062] Specifically, the image vertex is used as the extreme point, and the radius r is divided according to the logarithmic spacing (adapting to the sonar distance attenuation characteristics), and the equiangular division (uniform coverage in azimuth), build The size of the sampling grid.
[0063] Use the grid_sample function to project the P3, P4, and P5 feature maps from the rectangular coordinate system to the (r,θ) polar coordinate grid to generate polar coordinate domain feature maps (Polar-P3, Polar-P4, Polar-P5).
[0064] Multi-scale polar coordinate feature fusion module: bidirectional fusion is performed in the polar coordinate domain. In the top-down path (FPN), the high-level features (Polar-P5) are upsampled by 2 times and then spliced with the middle-level features (Polar-P4), and fused through the C2f convolution block (512 channels); then upsampled and spliced with the low-level features (Polar-P3) to generate the C2f (256) feature map.
[0065] In the bottom-up path (PAN), the low-level detail features (C2f(256)) are downsampled by 2 times and then fused with the middle-level features (Polar-P4) to generate C2f(512); they are further downsampled and fused with the high-level features (Polar-P5) to generate C2f(1024).
[0066] The final output is three multi-scale polar coordinate feature maps (Polar-P3 (1 / 8), Polar-P4 (1 / 16), and Polar-P5 (1 / 32)). These maps maintain the geometric consistency of curved obstacles within the polar coordinate domain. Pixel alignment at curved obstacles preserves both local texture and global semantics. This approach only adds three grid_samples and two C2f modules, increasing FLOPs by approximately 5% while improving the mean Intersection Over Union (MIoU) of curved objects by 2 to 3 percentage points. Furthermore, the remapping operation supports backpropagation, and pre-trained weights can be strictly loaded (strict=False), eliminating the need for additional training convergence.
[0067] The head network receives the multi-scale polar coordinate feature maps (Polar-P3, Polar-P4, and Polar-P5) output by Polar-PAN. It linearly combines these multi-scale feature maps through 1×1 convolution to generate instance masks, identifying the spatial distribution of obstacles. This generates a forward-looking sonar image of the target, including the obstacle mask. This accurately depicts the contours of irregular obstacles, which can subsequently guide the AUV in obstacle avoidance.
[0068] In practice, before inputting the forward-looking sonar image to be detected into the improved YOLOv8n-seg model, in order to improve the accuracy of the model prediction, the improved YOLOv8n-seg model can be trained in advance.
[0069] Optionally, the improved YOLOv8n-seg model described in the embodiments of this specification is run on the platform, and the method may include: The GPU utilization of the platform is obtained, and the number of mask prototypes and the input resolution of the improved YOLOv8n-seg model are adjusted according to the hysteresis threshold.
[0070] In practice, the improved YOLOv8n-seg model can run on the Jetson Orin NX 15W platform. Different platform hardware conditions result in different frame rates when running. Mask prototype dynamic scheduling can detect the platform's GPU utilization and ensure that the frame rate can be between 25 and 28 fps.
[0071] Figure 4 A logical diagram of mask scheduling provided for an embodiment of this specification.
[0072] like Figure 4 As shown, an independent thread is set to read the GPU utilization U every 0.5 seconds, and rising and falling thresholds are set to avoid performance fluctuations caused by frequent adjustments.
[0073] The core utilization U is read through the NVML interface. If the current state is K32 and U exceeds 70%, the number of mask prototypes is halved to K16. If it is still above the threshold, it is further reduced to K8, and the network input resolution is simultaneously reduced from 640px to 480px, thereby directly reducing the number of pixels in the feature extraction layer by 25%. When the utilization falls below the 50% hysteresis threshold, the network is restored in the reverse order from K8 to K16 to K32, and the resolution is reset to 640px when the K16 level is about to rise again. Resolution scaling performs a bilinear interpolation on the input frame.
[0074] Since the prototype mask of YOLOv8-Seg already captures 95% of semantic information in the first 16 channels, reducing K from 32 to 8 results in only a 0.4 percentage point loss in mIoU, but reduces the number of mask multiplications and additions and video memory usage to a quarter. Combined with resolution downscaling, total FLOPs decreases by approximately 60%. Running on a 15W Jetson Orin NX platform, GPU_util can oscillate from 35% to 85% under task parallelization. The scheduling mechanism ensures a consistent frame rate of 25–28 fps, while the unscheduled baseline model drops to 15 fps during high-load periods and experiences missed detections. Furthermore, a hysteresis design of U_hi / U_lo = 70% / 50% is used to prevent jitter around 60%.
[0075] Through multi-domain adaptive threshold noise reduction, polar coordinate feature fusion network and dynamic mask scheduling, the system achieves high precision, real-time performance and low operation and maintenance costs, which are difficult to achieve with similar sonar obstacle avoidance systems: multi-domain adaptive threshold noise reduction reduces speckle noise by an average of 5dB without relying on any prior parameters and fully preserves the edges of weak echoes; multi-scale feature fusion effectively enhances the network's perception of the boundaries of arc-shaped and fan-shaped obstacles; dynamic clipping of the number of mask prototypes and closed-loop scheduling of resolution ensures a stable frame rate of 25-28FPS for the detection platform with an accuracy loss of less than 0.4 percentage points, providing highly reliable, low-power real-time obstacle avoidance capabilities for AUV inspection, port security and emergency search and rescue.
[0076] The present invention achieves a real-time three-way balance among noise, computing power, and accuracy, providing AUVs with stable and reliable forward-looking sonar segmentation and obstacle avoidance capabilities in actual scenarios with long ranges and variable loads.
[0077] Figure 5 This is a schematic diagram of the structure of a forward-looking sonar image obstacle detection device provided in an embodiment of this specification.
[0078] Corresponding to the method embodiment, this embodiment further provides a forward-looking sonar image obstacle detection device, which may include: An acquisition module 502 is used to acquire a forward-looking sonar image to be detected; A processing module 504 is configured to perform noise reduction processing on the forward-looking sonar image to be detected by multi-domain adaptive threshold noise reduction; a determination module 506 configured to input the denoised forward-looking sonar image to be detected into an improved YOLOv8n-seg model to obtain a target forward-looking sonar image containing an obstacle mask output by the improved YOLOv8n-seg model; Among them, the improved YOLOv8n-seg model replaces the neck network in the original YOLOv8n-seg model with a polar coordinate feature fusion network, and the polar coordinate feature fusion network is used to correct fan-shaped distortion and output multi-scale polar coordinate features.
[0079] Optionally, the denoising process of the forward-looking sonar image to be detected by multi-domain adaptive threshold denoising in the embodiments of this specification may specifically include: The local noise variance is calculated by pixel sliding window, and the directional gradient is calculated by convolution to generate a pixel-by-pixel threshold map; The pixel-by-pixel threshold map is subjected to two-layer discrete wavelet decomposition, and non-local block matching filtering is performed according to a three-dimensional block matching algorithm to obtain the forward-looking sonar image to be detected after noise reduction.
[0080] Optionally, in the embodiment of this specification, the method of inputting the denoised forward-looking sonar image to be detected into the improved YOLOv8n-seg model to obtain the target forward-looking sonar image containing the obstacle mask output by the improved YOLOv8n-seg model may specifically include: Inputting the denoised forward-looking sonar image to be detected into the backbone network of the improved YOLOv8n-seg model, performing local feature extraction on the forward-looking sonar image to be detected through a convolution operation, and generating a feature map; The polar coordinate feature fusion network remaps the feature map to the polar coordinate network, and performs multi-scale feature fusion in the polar coordinate domain to generate a multi-scale feature map; The head network receives the multi-scale feature maps and generates a target forward-looking sonar image containing an obstacle mask through linear combination.
[0081] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.
[0082] Figure 6 This is a schematic diagram of the structure of a forward-looking sonar image obstacle detection device provided in an embodiment of this specification. Figure 6 As shown, an embodiment of this specification provides a forward-looking sonar image obstacle detection device 600, including a memory 630, a processor 610, and a computer program 620 stored in the memory. The processor 610 executes the computer program 620 to implement the forward-looking sonar image obstacle detection method described in any of the above embodiments.
[0083] An embodiment of this specification provides a forward-looking sonar image obstacle detection device, which may include a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the forward-looking sonar image obstacle detection method described in any of the above embodiments.
[0084] An embodiment of this specification provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the forward-looking sonar image obstacle detection method described in any of the above embodiments can be implemented.
[0085] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. Figure 6 As for the device shown, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0086] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0087] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0088] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0089] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0090] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0091] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0092] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0094] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0095] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0096] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0097] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0098] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0099] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0100] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for detecting obstacles in forward-looking sonar images, characterized in that: include: Acquire the forward-looking sonar image to be detected; Performing noise reduction processing on the forward-looking sonar image to be detected by multi-domain adaptive threshold noise reduction; Inputting the denoised forward-looking sonar image to be detected into an improved YOLOv8n-seg model to obtain a target forward-looking sonar image containing an obstacle mask output by the improved YOLOv8n-seg model; Among them, the improved YOLOv8n-seg model replaces the neck network in the original YOLOv8n-seg model with a polar coordinate feature fusion network, and the polar coordinate feature fusion network is used to correct fan-shaped distortion and output multi-scale polar coordinate features.
2. The method according to claim 1, characterized in that The denoising process of the forward-looking sonar image to be detected by multi-domain adaptive threshold denoising specifically includes: The local noise variance is calculated by pixel sliding window, and the directional gradient is calculated by convolution to generate a pixel-by-pixel threshold map; The pixel-by-pixel threshold map is subjected to two-layer discrete wavelet decomposition, and non-local block matching filtering is performed according to a three-dimensional block matching algorithm to obtain the forward-looking sonar image to be detected after noise reduction.
3. The method according to claim 1, characterized in that Inputting the noise-reduced forward-looking sonar image to be detected into the improved YOLOv8n-seg model to obtain the target forward-looking sonar image containing the obstacle mask output by the improved YOLOv8n-seg model specifically includes: Inputting the denoised forward-looking sonar image to be detected into the backbone network of the improved YOLOv8n-seg model, performing local feature extraction on the forward-looking sonar image to be detected through a convolution operation, and generating a feature map; The polar coordinate feature fusion network remaps the feature map to the polar coordinate network, and performs multi-scale feature fusion in the polar coordinate domain to generate a multi-scale feature map; The head network receives the multi-scale feature maps and generates a target forward-looking sonar image containing an obstacle mask through linear combination.
4. The method according to claim 1, wherein The improved YOLOv8n-seg model is run on the platform, and the method includes: The GPU utilization of the platform is obtained, and the number of mask prototypes and the input resolution of the improved YOLOv8n-seg model are adjusted according to the hysteresis threshold.
5. The method according to claim 2, characterized in that The calculation formula of the adaptive threshold of the pixel-by-pixel threshold map is: in, At the pixel point The adaptive threshold, To control the weight coefficient of noise suppression, The weight coefficient reserved for the control edge, Pixel The standard deviation of the local noise at Pixel The magnitude of the directional gradient at .
6. A forward-looking sonar image obstacle detection device, characterized in that: include: An acquisition module is used to acquire the forward-looking sonar image to be detected; A processing module, configured to perform noise reduction processing on the forward-looking sonar image to be detected by multi-domain adaptive threshold noise reduction; a determination module, configured to input the forward-looking sonar image to be detected after noise reduction processing into an improved YOLOv8n-seg model, and obtain a target forward-looking sonar image containing an obstacle mask output by the improved YOLOv8n-seg model; Among them, the improved YOLOv8n-seg model replaces the neck network in the original YOLOv8n-seg model with a polar coordinate feature fusion network, and the polar coordinate feature fusion network is used to correct fan-shaped distortion and output multi-scale polar coordinate features.
7. The device according to claim 6, characterized in that The denoising process of the forward-looking sonar image to be detected by multi-domain adaptive threshold denoising specifically includes: The local noise variance is calculated by pixel sliding window, and the directional gradient is calculated by convolution to generate a pixel-by-pixel threshold map; The pixel-by-pixel threshold map is subjected to two-layer discrete wavelet decomposition, and non-local block matching filtering is performed according to a three-dimensional block matching algorithm to obtain the forward-looking sonar image to be detected after noise reduction.
8. The device according to claim 6, characterized in that Inputting the noise-reduced forward-looking sonar image to be detected into the improved YOLOv8n-seg model to obtain the target forward-looking sonar image containing the obstacle mask output by the improved YOLOv8n-seg model specifically includes: Inputting the denoised forward-looking sonar image to be detected into the backbone network of the improved YOLOv8n-seg model, performing local feature extraction on the forward-looking sonar image to be detected through a convolution operation, and generating a feature map; The polar coordinate feature fusion network remaps the feature map to the polar coordinate network, and performs multi-scale feature fusion in the polar coordinate domain to generate a multi-scale feature map; The head network receives the multi-scale feature maps and generates a target forward-looking sonar image containing an obstacle mask through linear combination.
9. A forward-looking sonar image obstacle detection device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Fan blade surface damage detection method and system
CN119887790A
Embroidery stitch identification method based on deep learning
CN120014368A
Sonar image target detection method based on deep learning, electronic equipment and storage medium
CN120510359A
Deep convolutional neural network-based submerged oil sonar detection image recognition method
WO2021243743A1