Image sonar small target detection method and system based on improved attention mechanism

By inserting up-down sampling and convolutional operations of improved attention mechanism into YOLOv8's backbone network, the problems of low detection accuracy and high calculation overhead in image sonar small object detection are solved, and a more efficient small object detection effect is achieved.

CN120107768AActive Publication Date: 2025-06-06INST OF ACOUSTICS CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510145354.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-06
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The prior art has problems such as low detection accuracy, large calculation overhead and poor generalization in image sonar small object detection.

Method used

A small object detection method for image sonar based on improved attention mechanism is proposed. By inserting up-down sampling and convolution operations on layer 10 of YOLOv8's backbone network, combining point multiplication and splicing operations, the model's detection ability of small objects is optimized.

Benefits of technology

Through up-down sampling and convolution operations, the model can learn image features in different dimensions more effectively, reduce background interference, reduce calculation overhead, and improve the accuracy and efficiency of small-object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107768A_ABST
    Figure CN120107768A_ABST
Patent Text Reader

Abstract

The invention provides an image sonar small target detection method and system based on an improved attention mechanism, and the method is realized based on YOLOv8, and the method comprises the steps: inserting the following process in the tenth layer of a backbone network of YOLOv8: carrying out the up-sampling of an original image, sequentially carrying out the 5 * 5 point convolution, vertical convolution and horizontal convolution, and finally obtaining data F'through an activation function; performing down-sampling on an original image, then performing 1 * 1 point convolution, horizontal convolution and vertical convolution in sequence, and finally obtaining data F ''through an activation function; and carrying out dot product on the data F ', the data F' 'and the original image data, and splicing with the original image data to obtain a final output result. Compared with the prior art, the method has the advantages that the model learns image features of different dimensions through up-down sampling, convergence is accelerated through aggregation in the horizontal direction and the vertical direction and finally residual connection, a more refined result can be output, background interference is reduced, and calculation overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image sonar small target detection, and specifically relates to an image sonar small target detection method and system based on an improved attention mechanism. Background Art

[0002] Image sonar has important applications in underwater search and rescue and mine scanning, but the accuracy of small target detection is affected by factors such as the complex water environment. Deep learning methods can improve detection results, but they increase computational overhead and lack generalization research. Although the application of deep learning methods in image sonar small target detection has made some progress, there are still some defects and challenges. The following are the main aspects:

[0003] 1. Inherent challenges of small object detection:

[0004] Small objects occupy fewer pixels in an image, have limited feature information, and are easily submerged by background noise. Existing deep learning models (such as Faster R-CNN, YOLO, etc.) perform poorly when processing small objects and have difficulty capturing effective features.

[0005] 2. Model complexity and computational cost:

[0006] Deep learning models usually require a lot of computing resources, and sonar image processing has high real-time requirements. Complex models may lead to high computing costs and slow reasoning speeds, making it difficult to meet the needs of practical applications.

[0007] 3. Lack of dedicated models for sonar data:

[0008] Most of the existing deep learning models are designed for optical images, but the imaging principles and characteristics of sonar images are quite different from those of optical images. Directly applying existing models may lead to poor performance, and a dedicated model needs to be designed based on the characteristics of sonar images. Summary of the invention

[0009] The purpose of this application is to overcome the defects of the prior art, such as low detection accuracy, high computational overhead and poor generalization.

[0010] In order to achieve the above objectives, this application proposes an image sonar small target detection method based on an improved attention mechanism, which is implemented based on YOLOv8 and includes:

[0011] Insert the following process into the 10th layer of the YOLOv8 backbone network:

[0012] The original image is upsampled, and then 5x5 point convolution, vertical convolution and horizontal convolution are performed in sequence, and finally the data F' is obtained through the activation function;

[0013] The original image is downsampled, and then 1x1 point convolution, horizontal convolution and vertical convolution are performed in sequence, and finally the data F" is obtained through the activation function;

[0014] The data F', data F" and the original image data are point-multiplied, and then concatenated with the original image data as the final output result.

[0015] As an improvement of the above method, the up-sampling adopts a bilinear interpolation method.

[0016] As an improvement of the above method, the downsampling adopts an average pooling method.

[0017] As an improvement of the above method, the vertical convolution is a 1x7 convolution.

[0018] As an improvement of the above method, the horizontal convolution is a 7x1 convolution.

[0019] As an improvement of the above method, the activation function is Sigmoid.

[0020] The present application also provides an image sonar small target detection system based on an improved attention mechanism, which is implemented based on the above method, and the system includes:

[0021] The upsampling module is used to upsample the original image, and then perform 5x5 point convolution, vertical convolution and horizontal convolution in sequence, and finally obtain the data F' through the activation function;

[0022] The downsampling module is used to downsample the original image, and then perform 1x1 point convolution, horizontal convolution and vertical convolution in sequence, and finally obtain the data F" through the activation function;

[0023] The dot product module is used to perform dot product on the data F', data F" and the original image data, and then concatenate them with the original image data as the final output result.

[0024] Compared with the prior art, the advantages of this application are:

[0025] This method allows the model to learn image features of different dimensions through up and down sampling, and then aggregates them in the horizontal and vertical directions. Finally, it uses residual connections to accelerate convergence, which can output more refined results, reduce background interference, and reduce computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 The figure shows a flow chart of the image sonar small target detection method based on the improved attention mechanism;

[0027] Figure 2(a) shows a schematic diagram of vertical and horizontal decoupled convolution;

[0028] Figure 2(b) shows a schematic diagram of zoomed attention aggregation;

[0029] Figure 3 The figure shows the structure of YOLOv8 and the location where the attention mechanism is inserted. DETAILED DESCRIPTION

[0030] The technical solution of the present application is described in detail below with reference to the accompanying drawings.

[0031] Since small targets are sparse and the background accounts for a large part in sonar images, applying the attention mechanism helps focus on key areas and reduce background interference.

[0032] Example 1

[0033] like Figure 1 As shown, this application proposes an image sonar small target detection method (SAMS) based on an improved attention mechanism. This method first infers the upsampled attention map in parallel. and downsampling The attention map is then inferred through vertical and horizontal convolutions of different orders to obtain the final attention results F' and F", and finally the residual connection is used to accelerate the convergence (as shown in Formula 1, Figure 1 ). This application uses bilinear interpolation for upsampling and average pooling for downsampling, and enhances the model's attention to local and global information through up and down sampling. After upsampling, a 5x5 point convolution is performed first, and then 1x7 and 7x1 vertical and horizontal convolutions are used to optimize computational efficiency and performance. After downsampling, a 1x1 point convolution is performed first, and then 7x1 and 1x7 horizontal and vertical convolutions are used to effectively fuse information of different dimensions and improve robustness. The complexity is reduced by adjusting the order of convolutions to enhance the perception of small targets. As shown in Figures 2(a) and 2(b), Figure 2(a) is a schematic diagram of the structure of a depth-separable convolution, which consists of point convolution, horizontal convolution, and vertical convolution. The upper half of Figure 2(b) represents the convolution part extracted after upsampling, the lower half represents the part after downsampling convolution, and the rightmost side represents the multiplication of the upsampling result and the downsampling result, that is, a schematic diagram of aggregating the up and down sampling results.

[0034]

[0035] in, The method uses ups and downsampling to allow the model to learn image features of different dimensions, and then aggregates them horizontally and vertically. Finally, residual connections are used to accelerate convergence and output more refined results.

[0036] like Figure 3 As shown, the method of the present application is inserted into the 10th layer of the YOLOv8 backbone network, and small targets are detected based on the YOLOv8 network.

[0037] The present invention is deployed on a machine equipped with an RTX3090 graphics card and embeds the proposed attention mechanism into the tenth layer of the YOLOv8 backbone network (e.g. Figure 3 ), and used the UCPR2021 dataset for experimental verification.

[0038] The URPC2021 dataset contains 6,000 forward-looking sonar images, covering 8 types of objects. The dataset is divided into training set, validation set, and test set, with a division ratio of 7:1:2. The default training strategy of YOLOv8 was adopted, and 350 rounds of training were performed using the SGD optimizer. The batch size was set to 32, the initial learning rate was 0.01, and the experimental input resolution was 640×640 to carry out target detection experiments.

[0039] In the experiment, object detection models such as RetinaNet, SSD, YOLOv3, YOLOv5, YOLOv6 and YOLOv8 were used for comparison. The results showed that YOLOv8-SAMS achieved 59.2% mAP under the same conditions, an increase of 2.9 percentage points compared to the original YOLOv8, proving the effectiveness of SAMS.

[0040] Table 1 Experimental results of URPC2021 under different baseline models

[0041]

[0042] Example 2

[0043] The present application also provides an image sonar small target detection system based on an improved attention mechanism, which is implemented based on the above method, and the system includes:

[0044] The upsampling module is used to upsample the original image, and then perform 5x5 point convolution, vertical convolution and horizontal convolution in sequence, and finally obtain the data F' through the activation function;

[0045] The downsampling module is used to downsample the original image, and then perform 1x1 point convolution, horizontal convolution and vertical convolution in sequence, and finally obtain the data F" through the activation function;

[0046] The dot product module is used to perform dot product on the data F', data F" and the original image data, and then concatenate them with the original image data as the final output result.

[0047] The present application may also provide a computer device, comprising: at least one processor, a memory, at least one network interface and a user interface. The various components in the device are coupled together through a bus system. It is understood that the bus system is used to achieve connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus.

[0048] The user interface may include a display, a keyboard or a pointing device, such as a mouse, a trackball, a touch pad or a touch screen.

[0049] It is understood that the memory in the embodiments disclosed in the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0050] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset thereof, or an extended set thereof: an operating system and applications.

[0051] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program for implementing the method of the embodiment of the present disclosure can be included in the application.

[0052] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in an application program, and is used to:

[0053] Execute the steps of the above method.

[0054] The above method can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The above-disclosed methods, steps and logic block diagrams can be implemented or executed. The general processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the above-disclosed method can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined to execute. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0055] It is understood that the embodiments described in the present application can be implemented by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application or a combination thereof.

[0056] For software implementation, the technology of the present application can be implemented by executing the functional modules (such as procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0057] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application is described in detail with reference to the embodiments, a person skilled in the art should understand that any modification or equivalent replacement of the technical solution of the present application does not depart from the spirit and scope of the technical solution of the present application and should be included in the scope of the claims of the present application.

Claims

1. A small target detection method for image sonar based on an improved attention mechanism, implemented based on YOLOv8, including: Insert the following process into the 10th layer of the YOLOv8 backbone network: The original image is upsampled, and then 5x5 point convolution, vertical convolution and horizontal convolution are performed in sequence, and finally the data F' is obtained through the activation function; The original image is downsampled, and then 1x1 point convolution, horizontal convolution and vertical convolution are performed in sequence, and finally the data F" is obtained through the activation function; The data F', data F" and the original image data are point-multiplied, and then concatenated with the original image data as the final output result.

2. The image sonar small target detection method based on improved attention mechanism according to claim 1 is characterized in that: The up-sampling adopts a bilinear interpolation method.

3. The image sonar small target detection method based on improved attention mechanism according to claim 1 is characterized in that: The downsampling adopts an average pooling method.

4. The image sonar small target detection method based on improved attention mechanism according to claim 1 is characterized in that: The vertical convolution is a 1x7 convolution.

5. The image sonar small target detection method based on improved attention mechanism according to claim 1 is characterized in that: The horizontal convolution is a 7x1 convolution.

6. The image sonar small target detection method based on improved attention mechanism according to claim 1 is characterized in that: The activation function is Sigmoid.

7. An image sonar small target detection system based on an improved attention mechanism, implemented based on the method described in any one of claims 1 to 6, characterized in that: The system comprises: The upsampling module is used to upsample the original image, and then perform 5x5 point convolution, vertical convolution and horizontal convolution in sequence, and finally obtain the data F' through the activation function; The downsampling module is used to downsample the original image, and then perform 1x1 point convolution, horizontal convolution and vertical convolution in sequence, and finally obtain the data F" through the activation function; and The dot product module is used to perform dot product on the data F', data F" and the original image data, and then concatenate them with the original image data as the final output result.

Citation Information

Patent Citations

  • Forward-looking sonar image underwater multi-target tracking method based on deep learning

    CN116883766A

  • Traffic sign detection algorithm based on improved YOLOv8

    CN118470684A

  • Deep convolutional neural network with self-transfer learning

    US20190122360A1

Cited By

  • Underwater acoustic positioning self-adaptive smooth filtering method and system

    CN120847710A

  • An underwater acoustic positioning adaptive smoothing filtering method and system

    CN120847710B