Sonar target detection method based on improved YOLOv5 and edge end reasoning equipment
By improving the YOLOv5 sonar object detection method, combined with block cutting and improved C3_MIX_attention module, the problem of low accuracy of traditional underwater sonar object detection technology is solved, efficient and accurate underwater target recognition is achieved, and the detection ability of complex targets is enhanced.
Patent Information
- Application Number
- CN202411865662.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional underwater sonar target detection technology relies on manual feature extraction, making it difficult to adapt to complex and changeable underwater environments, resulting in low detection accuracy, frequent missed detection and misdetecting, making it difficult to achieve high accuracy identification of underwater targets of different sizes and forms.
The sonar object detection method based on improved YOLOv5 is adopted. By constructing an underwater forward-looking sonar image sample set, block cropping and target position label conversion are carried out, an improved YOLOv5 sonar object detection network is built, and an input feature processing module is added before the input layer, and the C3 module is improved into a C3_MIX_attention module to improve the target feature extraction capability.
It realizes fast, efficient and accurate underwater target recognition, improves the feature capture ability of complex targets with sparse information, enhances the detection ability of small forward sonar targets, reduces the missed detection rate, and improves detection continuity.
Smart Images

Figure CN120047659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image target detection, and particularly to a sonar target detection method and an edge-side inference device. Background Art
[0002] In many fields such as ocean engineering, port operation, and underwater detection, it is crucial to accurately and efficiently identify and monitor underwater targets. As a key means of underwater detection, sonar technology can obtain relevant information of underwater objects by emitting sound waves and receiving reflected signals. However, the original sonar data mostly appears in the form of complex images, and advanced image processing and target detection technologies are needed to achieve effective interpretation. The target detection technology in the field of computer vision has developed vigorously, and deep learning algorithms are particularly favored. Among them, the YOLO (You Only Look Once) series is widely used in various scenarios due to its efficient and real-time detection characteristics, from security monitoring to autonomous driving, and is now gradually penetrating into underwater target detection work. The edge computing platform, with certain computing power and portability, provides hardware support for processing sonar data nearby in complex actual environments and timely feedback of target detection results. The technical system integrating sonar imaging, deep learning target detection algorithms, and edge computing platforms has emerged as the times require.
[0003] Traditional underwater sonar target detection technologies often rely on manual feature extraction methods based on signal processing. In the initial stage, basic image processing means such as threshold segmentation and edge detection are mostly used to preprocess sonar images and attempt to separate targets from the background. Technicians set gray thresholds manually based on experience to distinguish the part with brightness higher than the threshold (presumably the target) from the part lower than the threshold (regarded as the background) in the image; or use edge detection operators such as Sobel and Canny to outline the target contour. In the target recognition stage, the template matching method is adopted. A standard template library of common underwater targets (such as ships, reefs, etc.) is pre-constructed, and each part of the processed sonar image is compared with the template one by one. Whether the target is matched is measured based on the similarity. At the algorithm operation platform level, it mostly relies on large onshore data centers or local high-performance computers. The collected sonar data is remotely transmitted back to the central computer room for centralized processing and analysis, and the results are then fed back to the on-site operation end after processing.
[0004] Traditional technologies have exposed many limitations. In terms of detection accuracy, manual feature extraction highly depends on human experience. The selection of thresholds and the design of templates are difficult to adapt to complex and changeable underwater environments (such as the sonar image being blurred due to water flow and light and shadow interference), resulting in frequent missed detections and false detections, and it is difficult to achieve high-accuracy recognition of underwater targets of different sizes and shapes. Summary of the Invention
[0005] In view of this, the present invention proposes a sonar target detection method and an edge-side inference device based on improved YOLOv5, which can quickly, efficiently, and accurately identify underwater targets.
[0006] To solve the above technical problems, the present invention is implemented as follows.
[0007] Further, a sonar target detection method based on improved YOLOv5 includes the following steps:
[0008] Step 1: Construct an underwater forward-looking sonar image sample set, and perform position annotation on the targets in the images to generate label files;
[0009] Step 2: Perform block cropping on the images in the sample set, and perform automatic conversion of the target position labels in the cropped image blocks;
[0010] Step 3: Construct an underwater forward-looking sonar image training, validation, and test data set;
[0011] Step 4: Construct an improved YOLOv5 sonar target detection network;
[0012] Step 5: Use the training set and validation set to train and perform real-time online evaluation on the sonar target detection network, and use the test set to verify the performance of the sonar target detection network;
[0013] Step 6: Obtain an underwater forward-looking sonar image to be recognized, and use the sonar target detection network that has passed the performance verification to obtain the type and position information of the target box, so as to achieve target detection.
[0014] Further, collect the original underwater forward-looking sonar image data, parse and form a fan-shaped image, store the image and name it according to the actual acquisition time and frame number; label the fan-shaped image, and the label is the relative position of the target in the image.
[0015] Further, in step S2, it includes: crop the fan-shaped image horizontally and vertically to obtain multiple images, and there is pixel overlap between adjacent images; perform coordinate conversion on the relative position of the target in the fan-shaped image before cropping to obtain the relative position of the target in the cropped image, and generate a new label file.
[0016] Further, compared with the YOLOv5 target detection network, the improved YOLOv5 sonar target detection network adds an input feature processing module before the input layer, and improves the C3 module in the backbone feature extraction network to a C3_MIX_attention module, where:
[0017] The input of the input feature processing module is 3 images with the same cropping position and consecutive in the time series, and the size is H 0 ×W0 The sonar images A, B, and C of 3×, the module performs the following operations on the input: convert the 3 pictures into grayscale images, and obtain 3 sonar pictures A1, B1, and C1 with a size of H 0 ×W 0 ×1; create a three-dimensional matrix O with a size of H 0 ×W 0 ×3, put the grayscale data in A1 into the first layer of O; perform point-by-point subtraction on A1 and B1, and send the grayscale data in the subtracted grayscale image A1 - B1 into the second layer of O; perform point-by-point subtraction on A1 and C1, and send the grayscale data in the subtracted grayscale image A1 - C1 into the third layer of O; the three-dimensional matrix O is the output of the input feature processing module;
[0018] The C3_MIX_attention module includes a C3 module, which performs average pooling, max pooling, and convolution operations on the output of the C3 module respectively to form three branches, and obtains 3 feature maps A2, M, and C2 with a size of H×W×1, where the output size of the C3 module is denoted as H×W×C; splice these three feature maps A2, M, and C2 on the channel to obtain a feature map MIX with a size of H×W×3; compress the feature map MIX into a feature map MIX_attention with a size of H×W×1 through a convolution module; after multiplying MIX_attention and the output of the C3 module bit by bit, and then adding them bit by bit with the output of the C3 module to obtain the output of this module.
[0019] A sonar target detection edge - end inference device based on improved YOLOv5, the device includes:
[0020] Data parsing module: configured to parse the forward - looking sonar data to generate a fan - shaped diagram;
[0021] Recognition network module: configured to construct an improved YOLOv5 target detection network;
[0022] Edge - end recognition module: configured to obtain the underwater forward - looking sonar image to be recognized, input it into the trained improved YOLOv5 sonar target detection network, and obtain the type and position information of the target box of the underwater forward - looking sonar image to be recognized.
[0023] Beneficial effects:
[0024] (1) In the present invention, by processing and fusing multiple consecutive sonar images in the time series, the target features in the sonar images can be better extracted. This enables more accurate capture of the features of complex targets with sparse information, thereby improving the accuracy of target detection. The present invention has a stronger detection ability for small forward - looking sonar targets and a lower missed - detection rate.
[0025] (2) The present invention has higher detection continuity for targets in continuous forward-looking sonar.
[0026] (3) The improved YOLOv5 target detection algorithm used in the present invention uses common operators, which are well supported on edge devices and are easy to implement. Description of the Drawings
[0027] Figure 1 is a schematic flow diagram of the sonar target detection method based on improved YOLOv5 provided by the present invention;
[0028] Figure 2 is a network structure diagram of the input feature processing module;
[0029] Figure 3 is a structure diagram of the C3_MIX_attention module;
[0030] Figure 4 is a network structure diagram of the original YOLOv5;
[0031] Figure 5 is a system block diagram of the edge-side inference device described in the present invention. Detailed Embodiments
[0032] The present invention will be described in detail below with reference to the drawings and embodiments.
[0033] As Figures 1-4 shown, the present invention proposes a sonar target detection method based on improved YOLOv5, including the following steps:
[0034] As Figure 1 shown, a sonar target detection method based on improved YOLOv5 includes the following steps:
[0035] Step S1: Construct an underwater forward-looking sonar image sample set, and generate a label file by annotating the positions of targets in the images;
[0036] Step S1 specifically includes: collecting original underwater forward-looking sonar image data, using a conventional sonar mapping algorithm to parse the original underwater forward-looking sonar image data to form a fan-shaped map, storing the pictures and naming them according to the actual collection time and frame number; using the labelImg tool to annotate the relative positions of targets in the fan-shaped pictures to generate a label file in txt format, and each line of the file includes the category of the target, the x coordinate, y coordinate of the target center in the image, the width and height information of the target, and packing the pictures and the corresponding label files to construct an underwater forward-looking sonar image sample set.
[0037] Step S2: Perform block cropping on the images in the underwater forward-looking sonar image sample set and automatically convert the target position labels in the cropped image blocks;
[0038] Step S2 specifically includes: cropping the images in the underwater forward-looking sonar image sample set. The original images with a resolution of 1200×2300 are cropped into 3 images horizontally and 2 images vertically. The adjacent images overlap by 60 pixels. Then, the cropped images are scaled, and finally 6 images with a resolution of 640×640 are obtained. Automatically convert the coordinate positions of the label files of the relative positions of the targets in the original images to obtain the relative positions of the targets in the cropped images, and generate 6 new label files representing the relative positions of the targets in the images.
[0039] Step S3: Construct an underwater forward-looking sonar image training, validation, and test data set;
[0040] Step S4: Construct an improved YOLOv5 sonar target detection network;
[0041] Figure 4 As shown in the structure diagram of the original YOLOv5 target detection network, compared with the original YOLOv5 target detection network, the improved YOLOv5 sonar target detection network adds an input feature processing module before the input layer and improves the C3 module in the backbone feature extraction network to the C3_MIX_attention module, where:
[0042] As Figure 2 shown, the input of the input feature processing module is 3 sonar images A, B, and C with the same cropping position and consecutive in the time series, and the size of 640×640×3. The module performs the following operations on the input: convert the 3 images into grayscale images to obtain 3 sonar images A1, B1, and C1 with the size of 640×640×1; create a three-dimensional matrix O with the size of 640×640×3, and put the data in A1 into the first layer of O; perform point-by-point subtraction on A1 and B1, and send the data in the subtracted grayscale image A1 - B1 into the second layer of O; perform point-by-point subtraction on A1 and C1, and send the data in the subtracted grayscale image A1 - C1 into the third layer of O; the three-dimensional matrix O is the output of the input feature processing module. After the original large image in Step S2 is cropped into 6 new images, for each cropped position, find the images consecutive in the time series. It is equivalent to performing six identifications on the original large image. By processing the input consecutive 3-frame sonar images through the input feature processing module, the image information of the current frame and the motion information of the target in the short term can be obtained, so as to assist the recognition model in detecting the target.
[0043] As Figure 3As shown, the C3_MIX_attention module includes a C3 module. The output of the C3 module is respectively subjected to average pooling, max pooling and convolution operations to form three branches, and three feature maps A2, M, and C2 with a scale of H×W×1 are obtained. Among them, the output scale of the C3 module is denoted as H×W×C. These three feature maps A2, M, and C2 are concatenated in the channel dimension to obtain a feature map MIX with a scale of H×W×3. A convolutional module is used to compress the feature map MIX into a feature map MIX_attention with a scale of H×W×1. After multiplying MIX_attention and the output of the C3 module bitwise, and then adding them bitwise to the output of the C3 module, the output of this module is obtained.
[0044] Step S5: Use the underwater forward-looking sonar image training set and validation set to train and perform real-time online evaluation on the improved YOLOv5 sonar target detection network to obtain the trained improved YOLOv5 sonar target detection network, and use the test set to verify the final algorithm performance;
[0045] Step S6: Deploy the recognition algorithm to the edge device, obtain the underwater forward-looking sonar image to be recognized, and input it into the trained improved YOLOv5 sonar target detection network to obtain the type and position information of the target box of the underwater forward-looking sonar image to be recognized.
[0046] As Figure 5 shown, a sonar target detection edge AI inference device based on improved YOLOv5, the device includes:
[0047] Data parsing module: configured to parse the forward-looking sonar data to generate a fan-shaped diagram;
[0048] Recognition network module: configured to construct an improved YOLOv5 target detection network;
[0049] Edge recognition module: configured to obtain the underwater forward-looking sonar image to be recognized, and input it into the trained improved YOLOv5 sonar target detection network to obtain the type and position information of the target box of the underwater forward-looking sonar image to be recognized.
[0050] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not limited. Therefore, those skilled in the art of the present invention can modify or equivalently replace the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the spirit and technical solutions of the present invention, and should all fall within the protection scope of the present invention.
Claims
1. A sonar target detection method based on improved YOLOv5, characterized in that: The steps include: Step 1: Construct a sample set of underwater forward-looking sonar images, mark the locations of targets in the images and generate label files; Step 2: Crop the image of the sample set into blocks, and automatically convert the target position label in the cropped image block; Step 3: Construct underwater forward-looking sonar image training, verification, and test datasets; Step 4: Build an improved YOLOv5 sonar target detection network; Step 5: Use the training set and the validation set to train and conduct real-time online evaluation on the sonar target detection network, and use the test set to verify the performance of the sonar target detection network; Step 6: Obtain the underwater forward-looking sonar image to be identified, and use the sonar target detection network that has passed performance verification to obtain the type and position information of the target frame to achieve target detection.
2. The method according to claim 1, characterized in that The step S1 includes: collecting underwater forward-looking sonar raw image data, parsing to form a fan-shaped image, storing the image and naming it according to the actual collection time and frame number; labeling the fan-shaped image, wherein the label is the relative position of the target in the image.
3. The method according to claim 1, characterized in that The step S2 includes: cropping the fan-shaped image horizontally and vertically to obtain multiple images, with pixels of adjacent images overlapping; performing coordinate transformation on the relative position of the target in the fan-shaped image before cropping, obtaining the relative position of the target in the cropped image, and generating a new label file.
4. The method according to any one of claims 1 to 3, characterized in that: Compared with the YOLOv5 target detection network, the improved YOLOv5 sonar target detection network adds an input feature processing module before the input layer, and improves the C3 module in the backbone feature extraction network to a C3_MIX_attention module, wherein: The input feature processing module inputs three sonar images A, B, C of size H0×W0×3 with the same cropping position and continuous in time series. The module performs the following operations on the input: convert the three images into grayscale images, and obtain three sonar images A1, B1, C1 of size H0×W0×1; create a three-dimensional matrix O of size H0×W0×3, and put the grayscale data in A1 into the first layer of O; perform point-by-point difference on A1 and B1, and send the grayscale data in the grayscale image A1-B1 after the difference to the second layer of O; perform point-by-point difference on A1 and C1, and send the grayscale data in the grayscale image A1-C1 after the difference to the third layer of O; the three-dimensional matrix O is the output of the input feature processing module; The C3_MIX_attention module includes a C3 module, which performs average pooling, maximum pooling and convolution operations on the output of the C3 module to form three branches, and obtains three feature maps A2, M, and C2 with a scale of H×W×1, where the output scale of the C3 module is recorded as H×W×C; the three feature maps A2, M, and C2 are spliced on the channel to obtain a feature map MIX with a scale of H×W×3; the feature map MIX is compressed into a feature map MIX_attention with a scale of H×W×1 through a convolution module; MIX_attention is bit-wise multiplied with the output of the C3 module, and then bit-wise added with the output of the C3 module to obtain the output of the module.
5. A sonar target detection edge inference device based on improved YOLOv5, characterized in that: The device comprises: Data parsing module: configured to parse forward-looking sonar data and generate a fan-shaped graph; Recognition network module: configured to build an improved YOLOv5 target detection network; The edge recognition module is configured to obtain an underwater forward-looking sonar image to be identified, input it into the trained improved YOLOv5 sonar target detection network, and obtain the type and position information of the target box of the underwater forward-looking sonar image to be identified.