Sonar target detection method based on rotating frame
Through the sonar object detection method based on rotary frames, combined with the multi-frame compression input layer and feature pyramid network, the problems of sonar detection are solved, and efficient underwater object detection and recognition are achieved.
Patent Information
- Application Number
- CN202311633632.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-08
AI Technical Summary
The existing sonar object detection methods are easily affected by the underwater sound propagation environment and are poorly robust. The horizontal frame-based detection method brings background information, resulting in insufficient detection accuracy.
A sonar object detection method based on rotation frame is adopted to build a multi-frame compressed input layer, combine partial networks and feature pyramid networks to perform feature fusion, add angle information, use rotation frames to perform object detection, and build a multi-task loss function optimization model.
Effectively reduce the background information of the detection frame, improve detection accuracy, adapt to the random direction of the target in the sonar image, and achieve fast and high-precision underwater target detection.
Smart Images

Figure CN120279398A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection and relates to a sonar target detection method based on a rotated bounding box. Background Art
[0002] Current sonar target detection methods include traditional sonar target detection methods and deep learning-based sonar target detection methods. Traditional sonar target detection methods extract features manually and are easily affected by the underwater sound propagation environment, with poor robustness. In the past, deep learning-based sonar target detection adopted a horizontal bounding box prediction method. However, due to the randomness of the target angle in sonar images affected by the imaging direction, the detection method based on horizontal bounding boxes inevitably brings in a lot of background information. In this paper, a detection method based on a rotated bounding box by adding angle information on the basis of the YOLOv5 horizontal bounding box is used for sonar target detection. This method can effectively reduce the background information of the detection bounding box and improve the accuracy of sonar target detection. Summary of the Invention
[0003] In view of this, the present invention provides a sonar target detection method based on a rotated bounding box, including:
[0004] Step S1: Construct a multi-frame compressed input layer and a sonar target detection model based on a rotated bounding box. The sonar target detection model at least includes a feature extraction module based on a cross-stage partial network, a feature fusion module based on a Feature Pyramid Network (FPN) and a Path Aggregation Network (PAN), and a target prediction module;
[0005] Step S2: Collect forward-looking sonar beam data and transform it to generate a forward-looking sonar image; preprocess and annotate the forward-looking sonar image to construct a forward-looking sonar dataset;
[0006] Step S3: Set training hyperparameters and train the sonar target detection model based on the rotated bounding box using the forward-looking sonar dataset;
[0007] Step S4: After training, test the sonar target detection model based on the rotated bounding box and record the test results;
[0008] Step S5: If the test results meet the index requirements, this model can be applied to real-time forward-looking sonar target detection. Otherwise, adjust the training hyperparameters until the test results meet the index requirements;
[0009] Specifically, in the step S1, constructing the multi-frame compressed input layer includes: setting n consecutive frames of sonar grayscale images {I j, j = 1, 2, 3, ...}, construct an n-channel image in sequence and use this n-channel image as the model input, effectively solving the problems of sonar image blurring and few target feature information; the feature extraction module based on the cross-stage partial network processes the input feature image in the following way: divide the feature image into two parts, one part outputs a new feature map through convolution operation, and the other part directly skips the connection through a bypass, and finally concatenates the two part feature maps together to achieve cross-stage feature fusion, and finally outputs feature images of three scales.
[0010] In particular, in step S1, it can be carried out from top to bottom through the Feature Pyramid Network (FPN) to transfer deep semantic
[0011] information to the shallow layer for feature fusion, and transfer the shallow layer's localization information to the deep layer through the Path Aggregation Network (PAN) for feature fusion from the shallow layer to the deep layer.
[0012] In particular, step S2 specifically includes: deploying an underwater target, using an autonomous underwater vehicle equipped with a forward-looking sonar to collect forward-looking sonar beam data of the deployed target under multiple working conditions, and mapping the forward-looking beam data in the plane coordinates into a fan-shaped picture in polar coordinates through coordinate transformation; performing single-sample enhancement and multi-sample enhancement on the fan-shaped picture; the single-sample enhancement includes: random cropping, scaling or color change; the multi-sample enhancement includes mixed sample data enhancement; during annotation, first compress continuous n-frame sonar gray images to generate a multi-channel image, with the target position and label based on the middle frame, and use a rotating target annotation tool to annotate the preprocessed picture to generate an annotation file; jointly form a forward-looking sonar dataset with the forward-looking sonar pictures and annotation files.
[0013] In particular, in step S3, a training loss function is constructed based on the three-scale feature maps output by the target prediction module, and the training loss function includes the localization loss of the bounding box, the confidence loss, the class loss, and the angle loss.
[0014] In particular, the localization loss includes:
[0015]
[0016]
[0017]
[0018] Among them: CIou represents the localization loss, ρ 2 (b, b gt ) is the distance between the predicted box and the ground truth box, c is the length of the minimum diagonal of the minimum bounding box composed of the predicted box and the ground truth box, α is the balance ratio coefficient, w and h are the width and height of the predicted box, wgt , h gt is the width and height of the ground truth box; Iou represents the intersection over union, which represents the overlap degree between the predicted box and the ground truth box;
[0019] The confidence loss includes:
[0020]
[0021] where: S×S is the number of grids, B is the number of anchor boxes in each grid, indicates whether there is an object in the j-th anchor box of the i-th grid. obj means the predicted box contains an object, and the corresponding value is 1. noobj means the predicted box does not contain an object, and the corresponding value is 0, C i is the confidence;
[0022] The class loss includes:
[0023]
[0024] where: indicates whether there is an object in the j-th anchor box of the i-th grid. If there is, the value is 1, otherwise it is 0. noobj means the predicted box does not contain an object, and the corresponding value is 0; classes is the class set, p i (c) is the probability of the i-th category;
[0025] The angle loss includes:
[0026]
[0027] where: θ is the predicted angle value, θ gt is the ground truth value.
[0028] Specifically, in the step S5, the sonar target detection model based on the rotated box is evaluated based on precision and recall.
[0029] Beneficial effects:
[0030] Through the sonar target detection method based on the rotated box of the present invention, the background information included in the detection box can be effectively reduced, making the detection more accurate;
[0031] Through the multi-frame compression input layer constructed by the detection model based on the rotated box of the present invention, the problems of sonar image blurring and less target feature information are effectively solved;
[0032] Through the detection model based on the rotated box constructed by the present invention, the rotation angle of the target can be detected to adapt to the random directions of the targets in the sonar image;
[0033] Through the network modules such as CSPDarknet and FPN adopted by the present invention, the model has stronger feature expression ability;
[0034] The multi-task loss function including rotation angle prediction designed by the present invention can jointly optimize the classification, localization and angle prediction capabilities of the model;
[0035] Through the training, verification and deployment processes of the present invention, a sonar target detection model with relatively high accuracy and recall rate can be obtained quickly;
[0036] Through the method of the present invention, the automatic detection and recognition of underwater targets can be realized, providing accurate input information for subsequent algorithms such as underwater target tracking and localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic flowchart of the sonar target detection method based on a rotated bounding box proposed in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.
[0039] The present invention provides a sonar target detection method based on a rotated bounding box, as Figure 1 shown, including:
[0040] Step S1: Construct a multi-frame compressed input layer and a sonar target detection model based on a rotated bounding box. The sonar target detection model at least includes a feature extraction module based on a cross-stage partial network, a feature fusion module based on a feature pyramid network FPN and a path aggregation network PAN, and a target prediction module. In step S1, the feature extraction module based on the cross-stage partial network processes the input feature image in the following manner: the feature image is divided into two parts, one part outputs a new feature map through convolutional operations, and the other part directly skips the connection through a bypass. Finally, the two part feature maps are concatenated together to achieve cross-stage feature fusion, and finally three scales of feature images are output. In this embodiment, the CSPDarknet network can be selected as the backbone network, which has good feature extraction ability and calculation speed.
[0041] In step S1, constructing the multi-frame compressed input layer includes: setting n consecutive frames of sonar grayscale images {I j , j = 1, 2, 3,...}, constructing an n-channel image in sequence, and using this n-channel image as the model input, effectively solving the problems of blurred sonar images and less target feature information; performing feature fusion from deep to shallow through the feature pyramid network FPN, and performing feature fusion from shallow to deep through the path aggregation network PAN.
[0042] The Feature Pyramid Network (FPN) can optionally extract feature maps of different scales from the last three stages of the backbone network, usually denoted as C3, C4, and C5, where C5 has the lowest resolution but the richest semantic information. The number of channels of all feature maps is adjusted to a unified size through a 1x1 convolutional layer. C5 is upsampled to the same size as C4 and added element-wise to C4 to obtain a semantic feature map P5 with a higher resolution; the upsampling and fusion operations are repeated to obtain P4 and P3, and the feature pyramid is constructed layer by layer.
[0043] The Path Aggregation Network (PAN) can optionally extract the feature map P5 from the bottom of the FPN pyramid as the input to the PAN module. The extracted feature map P5 undergoes a 1x1 convolution and upsampling to obtain a higher-resolution P5_up. P5_up is added element-wise to P4 of the pyramid to obtain P4' that is fused from the bottom and the top; the above operations are repeated and propagated layer by layer upward to achieve bottom-up path aggregation.
[0044] Finally, through the top-down operation of the Feature Pyramid Network (FPN), deep semantic information can be transmitted to the shallow layer for feature fusion, and through the Path Aggregation Network (PAN), shallow localization information can be transmitted to the deep layer for feature fusion from the shallow layer to the deep layer; by jointly using the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN), the feature fusion module can obtain a complete feature map that retains both global information and local details, thereby improving the performance of object detection.
[0045] Step S2: Collect forward-looking sonar beam data and transform it to generate forward-looking sonar images; preprocess and annotate the forward-looking sonar images to construct a forward-looking sonar dataset; specifically included in the step S2 are: deploying underwater targets, using an underwater unmanned vehicle equipped with a forward-looking sonar to collect forward-looking sonar beam data for the deployed targets under multiple working conditions, and mapping the forward-looking beam data in the plane coordinates into a fan-shaped image in polar coordinates through coordinate transformation; performing single-sample enhancement and multi-sample enhancement on the fan-shaped image; the single-sample enhancement includes: random cropping, scaling, or color change; the multi-sample enhancement includes mixed sample data enhancement; during annotation, first compress n consecutive frames of sonar grayscale images to generate a multi-channel image, with the target position and label based on the middle frame, and use a rotating target annotation tool to annotate the preprocessed images to generate an annotation file; jointly form a forward-looking sonar dataset with the forward-looking sonar images and the annotation file. In this embodiment, the rotating target annotation tool RoLabel can be used to annotate the preprocessed images to generate an annotation file, and the forward-looking sonar images and the annotation file jointly form a forward-looking sonar dataset. Randomly select the preprocessed images and annotation files as the training set according to a preset ratio, and the remaining images and annotation files as the test set.
[0046] Among them, the training set and the test set constitute the dataset of the sonar target detection model based on rotated bounding boxes. The preset ratio is a pre-set ratio, which can be 70%.
[0047] Step S3: Set training hyperparameters and train the sonar target detection model based on rotated bounding boxes using the forward-looking sonar dataset; in this step, the hyperparameters for training the sonar target detection model based on rotated bounding boxes are initialized.
[0048] For example, during training, the number of training times is 300, the batch size is 16, the initial learning rate is 0.05, the momentum decay is 0.897, and the weight decay is 0.0003. The training set data is used as the input of the sonar target detection model based on rotated bounding boxes, and the training loss function is used to perform 300 times of training. After training, the sonar target detection model based on rotated bounding boxes is obtained. The target prediction module outputs feature maps of three scales. In addition to the localization information, confidence information, and category information, the angle information is additionally added. Let the number of target categories be C, and the width and height of the feature map be W and H. Then the size of the feature map is (6 + C) * W * H, where 6 represents the center coordinates, width, height, angle, and confidence of the prediction box.
[0049] In step S3, a training loss function is constructed based on the feature maps of the three scales output by the target prediction module. The training loss function includes the localization loss, confidence loss, category loss, and angle loss of the bounding box.
[0050] The localization loss includes:
[0051]
[0052]
[0053]
[0054] Among them: CIou represents the localization loss, ρ 2 (b, b gt ) is the distance between the prediction box and the ground truth box, c is the length of the minimum diagonal of the minimum bounding box formed by the prediction box and the ground truth box, α is the balance ratio coefficient, w and h are the width and height of the prediction box, w gt 、h gt are the width and height of the ground truth box; Iou represents the intersection over union, representing the overlap degree between the prediction box and the ground truth box;
[0055] The confidence loss includes:
[0056]
[0057] Among them: S×S is the number of grids, B is the number of anchor boxes for each grid, Whether there is a target in the j-th anchor box of the i-th grid. "obj" indicates that the prediction box contains a target, and the corresponding value is 1; "noobj" indicates that the prediction box does not contain a target, and the corresponding value is 0. C i is the confidence;
[0058] The class loss includes:
[0059]
[0060] Among them: Whether there is a target in the j-th anchor box of the i-th grid. If there is, it is 1, otherwise it is 0. "noobj" indicates that the prediction box does not contain a target, and the corresponding value is 0; "classes" is the class set, p i (c) is the probability of the i-th category;
[0061] The angle loss includes:
[0062]
[0063] Among them: θ is the predicted angle value, θ gt is the true value.
[0064] Step S4: After training, test the sonar target detection model based on the rotated bounding box and record the test results.
[0065] Evaluate the sonar target detection model based on the rotated bounding box according to precision and recall;
[0066] As an example, according to the test results of S4, the sonar target detection model based on the rotated bounding box can be evaluated according to precision and recall; precision is the proportion of correctly detected targets among the detected targets, and recall is the proportion of correctly detected targets among the actual number of targets. The higher the two indicators, the better. For example, the precision is set to 80% and the recall is set to 90%. If the index requirements are not met, modify the training hyperparameters and retrain until the test results meet the index requirements.
[0067] Step S5: If the test results meet the index requirements, this model can be applied to real-time forward-looking sonar target detection. Otherwise, adjust the training hyperparameters until the test results meet the index requirements;
[0068] In the said step S5, the sonar target detection model based on the rotated bounding box is evaluated according to precision and recall.
[0069] In summary, the above is only the preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0070] For those skilled in the art, it is obvious that the embodiments of the present invention are not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the embodiments of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the embodiments of the present invention. Any reference signs in the claims should not be construed as limiting the claimed invention. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units, modules or devices described in the system, apparatus or terminal claims can also be implemented by the same unit, module or device through software or hardware. The words such as first, second, etc. are used to denote names and do not represent any particular order.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention and not to limit them. Although the technical solutions of the embodiments of the present invention have been described in detail with reference to the above preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the embodiments of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sonar target detection method based on a rotated bounding box, characterized in that, Including: Step S1: Construct a multi-frame compressed input layer and a sonar target detection model based on a rotated bounding box. The sonar target detection model at least includes a feature extraction module based on a cross-stage partial network, a feature fusion module based on a Feature Pyramid Network (FPN) and a Path Aggregation Network (PAN), and a target prediction module. Step S2: Collect forward-looking sonar beam data and transform it to generate forward-looking sonar images; preprocess and annotate the forward-looking sonar images to construct a forward-looking sonar dataset. Step S3: Set training hyperparameters and train the sonar target detection model based on the rotated bounding box using the forward-looking sonar dataset. Step S4: After training, test the sonar target detection model based on the rotated bounding box and record the test results. Step S5: If the test results meet the index requirements, this model can be applied to real-time forward-looking sonar target detection. Otherwise, adjust the training hyperparameters until the test results meet the index requirements.
2. The sonar target detection method based on a rotating box according to claim 1, characterized in that, In the step S1, constructing a multi-frame compressed input layer includes: setting consecutive n frames of sonar grayscale images {I j , j = 1, 2, 3,...}, constructing an n-channel image in sequence, and using the n-channel image as the model input, effectively solving the problems of sonar image blurring and less target feature information; the feature extraction module based on the cross-stage partial network processes the input feature image in the following way: dividing the feature image into two parts, one part outputs a new feature map through convolution operations, and the other part directly skips the connection through a bypass, and finally concatenating the two part feature maps together to achieve cross-stage feature fusion, and finally outputting feature images of three scales.
3. The method for detecting a sonar target based on a rotated bounding box according to claim 2, wherein in step S1, it is carried out from top to bottom through the Feature Pyramid Network (FPN) network, which can transfer deep semantic information to the shallow layer for feature fusion, and the Path Aggregation Network (PAN) transfers the shallow layer's localization information to the deep layer for feature fusion from the shallow layer to the deep layer.
4. The method for detecting a sonar target based on a rotated bounding box according to claim 3, wherein step S2 specifically includes: deploying underwater targets, using an underwater unmanned vehicle equipped with a forward-looking sonar to collect forward-looking sonar beam data of the deployed targets under multiple working conditions, and mapping the forward-looking beam data in the plane coordinates into a fan-shaped image in polar coordinates through coordinate transformation; performing single-sample enhancement and multi-sample enhancement on the fan-shaped image; the single-sample enhancement includes: random cropping, scaling, or color change; the multi-sample enhancement includes mixed sample data enhancement; during annotation, first compress consecutive n frames of sonar grayscale images to generate a multi-channel image, with the target position and label based on the middle frame, and use a rotated target annotation tool to annotate the preprocessed images to generate annotation files; jointly form a forward-looking sonar dataset with the forward-looking sonar images and annotation files.
5. The sonar target detection method based on a rotating box according to claim 4, wherein In step S3, a training loss function is constructed based on the feature maps of three scales output by the target prediction module. The training loss function includes the localization loss of the bounding box, the confidence loss, the class loss, and the angle loss.
6. The method for detecting a sonar target based on a rotated bounding box according to claim 5, wherein the localization loss includes: Among them: CIou represents the localization loss, and ρ 2 (b, b gt ) is the distance between the predicted bounding box and the ground truth bounding box, c is the length of the minimum diagonal of the minimum enclosing box formed by the predicted bounding box and the ground truth bounding box, α is the balance ratio coefficient, w and h are the width and height of the predicted bounding box, w gt 、h gt are the width and height of the ground truth bounding box; in the step S1, it represents the intersection over union, representing the overlap degree between the predicted bounding box and the ground truth bounding box; the confidence loss includes: Where: S×S is the number of grids, and B is the number of anchor boxes for each grid. indicates whether there is an object in the j-th anchor box of the i-th grid. "obj" means the prediction box contains an object, and the corresponding value is 1; "noobj" means the prediction box does not contain an object, and the corresponding value is 0. C i is the confidence level. the class loss includes: Wherein: Indicates whether there is a target in the j-th anchor box of the i-th grid. If there is, it is 1; otherwise, it is 0. noobj indicates that the prediction box does not contain a target, and the corresponding value is 0; classes is the class set, p i (c) is the probability of the i-th category; the angle loss includes: Where: θ is the predicted angle value, and θ gt is the true value.
7. The method for detecting a sonar target based on a rotated bounding box according to claim 5, wherein in step S5, the sonar target detection model based on the rotated bounding box is evaluated based on precision and recall.