Underwater sonar image recognition method and system based on improved YOLOv8
By introducing a multi-dimensional parallel attention mechanism and a comparative learning architecture in the YOLOv8 model, the YOLOv8-MDPA-CL network model was constructed, and the problem that underwater sonar image recognition technology is difficult to identify small targets in low resolution and high noise environments is solved, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510124197.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The existing underwater sonar image recognition technology is difficult to accurately identify small targets in low-resolution and high-noise environments, and the existing algorithms have shortcomings in problems such as excessive parameter volume and training degradation.
Based on the method of improving YOLOv8, a multi-dimensional parallel attention mechanism (MDPA) and a comparative learning architecture are introduced to build the YOLOv8-MDPA-CL network model, and the model's ability to capture global information through the multi-dimensional parallel attention mechanism is improved, and the feature extraction is optimized through the comparative learning architecture to enhance the generalization ability and stability of the model.
It significantly improves the model's adaptability to complex scenes in underwater sonar images, improves the accuracy and robustness of target details recognition, and enhances the recognition performance in low resolution and noise environments.
Smart Images

Figure CN119992305A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and in particular relates to an underwater sonar image recognition method and system based on improved YOLOv8. Background Art
[0002] Underwater sonar image recognition technology has wide application value in the fields of marine exploration, environmental protection and unmanned system operation. However, the complex and changeable underwater environment often causes sonar images to have low resolution and high noise, which greatly increases the difficulty of target recognition, especially in the positioning and recognition of small targets.
[0003] With the development of deep learning technology, convolutional neural networks (CNNs) have greatly promoted the progress in the field of target detection. Existing algorithms are mainly divided into two-stage methods (such as regional convolutional neural networks and fast regional convolutional neural networks) and one-stage methods (such as the YOLO model and single-stage target detection algorithm). The one-stage method is more suitable for underwater target detection tasks due to its higher real-time performance. In addition, the introduction of Transformer in recent years has also opened up new directions for target recognition. For example, the combination of Transformer modules and YOLO networks can further improve the ability to capture sparse target features, and at the same time optimize the recognition performance of complex targets by improving the attention mechanism.
[0004] However, the existing models still have some shortcomings. First, the low resolution and noise of sonar images make it difficult to present target details. Second, most algorithms improve feature extraction capabilities by deepening the network depth, but this brings about the problems of excessive parameters and training degradation. Finally, small target detection performs poorly in existing algorithms and is easily affected by background interference, resulting in missed detection or false detection. Summary of the invention
[0005] The present invention provides an underwater sonar image recognition method based on improved YOLOv8, which improves the representation ability and generalization performance of the network model, and makes the network model more accurate in identifying target details in complex underwater environments.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] The first aspect of the present invention provides an underwater sonar image recognition method based on improved YOLOv8, comprising:
[0008] Obtain the sonar monitoring image of the monitoring area collected by the underwater sonar equipment, and input the sonar monitoring image into the preset YOLOv8-MDPA-CL network model to obtain the underwater target detection result;
[0009] The YOLOv8-MDPA-CL network model construction and training process includes:
[0010] Construct the YOLOv8-MDPA-CL network model based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture;
[0011] Sonar training images of monitoring targets in various set underwater scenes are obtained, and the sonar training images are preprocessed and real labels are added to obtain training samples. The YOLOv8-MDPA-CL network model is trained with the training samples to obtain training recognition results. The training loss value between the training recognition results and the real labels is calculated using the contrast loss function and the YOLOv8 loss function. The weight parameters of the YOLOv8-MDPA-CL network model are optimized according to the training loss value. The training process of the YOLOv8-MDPA-CL network model is iterated repeatedly until the iteration termination condition is reached and the trained YOLOv8-MDPA-CL network model is output.
[0012] Furthermore, the sonar monitoring image of the monitoring area collected by the underwater sonar equipment is obtained, specifically including:
[0013] The underwater sonar device is set as a multi-beam forward-looking sonar detection device, and the multi-beam forward-looking sonar detection device is controlled to release a 720kHz sonar signal and a 1200kHz sonar signal simultaneously in the monitoring area; the 720kHz sonar signal is used for target detection within a first distance range, and the 1200kHz sonar signal is used for target detection within a second distance range, and the maximum distance value within the first distance range is greater than the maximum distance value within the second distance range;
[0014] When performing sonar imaging of a target at multiple positions and angles, a forward-looking sonar annotation tool is used to annotate the same target in real time to obtain a sonar monitoring image.
[0015] Furthermore, the process of preprocessing the sonar training image includes:
[0016] The sonar training image is denoised using a Gaussian filtering algorithm, wherein the Gaussian kernel size is set to 3×3 or 5×5, and the standard deviation is set to 0;
[0017] The sonar training images were rotated by ±15°; the sonar training images were randomly flipped horizontally and vertically; and the sonar training images were randomly cropped and scaled to obtain training samples.
[0018] Furthermore, the YOLOv8-MDPA-CL network model is constructed based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture, including:
[0019] The backbone network based on the YOLOv8 network model and the multi-dimensional parallel attention mechanism are used to obtain an improved backbone network;
[0020] The improved backbone network includes an initial layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer; the initial layer is set as a CBL module, the first feature extraction layer, the second feature extraction layer and the third feature extraction layer include a CBL module and a C2f module connected in sequence; an MDPA module is constructed based on a multi-dimensional parallel attention mechanism, and the fourth feature extraction layer includes a CBL module, an SPPF module and an MDPA module connected in sequence;
[0021] The MDPA module is connected to a contrastive learning framework, which calculates a contrastive loss value according to a contrastive loss function, and the contrastive loss value is used to constrain feature extraction of the improved backbone network;
[0022] The YOLOv8-MDPA-CL network model is obtained based on the improved backbone network, contrastive learning architecture, neck network and head network of the YOLOv8 network model.
[0023] Furthermore, the MDPA module extracts features from the intermediate feature map output by the SPPF module to obtain an output feature map of the improved backbone network, specifically including:
[0024] The intermediate feature map is divided into multiple independent subspaces and converted into input vector X. The input vector X in each independent subspace is respectively mapped to the mapping matrix W Q , W K , W V Perform point-by-point convolution to generate query vector matrix Q, key vector matrix K and value vector matrix V;
[0025] Multiply the query vector matrix Q and the key vector matrix K to obtain the content attention score matrix; perform position encoding on the independent subspaces from the height and width dimensions and add them to obtain the position encoding matrix R; multiply the position encoding matrix R and the query vector matrix Q to obtain the position encoding attention score matrix;
[0026] The position encoding attention score matrix is added to the content attention score matrix, and normalized by the softmax function to obtain the normalized feature result; the normalized feature result is multiplied by the value vector matrix V to obtain a single-dimensional attention output feature, and the single-dimensional attention outputs of each independent subspace are spliced to generate a multi-dimensional parallel attention output feature, which is the output feature map of the improved backbone network.
[0027] Furthermore, the contrastive learning architecture calculates the contrastive loss value according to the contrastive loss function, which specifically includes:
[0028]
[0029] Among them, N represents the number of sample boxes predicted by the YOLOv8-MDPA-CL network model, L con is the contrast loss value; z i and z j Represents the feature vector of training sample i and training sample j, represents the true label of training sample i; τ represents the temperature coefficient of the contrast loss function, z k is the feature vector of training sample k.
[0030] Furthermore, the contrast loss function and the YOLOv8 loss function are used to calculate the training loss value between the training recognition result and the true label, including:
[0031] L YOLOv8 =αL cls +βL bbox +γL conf
[0032] L total =μ1L con +μ2L YOLOv8
[0033] The formula is, L YOLOv8 is the YOLOv8 loss value; L con is the contrast loss value; α, β and γ are the weight coefficients in the YOLOv8 loss function; μ1 and μ2 are the weight coefficients of the contrast loss value and the YOLOv8 loss value respectively; L total is the training loss value between the training recognition result and the true label; L cls is the classification loss value, L bbox is the bounding box regression loss, L conf is the confidence loss.
[0034] A second aspect of the present invention provides an underwater sonar image recognition system based on improved YOLOv8, comprising:
[0035] The monitoring unit is used to obtain the sonar monitoring image of the monitoring area collected by the underwater sonar equipment, and input the sonar monitoring image into the preset YOLOv8-MDPA-CL network model to obtain the underwater target detection result;
[0036] Model building unit, used to build the YOLOv8-MDPA-CL network model based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture;
[0037] An acquisition unit is used to acquire sonar training images of monitoring targets in various set underwater scenes, preprocess the sonar training images and add real labels to obtain training samples;
[0038] The training unit is used to train the YOLOv8-MDPA-CL network model using the training samples to obtain the training recognition results, calculate the training loss value between the training recognition results and the true labels using the contrast loss function and the YOLOv8 loss function, optimize the weight parameters of the YOLOv8-MDPA-CL network model according to the training loss value, and repeatedly iterate the training process of the YOLOv8-MDPA-CL network model until the iteration termination condition is reached to output the trained YOLOv8-MDPA-CL network model.
[0039] Furthermore, the model building unit builds a YOLOv8-MDPA-CL network model based on the YOLOv8 network model, the multi-dimensional parallel attention mechanism and the contrastive learning architecture, specifically including:
[0040] The backbone network based on the YOLOv8 network model and the multi-dimensional parallel attention mechanism are used to obtain an improved backbone network;
[0041] The improved backbone network includes an initial layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer; the initial layer is set as a CBL module, the first feature extraction layer, the second feature extraction layer and the third feature extraction layer include a CBL module and a C2f module connected in sequence; an MDPA module is constructed based on a multi-dimensional parallel attention mechanism, and the fourth feature extraction layer includes a CBL module, an SPPF module and an MDPA module connected in sequence;
[0042] The MDPA module is connected to a contrastive learning framework, which calculates a contrastive loss value according to a contrastive loss function, and the contrastive loss value is used to constrain feature extraction of the improved backbone network;
[0043] The YOLOv8-MDPA-CL network model is obtained based on the improved backbone network, contrastive learning architecture, neck network and head network of the YOLOv8 network model.
[0044] The third aspect of the present invention provides an electronic device, comprising a storage medium and a processor; the storage medium is used to store instructions; and the processor is used to operate according to the instructions to execute the underwater sonar image recognition method described in the first aspect.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] In the present invention, by introducing the multidimensional parallel attention mechanism (MDPA), the model's ability to capture global information is significantly improved, the receptive field is expanded, and the model is more adaptable to complex scenes in underwater sonar images; the multidimensional parallel attention mechanism (MDPA) can more accurately extract and integrate deep-level features in sonar images, thereby better identifying the detailed information of the target.
[0047] The present invention uses the contrast loss function and the YOLOv8 loss function to calculate the training loss value between the training recognition result and the true label, optimizes the weight parameters of the YOLOv8-MDPA-CL network model according to the training loss value, and uses the similarity of positive and negative sample pairs to constrain the extraction of backbone network features through the contrast learning architecture, so that the feature representations of similar objects are more similar, and the feature representations of different objects are more different, thereby enhancing the generalization ability and stability of the model in complex underwater environments. In the case of low-resolution images and more noise interference, the YOLOv8-MDPA-CL network model can still accurately identify the target, thereby improving the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flow chart of the underwater sonar image recognition method provided in this embodiment 1;
[0049] Figure 2 is a structural diagram of the improved backbone network provided in this embodiment 1;
[0050] Figure 3 is a structural diagram of the MDPA module provided in this embodiment 1;
[0051] Figure 4 is a structural diagram of the YOLOv8-MDPA-CL network model provided in this embodiment 1;
[0052] Figure 5 It is an iterative graph of the YOLOv8-MDPA-CL network model provided in this embodiment 1 trained on a public dataset. DETAILED DESCRIPTION
[0053] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0054] Example 1
[0055] like Figures 1 to 4 As shown, this implementation provides an underwater sonar image recognition method based on improved YOLOv8, including:
[0056] Obtain sonar monitoring images of the monitoring area collected by underwater sonar equipment, including:
[0057] The underwater sonar device is set as a multi-beam forward-looking sonar detection device, and the multi-beam forward-looking sonar detection device is controlled to release a 720kHz sonar signal and a 1200kHz sonar signal simultaneously in the monitoring area; the 720kHz sonar signal is used for target detection within a first distance range, and the 1200kHz sonar signal is used for target detection within a second distance range, and the maximum distance value within the first distance range is greater than the maximum distance value within the second distance range;
[0058] When performing sonar imaging of a target at multiple positions and angles, a forward-looking sonar annotation tool is used to annotate the same target in real time to obtain a sonar monitoring image; the sonar monitoring image is input into the preset YOLOv8-MDPA-CL network model to obtain underwater target detection results;
[0059] The YOLOv8-MDPA-CL network model construction and training process includes:
[0060] The YOLOv8-MDPA-CL network model is constructed based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture, including:
[0061] The backbone network based on the YOLOv8 network model and the multi-dimensional parallel attention mechanism are used to obtain an improved backbone network;
[0062] The improved backbone network includes an initial layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer; the initial layer is set as a CBL module, the first feature extraction layer, the second feature extraction layer and the third feature extraction layer include a CBL module and a C2f module connected in sequence; an MDPA module is constructed based on a multi-dimensional parallel attention mechanism, and the fourth feature extraction layer includes a CBL module, an SPPF module and an MDPA module connected in sequence;
[0063] First, a series of feature maps are obtained after processing by the CBL module, C2f module and SPPF module, and low-level features such as edges, textures and other basic information are extracted from the sonar image. Then, the MDPA module is used to enhance the feature extraction of the sonar image.
[0064] The MDPA module extracts the intermediate feature map output by the SPPF module to obtain the output feature map of the improved backbone network, which specifically includes:
[0065] The intermediate feature map is divided into multiple independent subspaces and converted into input vector X. The input vector X in each independent subspace is respectively mapped to the mapping matrix W Q , W K , W V Perform point-by-point convolution to generate query vector matrix Q, key vector matrix K and value vector matrix V;
[0066] Multiply the query vector matrix Q and the key vector matrix K to obtain the content attention score matrix; perform position encoding on the independent subspaces from the height and width dimensions and add them to obtain the position encoding matrix R; multiply the position encoding matrix R and the query vector matrix Q to obtain the position encoding attention score matrix;
[0067] The position encoding attention score matrix is added to the content attention score matrix, and normalized by the softmax function to obtain the normalized feature result; the normalized feature result is multiplied by the value vector matrix V to obtain a single-dimensional attention output feature, and the single-dimensional attention outputs of each independent subspace are spliced to generate a multi-dimensional parallel attention output feature, which is the output feature map of the improved backbone network.
[0068] The MDPA module is connected to a contrastive learning framework, which calculates a contrastive loss value according to a contrastive loss function, and the contrastive loss value is used to constrain feature extraction of the improved backbone network;
[0069] The YOLOv8-MDPA-CL network model is obtained based on the improved backbone network, contrastive learning architecture, neck network and head network of the YOLOv8 network model.
[0070] The neck network of the YOLOv8 network model includes Feature Pyramid Networks (FPN) and Path Aggregation Network (PAN). The Feature Pyramid Network consists of a C2f module and an upsample module (Upsample) and a concatenation module (Concat). It starts upsampling from the deep level, and the upsampled features of each layer are fused with the corresponding low-level features to build a top-down feature pyramid to achieve preliminary fusion of multi-scale features. The Path Aggregation Network consists of a CBL module and a Concat module. It starts downsampling from the bottom-level features, and passes the low-level features to the high-level features layer by layer to further enrich the multi-scale features. The Concat module in the Feature Pyramid Network FPN is connected to the C2f module in the head network, the Upsampl module is connected to the MDPA module in the backbone network to realize the transmission of feature maps, and the Concant in the Path Aggregation Network is connected to the C2f module in FPN and the MDPA module in the backbone network to realize the transmission of feature maps.
[0071] The head network receives feature maps output from the FPN in the neck network and the C2f module in the PAN. First, the head network predicts the category of each grid cell in the feature map and outputs the target category and its probability distribution that each grid may contain. Next, for each grid cell, the head network performs bounding box regression and calculates the position parameters and size parameters of the box. In addition, the head network also calculates a confidence score for each predicted box to indicate whether the box contains the target. A high confidence score indicates that the box contains the target, while a low confidence score indicates that the box does not contain the target. Through this series of steps, the head network is able to convert the fused feature map into a complete target detection result, providing accurate information for the final target recognition task.
[0072] The sonar training images of monitoring targets in various set underwater scenes are obtained, and the process of preprocessing the sonar training images includes:
[0073] The sonar training image is denoised using a Gaussian filtering algorithm, wherein the Gaussian kernel size is set to 3×3 or 5×5, and the standard deviation is set to 0;
[0074] The sonar training images were rotated by ±15°; the sonar training images were randomly flipped horizontally and vertically; and the sonar training images were randomly cropped and scaled to obtain training samples.
[0075] Add true labels to the training samples, and use the training samples to train the YOLOv8-MDPA-CL network model to obtain training recognition results.
[0076] The positive and negative samples are distinguished by the similarity and difference of the real labels, and then the similarity and difference between the training samples are compared, so that the feature representations of the same objects are more similar, and the feature representations of different objects are more different. The contrast loss function calculates the feature similarity of the positive sample pair (through cosine similarity) and minimizes it, while maximizing the feature difference of the negative sample pair.
[0077] The contrastive learning architecture calculates the contrastive loss value according to the contrastive loss function, which includes:
[0078]
[0079] Among them, N represents the number of sample boxes predicted by the YOLOv8-MDPA-CL network model, L con is the contrast loss value; z i and z j Represents the feature vector of training sample i and training sample j, represents the true label of training sample i; z kis the feature vector of training sample k; τ represents the temperature coefficient of the contrast loss function, which is used to adjust the scale of similarity and will affect the similarity calculation between each sample in the contrast loss function.
[0080] The contrast loss function and the YOLOv8 loss function are used to calculate the training loss value between the training recognition result and the true label, including:
[0081] L YOLOv8 =αL cls +βL bbox +γL conf
[0082] L total =μ1L con +μ2L YOLOv8
[0083] The formula is, L YOLOv8 is the YOLOv8 loss value; L con is the contrast loss value; α, β and γ are the weight coefficients in the YOLOv8 loss function; μ1 and μ2 are the weight coefficients of the contrast loss value and the YOLOv8 loss value respectively; L total is the training loss value between the training recognition result and the true label; L cls is the classification loss value, L bbox is the bounding box regression loss, L conf is the confidence loss.
[0084] The weight parameters of the YOLOv8-MDPA-CL network model are optimized according to the training loss value, and the training process of the YOLOv8-MDPA-CL network model is repeatedly iterated until the iteration termination condition is reached and the trained YOLOv8-MDPA-CL network model is output. In this embodiment, the iteration termination condition is: the accuracy of the YOLOv8-MDPA-CL network model tested using the validation set reaches a set threshold A or the mean average precision (mAP) reaches a set threshold B.
[0085] In this embodiment, by introducing the multi-dimensional parallel attention mechanism (MDPA), the model's ability to capture global information is significantly improved, the receptive field is expanded, and it is more adaptable to complex scenes in underwater sonar images; the multi-dimensional parallel attention mechanism (MDPA) can more accurately extract and integrate deep-level features in sonar images, thereby better identifying detailed information of the target.
[0086] This embodiment uses the similarity of positive and negative sample pairs to constrain the extraction of backbone network features through a contrastive learning architecture, making the feature representations of similar objects more similar and the feature representations of different objects more different, thereby enhancing the generalization ability and stability of the model in complex underwater environments. In the case of low-resolution images and more noise interference, the YOLOv8-MDPA-CL network model can still accurately identify targets, improving the robustness of the model.
[0087] In this embodiment, training is performed on the public dataset UATD-2022 dataset. A stochastic gradient descent (SGD) optimizer with specific settings is used. The momentum of the SGD optimizer is set to 0.937, the weight decay is set to 0.0005, and the batch size is set to 32. The number of iterations is set to 150 cycles. The warm-up momentum is set to 0.8, and the initial learning rate is set to 0.01. In the first three cycles of training, the learning rate of the YOLOv8-MDPA-CL network model gradually increases from the initial learning rate. After three cycles, the learning rate becomes 0.01 and gradually decreases in each cycle to continue training the YOLOv8-MDPA-CL network model.
[0088] The YOLOv8-MDPA-CL network model, the YOLOv5n model, and the YOLOv8n model are compared, and the evaluation indicators are mean average precision (mAP), precision (P), recall (R), and F1 score. mAP is the average accuracy of all target categories under a fixed cross-merge ratio. mAP@0.5 represents the average accuracy of all target categories when the cross-merge ratio threshold is fixed to 0.5.
[0089] The formulas for AP and mAP are as follows:
[0090]
[0091] In the formula, r is the integral variable, p(r) is the precision and recall curve, AP is the area under the precision and recall curve, reflecting the average accuracy of a single category. n represents the number of target categories.
[0092] Precision and recall rate reflect the model's false positive rate Precision and missed positive rate Recall respectively. The formula is as follows:
[0093]
[0094] In the formula, TP refers to the number of positive samples correctly identified as positive samples, reflecting the number of targets accurately identified; FN refers to the number of positive samples misclassified as negative samples, reflecting the number of missed targets; FP refers to the number of negative samples misclassified as positive samples, reflecting the number of targets incorrectly identified.
[0095] The F1 value balances the accuracy and recall of the classification model, representing a harmonic average between precision and recall. The calculation formula of the F1 value is as follows:
[0096]
[0097] In addition, the number of parameters (Parameters), model size (Model Size) and floating-point operations (GFLOPs) are other indicators for evaluating models. Specifically, model size refers to the space occupied by the model on the storage device, while parameters refer to the number of adjustable parameters learned in the neural network model. FLOPs and parameters are usually used to evaluate the complexity of the model.
[0098] Table 1. Comparison of the performance of this embodiment and other target detection algorithms
[0099]
[0100] As can be seen from Table 1, YOLOv8-MDPA-CL performs well in most key indicators. In Test1, YOLOv8-MDPA-CL leads other models in precision (88.0%), F1 score (84.9%), and mAP@0.5 (84.5%), and the number of parameters (2.82M) and model size (5.8MB) are also the smallest, showing a good model compression effect. In Test2, although the recall rate is slightly lower than YOLOv8_mobilenetv3 (81.4% vs 83.4%), it is still the best in precision (88.6%) and mAP@0.5 (84.7%), and the overall performance is robust.
[0101] Compared with YOLOv8n and YOLOv8_mobilenetv3, YOLOv8-MDPA-CL has advantages in performance, number of parameters, computational efficiency, and model size. Although YOLOv8_mobilenetv3 performs better in recall, it has lower mAP and precision, so YOLOv8-MDPA-CL is the best choice, especially for underwater applications that need to balance accuracy, efficiency, and resource consumption.
[0102] from Figure 5 As shown in the figure, on the UATD-2022 dataset, the YOLOv8-MDPA-CL model can achieve the best performance after about 80 iterations, and after reaching the optimal state, the training results of the model remain stable without overfitting or performance degradation. Therefore, it can be concluded that YOLOv8-MDPA-SL can quickly find the optimal solution and maintain good stability during training.
[0103] Example 2
[0104] This embodiment discloses an underwater sonar image recognition system based on improved YOLOv8, the underwater sonar image recognition system is used to execute the underwater sonar image recognition method described in Example 1, and the underwater sonar image recognition system includes:
[0105] The monitoring unit is used to obtain the sonar monitoring image of the monitoring area collected by the underwater sonar equipment, and input the sonar monitoring image into the preset YOLOv8-MDPA-CL network model to obtain the underwater target detection result;
[0106] Model building unit, used to build the YOLOv8-MDPA-CL network model based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture;
[0107] An acquisition unit is used to acquire sonar training images of monitoring targets in various set underwater scenes, preprocess the sonar training images and add real labels to obtain training samples;
[0108] The training unit is used to train the YOLOv8-MDPA-CL network model using the training samples to obtain the training recognition results, calculate the training loss value between the training recognition results and the true labels using the contrast loss function and the YOLOv8 loss function, optimize the weight parameters of the YOLOv8-MDPA-CL network model according to the training loss value, and repeatedly iterate the training process of the YOLOv8-MDPA-CL network model until the iteration termination condition is reached to output the trained YOLOv8-MDPA-CL network model.
[0109] The model building unit builds a YOLOv8-MDPA-CL network model based on the YOLOv8 network model, the multi-dimensional parallel attention mechanism and the contrastive learning architecture, specifically including:
[0110] The backbone network based on the YOLOv8 network model and the multi-dimensional parallel attention mechanism are used to obtain an improved backbone network;
[0111] The improved backbone network includes an initial layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer; the initial layer is set as a CBL module, the first feature extraction layer, the second feature extraction layer and the third feature extraction layer include a CBL module and a C2f module connected in sequence; an MDPA module is constructed based on a multi-dimensional parallel attention mechanism, and the fourth feature extraction layer includes a CBL module, an SPPF module and an MDPA module connected in sequence;
[0112] The MDPA module is connected to a contrastive learning framework, which calculates a contrastive loss value according to a contrastive loss function, and the contrastive loss value is used to constrain feature extraction of the improved backbone network;
[0113] The YOLOv8-MDPA-CL network model is obtained based on the improved backbone network, contrastive learning architecture, neck network and head network of the YOLOv8 network model.
[0114] Example 3
[0115] This embodiment discloses an electronic device, including a storage medium and a processor; the storage medium is used to store instructions; it is characterized in that the processor is used to operate according to the instructions to execute the underwater sonar image recognition method described in Example 1.
[0116] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0117] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0118] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An underwater sonar image recognition method based on improved YOLOv8, characterized in that: include: Obtain the sonar monitoring image of the monitoring area collected by the underwater sonar equipment, and input the sonar monitoring image into the preset YOLOv8-MDPA-CL network model to obtain the underwater target detection result; The YOLOv8-MDPA-CL network model construction and training process includes: Construct the YOLOv8-MDPA-CL network model based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture; Sonar training images of monitoring targets in various set underwater scenes are obtained, and the sonar training images are preprocessed and real labels are added to obtain training samples. The YOLOv8-MDPA-CL network model is trained with the training samples to obtain training recognition results. The training loss value between the training recognition results and the real labels is calculated using the contrast loss function and the YOLOv8 loss function. The weight parameters of the YOLOv8-MDPA-CL network model are optimized according to the training loss value. The training process of the YOLOv8-MDPA-CL network model is iterated repeatedly until the iteration termination condition is reached and the trained YOLOv8-MDPA-CL network model is output.
2. The underwater sonar image recognition method according to claim 1, characterized in that: Obtain sonar monitoring images of the monitoring area collected by underwater sonar equipment, including: The underwater sonar device is set as a multi-beam forward-looking sonar detection device, and the multi-beam forward-looking sonar detection device is controlled to release a 720kHz sonar signal and a 1200kHz sonar signal simultaneously in the monitoring area; the 720kHz sonar signal is used for target detection within a first distance range, and the 1200kHz sonar signal is used for target detection within a second distance range, and the maximum distance value within the first distance range is greater than the maximum distance value within the second distance range; When performing sonar imaging of a target at multiple positions and angles, a forward-looking sonar annotation tool is used to annotate the same target in real time to obtain a sonar monitoring image.
3. The underwater sonar image recognition method according to claim 1, characterized in that: The process of preprocessing sonar training images includes: The sonar training image is denoised using a Gaussian filtering algorithm, wherein the Gaussian kernel size is set to 3×3 or 5×5, and the standard deviation is set to 0; The sonar training images were rotated by ±15°; the sonar training images were randomly flipped horizontally and vertically; and the sonar training images were randomly cropped and scaled to obtain training samples.
4. The underwater sonar image recognition method according to claim 1, characterized in that: The YOLOv8-MDPA-CL network model is constructed based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture, including: The backbone network based on the YOLOv8 network model and the multi-dimensional parallel attention mechanism are used to obtain an improved backbone network; The improved backbone network includes an initial layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer; the initial layer is set as a CBL module, the first feature extraction layer, the second feature extraction layer and the third feature extraction layer include a CBL module and a C2f module connected in sequence; an MDPA module is constructed based on a multi-dimensional parallel attention mechanism, and the fourth feature extraction layer includes a CBL module, an SPPF module and an MDPA module connected in sequence; The MDPA module is connected to a contrastive learning framework, which calculates a contrastive loss value according to a contrastive loss function, and the contrastive loss value is used to constrain feature extraction of the improved backbone network; The YOLOv8-MDPA-CL network model is obtained based on the improved backbone network, contrastive learning architecture, neck network and head network of the YOLOv8 network model.
5. The underwater sonar image recognition method according to claim 4, characterized in that: The MDPA module extracts the intermediate feature map output by the SPPF module to obtain the output feature map of the improved backbone network, which specifically includes: The intermediate feature map is divided into multiple independent subspaces and converted into input vector X. The input vector X in each independent subspace is respectively mapped to the mapping matrix W Q , W K , W V Perform point-by-point convolution to generate query vector matrix Q, key vector matrix K and value vector matrix V; Multiply the query vector matrix Q and the key vector matrix K to obtain the content attention score matrix; perform position encoding on the independent subspaces from the height and width dimensions and add them to obtain the position encoding matrix R; multiply the position encoding matrix R and the query vector matrix Q to obtain the position encoding attention score matrix; The position encoding attention score matrix is added to the content attention score matrix, and normalized by the softmax function to obtain the normalized feature result; the normalized feature result is multiplied by the value vector matrix V to obtain a single-dimensional attention output feature, and the single-dimensional attention outputs of each independent subspace are spliced to generate a multi-dimensional parallel attention output feature, which is the output feature map of the improved backbone network.
6. The underwater sonar image recognition method according to claim 1, characterized in that: The contrastive learning architecture calculates the contrastive loss value according to the contrastive loss function, which includes: Among them, N represents the number of sample boxes predicted by the YOLOv8-MDPA-CL network model, L con is the contrast loss value; z i and z j Represents the feature vector of training sample i and training sample j, represents the true label of training sample i; τ represents the temperature coefficient of the contrast loss function, z k is the feature vector of training sample k.
7. The underwater sonar image recognition method according to claim 6, characterized in that: The contrast loss function and the YOLOv8 loss function are used to calculate the training loss value between the training recognition result and the true label, including: L YOLOv8 =αL cls +βL bbox +γL conf L total =μ1L con +μ2L YOLOv8 The formula is, L YOLOv8 is the YOLOv8 loss value; L con is the contrast loss value; α, β and γ are the weight coefficients in the YOLOv8 loss function; μ1 and μ2 are the weight coefficients of the contrast loss value and the YOLOv8 loss value respectively; L total is the training loss value between the training recognition result and the true label, L cls is the classification loss value, L bbox is the bounding box regression loss, L conf is the confidence loss.
8. An underwater sonar image recognition system based on improved YOLOv8, characterized in that: include: The monitoring unit is used to obtain the sonar monitoring image of the monitoring area collected by the underwater sonar equipment, and input the sonar monitoring image into the preset YOLOv8-MDPA-CL network model to obtain the underwater target detection result; Model building unit, used to build the YOLOv8-MDPA-CL network model based on the YOLOv8 network model, multi-dimensional parallel attention mechanism and contrastive learning architecture; An acquisition unit is used to acquire sonar training images of monitoring targets in various set underwater scenes, preprocess the sonar training images and add real labels to obtain training samples; The training unit is used to train the YOLOv8-MDPA-CL network model using the training samples to obtain the training recognition results, calculate the training loss value between the training recognition results and the true labels using the contrast loss function and the YOLOv8 loss function, optimize the weight parameters of the YOLOv8-MDPA-CL network model according to the training loss value, and repeatedly iterate the training process of the YOLOv8-MDPA-CL network model until the iteration termination condition is reached to output the trained YOLOv8-MDPA-CL network model.
9. The underwater sonar image recognition system according to claim 8, characterized in that: The model building unit builds a YOLOv8-MDPA-CL network model based on the YOLOv8 network model, the multi-dimensional parallel attention mechanism and the contrastive learning architecture, specifically including: The backbone network based on the YOLOv8 network model and the multi-dimensional parallel attention mechanism are used to obtain an improved backbone network; The improved backbone network includes an initial layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a fourth feature extraction layer; the initial layer is set as a CBL module, the first feature extraction layer, the second feature extraction layer and the third feature extraction layer include a CBL module and a C2f module connected in sequence; an MDPA module is constructed based on a multi-dimensional parallel attention mechanism, and the fourth feature extraction layer includes a CBL module, an SPPF module and an MDPA module connected in sequence; The MDPA module is connected to a contrastive learning framework, which calculates a contrastive loss value according to a contrastive loss function, and the contrastive loss value is used to constrain feature extraction of the improved backbone network; The YOLOv8-MDPA-CL network model is obtained based on the improved backbone network, contrastive learning architecture, neck network and head network of the YOLOv8 network model.
10. An electronic device comprising a storage medium and a processor; the storage medium is used to store instructions; characterized in that: The processor is used to operate according to the instructions to execute the underwater sonar image recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Underwater sonar target detection system and method
CN114693982A
Underwater sonar image target detection method and system based on neural network
CN118628898A
Combined self-supervision target identification method suitable for underwater
CN118865082A
Underwater target detection method and device based on comparative learning, and medium
CN119206463A
Training method of underwater sea urchin image recognition model, and underwater sea urchin image recognition method and device
US20240020966A1
Cited By
Underwater building defect detection method based on adaptive deep learning model training
CN120823491A
Ground time sequence observation image wake cloud identification method and system
CN120913129A
Underwater side-scan sonar image recognition method and system based on double attention mechanism
CN122116109A
A method and system for underwater side-scan sonar image recognition based on dual attention mechanism
CN122116109B