A method for underwater target detection based on deep learning

By using the feature extraction module of SPDConv and GAM in underwater target detection, combined with the NWD-CIoU loss function and PRHead detection head, the problems of low underwater target detection accuracy, small target and fuzzy target detection difficulties in the prior art are solved, and higher detection accuracy and robustness are achieved.

CN118196396BActive Publication Date: 2025-05-09JINAN LETONG ENVIRONMENTAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410467026.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-18
Publication Date
2025-05-09
Estimated Expiration
2044-04-18

AI Technical Summary

Technical Problem

The existing underwater target detection technology has problems such as low detection accuracy, difficulty in detecting small targets and difficulty in detecting fuzzy targets.

Method used

A deep learning-based method is adopted to build a feature extraction module through non-strike convolution (SPDConv) and global attention mechanism (GAM), and a normalized Gaussastan distance (NWD) and CIoU design a positioning regression loss function, and a dynamic object detection head (PRHead) is designed to enhance the detection capability of the model.

Benefits of technology

The accuracy and robustness of underwater target detection are improved, especially when detecting small targets and fuzzy targets, the model's perception of spatial information and different scale features is enhanced, and the underwater target detection task has the requirements of lightweight and high real-time model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118196396B_ABST
    Figure CN118196396B_ABST
Patent Text Reader

Abstract

The present invention provides an underwater target detection method based on deep learning, comprising the following steps: collecting underwater target detection data sets, preprocessing the data sets, screening qualified image data, balancing the number of underwater target samples of each category, and annotating the constructed data sets; the present invention combines the attention mechanism and the PReLU activation function to construct a new target detection head, called PRHead. This PRHead has dynamic characteristics and can better adapt to the needs of different target detection tasks. The use of the attention mechanism can make the model focus more on important target areas, while the PReLU activation function can improve the nonlinear fitting ability of the model, further enhancing the expression ability of PRHead. The underwater target detection model trained by the method proposed by the present invention has higher detection accuracy, fewer model parameters and less calculation, and meets the requirements of underwater target detection tasks for lightweight models and high real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of methods and systems of computing models, and in particular to an underwater target detection method based on deep learning. Background Art

[0002] Chinese patent document number CN202410138335.3, published on March 8, 2024, proposed an underwater target detection model based on attention and multi-scale feature fusion [5]. The model is divided into four modules, namely, preprocessing module, feature extraction module based on focus self-attention, multi-scale feature fusion module and underwater target positioning module. First, the semantic feature representation of the image from low level to high level is extracted, and the importance distribution of the target in the image data is automatically learned through the attention mechanism. By achieving higher attention to the target area, the key texture, shape, color and other information in the image are finally successfully captured, providing a basis for subsequent target detection. The multi-scale feature information is fused through the feature fusion module to obtain a multi-dimensional data description of the target feature, thereby improving the detection rate and positioning accuracy of the target. After the above key steps, the model can complete the task of underwater target detection and positioning based on multi-scale fusion features.

[0003] The current underwater target detection technology has the following main shortcomings:

[0004] (1) Low model detection accuracy: The existing target detection algorithm has relatively low positioning accuracy for targets in underwater target images. Due to the complexity of the underwater environment and its significant differences from the terrestrial environment, the existing official model detection is directly used, and its single grid prediction box is difficult to accurately capture the position of the target, resulting in large positioning errors.

[0005] (2) Difficulty in detecting small targets: Underwater images involve a variety of small-sized targets, such as fish, sea urchins, scallops, etc., which poses a challenge to the detection accuracy of existing target detection algorithms. Existing target detection algorithms have poor performance in detecting small targets and are prone to missed detection and false detection.

[0006] (3) Difficulty in detecting blurred targets: Underwater images usually have low contrast and blurred details, which results in blurred features of some underwater targets. Existing technologies are difficult to capture the features of such targets, resulting in serious underdetection of blurred targets by the model. Summary of the invention

[0007] In view of this, an embodiment of the present invention hopes to provide an underwater target detection method based on deep learning to solve or alleviate the technical problems existing in the prior art, and at least provide a beneficial option.

[0008] The technical solution of the embodiment of the present invention is implemented as follows: A method for underwater target detection based on deep learning comprises the following steps:

[0009] S1. Collect underwater target detection datasets, preprocess the datasets, screen qualified image data, balance the number of underwater target samples in each category, and annotate the constructed datasets;

[0010] S2. Segment and concatenate the input feature map to obtain a feature map of the target size, divide the feature map into two parts along the channel dimension, and perform subsequent processing;

[0011] S3. Combining CIoU and NWD, design the corresponding NWD-CIoU loss function for positioning regression loss in the model training process;

[0012] S4. Dynamically adjust the weights of features of different scales in the feature map to enhance the model's perception of spatial information. At the same time, adjust the attention weights of different positions;

[0013] S5. Based on the constructed underwater target detection model, train and update the parameters of each layer. First, initialize the parameters and set the hyperparameters required in the training process; then, divide the processed underwater target dataset into training set, validation set and test set; then, train the model and adjust the model parameters according to the training loss and validation loss until the loss value converges;

[0014] S6. Use the trained model to detect new underwater target images.

[0015] Preferably, labelme annotation tool is used as annotation software in S1.

[0016] Preferably, the logic of converting the image data into the feature map F6 in step S2 is:

[0017] SCGConv is used to perform an SPDConv convolution on the input feature map F1 to obtain a new feature map F2, and then F2 is divided into two parts along the channel dimension, one part is the unprocessed feature map F4, and the other part F3 is used as the input feature map of GAM_Bottleneck. In GAM_Bottleneck, the feature map is processed by the GAM attention mechanism after two layers of CBS layers, and the output of GAM is concatenated with the original input of GAM_Bottleneck as the output of GAM_Bottleneck;

[0018] In SCGConv, the output of each GAM_Bottleneck is used as the input of the next GAM_Bottleneck, which is repeated n times in sequence. The output of each GAM_Bottleneck is concatenated with the unprocessed feature maps F3 and F4 to obtain the feature map F5. After F5 passes through a layer of SPDConv, the final output feature map F6 of SCGConv is obtained.

[0019] Preferably, in said S3, the following steps are also included:

[0020] The formulas of CIoU and NWD are shown in (1) and (2) respectively.

[0021]

[0022] In the formula, IoU represents the intersection over union of the predicted box and the real box, b and b gt are the center points of the predicted box and the real box respectively, ρ(b,b gt ) is the Euclidean distance between the two center points, c w and c h Respectively represent the width and height of the minimum bounding box composed of the predicted box and the true box, w gt and h gt Represents the width and height of the real box, w and h represent the width and height of the predicted box. Z is the number of data set categories, is a second-order Wasserstein distance definition, where N a and N b Respectively represent the bounding box A=(cx a ,cy a ,w a ,h a ) and bounding box B = (cx b ,cy b ,w b ,h b ), ||·|| F is the Frobenius norm;

[0023] CIoU and NWD are combined as a new positioning regression loss function NWD-CIoU. The calculation formula of NWD-CIoU is shown in formula (4);

[0024] L NWD-CIoU =α·L CIoU +(1-α)·L NWD (4)

[0025] Where α is the weight ratio of CIoU.

[0026] Preferably, the step S4 further includes the following steps:

[0027] Simultaneously implement scale-aware attention π L , spatial perception attention π S and task-aware attention π C Unified dynamic detection; for the input underwater target feature map F, PRHead inputs its scale-aware attention module π L Get the output feature map F1, and then input the feature map F1 into the spatial perception attention π S Get the output feature map F2, and finally input the feature map F2 into the task perception attention π C The final output feature map F3 of PRHead is obtained.

[0028] Given the underwater target detection layer 3D feature tensor F∈R L×S×C , where L represents the level of the feature map, S represents the product of the width and height of the feature map, and C represents the number of channels of the feature map. The attention calculation formula of PRHead is shown in formula (5).

[0029] W(F)=π C (π S (π L (F)·F)·F)·F (5)

[0030] In the formula, π C (·),π S (·),π L (·) are task-aware attention function, space-aware attention function and scale-aware attention function, which act on dimensions C, S and L respectively.

[0031] Preferably, scale-aware attention improves the recognition ability of underwater targets of different scales by dynamically adjusting the weights of features of different scales in the feature map. The calculation process is shown in formula (6).

[0032]

[0033] σ(x)=max(0,min(1,(x+1) / 2)) (7)

[0034] Where f(·) represents a linear function approximated by 1×1 convolution, σ(·) represents the Hard-sigmoid activation function, and for the input underwater target feature map F, the task-aware attention module π L It is subjected to global pooling, 1×1 convolution, PReLU activation and Hard Sigmoid activation in turn to obtain π L The output feature map F1, after being processed by the scale-aware attention module, the features become more sensitive to underwater targets of different scales.

[0035] Preferably, the calculation process of spatial perception attention is as follows:

[0036]

[0037] Where K represents the number of sparse sampling locations, p j +Δp j The spatial offset representing the self-learning is concentrated in the moving position of the unique region, which is used to focus on some discriminative positions, Δm j Indicates self-learning at position p j The important scalar at π is obtained by learning the intermediate level of the underwater target feature map F. For the input underwater target feature map F1, the spatial perception attention π S After re-arranging the index, perform a 3×3 convolution, and then perform offset shift and sigmoid activation operations on the convolved feature map, respectively. The output feature maps obtained by the above two operations are concatenated to obtain π S The output feature map F2.

[0038] Optionally, the task-aware attention is calculated by dynamically adjusting the attention weights of different positions in the underwater target feature map as follows:

[0039] π C (F)·F=max(α 1 (F)·F C +β 1 (F),α 2 (F)·F C +β 2 (F)) (9)

[0040] In the formula, F C is the feature slice of the Cth channel, α 1 (F),β 1 (F),α 2 (F),β 2 (F) are all parameters that depend on the input F and are used to learn and control the activation threshold. Task-aware attention uses the above parameters to activate different channels differently to achieve attention operation. For the input underwater target feature map F2, task-aware attention π C After global pooling, full connection, PReLU activation, full connection, and regularization operations, we can get π C The output feature map F3 of the module, the underwater target feature map F3 is the final output feature map of PRHead.

[0041] Preferably, in step S5, all neural network parameters are first initialized, and hyperparameters related to the steel bar binding quality detection model are set. Then, the training set data is divided into multiple batches, and the data of each batch is input into the underwater target detection model for training to obtain the training loss value loss of the batch; after completing a round of training for all batches of data in the entire training set, the verification set is input into the underwater target detection model in batches to obtain the verification loss value batch_loss of the corresponding batch; during the training and verification process, the underwater target detection model will automatically learn and adjust parameters according to the loss and batch_loss conditions each time; repeat multiple rounds of training until the batch_loss value converges, and the training of the underwater target detection model is completed.

[0042] The embodiment of the present invention has the following advantages due to the adoption of the above technical solution:

[0043] The present invention uses non-strided convolution SPDConv to replace traditional convolution, and combines the global attention mechanism (GAM) to construct a new feature extraction module, enhance global context information, and improve the ability of the model backbone network to extract fuzzy targets and small target features; by using the normalized Gauss-Wasserstein distance (NWD) and CIoU to construct a new positioning regression loss function, the positioning accuracy of small targets in complex underwater environments is improved; by combining the attention mechanism and the PReLU activation function, we have constructed a new target detection head, called PRHead. This PRHead has dynamic characteristics and can better adapt to the needs of different target detection tasks. The use of the attention mechanism can make the model focus more on important target areas, while the PReLU activation function can improve the nonlinear fitting ability of the model, further enhancing the expression ability of PRHead. The underwater target detection model trained by the method proposed by the present invention has higher detection accuracy, fewer model parameters and less calculation, and meets the requirements of underwater target detection tasks for lightweight models and high real-time performance.

[0044] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0046] Figure 1 It is a flow chart of the steps of the present invention.

[0047] Figure 2 This is the structural diagram of the SCG-NCPR algorithm of the present invention.

[0048] Figure 3 This is the SCGConv structure diagram of the present invention.

[0049] Figure 4 This is a structural diagram of the PRHead of the present invention. DETAILED DESCRIPTION

[0050] In the following, only some exemplary embodiments are briefly described. As those skilled in the art will appreciate, the described embodiments may be modified in various ways without departing from the spirit or scope of the present invention. Therefore, the drawings and descriptions are considered to be exemplary and non-restrictive in nature.

[0051] It should be noted that the terms "first", "second", "symmetrical", "array", etc. are only used to distinguish between description and position description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the definition of "first", "symmetrical", etc. can explicitly or implicitly include one or more of these features; similarly, when the quantity of certain features is not limited in the form of words such as "two" or "three", it should be noted that this feature also explicitly or implicitly includes one or more feature quantities;

[0052] In the present invention, unless otherwise clearly specified and limited, the terms such as "installation", "connection", "fixation" and the like should be understood in a broad sense; for example, it can be a fixed connection, a detachable connection, or an integral molding; it can be a mechanical connection, a direct connection, welding, or an indirect connection through an intermediate medium, or it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood based on the description and drawings combined with specific circumstances.

[0053] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0054] like Figure 1 As shown, an embodiment of the present invention provides an underwater target detection method based on deep learning, comprising the following steps:

[0055] S1. Collect underwater target detection datasets, preprocess the datasets, screen qualified image data, balance the number of underwater target samples in each category, and annotate the constructed datasets;

[0056] S2. Segment and concatenate the input feature map to obtain a feature map of the target size, divide the feature map into two parts along the channel dimension, and perform subsequent processing;

[0057] S3. Combining CIoU and NWD, design the corresponding NWD-CIoU loss function for positioning regression loss in the model training process;

[0058] S4. Dynamically adjust the weights of features of different scales in the feature map to enhance the model's perception of spatial information. At the same time, adjust the attention weights of different positions;

[0059] S5. Based on the constructed underwater target detection model, train and update the parameters of each layer. First, initialize the parameters and set the hyperparameters required in the training process; then, divide the processed underwater target dataset into training set, validation set and test set; then, train the model and adjust the model parameters according to the training loss and validation loss until the loss value converges;

[0060] S6. Use the trained model to detect new underwater target images. In actual engineering applications, the collected underwater target image data is input into the underwater target detection model, which automatically detects the image to obtain the detection result of the underwater target image, and displays the detection result on the image, including the category of the underwater target, the location of the underwater target, and the detection confidence.

[0061] Example

[0062] This embodiment uses an underwater target image dataset, which contains 14,000 images, 74,903 labeled objects and 10 categories, for training an underwater target detection model.

[0063] Dataset annotation: Annotate the collected underwater target dataset, including target category and target location.

[0064] Dataset division: The dataset is randomly divided into training set, validation set and test set in a ratio of 6:2:2, that is, the training set has 8400 images, and the validation set and test set have 2800 images each.

[0065] Model training: The data in the training set is input into the underwater target detection model for training. The entire training process is set to 300 epochs, and the parameters in the model are adjusted according to the loss value in each iteration.

[0066] Training results: After adjusting the relevant parameters, the accuracy of the underwater target detection model finally stabilized, and the average accuracy on the test set fluctuated between 85% and 90%.

[0067] In this embodiment, labelme annotation tool is used as annotation software in S1.

[0068] In this embodiment, the logic of converting the image data into the feature map F6 in step S2 is:

[0069] SCGConv is used to perform an SPDConv convolution on the input feature map F1 to obtain a new feature map F2, and then F2 is divided into two parts along the channel dimension, one part is the unprocessed feature map F4, and the other part F3 is used as the input feature map of GAM_Bottleneck. In GAM_Bottleneck, the feature map is processed by the GAM attention mechanism after two layers of CBS layers, and the output of GAM is concatenated with the original input of GAM_Bottleneck as the output of GAM_Bottleneck;

[0070] In SCGConv, the output of each GAM_Bottleneck is used as the input of the next GAM_Bottleneck, which is repeated n times in sequence. The output of each GAM_Bottleneck is concatenated with the unprocessed feature maps F3 and F4 to obtain the feature map F5. After F5 passes through a layer of SPDConv, the final output feature map F6 of SCGConv is obtained.

[0071] More specifically, the underwater target feature map F1 with a size of C×H×W is passed through the SPDConv layer. For F1, SPDConv divides it into a series of feature sub-maps f with a size of C×H / 2×W / 2 x,y , and then all feature subgraphs f x,y The feature map X1 of size 4C×H / 2×W / 2 is obtained by splicing along the channel dimension. Finally, the non-strided convolution layer reduces the dimension of the input feature map X1 through 1×1 convolution to retain all the discriminant information as much as possible. Finally, SPDConv outputs the output feature map F2 of size C / 2×H / 2×W / 2, and splits the feature map F2 into feature map F3 and feature map F4 of size C / 4×H / 2×W / 2 along the channel dimension.

[0072] The feature map F3 is input into GAM_Bottleneck, and GAM_Bottleneck performs two CBS convolution operations on F3 with a convolution kernel of 3×3 and a step size of 1, and then inputs the GAM attention mechanism. Finally, the output feature map of the GAM attention mechanism is concatenated with F3 to obtain the output feature map of GAM_Bottleneck. In the SCGConv module, the output of each GAM_Bottleneck is used as the input of the next GAM_Bottleneck, and this is repeated n times in sequence;

[0073] Split the feature map F2 along the channel dimension into feature map F3 and feature map F4 with sizes of C / 4×H / 2×W / 2;

[0074] The feature map F3 is input into GAM_Bottleneck, and GAM_Bottleneck performs two CBS convolution operations on F3 with a convolution kernel of 3×3 and a step size of 1, and then inputs the GAM attention mechanism. Finally, the output feature map of the GAM attention mechanism is concatenated with F3 to obtain the output feature map of GAM_Bottleneck. In the SCGConv module, the output of each GAM_Bottleneck is used as the input of the next GAM_Bottleneck, and this is repeated n times in sequence;

[0075] Concatenate the output of each GAM_Bottleneck in the previous step with feature maps F3 and F4 to obtain feature map F5 of size (n+2)·C / 4×H / 2×W / 2;

[0076] After inputting the feature map F5 into SPDConv for non-stride convolution, the output feature map F6 of SCGConv is obtained, and the size of F6 is C / 2×H / 4×W / 4.

[0077] The SCGConv module obtains underwater target feature maps through this series of steps. These feature maps have rich detail features, while enhancing global context information and improving cross-dimensional channel spatial correlation. This makes the model pay more attention to the characteristics of the target itself and reduces dependence on redundant information, making it more suitable for underwater blurred target and small target detection tasks.

[0078] Example: Assume that the size of the underwater target feature map is C×H×W, where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map. Example: For an underwater target feature map F1 with an input size of 64×320×320, SCGConv first inputs it into the SPDConv layer. For F1, SPDConv divides it into a series of feature sub-maps f with a size of 64×160×160. x,y , and then all feature subgraphs f x,yThe feature map X1 with a size of 256×160×160 is obtained by splicing along the channel dimension. Finally, the non-strided convolution layer reduces the dimension of the input feature map X1 through a convolution with a convolution kernel of 1×1. Finally, SPDConv outputs an output feature map F2 with a size of 32×160×160. For F2, SCGConv divides it into feature maps F3 and F4 with a size of 16×160×160 along the channel dimension, and then inputs the feature map F3 into GAM_Bottleneck. GAM_Bottleneck performs two consecutive CBS convolution operations on F3 with a convolution kernel of 3×3 and a stride of 1. Then, the GAM attention mechanism is input, and the output feature map of the GAM attention mechanism is spliced ​​with F3 to obtain a GAM_Bottleneck output feature map with a size of 16×160×160. In the SCGConv module, the output of each GAM_Bottleneck is used as the input of the next GAM_Bottleneck, which is repeated n times in sequence. The outputs of the above n GAM_Bottlenecks are then concatenated with the feature maps F3 and F4 to obtain a feature map F5 of size (n+2)×16×160×160, where (n+2)×16 is the number of channels. Finally, the feature map F5 is input into SPDConv for non-stride convolution to obtain an output feature map F6 of size 32×80×80. F6 is the output feature map of SCGConv.

[0079] The present invention uses the SCG-NCPR algorithm, including a non-strided-global attention convolution module (SCGConv), by using the non-strided convolution SPDConv, and combining the global attention mechanism (GAM, Global AttentionMechanism) to construct a new feature extraction module, enhance the global context information, and improve the ability of the model backbone network to extract fuzzy targets and small target features; design the NWD-CIoU loss function, by using the normalized Gauss-Wasserstein distance (NWD) and CIoU to construct a new positioning regression loss function, improve the positioning accuracy of small targets in complex underwater environments; design a detection head (PRHead), by using a dynamic target detection head combined with an attention mechanism and a PReLU activation function, construct a new target detection head PRHead, and improve the model's processing ability for small underwater targets. The above invention effectively solves the problem of insufficient detection accuracy of underwater small targets and fuzzy targets by the current target detection algorithm.

[0080] In this embodiment, in S3, the following steps are also included:

[0081] The formulas of CIoU and NWD are shown in equations (1) and (2) respectively;

[0082]

[0083] In the formula, IoU represents the intersection over union of the predicted box and the real box, b and b gt are the center points of the predicted box and the real box respectively, ρ(b,b gt ) is the Euclidean distance between the two center points, c w and c h Respectively represent the width and height of the minimum bounding box composed of the predicted box and the true box, w gt and h gt Represents the width and height of the real box, w and h represent the width and height of the predicted box. Z is the number of data set categories, is a second-order Wasserstein distance definition, where N a and N b Respectively represent the bounding box A=(cx a ,cy a ,w a ,h a ) and bounding box B = (cx b ,cy b ,w b ,h b ), ‖·‖ F is the Frobenius norm;

[0084] CIoU and NWD are combined as a new positioning regression loss function NWD-CIoU. The calculation formula of NWD-CIoU is shown in formula (4);

[0085] L NWD-CIoU =α·L CIoU +(1-α)·L NWD (4)

[0086] In the formula, α is the weight ratio of CIoU. The best effect is achieved when α = 0.5.

[0087] NWD-CIoU can maintain scale invariance to a certain extent and is fairer when measuring the similarity between bounding boxes of tiny objects, enabling the model to have better robustness when detecting small underwater targets and blurred targets.

[0088] In this embodiment, the step S4 further includes the following steps:

[0089] Simultaneously implement scale-aware attention π L , spatial perception attention π S and task-aware attention π C Unified dynamic detection; for the input underwater target feature map F, PRHead inputs its scale-aware attention module π LGet the output feature map F1, and then input the feature map F1 into the spatial perception attention π S Get the output feature map F2, and finally input the feature map F2 into the task perception attention π C The final output feature map F3 of PRHead is obtained.

[0090] Given the underwater target detection layer 3D feature tensor F∈R L×S×C , where L represents the level of the feature map, S represents the product of the width and height of the feature map, and C represents the number of channels of the feature map. The attention calculation formula of PRHead is shown in formula (5).

[0091] W(F)=π C (π S (π L (F)·F)·F)·F (5)

[0092] In the formula, π C (·),π S (·),π L (·) are task-aware attention function, space-aware attention function and scale-aware attention function, which act on dimensions C, S and L respectively. PRHead combines the three attention mechanisms of scale-aware, space-aware and task-aware, which can make full use of different aspects of underwater target feature maps and improve the model's understanding and detection capabilities of underwater targets. PRHead combines multiple attention mechanisms and has strong flexibility and versatility. It can be applied to different types of underwater target detection tasks, and can dynamically adjust the representation of feature maps according to specific task requirements, improving the adaptability and generalization ability of the model.

[0093] In this embodiment, the scale-aware attention improves the recognition capability of underwater targets of different scales by dynamically adjusting the weights of features of different scales in the feature map. The calculation process is shown in formula (6).

[0094]

[0095] σ(x)=max(0,min(1,(x+1) / 2)) (7)

[0096] Where f(·) represents a linear function approximated by 1×1 convolution, σ(·) represents the Hard-sigmoid activation function, and for the input underwater target feature map F, the task-aware attention module π L It is subjected to global pooling, 1×1 convolution, PReLU activation and Hard Sigmoid activation in turn to obtain π LThe output feature map F1 of the scale-aware attention module becomes more sensitive to underwater targets of different scales after being processed by the scale-aware attention module. Through the operation of the scale-aware attention module, the feature map can be made more sensitive to underwater targets of different scales, thereby enhancing the model's ability to identify and detect targets at different scales, which helps to improve the robustness and adaptability of the model. Operations such as global pooling, 1×1 convolution, PReLU activation, and Hard Sigmoid activation in the module can help the model perceive and adjust information of different scales in the feature map, so that the model can better adapt to underwater target detection tasks at different scales, improving the generalization ability and adaptability of the model. Through the operation of the scale-aware attention module, the scale error that may be generated by the model when processing targets of different scales can be reduced, and the model's positioning and recognition accuracy of targets at different scales can be improved, which is conducive to improving the overall detection performance.

[0097] In this embodiment, the calculation process of spatial perception attention is as follows:

[0098]

[0099] Where K represents the number of sparse sampling locations, p j +Δp j The spatial offset representing the self-learning is concentrated in the moving position of the unique region, which is used to focus on some discriminative positions, Δm j Indicates self-learning at position p j The important scalar at π is obtained by learning the intermediate level of the underwater target feature map F. For the input underwater target feature map F1, the spatial perception attention π S After re-arranging the index, perform a 3×3 convolution, and then perform offset shift and sigmoid activation operations on the convolved feature map, respectively. The output feature maps obtained by the above two operations are concatenated to obtain π S The output feature map F2 of the underwater target feature map becomes sparser after being processed by the spatial perception attention module, focusing on foreground targets at different positions. Through the operation of the spatial perception attention module, attention can be focused on unique areas in the underwater target feature map, thereby focusing on some discriminative positions, improving the attention and discriminability of the target position, and helping to improve the accuracy of target detection and recognition. Through the processing of the spatial perception attention module, the feature map can be made sparser, that is, only focusing on some important positions and features, reducing the impact of redundant information, and improving the expression efficiency and discriminability of features.

[0100] In this embodiment, the calculation process of task-aware attention by dynamically adjusting the attention weights of different positions in the underwater target feature map is as follows:

[0101] πC (F)·F=max(α 1 (F)·F C +β 1 (F),α 2 (F)·F C +β 2 (F)) (9)

[0102] In the formula, F C is the feature slice of the Cth channel, α 1 (F), β 1 (F), α 2 (F), β 2 (F) are all parameters that depend on F and are used to learn and control the activation threshold. Task-aware attention uses the above parameters to activate different channels differently to achieve attention operation. For the input underwater target feature map F2, task-aware attention π C After global pooling, full connection, PReLU activation, full connection, and regularization operations, we can get π C The output feature map F3 of the module, the underwater target feature map F3 is the final output feature map of PRHead. After the feature map is processed by the task-aware attention module, the underwater target features will form different activations based on different downstream tasks. Through global pooling, full connection, PReLU activation, full connection, regularization and other operations, the task-aware attention module can effectively enhance and extract useful information in the underwater target feature map, so that the model can better understand and extract features, thereby improving the accuracy of detection and recognition. Through the processing of the task-aware attention module, the redundant information in the underwater target feature map can be eliminated, and more critical and effective features can be extracted, thereby reducing the learning burden of the model and improving the efficiency and performance of the model.

[0103] In this embodiment, in step S5, all neural network parameters are first initialized, and hyperparameters related to the steel bar binding quality detection model are set. Then, the training set data is divided into multiple batches, and the data of each batch is input into the underwater target detection model for training to obtain the training loss value loss of the batch; after completing a round of training for all batches of data in the entire training set, the verification set is input into the underwater target detection model according to the batch, and the verification loss value batch_loss of the corresponding batch is obtained; during the training and verification process, the underwater target detection model will automatically learn and adjust the parameters according to the loss and batch_loss conditions each time; repeat multiple rounds of training until the batch_loss value converges, and the training of the underwater target detection model is completed. This application enables the model to be iteratively optimized on different data through multiple rounds of training, enhances the generalization ability of the model, and improves the accuracy of detection. Moreover, the parameter adjustment and the setting of the end conditions during the training process are automated, reducing the need for manual intervention. Further, by dividing the data into multiple batches for training, computing resources can be effectively utilized, and the speed and efficiency of training can be improved. At the same time, verification of the validation set can avoid overfitting of the model and improve the generalization ability of the model.

[0104] For the input underwater target image, it first passes through two convolutional layers (Conv2D_BN_SiLU) to obtain the channel feature map of the underwater target image. Then, the channel feature map is input into the SCGConv module, which uses non-strided convolution, which can more effectively retain information compared to traditional convolution operations, and uses the GAM attention mechanism to enhance global context information, thereby enhancing cross-dimensional channel spatial correlation. Subsequently, the steel bar binding feature map passes through the Conv2D_BN_SiLU module and the SCGConv module in turn, and the above operation is repeated twice.

[0105] Next, the underwater target feature maps extracted by the three SCGConv modules are sequentially input into the Feature Pyramid Network (FPN). FPN performs feature fusion on underwater target features of different scales extracted by SCGConv, and then sends the fused features to the PRHead detection head for detection. PRHead can greatly enhance the model's expressiveness for targets, and is particularly suitable for small target detection, making the model more flexible and accurate. Finally, the model's detection results for underwater target images are obtained.

[0106] During model training, the NWD-CIoU loss function is used as the localization regression loss function of SCG-NCPR. The NWD-CIoU loss function can maintain scale invariance to a certain extent and is more fair when measuring the similarity between bounding boxes of tiny objects, making the model more robust when detecting small underwater targets and blurred targets.

[0107] The present invention improves the network's detection accuracy for underwater blurred targets and small targets by using a non-stride-global attention convolution module (SCGConv). In SCGConv, we use the GAM attention mechanism to design the GAM_Bottleneck structure, and combine it with the non-stride convolution SPDConv to improve the model's ability to extract detailed features of underwater targets. Such a design helps to reduce the model's dependence on redundant information generated during the iteration process, so that the model can more accurately detect underwater blurred targets and small targets. In addition, the present invention combines the normalized Gaussian Wasserstan distance (NWD) with CIoU to construct a new positioning regression loss function to more accurately measure the similarity between the model's predicted bounding box and the true target bounding box, thereby improving the model's positioning accuracy, especially when dealing with small targets in complex underwater environments.

[0108] By combining the attention mechanism and the PReLU activation function, we built a new object detection head, called PRHead. This PRHead has dynamic characteristics and can better adapt to the needs of different object detection tasks. The use of the attention mechanism can make the model focus more on important target areas, while the PReLU activation function can improve the nonlinear fitting ability of the model, further enhancing the expression ability of PRHead.

[0109] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for underwater target detection based on deep learning, characterized in that: The following steps are involved: S1. Collect underwater target detection datasets, preprocess the datasets, screen qualified image data, balance the number of underwater target samples in each category, and annotate the constructed datasets; S2. Segment and concatenate the input feature map to obtain a feature map of the target size, divide the feature map into two parts along the channel dimension, and perform subsequent processing; S3. Combining CIoU and NWD, design the corresponding NWD-CIoU loss function for positioning regression loss in the model training process; The formulas of CIoU and NWD are shown in equations (1) and (2) respectively: In the formula, IoU represents the intersection over union of the predicted box and the real box, b and b gt are the center points of the predicted box and the real box respectively, ρ(b,b gt ) is the Euclidean distance between the two center points, c w and c h Respectively represent the width and height of the minimum bounding box composed of the predicted box and the true box, w gt and h gt Represent the width and height of the real box, w and h represent the width and height of the predicted box, Z is the number of data set categories, W2 2 (N a ,N b ) is a second-order Wasserstein distance definition, where N a and N b Respectively represent the bounding box A=(cx a ,cy a ,w a ,h a ) and bounding box B = (cx b ,cy b ,w b ,h b ), ||·|| F is the Frobenius norm; CIoU and NWD are combined as a new positioning regression loss function NWD-CIoU. The calculation formula of NWD-CIoU is shown in formula (4); L NWD-CIoU =α·L CIoU +(1-a)·L NWD (4) In the formula, α is the weight ratio of CIoU; S4, dynamically adjust the weights of features of different scales in the feature map to enhance the model's perception of spatial information, and at the same time, adjust the attention weights of different positions; S5. Based on the constructed underwater target detection model, train and update the parameters of each layer. First, initialize the parameters and set the hyperparameters required in the training process; then, divide the processed underwater target data set into a training set, a validation set, and a test set; then, train the model and adjust the model parameters according to the training loss and the validation loss until the loss value converges; S6. Use the trained model to detect new underwater target images.

2. The underwater target detection method based on deep learning according to claim 1, characterized in that: In S1, labelme annotation tool is used as the annotation software.

3. The underwater target detection method based on deep learning according to claim 2, characterized in that: The logic of converting the image data into the feature map F6 in step S2 is: SCGConv is used to perform an SPDConv convolution on the input feature map F1 to obtain a new feature map F2, and then F2 is divided into two parts along the channel dimension, one part is the unprocessed feature map F4, and the other part F3 is used as the input feature map of GAM_Bottleneck. In GAM_Bottleneck, the feature map is processed by the GAM attention mechanism after two layers of CBS layers, and the output of GAM is concatenated with the original input of GAM_Bottleneck as the output of GAM_Bottleneck; In SCGConv, the output of each GAM_Bottleneck is used as the input of the next GAM_Bottleneck, which is repeated n times in sequence. The output of each GAM_Bottleneck is concatenated with the unprocessed feature maps F3 and F4 to obtain the feature map F5. After F5 passes through a layer of SPDConv, the final output feature map F6 of SCGConv is obtained.

4. The underwater target detection method based on deep learning according to claim 1, characterized in that: The step S4 also includes the following steps: Simultaneously implement scale-aware attention π L , spatial perception attention π S and task-aware attention π C Unified dynamic detection; for the input underwater target feature map F, PRHead inputs its scale-aware attention module π L Get the output feature map F1, and then input the feature map F1 into the spatial perception attention π S Get the output feature map F2, and finally input the feature map F2 into the task perception attention π C Get the final output feature map F3 of PRHead; Given the underwater target detection layer 3D feature tensor F∈R L×S×C , where L represents the level of the feature map, S represents the product of the width and height of the feature map, and C represents the number of channels of the feature map. The attention calculation formula of PRHead is shown in formula (5); W(F)=π C (p S (p L (F)·F)·F)·F (5) In the formula, π C (·),π S (·),π L (·) are task-aware attention function, space-aware attention function and scale-aware attention function, which act on dimensions C, S and L respectively.

5. The underwater target detection method based on deep learning according to claim 4, characterized in that: Scale-aware attention improves the recognition ability of underwater targets of different scales by dynamically adjusting the weights of features of different scales in the feature map. Its calculation process is shown in formula (6); σ(x)=max(0,min(1,(x+1) / 2)) (7) Where f(·) represents a linear function approximated by 1×1 convolution, σ(·) represents the Hard-sigmoid activation function, and for the input underwater target feature map F, the task-aware attention module π L It is subjected to global pooling, 1×1 convolution, PReLU activation and Hard-sigmoid activation in turn to obtain π L The output feature map F1, after being processed by the scale-aware attention module, the features become more sensitive to underwater targets of different scales.

6. The underwater target detection method based on deep learning according to claim 5, characterized in that: The calculation process of spatial perception attention is as follows: Where K represents the number of sparse sampling locations, p j +Δp j The spatial offset representing the self-learning is concentrated in the moving position of the unique region, which is used to focus on some discriminative positions, Δm j Indicates self-learning at position p j The important scalar at π is obtained through the intermediate level learning of the underwater target feature map F; for the input underwater target feature map F1, the spatial perception attention π S After re-arranging the index, perform a 3×3 convolution, and then perform offset shift and sigmoid activation operations on the convolved feature map, respectively. The output feature maps obtained by the above two operations are concatenated to obtain π S The output feature map F2.

7. The underwater target detection method based on deep learning according to claim 6, characterized in that: The calculation process of task-aware attention by dynamically adjusting the attention weights of different positions in the underwater target feature map is: p C (F)·F=max(α 1 (F)·F C +b 1 (F),a 2 (F)·F C +b 2 (F)) (9) In the formula, F C is the feature slice of the Cth channel, α 1 (F), β 1 (F), α 2 (F), β 2 (F) are all parameters that depend on the input F. They are used to learn and control the activation threshold. Task-aware attention uses the above parameters to activate different channels differently to achieve attention operation. For the input underwater target feature map F2, task-aware attention π C After global pooling, full connection, PReLU activation, full connection, and regularization operations, we can get π C The output feature map F3 of the module, the underwater target feature map F3 is the final output feature map of PRHead.

8. The underwater target detection method based on deep learning according to claim 1, characterized in that: In step S5, all neural network parameters are first initialized, and hyperparameters related to the steel bar binding quality detection model are set; then, the training set data is divided into multiple batches, and the data of each batch is input into the underwater target detection model for training to obtain the training loss value loss of the batch; after completing a round of training for all batches of data in the entire training set, the verification set is input into the underwater target detection model according to batches to obtain the verification loss value batch_loss of the corresponding batch; during the training and verification process, the underwater target detection model will automatically learn and adjust parameters according to the loss and batch_loss conditions each time; repeat multiple rounds of training until the batch_loss value converges, and the training of the underwater target detection model is completed.

Citation Information

Patent Citations

  • Underwater target detection model and method based on attention and multi-scale feature fusion

    CN117671473B