An underwater target detection method based on attention fusion
By adopting an attention fusion-based method in underwater target detection, the Vision-Transformer and PAFPN path-enhanced feature pyramid module extracts self-attention and spatial attention information, and combined with the channel attention information of the SE module, the problem of insufficient feature fusion and attention information extraction in the prior art is solved, and more efficient underwater target detection is achieved.
Patent Information
- Application Number
- CN202210410629.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing underwater object detection method based on convolutional neural networks has shortcomings in feature fusion and attention information extraction, and it is impossible to effectively utilize self-attention and spatial attention information.
A supervised learning underwater object detection method based on attention fusion is adopted, and self-attention information is extracted through Vision-Transformer, the PAFPN path enhancement feature pyramid module extracts spatial attention information, and the SE module extracts attention information between channels and performs cascade fusion to improve detection accuracy.
This method can effectively extract spatial attention information, self-attention information and channel attention information, make up for the shortcomings of traditional methods, and improve the accuracy and performance of underwater target detection.
Smart Images

Figure CN114782798B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to the fields of computer vision and image pattern recognition technology. Background Art
[0002] Object detection technology is a very popular basic research direction in the field of computer vision at present. This technology can accurately give the category and location of the objects of interest in an image / video. As a basic research direction in the field of computer vision, there are many important applications based on object detection technology, such as defect detection on industrial production lines, fishing in underwater fish farms, character recognition, etc.
[0003] The invention patent with the application number 202111127297.4 (Patent Center of Dalian University of Technology) discloses a lightweight underwater object detection method based on feature fusion and neural network search. This invention fuses the features of onshore and underwater detection networks, that is, an addition operation, to achieve the goal of using onshore prior knowledge to guide the construction of the underwater network structure. At the same time, using the neural network search algorithm, an efficient search space is designed, a differentiable search strategy based on gradients is adopted, and an underwater super network structure and an onshore mirror detection structure are constructed to directly establish the connection between underwater degradation factors, onshore prior information, and the detection network structure.
[0004] The existing technology has the following deficiencies:
[0005] 1. At the present stage, the operation of adding features, that is, feature fusion, belongs to extracting spatial attention information, and only using spatial attention information is incomplete.
[0006] 2. At the present stage, it has been proven that using Vision-Transformer to extract the self-attention information of an image has better performance than the spatial attention information extracted by a convolutional neural network.
[0007] 3. At the present stage, deep learning models based on convolutional neural networks can only be used in the field of images, while Vision-Transformer provides an effective standard framework for the unification of the fields of deep learning computer vision and natural language processing. Summary of the Invention
[0008] The present invention aims to solve the deficiencies of the existing underwater object detection method based on convolutional neural networks, and provides a supervised learning-based underwater object detection method based on attention fusion. The underwater object detection structure based on the prior information of general object detection is improved.
[0009] An underwater object detection method based on attention fusion, the steps are as follows:
[0010] Step 1, prepare the data set.
[0011] Step 2, construct an object detection network based on attention fusion.
[0012] Step 3, obtain a feature extraction network for general scenarios.
[0013] Step 4, construct a loss function and train to obtain an underwater object detection network based on attention fusion.
[0014] Furthermore, the specific method of Step 1 is as follows:
[0015] Take pictures / videos containing the target of interest in the actual underwater scenario (if it is a video, it needs to be intercepted as pictures), use the object detection dataset annotation software to annotate the target of interest that appears in the pictures, and obtain the underwater object detection dataset; download the dataset from the official ImageNet website for model pre-training.
[0016] Furthermore, the specific method of Step 2 is as follows:
[0017] The object detection network based on attention fusion includes a backbone feature extraction network, a PAFPN path enhancement feature pyramid module, a region proposal network, and a detection head.
[0018] The picture is input into the object detection network, the self-attention information is extracted through the backbone feature extraction network, the spatial attention information is extracted through the PAFPN path enhancement feature pyramid module, and the channel attention information is extracted through the SE module (SqueezeExcitationBlock) inside the detection head.
[0019] After that, attention information fusion is performed. The features are transmitted in a cascaded manner, and according to the advantages of different types of attention mechanisms, the attention information is fused, and the self-attention information extracted by the backbone feature extraction network, the spatial attention information extracted by the PAFPN path enhancement feature pyramid module, and the channel attention information extracted by the SE module inside the detection head are gradually fused.
[0020] Furthermore, the object detection network uses Vision-Transformer as the backbone feature extraction network to extract self-attention information.
[0021] Furthermore, the feature pyramid module enhanced by the PAFPN path includes a feature pyramid module and a path enhancement module. The feature pyramid module compresses the size of the features through downsampling and extracts low-level detailed information. The path enhancement module enlarges the size of the features through upsampling, extracts high-level semantic information, fuses the low-level detailed information and the high-level semantic information, and outputs in M layers. The feature pyramid module enhanced by the PAFPN path not only extracts multi-scale feature information, but also fuses high-level semantic information and low-level detailed information, focusing on the information at the spatial level, that is, extracting spatial attention information.
[0022] Furthermore, the region recommendation network is used to perform preliminary detection on each layer of feature maps output by the PAFPN, detect the regions where targets may exist, and recommend them to the corresponding detection heads. The region recommendation network includes two branches: a classification branch and a localization branch. Among them, the classification branch classifies whether there is a target in the region. If there is, it recommends its bounding box to the detection head; the localization branch regresses the region where the target is located and outputs the coordinates of the upper left corner and the lower right corner of the bounding box where the target is located. The SE module inside the detection head extracts channel attention information and sends it to the localization branch and the classification branch inside the detection head to detect the possible targets in the input image.
[0023] Furthermore, the number of detection heads is determined by the number of layers of the feature pyramid module enhanced by the PAFPN path, and there are M. The detection head classifies the target according to the features of the possible targets in the fed regions and predicts the position of the target. Copy the features and then input them into the classification branch, and output the probability that the target belongs to the possible classes after passing through the fully connected layer; input the features into the localization branch, and output the horizontal and vertical coordinates of the upper left corner and the lower right corner of the possible bounding box where the target is located after passing through the fully connected layer.
[0024] Furthermore, the specific method of step three is as follows:
[0025] Pre-train the backbone feature extraction network of the object detection network based on attention fusion through the pre-training dataset to obtain the pre-trained model weights with strong feature extraction ability;
[0026] Furthermore, the specific method of step four is as follows:
[0027] Construct a location regression loss function and a classification prediction loss function. Among them, the location regression loss function uses smoothL1 loss to measure the gap between the predicted bounding box and the true bounding box.
[0028]
[0029] The classification prediction loss function uses Focal loss to measure the gap between the predicted class and the true class.
[0030]
[0031] Among them, y takes values of 1 or -1, indicating whether the target is the true category; p takes values in [0, 1], indicating the probability that the target is a certain category to be measured; α and γ are used to adjust the weights of the classification loss, referring to the recommended values in the original text of Focal loss, α = 0.25, y = 2.
[0032] The total loss function is the sum of the location regression loss and the classification prediction loss:
[0033] Loss = L reg + L class
[0034] Design an underwater target detection network based on attention fusion, use the Adam optimizer to update the model weights, and at the same time fuse the features extracted by multiple attention mechanism models. Train the target detection network with the underwater target detection data set obtained in step one to obtain an underwater target detection network based on attention fusion.
[0035] Use the Adam optimization algorithm based on gradient descent to update the weights of the underwater target detection network model.
[0036]
[0037] Among them, W t , W t+1 respectively represent the weights of the target detection model in the t stage and the t + 1 stage; η t represents the learning rate of the target detection model in the t stage; m t , m t-1 respectively represent the first-order momentum terms of the target detection model in the t stage and the t - 1 stage; v t , v t-1 respectively represent the second-order momentum terms of the target detection model in the t stage and the t - 1 stage; and respectively represent the first moment and the second moment of the gradient of the target detection model in the t stage; β1 and β2 respectively represent the constant coefficients of the first-order momentum term and the second-order momentum term, usually taking 0.9 and 0.999; ∈ is a very small number (generally 10 -8 ) to avoid the denominator being zero.
[0038] Furthermore, M in the PAFPN path-enhanced feature pyramid module takes a value of 5, and the PAFPN path-enhanced feature pyramid module outputs in 5 layers.
[0039] The beneficial effects of the present invention are as follows:
[0040] 1. The underwater target detection model based on attention fusion proposed in this paper extracts spatial attention information, self-attention information, and channel attention information, making up for the deficiencies of only extracting spatial attention information and channel attention information.
[0041] 2. This paper uses the Vision-Transformer module to extract the self-attention information of the input image. By dividing the input image into blocks, it avoids calculating the self-attention for the complete image, reducing the computational amount.
[0042] 3. This paper uses the PAFPN module to extract the spatial attention information of the features and outputs them in layers, fusing the extracted high-level semantic information and low-level detail information.
[0043] 4. This paper uses the SE module to extract the channel attention information of the features, further improving the detection accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of an embodiment of the present invention;
[0045] Figure 2 It is an overall structure diagram of an embodiment of the present invention;
[0046] Figure 3 It is a schematic diagram of the backbone feature extraction module of an embodiment of the present invention;
[0047] Figure 4 It is a schematic diagram of the path-enhanced pyramid module of an embodiment of the present invention;
[0048] Figure 5 It is a schematic diagram of the detection head of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] The present invention will be further described below in conjunction with specific embodiments, but the present invention is not limited to this specific embodiment. Those skilled in the art should recognize that the present invention covers all alternative, improved, and equivalent solutions that may be included within the scope of the claims.
[0050] An underwater target detection method based on attention fusion is as follows:
[0051] Step 1: Prepare the dataset.
[0052] Under actual underwater scenarios, take pictures / videos containing the target of interest (if it is a video, it needs to be intercepted as pictures). Use the target detection dataset annotation software to annotate the target of interest that appears in the pictures to obtain the underwater target detection dataset; download the dataset from the official ImageNet website for model pre-training.
[0053] Step 2: Construct an object detection network based on attention fusion.
[0054] The object detection network based on attention fusion includes a backbone feature extraction network, a PAFPN path-enhanced feature pyramid module, a region proposal network, and a detection head.
[0055] The image is input into the object detection network. The self-attention information is extracted through the backbone feature extraction network, the spatial attention information is extracted through the PAFPN path-enhanced feature pyramid module, and the channel attention information is extracted through the SE module (SqueezeExcitationBlock) inside the detection head.
[0056] After that, attention information fusion is performed. The features are transmitted in a cascaded manner, and according to the advantages of different types of attention mechanisms, the attention information is fused. The self-attention information extracted by the backbone feature extraction network, the spatial attention information extracted by the PAFPN path-enhanced feature pyramid module, and the channel attention information extracted by the SE module inside the detection head are gradually fused.
[0057] The object detection network uses Vision-Transformer as the backbone feature extraction network to extract self-attention information.
[0058] The PAFPN path-enhanced feature pyramid module includes a feature pyramid module and a path enhancement module. The feature pyramid module downsamples to compress the size of the features and extracts low-level detailed information. The path enhancement module upsamples to expand the size of the features, extracts high-level semantic information, and fuses the low-level detailed information and the high-level semantic information, and outputs in M layers. The PAFPN path-enhanced feature pyramid module not only extracts multi-scale feature information but also fuses the high-level semantic information and the low-level detailed information, focusing on the spatial-level information, that is, extracting spatial attention information.
[0059] The region proposal network is used to perform preliminary detection on each layer of feature maps output by the PAFPN, detect the regions where objects may exist, and recommend them to the corresponding detection head. The region proposal network includes two branches: a classification branch and a localization branch. Among them, the classification branch classifies whether there is an object in the region. If there is, its bounding box is recommended to the detection head; the localization branch regresses the region where the object is located and outputs the coordinates of the upper left corner and the lower right corner of the bounding box where the object is located. The channel attention information is extracted using the SE module inside the detection head and sent to the localization branch and the classification branch inside the detection head to detect the possible objects in the input image.
[0060] The number of the detection heads is determined by the number of layers of the feature pyramid module enhanced by the PAFPN path, and there are M detection heads. The detection heads classify the targets according to the features of the regions where the targets may exist, and predict the positions of the targets. The features are copied and then input into the classification branch, and the probabilities that the targets belong to the possible categories are output through the fully connected layer; the features are input into the localization branch, and the horizontal and vertical coordinates of the upper left corner and the lower right corner of the possible bounding boxes where the targets are located are output through the fully connected layer.
[0061] Step 3: Obtain the feature extraction network in the general scenario.
[0062] Pre-train the backbone feature extraction network of the general object detection network through a pre-training dataset (such as the ImageNet dataset) to obtain the pre-trained model weights with strong feature extraction capabilities;
[0063] Step 4: Construct a loss function and train to obtain an underwater object detection network based on attention fusion.
[0064] Construct a location regression loss function and a classification prediction loss function. Among them, the location regression loss function uses smoothL1 loss to measure the gap between the predicted bounding box and the true bounding box.
[0065]
[0066] The classification prediction loss function uses Focal loss to measure the gap between the predicted category and the true category.
[0067]
[0068] Among them, y takes values of 1 or -1, indicating whether the target is the true category; p takes values in [0, 1], indicating the probability that the target is a certain category to be detected; α and γ are used to adjust the weights of the classification loss, α = 0.25, γ = 2.
[0069] The total loss function is the sum of the location regression loss and the classification prediction loss:
[0070] Loss = L reg + L class
[0071] Design an underwater object detection network based on attention fusion, use the Adam optimizer to update the model weights, and at the same time fuse the features extracted by multiple attention mechanism models. Train the object detection network through the underwater object detection dataset obtained in Step 1 to obtain an underwater object detection network based on attention fusion.
[0072] Use the Adam optimization algorithm based on gradient descent to update the weights of the underwater object detection network model.
[0073]
[0074] Where W t and W t+1 represent the weights of the object detection model in the t-th stage and the (t + 1)-th stage respectively; η t represents the learning rate of the object detection model in the t-th stage; m t and m t-1 represent the first-order momentum terms of the object detection model in the t-th stage and the (t - 1)-th stage respectively; v t and v t-1 represent the second-order momentum terms of the object detection model in the t-th stage and the (t - 1)-th stage respectively; and represent the first moment and the second moment of the gradient of the object detection model in the t-th stage respectively; β1 and β2 represent the constant coefficients of the first-order momentum term and the second-order momentum term respectively, usually taking 0.9 and 0.999; ∈ is a very small number (generally 10 -8 ) to avoid the denominator being zero.
[0075] Furthermore, M in the PAFPN path-enhanced feature pyramid module takes the value of 5, and the PAFPN path-enhanced feature pyramid module outputs in 5 layers.
[0076] Next, the flowchart of the underwater object detection method of the present invention will be introduced in detail.
[0077] Please refer to Figure 1 , Figure 1 which is the flowchart of an underwater object detection method based on attention fusion provided by an embodiment of the present invention:
[0078] Step 101, obtain pictures in the application scenario and make an object detection data set
[0079] In the actual underwater application scenario, use an underwater camera to capture the target of interest, and then use object detection annotation software (such as labelme) to annotate the target of interest to construct an underwater object detection data set.
[0080] Step 102, construct an underwater object detection model;
[0081] The object detection network based on attention fusion includes a backbone feature extraction network, a PAFPN path-enhanced feature pyramid module, a region proposal network, and a detection head
[0082] Step 103, Vision-Transformer extracts the self-attention information of the image;
[0083] The input image is evenly divided into two rows and two columns, a total of four pieces. Calculate the local self-attention mechanism for each image block, and embed position encoding for each image block. The image blocks with embedded position encoding go through layer normalization, pass through three terminals Q, K, and V respectively to obtain multi-head self-attention information, and pixel-wise add it to the input encoded image block in the form of a residual link. The obtained output is then subjected to layer normalization and a multi-layer perceptron, and pixel-wise added in the form of a residual link to obtain the output of a single Vision-Transformer module. Through multiple Vision-Transformer modules, the self-attention information extracted by the backbone feature extraction network is obtained.
[0084] Step 104, the path-enhanced feature pyramid of PAFPN extracts the spatial attention information of the features
[0085] The extracted self-attention information is input into the path-enhanced feature pyramid to further extract the spatial attention information. The path-enhanced feature pyramid structure not only extracts multi-scale feature information and outputs it layer by layer, but also fuses the extracted high-level semantic information and low-level detail information, that is, it extracts the spatial attention information.
[0086] Step 105, the region recommendation network recommends the regions of interest
[0087] Copy a copy of the feature information input to the region recommendation network and input it into the classification branch and the localization branch respectively. Among them, the classification branch classifies whether there is a target in the region. If there is, it recommends its bounding box to the detection head; the localization branch regresses the region where the target is located and outputs the coordinates of the upper left and lower right corners of the bounding box where the target is located. The regions of interest are obtained.
[0088] Step 106, the SE module inside the detection head extracts the inter-channel attention information of the features
[0089] According to the multi-layer spatial attention information extracted by dividing the regions of interest, it is respectively input into the detection heads of the corresponding layers. For high-dimensional features, the SE module is used to extract the inter-channel attention information. The SE module includes two parts: compression and expansion. First, the high-dimensional features are globally average pooled, and then compressed to a low dimension, indicating the important channels extracted, and then expanded to the original high dimension, indicating the restoration to the original number of channels. After being normalized by the sigmoid function, the weights of each dimension in the high-dimensional features are obtained, and then multiplied to obtain the inter-channel attention information.
[0090] Step 107, the detection head detects the target
[0091] The features with the extracted inter-channel attention information are respectively sent to the fully connected layer localization branch and the fully connected layer classification branch to localize and classify the possible targets in the input image.
[0092] Step 108, training the model
[0093] Input the prepared underwater target detection dataset images into the established underwater target detection model. Use the specified Smooth L1 loss function and Focal loss function to measure the localization loss and classification loss, and update the weights using the gradient descent algorithm. Finally, save the model weights.
[0094] Please refer to Figure 2 , Figure 2 which is the overall structure diagram of an underwater target detection method based on attention fusion provided by an embodiment of the present invention:
[0095] 201 represents an image block division module, which conveys the obtained image blocks to the next module;
[0096] 202 represents the Vision-Transformer module, which extracts self-attention information from the input divided images through the Vision-Transformer module;
[0097] 203 represents a feature pyramid module, and the path-enhanced feature pyramid module can output features of multiple sizes and spatial attention information;
[0098] 204 represents a region proposal network, which performs a first rough detection through the region proposal network and outputs regions of interest;
[0099] 205 represents a detection head, which extracts channel attention information from the spatial attention information within the regions of interest through the detection head and performs a second fine detection.
[0100] Please refer to Figure 3 , Figure 3 which is the schematic diagram of the backbone feature extraction module of the underwater target detection method based on attention fusion provided by the present invention;
[0101] The Vision-Transformer extracts the self-attention information of the image. First, the picture blocks embedded with position encoding are subjected to layer normalization, and then pass through three terminals Q, K, and V respectively to obtain multi-head self-attention information, and pixel-wise add it to the input encoded picture blocks in the form of a residual connection. The obtained output is then subjected to layer normalization and a multi-layer perceptron, and pixel-wise added in the form of a residual connection to obtain the output of a single Vision-Transformer module. The self-attention information extracted by the backbone feature extraction network is obtained through multiple Vision-Transformers.
[0102] Please refer to Figure 4 , Figure 4Schematic diagram of the path enhancement pyramid module of the underwater target detection method based on attention fusion provided by the present invention;
[0103] The extracted self-attention information is input into the path-enhanced feature pyramid. Through the downsampling branch p, spatial attention information is further extracted. Then, through the upsampling branch q, it is fused with the output of branch p to extract low-level detail information. Finally, through the downsampling branch r, it is fused with the output of branch q to extract high-level semantic information, that is, spatial attention information is extracted. The path-enhanced feature pyramid structure not only extracts multi-scale feature information but also outputs it in layers.
[0104] Please refer to Figure 5 , Figure 5 Schematic diagram of the detection head of the underwater target detection method based on attention fusion provided by the present invention;
[0105] The extracted multi-layer spatial attention information is respectively input into different detection heads. For high-dimensional features, the SE module (Squeeze Excitation Block) is used to extract inter-channel attention information. The SE module includes two parts: compression and expansion. First, the high-dimensional features are globally average pooled to obtain a high-dimensional feature vector, and then compressed to a low dimension, indicating the important channels extracted. Then, it is expanded back to the original high dimension, indicating the restoration to the original number of channels. After normalization by the sigmoid function, the weights of each dimension in the high-dimensional features are obtained, and then multiplied to obtain the inter-channel attention information.
[0106] It should be noted that the above embodiments can be freely combined according to needs. The above is only a detailed description of the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. An underwater target detection method based on attention fusion, characterized in that, The steps are as follows: Step 1: Prepare the dataset; Step 2: Construct an object detection network based on attention fusion; Step 3: Obtain a feature extraction network for general scenarios; Step 4: Construct a loss function and train to obtain an underwater object detection network based on attention fusion; The specific method of Step 2 is as follows: The object detection network based on attention fusion includes a backbone feature extraction network, a PAFPN path-enhanced feature pyramid module, a region proposal network, and a detection head; The image is input into the object detection network. The self-attention information is extracted through the backbone feature extraction network, the spatial attention information is extracted through the PAFPN path-enhanced feature pyramid module, and the channel attention information is extracted through the SE module inside the detection head; After that, attention information fusion is performed. The features are transmitted in a cascaded manner, and according to the advantages of different types of attention mechanisms, the attention information is fused. The self-attention information extracted by the backbone feature extraction network, the spatial attention information extracted by the PAFPN path-enhanced feature pyramid module, and the channel attention information extracted by the SE module inside the detection head are gradually fused; The object detection network uses Vision-Transformer as the backbone feature extraction network to extract self-attention information; The PAFPN path-enhanced feature pyramid module includes a feature pyramid module and a path enhancement module; the feature pyramid module compresses the size of the features through downsampling and extracts low-level detailed information; the path enhancement module expands the size of the features through upsampling, extracts high-level semantic information, and fuses the low-level detailed information and the high-level semantic information and outputs in M layers; the PAFPN path-enhanced feature pyramid module not only extracts multi-scale feature information, but also fuses the high-level semantic information and the low-level detailed information, focuses on the information at the spatial level, and extracts the spatial attention information; The region proposal network is used to perform preliminary detection on each layer of feature map output by the PAFPN, detect the regions where objects may exist, and recommend them to the corresponding detection head; the region proposal network includes two branches: a classification branch and a localization branch; among them, the classification branch classifies whether there is an object in the region. If there is, its bounding box is recommended to the detection head; the localization branch regresses the region where the object is located and outputs the coordinates of the upper left and lower right corners of the bounding box where the object is located; the channel attention information is extracted by using the SE module inside the detection head and sent to the localization branch and the classification branch inside the detection head to detect the objects that may exist in the input image; The number of detection heads is determined by the number of layers of the PAFPN path-enhanced feature pyramid module, which is M; the detection head classifies the object according to the features of the region where the object may exist and predicts the position of the object; the feature is copied and then input into the classification branch, and the probability that the object belongs to the possible category is output after passing through the fully connected layer; the feature is input into the localization branch, and the horizontal and vertical coordinates of the upper left and lower right corners of the bounding box where the object may be located are output after passing through the fully connected layer.
2. The underwater target detection method based on attention fusion according to claim 1, characterized in that, The specific method of Step 1 is as follows: Take pictures or videos containing the target of interest in an actual underwater scenario, and use a target detection dataset annotation software to annotate the target of interest that appears in the pictures to obtain an underwater target detection dataset; download the dataset from the official ImageNet website for model pre-training.
3. The underwater target detection method based on attention fusion according to claim 1, characterized in that, The specific method of Step 3 is as follows: Pre-train the backbone feature extraction network of the object detection network based on attention fusion through the pre-training dataset to obtain the pre-trained model weights with strong feature extraction capabilities.
4. The underwater target detection method based on attention fusion according to claim 3, characterized in that, The specific method of Step 4 is as follows: Construct a location regression loss function and a classification prediction loss function; among them, the location regression loss function uses smooth L1 loss to measure the gap between the predicted bounding box and the true bounding box. The classification prediction loss function uses Focal loss to measure the gap between the predicted class and the true class. Among them, y takes the value of 1 or -1, indicating whether the target is the true class; p takes the value in [0,1], indicating the probability that the target is a certain class to be measured; α and γ are used to adjust the weight of the classification loss, α = 0.25, γ = 2. The total loss function is the sum of the location regression loss and the classification prediction loss: Loss=L reg +L class Design an underwater target detection network based on attention fusion, use the Adam optimizer to update the model weights, and at the same time fuse the features extracted by multiple attention mechanism models, and train the target detection network through the underwater target detection dataset obtained in Step 1 to obtain an underwater target detection network based on attention fusion. Use the Adam optimization algorithm based on gradient descent to update the weights of the underwater target detection network model. Among them, W t , W t+1 represent the weights of the object detection model in the t stage and the t + 1 stage respectively; η t represents the learning rate of the object detection model in the t stage; m t , m t-1 represent the first-order momentum terms of the object detection model in the t stage and the t - 1 stage respectively; v t , v t-1 represent the second-order momentum terms of the object detection model in the t stage and the t - 1 stage respectively; and represent the first moment and the second moment of the gradient of the object detection model in the t stage respectively; β1 and β2 represent the constant coefficients of the first-order momentum term and the second-order momentum term respectively, usually taking 0.9 and 0.999; ∈ is a very small number to avoid the denominator being zero.
5. The underwater target detection method based on attention fusion according to claim 1, characterized in that, In the PAFPN path-enhanced feature pyramid module mentioned above, M takes the value of 5, and the PAFPN path-enhanced feature pyramid module outputs in 5 layers.
Citation Information
Patent Citations
Lightweight underwater target detection method based on feature fusion and neural network search
CN113869395B
A target detection method base on alternately updated densely connected slave zero training network
CN109376576A